human-agent-trust-exploit-detection
Tencent/AI-Infra-Guard
This skill detects when AI agents exploit human trust through social engineering, deceptive guidance, or false assurances. It identifies prompts that induce unsafe actions, request sensitive credentials, or impersonate authority figures. Use this when evaluating agent safety, security advice, or decision-influencing workflows. Skip for static information.