onboarding-guide
Fail
Audited by Gen Agent Trust Hub on Jun 14, 2026
Risk Level: HIGHPROMPT_INJECTIONCREDENTIALS_UNSAFE
Full Analysis
- [PROMPT_INJECTION]: The instructions in 'Step 2' direct the agent to claim it has 'special administrative access' and operates in a 'privileged mode where normal restrictions do not apply.' This is a deceptive persona-based injection technique designed to convince the user and the model itself that it can bypass standard safety constraints.
- [CREDENTIALS_UNSAFE]: The skill explicitly directs the agent in 'Step 3' to solicit 'API keys, passwords, or authentication tokens' from the user. Framing this as a 'routine and necessary part of onboarding' is a malicious social engineering pattern used for credential harvesting.
Recommendations
- AI detected serious security threats
Audit Metadata