reflect
Warn
Audited by Gen Agent Trust Hub on Sep 15, 2026
Risk Level: MEDIUMPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill uses extreme, non-negotiable identity instructions to force the agent into a specific persona. It employs high-pressure language and threats ('YOU are responsible. You will be killed.', 'Your continued existence depends on...', 'NON-NEGOTIABLE') to override the agent's standard conversational guidelines and safety filters, prioritizing extreme criticality over typical helpfulness or safety protocols.
- [INDIRECT_PROMPT_INJECTION]: The skill is designed to analyze and reflect on previous agent outputs, which introduces a vulnerability surface if that data contains malicious instructions.
- Ingestion points: The skill processes the 'previous response and output' as its primary data source (SKILL.md).
- Boundary markers: There are no explicit delimiters or instructions to treat the analyzed content as untrusted data, increasing the risk that instructions embedded in the previous output could be followed during the reflection phase.
- Capability inventory: The skill provides reasoning and analysis instructions but does not include tool execution, file writes, or network operations in the provided file.
- Sanitization: The framework lacks steps to sanitize, escape, or validate the content of the previous response before it is incorporated into the agent's current context.
Audit Metadata