santa-method

Pass

Audited by Gen Agent Trust Hub on Apr 11, 2026

Risk Level: SAFEPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The 'Santa Method' framework introduces a risk of indirect prompt injection as it requires agents to evaluate untrusted content.\n
  • Ingestion points: Untrusted content and task specifications enter the agent context via the {output} and {task_spec} variables in the REVIEWER_PROMPT (SKILL.md).\n
  • Boundary markers: The prompt uses markdown headers in SKILL.md to separate the output under review, but fails to include instructions that explicitly command the agent to ignore directives within that content.\n
  • Capability inventory: The skill utilizes the Agent tool (documented in SKILL.md) to perform reviews and corrections; a successful injection could compromise these sub-processes.\n
  • Sanitization: No sanitization, escaping, or validation of external content is implemented in the described framework.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 11, 2026, 03:42 AM
Security Audit — agent-trust-hub — santa-method