judge-with-debate

Pass

Audited by Gen Agent Trust Hub on Sep 15, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill exhibits an indirect prompt injection surface because it processes external, untrusted solutions or code files via subagent workflows without sanitizing the input text.
  • Ingestion points: External data enters the system through the {path to solution file(s)} and {task description} placeholders in SKILL.md during Phase 1 (Independent Analysis) and Phase 2 (Debate Rounds).
  • Boundary markers: The system uses standard Markdown headings (e.g., ## Solution) to separate untrusted content from the instructions. It lacks explicit system instructions or isolation techniques to prevent the model from obeying embedded commands inside the solution files.
  • Capability inventory: The orchestrator utilizes a Task tool to launch subagents (sadd:meta-judge and sadd:judge) with the Opus model, which write text reports to the local filesystem under .specs/reports/.
  • Sanitization: There is no mechanism in place to sanitize, filter, or escape the content of the solution files before embedding them into the prompts passed to the subagents.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 15, 2026, 01:42 AM
Security Audit — agent-trust-hub — judge-with-debate