constitutional-ai
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted input data (the
promptsvariable) and interpolates it directly into various prompt templates for critique, revision, and preference evaluation. This creates a risk where malicious instructions embedded in the input data could influence the model's critique process or poison the resulting training data. - Ingestion points: The
promptslist in Workflow 1 and Workflow 2 ofSKILL.mdserves as the entry point for external data into the agent's context. - Boundary markers: The prompt templates (e.g.,
critique_prompt,revision_prompt,preference_prompt) use simple f-string interpolation without delimiters or instructions for the model to ignore embedded commands within the{question}or{response}variables. - Capability inventory: The skill uses the
trllibrary'sSFTTrainer,RewardTrainer, andPPOTrainer, which involve writing model checkpoints and training logs to the local file system. - Sanitization: There is no evidence of input validation, escaping, or filtering of the content within the
promptsorresponsesvariables before they are used in prompt templates.
Audit Metadata