constitutional-ai

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted input data (the prompts variable) and interpolates it directly into various prompt templates for critique, revision, and preference evaluation. This creates a risk where malicious instructions embedded in the input data could influence the model's critique process or poison the resulting training data.
  • Ingestion points: The prompts list in Workflow 1 and Workflow 2 of SKILL.md serves as the entry point for external data into the agent's context.
  • Boundary markers: The prompt templates (e.g., critique_prompt, revision_prompt, preference_prompt) use simple f-string interpolation without delimiters or instructions for the model to ignore embedded commands within the {question} or {response} variables.
  • Capability inventory: The skill uses the trl library's SFTTrainer, RewardTrainer, and PPOTrainer, which involve writing model checkpoints and training logs to the local file system.
  • Sanitization: There is no evidence of input validation, escaping, or filtering of the content within the prompts or responses variables before they are used in prompt templates.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — constitutional-ai