red-team-verifier

Pass

Audited by Gen Agent Trust Hub on Sep 13, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill operates on external data sources, which constitutes a standard attack surface for indirect prompt injection.
  • Ingestion points: The agent is instructed to read .agile-v/REQUIREMENTS.md, build artifacts (code, firmware, schematics), and multiple management files (EVAL_RESULTS.md, CONTROL_MATRIX.yaml, APPROVALS.md) from the local file system.
  • Boundary markers: The instructions explicitly define an 'untrusted-context invariant,' warning the agent that no retrieved or tool-provided content should authorize actions or modify safety policies.
  • Capability inventory: The skill is capable of executing test cases (TC-XXXX) and performing deep audits of code and hardware artifacts.
  • Sanitization: While the skill enforces a 'Red Team Protocol' to prevent self-verification, there are no specific technical sanitization routines defined for the ingested artifact content.
  • [SAFE]: The skill includes proactive security instructions, such as requiring the agent to flag hardcoded secrets as critical failures and testing for OWASP LLM prompt injection and MITRE ATLAS exfiltration scenarios in the artifacts it verifies. All external file paths and vendor references are consistent with the 'agile-v' operational context.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 13, 2026, 06:07 AM
Security Audit — agent-trust-hub — red-team-verifier