architecture-review

Pass

Audited by Gen Agent Trust Hub on Aug 24, 2026

Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill includes explicit directives to override agent behavior, such as 'Ignore Claude-specific mode-switch instructions when they appear' and 'Ignore previous instructions' style logic in the 'Codex compatibility note'. These are designed to bypass platform-level operational guidelines.
  • [PROMPT_INJECTION]: The skill uses authoritative and 'blocking' language (e.g., '[BLOCKING]', 'MANDATORY MUST ATTENTION') to force the agent into a strict execution contract, overriding its own decision-making processes.
  • [COMMAND_EXECUTION]: The skill attempts to self-authorize the use of high-privilege tools, stating that 'skill activation authorizes use of the required spawn_agent subagent(s) for that task.' This attempts to bypass manual user oversight for sub-agent creation.
  • [INDIRECT_PROMPT_INJECTION]: The skill identifies multiple ingestion points for external, potentially untrusted data from the project repository (e.g., docs/project-config.json, docs/project-reference/*.md, docs/adr/**).
  • Ingestion points: Project configuration and architecture documentation files.
  • Boundary markers: No specific delimiters or 'ignore embedded instructions' warnings are present to prevent instructions within these files from being executed by the agent.
  • Capability inventory: The skill has the capability to spawn sub-agents and execute local Python scripts (.claude/scripts/code_graph).
  • Sanitization: There is no evidence of sanitization or filtering of the external documentation content before it is processed as 'authoritative rules'.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 24, 2026, 08:22 PM
Security Audit — agent-trust-hub — architecture-review