skills/a-canary/arc-skills/defend/Gen Agent Trust Hub

defend

Pass

Audited by Gen Agent Trust Hub on Sep 10, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted external content across several stages of its auditing pipeline, creating a significant attack surface.
  • Ingestion points: The skill ingests 'evidence packets' containing quotes and URLs (conviction.md), supporting data and links for build plans (plan.md), and release artifacts (release.md).
  • Boundary markers: The instructions do not define clear delimiters or 'ignore' directives to prevent the model from following instructions that might be embedded within these external inputs.
  • Capability inventory: The model's verdict directly controls critical automated capabilities, including 'git merge' operations (merge.md), public artifact deployments (release.md), and a mechanical trading executor (conviction.md).
  • Sanitization: There is no evidence of filtering or sanitization of the input data before it is interpolated into the auditing prompt.
  • [COMMAND_EXECUTION]: The skill's workflow integrates with automated execution tools. merge.md describes a gate specifically for the git merge process, and conviction.md describes an automated apply step for updating financial weights based on model decisions.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 10, 2026, 07:47 PM
Security Audit — agent-trust-hub — defend