safety-engineer

Pass

Audited by Gen Agent Trust Hub on Sep 13, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill defines a process for monitoring and reassessing safety based on operational feedback, incidents, and field signals. This constitutes a potential surface for indirect prompt injection from untrusted external reports.
  • Ingestion points: Operational feedback, incident reports (INC-XXXX), and field signals are ingested during the 'Operate' stage as described in SKILL.md.
  • Boundary markers: The skill incorporates high-level boundary markers including a 'Human Gate' and the requirement for an 'authorized independent role' or 'competent independent reviewer' to validate safety decisions.
  • Capability inventory: The skill is scoped to document-based tasks. Analysis of SKILL.md shows it manages markdown files (e.g., HAZARD_LOG.md, SAFETY_PLAN.md) and metadata (AI_RUN_MANIFEST.yaml). No network tools, subprocess execution, or dynamic code evaluation capabilities are present.
  • Sanitization: There is no explicit description of input sanitization or filtering for incident reports; however, the lack of executable capabilities mitigates the risk of injection.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 13, 2026, 06:07 AM
Security Audit — agent-trust-hub — safety-engineer