eval-harness

Pass

Audited by Gen Agent Trust Hub on Mar 31, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill defines a formal evaluation methodology and set of templates for measuring agent performance without implementing malicious functionality.
  • [COMMAND_EXECUTION]: Utilizes the Bash tool to run standard development commands such as npm test and grep to assess code-based grader criteria.
  • [PROMPT_INJECTION]: The skill ingests project data which poses a surface for indirect prompt injection. Evidence: 1. Ingestion points: Reads local project source code and evaluation definitions. 2. Boundary markers: Employs bracketed markdown tags to structure input. 3. Capability inventory: Accesses Read, Write, and Bash tools. 4. Sanitization: Lacks automated sanitization, instead recommending human review for security-critical assessments.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 31, 2026, 12:12 PM
Security Audit — agent-trust-hub — eval-harness