eval-harness

Pass

Audited by Gen Agent Trust Hub on Sep 1, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill defines a model-based grading framework that processes code changes and open-ended AI outputs to assess quality and correctness. This creates a surface for indirect prompt injection if the content being graded contains adversarial instructions meant to influence the grader's evaluation.
  • Ingestion points: The [MODEL GRADER PROMPT] section in SKILL.md ingests code changes and open-ended text outputs for analysis.
  • Boundary markers: The provided templates do not currently include explicit delimiters or instructions to the model to ignore potential instructions embedded within the analyzed content.
  • Capability inventory: The skill employs tools including Bash, Write, Edit, Grep, and Glob to interact with the file system and run development commands like npm test and npm run build.
  • Sanitization: No explicit sanitization or filtering logic is defined for the content passed to the model-based grader, though the skill encourages human review for security-sensitive evaluations.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 1, 2026, 02:38 AM
Security Audit — agent-trust-hub — eval-harness