eval-harness
Pass
Audited by Gen Agent Trust Hub on Mar 31, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill defines a formal evaluation methodology and set of templates for measuring agent performance without implementing malicious functionality.
- [COMMAND_EXECUTION]: Utilizes the Bash tool to run standard development commands such as
npm testandgrepto assess code-based grader criteria. - [PROMPT_INJECTION]: The skill ingests project data which poses a surface for indirect prompt injection. Evidence: 1. Ingestion points: Reads local project source code and evaluation definitions. 2. Boundary markers: Employs bracketed markdown tags to structure input. 3. Capability inventory: Accesses Read, Write, and Bash tools. 4. Sanitization: Lacks automated sanitization, instead recommending human review for security-critical assessments.
Audit Metadata