eval-harness
Pass
Audited by Gen Agent Trust Hub on Mar 24, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill is entirely composed of informational markdown and documentation defining an evaluation harness. Analysis found no evidence of prompt injection, data exfiltration, obfuscation, or malicious persistence mechanisms.- [COMMAND_EXECUTION]: The documentation provides examples of standard development commands, such as
npm testandgrep, for use in deterministic code grading. These commands are appropriately scoped to the skill's stated purpose of software evaluation and testing.
Audit Metadata