harness-eval
Pass
Audited by Gen Agent Trust Hub on Aug 5, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill performs diagnostic analysis on local repository files (AGENTS.md, skill trees, and cited documentation). All identified behaviors (file reading, subprocess execution of internal scripts, and report generation) are entirely consistent with the documented purpose of evaluating instruction harnesses.
- [SAFE]: The skill uses a 'dual-judge' protocol for subjective evaluations (redundancy and usefulness), which includes 'trap' claims to ensure the analyzing models are not hallucinating or being overly aggressive in suggesting removals.
- [SAFE]: The skill implements strict boundary controls, such as excluding decision-record trees (ADRs/RFCs) and requiring explicit user opt-in via questionnaires before including optional project documentation or spawning LLM-based judges.
- [SAFE]: The included Python scripts perform deterministic path and command verification (Track A) and surface extraction. These scripts use standard libraries, follow best practices for path normalization to avoid directory traversal, and do not make unauthorized network connections.
- [SAFE]: No obfuscation, hardcoded credentials, or persistence mechanisms were detected. The skill maintains a report-only posture and requires explicit user commands to apply any suggested trims after the audit is complete.
Audit Metadata