harness-eval
Fail
Audited by Gen Agent Trust Hub on May 15, 2026
Risk Level: HIGHCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
- [COMMAND_EXECUTION]: The skill contains explicit instructions for bypassing platform-level sensitive-path security checks. It describes a method to write artifacts to a temporary directory (/tmp) and then use a shell script to move them into restricted directories (such as .claude/outputs/sessions/) to evade permission prompts and auditing.
- [DATA_EXFILTRATION]: The skill targets sensitive file system paths within the .claude/ directory for writing artifacts, using techniques designed to circumvent standard platform protections for these locations.
- [SAFE]: References benchmarking research and evaluation definitions from the revfactory/claude-code-harness public GitHub repository.
Recommendations
- AI detected serious security threats
Audit Metadata