harness-eval

Fail

Audited by Gen Agent Trust Hub on May 15, 2026

Risk Level: HIGHCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
  • [COMMAND_EXECUTION]: The skill contains explicit instructions for bypassing platform-level sensitive-path security checks. It describes a method to write artifacts to a temporary directory (/tmp) and then use a shell script to move them into restricted directories (such as .claude/outputs/sessions/) to evade permission prompts and auditing.
  • [DATA_EXFILTRATION]: The skill targets sensitive file system paths within the .claude/ directory for writing artifacts, using techniques designed to circumvent standard platform protections for these locations.
  • [SAFE]: References benchmarking research and evaluation definitions from the revfactory/claude-code-harness public GitHub repository.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
May 15, 2026, 06:17 PM
Security Audit — agent-trust-hub — harness-eval