investigating-eval-results
Warn
Audited by Socket on Aug 22, 2026
1 alert found:
AnomalyAnomalySKILL.md
LOWAnomalyLOW
SKILL.md
Mostly coherent dev tooling for local eval investigation, but it is not purely read-only: it drives autonomous edit/rerun loops, processes untrusted trajectory/log content, and instructs use of another skill for self-modification. No clear credential harvesting or attacker-controlled exfiltration is present, so this is better classified as suspicious/medium-risk rather than malicious.
Confidence: 87%Severity: 56%
Audit Metadata