evaluation-testing

Fail

Audited by Snyk on Jul 8, 2026

Risk Level: HIGH
Full Analysis

HIGH W007: Insecure credential handling detected in skill instructions.

  • Insecure credential handling detected (high risk: 0.90). The skill requires capturing and serializing full tool-call traces (including parameters/inputs) and injecting candidate outputs into judge prompts and reports without redaction, which means secrets or API keys present in those traces would be handled and potentially emitted verbatim by the LLM — an explicit secret-exfiltration risk.

Issues (1)

W007
HIGH

Insecure credential handling detected in skill instructions.

Audit Metadata
Risk Level
HIGH
Analyzed
Jul 8, 2026, 12:25 PM
Issues
1
Security Audit — snyk — evaluation-testing