llm-evaluation

Pass

Audited by Gen Agent Trust Hub on Aug 11, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides instructional content and code examples for LLM evaluation. All referenced Python packages are standard libraries in the AI/ML community. No evidence of prompt injection, obfuscation, data exfiltration, or unauthorized command execution was found. The logic is strictly focused on performance measurement and statistical analysis.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 11, 2026, 01:47 PM
Security Audit — agent-trust-hub — llm-evaluation