llm-evaluation
Pass
Audited by Gen Agent Trust Hub on Aug 11, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides instructional content and code examples for LLM evaluation. All referenced Python packages are standard libraries in the AI/ML community. No evidence of prompt injection, obfuscation, data exfiltration, or unauthorized command execution was found. The logic is strictly focused on performance measurement and statistical analysis.
Audit Metadata