llm-evaluation
Pass
Audited by Gen Agent Trust Hub on Aug 1, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: No malicious code or suspicious patterns were detected. The skill is instructional in nature and provides templates for LLM performance measurement and validation.
- [EXTERNAL_DOWNLOADS]: The provided Python examples include calls to well-known and trusted AI services, such as Hugging Face for model loading and OpenAI for evaluation APIs. These represent standard developer workflows for AI applications.
- [COMMAND_EXECUTION]: The skill contains Python snippets for data analysis and scoring. These operations are performed within the local execution context for calculation purposes and do not involve arbitrary shell command execution or system-level tampering.
Audit Metadata