llm-evals-audit
Pass
Audited by Gen Agent Trust Hub on Apr 23, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill's primary function is to perform a diagnostic audit of existing evaluation artifacts, such as judge prompts and alignment data, which are typically found within the user's project environment.
- [SAFE]: No malicious behaviors, such as unauthorized network access, data exfiltration, or hardcoded credentials, were detected in the skill's instructions or metadata.
- [SAFE]: The skill operates within the vendor's ecosystem, referencing complementary tools and well-known evaluation services in a manner consistent with its stated purpose.
- [SAFE]: While the skill ingests project-specific data for analysis, its capabilities are restricted to generating findings and recommendations, with no evidence of high-privilege execution or persistence mechanisms.
Audit Metadata