llm-evaluation-design
Pass
Audited by Gen Agent Trust Hub on Sep 4, 2026
Risk Level: SAFENO_CODE
Full Analysis
- [SAFE]: The skill is primarily instructional, composed of Markdown guidelines and YAML evaluation cases. It does not include any executable scripts or binary files.
- [SAFE]: There are no indicators of data exfiltration, hardcoded credentials, or unauthorized network activity.
- [INDIRECT_PROMPT_INJECTION]: The skill naturally processes user-supplied information such as task distributions and sample datasets. Although this is an entry point for external data, the prompt instructions mandate a rigorous input audit and the skill lacks any high-risk tools that could be exploited via injection.
Audit Metadata