llm-evaluation-design

Pass

Audited by Gen Agent Trust Hub on Sep 4, 2026

Risk Level: SAFENO_CODE
Full Analysis
  • [SAFE]: The skill is primarily instructional, composed of Markdown guidelines and YAML evaluation cases. It does not include any executable scripts or binary files.
  • [SAFE]: There are no indicators of data exfiltration, hardcoded credentials, or unauthorized network activity.
  • [INDIRECT_PROMPT_INJECTION]: The skill naturally processes user-supplied information such as task distributions and sample datasets. Although this is an entry point for external data, the prompt instructions mandate a rigorous input audit and the skill lacks any high-risk tools that could be exploited via injection.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 4, 2026, 06:26 AM
Security Audit — agent-trust-hub — llm-evaluation-design