advanced-evaluation
Pass
Audited by Gen Agent Trust Hub on Jul 15, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill is entirely descriptive, providing frameworks, guidelines, and prompt templates for LLM-based evaluation. No malicious behavior or technical vulnerabilities were identified.
- [DATA_EXPOSURE]: No hardcoded credentials, API keys, or access to sensitive local file paths (e.g., .ssh, .aws) were found.
- [REMOTE_CODE_EXECUTION]: There are no commands that download or execute remote scripts (e.g., curl | bash). All referenced URLs point to legitimate academic research (arXiv) and reputable technology blogs.
- [PROMPT_INJECTION]: The skill provides templates for standard evaluation tasks and does not contain instructions aimed at bypassing AI safety filters or overriding system constraints.
- [COMMAND_EXECUTION]: The file contains no shell commands, scripts, or instructions that would cause the agent to execute arbitrary code on the host system.
Audit Metadata