advanced-evaluation

Pass

Audited by Gen Agent Trust Hub on Aug 1, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill serves as a technical educational resource and implementation guide for LLM evaluation. No malicious code, prompt injections, or unauthorized data access patterns were found.\n- [EXTERNAL_DOWNLOADS]: The skill references established research papers on arxiv.org and reputable industry blogs. These references are strictly for documentation and research purposes and do not involve executable downloads.\n- [DATA_EXFILTRATION]: No network exfiltration or unauthorized file access was detected. The provided Python scripts perform local calculations for evaluation metrics such as weighted scores and bias indicators.\n- [PROMPT_INJECTION]: The prompt templates provided are specifically structured for evaluation tasks (direct scoring and pairwise comparison) and do not contain instructions to bypass safety guardrails or override system prompts.\n- [REMOTE_CODE_EXECUTION]: No remote code execution or dynamic command execution patterns were found. The Python examples demonstrate logic for data processing and statistical analysis using standard libraries.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 1, 2026, 03:31 AM
Security Audit — agent-trust-hub — advanced-evaluation