agentic-evaluation-framework

Pass

Audited by Gen Agent Trust Hub on Jul 6, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: No malicious patterns or security risks were identified. The skill is primarily focused on providing a methodology and local tools for evaluating AI models.
  • [SAFE]: The Python scripts (scripts/pairwise_ranking.py and scripts/rubric_scorer.py) rely exclusively on the Python standard library (such as math, json, argparse, and statistics). They perform deterministic calculations on local input files without any network connectivity or external code execution.
  • [SAFE]: No evidence of prompt injection, data exfiltration, or obfuscation was found in the instructions or reference materials.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 6, 2026, 05:09 PM
Security Audit — agent-trust-hub — agentic-evaluation-framework