agent-eval

Pass

Audited by Gen Agent Trust Hub on Aug 6, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill primarily consists of architectural documentation and best-practice guidelines for AI evaluation. The provided code snippets are instructional templates for scoring logic and regression gating, which do not contain malicious patterns, unauthorized network operations, or persistence mechanisms.
  • [COMMAND_EXECUTION]: The scripts/verify.sh script executes local file discovery and JSON validation commands to ensure golden-set integrity. It uses secure shell practices such as set -euo pipefail and restricts its scope to project-specific files, explicitly ignoring sensitive directories like .git, node_modules, and .venv.
  • [EXTERNAL_DOWNLOADS]: The skill references established evaluation frameworks including DeepEval, Inspect AI, and promptfoo. These are recognized tools in the AI development ecosystem. The skill suggests standard dependency management via pip install -r requirements.txt within CI environments.
  • [INDIRECT_PROMPT_INJECTION]: The skill defines a surface for processing external datasets in JSONL format.
  • Ingestion points: Dataset loading occurs in references/runner-and-gate.md via the load_cases function which parses *.jsonl files.
  • Boundary markers: The LLM-as-judge rubric templates in references/judge-design.md use structured headers (e.g., Context:, Answer:) to provide context to the judge model.
  • Capability inventory: Capabilities are limited to local file parsing, metric calculation, and LLM-based scoring; the skill does not grant the agent unsafe OS or network privileges.
  • Sanitization: While no explicit sanitization of the graded data is detailed, the use of rationale-forcing rubrics and pairwise comparisons helps mitigate accidental misinterpretation of adversarial content.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 6, 2026, 09:36 PM
Security Audit — agent-trust-hub — agent-eval