skills/lev-os/agents/eval-builder/Gen Agent Trust Hub

eval-builder

Pass

Audited by Gen Agent Trust Hub on Jul 16, 2026

Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill instructs the agent to use various CLI tools including lev eval, lev flowmind, and lev exec. It also provides validation shell commands using test, find, and rg to verify the presence of project files and configuration.
  • [DATA_EXFILTRATION]: The skill requires access to sensitive project files such as .lev/config.yaml, dna/gates.yaml, and AGENTS.md to establish the necessary context for building evaluation suites. This access is localized to the project environment for configuration resolution.
  • [PROMPT_INJECTION]: The skill is inherently designed to facilitate the creation of adversarial tests, including prompt injection probes, for other agentic features. This creates an indirect prompt injection surface as it processes external system descriptions.
  • Ingestion points: Processes 'any topic or system from the supplied context', project DNA, and configuration files.
  • Boundary markers: Utilizes structured YAML (LFD contracts) to define data, but lacks explicit delimiters or instructions to ignore embedded prompts within the systems it evaluates.
  • Capability inventory: Possesses the ability to write to the local filesystem (receipts/proofs) and execute shell commands via the lev tool suite.
  • Sanitization: Promotes the use of 'typed observation payloads' and 'deterministic scoring' to validate agent outputs, though it does not provide specific sanitization for the inputs it ingests during the design phase.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 16, 2026, 04:48 PM
Security Audit — agent-trust-hub — eval-builder