evaluating-llms-harness

Warn

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: MEDIUMDYNAMIC_EXECUTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONCREDENTIALS_UNSAFEEXTERNAL_DOWNLOADS
Full Analysis
  • [DYNAMIC_EXECUTION]: The skill utilizes a dynamic loading mechanism described in custom-tasks.md, where the !function tag in YAML configuration files allows the evaluation harness to load and execute arbitrary Python functions from a local utils.py file at runtime.
  • [COMMAND_EXECUTION]: The documentation includes multiple Python snippets (e.g., within SKILL.md training callbacks) that use os.system() to execute shell commands for model evaluation. Additionally, custom-tasks.md provides examples of using the subprocess module to execute code against test cases.
  • [INDIRECT_PROMPT_INJECTION]: The skill describes the use of the --allow_code_execution flag for benchmarks like HumanEval. This enables the harness to execute code generated by the model under test, creating a potential vector where a model producing malicious code could compromise the execution environment if the flag is enabled.
  • Ingestion points: Model-generated code outputs (HumanEval) and external datasets loaded from HuggingFace.
  • Boundary markers: Standard prompt delimiters used by the harness, though the skill does not explicitly define custom sanitization logic.
  • Capability inventory: Shell command execution (os.system), arbitrary code execution (allow_code_execution), network access (API providers), and file system writes (result logging).
  • Sanitization: Relies on the underlying harness logic; the allow_code_execution flag explicitly bypasses typical safety constraints for code generation tasks.
  • [CREDENTIALS_UNSAFE]: The api-evaluation.md guide instructs users to export sensitive API keys (OpenAI, Anthropic) into environment variables. While standard for API usage, this exposes credentials to the shell environment and any processes with access to environment variables.
  • [EXTERNAL_DOWNLOADS]: The skill facilitates downloading datasets and model weights from HuggingFace and installing dependencies from PyPI (e.g., lm-eval, vllm, human-eval). These originate from well-known and trusted platforms.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — evaluating-llms-harness