evaluating-llms-harness
Warn
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: MEDIUMDYNAMIC_EXECUTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONCREDENTIALS_UNSAFEEXTERNAL_DOWNLOADS
Full Analysis
- [DYNAMIC_EXECUTION]: The skill utilizes a dynamic loading mechanism described in
custom-tasks.md, where the!functiontag in YAML configuration files allows the evaluation harness to load and execute arbitrary Python functions from a localutils.pyfile at runtime. - [COMMAND_EXECUTION]: The documentation includes multiple Python snippets (e.g., within
SKILL.mdtraining callbacks) that useos.system()to execute shell commands for model evaluation. Additionally,custom-tasks.mdprovides examples of using thesubprocessmodule to execute code against test cases. - [INDIRECT_PROMPT_INJECTION]: The skill describes the use of the
--allow_code_executionflag for benchmarks like HumanEval. This enables the harness to execute code generated by the model under test, creating a potential vector where a model producing malicious code could compromise the execution environment if the flag is enabled. - Ingestion points: Model-generated code outputs (HumanEval) and external datasets loaded from HuggingFace.
- Boundary markers: Standard prompt delimiters used by the harness, though the skill does not explicitly define custom sanitization logic.
- Capability inventory: Shell command execution (
os.system), arbitrary code execution (allow_code_execution), network access (API providers), and file system writes (result logging). - Sanitization: Relies on the underlying harness logic; the
allow_code_executionflag explicitly bypasses typical safety constraints for code generation tasks. - [CREDENTIALS_UNSAFE]: The
api-evaluation.mdguide instructs users to export sensitive API keys (OpenAI, Anthropic) into environment variables. While standard for API usage, this exposes credentials to the shell environment and any processes with access to environment variables. - [EXTERNAL_DOWNLOADS]: The skill facilitates downloading datasets and model weights from HuggingFace and installing dependencies from PyPI (e.g.,
lm-eval,vllm,human-eval). These originate from well-known and trusted platforms.
Audit Metadata