evaluating-code-models

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: Fetches the evaluation harness and associated Docker images from the BigCode Project's official GitHub and GitHub Container Registry. As an industry-standard project, these sources are recognized as well-known repositories.
  • [DYNAMIC_EXECUTION]: The skill's primary function is to execute code generated by AI models to evaluate functional correctness (pass@k metrics). This requires the --allow_code_execution flag, which runs arbitrary model output on the host system or within a container.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted model generations and executes them. This creates a surface for indirect prompt injection where a model could be influenced to generate malicious code that performs unauthorized actions during the evaluation process.
  • Ingestion points: Model outputs generated during benchmarking tasks (SKILL.md, benchmarks.md).
  • Boundary markers: Uses task-specific stop words (e.g., \ndef, \nclass) to delimit generated code (custom-tasks.md).
  • Capability inventory: Arbitrary code execution via main.py when --allow_code_execution is enabled.
  • Sanitization: Recommends the use of Docker containers to isolate the execution environment and mitigate host system impact (SKILL.md, issues.md).
  • [COMMAND_EXECUTION]: Provides instructions for running shell commands to clone repositories, install dependencies via pip, and launch evaluation scripts using accelerate.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — evaluating-code-models