evaluating-code-models
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: Fetches the evaluation harness and associated Docker images from the BigCode Project's official GitHub and GitHub Container Registry. As an industry-standard project, these sources are recognized as well-known repositories.
- [DYNAMIC_EXECUTION]: The skill's primary function is to execute code generated by AI models to evaluate functional correctness (pass@k metrics). This requires the
--allow_code_executionflag, which runs arbitrary model output on the host system or within a container. - [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted model generations and executes them. This creates a surface for indirect prompt injection where a model could be influenced to generate malicious code that performs unauthorized actions during the evaluation process.
- Ingestion points: Model outputs generated during benchmarking tasks (SKILL.md, benchmarks.md).
- Boundary markers: Uses task-specific stop words (e.g.,
\ndef,\nclass) to delimit generated code (custom-tasks.md). - Capability inventory: Arbitrary code execution via
main.pywhen--allow_code_executionis enabled. - Sanitization: Recommends the use of Docker containers to isolate the execution environment and mitigate host system impact (SKILL.md, issues.md).
- [COMMAND_EXECUTION]: Provides instructions for running shell commands to clone repositories, install dependencies via
pip, and launch evaluation scripts usingaccelerate.
Audit Metadata