thunderdome
Fail
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: HIGHCOMMAND_EXECUTIONREMOTE_CODE_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: In
scripts/dispatch.py, the function_build_benchmark_cmdconstructs shell commands using f-strings that interpolateStrategyfields such asnameandbackbones. These fields are loaded directly from a YAML manifest. Because these values are not escaped or sanitized before being passed toasyncio.create_subprocess_execviabash -c, an attacker can provide a malicious manifest to execute arbitrary shell commands. - [COMMAND_EXECUTION]: In
scripts/research.py, thedogpile_researchfunction builds a shell command usingmanifest.descriptionormanifest.name. This command is executed usingsubprocess.runwithbash -lc. A crafted description in the manifest (e.g., using backticks or semicolons) would lead to command injection and execution in the host environment. - [REMOTE_CODE_EXECUTION]: The identified command injection vulnerabilities represent a high risk of remote code execution. If the agent processes a manifest provided by an external or untrusted source, the underlying system could be fully compromised.
- [COMMAND_EXECUTION]: In
scripts/dispatch.py, the_build_benchmark_cmdfor text modality uses string interpolation for several parameters includingstrategy.backbones,strategy.epochs, andstrategy.namewithin a complex shell string. Thestrategy.nameis also used to construct a temporary model directory path, which is then used in apython3 -cone-liner, creating additional injection points.
Recommendations
- AI detected serious security threats
Audit Metadata