thunderdome

Fail

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: HIGHCOMMAND_EXECUTIONREMOTE_CODE_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: In scripts/dispatch.py, the function _build_benchmark_cmd constructs shell commands using f-strings that interpolate Strategy fields such as name and backbones. These fields are loaded directly from a YAML manifest. Because these values are not escaped or sanitized before being passed to asyncio.create_subprocess_exec via bash -c, an attacker can provide a malicious manifest to execute arbitrary shell commands.
  • [COMMAND_EXECUTION]: In scripts/research.py, the dogpile_research function builds a shell command using manifest.description or manifest.name. This command is executed using subprocess.run with bash -lc. A crafted description in the manifest (e.g., using backticks or semicolons) would lead to command injection and execution in the host environment.
  • [REMOTE_CODE_EXECUTION]: The identified command injection vulnerabilities represent a high risk of remote code execution. If the agent processes a manifest provided by an external or untrusted source, the underlying system could be fully compromised.
  • [COMMAND_EXECUTION]: In scripts/dispatch.py, the _build_benchmark_cmd for text modality uses string interpolation for several parameters including strategy.backbones, strategy.epochs, and strategy.name within a complex shell string. The strategy.name is also used to construct a temporary model directory path, which is then used in a python3 -c one-liner, creating additional injection points.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Aug 26, 2026, 06:00 PM
Security Audit — agent-trust-hub — thunderdome