benchmark-models

Fail

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: HIGHCOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill executes shell commands using subprocess.run with a command string constructed via f-string interpolation. Specifically, the model parameter, which is a user-provided command-line argument, is directly embedded into a bash -c command without sanitization. This allows an attacker to inject arbitrary shell commands (e.g., by providing a model name like 'dummy; malicious_command;').
  • Evidence: In run.sh, the embedded Python script contains the run_live function which executes result = subprocess.run(["bash", "-c", f"... --model {model} ..."], ...).
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Mar 17, 2026, 06:34 AM
Security Audit — agent-trust-hub — benchmark-models