benchmark-models
Fail
Audited by Gen Agent Trust Hub on Mar 17, 2026
Risk Level: HIGHCOMMAND_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The skill executes shell commands using
subprocess.runwith a command string constructed via f-string interpolation. Specifically, themodelparameter, which is a user-provided command-line argument, is directly embedded into abash -ccommand without sanitization. This allows an attacker to inject arbitrary shell commands (e.g., by providing a model name like 'dummy; malicious_command;'). - Evidence: In
run.sh, the embedded Python script contains therun_livefunction which executesresult = subprocess.run(["bash", "-c", f"... --model {model} ..."], ...).
Recommendations
- AI detected serious security threats
Audit Metadata