run-experiment
Pass
Audited by Gen Agent Trust Hub on Sep 15, 2026
Risk Level: SAFECOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONCREDENTIALS_UNSAFE
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes configuration data from external files located in the project workspace, which acts as a vulnerability surface for indirect instructions. * Ingestion points: Reads
CLAUDE.mdandvast-instances.jsonfrom the project root to extract server aliases, paths, environment variables, and GPU configurations. * Boundary markers: None. The skill does not implement delimiters or instruct the agent to ignore instructions embedded within these files. * Capability inventory: UsesBash(*),Write,Edit, andssh(via shell) to perform sensitive actions based on the ingested data. * Sanitization: No validation or sanitization is performed on values retrieved from these files before they are interpolated into shell commands. - [COMMAND_EXECUTION]: Shell commands are dynamically constructed using variables sourced from untrusted workspace files. * Evidence: The
<server>,<conda_path>, and<script>variables extracted fromCLAUDE.mdare interpolated directly intossh,rsync, andscreencommand strings, enabling potential command injection. - [DYNAMIC_EXECUTION]: The skill generates and modifies executable source code prior to runtime execution. * Evidence: Step 3.5 instructs the agent to modify existing training scripts by injecting
wandbinitialization and logging code via theEdittool. * Evidence: Step 4 (Modal) generates amodal_launcher.pyscript which is subsequently executed using themodal runcommand. - [CREDENTIALS_UNSAFE]: The skill manages sensitive configuration and authentication information in potentially insecure ways. * Evidence: It reads notification settings from
~/.claude/feishu.json. * Evidence: It suggests retrieving aWANDB_API_KEYfromCLAUDE.mdforwandb login, which is a risky location for secrets as it may be committed to shared version control systems.
Audit Metadata