monitor-experiment

Warn

Audited by Gen Agent Trust Hub on Sep 15, 2026

Risk Level: MEDIUMDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [DYNAMIC_EXECUTION]: The skill assembles and executes Python scripts on remote servers at runtime using the ssh <server> "python3 -c ..." pattern. This is primarily used in Step 3.5 to interface with the Weights & Biases API.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes data from external, potentially untrusted sources including remote screen logs and experiment result files.
  • Ingestion points: Remote screen hardcopy files and JSON result files retrieved via SSH in Step 2 and Step 3.
  • Boundary markers: The instructions lack delimiters or explicit boundary markers to prevent the agent from following instructions embedded in experiment logs or results.
  • Capability inventory: The skill is granted Bash(ssh *) and Write capabilities, which could be leveraged if the agent attempts to execute commands found in the ingested data.
  • Sanitization: There is no evidence of sanitization, filtering, or validation performed on the remote output before it is summarized or interpreted.
  • [COMMAND_EXECUTION]: The skill uses SSH to execute various commands on remote infrastructure, including screen management, file system operations (ls, cat, tail), and Python script execution.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 15, 2026, 02:07 PM
Security Audit — agent-trust-hub — monitor-experiment