rock-eval

Fail

Audited by Gen Agent Trust Hub on Jun 30, 2026

Risk Level: HIGHREMOTE_CODE_EXECUTIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
  • [REMOTE_CODE_EXECUTION]: The skill's documentation and instructions promote the use of remote shell scripts for tool installation. Specifically, it provides commands to fetch and execute scripts from xrl.alibaba-inc.com using patterns like bash -c "$(curl -fsSL http://xrl.alibaba-inc.com/install.sh)". Although this involves direct execution of remote code, the domain belongs to a well-known technology company and represents the tool's official infrastructure.
  • [COMMAND_EXECUTION]: The skill facilitates the execution of arbitrary commands and scripts within isolated sandbox environments via the rc sandbox exec command. This is a core functionality required to run the benchmarks and regression tests described in the skill.
  • [EXTERNAL_DOWNLOADS]: The skill makes several requests to external domains to download configuration files, benchmarks, and installation scripts, primarily targeting the alibaba-inc.com domain.
  • [PROMPT_INJECTION]: The skill implements a "Deep Analysis" phase where it ingests and processes "agent trajectories"—logs of interactions from other AI agents. This creates an indirect prompt injection surface because the analyzed logs could contain malicious instructions designed to influence the analyzing agent. The ingestion points include trajectory JSON files and result summaries. The capability inventory of the skill includes starting sandboxes and executing arbitrary commands, which could be exploited if an injection is successful. While the skill uses structured report templates to manage this data, it does not include explicit safety guidelines for handling potentially adversarial content within these logs.
Recommendations
  • HIGH: Downloads and executes remote code from: http://xrl.alibaba-inc.com/install.sh, http://xrl.alibaba-inc.com/install_beta.sh - DO NOT USE without thorough review
Audit Metadata
Risk Level
HIGH
Analyzed
Jun 30, 2026, 12:05 PM
Security Audit — agent-trust-hub — rock-eval