openrlhf-training

Warn

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: MEDIUMPRIVILEGE_ESCALATIONCOMMAND_EXECUTIONDYNAMIC_EXECUTIONEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PRIVILEGE_ESCALATION]: The skill instructs the user to launch a Docker container with the SYS_ADMIN capability (--cap-add=SYS_ADMIN), which grants the container extensive privileges on the host system, typically used for managing GPU resources but increasing the attack surface.
  • [COMMAND_EXECUTION]: Instructions include the use of sudo pip uninstall, which requires administrative privileges to modify system-level Python environments.
  • [DYNAMIC_EXECUTION]: The framework relies on loading and executing arbitrary local Python scripts to define custom reward functions (--remote_rm_url) and multi-step agent interactions (--agent_func_path).
  • [DYNAMIC_EXECUTION]: Documentation in references/custom-rewards.md includes an example (reward_func_code_gen.py) that writes model-generated responses to a temporary file and executes them using subprocess.run. While a timeout is implemented, this pattern involves executing untrusted code locally.
  • [EXTERNAL_DOWNLOADS]: The skill fetches official Docker images from NVIDIA's registry (nvcr.io) and installs the openrlhf package and its dependencies from standard registries.
  • [INDIRECT_PROMPT_INJECTION]: The skill exhibits an attack surface where untrusted data from training datasets or model completions is processed by custom logic.
  • Ingestion points: Training datasets (e.g., OpenRLHF/preference_dataset_mixture2_and_safe_pku) and generated completions in reward_func or AgentInstance (found in references/custom-rewards.md).
  • Boundary markers: None present; the logic does not explicitly delimit model output from control instructions in the reward evaluation scripts.
  • Capability inventory: Access to shell command execution (subprocess.run), file system operations (tempfile), and distributed network communication via Ray and DeepSpeed.
  • Sanitization: Absent in provided examples; model-generated code is executed directly via pytest without isolation or sandboxing.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — openrlhf-training