grpo-rl-training

Fail

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: CRITICALDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [DYNAMIC_EXECUTION]: The file examples/reward_functions_library.py includes a run_test_cases function that calls exec() on strings extracted from the model's own completions. Executing AI-generated code at runtime is a critical vulnerability because LLM outputs are untrusted and could contain malicious logic or commands designed to compromise the host system. While the code contains a warning about the need for sandboxing, the provided implementation directly executes the untrusted code in the current process.
  • [INDIRECT_PROMPT_INJECTION]: The skill architecture involves processing external datasets and model generations, which creates a vulnerability to indirect prompt injection. 1. Ingestion points: External training data loaded via datasets.load_dataset and completions generated by the model during training. 2. Boundary markers: The skill uses XML tags for structure, but lacks security-focused delimiters to separate instructions from data. 3. Capability inventory: The training process includes a high-risk capability via the exec() function in the reward library. 4. Sanitization: The skill does not perform sanitization of model outputs before execution, relying only on basic pattern matching.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
CRITICAL
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — grpo-rl-training