torchforge-rl-training
Fail
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: HIGHMETADATA_POISONINGEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [METADATA_POISONING]: The skill's description and body claim that 'torchforge is Meta's PyTorch-native RL library.' However, the referenced resources point to 'meta-pytorch.org' and 'github.com/meta-pytorch,' which are not recognized as official domains or repositories for Meta or the PyTorch project. This indicates potential impersonation designed to mislead users regarding the software's origin and safety.
- [EXTERNAL_DOWNLOADS]: The instructions encourage users to download code from unverified external sources (e.g., github.com/meta-pytorch/torchforge) and execute local installation scripts (
./scripts/install.sh) that rely on these untrusted repositories. - [COMMAND_EXECUTION]: The skill provides instructions for the agent to perform high-privilege operations, such as executing installation scripts, running distributed training processes, and submitting Slurm jobs (
sbatch). These actions are performed using a library from an unverified source that is masquerading as a major vendor. - [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and process external datasets from Hugging Face (
openai/gsm8k) within its training workflows. The instructions lack boundary markers or sanitization requirements for this untrusted data, creating a surface where malicious content in the training data could attempt to influence the agent's behavior during high-privilege execution phases. - Ingestion points: Dataset defined in
SKILL.md(e.g., path: 'openai/gsm8k'). - Boundary markers: Absent; no instructions for using delimiters or ignoring embedded commands.
- Capability inventory: Shell script execution, multi-process Python training, and Slurm job submission (SKILL.md).
- Sanitization: Absent; no instructions for validating or escaping dataset content.
Recommendations
- AI detected serious security threats
Audit Metadata