evaluating-cosmos-policy

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFE
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill instructs the user to clone official and public repositories for its functionality.
  • Fetches the official cosmos-policy repository from NVIDIA's GitHub (github.com/NVlabs/cosmos-policy.git).
  • Fetches a compatible fork of robocasa for policy evaluation (github.com/moojink/robocasa-cosmos-policy.git).
  • [COMMAND_EXECUTION]: The skill uses standard CLI tools (uv, python, git, srun, sbatch) to manage dependencies and execute evaluation scripts. These commands are localized to the project environment and do not demonstrate unauthorized behavior.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes simulation data (visual observations and robot proprioception) which is an ingestion surface. However, the use case is restricted to robot manipulation evaluation within controlled simulation environments, and the skill includes guidance for validating results.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:49 AM
Security Audit — agent-trust-hub — evaluating-cosmos-policy