physicalai-train-benchmarking-a-policy

Pass

Audited by Gen Agent Trust Hub on Jul 11, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides documentation for benchmarking trained policies using the physicalai library. All operations and code snippets follow standard development practices for model evaluation.
  • [COMMAND_EXECUTION]: The skill involves the execution of the physicalai benchmark CLI tool and the pytest testing framework. These commands are used for their intended purpose of running benchmarks and unit tests on local project files.
  • [SAFE]: No evidence of data exfiltration, hardcoded credentials, or malicious network activity was detected. The skill operates on local directories (e.g., experiments/, results/) for loading checkpoints and saving metrics.
  • [SAFE]: No obfuscation techniques, such as Base64 encoding of commands or hidden Unicode characters, are present in the documentation or code examples.
  • [SAFE]: External dependencies mentioned (like libero or robocasa) are described as optional extras for specific simulation environments, which is standard for modular ML libraries.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 11, 2026, 12:18 AM
Security Audit — agent-trust-hub — physicalai-train-benchmarking-a-policy