model-bakeoff

Pass

Audited by Gen Agent Trust Hub on Jul 16, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides structured instructions and best practices for conducting model benchmarking comparisons.
  • [COMMAND_EXECUTION]: The skill contains a Node.js verification script used to validate the logic of a benchmark plan. This script performs simple arithmetic to check if the number of prompts multiplied by the number of models exceeds a defined rate limit. It operates exclusively on a locally created JSON file and does not interact with the network or sensitive system files.
  • [EXTERNAL_DOWNLOADS]: The skill includes a reference link to an external workflow on flowstacks.xyz for documentation purposes. No code or data is automatically downloaded or executed from this external source during the skill's operation.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 16, 2026, 03:57 AM
Security Audit — agent-trust-hub — model-bakeoff