model-bakeoff
Pass
Audited by Gen Agent Trust Hub on Jul 16, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides structured instructions and best practices for conducting model benchmarking comparisons.
- [COMMAND_EXECUTION]: The skill contains a Node.js verification script used to validate the logic of a benchmark plan. This script performs simple arithmetic to check if the number of prompts multiplied by the number of models exceeds a defined rate limit. It operates exclusively on a locally created JSON file and does not interact with the network or sensitive system files.
- [EXTERNAL_DOWNLOADS]: The skill includes a reference link to an external workflow on flowstacks.xyz for documentation purposes. No code or data is automatically downloaded or executed from this external source during the skill's operation.
Audit Metadata