leaderboard
Pass
Audited by Gen Agent Trust Hub on Aug 3, 2026
Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
- [SAFE]: The skill enforces robust data integrity by requiring the agent to hash submitted benchmark files against reference solutions and re-run grading scripts to verify that results are reproducible.
- [SAFE]: Explicit instructions are provided to sanitize submissions for Personally Identifiable Information (PII), such as checking for local file paths (/Users/, /home/) or usernames in transcripts.
- [COMMAND_EXECUTION]: Uses standard development tools including
git,uv, andpytest. These tools are employed for routine repository management, package synchronization, and testing within the local project context.
Audit Metadata