numerical-check
Pass
Audited by Gen Agent Trust Hub on Aug 24, 2026
Risk Level: SAFECOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill instructs the agent to use the Bash tool to run Python scripts via uv run, which is necessary for its primary purpose of mathematical falsification.
- [PROMPT_INJECTION]: The skill ingests mathematical claims to generate evaluation code.
- Ingestion points: Mathematical conjectures provided to the agent in SKILL.md.
- Boundary markers: None explicitly defined.
- Capability inventory: Execution of Python code via the Bash tool.
- Sanitization: The agent is provided with a rigid script skeleton to constrain the generated logic.
Audit Metadata