create-evidence-case
Pass
Audited by Gen Agent Trust Hub on Mar 17, 2026
Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADS
Full Analysis
- [COMMAND_EXECUTION]: The skill orchestrates complex tasks by executing shell scripts in sibling skill directories (e.g.,
memory,assistant,extract-entities) usingsubprocess.runandsubprocess.Popenwithin therunner.py,storage.py, andreport.pymodules. - [EXTERNAL_DOWNLOADS]: The
runner.pymodule communicates with a local Lean4 compilation service athttp://localhost:8604/compilevia HTTP POST requests to verify formal claims. - [SAFE]: The skill incorporates an internal 'Adversarial Reasoning Framework' via a 7-gate decision process. This includes 'Gate 2b: Grounding', which specifically identifies fabricated identifiers and misspelled framework terms to prevent the agent from accepting and reasoning over malicious or hallucinated data.
Audit Metadata