lean4-prove
Fail
Audited by Gen Agent Trust Hub on Mar 17, 2026
Risk Level: HIGHCREDENTIALS_UNSAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSDATA_EXFILTRATION
Full Analysis
- [CREDENTIALS_UNSAFE]: The skill contains hardcoded default credentials for ArangoDB.
- In
pilot_formalize.py, the variableARANGO_PASSis hardcoded to"openSesame". - In
run_formalization_benchmark.py, the ArangoDB client is initialized withusername="root"andpassword="openSesame". - [DATA_EXFILTRATION]: The skill accesses highly sensitive authentication files.
- In
sanity.sh, the script reads~/.claude/.credentials.jsonto verify and inspect Claude OAuth tokens, specifically extracting theexpiresAtfield. While used here for health checks, this provides a pathway for unauthorized access to the user's Claude account credentials. - [COMMAND_EXECUTION]: The skill performs extensive arbitrary command execution and container manipulation.
- Multiple files (
prove.py,extract_lemma_deps.py,ingest_prover_v1.py) usesubprocess.runto executedocker execcommands, allowing for the execution of code inside containers. prove.pycalls theclaudeCLI usingsubprocess.runin a non-interactive mode.formalize.pyandfull_formalize.pyexecute/scillm/run.shvia subprocess to generate code.- [EXTERNAL_DOWNLOADS]: The skill fetches data and scripts from remote sources.
- The
Dockerfileexecutes a remote script directly from GitHub:curl -sSf https://raw.githubusercontent.com/leanprover/elan/master/elan-init.sh | bash. - Multiple ingestion scripts (
ingest_autoformalization.py,ingest_prover_v1.py,ingest_prover_v2.py) download datasets from HuggingFace using thedatasetslibrary. - [INDIRECT_PROMPT_INJECTION]: The skill exhibits a significant attack surface for indirect prompt injection.
- Ingestion points: Natural language requirements are accepted via CLI arguments or
stdininprove.pyandrun.sh. - Capability inventory: The skill uses
subprocess.runanddocker execto compile and execute Lean4 code generated from these requirements. - Sanitization: There is no meaningful sanitization of the input text before it is interpolated into LLM prompts in
formalize.pyandprove.py.
Recommendations
- HIGH: Downloads and executes remote code from: unknown (check file) - DO NOT USE without thorough review
- AI detected serious security threats
Audit Metadata