evidence-case-lab

Pass

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: SAFECREDENTIALS_UNSAFECOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [CREDENTIALS_UNSAFE]: The file generate_expansion.py contains a hardcoded API key sk-dev-proxy-123 used as a default value for the SCILLM_PROXY_KEY configuration.
  • [COMMAND_EXECUTION]: Multiple Python scripts within the skill execute shell commands using subprocess.run to interact with local development tools and monitoring scripts:
  • backfill_qra_evidence.py executes a sibling script at task-monitor/run.sh via the _tm function.
  • generate_stress_questions.py uses subprocess.run to invoke the scillm utility through uv run.
  • [PROMPT_INJECTION]: The skill ingests and processes question banks in JSON format which are interpolated into prompts for evaluation by an LLM-based plausibility gate. This creates a potential surface for indirect prompt injection.
  • Ingestion points: Question banks are loaded from questions.json and state/questions_combined.json in evidence_case_lab.py and run_stress_test.py.
  • Boundary markers: There is no evidence of explicit prompt delimiters or instructions to prevent the agent from obeying directives embedded within the ingested question text.
  • Capability inventory: The skill has permissions to write to local state and log files, execute shell commands through subprocesses, and interact with the ArangoDB and Memory daemon services via Unix Domain Sockets.
  • Sanitization: The skill lacks input sanitization or filtering to identify and remove prompt injection patterns from the ingested question data.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 26, 2026, 06:01 PM
Security Audit — agent-trust-hub — evidence-case-lab