evidence-case-lab
Pass
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: SAFECREDENTIALS_UNSAFECOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
- [CREDENTIALS_UNSAFE]: The file
generate_expansion.pycontains a hardcoded API keysk-dev-proxy-123used as a default value for theSCILLM_PROXY_KEYconfiguration. - [COMMAND_EXECUTION]: Multiple Python scripts within the skill execute shell commands using
subprocess.runto interact with local development tools and monitoring scripts: backfill_qra_evidence.pyexecutes a sibling script attask-monitor/run.shvia the_tmfunction.generate_stress_questions.pyusessubprocess.runto invoke thescillmutility throughuv run.- [PROMPT_INJECTION]: The skill ingests and processes question banks in JSON format which are interpolated into prompts for evaluation by an LLM-based plausibility gate. This creates a potential surface for indirect prompt injection.
- Ingestion points: Question banks are loaded from
questions.jsonandstate/questions_combined.jsoninevidence_case_lab.pyandrun_stress_test.py. - Boundary markers: There is no evidence of explicit prompt delimiters or instructions to prevent the agent from obeying directives embedded within the ingested question text.
- Capability inventory: The skill has permissions to write to local state and log files, execute shell commands through subprocesses, and interact with the ArangoDB and Memory daemon services via Unix Domain Sockets.
- Sanitization: The skill lacks input sanitization or filtering to identify and remove prompt injection patterns from the ingested question data.
Audit Metadata