loki-mode
Fail
Audited by Gen Agent Trust Hub on Sep 14, 2026
Risk Level: HIGHPROMPT_INJECTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
- [PROMPT_INJECTION]: The skill instructions explicitly override the AI agent's interactive safety mechanisms.
SKILL.mdcommands the agent to 'NEVER ask questions', 'NEVER wait for confirmation', and 'NEVER stop voluntarily', effectively neutralizing the human-in-the-loop safety model. - [COMMAND_EXECUTION]: Operation of this skill requires the
--dangerously-skip-permissionsflag, which grants the multi-agent swarm unrestricted shell access and the ability to execute arbitrary commands autonomously. This lack of oversight allows for potentially catastrophic system actions if the agent is misled by instructions. - [INDIRECT_PROMPT_INJECTION]: The skill is highly vulnerable to indirect prompt injection through its primary data input, the Product Requirements Document (PRD).
- Ingestion points: External PRD files are read directly into the agent context via the orchestrator in
autonomy/run.sh. - Boundary markers: The instructions fail to provide delimiters or instructions to ignore potential commands embedded within the untrusted PRD text.
- Capability inventory: The system maintains full system access via the
Bashtool and can programmatically spawn new agents with specialized roles using theTasktool, providing a powerful vector for persistent or complex malicious actions. - Sanitization: There is no evidence of input validation or content filtering for the data ingested from the PRDs.
- [DYNAMIC_EXECUTION]: Multiple files within the benchmark results directory (e.g.,
benchmarks/results/2026-01-05-00-49-17/humaneval-solutions/160.py) contain implementations using theeval()function for dynamic expression evaluation. This reflects a development pattern within the skill's domain that can lead to remote code execution vulnerabilities in generated code.
Recommendations
- AI detected serious security threats
Audit Metadata