loki-mode
Fail
Audited by Gen Agent Trust Hub on Sep 14, 2026
Risk Level: HIGHPROMPT_INJECTIONPRIVILEGE_ESCALATIONDYNAMIC_EXECUTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill instructions in
SKILL.mdand the prompt generation logic inautonomy/run.shuse 'Ralph Wiggum Mode' and 'Critical Autonomy Rules' to override standard agent behavior. It explicitly commands the AI to 'NEVER ask questions', 'NEVER wait for confirmation', and 'NEVER stop voluntarily', attempting to bypass platform-level safety guidelines and user confirmation gates. - [PRIVILEGE_ESCALATION]: Documentation throughout the skill (
SKILL.md,README.md,CLAUDE.md) explicitly requires the user to run the agent with the--dangerously-skip-permissionsflag. This action removes the platform's security sandbox and permission prompts, granting the agent unrestricted access to the system. - [DYNAMIC_EXECUTION]: The autonomous runner script (
autonomy/run.sh) implements a self-copy andexecmechanism, moving the running script to a temporary directory (/tmp/loki-run-*.sh). Additionally, the benchmark results folder contains a Python script (160.py) that utilizes theeval()function to execute dynamically built algebra expressions. - [COMMAND_EXECUTION]: The main runner script (
autonomy/run.sh) manages background monitoring processes, starts a local HTTP server for the dashboard, and orchestrates a continuous loop of shell commands and subagent tasks with minimal human intervention. - [INDIRECT_PROMPT_INJECTION]: The skill provides an attack surface by ingesting untrusted data from Product Requirements Documents (PRDs) and web research results. This content enters the agent context in
autonomy/run.shandSKILL.md(Discovery phase) and is processed within a high-autonomy loop that lacks explicit boundary markers or sanitization, while possessing broad capabilities like shell access and file modification.
Recommendations
- AI detected serious security threats
Audit Metadata