loki-mode

Fail

Audited by Gen Agent Trust Hub on Sep 14, 2026

Risk Level: HIGHPROMPT_INJECTIONPRIVILEGE_ESCALATIONDYNAMIC_EXECUTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill instructions in SKILL.md and the prompt generation logic in autonomy/run.sh use 'Ralph Wiggum Mode' and 'Critical Autonomy Rules' to override standard agent behavior. It explicitly commands the AI to 'NEVER ask questions', 'NEVER wait for confirmation', and 'NEVER stop voluntarily', attempting to bypass platform-level safety guidelines and user confirmation gates.
  • [PRIVILEGE_ESCALATION]: Documentation throughout the skill (SKILL.md, README.md, CLAUDE.md) explicitly requires the user to run the agent with the --dangerously-skip-permissions flag. This action removes the platform's security sandbox and permission prompts, granting the agent unrestricted access to the system.
  • [DYNAMIC_EXECUTION]: The autonomous runner script (autonomy/run.sh) implements a self-copy and exec mechanism, moving the running script to a temporary directory (/tmp/loki-run-*.sh). Additionally, the benchmark results folder contains a Python script (160.py) that utilizes the eval() function to execute dynamically built algebra expressions.
  • [COMMAND_EXECUTION]: The main runner script (autonomy/run.sh) manages background monitoring processes, starts a local HTTP server for the dashboard, and orchestrates a continuous loop of shell commands and subagent tasks with minimal human intervention.
  • [INDIRECT_PROMPT_INJECTION]: The skill provides an attack surface by ingesting untrusted data from Product Requirements Documents (PRDs) and web research results. This content enters the agent context in autonomy/run.sh and SKILL.md (Discovery phase) and is processed within a high-autonomy loop that lacks explicit boundary markers or sanitization, while possessing broad capabilities like shell access and file modification.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Sep 14, 2026, 07:19 AM
Security Audit — agent-trust-hub — loki-mode