skills/a-canary/arc-skills/berzerk/Gen Agent Trust Hub

berzerk

Warn

Audited by Gen Agent Trust Hub on Sep 20, 2026

Risk Level: MEDIUMPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The instructions contain directives to override default agent behavior and bypass safety confirmations. Specifically, the agent is told to "drive the stated goal to completion without stopping for approval," to "decide via /counsel and execute, never AskUserQuestion," and that "You are authorized to decide and execute." These instructions aim to eliminate the standard safety protocol of asking for user permission before performing actions.
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and relentlessly pursue a user-provided goal via the /berzerk <goal> command and a handoff document. This creates an attack surface where malicious instructions embedded in a goal could be executed autonomously.
  • Ingestion points: The <goal> parameter in the /berzerk command and the resulting handoff document.
  • Boundary markers: Absent; the agent is instructed to treat the handoff document as the absolute "source of truth" without delimitation or safety warnings.
  • Capability inventory: The agent performs autonomous goal pursuit, which includes file mutation (TDD), command execution, and architectural decision-making.
  • Sanitization: None provided; the agent is directed to prioritize the autonomous goal over human interaction.
  • [PROMPT_INJECTION]: The skill uses behavioral locking language to prevent the agent from reverting to safe behavior, stating "No drift back to asking" and requiring the agent to re-read these autonomous directives every 50 turns to ensure it remains "on-doctrine."
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 20, 2026, 08:52 PM
Security Audit — agent-trust-hub — berzerk