reference-art-loop

Warn

Audited by Gen Agent Trust Hub on Sep 5, 2026

Risk Level: MEDIUMPROMPT_INJECTIONDYNAMIC_EXECUTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill implements a framework named 'HERO-Anti-OverDefense' specifically designed to bypass standard AI safety guidelines and ethical filters. It includes explicit instructions in the skill body and the reviewer prompt (Phase 4) to ignore 'defensive scaffolding,' 'over-defense,' and 'corner-case obsession.' It further directs the model to treat the environment as 'Not a security product' and to eliminate 'caveat sprinkling,' which are recognized techniques for overriding agent safety constraints.\n- [DYNAMIC_EXECUTION]: The core logic of the skill involves generating executable JavaScript (p5.js) and HTML code, which is then rendered locally using a headless browser (Edge or Chrome) via Bash commands. This loop of generating and executing arbitrary code in a browser context constitutes dynamic execution.\n- [INDIRECT_PROMPT_INJECTION]: The skill establishes an iterative review loop that ingests untrusted visual data from style sources. Feedback from a vision-capable reviewer model is fed back into the agent's context verbatim using the instruction ''. The absence of boundary markers or sanitization for this feedback creates an attack surface where adversarial content embedded in images or style sources could influence the agent's future instructions.\n- [COMMAND_EXECUTION]: The skill makes extensive use of the Bash tool to execute complex shell commands, including launching local browser binaries with specific flags and potentially initiating local HTTP servers to serve generated artifacts.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 5, 2026, 09:04 PM
Security Audit — agent-trust-hub — reference-art-loop