visual-regression

Pass

Audited by Gen Agent Trust Hub on Mar 23, 2026

Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSREMOTE_CODE_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill executes shell commands to install dependencies and run tests using tools such as npm, npx, and the Flutter CLI (e.g., npm install -D backstopjs, npx playwright test).- [EXTERNAL_DOWNLOADS]: Fetches and installs third-party testing frameworks including Playwright, BackstopJS, and Chromatic from the official npm registry.- [REMOTE_CODE_EXECUTION]: Generates test scripts dynamically based on the detected project structure and routes, which are subsequently executed by the agent to perform visual comparisons.- [PROMPT_INJECTION]: Employs directive language such as "You are in AUTONOMOUS MODE" and "Do NOT ask questions" to override default interactive agent behavior during task execution.- [INDIRECT_PROMPT_INJECTION]: Vulnerability surface identified: The skill ingests untrusted project data (route definitions, package.json) and interpolates it into generated test code without explicit sanitization or boundary markers. Capability inventory includes file system writes and shell command execution.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 23, 2026, 10:58 AM
Security Audit — agent-trust-hub — visual-regression