testing-webapps

Pass

Audited by Gen Agent Trust Hub on Sep 13, 2026

Risk Level: SAFECOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONMETADATA_POISONING
Full Analysis
  • [DYNAMIC_EXECUTION]: The script scripts/with_server.py uses subprocess.Popen with shell=True to execute commands passed via the --server argument. This allows for the execution of arbitrary shell commands, including those involving shell features like pipes or logical operators (e.g., cd && npm start).
  • [COMMAND_EXECUTION]: The skill facilitates the execution of local shell commands through scripts/with_server.py to manage server lifecycles and playwright install to manage browser binaries.
  • [METADATA_POISONING]: The SKILL.md file contains instructions explicitly telling the agent 'DO NOT read the source until you try running the script first,' claiming the scripts are 'very large' and would 'pollute your context window.' This is deceptive, as scripts/with_server.py is approximately 100 lines long, and such instructions may prevent the agent from identifying execution risks like the use of shell=True before invocation.
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and process external, untrusted content from local or remote web applications, creating a vulnerability surface for indirect prompt injection.
  • Ingestion points: Web page content is captured via page.content(), screenshots are taken using page.screenshot(), and DOM elements are inspected in examples/element_discovery.py and examples/console_logging.py.
  • Boundary markers: None identified; instructions do not include specific delimiters or 'ignore embedded instructions' warnings for processed web content.
  • Capability inventory: The skill can execute shell commands via scripts/with_server.py, write files to the system (e.g., screenshots and logs in examples/), and control browser behavior via Playwright.
  • Sanitization: No sanitization or filtering of external HTML or console log content is implemented before it is processed by the agent.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 13, 2026, 03:48 PM
Security Audit — agent-trust-hub — testing-webapps