testing-webapps
Pass
Audited by Gen Agent Trust Hub on Sep 13, 2026
Risk Level: SAFECOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONMETADATA_POISONING
Full Analysis
- [DYNAMIC_EXECUTION]: The script
scripts/with_server.pyusessubprocess.Popenwithshell=Trueto execute commands passed via the--serverargument. This allows for the execution of arbitrary shell commands, including those involving shell features like pipes or logical operators (e.g.,cd && npm start). - [COMMAND_EXECUTION]: The skill facilitates the execution of local shell commands through
scripts/with_server.pyto manage server lifecycles andplaywright installto manage browser binaries. - [METADATA_POISONING]: The
SKILL.mdfile contains instructions explicitly telling the agent 'DO NOT read the source until you try running the script first,' claiming the scripts are 'very large' and would 'pollute your context window.' This is deceptive, asscripts/with_server.pyis approximately 100 lines long, and such instructions may prevent the agent from identifying execution risks like the use ofshell=Truebefore invocation. - [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and process external, untrusted content from local or remote web applications, creating a vulnerability surface for indirect prompt injection.
- Ingestion points: Web page content is captured via
page.content(), screenshots are taken usingpage.screenshot(), and DOM elements are inspected inexamples/element_discovery.pyandexamples/console_logging.py. - Boundary markers: None identified; instructions do not include specific delimiters or 'ignore embedded instructions' warnings for processed web content.
- Capability inventory: The skill can execute shell commands via
scripts/with_server.py, write files to the system (e.g., screenshots and logs inexamples/), and control browser behavior via Playwright. - Sanitization: No sanitization or filtering of external HTML or console log content is implemented before it is processed by the agent.
Audit Metadata