webapp-testing

Warn

Audited by Gen Agent Trust Hub on Aug 19, 2026

Risk Level: MEDIUMCOMMAND_EXECUTIONDATA_EXFILTRATIONPROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The script scripts/with_server.py accepts command-line arguments via the --server flag and executes them using subprocess.Popen with shell=True. This design allows for the execution of arbitrary shell commands, including those with shell features like cd and &&. If the agent passes unvalidated user-controlled input or data from a web page into this script, it could lead to arbitrary command execution on the host system.
  • [DATA_EXFILTRATION]: The skill uses the Playwright library for browser automation, which includes the capability to access the local filesystem using the file:// protocol. The script examples/static_html_automation.py demonstrates this by loading local HTML files. This capability could be abused to read sensitive files (e.g., credentials or configuration) if the agent is directed to a malicious file path.
  • [PROMPT_INJECTION]: The skill interacts with dynamic web content, which presents a surface for indirect prompt injection attacks where malicious instructions are embedded in the pages being tested.
  • Ingestion points: The agent ingests external data by reading page contents via page.content(), inspecting elements with page.locator(), and capturing visual state with page.screenshot(), as described in SKILL.md and examples/element_discovery.py.
  • Boundary markers: There are no explicit instructions or delimiters provided to the agent to treat rendered HTML content as untrusted or to ignore instructions contained within it.
  • Capability inventory: The agent has the ability to execute shell commands via scripts/with_server.py, write files to the local system (e.g., screenshots and logs), and perform outbound network requests via the browser.
  • Sanitization: The skill does not implement any filtering, sanitization, or validation of the HTML content before it is processed by the agent.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Aug 19, 2026, 05:31 PM
Security Audit — agent-trust-hub — webapp-testing