application-quality-assurance

Warn

Audited by Gen Agent Trust Hub on Sep 23, 2026

Risk Level: MEDIUMCOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The helper script scripts/with_server.py utilizes subprocess.Popen with shell=True to execute server start commands provided via the --server argument. This implementation allows for arbitrary shell command execution, including command chaining (e.g., using ;, &&, or |) with parameters supplied by the agent.
  • [DYNAMIC_EXECUTION]: The skill facilitates the execution of arbitrary commands. The script scripts/with_server.py is designed to take a sequence of arguments representing a command and execute it via subprocess.run after the servers have started.
  • [INDIRECT_PROMPT_INJECTION]: The skill is intended to interact with and extract data from external web applications, creating an attack surface where a malicious site could inject instructions to manipulate the agent's behavior.
  • Ingestion points: The skill reads application state via page.content() and page.locator().all() as described in examples/element_discovery.py and the SKILL.md decision tree.
  • Boundary markers: Absent. The instructions do not provide delimiters or warnings to the agent to treat external page content as potentially untrusted data.
  • Capability inventory: The skill has access to shell execution via scripts/with_server.py and filesystem writes in the examples/ directory.
  • Sanitization: Content retrieved from web applications is processed directly without sanitization or validation logic.
  • [PROMPT_INJECTION]: The SKILL.md file contains instructions that discourage the agent from inspecting the source code of the provided helper scripts ("DO NOT read the source until you try running the script first"). This functions as a concealment tactic that attempts to limit the agent's awareness of the script's internal execution logic, such as the use of shell=True.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 23, 2026, 04:59 PM
Security Audit — agent-trust-hub — application-quality-assurance