computer-use-agents
Pass
Audited by Gen Agent Trust Hub on Sep 14, 2026
Risk Level: SAFECOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill implements desktop automation using the
pyautoguilibrary, allowing the agent to perform actions such as clicking, typing text, and pressing keys on the host or virtual system. - [DYNAMIC_EXECUTION]: The implementation includes tool definitions for a vision-language model to execute shell commands (
bash) and perform file system operations via a text editor, which involves runtime command generation and execution. - [INDIRECT_PROMPT_INJECTION]: The Perception-Reasoning-Action loop ingests screenshots of the desktop environment to drive agent decisions. This architecture is susceptible to indirect prompt injection if the agent processes screens containing malicious instructions (e.g., from a browser window or untrusted document).
- Ingestion points:
pyautogui.screenshot()inSKILL.mdprovides visual input of the current desktop state to the model. - Boundary markers: No specific prompt delimiters or instructions to ignore embedded visual commands are implemented in the provided logic.
- Capability inventory: The skill exposes full GUI control (
pyautogui) and shell command execution (subprocess) capabilities to the agent. - Sanitization: There is no evidence of filtering, validation, or sanitization of the visual data before it is interpreted by the reasoning model.
Audit Metadata