langgraph
Fail
Audited by Gen Agent Trust Hub on Apr 11, 2026
Risk Level: HIGHCOMMAND_EXECUTIONREMOTE_CODE_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The provided code example for a
calculatortool uses the Pythoneval()function on an input string directly supplied by the agent. This allows for arbitrary command execution within the application environment if the LLM generates or is manipulated into providing malicious code instead of a simple math expression. - [REMOTE_CODE_EXECUTION]: The combination of processing untrusted user input and a tool that executes arbitrary strings via
eval()creates a high risk for remote code execution. An attacker could use indirect prompt injection techniques to force the agent to call the calculator tool with malicious Python logic. - Ingestion points: Untrusted data enters the agent context through the
messageslist inAgentStateand thequeryfield inRouterState(SKILL.md). - Boundary markers: Absent; the examples do not use delimiters or system instructions to prevent the agent from obeying instructions embedded in user-supplied messages.
- Capability inventory: The skill includes the capability to execute arbitrary Python code via the
eval()call in thecalculatortool defined inSKILL.md. - Sanitization: Absent; the calculator implementation lacks any sanitization, validation, or sandboxing of the input before it is passed to the execution function.
Recommendations
- AI detected serious security threats
Audit Metadata