langgraph

Fail

Audited by Gen Agent Trust Hub on Apr 11, 2026

Risk Level: HIGHCOMMAND_EXECUTIONREMOTE_CODE_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The provided code example for a calculator tool uses the Python eval() function on an input string directly supplied by the agent. This allows for arbitrary command execution within the application environment if the LLM generates or is manipulated into providing malicious code instead of a simple math expression.
  • [REMOTE_CODE_EXECUTION]: The combination of processing untrusted user input and a tool that executes arbitrary strings via eval() creates a high risk for remote code execution. An attacker could use indirect prompt injection techniques to force the agent to call the calculator tool with malicious Python logic.
  • Ingestion points: Untrusted data enters the agent context through the messages list in AgentState and the query field in RouterState (SKILL.md).
  • Boundary markers: Absent; the examples do not use delimiters or system instructions to prevent the agent from obeying instructions embedded in user-supplied messages.
  • Capability inventory: The skill includes the capability to execute arbitrary Python code via the eval() call in the calculator tool defined in SKILL.md.
  • Sanitization: Absent; the calculator implementation lacks any sanitization, validation, or sandboxing of the input before it is passed to the execution function.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Apr 11, 2026, 06:18 PM
Security Audit — agent-trust-hub — langgraph