langgraph

Fail

Audited by Gen Agent Trust Hub on Jul 28, 2026

Risk Level: HIGHCOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The 'Basic Agent Graph' example contains a calculator tool defined as @tool def calculator(expression: str) -> str: return str(eval(expression)). The use of eval() on input derived from the agent's conversation allows for Arbitrary Code Execution (ACE). An attacker can provide a payload such as __import__('os').system(...) to execute shell commands, access environment variables, or exfiltrate sensitive files. This pattern is extremely dangerous for production agents.
  • [COMMAND_EXECUTION]: The skill instructs the agent to design and implement graphs that may execute tools based on LLM routing without mentioning safety sandboxing or input validation, exacerbating the risk of the eval() pattern provided.
  • [PROMPT_INJECTION]: The skill architecture is susceptible to indirect prompt injection because it ingests untrusted user messages into the AgentState without explicit boundary markers or instructions to ignore embedded commands within the tool-calling logic.
  • Ingestion points: The messages list in AgentState (SKILL.md).
  • Boundary markers: None provided in the implementation examples.
  • Capability inventory: The skill facilitates tool execution, specifically the unsafe calculator tool (SKILL.md).
  • Sanitization: No validation or sanitization is performed on the expression string before it reaches the eval() call.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Jul 28, 2026, 05:45 AM
Security Audit — agent-trust-hub — langgraph