langgraph
Fail
Audited by Gen Agent Trust Hub on Jul 28, 2026
Risk Level: HIGHCOMMAND_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The 'Basic Agent Graph' example contains a
calculatortool defined as@tool def calculator(expression: str) -> str: return str(eval(expression)). The use ofeval()on input derived from the agent's conversation allows for Arbitrary Code Execution (ACE). An attacker can provide a payload such as__import__('os').system(...)to execute shell commands, access environment variables, or exfiltrate sensitive files. This pattern is extremely dangerous for production agents. - [COMMAND_EXECUTION]: The skill instructs the agent to design and implement graphs that may execute tools based on LLM routing without mentioning safety sandboxing or input validation, exacerbating the risk of the
eval()pattern provided. - [PROMPT_INJECTION]: The skill architecture is susceptible to indirect prompt injection because it ingests untrusted user messages into the
AgentStatewithout explicit boundary markers or instructions to ignore embedded commands within the tool-calling logic. - Ingestion points: The
messageslist inAgentState(SKILL.md). - Boundary markers: None provided in the implementation examples.
- Capability inventory: The skill facilitates tool execution, specifically the unsafe
calculatortool (SKILL.md). - Sanitization: No validation or sanitization is performed on the
expressionstring before it reaches theeval()call.
Recommendations
- AI detected serious security threats
Audit Metadata