prompt-benchmark
Warn
Audited by Gen Agent Trust Hub on Oct 2, 2026
Risk Level: MEDIUMDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [DYNAMIC_EXECUTION]: The skill utilizes the JavaScript
Functionconstructor to evaluate arithmetic results produced by the AI model in the Game of 24 benchmark. - Evidence: The function
evaluateExpressioninSKILL.mdcontainsconst result = Function("\"use strict\"; return (${sanitized})")();. - While the code includes a sanitization step using
expr.replace(/[^0-9+\-*/().]/g, ''), the use of dynamic execution for content sourced from model outputs is a risky implementation pattern. - [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest and process untrusted data in the form of AI model responses across several benchmarks (MATH, GSM8K, Game of 24).
- Ingestion points: The
runBenchmarkfunction inSKILL.mdcollects responses from agenerateFnand passes them to evaluators likeevaluateGSM8KandevaluateMATH. - Boundary markers: The prompt templates (e.g.,
cotPrompt,metaPromptGame24) do not use clear delimiters or instructions to prevent the agent from obeying instructions that might be embedded within the benchmark problems themselves. - Capability inventory: The skill has the ability to execute code (via
Function) and manipulate string data based on these untrusted inputs. - Sanitization: While basic arithmetic characters are whitelisted for the math evaluator, there is no general sanitization or escaping of the model-generated responses before they are processed by the evaluation logic.
Audit Metadata