ln-45-benchmark-comparator
Pass
Audited by Gen Agent Trust Hub on Oct 5, 2026
Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill's execution contract includes the directive "do not invent approval gates from caution." This instruction explicitly tells the agent to disregard its own safety-oriented heuristics or procedural filters that might otherwise restrict its actions, which is a common pattern for attempting to bypass platform guardrails.
- [INDIRECT_PROMPT_INJECTION]: The skill is designed to process and execute "candidates" (external implementations or tools) and read "external semantics" from documentation, which introduces a surface where malicious instructions in that content could influence the agent.
- Ingestion points: The skill reads repository instructions, candidate source code, task-specific inputs, and external documentation during the benchmark definition and grading phases.
- Boundary markers: The skill requires the use of clean Git worktrees and temporary directories to isolate candidate execution environments and prevent cross-run contamination.
- Capability inventory: The skill utilizes a shell runner and execution wrapper to run arbitrary candidate code and capture stdout, stderr, and exit statuses for analysis.
- Sanitization: It mandates the use of an independent oracle and blind correctness grading to ensure outcomes are verified against expected results rather than relying on potentially malicious candidate self-reporting.
Audit Metadata