llm-judge

Pass

Audited by Gen Agent Trust Hub on Sep 15, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill ingests untrusted data from external repositories and specification files, which are processed by subagents. Maliciously crafted content in these files could attempt to influence the evaluation logic or subagent behavior. Ingestion points: The skill reads a user-provided spec file in SKILL.md and repository contents in references/repo-agent.md. Boundary markers: The skill uses markdown headers like 'Spec Document:' but lacks explicit instructions for subagents to ignore potentially malicious commands embedded within that data. Capability inventory: The subagents have the ability to execute git commands and trigger local test runners (pytest, npm test, go test) as defined in references/repo-agent.md. Sanitization: No explicit sanitization or filtering of the ingested content is performed before it is passed to the subagents.
  • [COMMAND_EXECUTION]: The skill executes shell commands (git) and test runners (pytest, npm test, go test) against the source code in the repositories being evaluated. While this is the intended purpose of the skill, it involves running code derived from untrusted repositories.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 15, 2026, 03:09 PM
Security Audit — agent-trust-hub — llm-judge