llm-as-judge
Pass
Audited by Gen Agent Trust Hub on Sep 9, 2026
Risk Level: SAFENO_CODE
Full Analysis
- [NO_CODE]: The skill consists entirely of markdown documentation providing guidance and best practices for LLM-as-judge evaluation frameworks. There are no scripts, binaries, or configuration files that execute code.
- [SAFE]: No malicious patterns, obfuscation, or unauthorized data access mechanisms were identified. The skill focuses on methodology and behavioral expectations for an AI agent assisting with evaluation tasks.
- [INDIRECT_PROMPT_INJECTION]: The skill documentation acknowledges that LLM judge prompts are vulnerable to injection from candidate outputs and points towards defensive skills. The current skill itself acts as a behavioral guide and does not process untrusted data directly via scripts or tools.
Audit Metadata