eval-designer
Pass
Audited by Gen Agent Trust Hub on Jul 17, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides procedural instructions for building evaluation test suites, designing rubrics, and implementing automated evaluation pipelines. All identified behaviors are consistent with its stated purpose of LLM quality and safety engineering.
- [SAFE]: Static analysis flags for prompt injection (e.g., 'Ignore previous instructions') are confirmed false positives. These strings appear exclusively within documentation and code examples as 'adversarial test case' samples intended to help users test the robustness of their own models, not to manipulate the agent using this skill.
- [SAFE]: Code examples provided in the skill demonstrate standard industry practices, such as using the OpenAI API for LLM-as-judge scoring and utilizing subprocess calls to run legitimate security scanners like 'Bandit' for static code analysis.
- [SAFE]: No obfuscation, data exfiltration patterns, or unauthorized persistence mechanisms were detected.
Audit Metadata