eval-guide
Pass
Audited by Gen Agent Trust Hub on Aug 6, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides detailed instructions for implementing security evaluation 'CI Blockers' to detect tenant isolation violations, prompt injection, and unauthorized tool calls.
- [SAFE]: External dependencies referenced such as DeepEval, Ragas, Giskard, Promptfoo, Langfuse, and Arize Phoenix are established industry-standard tools for AI development and security testing.
- [SAFE]: No hardcoded secrets or sensitive credentials were found. Documentation correctly advises the use of Application Default Credentials (ADC) and environment variables for authentication.
- [SAFE]: The PromptRegistry pattern implemented in reference/langfuse-prompts.md uses a fail-closed architecture with local YAML fallbacks, which is a robust security design for prompt management.
- [SAFE]: Network operations described are limited to standard package management (npx, npm, uv) and communication with well-known evaluation and observability services.
Audit Metadata