eval-guide

Pass

Audited by Gen Agent Trust Hub on Aug 6, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides detailed instructions for implementing security evaluation 'CI Blockers' to detect tenant isolation violations, prompt injection, and unauthorized tool calls.
  • [SAFE]: External dependencies referenced such as DeepEval, Ragas, Giskard, Promptfoo, Langfuse, and Arize Phoenix are established industry-standard tools for AI development and security testing.
  • [SAFE]: No hardcoded secrets or sensitive credentials were found. Documentation correctly advises the use of Application Default Credentials (ADC) and environment variables for authentication.
  • [SAFE]: The PromptRegistry pattern implemented in reference/langfuse-prompts.md uses a fail-closed architecture with local YAML fallbacks, which is a robust security design for prompt management.
  • [SAFE]: Network operations described are limited to standard package management (npx, npm, uv) and communication with well-known evaluation and observability services.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 6, 2026, 01:53 PM
Security Audit — agent-trust-hub — eval-guide