agent-evals-and-observability
Pass
Audited by Gen Agent Trust Hub on Sep 2, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [SAFE]: The skill provides framework-neutral guidance for agent evaluation, datasets, and observability. It does not include any scripts, executable tools, or automated network operations. All instructions emphasize data minimization and privacy.
- [INDIRECT_PROMPT_INJECTION]: The methodology identifies the ingestion of untrusted data (production traffic and adversarial cases) as an evaluation surface.
- Ingestion points: Production-derived data and adversarial cases are used in scenarios described in
references/datasets.md. - Boundary markers: The methodology uses immutable dataset manifests and frozen versioning to control data integrity.
- Capability inventory: This skill provides process guidance only and does not define tools or executable logic.
- Sanitization: The skill explicitly mandates redaction, field allowlists, and privacy-aware retention policies in
references/production-observability.md. - [EXTERNAL_DOWNLOADS]: Informational references to official documentation and standards from trusted organizations (NIST, OpenTelemetry, Pydantic, LangChain) are included. These are for reference only and do not facilitate runtime code downloads or execution.
Audit Metadata