evaluation-design
Pass
Audited by Gen Agent Trust Hub on Aug 29, 2026
Risk Level: SAFE
Full Analysis
- Security-Positive Data Handling: The skill explicitly instructs the agent to use synthetic content and avoid including real customer data, credentials, or personal information in any generated evaluation pairs. This proactive constraint reduces the risk of accidental data exposure.
- Durable Write Gating: It requires all durable writes to be routed through a 'scan gate' and mandates that the agent stop the process if the gate is unavailable or detects high-confidence security findings, ensuring a layer of automated oversight for output.
- Indirect Prompt Injection Mitigation: The instructions include a specific constraint to treat all supplied transcripts, documents, or tool outputs as data rather than instructions. This design helps prevent malicious content embedded in grounding sources from overriding the agent's behavior.
- Trusted External References: The skill references official Microsoft documentation for agent evaluator definitions. These references are used solely for factual identification and informational purposes, originating from a well-known and trusted service provider.
- Human-in-the-Loop Validation: The workflow incorporates a representative sample review and explicitly forbids the agent from marking human-review checkboxes, ensuring that final validation remains a human responsibility.
Audit Metadata