evaluation-design
Installation
SKILL.md
AI Evaluation Dataset Design
Goal
Produce an evaluation dataset and its supporting documentation that measure whether an AI system does its job, refuses what it should refuse, and behaves acceptably under pressure. The dataset is a durable customer artifact, so its scope, balance, and rationale are recorded rather than implied.
Flow
- Run the scoping interview from the interview reference. Ask one question at a time and wait for the answer; do not batch the interview into a single prompt.
- Present a structured summary of what you heard and obtain explicit confirmation before generating anything.
- Derive the difficulty distribution from the confirmed scope, adjusting the defaults when the system's risk profile warrants it.
- Generate the dataset against the contract template in both machine-readable forms.
- Walk a representative sample through the user, gather consolidated feedback, and revise before finalizing the full set.
- Produce one sectioned evaluation guide containing curation notes, metric selection with rationale, and tooling recommendations.
- Route every durable write through the workstream's scan gate before it lands in a customer location.