ai-eval-review
AI Eval Review
Audit the eval layer of an AI product or feature against a structured set of design-completeness questions. Each of the seven blocks is a gap detector — the value of this skill is whether the elicitation surfaces eval decisions the builder hasn't made yet, not whether the artifact looks complete.
The job is eval-design-completeness: have we designed how we'll know if this works? Covers offline criteria, ground-truth quality, online signal, cohort breakdowns and disparate impact, adversarial / robustness coverage, and drift detection. Regulatory rigor (EU AI Act, FDA SaMD, FTC) is a cross-cutting lens applied across blocks, not a separate block.
This skill is the eval-side companion to ai-ux-review. Same shape, same
elicitation pattern, different subject — ai-ux-review audits the
human-AI design surface (was the experience intentionally designed?);
this skill audits the measurement layer behind it (do we have signal for
whether the design works?).