ai-eval-playbook
AI Evaluation Playbook — live-fetch guide
This skill is a stable index over a living document, not the content itself. The canonical playbook lives at https://eval.playbook.org.ai/ and its maintainers publish quarterly releases. The cached summaries below are orientation only. Fetch the live page before composing any substantive answer, and say which parts of the reply come from the live page versus cached orientation.
Adapted 2026-08-05 from the maintainers' own skill file (github.com/IDinsight/ai-eval-playbook, skills/), with tool references fixed for Ane's harness, the corrected /user-experience path (the upstream typo path now redirects), and routing scoped to Ane's skill lanes.
Step 0 — local first
- Read
mel_wiki/wiki/frameworks/ai-evaluation-playbook.md(work folder). It carries the verified citation block, the 4-level and MVE tables, the IPPF EN application, ECA calibration, and refresh state. For a quick orientation answer, that page plus this file is often enough — say so and name the rung. - If the AI tool touches SRHR, apply
mel_wiki/wiki/frameworks/ai-srhr-mel-framework.md(WHO / UNESCO / EU AI Act compliance frame) alongside the playbook's evaluation design.
How to fetch fresh content (priority order)
- WebFetch the canonical URL. If context-mode redirects the call, use
ctx_fetch_and_indexand query the indexed page. ?ask=endpoint for narrow lookups: append.md?ask=<URL-encoded-question>to any page URL for a direct answer with citations. Example:https://eval.playbook.org.ai/model-behaviour/level-1-module-evaluation/overview.md?ask=what+is+a+golden+dataset.- claude-in-chrome as fallback for pages that block automated fetch.
- Avoid old GitBook-ID URLs (
/spaces/<id>/pages/<id>); use the semantic paths below.
URL notes (as of 2026-08-05): /user-expereince (upstream's documented typo path) now redirects to /user-experience — use the corrected spelling. Section roots (e.g. /model-behaviour) duplicate at …/overview.md. Two Level-3 sub-pages have slug/title mismatches (descriptive-analysis renders "Identify outcome metrics"; user-privacy-and-security renders the process-evaluation page) — quote the rendered H1, not the slug.