ragcheck
Installation
SKILL.md
ragcheck
ragcheck is a Go CLI that scores retrieval runs against qrels (query relevance judgments) and judges RAG answer quality offline — no notebooks, no external eval service, no network calls. It ships two direct subcommands (score, judge) plus an interactive Bubble Tea TUI when invoked with no arguments.
Workflow
-
Confirm
ragcheckis available. If not installed (brew install itamaker/tap/ragcheckor a GitHub release binary), and you have the tool's source checked out (github.com/itamaker/ragcheck-skill), build it locally withgo build -o ragcheck .or run subcommands viago run . <subcommand> .... -
To score a retrieval run against ground-truth relevance judgments:
- Ensure you have a qrels file (array of
{query_id, relevant[]}or{query_id, grades{doc_id: grade}}) and a run file (array of{query_id, results[]}, ranked document IDs) that sharequery_idvalues. - Run
ragcheck score -qrels <qrels.json> -run <run.json> -k <N>(-kdefaults to 5). - Add
-jsonfor machine-readable output when you need to parse or diff the numbers programmatically. - Output includes: query count, missing-run count (qrels with no matching run), and Precision@k, Recall@k, HitRate@k, MRR@k, MAP@k, nDCG@k.
- Ensure you have a qrels file (array of
-
To judge RAG answer quality offline (no qrels needed):
- Build a judge input file: an array of
{query_id, question, answer, contexts[], reference?}records (referenceis optional). - Run
ragcheck judge -input <judge.json>, optionally with-json. - Output includes per-query and averaged answer relevance, context relevance, groundedness, and reference coverage, plus any unsupported answer tokens (answer content not backed by the retrieved contexts) and missing-context hints (reference content absent from contexts).
- Build a judge input file: an array of