datasetlint
datasetlint
datasetlint is a Go CLI that scans a pair of JSONL dataset files (train + eval) for quality, overlap, and semantic duplication issues before they quietly corrupt training or benchmark results. Prefer the scan subcommand for scripted/agent use — it is non-interactive and scriptable, unlike the default Bubble Tea TUI.
Workflow
-
Resolve the binary.
- Check
command -v datasetlint. If it's on PATH (e.g. viabrew install itamaker/tap/datasetlint), use it directly. - Otherwise build from source in the tool's repository (github.com/itamaker/datasetlint-skill):
go build -o dist/datasetlint .(ormake build), then invoke./dist/datasetlint.
- Check
-
Confirm the input shape. Each JSONL line is one JSON object with string fields
id,input,output,label(onlyinputandoutputare meaningfully compared;idandlabelare optional but drive missing-ID and label-conflict checks). Seeexamples/train.jsonlandexamples/eval.jsonlfor real, minimal samples. -
Run the baseline scan. Both
-trainand-evalare required flags — there is no single-file mode:datasetlint scan -train path/to/train.jsonl -eval path/to/eval.jsonlThis prints, per split: row count, missing-ID count, empty-input/empty-output counts, duplicate-input count, near-duplicate pair count, label-conflict count, label distribution, and input/output token-length avg + p95. It also prints cross-split literal overlap and cross-split semantic overlap counts (train/eval leakage signals).
-
For machine-readable output (CI gates, further parsing), add
-json. This emits the full report as JSON (train,eval,overlap_count,overlap_examples,semantic_overlap_count,semantic_overlap_samples— seereferences/REFERENCE.mdfor the full field list).