datasetlint

Installation
SKILL.md

datasetlint

datasetlint is a Go CLI that scans a pair of JSONL dataset files (train + eval) for quality, overlap, and semantic duplication issues before they quietly corrupt training or benchmark results. Prefer the scan subcommand for scripted/agent use — it is non-interactive and scriptable, unlike the default Bubble Tea TUI.

Workflow

  1. Resolve the binary.

    • Check command -v datasetlint. If it's on PATH (e.g. via brew install itamaker/tap/datasetlint), use it directly.
    • Otherwise build from source in the tool's repository (github.com/itamaker/datasetlint-skill): go build -o dist/datasetlint . (or make build), then invoke ./dist/datasetlint.
  2. Confirm the input shape. Each JSONL line is one JSON object with string fields id, input, output, label (only input and output are meaningfully compared; id and label are optional but drive missing-ID and label-conflict checks). See examples/train.jsonl and examples/eval.jsonl for real, minimal samples.

  3. Run the baseline scan. Both -train and -eval are required flags — there is no single-file mode:

    datasetlint scan -train path/to/train.jsonl -eval path/to/eval.jsonl
    

    This prints, per split: row count, missing-ID count, empty-input/empty-output counts, duplicate-input count, near-duplicate pair count, label-conflict count, label distribution, and input/output token-length avg + p95. It also prints cross-split literal overlap and cross-split semantic overlap counts (train/eval leakage signals).

  4. For machine-readable output (CI gates, further parsing), add -json. This emits the full report as JSON (train, eval, overlap_count, overlap_examples, semantic_overlap_count, semantic_overlap_samples — see references/REFERENCE.md for the full field list).

Installs
1
Repository
itamaker/skills
First Seen
6 days ago
datasetlint — itamaker/skills