data-cleaning

Installation
SKILL.md

Data cleaning — make dirty data trustworthy, and make the cleaning auditable

A clean table is typed + deduped + normalized + validated + reproducible. The deliverable here is never "I opened a notebook and fixed some rows by hand." It is a re-runnable function clean(raw) -> df plus a schema gate that fails loud when next month's file violates the contract. Reproducible means the same input always yields the same output: versions pinned, sorts deterministic, nothing random without a seed. If you can't re-run it tomorrow and get the identical result, you haven't cleaned the data — you've edited a snapshot.

Cleaning starts once you hold tabular rows and ends at a validated table/DataFrame/Parquet. Before that boundary the job is acquisition (data-scraper, structured-extraction); after it, consumption (spreadsheet-ops, analytics, business-intelligence, forecasting). Multi-GB analytical SQL is an engine choice, not a cleaning one — duckdb.

Installs
4
GitHub Stars
116
First Seen
Aug 6, 2026
data-cleaning — ericrisco/rsc-harness