data-cleaning
Data cleaning — make dirty data trustworthy, and make the cleaning auditable
A clean table is typed + deduped + normalized + validated + reproducible. The deliverable here is
never "I opened a notebook and fixed some rows by hand." It is a re-runnable function clean(raw) -> df
plus a schema gate that fails loud when next month's file violates the contract. Reproducible means
the same input always yields the same output: versions pinned, sorts deterministic, nothing random without
a seed. If you can't re-run it tomorrow and get the identical result, you haven't cleaned the data — you've
edited a snapshot.
Cleaning starts once you hold tabular rows and ends at a validated table/DataFrame/Parquet. Before
that boundary the job is acquisition (data-scraper,
structured-extraction); after it, consumption
(spreadsheet-ops, analytics,
business-intelligence,
forecasting). Multi-GB analytical SQL is an engine choice, not a cleaning one
— duckdb.