machine-learning
Installation
SKILL.md
Machine learning — classic/tabular models, done without lying to yourself
Tabular ML is easy to run and easy to fool yourself with. The deliverable is never "the notebook
printed 0.99" — it is an honest estimate of how the model behaves on data it has never seen: a
Pipeline that fits every transform on train only, a metric that survives class imbalance, and a
DummyClassifier baseline it beats. A number you can't reproduce on a sacred test set you touched exactly
once isn't a result — it's a leak you haven't found yet.
Is this the right skill? (decide first)
| Your situation | Reach for |
|---|---|
| Rows of features, predict a column, with trees / linear models / sklearn | machine-learning (this skill) |
| Images, audio, long text, sequences, or you need a neural net / PyTorch | deep-learning |
| The table is still dirty (nulls, dupes, mixed types, bad dates) | data-cleaning first — it hands you a validated table |
| Text/token classification, NER, tokenization, LLM-adjacent NLP metrics | nlp (a TF-IDF + linear/GBDT baseline still lives happily in this skill's pipeline) |
| KPIs, dashboards, "explain the business" | analytics / business-intelligence |
| Forecast a dated series forward (revenue next quarter) | forecasting |
| Building a training corpus of JSONL messages / preference pairs for an LLM | training-data |