batch-pipeline-designer
Installation
SKILL.md
Batch Pipeline Designer
When to Use
Use this skill when you are designing or evaluating a batch processing pipeline — a system that reads a bounded, fixed-size dataset, runs a job over it, and produces output. Batch jobs are scheduled periodically (hourly, daily), tolerate high latency, and are measured by throughput rather than response time.
This skill applies when:
- Processing large datasets offline (logs, clickstream, database snapshots, file dumps)
- Building ETL workflows to move data between systems
- Generating derived datasets: search indexes, ML training features, recommendation model inputs, aggregated reports
- Implementing multi-step data transformations across distributed storage
- Joining two or more large datasets that cannot fit in memory on a single machine
- Processing graph-structured data (social graphs, link graphs, dependency graphs) in bulk
This skill does NOT apply to:
- Stream processing where input is unbounded (see
stream-processing-designer) - Online query serving where latency < 1 second is required
- Transactional workloads (see
oltp-olap-workload-classifierto confirm the workload type)