dataset-processing-multiprocessing
Installation
SKILL.md
Dataset Processing with Multiprocessing
This skill provides comprehensive guidance for processing large datasets (HuggingFace, Arrow, JSONL) in parallel using speedy_utils.multi_process. It covers the complete architectural pattern including worker design, data sharding, temporary file management, and robust error handling.
When to Use This Skill
Use this skill when you need to:
- Process large HuggingFace datasets that don't fit efficiently in memory
- Apply expensive transformations (tokenization, format conversion, packing)
- Leverage multiple CPU cores for data preprocessing pipelines
- Combine dataset loading, transformation, and saving in one pipeline
- Integrate external tools (tokenizers, Megatron, format converters) with parallel processing
- Manage intermediate temporary files safely across multiple workers