skills/smithery.ai/dataset-processing-multiprocessing

dataset-processing-multiprocessing

Installation
SKILL.md

Dataset Processing with Multiprocessing

This skill provides comprehensive guidance for processing large datasets (HuggingFace, Arrow, JSONL) in parallel using speedy_utils.multi_process. It covers the complete architectural pattern including worker design, data sharding, temporary file management, and robust error handling.

When to Use This Skill

Use this skill when you need to:

  • Process large HuggingFace datasets that don't fit efficiently in memory
  • Apply expensive transformations (tokenization, format conversion, packing)
  • Leverage multiple CPU cores for data preprocessing pipelines
  • Combine dataset loading, transformation, and saving in one pipeline
  • Integrate external tools (tokenizers, Megatron, format converters) with parallel processing
  • Manage intermediate temporary files safely across multiple workers
Installs
1
First Seen
Apr 14, 2026
dataset-processing-multiprocessing from smithery.ai