large-document-processing
Installation
SKILL.md
Large Document Processing & Intelligent Text Chunking
Overview
Two tightly related concerns combined here:
- Large document parsing — DOCX/PDF/EPUB ingestion with structure preservation
- Intelligent text chunking — splitting parsed text into semantically coherent pieces for AI training or RAG
Source Files
| File | Purpose |
|---|---|
src/utils/nwt_epub_parser.py |
EPUB parser for NWT Bible (English + Chuukese) |
scripts/extract_jwpub.py |
Extract JW publication .jwpub archives |
scripts/setup_large_document_processing.py |
One-time document pipeline setup |
output/processed_document/ |
Output directory for processed content |