tesseract-ocr
Installation
SKILL.md
Tesseract OCR
Workflow
- Confirm document classes, language mix, and quality targets.
- Build preprocessing pipeline (rescale, denoise, binarize, deskew).
- Select OCR engine mode and page segmentation mode per document type.
- Tune dictionaries, whitelists, and language model packs.
- Add post-processing and confidence-based review logic.
- Benchmark accuracy and latency on representative datasets.
- Operationalize monitoring and failure triage.
Preflight (Ask / Check First)
- Tesseract version and installed language data.
- Input quality (resolution, skew, compression artifacts).
- Document classes (receipts, IDs, forms, books).
- Accuracy target and acceptable manual-review rate.
- Runtime constraints (CPU budget, throughput, batch size).