tesseract-ocr

Installation
SKILL.md

Tesseract OCR

Workflow

  1. Confirm document classes, language mix, and quality targets.
  2. Build preprocessing pipeline (rescale, denoise, binarize, deskew).
  3. Select OCR engine mode and page segmentation mode per document type.
  4. Tune dictionaries, whitelists, and language model packs.
  5. Add post-processing and confidence-based review logic.
  6. Benchmark accuracy and latency on representative datasets.
  7. Operationalize monitoring and failure triage.

Preflight (Ask / Check First)

  • Tesseract version and installed language data.
  • Input quality (resolution, skew, compression artifacts).
  • Document classes (receipts, IDs, forms, books).
  • Accuracy target and acceptable manual-review rate.
  • Runtime constraints (CPU budget, throughput, batch size).
Installs
2
GitHub Stars
2
First Seen
Mar 21, 2026
tesseract-ocr — kittne/codex-skills-by-codex