extracting-with-ocr

Installation
SKILL.md

Extracting with OCR

Use this when a document is image-based: scanned PDFs, photographed pages, screenshots, JPEG/PNG/TIFF with text. Xberg auto-OCRs raster images and auto-detects PDFs that lack a text layer. Force it on when extraction returned empty/garbled text from a PDF that "looks" textual.

When to force OCR

  • Extraction returned an empty content field, but the file opens visually.
  • The PDF text layer is junk (copy-paste from a viewer produces gibberish).
  • You want consistent output across mixed scanned + digital PDFs.
Installs
10
Repository
xberg-io/xberg
GitHub Stars
9.3K
First Seen
Aug 2, 2026
extracting-with-ocr — xberg-io/xberg