document-processing
Document processing
File in, content out — or data in, file out. You open a byte stream (PDF, DOCX, scan) and either pull the content out, or you build a new document from a template and a data dict. That is the whole job: the deliverable is bytes of a document or the literal content of one.
The boundary test, apply it first:
- Deliverable is raw text / Markdown / table cells / a generated file → you are in the right place.
- Deliverable is a typed object matching a schema (
{parties: [...], total: 1234.50}) → that isstructured-extraction. This skill stops at "clean Markdown out of the file"; the schema-constrained extraction runs on that Markdown.
Everything else routes too: signing with an audit trail → e-signature, spreadsheet grids/formulas/XLSX-as-data → spreadsheet-ops, indexing for cross-document Q&A → rag (this skill produces the text rag ingests, it does not index it), downloading the files off a site → data-scraper.
Step 0 — does the PDF have a text layer?
The most expensive mistake in this skill is OCR'ing a PDF that already has a text layer. A digital PDF (exported from Word, a browser, a report tool) carries selectable text — extracting it is free, instant, and lossless. OCR is slow, costs money or GPU, and introduces errors. Never OCR a PDF you can extract.
Check before you pick an engine:
import pdfplumber