document-processing

Installation
SKILL.md

Document processing

File in, content out — or data in, file out. You open a byte stream (PDF, DOCX, scan) and either pull the content out, or you build a new document from a template and a data dict. That is the whole job: the deliverable is bytes of a document or the literal content of one.

The boundary test, apply it first:

  • Deliverable is raw text / Markdown / table cells / a generated file → you are in the right place.
  • Deliverable is a typed object matching a schema ({parties: [...], total: 1234.50}) → that is structured-extraction. This skill stops at "clean Markdown out of the file"; the schema-constrained extraction runs on that Markdown.

Everything else routes too: signing with an audit trail → e-signature, spreadsheet grids/formulas/XLSX-as-data → spreadsheet-ops, indexing for cross-document Q&A → rag (this skill produces the text rag ingests, it does not index it), downloading the files off a site → data-scraper.

Step 0 — does the PDF have a text layer?

The most expensive mistake in this skill is OCR'ing a PDF that already has a text layer. A digital PDF (exported from Word, a browser, a report tool) carries selectable text — extracting it is free, instant, and lossless. OCR is slow, costs money or GPU, and introduces errors. Never OCR a PDF you can extract.

Check before you pick an engine:

import pdfplumber
Installs
2
GitHub Stars
116
First Seen
Aug 6, 2026
document-processing — ericrisco/rsc-harness