statement-extract-and-prove
Statement Extract and Prove
A model reading a statement PDF by eye drops rows on long tables, flips signs, reads 1.234,56 as 1.234, and invents numbers where a scan is faint. Roughly one statement in seven fails to tie on the first pass. The fix is deterministic extraction from the PDF text layer plus arithmetic that the statement itself supplies: the bank already printed the opening balance, the closing balance, often a running balance on every row, and a summary box. If the extracted rows reproduce all of those, the extraction is proven. If not, the arithmetic points at the exact row.
Scope: this skill gets statement and table rows out of PDFs and proves them. For categorizing and reconciling against the books, hand the CSV to bookkeeping-close (its scripts/reconcile.py reads this CSV directly). For pulling header fields out of receipts and invoices (vendor, total, due date), use financial-parser; that skill does field extraction, while this one handles row tables that must tie out.
Setup
pip install pdfplumber openpyxl # required; openpyxl only for --xlsx
# optional, scanned PDFs only: ocrmypdf + tesseract (see references/ocr.md)
Workflow
1. Check for a text layer
Run the extractor. It checks the text layer first and exits with code 3 and NO TEXT LAYER when the PDF is an image. In that case, OCR it as described in references/ocr.md and continue with the OCR'd file. Do not start reading the page images yourself: OCR plus the proof below catches misreads, and eyeballing catches none.