facsimile-edition
Facsimile Edition Pipeline
Reference implementation lives in Creative Intelligence Agency/publishing/CLM Publishing/Cross & Plough/_facsimile_samples/experiments/prod/ (and …/ed52/ for a second-issue example). Scripts have hard-coded C&P paths — copy + repoint for a new title.
The one load-bearing fact
A generative image model redraws text from its language prior — it cannot OCR-faithfully reproduce dense body type on its own. Every page MUST be grounded with its exact transcript, and the model must be told to trust the transcript over the scan. Proven: ungrounded nano-banana/Gemini corrupt ~15 words/page (invisible at a glance); gpt-image-2 + per-page transcript = ~0 errors. See [[feedback_gpt_image2_vs_gemini_facsimile]], [[feedback_vlm_fullpage_autocorrects_text]].
Pipeline (6 stages)
1. Render + split source folios
The archive scans are often 2-up spreads on a dark scanner bed, sometimes skewed. Per page:
pdftoppm -r 300 -png source.pdf out→ render.- Split spreads at the center; crop to the bright paper region (rows/cols where
>120brightness covers >35%) to drop the dark bed. - Badly-rotated folio (e.g. a 45° scan): rotate to level, then remove the black scanner-bed corners with a PIL BoxBlur dark-density filter (
density>0.72= bed; text columns ~0.3 survive), crop tight to the ink. Don't crop to the diamond bbox — that keeps the black corners. - Gutter split (2026-07,
split_spreads.py): two strategies — (a) if a dark gutter shadow exists, split at the darkest central valley; (b) if the scan is flat with no shadow (paper fraction ~0.8 straight across), split at the midpoint of the paper region. A pure "darkest column" heuristic fails on flat scans (picks a text column or edge → one half keeps the whole spread). Map PDF spread page → (left folio, right folio) explicitly.
2. Transcribe (the grounding)
Spawn one agent per folio (parallel). Each reads the rendered folio + the issue's archive markdown (Sources/OCR_Transcriptions/NN_*.md) and writes the verbatim text of that folio only, in reading order, with light bracketed layout hints ([Centered heading:], [Two columns, drop-cap "I"], [Three EQUAL-width columns]). Markdown is the authority for spelling. → transcripts/page_NN.txt.
- If the scan is cut off (missing words at an edge), add a note:
[IMPORTANT — source scan is cut at the upper-right; render the COMPLETE text below, nothing truncated.]The grounding then fills the gap.