markitright
MarkItRight
Convert any PDF to clean markdown with figures intact. Four phases: preflight → extract → QA → finalize.
Implementation lives at: Creative Intelligence Agency/crux/document-tools/markitright/ — that's where the markitright CLI, pipeline.py, .venv, and scripts/ are. This SKILL.md in .claude/skills/ is the registered triggering entrypoint; when invoked, cd to the implementation directory and activate the venv.
Core Principle: Screenshot Regions, Don't Extract Bytes
The single most important lesson from this pipeline: for any figure region, render a pixmap of that bbox — don't pull block["image"] raw bytes.
rect = fitz.Rect(block["bbox"]) + (-5, -5, 5, 5) # small margin for labels
mat = fitz.Matrix(2, 2) # 2x zoom
pix = page.get_pixmap(matrix=mat, clip=rect)
pix.save(img_path)
Why: raw embedded image bytes miss everything drawn on top of the image (axis labels, legends, overlaid captions) and miss vector charts entirely. A clip-based pixmap captures exactly what a human sees. This works for raster images AND vector drawings — no detection needed.