pdf-vision
Installation
SKILL.md
PDF Vision Extraction Skill (Enhanced)
Overview
This skill handles image-based or scanned PDFs that contain no selectable text. It supports multiple vision APIs with automatic fallback:
Primary Models
- Xflow:
qwen3-vl-plus(your primary vision model) - ZhipuAI:
glm-4.6v-flash(free vision model with fallback support) - Fallback:
glm-5(text-only, but may work with some image prompts)
Unlike traditional PDF text extraction tools (pdftotext, pdfplumber) which only work on text-based PDFs, this skill can process:
- Scanned documents
- Image-only PDFs
- Photographed documents
- Handwritten notes (with limitations)
- Complex layouts with tables and formatting