extract-html
Installation
SKILL.md
Extract-HTML Skill
A robust skill for converting HTML documents into strictly valid JSON based on a user-provided JSON Schema.
Capabilities
- Schema Compliance: Guarantees output conforms to the provided JSON Schema (using Schematron-3B + validation loop).
- Deterministic Tables: Extracts HTML tables using
pandas.read_htmland injects them as context, preventing hallucination of data. - Media Text Extraction: Identifies images, filters by pixel size, and optionally uses a Vision API (OpenAI-compatible) to extract text/OCR.
- Self-Correction: Validates model output and retries with error feedback if schema validation fails.