regulation-extractor
Installation
SKILL.md
Regulation Extractor v3.0
从建筑工程规范PDF中提取条文,支持文字层提取和纯图片PDF的OCR识别,经过多轮清洗后输出高质量结构化数据。
工作流(完整Pipeline)
Step 1: 提取条文(文字层)
python scripts/extract_regulation.py <pdf_path> -o <output.json>
支持 --chapters <chapters.json> 自定义章节映射。
支持编号格式:3.1.2、3. 1. 2、4级编号 6.1.2.4,自动去除空格。