structured-extraction
Structured extraction — text in, a typed object you can trust out
The deliverable is a typed object that conforms to a schema you defined — not prose, not "roughly JSON." The whole skill rests on one distinction the rest of the file keeps returning to:
Native structured outputs make the JSON valid and typed. They never make the values correct.
Constrained decoding guarantees the model cannot emit a token that breaks your schema, so JSON.parse
errors, missing keys, wrong types, and stray markdown fences disappear at the source. It does nothing
to stop the model from putting a plausible-but-wrong email in a string field, snapping a fuzzy category to
the wrong enum, or coercing "$1,200" into 1200.0 when the currency mattered. Owning both halves — the
shape (decoding) and the values (validation) — is this skill. If you only do the first half you ship a
database full of well-typed lies.
Boundary test (bytes vs. schema). If the input is a PDF, scan, DOCX, or HTML and the deliverable is the
raw text/Markdown/cells of that document, that is upstream: document-processing
produces the text, this skill turns that text into typed fields. If you're holding text and want it shaped,
you're in the right place.