edit-llm-inference-style
Installation
SKILL.md
Edit LLM Inference Style (speedy_utils)
Use this skill to standardize or modify how speedy_utils builds prompts and consumes generation outputs. It focuses on chat templating, reasoning-style prefixes, and safe stopping rules for structured answer extraction.
When to Use This Skill
Use this skill when you need to:
- Insert or enforce a reasoning prefix (for example,
<think>\n). - Switch a flow to
LLM.generate()instead of chat completion helpers. - Apply a tokenizer chat template before generating.
- Stop generations on boxed answers (
\boxed{}) or<|im_end|>tokens. - Normalize outputs before evaluation (for example, GSM8K or math tasks).
Prerequisites
- A model-backed tokenizer available via
transformers.AutoTokenizer. - A speedy_utils
LLMinstance configured to point at the correct backend.