grpo-rlvr-training
Installation
SKILL.md
GRPO & RLVR Training
This skill assumes finetuning-method-selection
already routed here because the target behavior
has a verifiable pass/fail signal — not
demonstrations (lora-qlora-recipes) or
preference pairs (preference-optimization).
What follows is when RL is the right tool, the
reference recipe, the mandatory reward-inspection
gate, and how to pick a GRPO variant when the
base recipe misbehaves.