vision-sft

Installation
SKILL.md

Vision-Language SFT

This skill assumes finetuning-method-selection already routed here: the data shape is image+text demonstrations, not preference pairs or a verifiable reward signal, and the base is a vision-language model rather than a text-only one. lora-qlora-recipes covers the text-only LoRA/QLoRA recipe this skill specializes for the vision tower and projector; read that skill first if the LoRA fundamentals (rank, alpha, target modules) aren't already familiar.

Installs
429
Repository
wshobson/agents
GitHub Stars
38.4K
First Seen
Jul 14, 2026
vision-sft — wshobson/agents