voice-segment-selector
Installation
SKILL.md
voice-segment-selector
Prepare ranked, gender-bucketed, transcript-backed voice clips from audio or video for /tts-train.
This skill is not speaker diarization. Gender labels are vocal characteristic guesses, not identity. Human review is recommended before training export.
Agent playbook (read this first)
Project agents should follow this exact sequence:
- Acquire source media (if needed)
- Local file: use path directly
- YouTube URL: run
/ingest-youtubefirst, or download audio yourself - Long audiobook: get
ffprobe_chapters.jsonfrom/extract-audiobookfirst
- Prepare job with an explicit
--job-dir(do not rely on memory of/tmpnames) - Review top candidates with the human (
reviewAPI ordecide) - Export accepted clips to
metadata.jsonl - Optional: build one
--target-sec 30bundle WAV - Hand off exported dataset path to
/tts-train