voice-segment-selector

Installation
SKILL.md

voice-segment-selector

Prepare ranked, gender-bucketed, transcript-backed voice clips from audio or video for /tts-train.

This skill is not speaker diarization. Gender labels are vocal characteristic guesses, not identity. Human review is recommended before training export.

Agent playbook (read this first)

Project agents should follow this exact sequence:

  1. Acquire source media (if needed)
    • Local file: use path directly
    • YouTube URL: run /ingest-youtube first, or download audio yourself
    • Long audiobook: get ffprobe_chapters.json from /extract-audiobook first
  2. Prepare job with an explicit --job-dir (do not rely on memory of /tmp names)
  3. Review top candidates with the human (review API or decide)
  4. Export accepted clips to metadata.jsonl
  5. Optional: build one --target-sec 30 bundle WAV
  6. Hand off exported dataset path to /tts-train
Installs
1
GitHub Stars
7
First Seen
Aug 26, 2026
voice-segment-selector — grahama1970/agent-skills