twelvelabs-video-understanding

Installation
SKILL.md

TwelveLabs video understanding

TwelveLabs builds video foundation models that read footage the way a human editor does — across visuals, on-screen text, motion, sound, speech, and music — and expose that understanding through a REST API, SDKs (Python, Node), an MCP server, and NLE plugins. This skill is for driving that platform to log, search, describe, segment, and tag video for production work. It is not a video generator; TwelveLabs does not synthesize or edit pixels.

All volatile facts below carry a verification date. Everything moves fast here — re-verify model names, limits, and pricing against docs.twelvelabs.io before quoting them to a user as current.

When this skill applies

Reach for TwelveLabs when the job is any of:

  • Archive / footage search — "find every shot where the CEO is on stage," "clips with a red car at night," "where does someone say 'quarterly earnings'." Natural-language, image, or combined queries against indexed video.
  • Logging & tagging — auto-generate loggable metadata (who/what/where/action) for raw footage or dailies.
  • Segmentation — chapters, scene breaks, highlights, speaker changes, sports plays, ad-break points.
  • Video-to-text — summaries, descriptions, captions, Q&A over a clip, structured JSON extraction (e.g. shot lists, compliance flags).
  • Embeddings — multimodal vectors for a custom recommender, dedup, similarity, or a RAG-over-video store.
  • Compliance / brand-safety review — locate logos, on-screen text, spoken phrases, or sensitive content across a library.

Do not use it to create video, apply visual effects, transcode, or as a general speech-to-text tool (it does speech understanding for search/analysis, but a dedicated ASR is cheaper if plain transcripts are all you need).

Installs
31
GitHub Stars
131
First Seen
Jul 11, 2026
twelvelabs-video-understanding — calesthio/generative-media-skills