media-tag

Installation
SKILL.md

Media Tag

Context: $ARGUMENTS

Four open-source vision-language families, four different jobs:

  • CLIP / SigLIP — "is this image closer to prompt A or prompt B?" Zero-shot classification, similarity scoring, semantic search index.
  • BLIP-2 — pure image captioning. One image in, one sentence out.
  • LLaVA — full vision-language model. Captioning, VQA, dense description, video-frame narration.

All commercial-safe.

Quick start

Installs
5
GitHub Stars
17
First Seen
May 27, 2026
media-tag — damionrashford/media-os