groq-core-workflow-b
Installation
SKILL.md
Groq Core Workflow B: Audio, Vision & Speech
Overview
Beyond chat completions, Groq provides ultra-fast audio transcription (Whisper at 216x real-time), multimodal vision (Llama 4 Scout/Maverick), and text-to-speech. These endpoints use the same groq-sdk client.
Prerequisites
groq-sdkinstalled,GROQ_API_KEYset- For audio: audio files in supported formats
- For vision: image URLs or base64 images
Audio Models
| Model ID | Languages | Speed | Best For |
|---|---|---|---|
whisper-large-v3 |
100+ | 164x real-time | Best accuracy, multilingual |
whisper-large-v3-turbo |
100+ | 216x real-time | Best speed/accuracy balance |
Supported audio formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm