openai-realtime-voice

Installation
SKILL.md

OpenAI Realtime voice

Use this skill when the deliverable is a live spoken interaction: a browser voice assistant, phone-like agent, low-latency speech-to-speech support flow, guided interview, tutoring conversation, live kiosk, or realtime voice UI that listens, reasons, speaks, interrupts, and may call tools.

Do not use this skill for batch speech-to-text, offline TTS, voiceover generation, audio editing, dubbing, music, or file-based audio tasks. Those belong to the separate OpenAI audio/request APIs or other speech/audio skills. The boundary is simple: if a persistent realtime session and turn-taking behavior are part of the product, use this skill; if the job is "turn this file/text into text/audio," do not.

All OpenAI API facts below were verified against first-party OpenAI documentation on 2026-07-10. Re-check model identifiers, prices, session limits, voices, API object fields, data retention, and policy requirements before production release because Realtime surfaces change quickly.

Source-grounded facts to carry into every design

Documented facts:

Installs
33
GitHub Stars
131
First Seen
Jul 11, 2026
openai-realtime-voice — calesthio/generative-media-skills