openai-audio
OpenAI audio
Use this skill when a user wants OpenAI audio production or understanding in a bounded request: generated narration, spoken product demos, podcast or interview transcription, subtitle timing, translation to English, audio QA, or adding audio input/output to an existing chat workflow.
Do not use this as the primary guide for continuous low-latency speech-to-speech agents, live interpreting, SIP/WebRTC sessions, or voice activity detection. Those belong to the separate realtime voice skill. This skill may still explain the boundary: request-based audio APIs fit files, scripts, and bounded requests; realtime sessions fit open connections with live audio events. Documented source: OpenAI’s audio guide distinguishes request-based APIs, realtime sessions, and multimodal chat completions; verified 2026-07-10: https://developers.openai.com/api/docs/guides/audio
Source status and evidence labels
Treat the following as documented facts verified on 2026-07-10 from OpenAI public docs: