Introducing next-generation audio models in the API
What changed
For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Gemini 3.1 Flash TTS: the next generation of expressive AI speech
- Advancing voice intelligence with new models in the API
- Introducing the Realtime API
Sources
- Introducing next-generation audio models in the API (openai-blog)primary