Introducing next-generation audio models in the API

Practical AI: Tools, Models & Frameworksapi

What changed

For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Gemini 3.1 Flash TTS: the next generation of expressive AI speech
  • Advancing voice intelligence with new models in the API
  • Introducing the Realtime API

Sources