Key takeaways

  • Your browser does not support the audio element.
  • They also make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative — helping you tackle…
  • In addition to this performance, it remains highly cost-effective — providing developers and enterprises with a capable and efficient model…

What happened

Your browser does not support the audio element. This content is generated by Google AI. Generative AI is experimental Today, we’re introducing two new models that bring advancements in near real-time reasoning to more effectively enable voice agents and make conversing with AI feel more intuitive and intelligent. For developers and enterprises, these models deliver the building blocks for reliable, production-ready voice agents.

It delivers increased intelligence for complex workflows while maintaining an uninterrupted conversational flow — using early verbal cues like “Let me check that…” to acknowledge prompts naturally, and live progress narration to walk users through multi-step background tasks as they progress. Across Google Workspace and Search, our Live models deliver more intuitive, collaborative experiences — especially when tackling your most complex tasks.

Why it matters

They also make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative — helping you tackle complex tasks using just your voice. 1% on Sierra’s τ-Voice-banking benchmark. 7% on Big Bench Audio, while maintaining a highly competitive price point compared to other frontier models. 8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena.

In addition to this performance, it remains highly cost-effective — providing developers and enterprises with a capable and efficient model built for scale. On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, our models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality. 8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses.

It automatically detects and transitions between 97 supported languages mid-conversation. It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background. 8 Live Extended Thinking reasons and speaks simultaneously.

What to watch

By using the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience. 8 Live Extended Thinking, highlighting its impressive latency, fluidity, and tool-calling capabilities.

All audio generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card. 8 Live Extended Thinking is rolling out starting today: Your information will be used in accordance with Google's privacy policy. You may opt out at any time.