Google has officially announced its latest breakthrough in the field of native end-to-end voice interaction, launching a new generation of real-time speech large model Gemini3.8Live and Gemini3.8Live Extended Thinking, which is customized for high-complexity tasks. These two models represent a significant leap in Google's native audio technology system. They not only push the real-time conversation delay and naturalness to a new level but also, for the first time, achieve deep multi-step reasoning capabilities in low-latency voice streams, completely changing the application paradigm of voice artificial intelligence in complex professional scenarios.

In previous mainstream voice interaction architectures, AI systems usually had to compromise between fast response shallow conversations and time-consuming deep thinking. When facing difficult questions, users often had to endure long periods of silence. The newly launched Gemini3.8Live by Google focuses on optimizing ultra-fast conversations and daily streaming interactions. It can provide users with extremely smooth and human-like real-time voice communication with more sensitive tone fluctuations, natural interruptions, and context understanding capabilities.

image.png

Building upon this, Gemini3.8Live Extended Thinking offers a fundamental upgrade for high-complexity business flows. This model gives AI the ability to "think while speaking" in the background, allowing the system to perform complex multi-step logical decomposition, code debugging, or technical diagnostics in the background while maintaining real-time continuous conversation, without abruptly interrupting the conversation flow when encountering difficult questions.

In terms of practical applications, the new Live series models will greatly expand the service boundaries of voice assistants. In addition to daily natural conversations, Gemini3.8Live and its extended reasoning version have shown significantly better performance than previous versions in high-intelligence demand scenarios such as real-time step-by-step troubleshooting, complex mathematical derivation, real-time foreign language simultaneous tutoring, and multi-task streaming instruction execution, ranking at the top in multiple multimodal and speech reasoning industry benchmark evaluations.

For developers and enterprise users, Google announced that the relevant model capabilities are now officially available through official APIs and development workstations starting today. Developers can directly call this ultra-low latency native audio interface to quickly build next-generation full-duplex voice agent services with advanced reasoning capabilities in product forms such as smart hardware, customer service, real-time programming assistance, and collaborative office.