Google has officially released the latest speech-to-text model in the Gemini series, Gemini 3.5 Transcribe, and positioned it as the most accurate speech recognition model in the series to date. The model is now fully open to developers, enterprises, and general users, with a focus on overcoming long-standing industry challenges such as noisy environments, complex technical terms, and natural spoken expressions.
In terms of core technology and application scenarios, Gemini 3.5 Transcribe demonstrates strong generalization capabilities. The new model not only effectively filters background noise but also provides excellent support for complex vocabulary in specialized fields. Currently, the model has significantly improved performance for the Gemini app on macOS and the Rambler feature in Gboard keyboard on Android. For developers, this tool can efficiently build innovative applications such as voice-enabled agents, real-time captioning tools, and post-call analysis workflows.

Regarding the complexity of human daily conversations, the new model has undergone multiple refined optimizations in its functionality. It better captures the natural semantic intent of speakers, identifies custom vocabulary, and assists users in completing various tasks smoothly through voice. Additionally, the model has several built-in practical features, including automatic self-correction, intelligent filtering and deletion of filler words like "um" and "uh," and automatic formatting of output text. More impressively, it can delegate more complex chain tasks such as file analysis and image generation to other Gemini models, providing users with an end-to-end multi-model collaboration experience.
In terms of objective transcription accuracy, Google's public data shows that the average word error rate (WER) of Gemini 3.5 Transcribe in streaming speech recognition scenarios is as low as 4.0%, and in non-streaming scenarios, it drops significantly to 2.6%. Even in noisy and complex environments, it maintains high recognition stability and is more accurate and error-free when identifying critical information composed of letters and numbers, such as order numbers and postal codes.
In terms of multilingual support and personalized adaptability, the model also performs outstandingly. Currently, it supports a total of 85 languages and can perfectly adapt to different regional accents and dialect variations. For preprocessed audio, it supports speaker identification with timestamps for up to three different speakers. In addition, the model highly supports user-defined vocabulary lists, which can perfectly handle professional terminology, special spellings, and exclusive names across various industries.
In terms of openness and product ecosystem integration, developers can currently experience Gemini 3.5 Transcribe in preview form through the Gemini API in Google AI Studio and Google Antigravity, while general users can directly experience it on Android and macOS platforms. In the future, Google plans to seamlessly integrate it into the Chrome browser, allowing users to conveniently input voice directly in any web page input field. Alongside the previous launch of Gemini 3.7 Flash, the full rollout of Gemini in Chrome to all U.S. Android users, and the integration of the assistant into Waymo autonomous taxis, Google is continuously improving its full-scenario AI product matrix at a high pace.


