Speech synthesis technology is entering an unprecedented era of refinement. On September 23 local time, Google officially announced the launch of two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, fully meeting the diverse needs of developers and creators in personalized character voice acting and large-scale audio generation.
In terms of product positioning and core functions, the two models have different focuses. The Gemini 3.8 Flash TTS mainly targets voice and character design, supporting users to create customized voices through intuitive natural language descriptions, and accurately control core elements such as tone, speech rate, and accent sentence by sentence. In addition, this model also supports efficient voice replication based on a short 30-second audio sample. The other model, Gemini 3.8 Flash-Lite TTS, focuses on large-scale audio generation and efficient voice acting scenarios.
In terms of ecosystem and multilingual support, Google stated that these two new models fully support more than 100 languages and dialects, and come with 2000 pre-made voices ready to use, providing strong underlying computing power for voice application development and content creation worldwide.



