The next-generation real-time translation large model Qwen3.8-LiveTranslate has been comprehensively upgraded in four dimensions: translation quality, latency, speaker identification, and speech synthesis.
This model adopts an audio-text interleaved Interleave architecture, integrating streaming understanding, text output, and speech generation into a single causal sequence. The average latency is reduced from 2.8 seconds in the previous generation to 2.3 seconds.

The three new capabilities are the core highlights of this upgrade: real-time speaker separation clearly identifies who speaks each sentence during multi-person alternating speeches, and the translated speech can stably retain the original speaker's voice timbre; simultaneous display of source and target text achieves bilingual synchronization, catering to both immediate understanding and source verification; long-context disambiguation reduces translation ambiguity caused by proper nouns and pronouns through cross-turn contextual association.

The model uses a Hybrid MoE Thinker–Talker dual-module design. The Thinker is responsible for understanding and translation, while the Talker generates the translated speech with the original voice timbre. On the Omnilingua-MSpeaker multi-speaker long audio evaluation set covering 14 language directions, the model outperforms current mainstream real-time translation systems in translation fidelity, fluency, conciseness, and speaker separation error rate.
The model currently supports 60 languages, and the API is now available on the Qwen AI platform. The team stated that the next step will focus on further reducing end-to-end latency, cross-conversation long-term memory, and expanding support for more languages.




