On July 31, Alibaba Tongyi Qwen officially released the speech recognition large model Qwen-Audio-3.0-ASR, featuring "no word loss in long audio and no need to teach industry-specific terms," and is now available for use on the Alibaba Cloud BaiLian platform.
The new model has been upgraded in five key dimensions:
First, long audio context memory, which allows the transcription to refer to previous context, keeping names and technical concepts consistent throughout several hours of meetings, eliminating the problem of "fragmentation" across segments;
Second, built-in industry-specific dictionaries, with a recall rate of 95.36% for medical terminology and 91.87% for IT programming, allowing accurate identification of obscure terms without manual configuration;
Third, tiered hot word customization, where enterprise-specific vocabulary takes effect immediately, achieving a recall rate of over 99% in most scenarios, without false triggers due to an increase in hot words; fourth, integrated voice polishing, completing the removal of filler words, cleaning up repetitions, handling self-corrections, and semantic reorganization in one step, producing content with readability close to the two-step approach of "ASR + large model polishing";
Fifth, a single model covering more than 30 languages, with an average semantic error rate of 17.09% for seven languages, outperforming industry models such as Azure and Gemini, suitable for international meetings and overseas customer service scenarios.
The simultaneously launched Qwen-Audio-3.0-ASR-Streaming is designed for real-time scenarios, with a theoretical character output delay of 300 milliseconds, a character error rate of 7.80% in Chinese industrial scenarios and 11.52% in English, leading the industry. This series previously ranked first globally with a 1.7% error rate in the Artificial Analysis evaluation, and is widely applied in areas such as meeting minutes, real-time subtitles, educational recordings, and intelligent customer service.







