Alibaba officially released the large-scale speech recognition model Qwen-Audio-3.0-ASR-Flash, with a clear core goal—enabling AI to accurately understand specialized terminology. The model has made systematic optimizations in three dimensions: context consistency, industry-specific word recognition, and hot word customization. It also has voice polishing capabilities, allowing it to directly output structured text.

The research team has continuously mined professional vocabulary from fields such as healthcare, IT programming, stock markets, and social celebrities, building a high-quality vocabulary database covering multiple industries. In the latest internal evaluation of industry-specific terms, the new model showed significant improvements in "listening accuracy" across various industries, with medical scenarios reaching 95.36%.
Three versions cover five scenarios, and it has already taken the global top position
The Qwen-Audio-ASR-Flash series has been verified in scenarios such as meeting note organization, real-time subtitles, educational recording, and intelligent customer service. Previously, it achieved the number one ranking globally on the AI evaluation platform Artificial Analysis with an error rate of 1.7%. This means that the accuracy of AI speech-to-text in real office and educational environments has reached the ceiling of usability.
The model is now available for access through the Alibaba Cloud BaiLian platform, offering three versions: the Flash version for real-time speech recognition up to 5 minutes, the Filetrans version for offline file transcription, and the Streaming version for real-time speech recognition. When AI not only "can listen," but also understands those complicated technical terms in your industry, the last mile of voice interaction may soon be solved.






