Microsoft officially released the new speech recognition model MAI-Transcribe-2 on September 3. As the most powerful and efficient transcription product to date, this model achieved an average word error rate (WER) of 5.2% in the FLEURS benchmark test (covering 60 languages), and ranked second in the Artificial Analysis word error rate ranking.

While maintaining high accuracy, its long audio processing speed can be up to 10 times faster than that of competitors, with a processing time of about 10 seconds for one hour of audio. The model is now available on Microsoft Foundry, MAI Playground, and Open Router, and offers a limited-time discounted price of just $0.10 per hour of audio.
MAI-Transcribe-2 has been comprehensively upgraded for real-world complex scenarios, integrating features such as speaker segmentation, word-level timestamps, keyword preferences, and automatic language detection. It supports code-switching phenomena such as Hinglish (Indian English) and Spanglish (Spanish-English mix), and has strong noise resistance. Developers can also freely switch between "word-for-word transcription" and "concise transcription" styles to accurately adapt to diverse production environments such as legal, clinical, subtitles, and high-compliance analysis.

As AI applications evolve from single-modal interaction to multi-modal and real-time speech intelligence, high-throughput, low-latency, and low-cost speech infrastructure is becoming a key competitive barrier. MAI-Transcribe-2 redefines the accuracy and latency Pareto frontier of speech transcription, significantly lowering the deployment threshold for multilingual audio, and injecting new momentum into the development of end-to-end real-time speech ecosystems.




