Qwen-Audio-3.0-TTS: A Voice Synthesis Large Model from Qwen - Direct Command with Natural Language, Real-time Version First Packet Delay Reduced to 300 Milliseconds
Qwen released Qwen-Audio-3.0-TTS, a text-to-speech model featuring free-style natural language instruction control. Users can simply describe desired tone, pace, and style in everyday language, eliminating the need to adjust engineering parameters and drastically lowering the usage barrier.....