On the occasion of its anniversary, Fish Audio officially launches its most powerful production-level speech large model to date, S2.1Pro. Unlike traditional text-to-speech (TTS) systems designed for scripted narration, this model is specifically built for real-time dialogue scenarios and is now available via the Fish Audio API, with a free version offering reasonable limits for development and testing purposes.

In terms of technical performance, S2.1Pro achieves a first-frame audio playback delay of approximately 90 milliseconds, seamlessly supporting natural and smooth conversation turn-taking; the model supports 83 languages and uses a unified speech recognition architecture. In terms of voice control flexibility, the model breaks free from the limitations of pre-set emotion menus, allowing direct embedding of free-form bracket tags within the text to precisely implement in-line fine-grained instructions such as whispers and tense laughter.
Additionally, S2.1Pro natively supports multi-speaker dialogue generation, and can complete high-quality voice cloning with just 10 to 30 seconds of short reference samples, accurately replicating the target tone and speaking style without the need for additional fine-tuning.
The launch of S2.1Pro marks the accelerating evolution of speech AI from traditional script-based synthesis toward real-time dialogue modes with ultra-low latency and high-fidelity interaction, providing a lower barrier and highly available computing infrastructure for the global deployment of intelligent speech interaction.



