StepXingchen officially launched the new StepAudio3 series of voice large models on September 15th. This release includes five products: StepAudio3Realtime, StepAudio3ASR, StepAudio3TTS, StepAudio3Gen, and StepAudio3Music. Several models have successfully ranked first globally in the authoritative evaluation list Artificial Analysis. Currently, these five models are fully available on the StepXingchen open platform, targeting five core application scenarios: realistic voice generation, comprehensive audio content generation, real-time voice interaction, complex scenario voice understanding, and music creation.

In the field of real-time interaction, StepAudio3Realtime is dedicated to achieving full-duplex native real-time conversation. It ranked first globally with a comprehensive score of 98.9% in the Artificial Analysis Conversational Dynamics ranking, and it also ranked first globally in the Artificial Analysis Speech Reasoning ranking with a speech reasoning accuracy of 99.7%. The model can accurately grasp the rhythm of the conversation, process user interruptions and continuous feedback, and support comprehensive understanding of semantics, tone, emotion, paralanguage, and ambient sounds. In addition, it supports parallel execution of reasoning and voice generation, and asynchronous execution of tool calls (Tool Call) and long tasks, ensuring that voice conversations are not blocked.

In the area of speech recognition, StepAudio3ASR integrates high-precision speech recognition with the knowledge reasoning capabilities of large language models, surpassing traditional pure sound transcription. It supports various scenarios such as Chinese, English, dialects, mixed Chinese and English, long audio, and professional fields. It performs excellently in medical, legal, financial, automotive, and programming fields, and can even easily handle complex input environments such as low volume, fast speech rate, unclear pronunciation, singing, and background music. Its non-streaming speech recognition word error rate (WER) is only 1.7%, tying for first globally.

In the area of voice generation, StepAudio3TTS is a realistic voice generation model designed for real-time interaction. It perfectly aligns with human natural spoken patterns in terms of voice base, intonation levels, rhythm, and pauses. It not only can reproduce paralinguistic expressions such as laughter, hesitation, stammering, repetition, and correction, but can also autonomously adjust emotional tone based on semantic content. With a streaming generation architecture, the model can synchronize generation and playback.

In the field of comprehensive audio content generation, StepAudio3Gen changes the long production, editing, and mixing chain of traditional audio. Users need only natural language descriptions and reference audio to generate vocals, sound effects, ambient sounds, and background music in one go. The model supports fine control over paralinguistic features such as character voice, speaking style, emotion, dialects, and laughter. It can also specify the timing and order of dialogue, sound effects, and background music, directly participating in the entire content's sound design and time sequencing.

In the field of music creation, StepAudio3Music supports both zero-generation and multi-round interactive creation based on ABC notation. Users can input through lyrics, a cappella, reference songs, or notation, and use natural language to directly control genre, vocals, melody, rhythm, instruments, and emotions. The model has deep modeling of song structure, cross-paragraph melodic development, and energy changes. Especially in the a cappella accompaniment scenario, just a single a cappella can automatically supplement harmony, rhythm, instruments, and complete arrangements, truly entering the real music production process.