iFlytek Xinghuo Speech Base Large Model Released: 0.65B Encoder with 30B MoE, Entirely Domestic Computing Power Trained
iFlytek has launched the speech base large model Spark-Audio-1.0-Preview, trained using entirely domestic computing power. Traditional audio processing often adopts a cascading approach, first converting speech to text and then understanding it, which has two major shortcomings: information loss (ignoring tone, emotion, background sounds) and fragmented processes (separate stages such as speech transcription and voiceprint recognition), affecting the completeness of scenario judgment.