阶跃星辰发布 StepAudio 3,是涵盖实时语音、语音识别、语音生成、音频生成和音乐的 5 个音频模型家族。Realtime 在 Artificial Analysis 的 Conversational Dynamics 达 98.9%、Speech Reasoning 达 99.7%,均排第一;ASR 达到 1.7% WER。
Introducing StepAudio 3, our new family of 5 audio models for real-time voice, speech recognition, speech generation, audio generation and music.
Realtime ranks #1 on Artificial Analysis for both Conversational Dynamics (98.9%) and Speech Reasoning (99.7%). ASR reaches 1.7% WER, matching the best result on the leaderboard.
Build voice agents that handle interruptions, reason while speaking, and call tools. Transcribe speech, generate expressive voices, and create full audio scenes and music.
Available now:
Voice AI Lab: https://audio.stepfun.ai/
Blog: https://static.stepfun.com/blog/stepaudio3/
来源:StepFun · x.com