跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 1 天前AI 评分60
AI 导读

阶跃星辰(StepFun)与 ACE Studio 发布 StepAudio 3 Music,是其首个音乐生成基础模型,可根据提示词与歌词生成完整歌曲,支持歌曲生成、纯器乐生成、音乐翻唱和人声编曲四种工作流,基于 ABC-COT 先规划音乐结构与编曲再合成,交互 demo 见 https://static.stepfun.com/blog/stepaudio3/music/。

正文

StepFun has released StepAudio 3 Music, a model that turns a text description and lyrics into a finished song.

The interesting engineering result is that better audio reconstruction did not necessarily produce better music generation.

The technical report compares single-codebook and residual vector quantization approaches. The final tokenizer uses one token stream at 50 Hz with 65,536 entries, followed by a flow-matching DiT renderer.

Why this matters: the representation has to preserve sound AND give the autoregressive model a sequence it can predict reliably. Optimizing the codec in isolation can miss that trade-off.

(In comment you will find some audio examples to give this architecture some context)

引用StepFun@StepFun_ai
StepFun and ACE Studio present StepAudio 3 Music, StepFun’s first music generation foundation model — built to turn a prompt and lyrics into a complete song. Describe the sound you want: • Genre, mood and vocal character • Instruments, key and BPM • Song structure and arrangement Four workflows in one model: 🎤 Song generation 🎹 Instrumental generation 🔁 Music cover 🎙️ Vocal-to-song arrangement Powered by ABC-COT, it plans musical structure and arrangement before synthesis. Generate a version, rewrite the prompt, edit the ABC notation, and iterate toward the sound in your head. (ABC-COT API coming soon) Try the interactive demo: › https://static.stepfun.com/blog/stepaudio3/music/
在 X 查看被引用的帖子

来源:Rohan Paul · x.com