arXiv:cs.LG· Louis McCallum, Mick Grierson·· 6 小时前AI 评分30
RAVE 与 EnCodec 音频解码器如何编码音高、BPM 等音乐特征:输入、深度与分布的影响
Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution
AI 导读
研究对 RAVE 解码器激活做了逐层与跨层聚类分析,覆盖三个不同音乐领域训练的模型和四类刺激,并用通用 EnCodec 验证架构泛化。合成刺激编码良好(音高 |ρ|=0.45、BPM |ρ|=0.76),自然音频编码减弱但依然显著,非线性探针下平均 R²=0.56,比线性探针提升 0.152。
正文
Abstract:Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
| Comments: | This manuscript has been accepted for publishing in IEEE Transactions on Audio, Speech and Language Processing (TASLP) |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07966 [cs.SD] |
| (or arXiv:2610.07966v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07966 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Louis McCallum [view email]
[v1]
Tue, 6 Oct 2026 08:32:02 UTC (3,498 KB)
来源:arXiv:cs.LG · arxiv.org