arXiv:cs.AI· Marios Glytsos, Brian McFee·· 3 小时前
音频编码器的嵌入向量里还剩多少音频?一项音频编码器反演审计
How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders
AI 导读
研究通过配对源重建审计音频编码器保留的源信息:用共享的 Stable Audio Open 潜在扩散解码器,从 VGGish、ConvNeXt、CLAP 和 EnCodec 的冻结表征重建 5 秒、44.1-kHz 立体声音乐。
正文
Abstract:Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.
| Comments: | ICASSP 2027 submission; 5 pages, 2 figures, 2 tables |
| Subjects: | Sound (cs.SD); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.12250 [cs.SD] |
| (or arXiv:2610.12250v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12250 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Marios Glytsos [view email]
[v1]
Thu, 8 Oct 2026 16:24:41 UTC (280 KB)
来源:arXiv:cs.AI · arxiv.org