arXiv:cs.LG· Christos Plachouras, Emmanouil Benetos, Johan Pauwels·· 4 小时前AI 评分30
音乐标注模式扩展:零样本预测还是少样本适配?
Extending Music Annotation Schemas: Zero-Shot Prediction or Few-Shot Adaptation?
AI 导读
研究基于 MGPHot 流行音乐标注数据集提出基准,模拟不同标注预算下的音乐模式扩展。结果显示,即便标注预算很小,监督适配也优于音频语言模型的零样本预测,而在预算适中时,复用冻结表示是最有效的方法,无需深度适配的调参。论文已投稿 IEEE ICASSP 2027,正在审稿中。
正文
Abstract:Automatic music annotation is typically tackled under the assumption of a fixed annotation schema. In practice, commercial music catalogs often need to accommodate new musical attributes as needs evolve. Given that expert music annotation is expensive, it is not evident which methodological approach is most effective at accommodating new attributes and backfilling existing tracks; audio-language models promise zero-shot prediction, but at what annotation budget does supervised adaptation become more compelling?
We propose a benchmark based on the MGPHot popular music annotation dataset for simulating music schema extension across different annotation budgets. We investigate zero-shot prediction with audio-language models, learning new attributes from pretrained representations, and adapting models trained on existing annotations. Our results suggest that supervised adaptation is more effective than zero-shot prediction even with small annotation budgets, while frozen representation reuse remains the most effective approach for modest budgets without the tuning required by deeper adaptation.
| Comments: | 5 pages, 3 figures. Submitted to IEEE ICASSP 2027; under review |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.06920 [cs.SD] |
| (or arXiv:2610.06920v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06920 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Christos Plachouras [view email]
[v1]
Fri, 2 Oct 2026 21:37:24 UTC (228 KB)
来源:arXiv:cs.LG · arxiv.org