arXiv:cs.LG· Drew Edwards, Akira Maezawa, Simon Dixon·· 3 小时前AI 评分38
用交叉注意力条件化学习爵士钢琴家风格
Learning Jazz Pianist Style with Cross-Attention Conditioning
AI 导读
研究者基于预训练符号音乐 Transformer,通过交叉注意力对学习到的钢琴家身份嵌入进行条件化,实现按特定艺术家风格生成音乐。滑动窗口分类器能将条件化续写稳定归因到正确艺术家,远超无条件基线;仅用合成生成数据训练的分类器在 12 类真实钢琴家上达到 87% 片段级和 95% 整曲级准确率。该分类器还被用于定位演奏中最具个人风格的音乐片段。
正文
Abstract:Jazz pianists develop distinctive traits that experienced listeners can often identify within seconds, yet the features underlying this recognition resist formal description. We study jazz pianist style through the lens of a pretrained symbolic music transformer, showing that its learned representations already encode pianist identity well enough for highly accurate classification across two benchmarks. We then augment the transformer with cross-attention over learned pianist identity embeddings, enabling it to generate music conditioned on a specific artist's style. Two evaluation protocols confirm that the generator captures meaningful stylistic structure: a sliding-window classifier consistently attributes conditioned continuations to the correct artist, far above unconditioned baselines; and a classifier trained entirely on synthetic generations identifies real pianists across 12 classes with 87% chunk-level and 95% song-level accuracy. Finally, we repurpose the classifier to locate the most characteristic moments within a performance, surfacing the specific musical gestures that distinguish each pianist's voice.
| Comments: | 8 pages, 6 figures. Accepted at ISMIR 2026. Audio demos, code and checkpoints: this https URL |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS) |
| ACM classes: | I.2.6; H.5.5 |
| Cite as: | arXiv:2610.02918 [cs.SD] |
| (or arXiv:2610.02918v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02918 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Drew Edwards [view email]
[v1]
Fri, 2 Oct 2026 07:05:22 UTC (312 KB)
来源:arXiv:cs.LG · arxiv.org