跳到正文
arXiv:cs.AI· Junyoung Koh, Hao-Wen Dong·· 4 小时前AI 评分34

重新审视人声合奏多音高估计中的输入时频表示:线性 STFT 优于 HCQT

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

AI 导读

针对人声合奏多音高估计,研究对比了 HCQT 与线性 STFT 两种输入时频表示,发现线性 STFT 在性能上超过 HCQT,同时大幅降低特征提取成本。进一步分析表明,更长的分析窗口或更宽的频谱覆盖并未带来额外提升,而将输入限制在预测音高范围内反而削弱了线性 STFT 的优势。

正文

View PDF HTML (experimental)

Abstract:Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03656 [cs.SD]
  (or arXiv:2610.03656v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.03656

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Junyoung Koh [view email]
[v1] Fri, 2 Oct 2026 17:35:06 UTC (1,485 KB)

来源:arXiv:cs.AI · arxiv.org