跳到正文
arXiv:cs.CL· Mo Yu, Yang Liu, Jing Qian·· 3 小时前AI 评分34

流式说话人分离器适配中的指标遮蔽:遗忘为何看似提升,以及排练的代价

When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal

AI 导读

一项研究对流式说话人分离器(streaming diarizer)在 7.5 小时双方对话数据上做小数据适配,发现适配显著提升域内分离性能并可迁移到独立语料。但改进并非在所有评测场景一致,额外混淆主要来自时间身份一致性受损而非说话人数量错误,局部重映射诊断显示不同语料的身份退化模式不同。排练(rehearsal)可减轻退化,却降低跨域迁移性能。

正文

View PDF HTML (experimental)

Abstract:Small-data adaptation can improve speech detection while degrading speaker attribution. We study this discrepancy in a released streaming diarizer adapted on 7.5 h of two-party conversation and evaluated across six corpora. Adaptation substantially improves in-domain diarization performance and transfers to an independent corpus. However, this improvement is not consistent across evaluation scenarios as the additional confusion is mainly associated with impaired temporal identity consistency rather than speaker-count errors. A local-remapping diagnostic reveals different patterns of identity degradation across corpora, indicating that adaptation may alter how streaming models maintain speaker assignments over time. Rehearsal reduces the observed degradation but reduces the cross-domain transfer performance. These results highlight the need to jointly evaluate detection accuracy, identity consistency, and retention behavior when adapting streaming diarization systems.
Comments: 5 pages, 4 figures
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.08828 [cs.CL]
  (or arXiv:2610.08828v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08828

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mo Yu [view email]
[v1] Sat, 26 Sep 2026 04:07:01 UTC (458 KB)

来源:arXiv:cs.CL · arxiv.org