跳到正文
arXiv:cs.LG· Nikita Torgashov, Okan K\"op\"ukl\"u·· 3 小时前AI 评分37

FASTDIAR:面向流式说话人分离的帧级说话人编码器

FASTDIAR: Frame-level speaker encoder for Streaming Diarization

AI 导读

FASTDIAR 将 SOTA 说话人识别架构改造为因果帧级编码器,单次读取音频流,基于过去两秒窗口每 80ms 输出一个嵌入向量,并配合在线聚类。该系统仅通过 utterance 级教师模型的蒸馏训练,在低重叠基准上以亚秒级延迟成为最准确的流式说话人分离器,说话人数量增长时性能退化远小于基于缓存的系统,单 CPU 线程上运行速度达实时的 5 倍。

正文

View PDF HTML (experimental)

Abstract:Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.
Comments: 5 pages, submitted to IEEE ICASSP 2027
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
Cite as: arXiv:2610.02941 [eess.AS]
  (or arXiv:2610.02941v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2610.02941

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nikita Torgashov [view email]
[v1] Fri, 2 Oct 2026 07:32:11 UTC (20 KB)

来源:arXiv:cs.LG · arxiv.org