arXiv:cs.LG· Nikita Torgashov, Okan K\"op\"ukl\"u·· 3 小时前AI 评分37
FASTDIAR:面向流式说话人分离的帧级说话人编码器
FASTDIAR: Frame-level speaker encoder for Streaming Diarization
AI 导读
FASTDIAR 将 SOTA 说话人识别架构改造为因果帧级编码器,单次读取音频流,基于过去两秒窗口每 80ms 输出一个嵌入向量,并配合在线聚类。该系统仅通过 utterance 级教师模型的蒸馏训练,在低重叠基准上以亚秒级延迟成为最准确的流式说话人分离器,说话人数量增长时性能退化远小于基于缓存的系统,单 CPU 线程上运行速度达实时的 5 倍。
正文
Abstract:Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.
| Comments: | 5 pages, submitted to IEEE ICASSP 2027 |
| Subjects: | Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD) |
| Cite as: | arXiv:2610.02941 [eess.AS] |
| (or arXiv:2610.02941v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02941 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nikita Torgashov [view email]
[v1]
Fri, 2 Oct 2026 07:32:11 UTC (20 KB)
来源:arXiv:cs.LG · arxiv.org