跳到正文
arXiv:cs.CL· Shao-Chun Hu, Zi-Xiang Lin, Jeih-Weih Hung, Hung-Shin Lee·· 6 小时前AI 评分32

SEAL:面向高效语音分离的混合闭合加性重建与细化感知专家路由

SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation

AI 导读

SEAL(Sparse Expert routing with Additive Latent reconstruction)通过零和加性残差重建与专家路由,解决紧致时频分离器的掩码缩放和共享单元算力浪费问题。

正文

View PDF HTML (experimental)

Abstract:Compact time-frequency separators that mask the mixture and refine through a shared cell face two limits. First, a bounded multiplicative mask only scales a mixture bin, so where overlapping components cancel, the estimate stays small. Second, a shared cell applies the same weights to every time-frequency token at every step, so enlarging it adds compute everywhere. We present SEAL (Sparse Expert routing with Additive Latent reconstruction) to address both. For reconstruction, a zero-sum additive residual bounded by the local mixture amplitude lets estimates be nonzero where components cancel yet still sum to the mixture. For routing, a query built from acoustic and inter-step evidence sends each token to one of six residual experts, and a norm cap keeps the step cue from overriding clear acoustic evidence. On EchoSet, SEAL (small) surpasses TIGER (small) by 0.31 dB SI-SDRi with 28% fewer parameters and 2.9 times fewer MACs, and SEAL (large) is within 0.07 dB SI-SDRi of TIGER (large) at 3.1 times fewer MACs.
Comments: Submitted to ICASSP 2027
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as: arXiv:2610.07047 [eess.AS]
  (or arXiv:2610.07047v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2610.07047

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Hung-Shin Lee [view email]
[v1] Mon, 5 Oct 2026 02:22:41 UTC (222 KB)

来源:arXiv:cs.CL · arxiv.org