arXiv:cs.CL· Ming-Hsiang Hu, Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen·· 6 小时前AI 评分34
GIVE-KWS:用门控注入视觉证据实现抗噪的示例查询关键词识别
GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting
AI 导读
GIVE-KWS 通过门控交叉注意力将唇动视觉证据注入查询音频,在 -10 dB 下将未见关键词的 EER 降低 72.9%,平均降低 62.8%。研究发现视觉鲁棒性依赖两个条件:带音素信息的视觉表征,以及注入式融合而非缩放音频特征。在带音素编码器下,注入比掩码在 -10 dB 时带来 4.0-9.3 dB 的有效 SNR 增益;编码器缺乏音素信息时增益近乎消失。
正文
Abstract:Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder's lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bearing encoder, injection yields an effective SNR gain of 4.0-9.3 dB over masking at -10 dB, whereas under a phoneme-poor one it nearly vanishes. Relative to the benchmark system, GIVE-KWS reduces unseen-keyword EER by 72.9% at -10 dB and 62.8% on average.
| Comments: | Submitted to ICASSP 2027 |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD) |
| Cite as: | arXiv:2610.07046 [eess.AS] |
| (or arXiv:2610.07046v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07046 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hung-Shin Lee [view email]
[v1]
Mon, 5 Oct 2026 02:00:29 UTC (190 KB)
来源:arXiv:cs.CL · arxiv.org