跳到正文
arXiv:cs.CL· Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov·· 6 小时前AI 评分41

不操控只读取:用 Router Logits 为 MoE 视觉语言模型做多模态安全检测

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

AI 导读

研究者提出一种轻量级 router-logit 安全检测器,在 prompt prefill 阶段读取 MoE 视觉语言模型的 routing 信号,在不修改模型参数与专家路由的前提下,于生成前识别不安全请求。在 Qwen3-VL 和 Kimi-VL 上,该检测器大幅降低 HoliSafe 基准的安全错误,并泛化到 MISHard 和 MM-SafetyBench 等分布外安全基准。

正文

View PDF HTML (experimental)

Abstract:Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.
Comments: 15 pages, 4 tables, 11 figures
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07774 [cs.CL]
  (or arXiv:2610.07774v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07774

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ziyuan Yang [view email]
[v1] Tue, 6 Oct 2026 05:09:02 UTC (5,243 KB)

来源:arXiv:cs.CL · arxiv.org