跳到正文
arXiv:cs.AI· Md Asiful Islam, Fahmida Alam, Mihai Surdeanu·· 4 小时前

ARBITER:面向 LLM 护栏的双假设推理框架

A Dual-Hypothesis Reasoning Framework for LLM Guardrails

AI 导读

研究者提出 LLM 护栏框架 ARBITER,通过双假设推理在做出安全判定前同时考虑提示词的安全与不安全解读,并采用多组件监督微调(MC-SFT)将输出分解为逻辑组件按重要性加权。ARBITER 用自生成推理轨迹和 LoRA 参数高效微调替代昂贵的大模型或闭源教师模型方案,性能仍更优。在三个安全审核基准上,它超过现有推理与非推理护栏基线,域外评估提升明显,并能为不安全判定提供证据短语解释。

正文

View PDF HTML (experimental)

Abstract:We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly considers both safe and unsafe interpretations of a prompt before making a safety decision, and (ii) multi-component supervised fine-tuning (MC-SFT), a structured training loss for reasoning-based guardrails that decomposes LLM outputs into logical components and weights them according to their importance. Existing reasoning-based guardrails often rely on expensive procedures, such as generating reasoning traces using larger or closed-source teacher models and applying full-parameter fine-tuning. In contrast, ARBITER uses a cost-effective self-generation strategy for reasoning traces and LoRA-based parameter-efficient fine-tuning while still achieving better performance than these expensive approaches. Additionally, ARBITER provides faithful evidence-phrase explanations for unsafe decisions, enabling a more transparent and interpretable guardrail method. Experiments on three safety moderation benchmarks show that ARBITER outperforms existing reasoning-based and non-reasoning guardrail baselines, with clear gains in out-of-domain evaluations.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.17575 [cs.AI]
  (or arXiv:2607.17575v3 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.17575

arXiv-issued DOI via DataCite

Submission history

From: Md Asiful Islam [view email]
[v1] Mon, 20 Jul 2026 05:42:25 UTC (1,688 KB)
[v2] Mon, 5 Oct 2026 03:32:38 UTC (1,708 KB)
[v3] Wed, 7 Oct 2026 17:23:11 UTC (1,708 KB)

来源:arXiv:cs.AI · arxiv.org