arXiv:cs.CL· Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang, Yue Liu, Dong Huang, Huijun Liu, Bin Ji, Jie M. Zhang, See-Kiong Ng·· 4 小时前AI 评分33
通过模型路由与协作实现集体偏见缓解
Collective Bias Mitigation via Model Routing and Collaboration
AI 导读
研究提出 Collective Bias Mitigation(CBM)框架,通过学习细粒度模型行为并在多个 LLM 间共享知识来缓解偏见,是首个系统探索不同 LLM 有效选择与组织以生成更公平回答的工作。
正文
Abstract:Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model's intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained model behavior and fostering knowledge sharing among diverse LLMs. This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses. Experiments show CBM substantially outperforms standalone baselines (e.g., in the top-7 setting, Committee lowers the age bias score from 0.25 to 0.10). Our Debating and Committee topologies achieve substantial bias reduction, with the latter balancing mitigation effectiveness and inference cost, highlighting the potential of CBM for fairer LLMs.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03240 [cs.CL] |
| (or arXiv:2610.03240v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03240 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mingzhe Du [view email]
[v1]
Fri, 2 Oct 2026 12:49:44 UTC (1,016 KB)
来源:arXiv:cs.CL · arxiv.org