跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Shenyan Zheng, Jiayou Zhong, Anudeex Shetty, Heng Ji, Preslav Nakov, Usman Naseem·· 15 小时前AI 评分37

VISPA:通过自动价值选择与激活实现多元对齐

VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

AI 导读

研究者提出免训练的多元对齐框架 VISPA,通过动态选择与模型内部激活引导,直接控制大语言模型的价值观表达。在医疗等多个领域的多种模型和评测设置中,VISPA 在全部多元对齐模式下均表现良好,并可适配不同的引导初始化方式、模型与价值观。该工作已被 EMNLP 2026 主会接收。

正文

View PDF HTML (experimental)

Abstract:As large language models are increasingly used in high-stakes domains, it is essential that their outputs reflect not average} human preference, rather range of varying perspectives. Achieving such pluralism, however, remains challenging. Existing approaches consider limited values or rely on prompt-level interventions, lacking value control and representation. To address this, we introduce VISPA, a training-free pluralistic alignment framework, that enables direct control over value expression by dynamic selection and internal model activation steering. Across extensive empirical studies spanning multiple models and evaluation settings, we show VISPA is performant across all pluralistic alignment modes in healthcare and beyond. Further analysis reveals VISPA is adaptable with different steering initiations, model, and/or values. These results suggest that pluralistic alignment can be achieved through internal activation mechanisms, offering a scalable path toward language models that serves all.
Comments: Accepted to EMNLP 2026 (Main Proceedings)
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2601.12758 [cs.CL]
  (or arXiv:2601.12758v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2601.12758

arXiv-issued DOI via DataCite

Submission history

From: Anudeex Shetty [view email]
[v1] Mon, 19 Jan 2026 06:38:52 UTC (3,239 KB)
[v2] Thu, 1 Oct 2026 12:20:12 UTC (2,426 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org