跳到正文
arXiv:cs.AI· Cheol Woo Kim, Jai Moondra, Roozbeh Nahavandi, Andrew Perrault, Milind Tambe, Swati Gupta·· 4 小时前AI 评分33

PALM:面向多目标 LLM 对齐的紧凑策略组合

Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment

AI 导读

研究者提出 PALM(Portfolio of Aligned LLMs)算法,通过结构化权重向量网格、惰性搜索与剪枝,在给定近似容差下返回一个可证明对每个权重向量都包含近最优策略的 LLM 组合,并给出组合规模的显式上界。实验显示,PALM 的近似误差普遍小于同等规模下均匀间隔或随机采样权重构建的组合,并可扩展至更高维奖励空间。该组合可支持个性化、奖励权重探索与紧凑的解码时配置。

正文

View PDF HTML (experimental)

Abstract:Aligning large language models (LLMs) requires balancing competing objectives such as helpfulness, harmlessness, and conciseness. The appropriate balance varies across users and applications, yet training, evaluating, and deploying many policies across different reward weights is costly. We study how to identify a small portfolio of LLMs that preserves near-optimal performance across all reward weightings. We propose PALM (Portfolio of Aligned LLMs), an algorithm that combines a structured grid of weight vectors, a lazy search that optimizes policies only where needed, and pruning. Given target approximation tolerances, PALM returns a portfolio that provably contains a near-optimal policy for every weight vector, with an explicit upper bound on portfolio size. Such portfolios can support scalable personalization, reward-weight exploration during model development, and compact decoding-time configurations. Experiments show that PALM generally achieves smaller approximation gaps than same-size portfolios built from uniformly spaced or randomly sampled weights. We further demonstrate that PALM scales effectively to higher-dimensional reward spaces through efficient search and sparse preference structure.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
MSC classes: 68T07
Cite as: arXiv:2604.04144 [cs.CL]
  (or arXiv:2604.04144v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2604.04144

arXiv-issued DOI via DataCite

Submission history

From: Cheol Woo Kim [view email]
[v1] Sun, 5 Apr 2026 15:12:07 UTC (881 KB)
[v2] Fri, 10 Apr 2026 17:55:07 UTC (881 KB)
[v3] Thu, 1 Oct 2026 20:55:36 UTC (844 KB)

来源:arXiv:cs.AI · arxiv.org