arXiv:cs.CL· Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried·· 3 小时前
用价值表征预测对齐泛化
Predicting Alignment Generalization with Value Representations
AI 导读
研究者提出"对齐泛化预测"任务,即预测将模型微调以遵循某一价值后,其在大量留出价值上的行为变化,并对 66 种对齐目标中的价值做了大规模分析。基于模型激活的表征显著优于文本描述方法,最佳激活方法的相关系数达 0.45,而描述基线仅为 0.05。该表征还可用于衡量多价值对齐目标中各价值的相似度,并与模型鲁棒性显著相关。
正文
Abstract:LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12410 [cs.CL] |
| (or arXiv:2610.12410v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12410 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Andy Liu [view email]
[v1]
Thu, 8 Oct 2026 17:47:26 UTC (1,160 KB)
来源:arXiv:cs.CL · arxiv.org