arXiv:cs.CL· Jiachen Zhao, Zhengxuan Wu, David Bau, Weiyan Shi·· 3 小时前AI 评分50
Persona Hierarchy Model:理解 LLM 微调中的上下文泛化
The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs
AI 导读
研究者提出 Persona Hierarchy Model,用共享默认人格解释微调行为为何有时局限于特定上下文、有时泛化到未见上下文。在覆盖 4 种行为、15 种训练上下文的 120 个微调模型上,泛化狭窄度与训练上下文人格和默认人格的相似度正相关(Qwen3-4B 的 Pearson r = 0.72)。
正文
Abstract:Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a shared default persona influences behavior across contexts. Under this model, fine-tuning that modifies the shared persona promotes broader transfer, whereas changes to local personas remain more context-specific. Across 120 fine-tuned models spanning four behaviors and 15 training contexts, generalization narrowness positively correlates with the similarity between the training context's persona and the default persona (Pearson's r = 0.72 for Qwen3-4B). Prior fine-tuning under the default context can broaden generalization in subsequent training under other contexts. Aligning contextual responses with default-persona responses produces stronger effects. Finally, we propose persona-preserving regularization (PPR) to confine undesired contextual generalization. In RL, PPR cuts reward hacking from 42-55% to at most 0.2% under every evaluated prompt while retaining accuracy gains. These results support the Persona Hierarchy Model as an explanation for contextual generalization and can motivate future controls on unintended generalization for better alignment of LLMs.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09384 [cs.CL] |
| (or arXiv:2610.09384v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09384 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiachen Zhao [view email]
[v1]
Wed, 7 Oct 2026 03:40:44 UTC (1,786 KB)
来源:arXiv:cs.CL · arxiv.org