arXiv:cs.LG(机器学习,全量分类)· SeongHyeon Kim, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, Juho Lee·· 1 天前AI 评分44
Mixture-Trained Merging:面向统一多目标模型的混合训练合并方法
Mixture-Trained Merging for Unified Multi-Objective Models
AI 导读
研究者提出 Mixture-Trained Merging(MTM),让每个分支在目标偏置的数据混合上训练而非单一目标,从而在权重空间合并时保持兼容性。在代码、数学、指令遵循和 think/non-think 控制上,MTM 优于朴素合并,并避免了单目标合并导致的 think/non-think 模式坍缩。
正文
Abstract:Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on single objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.
| Comments: | Accepted at NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01238 [cs.LG] |
| (or arXiv:2610.01238v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01238 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: SeongHyeon Kim [view email]
[v1]
Thu, 1 Oct 2026 07:37:37 UTC (193 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org