arXiv:cs.LG· Weihan Li, Xinlei Chen, Junhao Wu, Tianshi Zheng·· 4 小时前AI 评分41
上下文学习器何时该扩展假设空间?arXiv 研究提出结构修正环境
When Should an In-Context Learner Expand Its Hypothesis Space?
AI 导读
研究把"是否扩展假设空间"建模为代价敏感的序列决策问题,并构建 Structural Revision Environment:同一历史在不同价格、剩余时间与查询下最优动作不同,修正是一条价值边界而非证据阈值。在该环境中训练的 Transformer 能从效用本身复现这一边界;三个后训练谱系的语言模型预测中含失败敏感信号,却未反映到修正决策中,允许推理的三个模型也未把历史转化为对扩展收益的估计。
正文
Abstract:Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference's proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.
| Comments: | 43 pages, 11 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09471 [cs.LG] |
| (or arXiv:2610.09471v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09471 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinlei Chen [view email]
[v1]
Wed, 7 Oct 2026 05:26:15 UTC (928 KB)
来源:arXiv:cs.LG · arxiv.org