arXiv:cs.LG· Shuyu Gan, Young-Jun Lee, Dongyeop Kang·· 4 小时前AI 评分39
SanSi:一种面向 System 1.5 思维的回环类型化决策模型
SanSi: A Looped Typed Decision Model for System 1.5 Thinking
AI 导读
SanSi 将预训练回环语言模型改造为类型化决策模型,每个循环用 proper scoring rule 训练并在每次循环后读取选项概率,单次运行即可覆盖 1 到 8 个循环的任意预算。
正文
Abstract:Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.
| Comments: | 43 pages, 15 figures, 42 tables. Project page: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07730 [cs.CL] |
| (or arXiv:2610.07730v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07730 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuyu Gan [view email]
[v1]
Tue, 6 Oct 2026 04:28:22 UTC (511 KB)
来源:arXiv:cs.LG · arxiv.org