arXiv:cs.CL· Thea Aviss·· 4 小时前
State Stream Transformer (SST) V2:用并行训练的非线性循环实现潜在空间推理
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
AI 导读
State Stream Transformer (SST) V2 通过每层 FFN 驱动的非线性循环,在连续潜在空间沿全序列横向流式传递状态,并支持推理时按位置进行连续潜在推演,两遍并行训练近似顺序循环。
正文
Abstract:Current transformers discard their rich latent residual stream between positions, reconstructing latent reasoning context at each new position and leaving potential reasoning capacity untapped. The State Stream Transformer (SST) V2 enables parameter-efficient reasoning in continuous latent space through an FFN-driven nonlinear recurrence at each decoder layer, where latent states are streamed horizontally across the full sequence via a learned blend. This same mechanism supports continuous latent deliberation per position at inference time, dedicating additional FLOPs to exploring abstract reasoning before committing to a token. A two-pass parallel training procedure approximates the sequential recurrence, making co-training computationally practical. Hidden state analysis shows that the state stream facilitates reasoning through sharp, content-dependent reorganisations in continuous latent space; the LM head exposes the resulting latent belief states through the output distribution, while the state stream carries them forward to influence future positions. A learned probe shows that at the first generated token, the latent state already predicts whether the eventual answer will survive or break under additional latent computation for every subsequent position. Co-trained into an existing 27B backbone using only a small dataset of GSM8K examples and evaluated using an oracle to route each question to a depth of one to four recurrent forward passes per token, the SST achieves an architectural capacity bound of 61.11% on out-of-distribution GPQA-Diamond, a +15.15 point gain over a fine-tuning-matched baseline, and cuts that same baseline's remaining GSM8K errors by 46%. Together, these results provide a method for efficiently training a nonlinear recurrence and show that the state stream provides additional reasoning capacity beyond that of an otherwise training-matched transformer.
| Comments: | 51 pages, 21 figures. Added exact-recurrence validation experiment, and more detailed training and inference costs; clarified architectural-capacity evaluation and corrected the flat-depth overthinking analysis. Headline results unchanged |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.00206 [cs.LG] |
| (or arXiv:2605.00206v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.00206 arXiv-issued DOI via DataCite |
Submission history
From: Thea Aviss [view email]
[v1]
Thu, 30 Apr 2026 20:30:28 UTC (5,472 KB)
[v2]
Wed, 7 Oct 2026 22:13:04 UTC (5,476 KB)
来源:arXiv:cs.CL · arxiv.org