arXiv:cs.CL· Kerui Huang, Shuhan Liu, Xing Hu, Tongtong Xu, Lingfeng Bao, Xin Xia·· 4 小时前AI 评分43
SEER:面向推理模型的自增强思维链压缩框架
SEER: Self-Enhancing Chain-of-Thought Compression for Reasoning Models
AI 导读
SEER 是一种无需外部压缩工具的自增强 CoT 压缩框架,在 DeepSeek-R1-Distill-Qwen-7B 上平均缩短思维链长度 34.6%,同时提升任务表现并减少截断与推理循环。
正文
Abstract:Chain-of-Thought (CoT) prompting can substantially improve the reasoning ability of large language models (LLMs), but it often comes with high inference cost due to long and poorly controlled reasoning traces. This overhead is particularly problematic in software engineering tasks (e.g., code generation), where both latency and output reliability matter. To better understand this trade-off, we conduct an empirical study on widely used code generation benchmarks and observe that many modern reasoning models produce excessively verbose CoTs (often thousands of tokens), which frequently leads to truncation and unstable generation. Using a strict n-gram repetition detector, we find that most observed truncations are associated with degenerate looping behaviors. In addition, a HumanEval/129 case study shows that failed generations can be longer than successful ones, suggesting limited returns from overlong reasoning. Motivated by these findings, we propose SEER (Self-Enhancing Efficient Reasoning), a self-enhancing framework for adaptive CoT compression. SEER improves the conciseness of reasoning while preserving output quality, without relying on external compression tools. SEER refines self-generated CoT data via Best-of-N sampling to suppress looping and redundant traces, then applies a lightweight, data-driven filter to encourage concise yet correct reasoning. It then fine-tunes the model on the filtered data to internalize concise reasoning behaviors. Across four software engineering benchmarks on the evaluated DeepSeek-R1-Distill-Qwen-7B backbone, SEER reduces CoT length by 34.6% on average while improving task performance, with reduced truncation and fewer reasoning loops.
| Comments: | 23 pages. Published in ISSTA 2026 |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2509.14093 [cs.SE] |
| (or arXiv:2509.14093v3 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2509.14093 arXiv-issued DOI via DataCite |
|
| Journal reference: | Proc. ACM Softw. Eng. 3, ISSTA, Article ISSTA029 (October 2026), 23 pages |
| Related DOI: | https://doi.org/10.1145/3832120
DOI(s) linking to related resources |
Submission history
From: Kerui Huang [view email]
[v1]
Wed, 17 Sep 2025 15:33:44 UTC (261 KB)
[v2]
Tue, 10 Mar 2026 10:11:06 UTC (431 KB)
[v3]
Wed, 7 Oct 2026 15:17:11 UTC (429 KB)
来源:arXiv:cs.CL · arxiv.org