跳到正文
arXiv:cs.CL· Tao Liu, Tao Feng, Xiangheng Li, Jinwang Song, Yifan Li, Xiaoqing Cheng, Dixuan Zhang, Siquan Li, Lin Lan, Hongying Zan, Kunli Zhang, Chao Wu·· 4 小时前

SERL-SQL:面向 Text-to-SQL 强化学习智能体的选择性事后蒸馏框架

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

AI 导读

SERL-SQL 通过选择性执行反馈强化学习框架,将教师-学生似然差转化为有界掩码权重,仅对 SQL 和工具动作 token 重加权 GRPO 优势。在 BIRD-Dev 上达到 76.56% 执行准确率,Spider-Test 上达 89.92%,其奖励选择策略接近 oracle Best-of-N 上界并持续优于一致性选择方法。

正文

Authors:Tao Liu, Tao Feng, Xiangheng Li, Jinwang Song, Yifan Li, Xiaoqing Cheng, Dixuan Zhang, Siquan Li, Lin Lan, Hongying Zan, Kunli Zhang, Chao Wu

View PDF HTML (experimental)

Abstract:Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory-level reward, which provides limited guidance for identifying the SQL decisions responsible for success or failure. We propose SERL-SQL, a selective execution-grounded reinforcement learning framework for multi-turn Text-to-SQL agents. SERL-SQL samples on-policy SQL interaction trajectories and uses a training-only teacher to re-score student actions with execution feedback. The resulting teacher--student likelihood gap is converted into bounded, masked weights that reweight GRPO advantages only on SQL and tool-action tokens. In this way, task rewards preserve the optimization direction, while execution hindsight provides localized credit assignment. Experiments on BIRD, Spider, and cross-domain benchmarks show that SERL-SQL achieves competitive performance, reaching 76.56% execution accuracy on BIRD-Dev and 89.92% on Spider-Test. Moreover, our reward-based selection strategy closely approaches the oracle Best-of-N upper bound and consistently outperforms consistency-based selection, showing that SERL-SQL produces high-quality candidates that can be reliably identified by lightweight execution-grounded rewards. Our code will be released at this https URL.
Comments: 18 pages,19 figures, Underreview
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2608.00485 [cs.CL]
  (or arXiv:2608.00485v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2608.00485

arXiv-issued DOI via DataCite

Submission history

From: Tao Liu [view email]
[v1] Sat, 1 Aug 2026 07:24:12 UTC (916 KB)
[v2] Tue, 4 Aug 2026 08:32:19 UTC (1,689 KB)
[v3] Sun, 27 Sep 2026 10:15:08 UTC (2,369 KB)
[v4] Thu, 8 Oct 2026 03:05:40 UTC (1,596 KB)

来源:arXiv:cs.CL · arxiv.org