跳到正文
arXiv:cs.LG· Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo·· 4 小时前AI 评分49

ReLoop:面向可靠 LLM 优化的结构化建模与行为验证

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

AI 导读

针对 LLM 生成优化代码时"可求解但建模语义错误"的静默失败问题,研究提出 ReLoop,结合四阶段推理链的结构化生成与基于求解器参数扰动的行为验证两种机制,在组合问题上可行性-正确性差距可达 90 个百分点。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes and returns a solver-feasible solution while encoding a semantically incorrect formulation. On compositional problems, the resulting feasibility-correctness gap reaches 90 percentage points. We introduce ReLoop, which combines two mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify) to reduce formulation errors during generation. Behavioral verification detects the errors that remain by testing whether the formulation responds correctly to solver-based parameter perturbation, a signal that comes from the solver rather than from LLM self-review and requires no ground truth. The two mechanisms address different error structures: structured generation gives the largest gain on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), and behavioral verification gives its largest gain on localized defects (+4.4pp on MAMO-ComplexLP). With diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6, and relative to direct generation it raises or preserves every reported metric of the three chat-tuned foundation models on all three benchmarks. For the narrowly fine-tuned SFT model we test, the chain-of-thought prompt conflicts with its learned output format and lowers its accuracy on MAMO-ComplexLP; we document and analyze this interaction. We release RetailOpt-190, 190 compositional retail optimization scenarios in which several constraints interact.
Comments: Code and benchmark: this https URL
Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Optimization and Control (math.OC)
Cite as: arXiv:2602.15983 [cs.SE]
  (or arXiv:2602.15983v5 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2602.15983

arXiv-issued DOI via DataCite

Journal reference: NeurIPS 2026

Submission history

From: Junbo Jacob Lian [view email]
[v1] Tue, 17 Feb 2026 20:20:33 UTC (192 KB)
[v2] Wed, 29 Apr 2026 13:39:41 UTC (198 KB)
[v3] Sun, 16 Aug 2026 14:07:55 UTC (202 KB)
[v4] Sat, 26 Sep 2026 12:14:35 UTC (205 KB)
[v5] Tue, 6 Oct 2026 03:07:54 UTC (205 KB)

来源:arXiv:cs.LG · arxiv.org