跳到正文
原文
Thinking Machines· @thinkymachines · X·· 2026-08-28AI 评分45
AI 导读

为 RLVR 清洗数据并对齐奖励函数需要前期投入专业知识和精力,但结果是得到一个在复杂任务上达到 SOTA 的模型。 UIUC 和 Bridgewater 研究人员的客座文章,与我们的团队合作完成。 https://thinkingmachines.ai/news/putting-task-expertise-into-rl

正文

Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task.

Guest post by researchers at UIUC and Bridgewater, in collaboration with our team.
https://thinkingmachines.ai/news/putting-task-expertise-into-rl

引用Tinker@tinkerapi
LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark. https://thinkingmachines.ai/news/putting-task-expertise-into-rl
在 X 查看被引用的帖子

来源:Thinking Machines · x.com