Salesforce AI Research 提出 Critical-State RL,用于多轮工具调用的强化学习训练。该方法通过嵌套采样分离当前动作带来的奖励变化与下游噪声,只对改变结果的那一次调用做上下文赌博机更新,而非把奖励摊到整条轨迹。在 BFCL v4 缺失函数任务上,训练被选中的那一轮带来约 14 个百分点的提升,训练其他候选轮则准确率持平或下降。
Great paper from Salesforce AI Research on RL for multi-turn tool use.
The finding is that you want to train the one call where the action changes the outcome, instead of spreading reward across the whole trajectory.
(bookmark it)
When reward depends on later turns, much of its variation comes from what happens downstream.
Critical-State RL uses nested sampling to separate the reward variation caused by the current action from that noise, then trains only the selected call with contextual-bandit updates.
On BFCL v4 missing-function tasks, training the selected turn adds about 14 points. Training the other candidate turn leaves accuracy flat or lower.
Paper: https://academy.dair.ai/papers/critical-state-rl-diagnosing-trainable-states-for-multi-turn-tool-use-2609.24985
来源:DAIR.AI · x.com