跳到正文
原文
Perplexity· @perplexity_ai · X·· 8 天前AI 评分40
AI 导读

新研究:我们通过提示引导的自蒸馏,对 Computer 模型进行后训练,让它从自身错误中学习。 在一次线上 A/B 测试中,较晚训练的 checkpoint 相比早期 checkpoint 将工具调用失败率降低了 21.2%。

正文

New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation.

In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint.

来源:Perplexity · x.com