Adobe 论文提出把标准答案写成可执行代码,每次测试运行时重新获取当前结果,再由 AI 评分器对照检查智能体回答。在 53 个测试用例上,这种方式的 AI 评分与人类专家的一致率比文字描述答案高 29%,且少用 16% token;若没有可对照的答案,AI 评分器表现比随机还差。
Standard agent tests assume the right answer never changes, but on live data it does.
So this Adobe paper checks against code that recalculates it and gets more accurate grades.
Adobe's fix is to write the right answer as code that fetches the current result each time the test runs. An AI grader then checks the agent's reply against it.
On 53 test cases, the AI grader matched human experts 29% better this way than with a written description, and used 16% fewer tokens. With no answer to check against, the AI grader did worse than random.
– arxiv. org/abs/2609.16487
Title: "Skill-based Agentic Evaluation for Real-time Data Science Tasks"
来源:Rohan Paul · x.com