arXiv:cs.AI· Mohammadreza Sediqin, Shivali Dalmia, Srinivasa Karthikeya Reddy Kovvuri, Abhishek Mukherji·· 5 小时前AI 评分54
面向 Agent 评测的 Trust Layer:在基准分数旁报告分数是否可信
A Trust Layer for Agent Evaluation
AI 导读
论文提出 Trust Layer for Agent Evaluation,一个事后附加框架,在每条基准分数旁报告该分数是否可信,核验四项性质:结果是否被基准评分逻辑支持、通过是否来自可追溯计算、完成声明是否与实际一致、结果在重复执行下是否稳定。
正文
Abstract:Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc framework that reports, beside each recorded score, whether it should be believed. It verifies four properties: whether the result is supported by the benchmark's own grading logic, whether a passing answer was earned through traceable computation, whether the agent's completion claim matches what occurred, and whether the result is stable under repeated execution. The first three use only saved artifacts; the fourth re-runs the agent. Model judgments only label evidence under majority voting; all verdicts follow deterministic rules and never modify the recorded score. Applied to five agent configurations on 108 tasks from Agents' Last Exam, every model shows passing runs with no traceable computation (at rates varying tenfold), confirmed false completion claims, and unstable results: 18-46% of tasks do not stay in one score band over five runs. Only 22.6% of recorded passes clear all four checks (95% CI 15.0-32.6, n=84). Measuring what an agent can do and verifying that it did it are different problems, and current benchmarks address only the first.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07274 [cs.AI] |
| (or arXiv:2610.07274v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07274 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mohammadreza Sediqin [view email]
[v1]
Mon, 5 Oct 2026 19:15:37 UTC (90 KB)
来源:arXiv:cs.AI · arxiv.org