arXiv:cs.LG(机器学习,全量分类)· Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao·· 14 小时前AI 评分48
Actions with Receipts:为可回放工具智能体审计联合绑定声明、证据与执行
Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing
AI 导读
研究者提出 claim-anchored execution contract,将智能体输出的声明、其精确来源片段、产生该声明的有序执行前缀,以及执行时观察到的来源版本与访问状态联合绑定,并用确定性完整性验证器在语义标注前重建这些绑定。
正文
Abstract:Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being transplanted across claims, actions, runs, or source versions. We introduce a claim-anchored execution contract that jointly binds the emitted claim, its exact source span, the ordered execution prefix that produced it, and the source version and access state observed by that execution. Each receipt contains an emission anchor that deterministically locates the claim inside a committed answer or claim-bearing action, together with source identifiers, offsets, hashes, quotes, and a domain-separated execution commitment. A deterministic integrity verifier reconstructs these bindings before semantic or task labels are joined. We separate this integrity plane from a pluggable support plane, so structural validity is not used as a proxy for entailment. The contract exposes seven independently testable properties: claim-emission binding, source binding, ordered-execution binding, oracle separation, persisted-object replay, execution-rerun consistency, and version/access binding. Across 1,280 cross-object attacks, the joint contract detects 1,275 substitutions (0.9961). Removing a targeted property reduces its attack-detection rate to 0.0156-0.0625. On an independently adjudicated 384-pair split, the conflict-aware support guard reaches F1 0.8865 and false acceptance 0.0729; on unseen failure families, these rates are 0.8679 and 0.0938.
| Comments: | 35 pages, 8 figures |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00327 [cs.CR] |
| (or arXiv:2610.00327v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00327 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Miaobo Hu [view email]
[v1]
Tue, 29 Sep 2026 08:36:16 UTC (1,548 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org