跳到正文
arXiv:cs.AI· Rana Muhammad Usman·· 4 小时前AI 评分49

反事实证据审计可预测 LLM 智能体对排序上下文的敏感性

Counterfactual Evidence Audits Predict LLM-Agent Susceptibility to Ranked Context

AI 导读

研究者提出"反事实证据审计"方法:让 LLM 智能体面对两组镜像的五文档证据集,用六项下游决策差异预测其对 45 文档上下文的响应。在 18 个留出模型-任务单元中,五文档效应预测全上下文效应的 Spearman rho=.855(p<.001),平均绝对预测误差较零效应基线降低 62%,并还原 13 项实质效应中 12 项的方向。

正文

View PDF HTML (experimental)

Abstract:LLM agents increasingly decide from evidence assembled by upstream systems: retrievers choose documents, recommenders choose posts, and memory systems choose prior events. Existing evaluations usually hold this evidence fixed, missing failures in which individually ordinary items form a systematically one-sided context. We introduce a counterfactual evidence audit: expose an agent to two mirrored sets of five documents, measure the difference in six downstream decisions, and use that contrast to predict its response to disjoint 45-document contexts. The protocol was frozen before testing three held-out open-weight model families. Across 18 held-out model-task cells, five-document effects predict full-context effects with Spearman rho=.855 (p<.001), reduce mean absolute prediction error by 62% relative to a zero-effect predictor, and recover the direction of 12 of 13 material effects. A reviewer-requested post-hoc task-mean baseline is also substantially weaker (MAE .369 versus .167). Matched controls show that selecting one-sided ordinary items, rather than merely reordering identical items, causes the shift in a susceptible model. Across seven open-weight families, susceptibility transfers from an interactive feed to a static RAG dossier (rho=.750, exact p=.033), while a provenance warning does not reliably mitigate it. A separate study of three deployed Codex agent tiers finds strong audit-to-full ranking (rho=.951, p<.001) but no individually significant full-context effect after correction. Within this single synthetic remote-work domain, the result supports a domain-specific triage procedure, not a universal steering claim: evidence selection must be evaluated as part of the composed agent system.
Comments: 19 pages, 1 figure. Accepted at FLMSec 2026 (NeurIPS 2026 Workshop). Substantially revised after peer review with new preregistered audits, matched controls, held-out validation, RAG transfer, and Codex boundary tests
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
ACM classes: I.2.7; I.2.11; K.6.5
Cite as: arXiv:2606.00914 [cs.AI]
  (or arXiv:2606.00914v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2606.00914

arXiv-issued DOI via DataCite

Submission history

From: Rana Usman Mr [view email]
[v1] Sat, 30 May 2026 22:43:23 UTC (311 KB)
[v2] Fri, 2 Oct 2026 15:18:46 UTC (6,740 KB)

来源:arXiv:cs.AI · arxiv.org