跳到正文
arXiv:cs.CL· Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)·· 6 小时前AI 评分27

Jev 作为金融证据判定器的压力测试:同数引用替换评估

Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

AI 导读

研究以 Jev 作为 GPT-4.1-mini 计算轨迹的来源支撑验证器,用"同数引用替换"方法在固定操作数与算术的前提下移动引用位置,检验概率性证据验证在数字匹配之外的增量价值。

正文

Authors:Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)

View PDF HTML (experimental)

Abstract:Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.
Comments: counterfactual citation perturbation, evidence attribution verification, financial document question answering, Jev, LLM-as-a-judge, probabilistic source verification, tabular numerical reasoning
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.08675 [cs.CL]
  (or arXiv:2610.08675v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08675

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinyu Qiu [view email]
[v1] Tue, 6 Oct 2026 16:57:47 UTC (90 KB)

来源:arXiv:cs.CL · arxiv.org