跳到正文
arXiv:cs.AI· Tae-Eun Song·· 6 小时前AI 评分45

跨模型审查能否提升 LLM 验证效果?一项受控实验研究

When Does a Second Model Help? Cross-Model Review in LLM Verification

AI 导读

一项受控实验用 30 个工件、150 个植入错误和 900 次审查会话,测试不同模型对 LLM 输出的二次审查是否更有效。结果显示顶级跨模型审查者与同模型新会话审查的 F1 无显著差异,但两者发现的错误仅部分重叠(Jaccard 41.2%);一次同模型加一次跨模型审查命中 56.7% 植入错误,优于两次同模型审查的 42.7%(Holm 校正 p=.006)。

正文

View PDF HTML (experimental)

Abstract:Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CCR plus one cross-model review matches more planted errors than two CCR reviews (56.7% vs. 42.7%; Holm-adjusted p=.006), but not significantly more than two reviews by the top-tier cross-model reviewer, so model difference and reviewer capability are not separated. A lightweight cross-model reviewer scores no higher than same-model review. Withholding requirements from the reviewer raises F1 for the two lower tiers but not the top tier, in untested point estimates whose pattern depends on how failed sessions are scored. Before analysis we audited all session records, excluding one baseline run of uncertain provenance and 14 failed calls; results with all sessions are also reported. A partial check on public detector outputs from another benchmark neither replicates nor contradicts the main comparison. Records, artifacts, and scripts are available from the author on request.
Comments: 16 pages, 2 figures, 7 tables. Follow-up to arXiv:2603.12123 and arXiv:2603.21454. v2: corrects two condition labels in Table 1 (CCR sees the artifact only; SA runs in a new session) and dependent interpretations; adds review prompts, a TP/FP breakdown by severity, and limitations; states how each reviewer was run; softens case studies. Numbers unchanged except removed B5 percentages
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
Cite as: arXiv:2610.01471 [cs.CL]
  (or arXiv:2610.01471v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01471

arXiv-issued DOI via DataCite

Submission history

From: Tae-Eun Song [view email]
[v1] Thu, 1 Oct 2026 11:12:36 UTC (225 KB)
[v2] Tue, 6 Oct 2026 01:50:15 UTC (227 KB)

来源:arXiv:cs.AI · arxiv.org