跳到正文
arXiv:cs.CL· Leran Chen, Lingnan Kong, Zile Cai·· 3 小时前

MiniVer-V:为短视频事实核查识别最小充分证据

MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification

AI 导读

研究者提出 MiniVer-V 基准,包含 195 条短视频、三分类判定标注(supported、refuted、insufficient)和 5,510 个多模态证据单元,覆盖视觉关键帧、语音转写与网络检索外部来源。

正文

View PDF HTML (experimental)

Abstract:A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.
Comments: 33 pages, 2 figures, 20 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Multimedia (cs.MM)
ACM classes: I.2.7; H.3.3
Cite as: arXiv:2610.11233 [cs.CV]
  (or arXiv:2610.11233v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.11233

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Leran Chen [view email]
[v1] Thu, 8 Oct 2026 04:33:14 UTC (306 KB)

来源:arXiv:cs.CL · arxiv.org