跳到正文
arXiv:cs.LG· Dan Ben-Ami, Kobi Cohen, Chaim Baskin·· 7 小时前AI 评分46

冻结的视频语言模型是否已编码“证据就绪”信号?

Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

AI 导读

研究发现,未经修改的冻结 VideoLLM 已线性编码“证据就绪”信号,在七款模型共享的字节级一致评测中 AUROC 达 0.733-0.905。该信号以问题为条件,仅改变问题即可让 66.1% 的样本对读数反转,且在模型回答错误时 AUROC 仍保持 0.722。基于该读数构建的 Readiness Gating 策略在相同时长下最高提升准确率 +9.75 个百分点,计算开销可忽略。

正文

View PDF HTML (experimental)

Abstract:Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
ACM classes: I.2.10; I.2.6
Cite as: arXiv:2610.08560 [cs.CV]
  (or arXiv:2610.08560v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.08560

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dan Ben Ami [view email]
[v1] Tue, 6 Oct 2026 15:43:04 UTC (1,336 KB)

来源:arXiv:cs.LG · arxiv.org