跳到正文
arXiv:cs.AI· Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng·· 5 小时前AI 评分36

AdsCVR:电商跨视频推理基准与 AdSeek 主动证据获取框架

Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

AI 导读

研究者提出首个电商跨视频推理基准 AdsCVR,包含 2,483 个视频和 6,110 个问答对,覆盖六个推理维度。配套的 AdSeek 智能体框架在多轮探索中动态选择视觉与音频工具,以主动证据获取替代静态均匀采样,在 AdsCVR 测试集上达到 74.30% 准确率,较其 Qwen3-VL-8B-Instruct 主干高出 27.90 个百分点,并泛化至开放域 CrossVid 基准。

正文

View PDF HTML (experimental)

Abstract:E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. Cross- video reasoning requires models to locate fine-grained evidence among many redundant frames and integrate visual details, speech, and on-screen text. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. To address the sparse credit assignment of reinforcement learning, we develop an offline trajectory rectification mechanism that identifies reasoning errors and missing multimodal evidence in RL-generated trajectories. The corrected trajectories provide supervised fine-tuning signals that reduce biases learned during RL. This mechanism supports a rectified bootstrapping pipeline in which initial RL exposes reasoning bottlenecks, supervised fine-tuning corrects them, and a final RL stage further improves the policy. AdSeek achieves 74.30 percent accuracy on the AdsCVR test split, outperforming its Qwen3-VL-8B-Instruct backbone by 27.90 percentage points. It also generalizes to the open- domain CrossVid benchmark, demonstrating effective active evidence gathering.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03099 [cs.CV]
  (or arXiv:2610.03099v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.03099

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jinghan Zhao [view email]
[v1] Fri, 2 Oct 2026 10:19:49 UTC (11,683 KB)

来源:arXiv:cs.AI · arxiv.org