OneStreamer:统一流式视频交互中的感知、记忆与主动响应
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
OneStreamer 通过共享的主动生成过程,联合学习与查询无关的证据记录和任务响应,其 PHCM 生成带时间锚定的局部细节描述与已完成事件摘要,PSTL 仅监督 27.5% 的状态 token 即超越密集状态监督。团队还构建了超百万条记录的 OneStreamer-1M 数据集,4B 模型在全部八个流式视频理解基准上取得对比方法中的最佳结果。
Published on Oct 1
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
View arXiv page View PDF Project page GitHub 6 Add to collection
Get this paper in your agent:
hf papers read 2610.01762
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 1
Image-Text-to-Text • 4B • Updated 22 minutes ago • 2MCG-NJU/OneStreamer-4B
Datasets citing this paper 1
MCG-NJU/OneStreamer-1M
Preview • Updated 19 minutes ago • 59 • 2
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2610.01762 in a Space README.md to link it from this page.
Collections including this paper 1
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co