跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 3 小时前AI 评分36

OmniSeek:面向多轮音视频推理的原生工具集成框架

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

AI 导读

OmniSeek 是一个让 Omni-LLM 具备原生工具调用能力的智能体框架,可在长上下文中动态决定看或听、以及检索哪一时间窗口,将稀疏但关键的音视频片段追加回上下文以支持多轮推理。

正文

Published on Oct 1

Authors:

,

,

,

,

,

,

Abstract

We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.

View arXiv page View PDF Add to collection

Get this paper in your agent:

hf papers read 2610.02181

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.02181 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.02181 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.02181 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co