跳到正文
arXiv:cs.LG· Payal Mohapatra, Haodong Yang, Yueyuan Sui, Stephen Xia, Benjamin Lundell, Qi Zhu·· 4 小时前AI 评分41

SemARC:通过自适应获取与顺序融合实现高效多模态推理

Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion

AI 导读

研究者提出 SemARC,将顺序模态聚合器(SeMA)与自适应运行时控制器(ARC)结合,在编码器运行前用已获取证据选择下一个模态,并决定何时停止。在六个多模态分类数据集和十一个基线上,SemARC 平均宏 F1 提升 3.2%、总推理 GFLOPs 降低 61.4%,端到端延迟在 GPU 和 CPU 上降低 44.0%、Android INT8 上降低 47.2%。

正文

View PDF HTML (experimental)

Abstract:Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop. However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set. We introduce SemARC, which couples a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) and uses acquired evidence to select each modality before its encoder runs. SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches. We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders. ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order. Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset's most accurate baseline. End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average. Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average. SemARC thus offers a practical path toward efficient multimodal inference across heterogeneous devices.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07466 [cs.LG]
  (or arXiv:2610.07466v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07466

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Payal Mohapatra [view email]
[v1] Mon, 5 Oct 2026 22:25:56 UTC (968 KB)

来源:arXiv:cs.LG · arxiv.org