arXiv:cs.AI· Jiangong Xiao, Zhe Sun, Kanzhong Yao, Yuanbo Bi, Haofei Zhao, Ruixuan Hu, Guan Huang, Xuelong Li·· 3 小时前
MixFormerV2 跨介质恢复策略 CMRP:水空界面机器人视觉跟踪的连续真值构建
Continuous Ground-Truth Construction and a Recovery Policy for Air--Water Robotic Tracking
AI 导读
针对跨水空界面视觉跟踪中飞溅、气泡与折射干扰,研究者构建了同步相机帧与动捕位姿的标注流程,产出 22,346 帧的评测专用跨介质测试集。提出的跨介质恢复策略 CMRP 以置信度触发模板选择,无需重训视觉主干即可为 MixFormerV2 提供初始模板、窗口最优预触发模板和卡尔曼引导裁剪。
正文
Abstract:Visual tracking across the air-water interface is challenged by splashes, bubbles, refraction, reflections, and abrupt appearance changes that can temporarily invalidate observations. This setting poses two coupled difficulties: first, for evaluation, image-only annotation cannot reliably describe the target's physical location during visual blindness; second, for online tracking, corrupted observations can contaminate motion estimates and appearance templates. We address the first difficulty with a construction pipeline that synchronizes camera frames with motion-capture poses, projects known target geometry, corrects underwater projection with a medium-gated residual, and subjects the annotations to manual review. This yields an evaluation-only cross-medium test set of 22,346 frames. We further introduce a Cross-Medium Recovery Policy (CMRP) centered on confidence-triggered template selection. It supplies MixFormerV2 with the fixed initial template, a window-best pre-trigger template, and a trigger-frame Kalman-guided image crop, together with their associated weights, without retraining the visual backbone. In the accuracy evaluation, CMRP achieves 49.90 Macro Success AUC, 2.95 points above MixFormerV2 Official. On selected cross-medium transition and occlusion-recovery intervals, CMRP increases MixFormerV2 tracking coverage from 47.91\% to 50.43\% relative to Official updating, while mean loss-to-recovery latency over successfully recovered videos decreases from 55.3 to 49.3 frames.
| Comments: | 8 pages, 6 figures. Submitted to IEEE International Conference on Robotics and Automation (ICRA 2027) |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11096 [cs.CV] |
| (or arXiv:2610.11096v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11096 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiangong Xiao [view email]
[v1]
Thu, 8 Oct 2026 02:11:43 UTC (13,195 KB)
来源:arXiv:cs.AI · arxiv.org