跳到正文
arXiv:cs.LG· JiaCheng Ge, SiYu Zhang, ShengJie Li, XinTong Yang·· 4 小时前AI 评分29

EVFormer:融合自我中心视觉与 sEMG 的双向注意力双手姿态估计模型

EVFormer: An Egocentric Vision-EMG Bidirectional Attention Model for Bimanual Hand Pose Estimation

AI 导读

EVFormer 是一种融合当前 RGB 帧与过去 200 ms 双侧手腕 sEMG 的多模态框架,用于估计 44 个手指与手腕关节角度,通过顺序双向交叉注意力和特征级门控融合实现跨模态信息交互。在 296 个测试样本上,其平均绝对误差为 11.482 度,较纯视觉模型和晚期融合分别降低 13.20% 和 14.23%,并在五类手势中的四类取得最低误差。

正文

View PDF

Abstract:Egocentric bimanual hand pose estimation is important for virtual interaction, wearable control, and rehabilitation, but visual observations are often degraded by self-occlusion, hand-hand contact, and object manipulation. We propose EVFormer, a multimodal framework that combines the current RGB frame with the preceding 200 ms of bilateral wrist surface electromyography (sEMG) to estimate 44 finger and wrist joint angles. EVFormer separately encodes visual spatial features and sEMG temporal features, enables cross-modal information exchange through sequential bidirectional cross-attention, and integrates the two modalities using feature-wise gated fusion. We evaluate EVFormer in a single-participant feasibility study using one synchronized public EgoEMG recording with chronologically separated training, validation, and test splits. On 296 test samples, EVFormer achieves a mean absolute error of 11.482 degrees, compared with 13.228-13.610 degrees for vision-only, sEMG-only, late-fusion, and training-mean baselines. This corresponds to relative error reductions of 13.20% compared with the vision-only model and 14.23% compared with late fusion. EVFormer also achieves the lowest error in four of the five evaluated gesture classes. These results provide preliminary evidence that feature-level interaction between egocentric vision and sEMG can improve bimanual hand pose estimation. Further evaluation across participants, recording sessions, sensor placements, and real-world interaction conditions is required to establish the generalizability of the approach.
Subjects: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
Cite as: arXiv:2610.06970 [cs.LG]
  (or arXiv:2610.06970v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.06970

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: JiaCheng Ge [view email]
[v1] Sat, 3 Oct 2026 19:03:32 UTC (9,758 KB)

来源:arXiv:cs.LG · arxiv.org