视频中的多模态矛盾/犹豫识别:面向个性化数字健康干预
Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
研究探索用深度学习在视频中识别矛盾与犹豫(A/H)情绪,以支撑个性化数字健康干预。工作覆盖监督学习、无监督域适应个性化以及基于 LLM 的零样本推理三种设置,并在新发布的 BAH 视频数据集上实验。结果显示性能有限,需更适合的多模态模型及更好的时空与多模态融合方法来利用模态内/间的冲突。
Authors:Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger
Abstract:Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has recently gained considerable attention. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities.
| Comments: | 11 pages, 4 figures, ACII 2026. arXiv admin note: substantial text overlap with arXiv:2505.19328 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.11730 [cs.CV] |
| (or arXiv:2604.11730v5 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2604.11730 arXiv-issued DOI via DataCite |
Submission history
From: Soufiane Belharbi [view email]
[v1]
Mon, 13 Apr 2026 17:05:38 UTC (807 KB)
[v2]
Tue, 14 Apr 2026 11:00:18 UTC (807 KB)
[v3]
Mon, 4 May 2026 17:25:53 UTC (807 KB)
[v4]
Sun, 5 Jul 2026 12:30:14 UTC (1,052 KB)
[v5]
Fri, 2 Oct 2026 08:08:15 UTC (1,052 KB)
来源:arXiv:cs.LG · arxiv.org