arXiv:cs.AI· Yuekai Xu, Zitao Shuai, Yuzhe Yang·· 10 小时前AI 评分44
Sensor-Language-Action 模型:OpenSLA 统一传感器、语言与动作建模
Sensor-Language-Action Models
AI 导读
研究者提出 Sensor-Language-Action(SLA)建模框架,以语言作为传感与行动的语义接口,在统一模型中连接多模态传感器观测、自然语言与动作。团队构建了覆盖超 116,000 名个体、79 种传感器模态和 60 个动作组的大规模 SLA benchmark,并推出统一模型 OpenSLA,支持分层动作预测、状态理解与动作解释。
正文
Abstract:Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while remaining grounded in the underlying sensor evidence. We build a large-scale SLA benchmark consisting of datasets that span more than 116,000 individuals, 79 sensor modalities, and 60 action groups, together with a multi-faceted captioning pipeline that aligns user context, sensor dynamics, and action evidence. Building on this framework, we present OpenSLA, a unified SLA model for hierarchical action prediction, state understanding, and action explanation. Extensive experiments on real-world tasks in clinical prediction, operating rooms, and metabolic health verify its superior performance over the state-of-the-art. OpenSLA also demonstrates intriguing capabilities including language-guided evidence grounding and zero-shot generalization to unseen actions and cohorts.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08244 [cs.AI] |
| (or arXiv:2610.08244v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08244 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zitao Shuai [view email]
[v1]
Tue, 6 Oct 2026 12:26:21 UTC (6,442 KB)
来源:arXiv:cs.AI · arxiv.org