跳到正文

#具身智能

今日 5 条
今天10月1日周四
  1. HuggingFace Daily Papers(社区热门论文)37

    GGSD:通过游戏自对弈发现人类可直接操控的智能体技能

    研究者提出 Game-Guided Skill Discovery(GGSD)框架,利用游戏中的自对弈发现人类可直接操控的运动技能,由高层策略从少量离散技能中选择、低层策略学习对应行为。在 Ant、Franka-arm 和 Unitree G1 环境中,人类可替换高层策略直接操控智能体,并通过技能组合解决 Maze、CubePush 等未见任务,无需额外训练。

  2. HuggingFace Daily Papers(社区热门论文)34

    生成式行为克隆中的多模态性研究:确定性回归为何在仿真基准上依然有效

    一项研究重新审视生成式行为克隆(BC)中的多模态假设,发现确定性动作 Transformer 在 LIBERO、Meta-World 等仿真基准上仍能匹配此前报告的 SOTA 结果,且推理延迟往往更低。作者通过 GMM 聚类构建条件模态估计器,估计结果显示所测仿真基准在采样状态下几乎都是单模态的;只有在机器人双模态操作概念验证任务上才观察到模态坍缩。

  3. HuggingFace Daily Papers(社区热门论文)49

    WorldAuditBench:用多模态智能体审计交互式 3D 世界

    研究者推出 WorldAuditBench,一个面向 3D 世界审计的基准,包含 13 个环境中的 213 个异常任务,覆盖五类异常。在固定探索预算下评测五个前沿模型,两种审计范式的成功率仅 6.6% 至 42.3%,远低于人类的 83.4%。该基准用于研究多模态智能体如何在交互式 3D 环境中耦合动作与视觉推理。

9月30日周三
  1. HuggingFace Daily Papers(社区热门论文)41

    WorldAttention:面向交互式视频世界模型的高效注意力架构

    研究者提出 WorldAttention,一种通过专用注意力内核与分层 KV cache 协同设计实现高效推理的注意力架构,用于文本条件交互式视频世界模型。其 Hybrid Sparse Attention 结合线性全局注意力与头自适应稀疏注意力,Hierarchical KV Cache 将历史 KV 对按语义索引分页存放于多级内存。

  2. HuggingFace Daily Papers(社区热门论文)44

    EVO-WAM:通过视频-动作验证进化世界动作模型

    EVO-WAM 让世界动作模型无需额外专家演示即可适应新任务,通过视觉语言模型筛选完成任务的前缀、用逆动力学模型验证视频-动作一致性,再迭代训练。在 RoboTwin 2.0 的 7 个未见任务上,Cosmos3 平均成功率从 26.9% 提升至 68.0%,DreamZero 从 28.5% 提升至 46.4%;真实世界 3 个长程复合任务中 Cosmos3 从 20.0% 提升至 76.7%。

  3. HuggingFace Daily Papers(社区热门论文)33

    JEPA 世界模型新方法 AnisoWM:各向异性表示改进规划

    研究者提出 AnisoWM 与 ΛReg,用可学习的对角协方差替换固定各向同性高斯目标,并施加固定迹与各向异性约束,预测目标、预测器架构和欧氏规划器均不变,目标仅在训练时使用。在四个视觉控制环境中,AnisoWM 的规划成功率全部优于 LeWorldModel,其潜在规划代价与任务结果也更一致。

9月28日周一
9月27日周日
  1. Thomas Wolf37

    原来你不需要 IMU 🤯

    引用Aditya Bhatt ✈️ IROS 2026@aditya_bhatt

    In my latest PhD paper, we declare WAR on sensor-maxxing. For the first time, push-resilient humanoid walking, just with joint encoders! No IMU, no F/T sensors. 🥁 Introducing Blind Dexterity 🧵 👇 w/ @OKaidanov @liu_puze @Jan_R_Peters at @DFKI @ias_tudarmstadt

9月26日周六
9月24日周四
  1. X Square Robot42

    很高兴介绍 X-Planner:面向具身智能的事件结构化任务规划——一个将语义事件作为规划单元的前端,通过 Staircase Decoding 暴露离散+隐式接口,在真实机器人上超越基线。

    引用wisdom pan@wisdompan_ai

    To address long-horizon robotic manipulation tasks, we propose X-Planner, an embodied long-horizon task planner: It decomposes high-level natural language instructions into event-level subtasks, and can also output continuous implicit Chain-of-Thought (CoT). It directly interfaces with downstream VLA or WAM models to operate robots. GitHub: https://github.com/X-Square-Robot/Xplanner Checkpoint: https://huggingface.co/x-square-robot/X-Planner-9B-0916 Benchmark: https://huggingface.co/datasets/x-square-robot/xplanner-benchmark

  2. Microsoft Research 博客(RSS)62

    微软研究:把物理 AI 推理卸载出机器人可提升任务成功率与续航

    微软研究院对移动机器人操作负载做系统性测量,发现把物理 AI 推理从机载 GPU 卸载到边缘或云端 GPU 可提升任务表现与电池续航。

    推荐理由:微软对移动操作机器人推理负载的系统性测量,给出了卸载到边缘或云端 GPU 在任务成功率与续航上的量化差异。

9月17日周四
9月14日周一
  1. X Square Robot23

    期待参加在匹兹堡举办的 Saturday Robotics × IROS 2026!🤖 我们将展示 X2Real,这是我们用于评估真实世界通用机器人策略的大规模仿真基准——涵盖 10 个能力维度的 44 个分层长时程任务。 期待分享我们的最新工作,并与推进机器人学习、仿真到真实迁移和具身 AI 的研究者和开发者交流。 📍 匹兹堡 📅 2026 年 9 月 28 日 👉🏻 https://luma.com/tzbw7n61 到时见!

    引用Junfan Zhu 朱俊帆 ✈️ IROS@junfanzhu98

    🍾🍲 Saturday Robotics x IROS 2026 — Robotics Research Night 👉🏻 https://luma.com/tzbw7n61 We’re bringing a high-signal evening of robotics research to Pittsburgh on September 28. After a full day at IROS, we’ll bring together researchers, engineers, founders, students, and investors for technical discussions, networking, and a series of ~10-minute lightning talks. Tentative preview of the current lineup: 🤖 1. PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball Gary Yang @lzyang2000 (@Caltech) Perception-aware reinforcement learning + Control Barrier Functions for whole-body humanoid safety. Demonstrated on a Unitree G1, with 19/20 successful dodges and zero falls in real-world experiments. 🧠 2. How In-Context Learning Is Reshaping Robot Learning Data at Scale AaronLi (@RhodaAI) Exploring how in-context learning can change the way we think about robot learning data, scaling, and generalization. 🧪 3. X2Real: An eXtensive Simulation Benchmark for Real-World Generalist Policies Liangwang Ruan (@XSquareRobot) A new simulation benchmark built around faithfulness, diversity, and fairness, with 44 hierarchical long-horizon tasks across 10 capability dimensions and a reported 0.84 simulation-to-real correlation. 🦾 4. Rethinking Generalist Robotic Manipulation: Architecture, Data and Inference for Real-World Deployment Peiyan Li (Chinese Academy of Sciences, @CAS__Science) 3D VLA architectures, memory augmentation, ego/UMI human priors, large-scale robot pretraining, and inference-time contextual learning for deployable generalist manipulation. 🎯 5. HiRE: Hindsight Reward Editing for Policy Finetuning Haoyi Niu @t641769919 (@UCBerkeley) Accepted at CoRL 2026. A training-free approach to reward editing that uses successful and failed trajectories to identify “trap states” and provide denser, control-aware feedback for RL. 🔥 6. Lightning Talk — Open Slot We’re opening one additional slot for a technically deep research talk, new project, frontier paper, demo, open problem, or startup technical insight. 10 minutes. A few slides. One sharp technical idea. No fluff. Topics include World Models, Physical AI, Humanoids, VLAs, Robot Foundation Models, Manipulation, RL, Simulation & Sim-to-Real, Spatial Intelligence, Computer Vision, and Embodied AI. 📍 Pittsburgh 📅 September 28, 2026 🕠 5:30–9:30 PM 🍾 Networking + Technical Talks + Research Discussion 📩 junfanzhu98@gmail.com See you in Pittsburgh. 🤖 #IROS2026 #Robotics #PhysicalAI #RobotLearning #WorldModels #HumanoidRobotics #VLA #EmbodiedAI #RobotFoundationModels

9月3日周四
  1. X Square Robot35

    自变量 X Square 发布 TwinDEX,一套从人类指尖到机器人指尖的高保真灵巧操作框架,通过可穿戴外骨骼采集人手技能,并以匹配的硬件一致性在机器人手上复现。TwinDEX 采用三指九自由度架构,走"减法"路线保留拇指的核心灵巧作用,而非在夹爪上叠加手指。其核心主张是:数据生成阶段引入的系统性误差无法靠扩大数据集消除,保真度决定了学习性能的上限。

9月2日周三
  1. MIT News(RSS)49

    MIT 与 Motional 提出 CW-Net,帮人类预判自动驾驶汽车何时出错

    MIT 与自动驾驶公司 Motional 提出 Concept-Wrapper Network(CW-Net),将自动驾驶深度学习规划器的内部推理翻译为"接近停驶车辆""靠近骑行者"等可理解概念,且不改变原有驾驶性能。该模块用 1.3 亿个自动驾驶场景样本训练,在私人测试跑道的实车测试中帮助安全员更准确预判车辆行为,大规模模拟实验也得到类似结果,相关研究已发表于 Nature。

8月22日周六
8月11日周二
  1. MIT News(RSS)44

    MIT CSAIL 与清华提出 GeoPT:让 AI 模型学会物理,仿真提速 2 倍、数据省 60%

    MIT CSAIL 与清华研究人员提出预训练方法 GeoPT,通过 130 万条"合成动力学"样本让仿真模型学习物理规律,达到峰值性能的速度比领先模型快 2 倍,所需数据最多减少 60%。在工业基准上,GeoPT 在速度、精度和效率上超越 SOTA 模型,模拟船体受风浪时用 60% 更少标注数据、达到峰值精度快 4 倍,并能在数秒内完成超 1 亿网格点的高保真仿真。

7月27日周一
7月15日周三
7月9日周四
  1. MIT News(RSS)48

    MIT 推出 FloatForm:小型机器人船群自组装水上漂浮结构

    MIT 团队发布 FloatForm 系统,由 21 厘米见方的方形机器人船组成,可自主组装成刚性结构、拆解并重组,单次运行耗时 4 至 8 分钟。该系统仅需轻量中央规划器分配最终位置,其余导航与避碰由机器人自身完成,仿真显示可平滑扩展至 64 艘的集群。相关论文已发表于 Nature Communications。

7月1日周三
  1. Jim Fan49

    ENPIRE -> ASPIRE,我们 Physical AutoResearch 系列的第二项工作。我们正在构建机器人自我改进的组件,一次一个 /skill。

    引用Jim Fan@DrJimFan

    Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an evolutionary search over control programs, and distill the best know-how into an ever-expanding library. ASPIRE is a new type of continual learning: "training" is skill refinement instead of gradient descent. "Trained model" is a repo of sensorimotor skills instead of floating weights. “Distributed training” is a panel of agents each practicing a different skill instead of sharded minibatches. Here's the beauty: ASPIRE gives the tired terms "sim2real transfer" and "cross-embodiment transfer" a whole new meaning. Bridging the sim-to-real gap is notoriously brutal. An end-to-end policy has to swallow both the visual shift (sim looks toyish next to a real camera) and the subtle contact physics it never quite gets right. ASPIRE sidesteps the mess, because it doesn't ship pixels or weights across the gap, but ships the know-how. The robot still has to practice in the real world, not zero-shot, but it gets there way faster because it isn't rediscovering the strategy from scratch. Same for going single-arm to bimanual hardware, which usually requires new data and retraining from zero. ASPIRE achieves up to ~10x cut in "transfer learning” tokens (yes, tokens are the new unit of *training* compute ;) Check out our gallery of 150+ tasks and 90+ skills the robots taught themselves, all on the website! Kind of wild that we can ship the "learned weights" as an HTML page rather than a GGUF. We'll open-source the full stack so your own robot library starts compounding from ours! Deep dive in thread:

  2. Jim Fan47

    Jim Fan 团队发布 ASPIRE,让编码智能体观察仿真与真实机器人的多模态感知轨迹,通过进化搜索控制程序并将最优经验蒸馏进持续扩张的技能库。ASPIRE 把"训练"重新定义为技能精炼而非梯度下降,跨仿真到真实与跨本体迁移时只传技能知识而非像素或权重,迁移学习 token 消耗最多降低约 10x。团队展示了 150+ 任务、90+ 技能,并将开源全栈。

6月23日周二
  1. MIT News(RSS)43

    MIT 新型芯片 Gleanmer 可助微型机器人实时构建 3D 地图

    MIT 研究人员推出名为 Gleanmer 的系统级芯片,能让小型自主机器人仅用约 6 毫瓦功耗实时构建详细 3D 环境地图,功耗仅相当于一颗 LED。该芯片结合 GMMap 算法与专用硬件,用高斯椭球体替代传统体素表示障碍物与自由空间,单次处理深度图像即可生成高斯分布,无需存储整张图像。该成果已在 IEEE VLSI Symposium 上展示,也有望用于轻量 AR 头显。

10月13日周一