跳到正文

#论文/研究

今日 6 条
9月16日周三
  1. Apple Machine Learning Research(RSS)38

    Glyph:面向企业数据目录列描述与敏感本体标注的多策略智能体系统

    Apple 研究团队提出 Glyph,一个将列描述生成与列类型标注建模为有状态图编排的多智能体 LLM 生产系统。其 Descriptor 通过推理-行动工具循环从企业 GitHub 按需检索管道源码来支撑生成,Tagger 并行运行描述、业务线正则与元数据三种策略,并用 RRF 融合排序结果,从 275 叶节点的数据分类本体中打标。

  2. Google Research29

    Google Research 提出 Retrieve-for-Train:用 RL 编译扩散模型绕过推理瓶颈,加速复杂 AI 搜索

    Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现推理时单次非自回归的查询扇出。该方法在 Gemma3-4B 和 Qwen3-4B 上微调扇出语言模型,用集合级属性奖励评估整组结果,无需人工标注,也无需推理时的 CoT 思考 token。

9月14日周一
  1. X Square Robot23

    期待参加在匹兹堡举办的 Saturday Robotics × IROS 2026!🤖 我们将展示 X2Real,这是我们用于评估真实世界通用机器人策略的大规模仿真基准——涵盖 10 个能力维度的 44 个分层长时程任务。 期待分享我们的最新工作,并与推进机器人学习、仿真到真实迁移和具身 AI 的研究者和开发者交流。 📍 匹兹堡 📅 2026 年 9 月 28 日 👉🏻 https://luma.com/tzbw7n61 到时见!

    引用Junfan Zhu 朱俊帆 ✈️ IROS@junfanzhu98

    🍾🍲 Saturday Robotics x IROS 2026 — Robotics Research Night 👉🏻 https://luma.com/tzbw7n61 We’re bringing a high-signal evening of robotics research to Pittsburgh on September 28. After a full day at IROS, we’ll bring together researchers, engineers, founders, students, and investors for technical discussions, networking, and a series of ~10-minute lightning talks. Tentative preview of the current lineup: 🤖 1. PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball Gary Yang @lzyang2000 (@Caltech) Perception-aware reinforcement learning + Control Barrier Functions for whole-body humanoid safety. Demonstrated on a Unitree G1, with 19/20 successful dodges and zero falls in real-world experiments. 🧠 2. How In-Context Learning Is Reshaping Robot Learning Data at Scale AaronLi (@RhodaAI) Exploring how in-context learning can change the way we think about robot learning data, scaling, and generalization. 🧪 3. X2Real: An eXtensive Simulation Benchmark for Real-World Generalist Policies Liangwang Ruan (@XSquareRobot) A new simulation benchmark built around faithfulness, diversity, and fairness, with 44 hierarchical long-horizon tasks across 10 capability dimensions and a reported 0.84 simulation-to-real correlation. 🦾 4. Rethinking Generalist Robotic Manipulation: Architecture, Data and Inference for Real-World Deployment Peiyan Li (Chinese Academy of Sciences, @CAS__Science) 3D VLA architectures, memory augmentation, ego/UMI human priors, large-scale robot pretraining, and inference-time contextual learning for deployable generalist manipulation. 🎯 5. HiRE: Hindsight Reward Editing for Policy Finetuning Haoyi Niu @t641769919 (@UCBerkeley) Accepted at CoRL 2026. A training-free approach to reward editing that uses successful and failed trajectories to identify “trap states” and provide denser, control-aware feedback for RL. 🔥 6. Lightning Talk — Open Slot We’re opening one additional slot for a technically deep research talk, new project, frontier paper, demo, open problem, or startup technical insight. 10 minutes. A few slides. One sharp technical idea. No fluff. Topics include World Models, Physical AI, Humanoids, VLAs, Robot Foundation Models, Manipulation, RL, Simulation & Sim-to-Real, Spatial Intelligence, Computer Vision, and Embodied AI. 📍 Pittsburgh 📅 September 28, 2026 🕠 5:30–9:30 PM 🍾 Networking + Technical Talks + Research Discussion 📩 junfanzhu98@gmail.com See you in Pittsburgh. 🤖 #IROS2026 #Robotics #PhysicalAI #RobotLearning #WorldModels #HumanoidRobotics #VLA #EmbodiedAI #RobotFoundationModels

9月11日周五
9月10日周四
9月9日周三
  1. Mark Chen60

    OpenAI 宣布给出 Navier-Stokes 千禧年大奖难题的一个证明,该问题关于三维光滑流体运动的描述是否会失稳,已悬置约 90 年。证明由一组智能体使用一个能力显著强于 GPT-6 Astra 的 OpenAI 下一代模型产出。OpenAI 首席研究官 Mark Chen 转发并表示,许多同事因相信 AI 是解决各领域大挑战的最快途径而离开原领域,看到第一个挑战落下令他感到不真实。

    引用OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

9月8日周二
  1. vLLM 官方博客(RSS)65

    vLLM 联合 AgentX 优化真实智能体推理服务,成本较 Opus 5 API 最多低 106×

    vLLM 团队发布针对智能体负载的全栈优化方案,覆盖 KV 缓存管理、并行策略与 P/D 分离配比,并在 SemiAnalysis AgentX 公开基准上验证。

    推荐理由:原文给出 vLLM 针对 AgentX 基准的全栈优化路径与可复现结果,读者可以借此理解智能体负载的服务优化思路。

9月4日周五
  1. Google Research62

    Google 与 HHMI Janelia 发布完整雄性果蝇脑连接组图谱

    Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。

    推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。

9月3日周四
  1. X Square Robot35

    自变量 X Square 发布 TwinDEX,一套从人类指尖到机器人指尖的高保真灵巧操作框架,通过可穿戴外骨骼采集人手技能,并以匹配的硬件一致性在机器人手上复现。TwinDEX 采用三指九自由度架构,走"减法"路线保留拇指的核心灵巧作用,而非在夹爪上叠加手指。其核心主张是:数据生成阶段引入的系统性误差无法靠扩大数据集消除,保真度决定了学习性能的上限。

9月2日周三
  1. MIT News(RSS)49

    MIT 与 Motional 提出 CW-Net,帮人类预判自动驾驶汽车何时出错

    MIT 与自动驾驶公司 Motional 提出 Concept-Wrapper Network(CW-Net),将自动驾驶深度学习规划器的内部推理翻译为"接近停驶车辆""靠近骑行者"等可理解概念,且不改变原有驾驶性能。该模块用 1.3 亿个自动驾驶场景样本训练,在私人测试跑道的实车测试中帮助安全员更准确预判车辆行为,大规模模拟实验也得到类似结果,相关研究已发表于 Nature。

  2. Hugging Face:Blog(RSS)49

    BenchMIRT:LLM 基准测试究竟在测什么?

    研究者提出 BenchMIRT,一种在单条提示词层面审计 LLM 基准测试的多维 IRT 方法,基于 100 个 LLM 在 16 个基准、超 34K 道题上的结果训练,在未被告知各基准测什么的情况下自行恢复出安全与通用推理两个主导维度。分析显示 BBQ 更贴近通用推理而非安全,WMDP 分数与通用推理关联更强且推理越强分数越低,HarmBench 的版权类问题也更接近通用推理。

8月28日周五
  1. Thinking Machines45

    为 RLVR 清洗数据并对齐奖励函数需要前期投入专业知识和精力,但结果是得到一个在复杂任务上达到 SOTA 的模型。 UIUC 和 Bridgewater 研究人员的客座文章,与我们的团队合作完成。 https://thinkingmachines.ai/news/putting-task-expertise-into-rl

    引用Tinker@tinkerapi

    LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

  2. MIT News(RSS)35

    MIT 团队提出 PottsMPNN:不再以还原天然序列衡量蛋白质设计

    MIT 生物学系团队开发出机器学习框架 PottsMPNN,通过引入支配蛋白质结构与稳定性的物理原理并建模氨基酸两两相互作用,提升序列生成与突变稳定性预测能力,成果发表于 PNAS。研究者指出,长期以来以能否复现进化选出的天然序列作为成功标准并非蛋白质设计的最佳指标,PottsMPNN 在减少对天然序列依赖的同时,结构兼容性与能量预测反而改善,可设计出序列不类似任何天然蛋白的结构可行蛋白。

8月27日周四
  1. Saining Xie43

    很高兴看到 RAE 扩展到视频!

    引用Minghui Guo@MinghuiGuo77

    🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE

8月26日周三
  1. MIT News(RSS)43

    MIT 提出 CrysVCD 框架:让 AI 生成的材料更稳定、更贴近真实应用

    MIT 研究人员提出 CrysVCD 框架,在材料生成前用语言模型约束价电子规则,使常用材料模型在近 70% 的生成结果中达到高晶格动力学稳定性。该方法比生成后再筛选的方案效率高一个数量级,微调后生成的晶体材料机械稳定性达 68%、亚稳性达 85%,并可定向生成高热导率、高介电常数等材料。

8月25日周二
8月22日周六
8月21日周五
  1. Microsoft Research 博客(RSS)41

    微软 Skala 1.1 发布:训练数据增 2.5 倍,已接入 CP2K 并推进 Psi4、FHI-aims、ORCA、VASP 集成

    微软研究院发布深度学习交换关联泛函 Skala 1.1,训练数据比首个公开版本多 2.5 倍,在 GMTKN55 的 55 个类别中 32 项排名第一,加权平均误差 2.8 kcal/mol,精度超过当前领先的全局(范围分离)杂化泛函而保持半局域泛函的计算成本。

8月19日周三
8月17日周一
8月13日周四
  1. Microsoft Research 博客(RSS)37

    MindTopo 揭示多模态大模型的空间推理能力短板

    微软研究院推出 MindTopo 基准,从连续性、分离、顺序、包围、绳结五类拓扑关系评估多模态大模型的推理与规划能力。测试显示,模型在静态图像识别上表现明显优于交互式规划任务,失败多发生在规划阶段而非感知阶段,且整体远低于人类水平。图像与视频生成仅在单帧关系可见时偶有帮助,跨多步动作时难以维持拓扑约束。