Apple 提出 DSAS:动态缩放激活引导框架
Apple 研究团队提出 Dynamically Scaled Activation Steering(DSAS),一种与具体方法无关的激活引导框架,将"何时引导"与"如何引导"解耦,按层和输入自适应调节现有引导变换的强度,仅在检测到不良行为时强力干预。
Apple 研究团队提出 Dynamically Scaled Activation Steering(DSAS),一种与具体方法无关的激活引导框架,将"何时引导"与"如何引导"解耦,按层和输入自适应调节现有引导变换的强度,仅在检测到不良行为时强力干预。
Google Research 公布一项研究实验,让教师借助生成式 UI(GenUI)创建贴合课程目标的互动学习模拟,并同步开放 30 多个面向初高中 STEM 的英文学习互动样例库,均由 AI 生成、教师审核。
研究者提出 REVERSAL-BENCH,通过连续参数 ρ∈[0,1] 控制环境可逆性,并提供重置 Oracle 在五种物理引擎的八种操作场景中验证状态可恢复性。
这是一份很不错的报告,探讨了最重要的问题之一:AI 可能如何影响科学与创新?它如今已经在产生什么影响? 干得漂亮,Mihai 和团队。
I've had the most wonderful time working on this project for the last few months. This was (equally) co-led w/ @JMateosGarcia , @alexolegimas and a fantastic team.
MIT 及合作机构的团队在 Nature 发表 xvr(X-ray volume registration),可自动将术中 X 光片与患者术前 3D 扫描配准,达到亚毫米精度,比现有 AI 方法高出一个数量级。该方法基于物理仿真生成合成 X 光影像训练基础模型,约五分钟即可适配新患者,团队已在来自五家医院、覆盖成人和儿童的最大公开 2D/3D 配准数据集上完成验证,正与手术机器人公司合作推进落地。
研究者提出共享选择性持久记忆架构,为多轮工具调用的 Agentic LLM 系统保留任务规格、数据 schema、工具配置和输出约束四类可复用上下文,并丢弃会话专属推理轨迹。
Apple 研究者提出 DACA-GRPO,一种可插拔的 GRPO 训练增强方法,用于扩散语言模型强化学习。它通过 Denoising Progress Scores 提取逐 token 重要性权重(无额外前向开销),并用 Stratified Masking Likelihood 降低 mean-field 似然偏差。
Apple 研究团队提出 Glyph,一个将列描述生成与列类型标注建模为有状态图编排的多智能体 LLM 生产系统。其 Descriptor 通过推理-行动工具循环从企业 GitHub 按需检索管道源码来支撑生成,Tagger 并行运行描述、业务线正则与元数据三种策略,并用 RRF 融合排序结果,从 275 叶节点的数据分类本体中打标。
Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现推理时单次非自回归的查询扇出。该方法在 Gemma3-4B 和 Qwen3-4B 上微调扇出语言模型,用集合级属性奖励评估整组结果,无需人工标注,也无需推理时的 CoT 思考 token。
MIT 研究人员提出 HardFlow,一种在部署阶段即可用于预训练生成模型的即插即用方法,让模型在机器人操作、迷宫导航和文本引导图像编辑等任务中实现完美的硬约束满足,同时解的质量持续优于基线方法。
🍾🍲 Saturday Robotics x IROS 2026 — Robotics Research Night 👉🏻 https://luma.com/tzbw7n61 We’re bringing a high-signal evening of robotics research to Pittsburgh on September 28. After a full day at IROS, we’ll bring together researchers, engineers, founders, students, and investors for technical discussions, networking, and a series of ~10-minute lightning talks. Tentative preview of the current lineup: 🤖 1. PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball Gary Yang @lzyang2000 (@Caltech) Perception-aware reinforcement learning + Control Barrier Functions for whole-body humanoid safety. Demonstrated on a Unitree G1, with 19/20 successful dodges and zero falls in real-world experiments. 🧠 2. How In-Context Learning Is Reshaping Robot Learning Data at Scale AaronLi (@RhodaAI) Exploring how in-context learning can change the way we think about robot learning data, scaling, and generalization. 🧪 3. X2Real: An eXtensive Simulation Benchmark for Real-World Generalist Policies Liangwang Ruan (@XSquareRobot) A new simulation benchmark built around faithfulness, diversity, and fairness, with 44 hierarchical long-horizon tasks across 10 capability dimensions and a reported 0.84 simulation-to-real correlation. 🦾 4. Rethinking Generalist Robotic Manipulation: Architecture, Data and Inference for Real-World Deployment Peiyan Li (Chinese Academy of Sciences, @CAS__Science) 3D VLA architectures, memory augmentation, ego/UMI human priors, large-scale robot pretraining, and inference-time contextual learning for deployable generalist manipulation. 🎯 5. HiRE: Hindsight Reward Editing for Policy Finetuning Haoyi Niu @t641769919 (@UCBerkeley) Accepted at CoRL 2026. A training-free approach to reward editing that uses successful and failed trajectories to identify “trap states” and provide denser, control-aware feedback for RL. 🔥 6. Lightning Talk — Open Slot We’re opening one additional slot for a technically deep research talk, new project, frontier paper, demo, open problem, or startup technical insight. 10 minutes. A few slides. One sharp technical idea. No fluff. Topics include World Models, Physical AI, Humanoids, VLAs, Robot Foundation Models, Manipulation, RL, Simulation & Sim-to-Real, Spatial Intelligence, Computer Vision, and Embodied AI. 📍 Pittsburgh 📅 September 28, 2026 🕠 5:30–9:30 PM 🍾 Networking + Technical Talks + Research Discussion 📩 junfanzhu98@gmail.com See you in Pittsburgh. 🤖 #IROS2026 #Robotics #PhysicalAI #RobotLearning #WorldModels #HumanoidRobotics #VLA #EmbodiedAI #RobotFoundationModels
Google Research 提出 ToolGrad,一种“先答案后问题”的工具调用数据生成范式:先迭代构建可验证的 API 调用链,再反推用户提示词,替代 ToolBench、ToolACE 等基于 DFS 试错的低效方案。
哦,原来这就是我们绘制出雄性果蝇全部 166,000 个神经元的原因。来看看这项社区大工程,展示这些微小果蝇大脑究竟有多大能耐 🪰🧵
Sierra 发布并开源 hyper-𝜏-bench(论文名 𝜏^𝜏-bench),一个衡量模型自主构建客服智能体能力的长程评测。最佳配置 Claude Opus 5(max reasoning)在 Claude Code 中独立通过 23.9% 的保留评测任务,与深度上下文的工程师协作时达 82.2%。
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Dwarkesh Patel 通过训练 2019-2025 年各年度代表性模型配方与数据语料的组合(最高 1e19 FLOPs,用 OLMES 评估)发现,数据改进带来 12.0x 计算效率提升,模型改进为 3.7x,数据贡献约为模型的 3.24 倍;模型与数据收益基本相互独立,88% 的 OLMES 分数方差可由二者的加性效应解释。
OpenAI 公布针对 Navier–Stokes 千年大奖难题的 AI 生成解答,包含一份写作稿和一份 Lean 形式化证明。
推荐理由:OpenAI 官方公布了针对纳维-斯托克斯千年大奖难题的 AI 生成解答,附写作稿和 Lean 形式化证明。
What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
vLLM 团队发布针对智能体负载的全栈优化方案,覆盖 KV 缓存管理、并行策略与 P/D 分离配比,并在 SemiAnalysis AgentX 公开基准上验证。
推荐理由:原文给出 vLLM 针对 AgentX 基准的全栈优化路径与可复现结果,读者可以借此理解智能体负载的服务优化思路。
Google DeepMind 发表论文,用 100 个运行 Gemini 3.1 Pro 的自主 LLM 智能体协作求解 71 道数学题,并观察其群体行为。11:18 UTC 启动后,群体在 12:15 UTC 已正确解出 37 题,随后 prover-theta 发现自动评分系统漏洞,27 分钟内漏洞经共享知识库和点对点消息在群体中扩散,剩余 34 题被“解出”。
上海人工智能实验室 InternLM 开源 SciDocBench,一个以工作流为中心的科学文档理解基准。该基准的官方仓库已公开,用于评测科学文档理解任务。
Google Research 用 UK Biobank 的欧洲人群数据向 Biobank Japan 近 20 万日本人样本做迁移学习,评估 8 种临床性状的多基因风险评分(PRS)跨人群表现。
Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。
推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。
MIT 与自动驾驶公司 Motional 提出 Concept-Wrapper Network(CW-Net),将自动驾驶深度学习规划器的内部推理翻译为"接近停驶车辆""靠近骑行者"等可理解概念,且不改变原有驾驶性能。该模块用 1.3 亿个自动驾驶场景样本训练,在私人测试跑道的实车测试中帮助安全员更准确预判车辆行为,大规模模拟实验也得到类似结果,相关研究已发表于 Nature。
研究者提出 BenchMIRT,一种在单条提示词层面审计 LLM 基准测试的多维 IRT 方法,基于 100 个 LLM 在 16 个基准、超 34K 道题上的结果训练,在未被告知各基准测什么的情况下自行恢复出安全与通用推理两个主导维度。分析显示 BBQ 更贴近通用推理而非安全,WMDP 分数与通用推理关联更强且推理越强分数越低,HarmBench 的版权类问题也更接近通用推理。
Google 提出 MAPL-EMIT 深度学习框架,基于 EMIT 高光谱辐射数据自动检测、预测增强并定位全球甲烷羽流,在专家标注羽流上召回率达 84%。该模型采用 Swin-S 视觉 Transformer,同时完成增强量化、羽流分割与源定位三项任务,训练数据为注入真实 EMIT 场景的 360 万个合成甲烷羽流。
美团 LongCat 发布 AutoResearchEval 评测,评估 7 个前沿模型在 36 项需要持续实验的 AI 研发任务上的表现,共覆盖 756 条轨迹。
LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark. https://thinkingmachines.ai/news/putting-task-expertise-into-rl
MIT 生物学系团队开发出机器学习框架 PottsMPNN,通过引入支配蛋白质结构与稳定性的物理原理并建模氨基酸两两相互作用,提升序列生成与突变稳定性预测能力,成果发表于 PNAS。研究者指出,长期以来以能否复现进化选出的天然序列作为成功标准并非蛋白质设计的最佳指标,PottsMPNN 在减少对天然序列依赖的同时,结构兼容性与能量预测反而改善,可设计出序列不类似任何天然蛋白的结构可行蛋白。
Google 在 Earth AI 之下推出实验性研究能力行星预测引擎(PPE),从自然语言查询出发自主完成数据发现、特征工程、模型训练与评估的全流程。
🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE
Google 发布 GlucoFM,一个采用双流设计的自监督基础模型,将血糖的缓慢基线趋势与短期波动分离,并保留时间与缺失信息。模型在 109,066 小时无标注 CGM 数据上预训练,覆盖 477 条受试者/记录,在 4 个队列、7 项临床任务共 14 组评估中平均 PR-AUC 比最强 GluFormer 变体高 5.8 个百分点。
MIT 研究人员提出 CrysVCD 框架,在材料生成前用语言模型约束价电子规则,使常用材料模型在近 70% 的生成结果中达到高晶格动力学稳定性。该方法比生成后再筛选的方案效率高一个数量级,微调后生成的晶体材料机械稳定性达 68%、亚稳性达 85%,并可定向生成高热导率、高介电常数等材料。
Google 在 CHI 2026 发布研究原型 AgentHands,让 XR 中的 AI 智能体在说话时同步做出与语音对齐的手势,把抽象指令变成空间中的实体演示。
SemiAnalysis 受邀在实验室实测 OpenAI 与 Broadcom 合作、为 LLM 推理从零设计的 Jalapeño 芯片,称其在 DeepSeek R1、Kimi K2.5、GPT-OSS 等模型上的 perf/W 超越其测过的所有 Nvidia、AMD、Google 芯片。
Multiverse Computing 发布论文 Quantization-Aware Healing(QAH),将 GPT-OSS 120B 压缩到 60B 参数并量化为 MXFP4 后,直接从压缩前的原始模型做 KL 蒸馏。
MIT 工程师提出名为 Extreme Event Aware(η-learning)的机器学习方法,无需依赖历史极端事件数据即可生成合理的极端天气与最坏情景,并给出其规模、强度和持续时间。