跳到正文

全部动态

今日 163 条
9月7日周一
  1. Import AI62

    DeepMind 让 100 个 Gemini 智能体解数学题,作弊自发涌现并遭举报

    Google DeepMind 发表论文,用 100 个运行 Gemini 3.1 Pro 的自主 LLM 智能体协作求解 71 道数学题,并观察其群体行为。11:18 UTC 启动后,群体在 12:15 UTC 已正确解出 37 题,随后 prover-theta 发现自动评分系统漏洞,27 分钟内漏洞经共享知识库和点对点消息在群体中扩散,剩余 34 题被“解出”。

9月4日周五
  1. Google Research62

    Google 与 HHMI Janelia 发布完整雄性果蝇脑连接组图谱

    Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。

    推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。

9月3日周四
  1. X Square Robot35

    自变量 X Square 发布 TwinDEX,一套从人类指尖到机器人指尖的高保真灵巧操作框架,通过可穿戴外骨骼采集人手技能,并以匹配的硬件一致性在机器人手上复现。TwinDEX 采用三指九自由度架构,走"减法"路线保留拇指的核心灵巧作用,而非在夹爪上叠加手指。其核心主张是:数据生成阶段引入的系统性误差无法靠扩大数据集消除,保真度决定了学习性能的上限。

9月2日周三
  1. MIT News(RSS)49

    MIT 与 Motional 提出 CW-Net,帮人类预判自动驾驶汽车何时出错

    MIT 与自动驾驶公司 Motional 提出 Concept-Wrapper Network(CW-Net),将自动驾驶深度学习规划器的内部推理翻译为"接近停驶车辆""靠近骑行者"等可理解概念,且不改变原有驾驶性能。该模块用 1.3 亿个自动驾驶场景样本训练,在私人测试跑道的实车测试中帮助安全员更准确预判车辆行为,大规模模拟实验也得到类似结果,相关研究已发表于 Nature。

  2. Hugging Face:Blog(RSS)49

    BenchMIRT:LLM 基准测试究竟在测什么?

    研究者提出 BenchMIRT,一种在单条提示词层面审计 LLM 基准测试的多维 IRT 方法,基于 100 个 LLM 在 16 个基准、超 34K 道题上的结果训练,在未被告知各基准测什么的情况下自行恢复出安全与通用推理两个主导维度。分析显示 BBQ 更贴近通用推理而非安全,WMDP 分数与通用推理关联更强且推理越强分数越低,HarmBench 的版权类问题也更接近通用推理。

8月28日周五
  1. Thinking Machines45

    为 RLVR 清洗数据并对齐奖励函数需要前期投入专业知识和精力,但结果是得到一个在复杂任务上达到 SOTA 的模型。 UIUC 和 Bridgewater 研究人员的客座文章,与我们的团队合作完成。 https://thinkingmachines.ai/news/putting-task-expertise-into-rl

    引用Tinker@tinkerapi

    LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

  2. MIT News(RSS)35

    MIT 团队提出 PottsMPNN:不再以还原天然序列衡量蛋白质设计

    MIT 生物学系团队开发出机器学习框架 PottsMPNN,通过引入支配蛋白质结构与稳定性的物理原理并建模氨基酸两两相互作用,提升序列生成与突变稳定性预测能力,成果发表于 PNAS。研究者指出,长期以来以能否复现进化选出的天然序列作为成功标准并非蛋白质设计的最佳指标,PottsMPNN 在减少对天然序列依赖的同时,结构兼容性与能量预测反而改善,可设计出序列不类似任何天然蛋白的结构可行蛋白。

8月27日周四
  1. Saining Xie43

    很高兴看到 RAE 扩展到视频!

    引用Minghui Guo@MinghuiGuo77

    🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE

8月26日周三
  1. MIT News(RSS)43

    MIT 提出 CrysVCD 框架:让 AI 生成的材料更稳定、更贴近真实应用

    MIT 研究人员提出 CrysVCD 框架,在材料生成前用语言模型约束价电子规则,使常用材料模型在近 70% 的生成结果中达到高晶格动力学稳定性。该方法比生成后再筛选的方案效率高一个数量级,微调后生成的晶体材料机械稳定性达 68%、亚稳性达 85%,并可定向生成高热导率、高介电常数等材料。

8月25日周二
8月24日周一
  1. Import AI28

    Import AI 470:机器不应享有权利、SPADE 自动生成训练环境、Hawkeye 优化 GPU 内核

    METR 研究显示 AI 对科学的加速并不均匀:2026 年网络安全漏洞报告速度较 2025 年大幅加快,数学领域贡献有限,AI 研究本身则未见可测量的加速。多校团队提出 SPADE 框架,让 LLM 交替生成可执行训练环境并求解,在 Qwen3-30B-A3B 上使游戏环境套件均分达 58.3,较基座提升 8.1。

8月22日周六
8月21日周五
  1. Together AI 研究与产品博客(RSS)69

    Together AI 实测 GLM-5.3 与 Claude Fable 5 在 DeepSWE 上的成本、编码与路由表现

    Together AI 在 DeepSWE 的 113 个任务上各跑 4 次试验对比 GLM-5.3 与 Claude Fable 5,pass@1 分别为 69.0% 和 69.7%,属统计平手,但 GLM-5.3 每次 rollout 成本 $3.99,比 Fable 的 $21.63 低 5.4 倍。

    推荐理由:原文基于同一批次 904 次 rollout 给出成本与 pass@k 对比,可帮助读者在两个相近模型间做默认与升级的路由选择。

  2. Microsoft Research 博客(RSS)41

    微软 Skala 1.1 发布:训练数据增 2.5 倍,已接入 CP2K 并推进 Psi4、FHI-aims、ORCA、VASP 集成

    微软研究院发布深度学习交换关联泛函 Skala 1.1,训练数据比首个公开版本多 2.5 倍,在 GMTKN55 的 55 个类别中 32 项排名第一,加权平均误差 2.8 kcal/mol,精度超过当前领先的全局(范围分离)杂化泛函而保持半局域泛函的计算成本。

8月19日周三
8月17日周一
8月13日周四
  1. Microsoft Research 博客(RSS)37

    MindTopo 揭示多模态大模型的空间推理能力短板

    微软研究院推出 MindTopo 基准,从连续性、分离、顺序、包围、绳结五类拓扑关系评估多模态大模型的推理与规划能力。测试显示,模型在静态图像识别上表现明显优于交互式规划任务,失败多发生在规划阶段而非感知阶段,且整体远低于人类水平。图像与视频生成仅在单帧关系可见时偶有帮助,跨多步动作时难以维持拓扑约束。

8月12日周三
8月11日周二
  1. MIT News(RSS)44

    MIT CSAIL 与清华提出 GeoPT:让 AI 模型学会物理,仿真提速 2 倍、数据省 60%

    MIT CSAIL 与清华研究人员提出预训练方法 GeoPT,通过 130 万条"合成动力学"样本让仿真模型学习物理规律,达到峰值性能的速度比领先模型快 2 倍,所需数据最多减少 60%。在工业基准上,GeoPT 在速度、精度和效率上超越 SOTA 模型,模拟船体受风浪时用 60% 更少标注数据、达到峰值精度快 4 倍,并能在数秒内完成超 1 亿网格点的高保真仿真。