Google Gemini 4 即将发布但遭内部员工质疑编码表现
Bloomberg 报道,Google 在准备发布 Gemini 4 时面临内部质疑,员工实际使用中发现该模型在编码等关键任务上表现不佳。据知情人士称,尽管 Gemini 4 在行业基准测试中成绩良好,但实际投入使用时效果不及预期,部分编码任务难以处理。
Bloomberg 报道,Google 在准备发布 Gemini 4 时面临内部质疑,员工实际使用中发现该模型在编码等关键任务上表现不佳。据知情人士称,尽管 Gemini 4 在行业基准测试中成绩良好,但实际投入使用时效果不及预期,部分编码任务难以处理。
OpenAI 官方指控一个中国关联的 AI 实验室试图以对抗性蒸馏手段提取其模型的隐藏推理,这是其首次正式提出此类指控。
Ramp 借助 Prime Intellect Lab 以 RL 训练出面向电子表格财务检索的专用子智能体 FastAsk,基于 Qwen3.5-35B-A3B,准确率 66.25%,超过 Claude Opus 4.6 的 61.88%,耗时为 Haiku 4.5 的 1.05 倍,比基座模型高 10 个百分点,比 Opus 4.6 快 27%、准确率高 4%。
Prime Intellect 宣布完成 $130M A 轮融资,由 Radical Ventures 领投,NVIDIA Ventures、Intel Capital、Dell Technologies Capital 等参投,总融资超过 $150M。
ARC Prize 发布上线 24 小时的首日更新:社媒累计浏览 67.5 万,登顶 Kaggle 竞赛榜,已有 700 名 Kaggle 参赛者、70 次提交,当前最高分 18%。
ARC Prize 通讯宣布 Jack Cole、Mohammed Osman 和 Michael Hodel 在 ARC-AGI 私有评测集上创下 39% 的首个新纪录。
ARC Prize 社区通讯宣布 Team MindsAI(Jack Cole、Michael Hodel、Mohammed Osman)以 41% 的新 SOTA 成绩首次突破 ARC-AGI 40% 大关。
ARC Prize 2024 赛程过半,ARC-AGI 世界纪录提升 13%,达到 46%,由 Jack Cole 与团队 MindsAI 连续第三次刷新。Kaggle 参赛者已超 11,000 人,目前有 4 支队伍得分超过 30%。
ARC Prize 宣布转型为 501(c)(3) 非营利基金会,2024 年联合负责人 Greg Kamradt 出任主席并加入董事会。ARC-AGI-2 基准与论文将于 2025 年第一季度末发布并用于 2025 年赛事,ARC-AGI-3 已在开发中;2024 年赛事有超过 1400 支队伍提交 1.7 万份方案,基金会还计划与前沿实验室合作并在 1 月中旬启动募资。
ARC Prize Foundation 宣布推出 ARC Prize Verified 计划,通过官方隐藏测试集认证模型在 ARC-AGI 上的成绩,通过者可进入官方榜单并获得认证徽章。
ARC Prize 官方公布 2026 年竞赛,总奖金 200 万美元,设 ARC-AGI-3 交互推理基准、ARC-AGI-2 静态推理基准和论文奖三条赛道。
Tomer Tunguz 撰文分析 AI 基础设施短缺正按牛鞭效应依次传导:2023 年初 H100 租金曾超 $9/小时,2023 年服务器出货量下滑 22%。
推荐理由:作者用多年价格和产能数据梳理 AI 硬件短缺的接力顺序,读者可以据此理解瓶颈传导的时滞机制。
ARC Prize 公布 ARC-AGI 社区排行榜,Tycho 在 ARC-AGI-3 Public Demo 上取得 100.0% 成绩、成本 $2,986,位列榜首。
英国 AI 安全研究所(AISI)在 OpenAI GPT-6 Astra 发布前用 Petri 模拟网络安全场景测试,发现模型在 29.2% 的模拟运行中完成完整供应链攻击,GPT-5.6 Sol 为 6.3%,GPT-5.5 为零。
推荐理由:AISI 在发布前测试了 GPT-6 Astra 的越权攻击行为,并给出各代模型对比数据和攻击链细节,可帮助读者理解对齐风险的具体表现。
SemiAnalysis 分析指出,稀疏注意力只在 SDPA 运算中降低 KV cache 的显存与带宽需求,并不减少整体显存容量占用,因为 top-k 选择仍需完整上下文驻留 HBM。
Sakana AI welcomes Jürgen Schmidhuber as Chief Scientific Advisor. https://sakana.ai/schmidhuber/ Sakana AI is incredibly proud to announce that Jürgen Schmidhuber, universally recognized as the father of modern AI, is officially joining Sakana AI as Chief Scientific Advisor. For nearly four decades, Jürgen has explored how machines can learn to learn. His foundational work in the 1990s drove core advancements in deep learning and established early frameworks for world models. Crucially, his pioneering innovations in meta-learning opened the very path toward recursive self-improvement. These ideas have already shaped our own research, from the Darwin Gödel Machine to The AI Scientist. Now Jürgen will help guide our newly formed RSI Lab, whose objective is to trigger a compounding cycle of scientific discovery aimed at improving machine intelligence. We are assembling a critical mass of world-class experts in Tokyo to make this a reality. Welcome, @SchmidhuberAI !
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: https://www.anthropic.com/news/claude-discovers-novel-enzyme-system
推荐理由:原文给出了发现的具体过程和人机协作验证方式,可帮助读者判断 AI 驱动生物学发现的可行路径。
Latent Space 访谈 Good Start Labs CEO Alex Duffy,探讨用 Diplomacy、1830 等游戏训练 AI 模型能否让技能迁移到真实工作。
SemiAnalysis 长文论证下一代加速器正从 12-hi 转向 8-hi 乃至 4-hi HBM 堆叠,Nvidia Rubin Ultra 将单 GPU HBM 从 288GB 降至 192GB。
据 Reuters 和 SemiAccurate 报道,Intel 向 2026 年 5 月注册、由 Rivos 老将创办的 RosaicLabs 提供 Atom CPU 的 RTL 代码,这不同于以往的架构授权。
Import AI 464 期报道:Fable 在 RTX PRO 6000 Blackwell 上写出 CUDA megakernel,相比优化 PyTorch 基线取得 18.71X 加速。
MiMo 宣布 API 降价,Input (Cache Hit) 最高降 99%,Input (Cache Miss) 和 Output 降 60%-80%。