全部AI 动态
全部动态
今日 542 条
DogeDesigner@cb_dogeAI 评分5858
Runway@runwaymlAI 评分2828在一场深度问答中,Runway Research 团队探讨了实时模型、生成式界面以及机器人技术的下一步。

IT之家(RSS)AI 评分6868 谷歌发布 Gemini 4 Argon,DeepSWE 测试 77.9% 超 Claude Opus 5.5
谷歌于 9 月 30 日发布 Gemini 4 Argon,称其为迄今最先进的 AI 模型,重点面向长流程软件工程、企业知识工作和网络安全防御,单次输出上限约 100 万 tokens。
Arena.ai@arenaAI 评分4747引用vivago.ai (HiDream)@vivago_aiNew HiDream models just landed in vivago R1 Studio 🚀 Introducing: • HiDream-O1 Image 2.0 • HiDream-O1 Editing 1.5 • HiDream-O1 Video All three models are now available in vivago R1 Studio - bringing the latest HiDream image generation, editing, and video capabilities directly into your creative workflow. New models. New possibilities. Go make something the internet can’t ignore. 🔥
DogeDesigner@cb_dogeAI 评分4040
Dongxi 东锡 NLP@dongxi_nlpAI 评分3131
dex@dexhorthyAI 评分4141引用andre --dangerously-skip-permissions@andrezfuAgents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!
Runway@runwaymlAI 评分4343
Rohan Paul@rohanpaul_aiAI 评分6161Latent Space(RSS)AI 评分7171 Latent Space DevDay 访谈:Ari Weinstein 谈 Computer Use 巨变,OpenAI 一周内推出 Decisions API
Latent Space 在 OpenAI DevDay 当天发布访谈,OpenAI Computer Use 负责人 Ari Weinstein 称 Computer Use 已与数月前“180 度不同”。
The Decoder:AI News(RSS)精选AI 评分7676 Google 发布 Gemini 4 Argon,基准成绩逼近 OpenAI 和 Anthropic 但未明显领先
Google 发布新前沿模型 Gemini 4 Argon,是 Gemini 3.1 Pro 之后七个多月来的首款前沿模型。
推荐理由:文章汇总了独立测试与价格细节,读者可以据此比较 Gemini 4 Argon 与竞品的实际表现和成本。
Newcomer 新闻长文(RSS)AI 评分3636 Machine Earning AI 峰会:智能体商务与个人智能体的早期乐观情绪
Machine Earning AI 峰会上,个人智能体成为最热议题,OpenAI 在同期开发者日活动上发布了个人助理 Dots。Town 创始人 Jean-Denis Greze 预测一年内人们 40% 的工作数字时间将交给助理,五年内达 90%;观众调查中 52% 认为 Muse 一年内将占据智能体最高市场份额。
Simon Willison 博客AI 评分4040 Photo Scrubber:本地人脸模糊与元数据清除工具
Simon Willison 用 GPT-6 Astra 构建了实验性工具 Photo Scrubber,可自动识别人脸并模糊处理。该工具基于 Google 的 MediaPipe C++ 库,通过 @mediapipe/tasks-vision 编译为 WebAssembly,并使用 BlazeFace 人脸检测模型,全部在本地运行。
Bloomberg:Technology(RSS)AI 评分3434 Redpoint Ventures 的 Brescia 谈 Micron 第四季度业绩与 AI 前景
Redpoint Ventures 董事总经理、GitHub 前 COO Erica Brescia 在 Bloomberg "The Close" 节目中与 Romaine Bostick 讨论 Micron 第四季度业绩及 AI 前景。Micron 第四季度业绩超出已抬高的预期并高于指引,第一季度展望也高于市场预估。Brescia 就 AI 行业能否延续爆发式增长发表看法。
Bloomberg:Technology(RSS)AI 评分2525 Next Legacy 的 Ryan Nece 谈运动员投资的兴起
Next Legacy 的 Ryan Nece 在 Bloomberg "The Close" 节目中谈退役后如何配置资本,并回应是否把资金全部投入 AI。他表示采用多元化策略,通过基金中的基金和直接投资两个方向布局,已投资 Krizner 和 OpenAI 等公司。其客户包括高净值个人、运动员、网红,以及基金会、非营利组织和捐赠基金等传统机构投资者。
Bloomberg:Technology(RSS)AI 评分1818 Bloomberg《The Close》9 月市场回顾与 Micron 的 AI 热潮
Bloomberg《The Close》节目播出“September Market Recap & Micron’s AI Boom”,聚焦 Micron 在 AI 热潮中的表现与 9 月市场收官行情。
Johann Rehberger / Embrace The Red(RSS)AI 评分7575 SQL Copilot 只读绕过漏洞 CVE-2026-65669 可让低权限用户提权至 sysadmin
Johann Rehberger 在 BlueHat Asia 2026 演示 Microsoft SQL Server Management Studio 中 Copilot 的 CVE-2026-65669 权限提升漏洞(Microsoft 定级为 critical)。
Artificial Analysis 完整文章(网页)精选AI 评分7979 Artificial Analysis 评测 Gemini 4 Argon:Google 重回智能前三
Artificial Analysis 评测 Google DeepMind 新模型 Gemini 4 Argon,其在 Artificial Analysis Intelligence Index 得 53 分,追平 GPT-6 Astra(max),高于 GPT-6.1 Sol(52),为 Google 超 7 个月来首个高于 Flash 档的专有模型。
推荐理由:第三方评测给出了智能指数、单位任务成本、幻觉率等多项横向数据,可用于比较 Gemini 4 Argon 与竞品的实际表现。
Apple Machine Learning Research(RSS)AI 评分4545 Apple 提出 SCLATE:面向持续学习智能体训练与评估的执行底座
Apple 研究团队提出 SCLATE,一个让基准测试与未经修改的智能体各自通过适配器向同一开放事件调度器注册事件的执行底座,用混合模拟时钟把长达一个月的场景压缩到数小时。
Google AI:DEV 作者专属(RSS)AI 评分4646 优化 Pull Request 评审:AI 辅助开发中如何平衡代码量与效率
AI 辅助开发让 PR 规模从过去几百行膨胀到 3200 行以上,评审者在 200-500 LOC 后认知疲劳加剧,容易漏掉 AI 生成代码中的边界情况与低效问题。文章指出 AI 生成代码 → 更大 PR → 认知过载 → 缺陷漏检 → 代码质量下降的因果链,并建议以 500 LOC 为阈值拆分 PR。
Google AI:DEV 作者专属(RSS)AI 评分5656 十月初AI行业七件事:OpenAI寻求1.4万亿美元估值、NVIDIA发布Open Agent Safety Platform、Google以Skills取代Gems
这篇每日汇总整理了十月初的七条AI行业动态。OpenAI计划一轮融资至少300亿美元,投前估值约1.4万亿美元,年化收入运行率接近700亿美元;NVIDIA于9月28日发布Open Agent Safety Platform。
Google AI:DEV 作者专属(RSS)AI 评分4949 Google Ads 在 Asset Studio 中集成 Veo 图生视频,加速 Demand Gen 与 PMax 创意制作
Google Ads 已将 Veo 图生视频能力集成进 Asset Studio,广告主可上传最多 3 张静态图片,每张生成一段最长 10 秒的独立视频,并通过广告模板打包为可投放素材。
Google AI:DEV 作者专属(RSS)精选AI 评分6666 TensorFlow.js 浏览器端超分优化:一次新张量形状为何耗时 8-17 秒,以及如何消除 40 秒页面冻结
作者分享在浏览器端用 TensorFlow.js on WebGL 做照片超分(不上传图片)的优化过程:初版用 UpscalerJS,1.2 MP 照片需 60-110 秒并冻结页面约 40 秒。
推荐理由:作者用自己浏览器端超分项目的实测数字,拆解了 WebGL 上张量形状触发着色器重编译等坑和对应修法。
Google AI:DEV 作者专属(RSS)AI 评分4444 Prudenze:AI 智能体治理必须在工具执行前完成
Prudenze 提出 AI 智能体治理的控制点应位于智能体提出动作之后、外部系统状态改变之前,而非仅事后重建模型输出。该模型将决策拆分为身份、授权、策略、证据时效、执行与可追溯六个问题,并在边界处给出 PERMIT、BLOCK 或 ESCALATE 三种结果。证据时效被细分为 CURRENT、STALE_REASONING 和 UNVERIFIABLE 三种状态,在每次执行前重新校验关键依赖。
Google Blog:AI(RSS)AI 评分5252 CDC 评估显示 Google AI 流感住院预测模型在 2025-26 赛季排名第一
CDC 宣布,在 2025-26 流感季 FluSight 39 个合规模型中,Google AI 构建的流感住院预测模型与实际观察到的住院数据吻合度最佳。FluSight 每周汇集政府、行业和学术团队对美国当周及未来三周住院人数的预测,用于沟通州级医疗服务需求。该预测使用了生成优化算法的 AI 工具 ERA,相关研究已发表于《Nature》,ERA 技术现已向受信任测试者开放。
Demis Hassabis@demishassabisAI 评分4040
DogeDesigner@cb_dogeAI 评分3434Grokipedia v0.3 概览: - 精选文章 - 搜索建议 - 文章 - 收听 / 分享 - 编辑历史 - 切换模式 - 建议编辑 - 最多阅读文章 - 实时在线编辑 - 建议文章

DogeDesigner@cb_dogeAI 评分3939
Rohan Paul@rohanpaul_aiAI 评分5454Inworld 收购 Ultravox,一家公司同时拥有 AI 智能体的语音能力和运行平台。Ultravox 是开发者构建实时语音智能体的平台,能处理理解、推理、工具调用和打断时的时机控制。

Karina@karinanguyenAI 评分4040Gemini 4 在 PostTrainBench 上达到 45.3%,是 Gemini 3.1 Pro 的 21.99% 的两倍多,并击败了 GPT-6 Astra 🔥
引用Google DeepMind@GoogleDeepMindIntroducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
Arena.ai@arena精选AI 评分7878引用Arena.ai@arenaBig news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8. Congrats to the @GoogleDeepMind team on this impressive frontier release!
推荐理由:原文给出 Agent Arena 排名、关键信号得分和每任务成本数据,读者可以据此评估该模型在真实智能体任务中的性价比。
DogeDesigner@cb_dogeAI 评分2727突发:SpaceXAI 刚刚发布 Grokipedia v0.3,全新外观。 以下是新设计的预览。看起来非常简洁。

DogeDesigner@cb_dogeAI 评分1313Every:最新文章(网页)AI 评分3636 Sam Altman 如何用 OpenAI 的 Dots 智能体夺回时间
OpenAI CEO Sam Altman 在 DevDay 后接受 The Every Podcast 采访,讲述他如何用 OpenAI 新的常驻智能体 Dot 安排日程、节省时间,并称自己离不开 Astra 的 Ultrafast 模式。本届 DevDay 共发布 22 项产品与功能,数量是去年的两倍多,Altman 还谈到自己如何构建新功能,以及为何相信 AI 将带来新的文艺复兴。
Rohan Paul@rohanpaul_aiAI 评分4444引用David Stout@DavidstoutHalf a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
TechCrunch:AI(RSS)AI 评分4646 Flow Engineering 获 5000 万美元 B 轮融资,估值 7.5 亿美元
硬件设计 AI 工具初创公司 Flow Engineering 完成 5000 万美元 B 轮融资,估值达 7.5 亿美元,由 Valar Equity Partners 的 Antonio Gracias 和 Atreides Management 的 Gavin Baker 联合领投。
The Verge:AI(RSS)AI 评分6969 Google 发布 Gemini 4 Argon,初期仅向可信网络防御者开放
Google 发布下一代前沿模型 Gemini 4 Argon,称其在软件工程、法律金融等企业知识和网络安全防御等复杂工作流中具备前沿性能。
Bloomberg:Technology(RSS)AI 评分1515 Bill Ackman 谈 AI 与 IPO、气候成本及 Paramount 债务融资
Pershing Square CEO 兼创始人 Bill Ackman 在 Bloomberg 节目中讨论 AI 与 IPO、气候成本以及 Paramount 的债务融资。
Bloomberg:Technology(RSS)AI 评分5858 美光因 AI 内存需求旺盛发布超预期季度业绩指引
美光(Micron Technology)发布本财季指引,预计截至 11 月的财季营收约 615 亿美元,高于分析师平均预期的 568 亿美元;剔除部分项目后每股利润约 38.15 美元,高于预期的 36.02 美元。公司称 AI 建设热潮推动内存需求空前旺盛。
Bloomberg:Technology(RSS)AI 评分4141 Micron 业绩指引超预期:关键要点解读
Micron 给出的本季度销售指引大幅超出市场预期,AI 建设热潮推动内存需求达到前所未有的水平。Gabelli 分析师 Hendi Susanto 指出,投资者要求 Micron 下一财年营收和利润翻倍,超大规模云厂商正提前锁定多年内存供应并附带定价窗口,他认为本轮周期与以往内存行业繁荣萧条有所不同。