推荐理由:官方给出性能对标与 40% 成本降幅两个具体指标,读者可以据此评估替换现有 Opus 方案的性价比。
全部AI 动态
全部动态
今日 46 条
Claude@claudeai精选AI 评分6565
Unsloth AI@UnslothAIAI 评分6363引用Qwen@Alibaba_QwenMeet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: https://qwen.ai/blog?id=qwen-image-2.1 - GitHub: https://github.com/QwenLM/Qwen-Image-2.1 - Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1 - Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1
OpenAI:官网动态(RSS · 排除企业/客户案例)AI 评分3232 Parallel 用 GPT‑6 Astra 将研究时间和成本减半
Parallel 借助 GPT‑6 Astra 让智能体研究并综合劳动力市场数据,耗时和成本均降至此前模型的一半。
karminski-牙医@karminski3AI 评分4040引用Max For AI@MaxForAI🚨Qwen4家族首次曝光!! 刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族! 包含Qwen4-Max Qwen4-Flash&Qwen4-Plus 还有Qwen4-27B!!! 未来Qwen会训5-10T的模型
Latent Space(RSS)精选AI 评分7979 Xiaomi MiMo-V2.6-Pro 1T-A42B 登顶开源权重模型,训练仅花费约 $3M
Latent Space AINews 汇总 2026/9/19-9/21 AI 动态,核心是 Xiaomi 发布 MiMo-V2.6-Pro(1.02T 总参数/42B 激活,MIT 许可),以 Artificial Analysis Intelligence Index 46 分成为新的开源权重榜首,成本为 $0.435/M 输入、$0.87/M 输出 token。
推荐理由:除发布信息外还汇总了 RL 成本与训练细节,读者可以看到开源权重模型追赶闭源的具体路径。
Microsoft:GitHub 新仓库AI 评分2626 Microsoft 开源 ai_night_scientist:用 Night Science 强化科学构想中的智能体创造力
Microsoft 在 GitHub 发布新仓库 ai_night_scientist,主题为用 Night Science 强化科学构想中的智能体创造力。该仓库聚焦 AI 智能体在科研创意生成环节的能力增强,目前公开信息仅包含项目名称与方向,未披露模型、参数或评测数据。
Qwen@Alibaba_QwenAI 评分3535来自 @inteldevs 的 Day-0 OpenVINO 支持!🥳 Qwen-Image-2.1 已可在 Intel 硬件上优化运行。一个开放权重 checkpoint,同时支持生成与编辑。👇
引用Intel Devs@inteldevsWe're excited to offer Day0 OpenVINO support for Qwen-Image-2.1 Read more about what you can accomplish here: https://ms.spr.ly/6019a5nm5
elsewhere:文章(RSS)AI 评分6868 阶跃 Step 5 Preview 实测评测:数据可视化与金融分析亮眼,泛化和审美仍有短板
阶跃发布 Step 5 Preview,总参数量 600B、激活参数 27B,有视觉输入,官方称在 Artificial Analysis 上涨 44 分、单任务成本仅为 Claude Opus 5 的 1/8。作者与友人实测发现其在数据可视化、金融分析上表现不错,但泛化性、领域知识和审美偏弱,思考过程过长导致长任务耗时且易中断,且长上下文下安全指令遵循可被绕过。
StepFun@StepFun_aiAI 评分6363引用Artificial Analysis@ArtificialAnlysStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th
Simon Willison 博客AI 评分7373 TypeSafe AI 发布决策模型 Jev,只返回带置信度的数值输出
TypeSafe AI 上周发布 Jev,称为 System One 模型,作者更倾向叫决策模型。它接受文本输入但输出浮点数形式的分类、是非判断、评分及置信度,只按输入收费,价格为 $0.042 per million tokens,低于 OpenAI GPT-5 Nano 的 $0.05/million,输出免费。
Fuli Luo@_LuoFuliAI 评分6666小米 MiMo 团队发布 MiMo-V2.6,称其可能是开源模型团队迄今计算量最大的单次 RL 运行之一,通过 mid-training 与高强度 RL 打造,现为排名第一的开源模型。
Xiaomi MiMo@XiaomiMiMoAI 评分4040引用Design Arena@DesignArenaBREAKING: MiMo-V2.6-Pro by @XiaomiMiMo lands at #8 overall (#3 open-weight) on Design Arena with an Elo of 1338. This is an impressive 54-point and 22-position increase from MiMo-V2.5-Pro. MiMo-V2.6-Pro also reaches #4 overall in Website (#2 open-weight) and #6 overall in Agentic Frontend Development (#2 open-weight), showing strong performance across both direct generation and agentic coding. This places @XiaomiMiMo's new model among leading models such as GPT-5.6 Sol by @OpenAI, Claude Opus 5 by @Anthropic, and Kimi K3 by @MoonshotAI. Congratulations to the @XiaomiMiMo team on returning to a top-10 placement on Design Arena!
Xiaomi MiMo@XiaomiMiMoAI 评分5656引用Arena.ai@arenaMiMo-V2.6-Pro just landed @XiaomiMiMo back in the top 10 on Code Arena: WebDev, debuting at ~#10 overall, and ~#3 among open-weights models with an MIT license. It scores 1628 pts (AutoEval), tying Claude Fable 5 (High) and just ahead of Hy4-preview (1624 pts). That's a +153 pt jump from the previous MiMo-V2.5-Pro (1475 → 1628 pts). Among open-weights models, it lands at ~#3. Impressively only 7 pts behind Qwen3.8 Flash Next (#2) and 46 pts behind Kimi K3 Max (#1). Note: this is an early AutoEval score, in which a Reward Model trained on Arena’s human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. Congrats to the @XiaomiMiMo team on this release!
Xiaomi MiMo@XiaomiMiMo精选AI 评分6666推荐理由:原文给出 Pro 与 Claude Opus 5、GPT-5.6 Sol 的 agent 基准对比和开源范围,可据此评估其相对位置。
eric zakariasson@ericzakariassonAI 评分7575
Michael Truell@mntruellAI 评分4040引用SpaceXAI@SpaceXAIGrok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
SpaceXAI@SpaceXAIAI 评分4444
小米 MiMo:GitHub 新仓库(模型发布)AI 评分2525 小米 MiMo 发布 uni-agent:面向长程智能体训练的框架
小米 MiMo 在 GitHub 新建仓库 uni-agent,这是一个用于训练长程(long-horizon)智能体的框架。目前公开信息仅包含框架定位,未披露模型规模、训练数据或评测结果。
小米 MiMo:GitHub 新仓库(模型发布)AI 评分2828 小米 MiMo 的 verl/HybridFlow:灵活高效的 RL 后训练框架
小米 MiMo 在 GitHub 新建仓库 verl,其 HybridFlow 是一个灵活高效的强化学习后训练框架。该仓库定位为 RL post-training 框架,强调灵活性与效率,目前原文未披露参数规模、benchmark 分数或开源许可等细节。
swyx@swyxAI 评分2323引用Diogo Almeida@CompleteSkepticAfter co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
Qwen@Alibaba_Qwen精选AI 评分6666引用Hugging Apps@HuggingAppsQwen Image 2.1 is here! 🖼️ A 7B params native image generation and editing model, with up to 10 image references The model comes with it's own prompt enhancement LLMs, integrated with diffusers 🧨 and ComfyUI ▶️ on Spaces https://huggingface.co/spaces/hugging-apps/qwen-image-2-1
推荐理由:原文给出了模型的参数量、参考图能力与免配置体验入口,读者可以直接在浏览器试用判断适用性。
Mustafa Suleyman@mustafasuleymanAI 评分4949引用Artificial Analysis@ArtificialAnlysGenerating high-quality images is cheaper and faster than ever. Muse Image, MAI-Image-2.6 and GPT Images 2.5 have substantially shifted the Text to Image Pareto frontiers for both price and speed in recent weeks.
SenseTime@SenseTime_AIAI 评分5353商汤发布 SenseNova U1.5 技术报告,这是一个开源的 8B-MoT 原生统一模型,通过共享注意力连接理解与生成。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分4949 蚂蚁 inclusionAI 发布 Ming-Image-0.1-Design-Layer 设计图层分解模型
蚂蚁 inclusionAI 在 Hugging Face 发布 Ming-Image-0.1-Design-Layer,可将扁平化设计图按指定层数拆解为 RGBA 图层,采用 MIT 许可。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分5151 蚂蚁 inclusionAI 发布 Ming-Image-0.1-Design 文生图模型
蚂蚁 inclusionAI 在 Hugging Face 发布 Ming-Image-0.1-Design,是一个面向 UI、信息图、海报等文字密集视觉设计的 6B 文生图模型,支持完整视觉构图和带透明背景的 RGBA 输出。
Fuli Luo@_LuoFuliAI 评分3838
Epoch AI@EpochAIResearchAI 评分3232
StepFun@StepFun_aiAI 评分5757阶跃星辰与 ACE Studio 发布 StepAudio 3 Music,是 StepFun 首个音乐生成基础模型,可根据提示词和歌词生成完整歌曲。
NVIDIA Blog(RSS)AI 评分5454 NVIDIA Vera Rubin NVL72 首次参加 MLPerf Inference v6.1,吞吐最高达 GB300 NVL72 的 3.7x
NVIDIA Vera Rubin NVL72 首次提交 MLPerf Inference v6.1 预览结果,在 Qwen3-VL 上吞吐最高达 GB300 NVL72 的 3.7x,在 DeepSeek-R1 上达 2.5x,分别使用 vLLM 搭配 NVIDIA Dynamo 和 TensorRT-LLM。
蚂蚁 inclusionAI:HuggingFace 新模型精选AI 评分6262 蚂蚁 inclusionAI 开源 Realtime-Venus 9B 全双工音视频交互系统
蚂蚁 inclusionAI 发布 Realtime-Venus,一个支持主动音视频交互、异步委托和可打断全双工对话的开源系统,包含 Realtime-Venus-Omni 和 Realtime-Venus-Audio 两个 9B 检查点。
推荐理由:模型卡给出了 Realtime-Venus 的架构组成、全双工与异步委托能力及完整用法,读者可以据此评估它在实时音视频交互场景的可用性。
Jeff Dean@JeffDeanAI 评分4242令人振奋的成果,@LiamFedus!祝贺 Periodic Labs 的整个团队!
引用Liam Fedus@LiamFedusWe built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
Hao AI Lab@haoailabAI 评分4242
ViggleAI@ViggleAIAI 评分5353
StepFun@StepFun_aiAI 评分5757
Logan Kilpatrick@OfficialLoganKAI 评分5757
Google AI@GoogleAI精选AI 评分6666Google 发布迄今最先进的 Gemini Audio 模型:Gemini 3.8 Live 与 3.8 Live Extended Thinking。

推荐理由:官方说明了两款新音频模型在速度成本与推理能力上的分工,读者可据此判断语音交互场景如何选型。
Google DeepMind:Blog(RSS)精选AI 评分8282 Google DeepMind 发布 Gemini 3.8 Live 与 3.8 Live Extended Thinking 语音模型
Google DeepMind 发布 Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking 两款实时对话模型,前者面向规模化与成本效率,后者面向高复杂度任务与多步推理。
推荐理由:官方给出两款语音模型的基准成绩与开放渠道,可据此判断实时语音智能体的能力与成本取舍。
Google DeepMind@GoogleDeepMindAI 评分5252
Odyssey@odysseymlAI 评分2828