跳到正文

全部动态

今日 46 条
9月23日周三
  1. Unsloth AI63

    千问(Qwen)发布开源图像生成与编辑模型 Qwen-Image-2.1,7B 参数,官方称基准表现与 Nano Banana 2.0 相当。Unsloth 发布 GGUF 量化版,支持 12GB 显存本地运行,也可通过 offloading 在 6GB 显存运行 Dynamic FP8;量化文件见 https://huggingface.co/unsloth/Qwen-Image-2.1-GGUF,指南见 https://unsloth.ai/docs/models/qwen-image-2.1。原模型统一支持生成与编辑,可原生生成和编辑 RGBA 透明图层,支持最多 10 张参考图,链接包括 https://qwen.ai/blog?id=qwen-image-2.1 和 https://github.com/QwenLM/Qwen-Image-2.1。

    引用Qwen@Alibaba_Qwen

    Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: https://qwen.ai/blog?id=qwen-image-2.1 - GitHub: https://github.com/QwenLM/Qwen-Image-2.1 - Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1 - Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

9月22日周二
  1. karminski-牙医40

    卧槽5-10T???

    引用Max For AI@MaxForAI

    🚨Qwen4家族首次曝光!! 刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族! 包含Qwen4-Max Qwen4-Flash&Qwen4-Plus 还有Qwen4-27B!!! 未来Qwen会训5-10T的模型

  2. Latent Space(RSS)79

    Xiaomi MiMo-V2.6-Pro 1T-A42B 登顶开源权重模型,训练仅花费约 $3M

    Latent Space AINews 汇总 2026/9/19-9/21 AI 动态,核心是 Xiaomi 发布 MiMo-V2.6-Pro(1.02T 总参数/42B 激活,MIT 许可),以 Artificial Analysis Intelligence Index 46 分成为新的开源权重榜首,成本为 $0.435/M 输入、$0.87/M 输出 token。

    推荐理由:除发布信息外还汇总了 RL 成本与训练细节,读者可以看到开源权重模型追赶闭源的具体路径。

  3. elsewhere:文章(RSS)68

    阶跃 Step 5 Preview 实测评测:数据可视化与金融分析亮眼,泛化和审美仍有短板

    阶跃发布 Step 5 Preview,总参数量 600B、激活参数 27B,有视觉输入,官方称在 Artificial Analysis 上涨 44 分、单任务成本仅为 Claude Opus 5 的 1/8。作者与友人实测发现其在数据可视化、金融分析上表现不错,但泛化性、领域知识和审美偏弱,思考过程过长导致长任务耗时且易中断,且长上下文下安全指令遵循可被绕过。

  4. StepFun63

    阶跃星辰发布 Step 5 Preview,在 Artificial Analysis Intelligence Index 得 44 分,作者称其将智能-成本 Pareto 前沿外推,每任务成本约 $0.71。引用的 Artificial Analysis 评测称其为 600B 总参数、27B 激活的 MoE 模型,定价 $1/$2.70 每百万输入/输出 token,上下文窗口 1M token,支持文本、图像和视频输入,当前闭源权重,计划 10 月 15 日开放权重。

    引用Artificial Analysis@ArtificialAnlys

    StepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th

  5. Xiaomi MiMo40

    💗

    引用Design Arena@DesignArena

    BREAKING: MiMo-V2.6-Pro by @XiaomiMiMo lands at #8 overall (#3 open-weight) on Design Arena with an Elo of 1338. This is an impressive 54-point and 22-position increase from MiMo-V2.5-Pro. MiMo-V2.6-Pro also reaches #4 overall in Website (#2 open-weight) and #6 overall in Agentic Frontend Development (#2 open-weight), showing strong performance across both direct generation and agentic coding. This places @XiaomiMiMo's new model among leading models such as GPT-5.6 Sol by @OpenAI, Claude Opus 5 by @Anthropic, and Kimi K3 by @MoonshotAI. Congratulations to the @XiaomiMiMo team on returning to a top-10 placement on Design Arena!

  6. Xiaomi MiMo56

    小米 MiMo 发布 MiMo-V2.6-Pro,据 Arena 数据以 1628 分(AutoEval)位列 Code Arena: WebDev 总榜约第 10 名、MIT 许可开源权重模型中约第 3 名,较上一代 MiMo-V2.5-Pro 的 1475 分提升 153 分。该成绩为早期 AutoEval 分数,由基于人类偏好数据训练的 Reward Model 自动投票产生,随真实人类投票增加可能变化。

    引用Arena.ai@arena

    MiMo-V2.6-Pro just landed @XiaomiMiMo back in the top 10 on Code Arena: WebDev, debuting at ~#10 overall, and ~#3 among open-weights models with an MIT license. It scores 1628 pts (AutoEval), tying Claude Fable 5 (High) and just ahead of Hy4-preview (1624 pts). That's a +153 pt jump from the previous MiMo-V2.5-Pro (1475 → 1628 pts). Among open-weights models, it lands at ~#3. Impressively only 7 pts behind Qwen3.8 Flash Next (#2) and 46 pts behind Kimi K3 Max (#1). Note: this is an early AutoEval score, in which a Reward Model trained on Arena’s human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. Congrats to the @XiaomiMiMo team on this release!

  7. Xiaomi MiMo66

    小米 MiMo 发布 MiMo-V2.6 Pro 与 Flash 两款全模态模型,通过规模化强化学习训练。官方称 Pro 在多数 agent 基准上与 Claude Opus 5 和 GPT-5.6 Sol 相当,Artificial Analysis Intelligence Index 得分 46,为开源模型中最高;能力覆盖编码、computer use、3D 推理和创作。

    推荐理由:原文给出 Pro 与 Claude Opus 5、GPT-5.6 Sol 的 agent 基准对比和开源范围,可据此评估其相对位置。

9月21日周一
  1. swyx23

    Jev 播客明天上线 在 apple /youtube 订阅 @latentspacepod 感谢 @allenpark 和 Ke 促成此事! 引用推文核心要点:在联合发明 ChatGPT 之后,我一直在问自己:为什么超人类对话模型没有带来 AGI? 过去 2 年我一直在隐身模式下构建一种新的模型训练方式(RLCD),以及一种新型前沿 AI 模型,今天正式发布:Jev • 快 20-200 倍 • 便宜 40-400 倍(输出 token 免费) • 为决策优化的前沿可组合智能 据我所知,这是通往 AI 驱动的经济革命的最短路径

    引用Diogo Almeida@CompleteSkeptic

    After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution

  2. Qwen66

    千问(Qwen)发布 Qwen-Image-2.1,并在 Hugging Face Spaces 上线可浏览器直接试用的演示。该模型为 7B 参数的图像生成与编辑一体模型,单一 checkpoint 同时支持两种任务,最多可用 10 张参考图,自带提示词增强 LLM,并集成 diffusers 与 ComfyUI。

    引用Hugging Apps@HuggingApps

    Qwen Image 2.1 is here! 🖼️ A 7B params native image generation and editing model, with up to 10 image references The model comes with it's own prompt enhancement LLMs, integrated with diffusers 🧨 and ComfyUI ▶️ on Spaces https://huggingface.co/spaces/hugging-apps/qwen-image-2-1

    推荐理由:原文给出了模型的参数量、参考图能力与免配置体验入口,读者可以直接在浏览器试用判断适用性。

9月20日周日
9月19日周六
9月17日周四
9月16日周三
  1. 蚂蚁 inclusionAI:HuggingFace 新模型62

    蚂蚁 inclusionAI 开源 Realtime-Venus 9B 全双工音视频交互系统

    蚂蚁 inclusionAI 发布 Realtime-Venus,一个支持主动音视频交互、异步委托和可打断全双工对话的开源系统,包含 Realtime-Venus-Omni 和 Realtime-Venus-Audio 两个 9B 检查点。

    推荐理由:模型卡给出了 Realtime-Venus 的架构组成、全双工与异步委托能力及完整用法,读者可以据此评估它在实时音视频交互场景的可用性。

  2. Jeff Dean42

    令人振奋的成果,@LiamFedus!祝贺 Periodic Labs 的整个团队!

    引用Liam Fedus@LiamFedus

    We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.