跳到正文

全部动态

今日 46 条
7月16日周四
  1. Lilian Weng71

    Thinking Machines 推出开放权重模型 Inkling,作者 Lilian Weng 称其定位为在广泛能力类别上表现扎实、便于实际使用和定制的基础模型。Inkling 可跨文本、图像和音频模态高效推理,完整权重已开放下载,今日起可在 Tinker 上微调,并在 Inkling Playground 中体验。

    引用Thinking Machines@thinkymachines

    Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵

  2. Mira Murati72

    Mira Murati 宣布 Thinking Machines 的首个模型 Inkling,从零训练且权重开放。Inkling 可跨文本、图像和音频模态高效推理,官方放出全部权重,今天起可在 Tinker 上微调,并可在 Inkling Playground 体验。

    引用Thinking Machines@thinkymachines

    Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵

    推荐理由:Thinking Machines 发布首个模型 Inkling,开放全部权重并支持 Tinker 微调,可关注其跨模态推理的落地方式。

7月15日周三
  1. Hugging Face:Blog(RSS)81

    Thinking Machines 发布 1T 参数多模态开源模型 Inkling 及 Inkling-Small

    Thinking Machines 在 Hugging Face 上发布开源多模态模型 Inkling,共 975B 总参数、41B 激活参数,支持 1M 上下文,原生接收图像、文本和音频输入,训练数据为 45 万亿 token,并同时发布 276B 总参数、12B 激活参数的 Inkling-Small。

    推荐理由:文章给出了 Inkling 的架构细节、各档位 VRAM 需求和多条部署路径,便于读者评估在自己的硬件上运行哪种变体。

7月8日周三
7月5日周日
7月3日周五
7月2日周四
7月1日周三
6月30日周二
6月26日周五
6月11日周四
  1. Midjourney:Updates(RSS)58

    Midjourney 将默认模型从 V7 切换为 V8.1

    Midjourney 宣布经过测试和反馈后,默认模型已从 V7 更新为 V8.1。V8.1 更聪明、更连贯,更好地遵循详细提示词并渲染文字;开启 HD 模式后图像尺寸为 V7 的两倍、分辨率 4 倍,SD 模式 4 秒、HD 模式 12 秒出图。风格参考、个性化与美学在 V7 与 V8.1 间保持一致;V7 的 omni-reference 仍可使用,V8.0 alpha 将在两周后弃用。

  2. Google DeepMind:Blog(RSS)76

    Google DeepMind 发布开源实验模型 DiffusionGemma,GPU 文本生成最高提速 4 倍

    Google DeepMind 发布实验性开源模型 DiffusionGemma,基于文本扩散方法,在专用 GPU 上实现最高 4 倍的文本生成提速。该模型为 26B 总参数的 MoE,推理时仅激活 3.8B 参数,量化后可装入 18GB 显存,单张 NVIDIA H100 上超过 1000 tokens 每秒,RTX 5090 上超过 700 tokens 每秒。

    推荐理由:官方给出具体吞吐数字、显存占用和质量取舍,读者可据此判断它适不适合本地交互式工作流。

6月10日周三
6月9日周二
  1. Google DeepMind:Blog(RSS)77

    Google DeepMind 发布 Gemma 4 12B:无编码器统一多模态模型,可在 16GB 内存笔记本本地运行

    Google DeepMind 发布 Gemma 4 12B,定位于在笔记本上本地运行的智能体多模态模型,性能接近其 26B MoE 模型而内存占用不到一半。

    推荐理由:官方介绍了无编码器统一架构和 16GB 内存本机运行等细节,适合关注端侧多模态部署的开发者参考。

6月4日周四
  1. vLLM 官方博客(RSS)67

    NVIDIA Nemotron 3 Ultra 在 vLLM 获 Day-0 支持

    NVIDIA 宣布 Nemotron 3 Ultra 在 vLLM 上获得 Day-0 支持。该开源模型采用混合 Transformer-Mamba 潜在 MoE 架构,总参数 550B、激活 55B,上下文最长 1M tokens,支持 NVFP4 与 BF16 精度,面向长时程自主智能体工作流。

    推荐理由:官方给出架构细节、显存配置和 vLLM 部署命令,读者可以直接照此在自有环境里跑通该模型。

5月28日周四
5月27日周三
5月23日周六
5月22日周五
  1. Saining Xie45

    来看看 Jas 主导的 RAEv2。通过大量实验,我们发现了一些非常有趣的行为,说明为什么强大的表征编码器对像素解码器至关重要。 剧透:关键不在于爬 FID 的坡;像 ep@fid-k/fdr^k 这样的新指标表明,还有更多值得探索的空间!

    引用Jaskirat Singh@1jaskiratsingh

    In Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation. Today we make them a bit more simple. a bit more general. Result: >10x faster convergence, better reconstruction, better generation. And yes we test them on T2I and world models :) Introducing RAEv2

5月21日周四
5月16日周六
  1. Google DeepMind:Blog(RSS)85

    Google DeepMind 发布 Gemini 3.5,率先推出 3.5 Flash

    Google DeepMind 发布 Gemini 3.5 模型系列,首发的 3.5 Flash 已在 Gemini 应用、Search AI Mode、Google Antigravity、Gemini API 和 Gemini Enterprise 上线。

    推荐理由:官方发布稿给出了具体基准成绩、速度倍数和开放渠道,读者可以据此评估 3.5 Flash 在智能体和编码场景的可用性。

5月12日周二
  1. Lilian Weng39

    过去几个月,我们经历了许多乐趣(和压力😅),在训练运行日志中产出了 12 个版本(以及许多子版本)和 137 页内容。 事实证明,人与人之间的协作对于改善人机协作很重要。😊

    引用Thinking Machines@thinkymachines

    People talk, listen, watch, think, and collaborate at the same time, in real time. We've designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action. https://thinkingmachines.ai/blog/interaction-models

5月11日周一
5月7日周四
4月28日周二
  1. Fuli Luo65

    小米 MiMo 团队开源两款模型:MiMo-V2.5-Pro(Code Agent,总参数 1T)和 MiMo-V2.5(多模态 Agent,总参数 310B)。采用 MIT 许可证,支持商用部署、继续训练和微调,两款模型均支持 1M token 上下文窗口;官方称 MiMo-V2.5-Pro 在开源模型中 GDPVal-AA 和 ClawEval 排名第一。开发者还可申请 100T 免费 token(http://100t.xiaomimimo.com),权重在 Hugging Face 提供。

    引用Xiaomi MiMo@XiaomiMiMo

    Xiaomi MiMo-V2.5 is now officially open-sourced! MIT License, supporting commercial deployment, continued training, and fine-tuning - no additional authorization required. Two models, both supporting a 1M-token context window : • MiMo-V2.5-Pro: built for complex agent and coding tasks, ranking No.1 among open-source models on GDPVal-AA and ClawEval • MiMo-V2.5: a native omni-modal model with strong agent capabilities A model's value isn't measured by rankings alone — it's measured by the problems it solves. Let's build with MiMo now! 🤗 Weights: https://huggingface.co/collections/XiaomiMiMo/mimo-v25 📄 Blog: https://mimo.xiaomi.com/index#blog

4月24日周五
4月23日周四
4月22日周三
4月17日周五