Liquid AI LFM 模型文档:模型系列、部署格式与量化选型指南
Liquid AI 官方文档介绍其 Liquid Foundation Models (LFM) 多模态模型家族,面向快速推理和端侧部署,统一支持 32K 上下文(LFM2.5-8B-A1B 为 128K)。
Liquid AI 官方文档介绍其 Liquid Foundation Models (LFM) 多模态模型家族,面向快速推理和端侧部署,统一支持 32K 上下文(LFM2.5-8B-A1B 为 128K)。
Liquid AI 发布开源桌面智能体 LocalCowork,展示 LFM2-24B-A2B 完全在本地笔记本上执行工具调用,无云端、无 API 密钥、数据不出设备。
EmbeddingGemma 是 Google 基于 Gemma3 架构的 3 亿参数文本嵌入模型,在 MTEB 基准上为不到 5 亿参数的嵌入模型中质量最高,支持 100 多种语言、输入最长 2k token。页面给出使用 baseten_performance_client 的完整代码示例,包括各任务的 prompt 前缀和通过 PerformanceClient 调用生成嵌入的方法。
推荐理由:指南围绕 Sonnet 5.5 与 Opus 5.5 的选型、迁移调参和 Claude Code 使用给出实操建议,便于开发者上手。
推荐理由:作者给出 Opus 5.5 相对 Opus 5 的价格降幅,并提供成本计算器,读者可自行估算 Claude Code 任务成本变化。
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai
推荐理由:原文给出本地推理提速的具体配置和硬件门槛,读者可以直接复用到自己的 GLM-5.3-Flash 部署中。
Sebastian Raschka 在 Ahead of AI 发表长文,解读 DeepSeek V3.2 相比 V3/R1 的主要变化。