跳到正文
原文
Unsloth AI· @UnslothAI · X·· 27 天前精选AI 评分71
AI 导读

Unsloth 发布 GLM-5.3-Flash 的本地 GGUF 优化方案,通过更快解码和 MTP 支持使本地推理提速 1.6 至 3.4 倍,长上下文下最高达 3.3 倍。3-bit 量化可在 128GB 内存设备上通过 Unsloth Desktop 或 llama.cpp 运行,GGUF 权重和指南已发布,开箱即用且无需额外模块或 MTP 文件。

推荐理由

原文给出本地推理提速的具体配置和硬件门槛,读者可以直接复用到自己的 GLM-5.3-Flash 部署中。

正文

We made GLM-5.3-Flash run 3.3x faster locally!

Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction.

Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp.

Guide: https://unsloth.ai/docs/models/glm-5.3-flash#faster-inference-and-mtp-support
GGUF: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF

引用Z.ai@Zai_org
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai
在 X 查看被引用的帖子

来源:Unsloth AI · x.com