Stable Diffusion 3.5 Large 上架 Hugging Face 并支持 Diffusers
Diffusers welcomes Stable Diffusion 3.5 Large
Stable Diffusion 3.5 Large(8B)及 timestep-distilled 版本现已上架 Hugging Face Hub,可通过 Diffusers 使用,蒸馏版可在 4-8 步内生成图像。
Diffusers 团队亲述 SD3.5 Large 的用法,覆盖量化推理和 24GB 显存训练 LoRA,方法可直接复用。
🧨Diffusers 迎来 Stable Diffusion 3.5 Large
发布于 2024 年 10 月 22 日
Apolinário from multimodal AI art multimodalart 关注
Aritra Roy Gosthipaty ariG23498 关注
Stable Diffusion 3.5 是其前代 Stable Diffusion 3 的改进版本。从今天起,这些模型已在 Hugging Face Hub 上提供,并可与 🧨Diffusers 一起使用。 本次发布包含 两个检查点:
- 一个大型(8B)模型
- 一个大型(8B)时间步蒸馏模型,支持少步推理
在本文中,我们将重点介绍如何将 Stable Diffusion 3.5(SD3.5)与 Diffusers 一起使用,涵盖推理和训练两方面。
目录
架构变更
SD3.5(large)的 transformer 架构与 SD3(medium)非常相似,有以下变化:
- QK 归一化:对于训练大型 transformer 模型,QK 归一化现已成为标准做法,SD3.5 Large 也不例外。
- 双重注意力层:SD3.5 不再为 MMDiT 块中的每个模态流使用单一注意力层,而是使用双重注意力层。
文本编码器、VAE 和噪声调度器方面的其余细节与 SD3 Medium 完全相同。有关 SD3 的更多信息,我们建议查看原始论文。
在 Diffusers 中使用 SD3.5
请确保安装最新版本的 diffusers:
pip install -U diffusers
由于该模型受门控限制,在将其与 diffusers 一起使用之前,您首先需要前往 Stable Diffusion 3.5 Large Hugging Face 页面,填写表单并接受门控。进入后,您需要登录,以便系统知道您已接受门控。使用以下命令登录:
huggingface-cli login
以下代码片段将下载 8B 参数版本的 SD3.5,精度为 torch.bfloat16。这是 Stability AI 发布的原始 checkpoint 所使用的格式,也是推荐的推理运行方式。
import torch
from diffusers import StableDiffusion3Pipeline
pipe = StableDiffusion3Pipeline.from_pretrained(
"stabilityai/stable-diffusion-3.5-large", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="a photo of a cat holding a sign that says hello world",
negative_prompt="",
num_inference_steps=40,
height=1024,
width=1024,
guidance_scale=4.5,
).images[0]
image.save("sd3_hello_world.png")
该版本还附带一个“timestep-distilled”模型,它消除了无分类器引导,让我们能用更少的步数生成图像(通常为 4-8 步)。
import torch
from diffusers import StableDiffusion3Pipeline
pipe = StableDiffusion3Pipeline.from_pretrained(
"stabilityai/stable-diffusion-3.5-large-turbo", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="a photo of a cat holding a sign that says hello world",
num_inference_steps=4,
height=1024,
width=1024,
guidance_scale=1.0,
).images[0]
image.save("sd3_hello_world.png")
我们在 SD3 博客文章和官方 Diffusers 文档中展示的所有示例应该已经可以在 SD3.5 上运行。特别是,这两份资源都深入探讨了如何优化运行推理所需的内存。由于 SD3.5 Large 比 SD3 Medium 大得多,内存优化对于在消费级设备上运行推理变得至关重要。
使用量化运行推理
Diffusers 原生支持使用 bitsandbytes 量化,这能进一步优化内存。
首先,请确保安装所有必要的库:
pip install -Uq git+https://github.com/huggingface/transformers@main
pip install -Uq bitsandbytes
然后以 “NF4”精度加载 transformer:
from diffusers import BitsAndBytesConfig, SD3Transformer2DModel
import torch
model_id = "stabilityai/stable-diffusion-3.5-large"
nf4_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16
)
model_nf4 = SD3Transformer2DModel.from_pretrained(
model_id,
subfolder="transformer",
quantization_config=nf4_config,
torch_dtype=torch.bfloat16
)
现在,我们可以运行推理了:
from diffusers import StableDiffusion3Pipeline
pipeline = StableDiffusion3Pipeline.from_pretrained(
model_id,
transformer=model_nf4,
torch_dtype=torch.bfloat16
)
pipeline.enable_model_cpu_offload()
prompt = "A whimsical and creative image depicting a hybrid creature that is a mix of a waffle and a hippopotamus, basking in a river of melted butter amidst a breakfast-themed landscape. It features the distinctive, bulky body shape of a hippo. However, instead of the usual grey skin, the creature's body resembles a golden-brown, crispy waffle fresh off the griddle. The skin is textured with the familiar grid pattern of a waffle, each square filled with a glistening sheen of syrup. The environment combines the natural habitat of a hippo with elements of a breakfast table setting, a river of warm, melted butter, with oversized utensils or plates peeking out from the lush, pancake-like foliage in the background, a towering pepper mill standing in for a tree. As the sun rises in this fantastical world, it casts a warm, buttery glow over the scene. The creature, content in its butter river, lets out a yawn. Nearby, a flock of birds take flight"
image = pipeline(
prompt=prompt,
negative_prompt="",
num_inference_steps=28,
guidance_scale=4.5,
max_sequence_length=512,
).images[0]
image.save("whimsical.png")
你可以在 BitsAndBytesConfig 中控制其他参数。详情请参阅文档。
也可以直接加载使用与上述相同 nf4_config 量化的模型。这对于内存较低的机器特别有帮助。端到端示例请参阅这个 Colab Notebook。
使用量化在 SD3.5 Large 上训练 LoRA
得益于 bitsandbytes 和 peft 等库,可以在拥有 24GB 显存的消费级 GPU 上微调 SD3.5 Large 这样的大模型。现在已经可以利用我们现有的 SD3 训练脚本来训练 LoRA。下面的训练命令已经可以正常工作:
accelerate launch train_dreambooth_lora_sd3.py \
--pretrained_model_name_or_path="stabilityai/stable-diffusion-3.5-large" \
--dataset_name="Norod78/Yarn-art-style" \
--output_dir="yart_art_sd3-5_lora" \
--mixed_precision="bf16" \
--instance_prompt="Frog, yarn art style" \
--caption_column="text"\
--resolution=768 \
--train_batch_size=1 \
--gradient_accumulation_steps=1 \
--learning_rate=4e-4 \
--report_to="wandb" \
--lr_scheduler="constant" \
--lr_warmup_steps=0 \
--max_train_steps=700 \
--rank=16 \
--seed="0" \
--push_to_hub
然而,要使其与量化配合工作,我们需要调整几个参数。下面我们提供相关指引:
- 我们用量化配置初始化
transformer,或者直接加载量化后的 checkpoint。 - 然后,我们使用来自
peft的prepare_model_for_kbit_training()来准备它。 - 得益于
peft对bitsandbytes的强大支持,其余流程保持不变!
更完整的示例请参阅这个示例脚本。
使用 Stable Diffusion 3.5 Transformer 的单文件加载
你可以使用 Stability AI 发布的原始 checkpoint 文件,通过 from_single_file 方法加载 Stable Diffusion 3.5 Transformer 模型:
import torch
from diffusers import SD3Transformer2DModel, StableDiffusion3Pipeline
transformer = SD3Transformer2DModel.from_single_file(
"https://huggingface.co/stabilityai/stable-diffusion-3.5-large-turbo/blob/main/sd3.5_large.safetensors",
torch_dtype=torch.bfloat16,
)
pipe = StableDiffusion3Pipeline.from_pretrained(
"stabilityai/stable-diffusion-3.5-large",
transformer=transformer,
torch_dtype=torch.bfloat16,
)
pipe.enable_model_cpu_offload()
image = pipe("a cat holding a sign that says hello world").images[0]
image.save("sd35.png")
重要链接
- Hub 上的 Stable Diffusion 3.5 Large 合集
- Stable Diffusion 3.5 的官方 Diffusers 文档
- 使用量化运行推理的 Colab Notebook
- 训练 LoRA
- Stable Diffusion 3 论文
- Stable Diffusion 3 博客文章
致谢:感谢 Daniel Frank 提供本博客文章缩略图所用的背景照片。感谢 Pedro Cuenca 和 Tom Aarsen 对文章草稿的审阅。
本文提到的模型 1
#### stabilityai/stable-diffusion-3.5-large 文本到图像 • 8B•更新于 2024年10月22日• 102k• 4.07k
本文提到的合集 1
Stable Diffusion 3.5 合集 6 个项目•更新于 2025年1月9日• 190
来自我们博客的更多文章
指南 diffusers 量化 ## 将 Nunchaku 4-bit 扩散推理引入 Diffusers *
*
69 2026年7月23日
peft lora 指南 ## 超越 LoRA:你能击败最流行的微调技术吗? *
*
*
*
103 2026年6月18日
社区
编辑预览
在文本输入框中拖拽、粘贴或点击此处来上传图像、音频和视频。
点击或粘贴此处上传图像
本文提及的模型 1
#### stabilityai/stable-diffusion-3.5-large 文本到图像 • 8B•更新于 2024年10月22日• 102k• 4.07k
本文提及的合集 1
Stable Diffusion 3.5 合集 6 个项目•更新于 2025年1月9日• 190
系统主题
公司
网站
来源:Hugging Face:Blog(RSS) · huggingface.co


