跳到正文
原文
Hugging Face:Blog(RSS)·· 2024-10-22精选AI 评分78

Stable Diffusion 3.5 Large 上架 Hugging Face 并支持 Diffusers

Diffusers welcomes Stable Diffusion 3.5 Large

AI 导读

Stable Diffusion 3.5 Large(8B)及 timestep-distilled 版本现已上架 Hugging Face Hub,可通过 Diffusers 使用,蒸馏版可在 4-8 步内生成图像。

推荐理由

Diffusers 团队亲述 SD3.5 Large 的用法,覆盖量化推理和 24GB 显存训练 LoRA,方法可直接复用。

正文 · AI 翻译

Image 1: Hugging Face's logoHugging Face


返回文章

🧨Diffusers 迎来 Stable Diffusion 3.5 Large

发布于 2024 年 10 月 22 日

在 GitHub 上更新

- [x] 点赞 55

  • Image 2
  • Image 3
  • Image 4
  • Image 5
  • Image 6
  • Image 7
  • +49

Image 8: YiYi Xu's avatar

YiYi Xu YiYiXu 关注

Image 9: Aryan V S's avatar

Aryan V S a-r-r-o-w 关注

Image 10: Dhruv Nair's avatar

Dhruv Nair dn6 关注

Image 11: Sayak Paul's avatar

Sayak Paul sayakpaul 关注

Image 12: Linoy Tsaban's avatar

Linoy Tsaban linoyts 关注

Image 13: Apolinário from multimodal AI art's avatar

Apolinário from multimodal AI art multimodalart 关注

Image 14: Alvaro Somoza's avatar

Alvaro Somoza OzzyGT 关注

Image 15: Aritra Roy Gosthipaty's avatar

Aritra Roy Gosthipaty ariG23498 关注

Stable Diffusion 3.5 是其前代 Stable Diffusion 3 的改进版本。从今天起,这些模型已在 Hugging Face Hub 上提供,并可与 🧨Diffusers 一起使用。 本次发布包含 两个检查点:

  • 一个大型(8B)模型
  • 一个大型(8B)时间步蒸馏模型,支持少步推理

在本文中,我们将重点介绍如何将 Stable Diffusion 3.5(SD3.5)与 Diffusers 一起使用,涵盖推理和训练两方面。

目录

架构变更

SD3.5(large)的 transformer 架构与 SD3(medium)非常相似,有以下变化:

  • QK 归一化:对于训练大型 transformer 模型,QK 归一化现已成为标准做法,SD3.5 Large 也不例外。
  • 双重注意力层:SD3.5 不再为 MMDiT 块中的每个模态流使用单一注意力层,而是使用双重注意力层。

文本编码器、VAE 和噪声调度器方面的其余细节与 SD3 Medium 完全相同。有关 SD3 的更多信息,我们建议查看原始论文。

在 Diffusers 中使用 SD3.5

请确保安装最新版本的 diffusers:

pip install -U diffusers

由于该模型受门控限制,在将其与 diffusers 一起使用之前,您首先需要前往 Stable Diffusion 3.5 Large Hugging Face 页面,填写表单并接受门控。进入后,您需要登录,以便系统知道您已接受门控。使用以下命令登录:

huggingface-cli login

以下代码片段将下载 8B 参数版本的 SD3.5,精度为 torch.bfloat16。这是 Stability AI 发布的原始 checkpoint 所使用的格式,也是推荐的推理运行方式。

import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_pretrained(
    "stabilityai/stable-diffusion-3.5-large", torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt="a photo of a cat holding a sign that says hello world",
    negative_prompt="",
    num_inference_steps=40,
    height=1024,
    width=1024,
    guidance_scale=4.5,
).images[0]

image.save("sd3_hello_world.png")

Image 16: hello_world_cat

该版本还附带一个“timestep-distilled”模型,它消除了无分类器引导,让我们能用更少的步数生成图像(通常为 4-8 步)。

import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_pretrained(
    "stabilityai/stable-diffusion-3.5-large-turbo", torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt="a photo of a cat holding a sign that says hello world",
    num_inference_steps=4,
    height=1024,
    width=1024,
    guidance_scale=1.0,
).images[0]

image.save("sd3_hello_world.png")

Image 17: hello_world_cat_2

我们在 SD3 博客文章和官方 Diffusers 文档中展示的所有示例应该已经可以在 SD3.5 上运行。特别是,这两份资源都深入探讨了如何优化运行推理所需的内存。由于 SD3.5 Large 比 SD3 Medium 大得多,内存优化对于在消费级设备上运行推理变得至关重要。

使用量化运行推理

Diffusers 原生支持使用 bitsandbytes 量化,这能进一步优化内存。

首先,请确保安装所有必要的库:

pip install -Uq git+https://github.com/huggingface/transformers@main
pip install -Uq bitsandbytes

然后以 “NF4”精度加载 transformer:

from diffusers import BitsAndBytesConfig, SD3Transformer2DModel
import torch

model_id = "stabilityai/stable-diffusion-3.5-large"
nf4_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16
)
model_nf4 = SD3Transformer2DModel.from_pretrained(
    model_id,
    subfolder="transformer",
    quantization_config=nf4_config,
    torch_dtype=torch.bfloat16
)

现在,我们可以运行推理了:

from diffusers import StableDiffusion3Pipeline

pipeline = StableDiffusion3Pipeline.from_pretrained(
    model_id, 
    transformer=model_nf4,
    torch_dtype=torch.bfloat16
)
pipeline.enable_model_cpu_offload()

prompt = "A whimsical and creative image depicting a hybrid creature that is a mix of a waffle and a hippopotamus, basking in a river of melted butter amidst a breakfast-themed landscape. It features the distinctive, bulky body shape of a hippo. However, instead of the usual grey skin, the creature's body resembles a golden-brown, crispy waffle fresh off the griddle. The skin is textured with the familiar grid pattern of a waffle, each square filled with a glistening sheen of syrup. The environment combines the natural habitat of a hippo with elements of a breakfast table setting, a river of warm, melted butter, with oversized utensils or plates peeking out from the lush, pancake-like foliage in the background, a towering pepper mill standing in for a tree.  As the sun rises in this fantastical world, it casts a warm, buttery glow over the scene. The creature, content in its butter river, lets out a yawn. Nearby, a flock of birds take flight"
image = pipeline(
    prompt=prompt,
    negative_prompt="",
    num_inference_steps=28,
    guidance_scale=4.5,
    max_sequence_length=512,
).images[0]
image.save("whimsical.png")

Image 18: happy_hippo

你可以在 BitsAndBytesConfig 中控制其他参数。详情请参阅文档。

也可以直接加载使用与上述相同 nf4_config 量化的模型。这对于内存较低的机器特别有帮助。端到端示例请参阅这个 Colab Notebook。

使用量化在 SD3.5 Large 上训练 LoRA

得益于 bitsandbytes 和 peft 等库,可以在拥有 24GB 显存的消费级 GPU 上微调 SD3.5 Large 这样的大模型。现在已经可以利用我们现有的 SD3 训练脚本来训练 LoRA。下面的训练命令已经可以正常工作:

accelerate launch train_dreambooth_lora_sd3.py \
  --pretrained_model_name_or_path="stabilityai/stable-diffusion-3.5-large"  \
  --dataset_name="Norod78/Yarn-art-style" \
  --output_dir="yart_art_sd3-5_lora" \
  --mixed_precision="bf16" \
  --instance_prompt="Frog, yarn art style" \
  --caption_column="text"\
  --resolution=768 \
  --train_batch_size=1 \
  --gradient_accumulation_steps=1 \
  --learning_rate=4e-4 \
  --report_to="wandb" \
  --lr_scheduler="constant" \
  --lr_warmup_steps=0 \
  --max_train_steps=700 \
  --rank=16 \
  --seed="0" \
  --push_to_hub

然而,要使其与量化配合工作,我们需要调整几个参数。下面我们提供相关指引:

  • 我们用量化配置初始化 transformer,或者直接加载量化后的 checkpoint。
  • 然后,我们使用来自 peft 的 prepare_model_for_kbit_training() 来准备它。
  • 得益于 peft 对 bitsandbytes 的强大支持,其余流程保持不变!

更完整的示例请参阅这个示例脚本。

使用 Stable Diffusion 3.5 Transformer 的单文件加载

你可以使用 Stability AI 发布的原始 checkpoint 文件,通过 from_single_file 方法加载 Stable Diffusion 3.5 Transformer 模型:

import torch
from diffusers import SD3Transformer2DModel, StableDiffusion3Pipeline

transformer = SD3Transformer2DModel.from_single_file(
    "https://huggingface.co/stabilityai/stable-diffusion-3.5-large-turbo/blob/main/sd3.5_large.safetensors",
    torch_dtype=torch.bfloat16,
)
pipe = StableDiffusion3Pipeline.from_pretrained(
    "stabilityai/stable-diffusion-3.5-large",
    transformer=transformer,
    torch_dtype=torch.bfloat16,
)
pipe.enable_model_cpu_offload()
image = pipe("a cat holding a sign that says hello world").images[0]
image.save("sd35.png")

重要链接

致谢:感谢 Daniel Frank 提供本博客文章缩略图所用的背景照片。感谢 Pedro Cuenca 和 Tom Aarsen 对文章草稿的审阅。

本文提到的模型 1

Image 19 #### stabilityai/stable-diffusion-3.5-large 文本到图像 • 8B•更新于 2024年10月22日• 102k• 4.07k

本文提到的合集 1

Stable Diffusion 3.5 合集 6 个项目•更新于 2025年1月9日• 190

来自我们博客的更多文章

Image 20 指南 diffusers 量化 ## 将 Nunchaku 4-bit 扩散推理引入 Diffusers * Image 21 * Image 22 69 2026年7月23日

Image 23 peft lora 指南 ## 超越 LoRA:你能击败最流行的微调技术吗? * Image 24 * Image 25 * Image 26 * Image 27 103 2026年6月18日

社区

编辑预览

在文本输入框中拖拽、粘贴或点击此处来上传图像、音频和视频。

点击或粘贴此处上传图像

评论 ·注册或登录以发表评论

- [x] 点赞 55

  • Image 28
  • Image 29
  • Image 30
  • Image 31
  • Image 32
  • Image 33
  • Image 34
  • Image 35
  • Image 36
  • Image 37
  • Image 38
  • Image 39
  • +43

本文提及的模型 1

Image 40 #### stabilityai/stable-diffusion-3.5-large 文本到图像 • 8B•更新于 2024年10月22日• 102k• 4.07k

本文提及的合集 1

Stable Diffusion 3.5 合集 6 个项目•更新于 2025年1月9日• 190

系统主题

公司

服务条款隐私关于招聘

网站

模型数据集空间定价文档

来源:Hugging Face:Blog(RSS) · huggingface.co