跳到正文

#图像生成

今日 16 条
9月23日周三
  1. ViggleAI50

    Viggle 发布面向开源社区的 Qwen-Image-2.1 turbo,名为 Viggle-Turbo,采用 DMD 蒸馏,可在 4 个采样步内完成生成和编辑,且无需 classifier-free guidance。权重已在 Hugging Face 开放,并提供 Spaces 在线体验;据称速度比完整模型快 6 倍。

    引用Hugging Apps@HuggingApps

    Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo

  2. Tencent Hy59

    Hy Image3.5 preview 现已在 ComfyUI 中可用。人工评测胜率相比 Hy Image3.0 提升 30%;单一模型支持文生图与图生图,最高 2K 分辨率,支持多语言文字与小字渲染,覆盖电影、漫画、商业摄影、插画风格,人物身份与产品特征可在场景、服装和风格切换中保持一致。

    引用ComfyUI@ComfyUI

    Hy Image3.5 preview is now available in ComfyUI. Professional-grade image generation, +30% win rate in human eval vs Hy Image3.0 → Text to image and Image to image in one model, up to 2K → Multilingual text, symbols, and small print that render correctly → Cinematic, comic, commercial photography, and illustration styles → Identity and product features that hold through scene, outfit, and style changes

  3. Qwen59

    Qwen 官方宣布 Qwen-Image-2.1 在 Arena 的 Image Edit 和 Text-to-Image 两个榜单均排名第一的开源模型。据 @Arena 引用内容,其在 Image Edit Arena 得分 1367,总排名第 16,距第 15 名 GPT-Image-1.5-high-fidelity 仅 3 分。

    引用Arena.ai@arena

    Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!

  4. Unsloth AI63

    千问(Qwen)发布开源图像生成与编辑模型 Qwen-Image-2.1,7B 参数,官方称基准表现与 Nano Banana 2.0 相当。Unsloth 发布 GGUF 量化版,支持 12GB 显存本地运行,也可通过 offloading 在 6GB 显存运行 Dynamic FP8;量化文件见 https://huggingface.co/unsloth/Qwen-Image-2.1-GGUF,指南见 https://unsloth.ai/docs/models/qwen-image-2.1。原模型统一支持生成与编辑,可原生生成和编辑 RGBA 透明图层,支持最多 10 张参考图,链接包括 https://qwen.ai/blog?id=qwen-image-2.1 和 https://github.com/QwenLM/Qwen-Image-2.1。

    引用Qwen@Alibaba_Qwen

    Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: https://qwen.ai/blog?id=qwen-image-2.1 - GitHub: https://github.com/QwenLM/Qwen-Image-2.1 - Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1 - Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

9月22日周二
9月21日周一
  1. Qwen42

    感谢 @sgl_project 的 day-0 支持!🙌 SGLang-Diffusion 现已支持 Qwen-Image-2.1:文生图、多图编辑,以及透明 RGBA 输出。快来试试!🎨

    引用SGLang@sgl_project

    Day-0 support for @Alibaba_Qwen’s Qwen-Image 2.1 is here in SGLang-Diffusion! 🖥️ Native precision on a single RTX 4090 24GB with CPU offload - 1024×1024 generation in 18.7s and image editing in 21.7s with 22.7 GiB peak GPU memory during requests. - On an RTX PRO 6000 96GB: 8.0s generation and 9.6s editing. 🎨 Text-to-image, multi-image editing, and transparent RGBA output—all with one checkpoint. ⚡ Native inference with TP/SP, LoRA, and OpenAI-compatible APIs. 40 denoising steps, one image per request, warmed HTTP latency including PNG output. No quantization. Cookbook and GPU-specific commands below 👇

  2. Qwen66

    千问(Qwen)发布 Qwen-Image-2.1,并在 Hugging Face Spaces 上线可浏览器直接试用的演示。该模型为 7B 参数的图像生成与编辑一体模型,单一 checkpoint 同时支持两种任务,最多可用 10 张参考图,自带提示词增强 LLM,并集成 diffusers 与 ComfyUI。

    引用Hugging Apps@HuggingApps

    Qwen Image 2.1 is here! 🖼️ A 7B params native image generation and editing model, with up to 10 image references The model comes with it's own prompt enhancement LLMs, integrated with diffusers 🧨 and ComfyUI ▶️ on Spaces https://huggingface.co/spaces/hugging-apps/qwen-image-2-1

    推荐理由:原文给出了模型的参数量、参考图能力与免配置体验入口,读者可以直接在浏览器试用判断适用性。

9月20日周日
  1. Qwen65

    Qwen 官方宣布 Qwen-Image-2.1 已获得 ComfyUI 支持,开放权重,单一 7B checkpoint 同时支持生成和编辑。引用内容提到该模型支持原生 2K 图像生成、单次最多基于 10 张参考图进行指令编辑,以及带 alpha 通道的 RGBA 输出。

    引用ComfyUI@ComfyUI

    Qwen-Image-2.1 is now supported in ComfyUI! Open weights. One 7B checkpoint that generates and edits. → Image generation at native 2K → Instruction editing from up to 10 reference images in a single pass → RGBA output, alpha included

    推荐理由:正文点出该模型已在 ComfyUI 支持,读者可以据此更新本地图像生成与编辑工作流。

9月19日周六
  1. Grok55

    Grok 官方推广其 Imagine 功能,称用户可以用 Grok Imagine 制作自己的电影场景。其引用的案例称 PJaccetturo 及其工作室在 9 天内制作了一个试播集,用了 3,627 张图像、3,043 个视频,生成成本 $2,677。

    引用Grok Imagine@imagine

    Hollywood would spend 7 figures on this Odyssey scene. @PJaccetturo and his studio made this pilot in 9 days, with 3,627 images, 3,043 videos, and generations costing $2,677. Here’s their breakdown of how they did it 🧵

9月18日周五
9月17日周四
9月16日周三
9月15日周二
9月14日周一
9月11日周五
9月10日周四
  1. Hugging Face:Blog(RSS)61

    Hugging Face 用 Gradio Workflow 重建 AUTOMATIC1111,推出 Workflow1111

    Hugging Face 发布 Workflow1111,用 gr.Workflow 在单个画布上以 73 个节点重建了 AUTOMATIC1111 的 11 条媒体管线,涵盖文生图、hi-resolution fix、图生图、VLM 反推提示词、检测生成 inpaint 蒙版、ControlNet 风格 annotator、背景去除、PNG Info 和图生视频。

    推荐理由:官方用 Gradio Workflow 在单个画布上复刻了 AUTOMATIC1111 的主要功能,读者可以对照它了解节点式工作流与 ComfyUI 的差异。

9月8日周二
9月4日周五
9月1日周二
8月29日周六
8月28日周五
  1. Midjourney:Updates(RSS)65

    Midjourney 开放测试首个 V8.2 图像编辑模型

    Midjourney 开始让所有用户测试首个 V8.2 图像编辑模型。该模型支持用指令编辑图像、以最多 4 张图像参考生成新图(替代 omni-reference)、局部重绘(inpainting)与扩图(outpainting),并可用 personalization、moodboards 和 srefs。

    推荐理由:官方宣布 V8.2 图像编辑模型开放测试,列出指令编辑、多图参考、局部重绘等能力变化和具体入口。

8月22日周六
8月19日周三