跳到正文

#图像生成

今日 3 条
今天10月1日周四
  1. Arena.ai47

    Hidream-O1-Video-1.0 由 @HiDream_AI 刚刚登陆 Image-to-Video Arena,以 1456 分位列第 7! 该模型已在 @vivago_ai 上线,距 gemini-omni-flash 仅差 8 分,距 dreamina-seedance-2.0 和 2.5 均在 20 分以内。 恭喜 @HiDream_AI 发布!

    引用vivago.ai (HiDream)@vivago_ai

    New HiDream models just landed in vivago R1 Studio 🚀 Introducing: • HiDream-O1 Image 2.0 • HiDream-O1 Editing 1.5 • HiDream-O1 Video All three models are now available in vivago R1 Studio - bringing the latest HiDream image generation, editing, and video capabilities directly into your creative workflow. New models. New possibilities. Go make something the internet can’t ignore. 🔥

9月30日周三
  1. Black Forest Labs:Blog(网页)71

    Black Forest Labs 发布多模态基础模型 FLUX 3 并开放 Early Access

    Black Forest Labs 发布多模态基础模型 FLUX 3,基于 Self-Flow 方法在统一架构内联合学习图像、视频和音频,现已开放 Early Access。

    推荐理由:原文给出 FLUX 3 的多模态架构、早期对比评测数字和分阶段开放计划,读者可据此评估其对创作与具身场景的适用性。

9月28日周一
9月26日周六
9月24日周四
9月23日周三
  1. ViggleAI50

    Viggle 发布面向开源社区的 Qwen-Image-2.1 turbo,名为 Viggle-Turbo,采用 DMD 蒸馏,可在 4 个采样步内完成生成和编辑,且无需 classifier-free guidance。权重已在 Hugging Face 开放,并提供 Spaces 在线体验;据称速度比完整模型快 6 倍。

    引用Hugging Apps@HuggingApps

    Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo

  2. Qwen59

    Qwen 官方宣布 Qwen-Image-2.1 在 Arena 的 Image Edit 和 Text-to-Image 两个榜单均排名第一的开源模型。据 @Arena 引用内容,其在 Image Edit Arena 得分 1367,总排名第 16,距第 15 名 GPT-Image-1.5-high-fidelity 仅 3 分。

    引用Arena.ai@arena

    Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!

  3. Unsloth AI63

    千问(Qwen)发布开源图像生成与编辑模型 Qwen-Image-2.1,7B 参数,官方称基准表现与 Nano Banana 2.0 相当。Unsloth 发布 GGUF 量化版,支持 12GB 显存本地运行,也可通过 offloading 在 6GB 显存运行 Dynamic FP8;量化文件见 https://huggingface.co/unsloth/Qwen-Image-2.1-GGUF,指南见 https://unsloth.ai/docs/models/qwen-image-2.1。原模型统一支持生成与编辑,可原生生成和编辑 RGBA 透明图层,支持最多 10 张参考图,链接包括 https://qwen.ai/blog?id=qwen-image-2.1 和 https://github.com/QwenLM/Qwen-Image-2.1。

    引用Qwen@Alibaba_Qwen

    Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: https://qwen.ai/blog?id=qwen-image-2.1 - GitHub: https://github.com/QwenLM/Qwen-Image-2.1 - Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1 - Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

9月22日周二
9月21日周一
  1. Qwen66

    千问(Qwen)发布 Qwen-Image-2.1,并在 Hugging Face Spaces 上线可浏览器直接试用的演示。该模型为 7B 参数的图像生成与编辑一体模型,单一 checkpoint 同时支持两种任务,最多可用 10 张参考图,自带提示词增强 LLM,并集成 diffusers 与 ComfyUI。

    引用Hugging Apps@HuggingApps

    Qwen Image 2.1 is here! 🖼️ A 7B params native image generation and editing model, with up to 10 image references The model comes with it's own prompt enhancement LLMs, integrated with diffusers 🧨 and ComfyUI ▶️ on Spaces https://huggingface.co/spaces/hugging-apps/qwen-image-2-1

    推荐理由:原文给出了模型的参数量、参考图能力与免配置体验入口,读者可以直接在浏览器试用判断适用性。

9月19日周六
9月17日周四
9月14日周一
9月1日周二
8月17日周一
7月25日周六
  1. Midjourney:Updates(RSS)47

    Midjourney 发布 V8.2 图像模型

    Midjourney 推出 V8.2 图像模型,本次更新聚焦美学、图像质量与个性化。官方称图像将更具创意、大胆、精致、前卫和新颖,低质量图像的随机情况应会大幅减少。V8.2 的个性化档案拥有更大且更优的图像池,能更好理解用户个人品味,尤其适合在个人档案下积累了大量评分的用户。

6月11日周四
  1. Midjourney:Updates(RSS)58

    Midjourney 将默认模型从 V7 切换为 V8.1

    Midjourney 宣布经过测试和反馈后,默认模型已从 V7 更新为 V8.1。V8.1 更聪明、更连贯,更好地遵循详细提示词并渲染文字;开启 HD 模式后图像尺寸为 V7 的两倍、分辨率 4 倍,SD 模式 4 秒、HD 模式 12 秒出图。风格参考、个性化与美学在 V7 与 V8.1 间保持一致;V7 的 omni-reference 仍可使用,V8.0 alpha 将在两周后弃用。

5月22日周五
  1. Saining Xie45

    来看看 Jas 主导的 RAEv2。通过大量实验,我们发现了一些非常有趣的行为,说明为什么强大的表征编码器对像素解码器至关重要。 剧透:关键不在于爬 FID 的坡;像 ep@fid-k/fdr^k 这样的新指标表明,还有更多值得探索的空间!

    引用Jaskirat Singh@1jaskiratsingh

    In Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation. Today we make them a bit more simple. a bit more general. Result: >10x faster convergence, better reconstruction, better generation. And yes we test them on T2I and world models :) Introducing RAEv2

5月7日周四
3月25日周三
2月11日周三
2月3日周二
12月18日周四
12月5日周五
  1. 美团 LongCat:HuggingFace 新模型54

    美团 LongCat 发布图像编辑模型 LongCat-Image-Edit

    美团 LongCat 发布图像编辑模型 LongCat-Image-Edit,为 Longcat-Image 的编辑版本,支持中英双语编辑。官方称其在开源图像编辑模型中达到 SOTA,支持全局编辑、局部编辑、文本修改和参考引导编辑,能保持未编辑区域的布局、纹理、色调和主体身份一致,适合多轮编辑。

12月4日周四
10月25日周六
  1. 美团 LongCat:HuggingFace 新模型61

    美团 LongCat 发布 13.6B 开源视频生成模型 LongCat-Video

    美团 LongCat 团队发布 13.6B 参数的开源视频生成基础模型 LongCat-Video,统一支持文生视频、图生视频和视频续写三类任务。模型在 Video-Continuation 上预训练,可生成长达数分钟的视频而不出现色彩漂移或质量下降,并通过粗到细策略和 Block Sparse Attention 在数分钟内生成 720p、30fps 视频。

8月3日周日
9月2日周一
8月21日周三