New HiDream models just landed in vivago R1 Studio 🚀 Introducing: • HiDream-O1 Image 2.0 • HiDream-O1 Editing 1.5 • HiDream-O1 Video All three models are now available in vivago R1 Studio - bringing the latest HiDream image generation, editing, and video capabilities directly into your creative workflow. New models. New possibilities. Go make something the internet can’t ignore. 🔥
#图像生成
#图像生成
今日 3 条
Arena.ai@arenaAI 评分4747引用vivago.ai (HiDream)@vivago_ai
🚨 AI News | TestingCatalog@testingcatalogAI 评分5454
Arena.ai@arenaAI 评分5959Ideogram 4.5 进入 Image Edit Arena 前 20,总分 1351 分,排名第 18。
引用Ideogram@ideogram_aiIntroducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon.
Suno:Blog(网页)精选AI 评分6060 Suno 发布 v5.5,推出 Voices、Custom models 和 My Taste
Suno 发布 v5.5 模型,同时推出 Voices、Custom models 和 My Taste 三项个性化功能。
推荐理由:原文给出 Voices 的验证与隐私细节和两项新个性化功能,读者可以判断这些能力如何改变自己的创作流程。
World Labs:官网AI 评分6969 World Labs 发布下一代世界模型 Atlas
World Labs 发布下一代世界模型 Atlas,一个从零预训练的全能模型,可原生处理文本、图像、视频和 3D。
Meta AI:Blog(网页)AI 评分7474 Meta 发布 Muse Image 并预览 Muse Video
Meta Superintelligence Labs 发布 Muse Image 并预览 Muse Video,均为该实验室首批媒体生成模型。
Black Forest Labs:Blog(网页)精选AI 评分7171 Black Forest Labs 发布多模态基础模型 FLUX 3 并开放 Early Access
Black Forest Labs 发布多模态基础模型 FLUX 3,基于 Self-Flow 方法在统一架构内联合学习图像、视频和音频,现已开放 Early Access。
推荐理由:原文给出 FLUX 3 的多模态架构、早期对比评测数字和分阶段开放计划,读者可据此评估其对创作与具身场景的适用性。
Artificial Analysis@ArtificialAnlysAI 评分5656
ModelScope@ModelScope2022AI 评分5959ModelScope 发布 DiffSynth-Studio 下的 Qwen-Image-2.1 LayerExtract 和 LayerRemove 两个 LoRA,组成图层编辑工作流。
AK@_akhaliqAI 评分5151
ViggleAI@ViggleAIAI 评分4343引用Yun Chen@t_muxviggle-turbo for Qwen-Image-2.1 isn't just faster — for most prompts, it's just as good as base.
Tencent Hy@TencentHunyuanAI 评分3030引用GMI Cloud@gmi_cloudHy Image 3.5 preview is now live on GMI Cloud @TencentHunyuan $0.024 per image 8.8x cheaper than GPT Image 2 5.6x cheaper than Nano Banana Pro Create now 👇
Ant Ling@AntLingAGIAI 评分5050通过我们的 Day0 合作伙伴 @novita_labs,畅享市面上最佳的文本到图像 UI/UX 设计模型 🤠🥰
引用Novita AI@novita_labsMing-Image-0.1-Design from @AntLingAGI is now available via Novita on @OpenRouter. Launching as a Day-0 partner ⚡ 🎁 Free for 14 days. A text-to-image model built for graphic-design output and legible text rendering.
ViggleAI@ViggleAIAI 评分5050引用Hugging Apps@HuggingAppsQwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo
Qwen@Alibaba_QwenAI 评分5959引用Arena.ai@arenaQwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!
Ant Ling@AntLingAGIAI 评分4949
Unsloth AI@UnslothAIAI 评分6363引用Qwen@Alibaba_QwenMeet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: https://qwen.ai/blog?id=qwen-image-2.1 - GitHub: https://github.com/QwenLM/Qwen-Image-2.1 - Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1 - Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1
Qwen@Alibaba_QwenAI 评分3535来自 @inteldevs 的 Day-0 OpenVINO 支持!🥳 Qwen-Image-2.1 已可在 Intel 硬件上优化运行。一个开放权重 checkpoint,同时支持生成与编辑。👇
引用Intel Devs@inteldevsWe're excited to offer Day0 OpenVINO support for Qwen-Image-2.1 Read more about what you can accomplish here: https://ms.spr.ly/6019a5nm5
Qwen@Alibaba_Qwen精选AI 评分6666引用Hugging Apps@HuggingAppsQwen Image 2.1 is here! 🖼️ A 7B params native image generation and editing model, with up to 10 image references The model comes with it's own prompt enhancement LLMs, integrated with diffusers 🧨 and ComfyUI ▶️ on Spaces https://huggingface.co/spaces/hugging-apps/qwen-image-2-1
推荐理由:原文给出了模型的参数量、参考图能力与免配置体验入口,读者可以直接在浏览器试用判断适用性。
Mustafa Suleyman@mustafasuleymanAI 评分4949引用Artificial Analysis@ArtificialAnlysGenerating high-quality images is cheaper and faster than ever. Muse Image, MAI-Image-2.6 and GPT Images 2.5 have substantially shifted the Text to Image Pareto frontiers for both price and speed in recent weeks.
SenseTime@SenseTime_AIAI 评分5353商汤发布 SenseNova U1.5 技术报告,这是一个开源的 8B-MoT 原生统一模型,通过共享注意力连接理解与生成。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分4949 蚂蚁 inclusionAI 发布 Ming-Image-0.1-Design-Layer 设计图层分解模型
蚂蚁 inclusionAI 在 Hugging Face 发布 Ming-Image-0.1-Design-Layer,可将扁平化设计图按指定层数拆解为 RGBA 图层,采用 MIT 许可。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分5151 蚂蚁 inclusionAI 发布 Ming-Image-0.1-Design 文生图模型
蚂蚁 inclusionAI 在 Hugging Face 发布 Ming-Image-0.1-Design,是一个面向 UI、信息图、海报等文字密集视觉设计的 6B 文生图模型,支持完整视觉构图和带透明背景的 RGBA 输出。
通义 QwenAudio:原创语音项目AI 评分4848 通义 Qwen-Image-2.1 发布:Qwen 最强开源图像生成模型
通义(Qwen)发布开源图像生成模型 Qwen-Image-2.1,官方称其为 Qwen 目前最强的开源图像生成模型。该模型已在 QwenLM 项目下开源,具体参数规模与评测分数尚未在原文中披露。
上海人工智能实验室 InternLM:原创项目AI 评分4646 InternLM 发布 InternLumina-U2 多码本扩散大语言模型
上海人工智能实验室 InternLM 团队发布 InternLumina-U2,一个面向全视觉理解、图像生成与编辑的多码本扩散大语言模型。该模型将扩散生成能力与语言建模统一在同一框架内,可同时处理视觉理解与图像生成、编辑任务。
蚂蚁 inclusionAI:GitHub 新仓库AI 评分2323 蚂蚁 inclusionAI 发布 ConceptEdit:ConceptEdit-12M 数据集与 ConceptEdit-Bench 基准
蚂蚁 inclusionAI 在 GitHub 上线 ConceptEdit 项目,包含 ConceptEdit-12M 数据集与 ConceptEdit-Bench 基准。目前公开信息仅给出这两个名称,尚未披露模型规模、评测分数或使用方式等细节。
Midjourney:Updates(RSS)AI 评分4747 Midjourney 发布 V8.2 图像模型
Midjourney 推出 V8.2 图像模型,本次更新聚焦美学、图像质量与个性化。官方称图像将更具创意、大胆、精致、前卫和新颖,低质量图像的随机情况应会大幅减少。V8.2 的个性化档案拥有更大且更优的图像池,能更好理解用户个人品味,尤其适合在个人档案下积累了大量评分的用户。
Midjourney:Updates(RSS)AI 评分5858 Midjourney 将默认模型从 V7 切换为 V8.1
Midjourney 宣布经过测试和反馈后,默认模型已从 V7 更新为 V8.1。V8.1 更聪明、更连贯,更好地遵循详细提示词并渲染文字;开启 HD 模式后图像尺寸为 V7 的两倍、分辨率 4 倍,SD 模式 4 秒、HD 模式 12 秒出图。风格参考、个性化与美学在 V7 与 V8.1 间保持一致;V7 的 omni-reference 仍可使用,V8.0 alpha 将在两周后弃用。
Saining Xie@sainingxieAI 评分4545引用Jaskirat Singh@1jaskiratsinghIn Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation. Today we make them a bit more simple. a bit more general. Result: >10x faster convergence, better reconstruction, better generation. And yes we test them on T2I and world models :) Introducing RAEv2
上海人工智能实验室 InternLM:原创项目AI 评分3939 上海人工智能实验室 InternLM 推出 ETCHR:面向 MLLM 的推理感知图像编辑器
上海人工智能实验室 InternLM 项目发布 ETCHR,一个以问题为条件、具备推理感知能力的图像编辑器,定位为多模态大语言模型(MLLM)的解耦式视觉推理助手。该工具通过解耦设计为 MLLM 提供独立的视觉推理支持。
美团 LongCat:HuggingFace 新模型AI 评分5858 美团 LongCat 开源 LongCat-Next 原生多模态模型
美团 LongCat 团队开源原生多模态模型 LongCat-Next,基于 LongCat-Flash-Lite MoE(A3B),在单一自回归框架内统一文本、视觉和音频处理,采用 DiNA 离散原生自回归范式,结合 SAE 与 RVQ 构建语义完整的离散视觉表示,并提出 dNaViT 视觉接口。
FireRedTeam:原创语音与多模态项目AI 评分4545 FireRedTeam 开源 FireRed-Image-Edit 图像编辑基础模型
FireRedTeam 开源图像编辑基础模型 FireRed-Image-Edit,宣称达到开源 SOTA 水平。该模型具备精准指令跟随、高保真生成、优异的身份一致性和多元素无缝融合能力。
美团 LongCat:HuggingFace 新模型AI 评分5151 美团发布 LongCat-Image-Edit-Turbo 蒸馏版图像编辑模型
美团 LongCat 在 HuggingFace 发布 LongCat-Image-Edit-Turbo,是 LongCat-Image-Edit 的蒸馏版本,仅需 8 次 NFE 即可实现高质量图像编辑,推理延迟极低。
通义 QwenAudio:原创语音项目AI 评分3737 Qwen-Image-Layered:面向内在可编辑性的分层分解
QwenLM 推出 Qwen-Image-Layered,通过分层分解让图像具备内在可编辑性。该模型将图像拆解为多个图层,使各元素可独立编辑,而非在单一平面上修改。
美团 LongCat:HuggingFace 新模型AI 评分5454 美团 LongCat 发布图像编辑模型 LongCat-Image-Edit
美团 LongCat 发布图像编辑模型 LongCat-Image-Edit,为 Longcat-Image 的编辑版本,支持中英双语编辑。官方称其在开源图像编辑模型中达到 SOTA,支持全局编辑、局部编辑、文本修改和参考引导编辑,能保持未编辑区域的布局、纹理、色调和主体身份一致,适合多轮编辑。
美团 LongCat:HuggingFace 新模型AI 评分5757 美团 LongCat 发布开源双语图像生成模型 LongCat-Image
美团 LongCat 发布开源中英双语图像生成基础模型 LongCat-Image,仅 6B 参数,在多个基准上超越数倍于其规模的开源模型。
美团 LongCat:HuggingFace 新模型AI 评分6161 美团 LongCat 发布 13.6B 开源视频生成模型 LongCat-Video
美团 LongCat 团队发布 13.6B 参数的开源视频生成基础模型 LongCat-Video,统一支持文生视频、图生视频和视频续写三类任务。模型在 Video-Continuation 上预训练,可生成长达数分钟的视频而不出现色彩漂移或质量下降,并通过粗到细策略和 Block Sparse Attention 在数分钟内生成 720p、30fps 视频。
通义 QwenAudio:原创语音项目精选AI 评分7070 Qwen-Image 开源:支持复杂文字渲染与精确图像编辑的图像生成基础模型
QwenLM 发布 Qwen-Image,一个图像生成基础模型,主打复杂文字渲染和精确图像编辑能力,代码发布在 GitHub(https://github.com/QwenLM/Qwen-Image)。
FireRedTeam:原创语音与多模态项目AI 评分4444 FireRedTeam 的 StoryMaker:实现文生图角色一致性
FireRedTeam 的 StoryMaker 项目致力于解决文生图生成中的角色一致性问题。该项目聚焦于让同一角色在不同场景与提示词下保持外观统一,属于原创语音与多模态方向的工作。
FireRedTeam:原创语音与多模态项目AI 评分3838 FireRedTeam 发布 PhotoPoster:姿态驱动的图像生成开源项目
FireRedTeam 推出开源项目 PhotoPoster,用于姿态驱动的图像生成,以支持并推进人像动画领域的研究。