跳到正文

全部动态

今日 45 条
今天10月1日周四
  1. Artificial Analysis67

    Artificial Analysis 评测称 GPT-6.1 Sol 推高了 OpenAI 模型的成本效率前沿。max effort 下每 Intelligence Index 任务成本 $0.72,不到 GPT-6 Astra($3.26)的四分之一,比 GPT-6 Sol($1.04)低 31%,比 GPT-5.6 Sol($1.99)低 64%。

    推荐理由:第三方评测给出 GPT-6.1 Sol 与多款前代模型的具体成本对比数字,读者可据此评估 OpenAI 模型的性价比变化。

  2. IT之家(RSS)61

    谷歌推出 Gemini 4 Argon 旗舰模型,内部员工质疑其实战编码表现

    据彭博社报道,谷歌开始逐步推出旗舰模型 Gemini 4 Argon,先向一小批网络安全合作伙伴开放,之后优先面向付费订阅用户。谷歌称该模型多项基准测试靠前,安全测试成绩超过 OpenAI 的 Astra,但知情人士称其实际处理部分代码任务表现不佳,尤其前端设计能力参差不齐,且模型体量庞大、运行成本高。

  3. Arena.ai47

    Hidream-O1-Video-1.0 由 @HiDream_AI 刚刚登陆 Image-to-Video Arena,以 1456 分位列第 7! 该模型已在 @vivago_ai 上线,距 gemini-omni-flash 仅差 8 分,距 dreamina-seedance-2.0 和 2.5 均在 20 分以内。 恭喜 @HiDream_AI 发布!

    引用vivago.ai (HiDream)@vivago_ai

    New HiDream models just landed in vivago R1 Studio 🚀 Introducing: • HiDream-O1 Image 2.0 • HiDream-O1 Editing 1.5 • HiDream-O1 Video All three models are now available in vivago R1 Studio - bringing the latest HiDream image generation, editing, and video capabilities directly into your creative workflow. New models. New possibilities. Go make something the internet can’t ignore. 🔥

  4. Artificial Analysis 完整文章(网页)79

    Artificial Analysis 评测 Gemini 4 Argon:Google 重回智能前三

    Artificial Analysis 评测 Google DeepMind 新模型 Gemini 4 Argon,其在 Artificial Analysis Intelligence Index 得 53 分,追平 GPT-6 Astra(max),高于 GPT-6.1 Sol(52),为 Google 超 7 个月来首个高于 Flash 档的专有模型。

    推荐理由:第三方评测给出了智能指数、单位任务成本、幻觉率等多项横向数据,可用于比较 Gemini 4 Argon 与竞品的实际表现。

  5. Arena.ai78

    Arena 公布 Gemini 4 Argon (High) 在 Agent Arena 以 +7.92% 净改进分排名第 8,每任务成本 $0.62,重塑了 Pareto 前沿,比 Gemini 3.8 Flash (High) 高 4.96 个百分点。

    引用Arena.ai@arena

    Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8. Congrats to the @GoogleDeepMind team on this impressive frontier release!

    推荐理由:原文给出 Agent Arena 排名、关键信号得分和每任务成本数据,读者可以据此评估该模型在真实智能体任务中的性价比。

  6. Rohan Paul44

    美国 Austin 的 webAI 推出 3.66B 本地模型 TwIL-LM3-Pro,将 IBM Granite 的形式逻辑分数提升 28%,在形式逻辑上比 VibeThinker-3B 高约 35%、比 Qwen3.5-4B 高 24%、比 LFM2.5-8B-A1B 高 47%。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  7. Google DeepMind:Blog(RSS)77

    Google DeepMind 发布 Gemini 4 Argon 前沿模型

    Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。

    推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。

  8. Google Blog:AI(RSS)76

    Google 发布 Gemini 4 Argon 前沿模型

    Google 发布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 面向可信网络防御者开放,价格为每百万输入 token $2、输出 token $10,缓存输入 token 为输入价的 5%。

    推荐理由:官方公告给出定价、输出 token 上限和多个基准分数,读者可以据此评估它在编码与安全防御场景的落点。

  9. Karina50

    Gemini 的 PostTrainBench 得分翻了一倍多:21.99%(3.1 Pro)→ 45.3%(4)🔥🚀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  10. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  11. Rohan Paul61

    Google 发布新旗舰模型 Gemini 4 Argon,作者称其在多数基准上超过 GPT-6 Astra 和 Claude Opus 5.5,输出上限从 64K 提升到行业领先的 1M tokens,约为此前 128K 上限的近 8 倍。引用材料提到其在 Harvey's Legal Agent Benchmark 上领先,且目前仅限 Google 员工、经审核的网络安全相关机构和可信测试者使用。

    引用Rohan Paul@rohanpaul_ai

    MASSIVE reveal from Google. Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks. - beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks. - its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1. - output limit jumps from 64K to 1M tokens, an industry-leading ceiling, - Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback. - Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.

  12. Arena.ai66

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,通过 Fairwind Program 向部分受信任测试者开放。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

    推荐理由:榜单方公布了 Gemini 4 Argon (High) 在 Text Arena 的分项名次、1525 分和混合价格,读者可据此对比成本效率。

  13. 🚨 AI News | TestingCatalog47

    突发 🔥:Google 宣布 Gemini 4 Argon,一款新的前沿模型,面向"跨真实世界软件工程、法律和金融等企业知识工作、以及网络防御的复杂工作流"。 在 DeepSWE v1.1 上取得 77.9% 的分数,创下新 SOTA。在众多基准测试上表现优于 GPT-6 Astra、Opus 5.5 和 Fable 5.1。 即将推出,首先面向付费 API 客户和 Google AI Ultra 订阅用户。 很快!👀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  14. Aravind Srinivas48

    我们正在开源我们最先进的上下文嵌入模型,它在 turbopuffer 的 context-bench 中表现最佳。

    引用Perplexity@perplexity_ai

    We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage

  15. MiniMax (official)43

    基于 MiniMax H3 构建,@Creatify_Labs 的 Boreal-H3 是一款专为广告优化的视频模型,在更精准遵循创意简报的同时,保持产品和角色的一致性。 很高兴看到 MiniMax H3 成为更多面向特定行业的前沿模型的基础!✨

    引用Creatify Labs@Creatify_Labs

    Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.