跳到正文

全部动态

今日 563 条
今天10月1日周四
  1. 🚨 AI News | TestingCatalog56

    Inworld 宣布收购 Ultravox,一个用于构建实时语音智能体的平台,可听、推理、调用工具并在对话中处理轮次和打断。Ultravox 上内置的 Inworld 声音现已运行在 Realtime TTS-2 上,现有 voice ID 无需迁移或修改代码即可继续使用,TTS-2 还为 Ultravox 智能体增加自然语言引导能力。

    引用Inworld AI@inworld

    We’re excited to announce that @ultravox_dot_ai is now part of Inworld. Ultravox is the platform developers use to build real-time voice agents. We've worked with the team for a while through our TTS partnership, and today members of the team that built it are joining Inworld to keep developing it.

  2. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  3. Rohan Paul61

    Google 发布新旗舰模型 Gemini 4 Argon,作者称其在多数基准上超过 GPT-6 Astra 和 Claude Opus 5.5,输出上限从 64K 提升到行业领先的 1M tokens,约为此前 128K 上限的近 8 倍。引用材料提到其在 Harvey's Legal Agent Benchmark 上领先,且目前仅限 Google 员工、经审核的网络安全相关机构和可信测试者使用。

    引用Rohan Paul@rohanpaul_ai

    MASSIVE reveal from Google. Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks. - beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks. - its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1. - output limit jumps from 64K to 1M tokens, an industry-leading ceiling, - Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback. - Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.

  4. Arena.ai66

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,通过 Fairwind Program 向部分受信任测试者开放。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

    推荐理由:榜单方公布了 Gemini 4 Argon (High) 在 Text Arena 的分项名次、1525 分和混合价格,读者可据此对比成本效率。

  5. 🚨 AI News | TestingCatalog47

    突发 🔥:Google 宣布 Gemini 4 Argon,一款新的前沿模型,面向"跨真实世界软件工程、法律和金融等企业知识工作、以及网络防御的复杂工作流"。 在 DeepSWE v1.1 上取得 77.9% 的分数,创下新 SOTA。在众多基准测试上表现优于 GPT-6 Astra、Opus 5.5 和 Fable 5.1。 即将推出,首先面向付费 API 客户和 Google AI Ultra 订阅用户。 很快!👀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  6. Rohan Paul28

    “你的经济寿命实际上将在两年内终结。(因为 AI 正变得如此有能力)。 一个人类团队(在工作场景中)只会让情况更糟,因为他们要睡觉,他们会犯错,他们没喝咖啡醒来时还会有点脾气。所以你的认知价值会变成负数。 所以,你的经济寿命实际上将在两年内终结。” —— Emad Mostaque,Stability AI 创始人 --- 来自 YouTube 频道 “The Peter McCormack Show”,(链接见评论)

  7. Bloomberg:Technology(RSS)32

    Robinhood CEO 称新 AI 智能体应用安全可靠

    Robinhood 董事长兼 CEO Vlad Tenev 在 Houston 峰会上发布公司新的 AI 智能体应用,并称其安全可靠。他表示要让个人交易者获得对冲基金级别的工具,包括跨资产类别 24/7 交易、用户睡眠时仍可运行的自主“agent loops”,以及卫星影像和区块链分析等专用数据源。

  8. Bloomberg:Technology(RSS)54

    美光季度营收指引约 615 亿美元,超出分析师预期

    美光(Micron)给出的本财季(截至 11 月)营收指引约为 615 亿美元,高于分析师平均预期的 568 亿美元;剔除部分项目后每股利润预计约 38.15 美元,超过 36.02 美元的预期。公司称 AI 建设带来前所未有的需求,推动需求超过供给,但标题同时提到加薪压缩了利润率。

  9. IT之家(RSS)71

    美光科技 2026 财年归母净利润 849.69 亿美元,同比增长 895.07%

    美光科技发布 2026 财年年报,营业总收入 1331.88 亿美元,同比增长 256.33%,归母净利润 849.69 亿美元,同比增长 895.07%,毛利率 80.7%。第四财季营收 542.29 亿美元,环比增长 30.81%;公司预计 2027 财年第一季度营收 600 亿至 630 亿美元,并已量产 512GB DDR5 RDIMM 内存模块,速率可达 9200 MT/s。

  10. Ars Technica:AI(RSS)61

    RFK Jr. 称 AI 将摆脱医学专家的统治,Ars Technica 实测发现 AI 并不支持其观点

    美国卫生部长 Robert F. Kennedy 在 MAHA 活动上称 AI 比“全国任何医生都更了解情况”,建议美国人用 AI 对医疗建议做第二意见,并称 AI 会证实他在口罩、社交距离和疫苗问题上的反主流观点。Ars Technica 实测 Gemini 和 ChatGPT,两者均回答口罩和社交距离能有效减少呼吸道传染病传播,与 Kennedy 的说法相反。

  11. Rohan Paul50

    Inworld 宣布收购开发者用于构建实时语音智能体的平台 Ultravox,原团队成员加入 Inworld 继续开发。Ultravox 已将内置语音迁移到 Realtime TTS-2,该模型支持用内联指令在特定语段调整语速、语气和情感表达,例如在讲解复杂步骤时放慢语速再恢复正常。

    引用Inworld AI@inworld

    We’re excited to announce that @ultravox_dot_ai is now part of Inworld. Ultravox is the platform developers use to build real-time voice agents. We've worked with the team for a while through our TTS partnership, and today members of the team that built it are joining Inworld to keep developing it.

  12. Luma50

    Luma AI 宣布 Ideogram 4.5 已接入 Luma,供创意团队使用。该编辑模型据称能消除多轮编辑中的伪影累积和像素色彩偏移,可通过 https://app.lumalabs.ai 体验。

    引用Ideogram@ideogram_ai

    Introducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon.

  13. Aravind Srinivas48

    我们正在开源我们最先进的上下文嵌入模型,它在 turbopuffer 的 context-bench 中表现最佳。

    引用Perplexity@perplexity_ai

    We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage

  14. dex39

    这正是我们讨论杠杆的原因,也是 /show-me 的用途所在 https://www.youtube.com/watch?si=kPO3wH00gA3XQKhV&t=2959&v=-w6RDNBAI0E&feature=youtu.be

    引用Sureffi@Sureffi

    "a spec that is sufficiently detailed to generate code with a reliable degree of quality is roughly the same length and detail as the code itself" ^^ 100% don't believe in that. You should care about the code at the level of abstraction @dexhorthy is describing here. But no way in hell you can't compress it way smaller than the code itself would be (40:1 based on my measurements for a 30k LOC codebase). An LLM holds the priors for pretty much every single convention there is. Along with the cultures from Linus Torvalds to corporate Java. Two things - Understanding the model's priors for your choice of language/framework(s). - You holding those same priors. First screenshot is benchmark results from a few days ago. Second is what the 40:1 compression looks like. YMMW with typescript or python slop.