vLLM 支持基于 Gumbel-max 的无损文本水印
vLLM 现已支持基于 Gumbel-max 算法的无失真文本水印,通过 PRF 生成可复现的 keyed 噪声并把 PRNG、Gumbel 变换和 argmax 融合为单个 GPU kernel。
推荐理由:作者亲自实现了 vLLM 的水印功能,给出了算法原理、吞吐实测数据和启用命令,读者可以据此评估在现有推理服务中采用的成本。
vLLM 现已支持基于 Gumbel-max 算法的无失真文本水印,通过 PRF 生成可复现的 keyed 噪声并把 PRNG、Gumbel 变换和 argmax 融合为单个 GPU kernel。
推荐理由:作者亲自实现了 vLLM 的水印功能,给出了算法原理、吞吐实测数据和启用命令,读者可以据此评估在现有推理服务中采用的成本。
Meta 把用于 WhatsApp 和 Meta AI 应用的 Private Processing 机密计算基础设施扩展到 AI 眼镜,让流式转录、上下文搜索和长期回忆等云端 AI 负载在机密虚拟机(CVM)内运行,Meta 自身也无法读取用户数据。该方案基于 CPU 与 GPU 的 TEE 硬件隔离,客户端通过远程证明校验软件镜像,并叠加不可定向性与加密存储。
interactive media will be huge in the next year or two! this is incredible
Claude Opus 5.5 在浏览器中一次性构建出 Minecraft……结果太疯狂了! 这是 Claude Opus 5.5 x @Viggle_PINOC MCP 用于游戏开发的又一个例子。
Claude Opus 5.5 with PINOC MCP one shotted this: A playable Minecraft recreation Link, prompt and more examples below 🧵
YouTube 在 Made on YouTube 活动上公布自定义信息流 Custom Feeds,并预告面向观众和创作者的更多 AI 功能。用户输入描述即可生成信息流,通过修改描述、给推荐视频打分来调优,并可保存多个,该功能即将面向美国 web、移动端和 TV 用户推出。今年晚些时候 YouTube 还将支持评论发 GIF,并为私信加入群聊,群聊初期仅限美国、英国、新加坡、巴西及部分欧洲国家。
Claude Code 发布 v2.1.281,为 Claude apps gateway 新增 Claude Desktop 新密钥的桌面策略块支持、Bedrock 上游的 assume_role 与 Amazon Bedrock guardrail 配置,并支持在 settings.json 中用 "attribution": false 隐藏提交与 PR 署名。
Google 宣布所有 Google 或 Google Workspace 账号均可通过 Google Vids 使用最新的 Gemini Omni 1.1 Flash 模型免费生成高质量视频,入口为 vids.new 并选择 "Create AI videos"。
推荐理由:原文来自产品经理宣布,交代了免费开放入口、具体模型和新控制功能,读者可据此判断是否改变自己的视频制作流程。
Model Vault 现已在加拿大上线 🇨🇦 对于身处"真北"的用户,现在可以通过自动扩缩容工作负载、单租户部署和更低的 TCO,全面掌控你的 AI。Cohere 模型。完全私有。
We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.
推荐理由:OpenAI 官宣 ChatGPT Voice 可调用插件并进入 Work,读者可据此了解语音功能的实际使用边界。
Android Enterprise 公布 6 项更新,涵盖 Gemini 多步骤跨应用自动化、工作资料隔离与 IT 集中管控,以及 XREAL Aura 有线眼镜等 XR 设备统一管理。
Google Beam 正式扩展至美国、加拿大、英国、法国、德国和日本六国,由 18 家渠道合作伙伴支持部署,旗舰硬件为 HP Dimension with Google Beam。
勾勒一个世界。让 Atlas 赋予它生命。 抢先一睹 Chisel,现已在 Atlas beta 中上线。
Google 发布两个 Gemini 文本转语音模型 gemini-3.8-flash-tts 和 gemini-3.8-flash-lite-tts,提供超过 2,000 个声音,并支持用 30 秒音频样本创建自定义语音。
介绍 ACTx486,一种全新交互媒介的研究演示。 如果你能和任何视频对话、问任何问题呢? 我们的系统将一档现有播客变成了能倾听、回应并适应的东西。 研究由 @jakubzeg 完成:
Now Gemini can connect with 13 new apps like @adobe, @squarespace, @onepeloton, and more. Instead of switching between tabs, you can now grow your business, design assets, and plan your workouts all in Gemini. 🧵
Microsoft 在 GitHub 上线 agent365-runbook 仓库,提供将自有智能体接入 Microsoft Agent 365 的分步 runbook,覆盖身份、可观测性、Microsoft 365 数据访问和消息传递四部分。
免费两周!Hy Image3.5 预览版已在 OnSolo 上线。短剧角色设定图。全动态视频游戏素材。关键帧。角色在每一集中保持一致。编辑是精修,而非重来。
Hy Image3.5 preview is on OnSolo. Exclusive. 5 refs. 2K. 2 weeks Members free.
千问办公发布「企业上下文」,把群聊、文档、知识库等分散信息整理成结构化上下文,供 Agent 提取任务所需信息,思路是「连接与压缩」。同时推出由钉钉 A1 升级而来的 QwenNote A2 录音卡,默认不录音、云端 ASR 转写后永久删除音频,并上线多人协作功能。企业上下文依赖新模型 Qwen3.8-Omni-Flash,千问办公此前已与模型团队定制 Qwen 3.8 Flash。
消息已公布:DeepSeek-V4.1-Flash 在 WorkBuddy 上免费使用,截止到 10 月 9 日! 🗞️趁优惠还在,把它用在你下一个任务上吧。
Hy Image3.5 preview is now available in ComfyUI. Professional-grade image generation, +30% win rate in human eval vs Hy Image3.0 → Text to image and Image to image in one model, up to 2K → Multilingual text, symbols, and small print that render correctly → Cinematic, comic, commercial photography, and illustration styles → Identity and product features that hold through scene, outfit, and style changes
ChatGPT Ads 扩展至东南亚和台湾,覆盖超过 60 个国家和地区,符合条件的企业可借此触达更多用户。
汽车过去只是把你送到工作地点。现在它能在路上帮你做一部分工作。 你现在可以在你的特斯拉里使用 Grok Bot 和 Connectors。
.@Grok in your Tesla can now do meaningful work for you With Connectors, you can manage your inbox, clean up your calendar, or talk through existing files/chat/tasks – all hands-free
GPT-6 改进了提示词缓存机制,带来更高的缓存命中率、新的诊断工具、显式断点以及可降低延迟与成本的控制项。
Suno Studio 现已内置插件,包括压缩器。但压缩器到底有什么用? 在此阅读我们的概述: https://suno.com/blog/about-compression 或关注本线程中的要点🧵
我记得大约 1.5 年前我们刚做出这个东西的第一个粗糙版本,只是为了应对疯狂的量级 当时学到的道理现在依然成立:先爬,再走,后跑。找到能持续改进的飞轮。让它一直跑下去(适用于很多事情) 团队干得漂亮!
We rebuilt customer support around Grok Bot to scale our operations without adding headcount. It works autonomously to respond to customers, resolve tickets, and manage the queue. https://x.ai/news/grok-bot-customer-support
Perplexity 宣布 GPT-6 Sol 现已在 Perplexity 和 Computer 中可用。该模型成为 Computer 努力程度选择器中默认的 Light 选项。
推荐理由:官方宣布 GPT-6 Sol 上线并设为默认选项,用户可直接了解入口与默认设置变化。
一组 Claude Opus 5.5 的早期探索: @kevin_t_ngo 创作的一篇关于西瓜的短篇故事。
Simon Willison 发布 llm 0.36。该版本为其 LLM 命令行工具与 Python 库的更新,具体改动原文未列出。
Sierra 发布博客阐述其 AI 智能体平台如何实现全流程透明:用 Ghostwriter 以自然语言构建智能体,并在 Agent Studio 中查看和编辑每一次旅程、动作、策略与 persona。
Nous Research 在 GitHub 上线 hermes-plugin-blender,这是面向官方 Blender Lab MCP server 与工作流技能的 Hermes 可移植插件。该插件以 Hermes 插件形式接入 Blender Lab 的 MCP server,并提供对应的工作流技能。
I spent two years building this interactive tool to let you steer LLMs and agents at the token level. Introducing onPanda — a web app for token visualization & control, model inspection, data annotation, and more. Try it online (works on mobile): https://onpanda.diyer22.com/