跳到正文

全部动态

今日 87 条
9月24日周四
  1. vLLM 官方博客(RSS)66

    vLLM 支持基于 Gumbel-max 的无损文本水印

    vLLM 现已支持基于 Gumbel-max 算法的无失真文本水印,通过 PRF 生成可复现的 keyed 噪声并把 PRNG、Gumbel 变换和 argmax 融合为单个 GPU kernel。

    推荐理由:作者亲自实现了 vLLM 的水印功能,给出了算法原理、吞吐实测数据和启用命令,读者可以据此评估在现有推理服务中采用的成本。

  2. Meta Engineering Blog(RSS)49

    Meta 将 Private Processing 机密计算引入 AI 眼镜

    Meta 把用于 WhatsApp 和 Meta AI 应用的 Private Processing 机密计算基础设施扩展到 AI 眼镜,让流式转录、上下文搜索和长期回忆等云端 AI 负载在机密虚拟机(CVM)内运行,Meta 自身也无法读取用户数据。该方案基于 CPU 与 GPU 的 TEE 硬件隔离,客户端通过远程证明校验软件镜像,并叠加不可定向性与加密存储。

  3. Karina38

    当@jakubzeg和我在OpenAI相遇时,我们反复回到同一个困惑:太多AI产品都从一个空白聊天框开始。你几乎可以做任何事,但你必须决定从哪里开始,这让人不知所措。 有了ACTx486,我们从一段已经有故事的媒体出发,让你在观看的同时与之互动。这也让技术更加通用。 我们在这里写了更多关于这一选择的思考:https://www.actx486.com/

    引用Rehan Sheikh@rehan_shei

    interactive media will be huge in the next year or two! this is incredible

  4. Ars Technica:AI(RSS)36

    YouTube 承诺今年晚些时候推出自定义信息流与更多 AI 功能

    YouTube 在 Made on YouTube 活动上公布自定义信息流 Custom Feeds,并预告面向观众和创作者的更多 AI 功能。用户输入描述即可生成信息流,通过修改描述、给推荐视频打分来调优,并可保存多个,该功能即将面向美国 web、移动端和 TV 用户推出。今年晚些时候 YouTube 还将支持评论发 GIF,并为私信加入群聊,群聊初期仅限美国、英国、新加坡、巴西及部分欧洲国家。

  5. Google Blog:AI(RSS)68

    Google Vids 集成 Gemini Omni 1.1 Flash,所有账号可免费生成 1080p HD 视频

    Google 宣布所有 Google 或 Google Workspace 账号均可通过 Google Vids 使用最新的 Gemini Omni 1.1 Flash 模型免费生成高质量视频,入口为 vids.new 并选择 "Create AI videos"。

    推荐理由:原文来自产品经理宣布,交代了免费开放入口、具体模型和新控制功能,读者可据此判断是否改变自己的视频制作流程。

  6. Greg Brockman78

    OpenAI 宣布 ChatGPT Voice 重大升级,现可使用邮箱、日历、Slack 等插件,并可由 GPT-6 Astra、Sol 和 Luna 驱动。语音功能还登陆网页端和移动端的 ChatGPT Work,用户可仅靠语音在浏览器中创建文档、幻灯片、网站和表格或处理复杂任务,今日起随最新版 App 全球推出。

    引用OpenAI@OpenAI

    We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.

    推荐理由:OpenAI 官宣 ChatGPT Voice 可调用插件并进入 Work,读者可据此了解语音功能的实际使用边界。

  7. SiliconFlow42

    并非每次模型调用都需要一个答案。有时,它只需要一个决策。 👏 欢迎 Kev-4B 加入 SiliconFlow。 Kev-4B 是 Jev 的开源社区版,基于 Qwen3.5-4B 构建,用于结构化决策。 路由。排序。批准。升级——无需再生成一段回复。 无需部署或适配。一个 SiliconFlow API key,Kev 即可接入你的工作流。 特别感谢 @jaredpalmer 开源 Kev。❤️ 在 SiliconFlow 上试用 Kev-4B。⚡️

9月23日周三
  1. elsewhere:文章(RSS)50

    千问办公发布「企业上下文」与 QwenNote A2,要把企业内部连接起来

    千问办公发布「企业上下文」,把群聊、文档、知识库等分散信息整理成结构化上下文,供 Agent 提取任务所需信息,思路是「连接与压缩」。同时推出由钉钉 A1 升级而来的 QwenNote A2 录音卡,默认不录音、云端 ASR 转写后永久删除音频,并上线多人协作功能。企业上下文依赖新模型 Qwen3.8-Omni-Flash,千问办公此前已与模型团队定制 Qwen 3.8 Flash。

  2. Tencent Hy59

    Hy Image3.5 preview 现已在 ComfyUI 中可用。人工评测胜率相比 Hy Image3.0 提升 30%;单一模型支持文生图与图生图,最高 2K 分辨率,支持多语言文字与小字渲染,覆盖电影、漫画、商业摄影、插画风格,人物身份与产品特征可在场景、服装和风格切换中保持一致。

    引用ComfyUI@ComfyUI

    Hy Image3.5 preview is now available in ComfyUI. Professional-grade image generation, +30% win rate in human eval vs Hy Image3.0 → Text to image and Image to image in one model, up to 2K → Multilingual text, symbols, and small print that render correctly → Cinematic, comic, commercial photography, and illustration styles → Identity and product features that hold through scene, outfit, and style changes

  3. eric zakariasson38

    我记得大约 1.5 年前我们刚做出这个东西的第一个粗糙版本,只是为了应对疯狂的量级 当时学到的道理现在依然成立:先爬,再走,后跑。找到能持续改进的飞轮。让它一直跑下去(适用于很多事情) 团队干得漂亮!

    引用Grok Bot@bot

    We rebuilt customer support around Grok Bot to scale our operations without adding headcount. It works autonomously to respond to customers, resolve tickets, and manage the queue. https://x.ai/news/grok-bot-customer-support

  4. StepFun55

    阶跃星辰(StepFun)宣布开源内部使用的 LLM 数据标注与模型检查工具 onPanda,工作流为找到错误、修正 token、让模型继续生成。数据标注方面,标注时间中位数比人工后编辑降低 52%,SFT 与偏好数据可在同一流程完成(ΔPPL <1%),支持 token 级正负样本监督及图像、音频、视频上的 agent 轨迹标注。

    引用Lei Yang@diyerxx

    I spent two years building this interactive tool to let you steer LLMs and agents at the token level. Introducing onPanda — a web app for token visualization & control, model inspection, data annotation, and more. Try it online (works on mobile): https://onpanda.diyer22.com/