蚂蚁 inclusionAI 发布 SingProbe:基于 gemma-4-26B-A4B-it 的流式安全探针
蚂蚁 inclusionAI 发布 SingProbe,一个构建在 google/gemma-4-26B-A4B-it 上的内在流式护栏,复用基座模型隐藏状态逐 token 打分,探针参数仅 5.67M,解码开销低于 0.5%。
蚂蚁 inclusionAI 发布 SingProbe,一个构建在 google/gemma-4-26B-A4B-it 上的内在流式护栏,复用基座模型隐藏状态逐 token 打分,探针参数仅 5.67M,解码开销低于 0.5%。
🍾🍲 Saturday Robotics x IROS 2026 — Robotics Research Night 👉🏻 https://luma.com/tzbw7n61 We’re bringing a high-signal evening of robotics research to Pittsburgh on September 28. After a full day at IROS, we’ll bring together researchers, engineers, founders, students, and investors for technical discussions, networking, and a series of ~10-minute lightning talks. Tentative preview of the current lineup: 🤖 1. PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball Gary Yang @lzyang2000 (@Caltech) Perception-aware reinforcement learning + Control Barrier Functions for whole-body humanoid safety. Demonstrated on a Unitree G1, with 19/20 successful dodges and zero falls in real-world experiments. 🧠 2. How In-Context Learning Is Reshaping Robot Learning Data at Scale AaronLi (@RhodaAI) Exploring how in-context learning can change the way we think about robot learning data, scaling, and generalization. 🧪 3. X2Real: An eXtensive Simulation Benchmark for Real-World Generalist Policies Liangwang Ruan (@XSquareRobot) A new simulation benchmark built around faithfulness, diversity, and fairness, with 44 hierarchical long-horizon tasks across 10 capability dimensions and a reported 0.84 simulation-to-real correlation. 🦾 4. Rethinking Generalist Robotic Manipulation: Architecture, Data and Inference for Real-World Deployment Peiyan Li (Chinese Academy of Sciences, @CAS__Science) 3D VLA architectures, memory augmentation, ego/UMI human priors, large-scale robot pretraining, and inference-time contextual learning for deployable generalist manipulation. 🎯 5. HiRE: Hindsight Reward Editing for Policy Finetuning Haoyi Niu @t641769919 (@UCBerkeley) Accepted at CoRL 2026. A training-free approach to reward editing that uses successful and failed trajectories to identify “trap states” and provide denser, control-aware feedback for RL. 🔥 6. Lightning Talk — Open Slot We’re opening one additional slot for a technically deep research talk, new project, frontier paper, demo, open problem, or startup technical insight. 10 minutes. A few slides. One sharp technical idea. No fluff. Topics include World Models, Physical AI, Humanoids, VLAs, Robot Foundation Models, Manipulation, RL, Simulation & Sim-to-Real, Spatial Intelligence, Computer Vision, and Embodied AI. 📍 Pittsburgh 📅 September 28, 2026 🕠 5:30–9:30 PM 🍾 Networking + Technical Talks + Research Discussion 📩 junfanzhu98@gmail.com See you in Pittsburgh. 🤖 #IROS2026 #Robotics #PhysicalAI #RobotLearning #WorldModels #HumanoidRobotics #VLA #EmbodiedAI #RobotFoundationModels
vime 联合 RL-Kernel 在 AMD Instinct MI300X 上实现训练与 rollout 的逐位数值一致性,8× MI300X 跑 Qwen3-8B GRPO 实验连续 200 步 mismatch_count = 0、max_abs_diff = 0。
Perplexity 正使用 GPT-6 Astra 撰写沟通内容、修改软件并监控生产系统,且相比早期模型,人工检查频率大幅降低。
Gary Marcus 评 Dario Amodei 呼吁给 AI 发展减速的文章,Sam Altman 与 Elon Musk 已表态支持。Marcus 肯定其透明度承诺,但质疑其依赖与 AI 公司关系密切的 METR 做评估有监管捕获之嫌,指其拿中国当挡箭牌有损合作对话,并提出追责和产品召回等替代政策选项。文末提到特朗普反对减速,认为美国必须赢下 AI 竞赛。
推荐理由:Gary Marcus 对 Dario Amodei 的减速提案给出有保留的支持,并指出监管捕获、追责与召回等被绕开的政策选项。
vLLM 通过自适应投机 token 预算、KDA 前缀检查点、零拷贝混合 KDA 批次和延迟 MXFP4 收尾等优化,将 Kimi K3 服务性能从 v0.27.1 提升至 main:延迟降低 56%–60%,吞吐量提升 2.2–2.8 倍,TTFT 降低 72%–85%(并发 1、4、16,8K/1K 负载,TP8)。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
Nous Research 在 GitHub 发布 Hermes Agent 插件 hermes-plugin-snyk,通过 Snyk 官方 MCP server 实现代码、依赖、容器与 IaC 扫描,并内置 snyk-security-scan 技能(Agent Plugins v1)。
Nous Research 在 GitHub 发布 Hermes Agent 插件 hermes-plugin-touchdesigner,通过 twozero MCP server 驱动 TouchDesigner,并内置 touchdesigner-mcp 技能,采用 Agent Plugins v1 规范。
Our own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、工人转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵
Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"
GitHub 日本和韩国营销负责人把活动运营流程交给 GitHub Copilot,从单个 GitHub Issue 出发自动完成落地页复制、UTM 链接生成、报名名单清洗和会后报告。
Apple 发布 iPhone 18 Pro 和 iPhone 18 Pro Max,搭载 48MP 可变光圈 Fusion 主摄、A20 Pro 与新一代均热板,随 iOS 27 以 beta 推出 Siri AI。
Cognition 用 GPT-6 Astra 提升 Devin 测试软件并证明其可用的能力,目标是让工程师少审代码、多交付。该能力聚焦于 Devin 对自身工作的验证环节。
a16z 指出,许多 LP 对 SpaceX、Anthropic 和 OpenAI 三家前沿模型公司几乎零敞口,而 SpaceX 上市后市值约 2 万亿美元,成为规模达此前纪录 10 倍的史上最大 VC 背景 IPO,Anthropic 估值 965B 美元、OpenAI 最近估值 852B 美元。作者认为,传统把风投控制在整体组合 5-10% 的资产配置框架已经破裂,LP 需要重新调整风投仓位。
美国联邦实验室联盟将 2026 年技术转移卓越奖授予 MIT Lincoln Laboratory 与麻省总医院(MGH)联合开发的 AI-GUIDE 医疗设备。
OpenAI 将 Habitat 从 Python 库演进为全球分布式存储平台,支撑超 10 亿 ChatGPT 用户、每秒 2200 万次请求。该平台用于满足 ChatGPT 在线存储的规模化需求。
Together AI 扩展 Together Fine-Tuning 服务,新增 GLM 5.3、Kimi K2.7、DeepSeek-V4-Flash、Qwen 3.8-27B、Gemma 4 等开源权重模型支持,并通过 API、CLI 和 UI 提供逐步训练指标实时追踪。
Google Research 提出 ToolGrad,一种“先答案后问题”的工具调用数据生成范式:先迭代构建可验证的 API 调用链,再反推用户提示词,替代 ToolBench、ToolACE 等基于 DFS 试错的低效方案。
GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让开发者无需在编辑器、终端和浏览器之间切换即可完成 AI 编码闭环。diff 面板以绿色标注新增、红色标注删除,支持接受改动、留言或让 Copilot 继续修改;终端面板可直接运行项目命令并支持多窗口;浏览器面板可用 Pick & Polish 工具选中元素并让智能体调整。
a16z 发文指出,随着保费每年上涨 10% 以上,多数雇主正开始寻找替代方案,或转向低成本健康计划,或彻底放弃传统健康保险。这一规模达 1 万亿美元、覆盖 1.5 亿以上美国人的雇主医保市场,正因 AI 降低建计划与运营的固定成本门槛而出现代际替换机会,催生一批新型替代健康计划(AHP)、挑战者 PBM 和现代化基础设施平台。
Sierra 发布多模态智能体,将语音、文本和可视化整合进同一段对话,并自动判断何时切换模式,用户无需重来或重复表述。该智能体一次构建即可部署到所有渠道,视觉组件同样通用;其 MCP UI 集成支持把产品卡片、对比表格、日历和表单直接嵌入对话,组件由企业自行设计和托管,更新后自动同步,无需重新部署或为各平台维护不同版本。
NVIDIA 通过全栈 NIM 优化,在 Nemotron 3 Ultra 上实现 2.5 倍并发用户量。该优化针对生产环境部署大语言模型时,在现有 GPU 基础设施上提升并发服务能力并保持交互响应速度的需求,对提示词长、上下文跨步骤复用的智能体 AI 工作负载尤为关键。
César de la Fuente 的实验室使用 Codex 和 ChatGPT,在现存与已灭绝生物的基因组中搜寻抗菌候选分子,以对抗耐药性感染。
New from us: Anthropic just published scenarios for AI’s possible economic impacts, which range from minimal, to explosive GDP growth of 15% by 2030 as knowledge-worker unemployment hits 18%. I sat down with their co-founder Jack Clark to pick his brains on how they’re thinking about all of this.
编码器回归了,宝贝 (引用推文 @scaling01:这到底是什么外星架构)
what in the alien architecture is this
NVIDIA BioNeMo Inference Runtime(BioIR)可在 NVIDIA GPU 上加速受支持的生物分子结构预测模型,同时保留熟悉的 PyTorch 工作流。它通过优化 kernel 并在适用场景下使用 CUDA Graphs 提升模型速度,面向蛋白质组规模的高通量结构预测任务。
OpenAI 在 ChatGPT Work 中推出 Data agent,支持连接公司数据、发掘洞察,并用自然语言让 AI 构建交互式仪表盘。
推荐理由:上线方直接给出参数结构、上下文窗口、KV cache 对比和许可证信息,读者可据此评估实际部署选型。