跳到正文

全部动态

今日 46 条
9月14日周一
  1. X Square Robot23

    期待参加在匹兹堡举办的 Saturday Robotics × IROS 2026!🤖 我们将展示 X2Real,这是我们用于评估真实世界通用机器人策略的大规模仿真基准——涵盖 10 个能力维度的 44 个分层长时程任务。 期待分享我们的最新工作,并与推进机器人学习、仿真到真实迁移和具身 AI 的研究者和开发者交流。 📍 匹兹堡 📅 2026 年 9 月 28 日 👉🏻 https://luma.com/tzbw7n61 到时见!

    引用Junfan Zhu 朱俊帆 ✈️ IROS@junfanzhu98

    🍾🍲 Saturday Robotics x IROS 2026 — Robotics Research Night 👉🏻 https://luma.com/tzbw7n61 We’re bringing a high-signal evening of robotics research to Pittsburgh on September 28. After a full day at IROS, we’ll bring together researchers, engineers, founders, students, and investors for technical discussions, networking, and a series of ~10-minute lightning talks. Tentative preview of the current lineup: 🤖 1. PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball Gary Yang @lzyang2000 (@Caltech) Perception-aware reinforcement learning + Control Barrier Functions for whole-body humanoid safety. Demonstrated on a Unitree G1, with 19/20 successful dodges and zero falls in real-world experiments. 🧠 2. How In-Context Learning Is Reshaping Robot Learning Data at Scale AaronLi (@RhodaAI) Exploring how in-context learning can change the way we think about robot learning data, scaling, and generalization. 🧪 3. X2Real: An eXtensive Simulation Benchmark for Real-World Generalist Policies Liangwang Ruan (@XSquareRobot) A new simulation benchmark built around faithfulness, diversity, and fairness, with 44 hierarchical long-horizon tasks across 10 capability dimensions and a reported 0.84 simulation-to-real correlation. 🦾 4. Rethinking Generalist Robotic Manipulation: Architecture, Data and Inference for Real-World Deployment Peiyan Li (Chinese Academy of Sciences, @CAS__Science) 3D VLA architectures, memory augmentation, ego/UMI human priors, large-scale robot pretraining, and inference-time contextual learning for deployable generalist manipulation. 🎯 5. HiRE: Hindsight Reward Editing for Policy Finetuning Haoyi Niu @t641769919 (@UCBerkeley) Accepted at CoRL 2026. A training-free approach to reward editing that uses successful and failed trajectories to identify “trap states” and provide denser, control-aware feedback for RL. 🔥 6. Lightning Talk — Open Slot We’re opening one additional slot for a technically deep research talk, new project, frontier paper, demo, open problem, or startup technical insight. 10 minutes. A few slides. One sharp technical idea. No fluff. Topics include World Models, Physical AI, Humanoids, VLAs, Robot Foundation Models, Manipulation, RL, Simulation & Sim-to-Real, Spatial Intelligence, Computer Vision, and Embodied AI. 📍 Pittsburgh 📅 September 28, 2026 🕠 5:30–9:30 PM 🍾 Networking + Technical Talks + Research Discussion 📩 junfanzhu98@gmail.com See you in Pittsburgh. 🤖 #IROS2026 #Robotics #PhysicalAI #RobotLearning #WorldModels #HumanoidRobotics #VLA #EmbodiedAI #RobotFoundationModels

  2. Gary Marcus:The Road to AI We Can Trust(RSS)67

    Gary Marcus 点评 Dario Amodei 的 AI 减速提案:三分肯定、七分质疑

    Gary Marcus 评 Dario Amodei 呼吁给 AI 发展减速的文章,Sam Altman 与 Elon Musk 已表态支持。Marcus 肯定其透明度承诺,但质疑其依赖与 AI 公司关系密切的 METR 做评估有监管捕获之嫌,指其拿中国当挡箭牌有损合作对话,并提出追责和产品召回等替代政策选项。文末提到特朗普反对减速,认为美国必须赢下 AI 竞赛。

    推荐理由:Gary Marcus 对 Dario Amodei 的减速提案给出有保留的支持,并指出监管捕获、追责与召回等被绕开的政策选项。

9月13日周日
  1. Peter McCrory68

    Peter McCrory 转发并推荐 Dario Amodei 的新文章《We Must Pace the Frontier》,该文主张 AI 行业应放慢速度并提出三步计划,Anthropic 单方面承诺其中第一步,即向第三方评估者提供永久、员工级别的系统访问权限,以便核查安全措施、报告事故和评估训练中的模型对齐。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

  2. Aidan Gomez56

    Cohere CEO Aidan Gomez 引用 Sam Altman 关于认同 Dario 前沿限速、承诺接受独立评估机构的推文,并以讽刺口吻逐条批评:要求对手开放员工级访问、以安全为由关停不够安全的竞争者,以及中国不遵守就断供芯片。作者称这些想法是卡特尔式的做法。

    引用Sam Altman@sama

    I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.

  3. Jakub Pachocki71

    OpenAI 首席科学家 Jakub Pachocki 以一个爱心符号转发了 Dario Amodei 的新文章《We Must Pace the Frontier》,后者主张 AI 行业应放慢速度,并提出三部分计划,Anthropic 单方面承诺其中第一步,向第三方评估者提供永久、员工级别的系统访问权限,用于验证安全措施执行、报告事故并评估模型训练期间的对齐情况。全文见 https://darioamodei.com/post/we-must-pace-the-frontier。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

9月12日周六
  1. Thinking Machines42

    我们自己的 @johnschulman2 与 Dwarkesh 对话,讨论随着模型不断进步和自我改进,人类判断力在哪些方面仍然重要:教它们处理混乱的现实世界任务,用品味判断什么在长期内有效,以及最重要的——明确我们真正想要什么。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines

  2. Peter McCrory37

    这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、工人转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵

    引用modest proposal@modestproposal1

    Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"

9月11日周五
  1. a16z:News(RSS)43

    a16z:LP 为何错过 SpaceX、Anthropic 与 OpenAI 这一波 AI 浪潮

    a16z 指出,许多 LP 对 SpaceX、Anthropic 和 OpenAI 三家前沿模型公司几乎零敞口,而 SpaceX 上市后市值约 2 万亿美元,成为规模达此前纪录 10 倍的史上最大 VC 背景 IPO,Anthropic 估值 965B 美元、OpenAI 最近估值 852B 美元。作者认为,传统把风投控制在整体组合 5-10% 的资产配置框架已经破裂,LP 需要重新调整风投仓位。

  2. GitHub Blog30

    GitHub Copilot 应用新手指南:使用 diff、终端与浏览器面板

    GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让开发者无需在编辑器、终端和浏览器之间切换即可完成 AI 编码闭环。diff 面板以绿色标注新增、红色标注删除,支持接受改动、留言或让 Copilot 继续修改;终端面板可直接运行项目命令并支持多窗口;浏览器面板可用 Pick & Polish 工具选中元素并让智能体调整。

  3. a16z:News(RSS)32

    a16z:雇主开始寻找新型健康保险计划,AI 正在降低建计划门槛

    a16z 发文指出,随着保费每年上涨 10% 以上,多数雇主正开始寻找替代方案,或转向低成本健康计划,或彻底放弃传统健康保险。这一规模达 1 万亿美元、覆盖 1.5 亿以上美国人的雇主医保市场,正因 AI 降低建计划与运营的固定成本门槛而出现代际替换机会,催生一批新型替代健康计划(AHP)、挑战者 PBM 和现代化基础设施平台。

  4. Sierra:Blog(RSS)42

    Sierra 推出下一代多模态智能体:随对话形态变化的界面

    Sierra 发布多模态智能体,将语音、文本和可视化整合进同一段对话,并自动判断何时切换模式,用户无需重来或重复表述。该智能体一次构建即可部署到所有渠道,视觉组件同样通用;其 MCP UI 集成支持把产品卡片、对比表格、日历和表单直接嵌入对话,组件由企业自行设计和托管,更新后自动同步,无需重新部署或为各平台维护不同版本。

  5. NVIDIA Technical Blog(开发者技术博客 · RSS)29

    全栈 NIM 优化如何在 Nemotron 3 Ultra 上支撑 2.5 倍并发用户

    NVIDIA 通过全栈 NIM 优化,在 Nemotron 3 Ultra 上实现 2.5 倍并发用户量。该优化针对生产环境部署大语言模型时,在现有 GPU 基础设施上提升并发服务能力并保持交互响应速度的需求,对提示词长、上下文跨步骤复用的智能体 AI 工作负载尤为关键。

9月10日周四
  1. Peter McCrory52

    Anthropic 首席经济学家 Peter McCrory 与 Jack Clark 对谈其 AI 经济影响情景研究。他表示目标不是做预测,而是理解可能结果的区间及其出现的条件,希望厘清对不确定未来的分歧来源;引用内容提到研究情景从影响很小到 2030 年 GDP 增长 15%、知识工作者失业率达 18%。

    引用John Burn-Murdoch@jburnmurdoch

    New from us: Anthropic just published scenarios for AI’s possible economic impacts, which range from minimal, to explosive GDP growth of 15% by 2030 as knowledge-worker unemployment hits 18%. I sat down with their co-founder Jack Clark to pick his brains on how they’re thinking about all of this.

  2. SiliconFlow70

    DeepSeek-V4.1-Flash 在 SiliconFlow 上线,提供 Day 0 支持。模型为 552B MoE,prefill 阶段约激活 8B、decode 阶段约激活 16B,支持原生视觉与 1M 上下文窗口,KV cache 占用相比 V4 Flash 约缩小 4 倍,采用 MIT 许可证,主打高吞吐生产级推理。

    推荐理由:上线方直接给出参数结构、上下文窗口、KV cache 对比和许可证信息,读者可据此评估实际部署选型。