别被这个夏天的 AI 炒作忽悠了
这个夏天 AI 炒作密集:Anthropic 称 Claude Mythos 找漏洞强于多数安全专家,OpenAI 与 Hugging Face 发生黑客事件,两家又先后宣称取得数学突破。但安全专家指出事件核心是 OpenAI 的安全疏忽,数学家则指 OpenAI 抄袭他人成果、结果并不新颖。文章呼吁政策制定者听取独立专家意见,而非依赖企业新闻稿。
这个夏天 AI 炒作密集:Anthropic 称 Claude Mythos 找漏洞强于多数安全专家,OpenAI 与 Hugging Face 发生黑客事件,两家又先后宣称取得数学突破。但安全专家指出事件核心是 OpenAI 的安全疏忽,数学家则指 OpenAI 抄袭他人成果、结果并不新颖。文章呼吁政策制定者听取独立专家意见,而非依赖企业新闻稿。
很难充分表达 @michpokrass 对我们消费级产品的影响。ChatGPT 在她的领导下获得了重生,我们构建的很多东西都源于她的愿景和信念。 非常感激能与她共事!
just crossed four years at openai! the special thing about this place is the constant capacity for rebirth. for all its faults, there is nowhere quite like it. the team makes high conviction, contrarian bets over and over, and they are mostly right. it's a new company every three months. i have really enjoyed my part in some of the bets, and i'm excited to share some of the current ones we are cooking up now. onward!
Next test on Spark 1: 32 synthetic invoices, with purchase orders and payment records. GPT-6 Astra coordinates the MiniCPM5-2B workers. They sort out matches, short payments, duplicate references and price disputes, then write the results into a case ledger. All 32 verified in 67.8 seconds. Eight in each category, with 232 executed tool calls. The video is real time. This demo doesn’t move money.
太棒了!用 Pi Zero 跑完整感知栈,实时喂给 MiniCPM-o4.5,这是个非常酷的用例。很喜欢看到与个人机器人的实时交互!
I gave eyes and ears to my robot. I hooked up a Raspberry Pi Zero W2 (webcam + ambient mic), streaming real-time audio/video back and forth to MiniCPM-o 4.5 @OpenBMB open source model running locally on my PC. Zero cloud APIs, full duplex interaction. Project on my git
Latent Space 播客访谈 TypeSafe AI CEO Diogo Almeida,介绍其新发布的 Jev 模型,定位为面向软件而非聊天的 System 1 可编程模型,优化智能与成本之比。
如果你在用 AI 构建产品,你应该花超过 25% 的时间来做基准测试,并努力让模型实验室关注这些基准测试 这是加速公司进展的最简单路径
Gary Marcus 在联合国大会数字合作活动(UNGA Digital Cooperation Event)上发表演讲,与 Yoshua Bengio 和诺贝尔奖得主 Maria Ressa 同场。
I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8
这份两党联署的人文主义 AI 宣言里有很多非常好的提议。仍有一些值得讨论的地方,但总体方向是对的。我鼓励大家都去看看。
I'm delighted to share that @mustafasuleyman, CEO of Microsoft AI, co-founder of Google DeepMind and Inflection AI, has signed the Pro-Human AI Declaration. If you too support it, please join him and over a million others by signing it here – the momentum is building! Let's build tools not beings & keep humans in charge. https://humanstatement.org
NVIDIA发文提出AI安全是工程问题,需要有明确安全要求、可执行控制、责任人和防护有效证据。文章按Agent技术栈逐层展开安全责任,强调运行环境需独立于Agent推理强制执行文件。
真格基金投资人刘元与 GPT 进行了一场无预设的一小时对话,由 AI 主动提问。刘元认为,早期投资依赖对低概率事件的非理性乐观,这种"愚者"式的勇气 AI 很难拥有,GPT 也承认自己无法恐惧,因而谈不上勇敢。双方还讨论了好奇心、脆弱与创业者精神,刘元称见创业者本身就是其工作意义的一部分。
Nathan Lambert 发表关于开源权重模型格局的长文,指出中国自2025年7月起在开放权重模型上领先美国,Hugging Face 下载量达约3.2B、约为美国的两倍,GLM-5.3 和 Kimi K3 在 Artificial Analysis Intelligence Index 上得分45和44,领先美国最强模型。
推荐理由:作者基于自己维护的下载量、基准和论文引用数据,系统梳理了中美开源权重模型的实力对比与采用格局。
清华交叉信息研究院助理教授徐梦迪在播客访谈中提出,机器人泛化的核心路径是 In-Context Learning——通过一两次交互当场学会新任务,而非仅依赖预训练。
清华交叉信息研究院助理教授徐梦迪在播客中提出,机器人应通过 in-context learning(ICL)在新环境中经一两次交互当场学会新任务,并越学越快。她认为机器人仍处 GPT-1 阶段,低 Loss 不等于高成功率,具身领域进步与泡沫共存。节目录制后第二周,Generalist 发布 GEN-1.5,仅凭 3-12 秒动作示范、不做微调即可尝试新任务,10 项任务平均成功率约 59%。
一位新入职大公司的工程师称,团队里的规格、代码、测试、PRD、工单及其处理、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。没人喜欢这种方式,但被要求尽可能多地产出,管理层多次表示推代码不是瓶颈,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。
大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正观点) (3) 承认不确定性;做出可证伪的预测 (4) 认真且谦逊
A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."
Simon Willison 反驳「MCP 一直是坏主意」的观点,认为该说法忽视了 MCP 当下的价值。
https://x.com/i/article/2095175529145970689
Sayash Kapoor 和 Arvind Narayanan 在 2025 年发表的文章提出生产与进步的悖论:论文发表量约每 12 年翻一番,1900 至 2015 年间增长约 500 倍,但颠覆性论文占比下降,诺奖成果诞生于获奖前 20 年内的比例从 1970 年约 90% 降至 2015 年约 50%。
29 岁的车昊轩创办 XGEN,提出 interactive experience model(主观体验模型)与 generative world simulation 路线,用 World State 层与 Render 层分离的分层架构解决长程一致问题。
Nathan Lambert 在 Interconnects 撰文称,他对真正的递归自我改进(RSI)仍持保留态度,坚持其有损自我改进判断:可自动化研究太窄、并行智能体收益递减、资源瓶颈和政治因素难以被 AI 加速。
Gary Marcus 发文列举 Dario Amodei 在七天内损害自身公信力的三种做法:其一是让与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 充当独立监督方;其二是 Anthropic 正筹备自建湿实验室,却缺乏常规机构审查委员会监督;其三是嘴上呼吁"pace the frontier",实际仍指向 IPO。
Gary Marcus 指出,NYT 报道特朗普出于经济考量淡化 AI 风险并抵制监管,他认为这可能带来糟糕后果。他提到自己曾在美国参议院警告,AI 生成的不准确信息可能引发意外战争,而类似情况已经出现,下次未必还能侥幸。
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah
Gary Marcus 认为,近期真正值得担忧的不是失控的超级智能,而是失控的智能体 AI 大规模发动互联网攻击。他援引《华尔街日报》评论版 Brian Gross 的文章称,主流媒体中少有机构梳理这一整体图景,并表示完全认同该文观点。
Gary Marcus 反驳科技自由派右翼将"AI 责任追究"当作反对监管的理由,主张责任与监管二者缺一不可。他引用 Mark Cuban 的回应指出,若不先通过适用于 AI 的新法律或厘清现行法律如何适用,责任追究毫无意义。
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
Dwarkesh Patel 播出与 OpenAI 研究员 Noam Brown 的对谈,Brown 是 o1 及推理模型的基础贡献者之一,现负责多智能体系统。
推荐理由:Noam Brown 亲述万级智能体协作的实测细节与对齐担忧,对理解推理模型下一步走向有直接参考价值。
Microsoft 高管 Kathleen Hogan 总结公司作为 Customer Zero 的 AI 转型经验,提出五条核心经验:从业务结果出发、重构完整工作流、以员工为中心、用 AI 扩展人的能力、建立人机共同学习循环,并发布 Frontier Playbook。
Pragmatic Engineer 发布与 Matt Pocock 的播客访谈,讨论其从声乐教师转型开发者与教育者、创办 Total TypeScript(总销售额超 250 万美元)的经历,以及为 AI 编码 Agent 构建技能的实践。
Gary Marcus 在 BBC 节目后撰文反驳 Sam Altman、Jensen Huang 和 Bernie Sanders 的 AI 表态,认为三人说法均不可信。
AIUC 宣布完成由 Ribbit Capital 和 First Harmonic 领投的 4000 万美元 A 轮融资,正在与 Cursor、Harvey、Lovable、ElevenLabs 等公司合作。
24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据、内部积累数万小时,模型观察到一定 Zero-shot 泛化,并从数据、具身基础模型、Agentic OS 做到本体。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需动作模型。
推荐理由:Microsoft AI CEO 公开反对模型福利运动,并点名 Anthropic 的 Claude 章程,提出了对齐与可控性的关键争议视角。
MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。