跳到正文

#大佬观点

今日 39 条
9月22日周二
  1. MIT Technology Review · AI36

    别被这个夏天的 AI 炒作忽悠了

    这个夏天 AI 炒作密集:Anthropic 称 Claude Mythos 找漏洞强于多数安全专家,OpenAI 与 Hugging Face 发生黑客事件,两家又先后宣称取得数学突破。但安全专家指出事件核心是 OpenAI 的安全疏忽,数学家则指 OpenAI 抄袭他人成果、结果并不新颖。文章呼吁政策制定者听取独立专家意见,而非依赖企业新闻稿。

  2. Mark Chen20

    很难充分表达 @michpokrass 对我们消费级产品的影响。ChatGPT 在她的领导下获得了重生,我们构建的很多东西都源于她的愿景和信念。 非常感激能与她共事!

    引用Michelle Pokrass@michpokrass

    just crossed four years at openai! the special thing about this place is the constant capacity for rebirth. for all its faults, there is nowhere quite like it. the team makes high conviction, contrarian bets over and over, and they are mostly right. it's a new company every three months. i have really enjoyed my part in some of the bets, and i'm excited to share some of the current ones we are cooking up now. onward!

  3. OpenBMB23

    非常感谢分享!看到 MiniCPM5-2B 被用在这样实用的多智能体工作流中真的很酷——尤其是 workers 通过工具调用实际处理匹配、短付款、重复引用和争议。感谢你在测试和记录上付出的所有努力。对社区来说是个很好的案例 🙌

    引用Joey@aijoey

    Next test on Spark 1: 32 synthetic invoices, with purchase orders and payment records. GPT-6 Astra coordinates the MiniCPM5-2B workers. They sort out matches, short payments, duplicate references and price disputes, then write the results into a case ledger. All 32 verified in 67.8 seconds. Eight in each category, with 232 executed tool calls. The video is real time. This demo doesn’t move money.

  4. OpenBMB40

    太棒了!用 Pi Zero 跑完整感知栈,实时喂给 MiniCPM-o4.5,这是个非常酷的用例。很喜欢看到与个人机器人的实时交互!

    引用Mr Goodman@mrgoodmantweets

    I gave eyes and ears to my robot. I hooked up a Raspberry Pi Zero W2 (webcam + ambient mic), streaming real-time audio/video back and forth to MiniCPM-o 4.5 @OpenBMB open source model running locally on my PC. Zero cloud APIs, full duplex interaction. Project on my git

  5. Jeff Dean35

    感谢这场精彩的讨论,@dawnsongtweets!

    引用Dawn Song@dawnsongtweets

    I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8

9月21日周一
  1. Mustafa Suleyman39

    这份两党联署的人文主义 AI 宣言里有很多非常好的提议。仍有一些值得讨论的地方,但总体方向是对的。我鼓励大家都去看看。

    引用Max Tegmark@tegmark

    I'm delighted to share that @mustafasuleyman, CEO of Microsoft AI, co-founder of Google DeepMind and Inflection AI, has signed the Pro-Human AI Declaration. If you too support it, please join him and over a million others by signing it here – the momentum is building! Let's build tools not beings & keep humans in charge. https://humanstatement.org

  2. elsewhere:文章(RSS)48

    真格基金刘元与 GPT 对话一小时:AI 能否胜任早期投资

    真格基金投资人刘元与 GPT 进行了一场无预设的一小时对话,由 AI 主动提问。刘元认为,早期投资依赖对低概率事件的非理性乐观,这种"愚者"式的勇气 AI 很难拥有,GPT 也承认自己无法恐惧,因而谈不上勇敢。双方还讨论了好奇心、脆弱与创业者精神,刘元称见创业者本身就是其工作意义的一部分。

  3. Nathan Lambert:Interconnects(RSS)72

    Nathan Lambert 分析开源权重模型的中美力量对比

    Nathan Lambert 发表关于开源权重模型格局的长文,指出中国自2025年7月起在开放权重模型上领先美国,Hugging Face 下载量达约3.2B、约为美国的两倍,GLM-5.3 和 Kimi K3 在 Artificial Analysis Intelligence Index 上得分45和44,领先美国最强模型。

    推荐理由:作者基于自己维护的下载量、基准和论文引用数据,系统梳理了中美开源权重模型的实力对比与采用格局。

  4. elsewhere:文章(RSS)42

    对谈清华叉院助理教授徐梦迪:具身智能、世界模型与真正的泛化

    清华交叉信息研究院助理教授徐梦迪在播客中提出,机器人应通过 in-context learning(ICL)在新环境中经一两次交互当场学会新任务,并越学越快。她认为机器人仍处 GPT-1 阶段,低 Loss 不等于高成功率,具身领域进步与泡沫共存。节目录制后第二周,Generalist 发布 GEN-1.5,仅凭 3-12 秒动作示范、不做微调即可尝试新任务,10 项任务平均成功率约 59%。

  5. Simon Willison 博客37

    一位工程师爆料:大公司团队全靠 Claude Code 写代码,无人阅读

    一位新入职大公司的工程师称,团队里的规格、代码、测试、PRD、工单及其处理、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。没人喜欢这种方式,但被要求尽可能多地产出,管理层多次表示推代码不是瓶颈,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。

  6. Peter McCrory24

    大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正观点) (3) 承认不确定性;做出可证伪的预测 (4) 认真且谦逊

    引用Alex Imas@alexolegimas

    A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."

9月20日周日
9月19日周六
  1. Gary Marcus:The Road to AI We Can Trust(RSS)33

    Gary Marcus 批评 Dario Amodei 七天内三度失信

    Gary Marcus 发文列举 Dario Amodei 在七天内损害自身公信力的三种做法:其一是让与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 充当独立监督方;其二是 Anthropic 正筹备自建湿实验室,却缺乏常规机构审查委员会监督;其三是嘴上呼吁"pace the frontier",实际仍指向 IPO。

  2. Noam Brown52

    OpenAI 的 Noam Brown 引用他人对访谈的整理后补充四点看法:温度传感器通信只是学术例子,意在说明隔离无法给出绝对保证,应建立多层防御;他想讨论的是本应完全隔离的 agent 之间的协调,而非通过温度传感器窃取模型权重。他还表示从 HF 事件得到的教训是过度信任沙盒隔离而缺乏独立保障,air gap 是极强的防护,设计安全协议时宁可高估而非低估 AI。

    引用Fireside Alpha@firesidealpha

    OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah

  3. Gary Marcus:The Road to AI We Can Trust(RSS)24

    Gary Marcus:近期真正该担心的不是失控超级智能,而是失控的智能体 AI 大规模攻击互联网

    Gary Marcus 认为,近期真正值得担忧的不是失控的超级智能,而是失控的智能体 AI 大规模发动互联网攻击。他援引《华尔街日报》评论版 Brian Gross 的文章称,主流媒体中少有机构梳理这一整体图景,并表示完全认同该文观点。

9月18日周五
  1. Noam Brown40

    终于和 @dwarkesh_sp 一起做了一期关于多智能体的深度探讨!没有我在 @OpenAI 的队友们 @kevinleestone、@mikegmalek、@__eknight__、@amuellerml、@zhangir_azerbay、@CheukHeiChu 以及许多其他人在多智能体上的出色工作,这一切都不会发生。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?

9月17日周四
  1. elsewhere:文章(RSS)44

    对话王家伟:24 岁深朴智能首席科学家为何转身走向具身智能

    24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据、内部积累数万小时,模型观察到一定 Zero-shot 泛化,并从数据、具身基础模型、Agentic OS 做到本体。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需动作模型。

9月16日周三
  1. MIT News(RSS)23

    MIT 政治学者 Naoki Egami 如何用统计方法研究社会测量与 AI 工具误差

    MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。