跳到正文

全部动态

今日 67 条
9月22日周二
  1. Jeff Dean35

    感谢这场精彩的讨论,@dawnsongtweets!

    引用Dawn Song@dawnsongtweets

    I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8

9月21日周一
  1. Mustafa Suleyman39

    这份两党联署的人文主义 AI 宣言里有很多非常好的提议。仍有一些值得讨论的地方,但总体方向是对的。我鼓励大家都去看看。

    引用Max Tegmark@tegmark

    I'm delighted to share that @mustafasuleyman, CEO of Microsoft AI, co-founder of Google DeepMind and Inflection AI, has signed the Pro-Human AI Declaration. If you too support it, please join him and over a million others by signing it here – the momentum is building! Let's build tools not beings & keep humans in charge. https://humanstatement.org

  2. Import AI22

    Import AI 473:美国超级智能战略、人脑组织植入小鼠大脑与机器诠释学

    RAND 发布报告,建议美国在通往超级智能的不确定路径上采取"自由行动"战略,通过构建人-AI 生态系统、AI 安全架构、改造国家安全体系及提升公民应对能力来保留所有选项。报告还梳理了共存、拒止、加速三大类共七种原型战略,并列出危险临近程度、共存可行性、约束可行性、决定性战略优势、压制可行性五项关键不确定性。此外,研究人员将人类脑组织培育进小鼠大脑,用于实验研究。

  3. elsewhere:文章(RSS)48

    真格基金刘元与 GPT 对话一小时:AI 能否胜任早期投资

    真格基金投资人刘元与 GPT 进行了一场无预设的一小时对话,由 AI 主动提问。刘元认为,早期投资依赖对低概率事件的非理性乐观,这种"愚者"式的勇气 AI 很难拥有,GPT 也承认自己无法恐惧,因而谈不上勇敢。双方还讨论了好奇心、脆弱与创业者精神,刘元称见创业者本身就是其工作意义的一部分。

  4. Nathan Lambert:Interconnects(RSS)72

    Nathan Lambert 分析开源权重模型的中美力量对比

    Nathan Lambert 发表关于开源权重模型格局的长文,指出中国自2025年7月起在开放权重模型上领先美国,Hugging Face 下载量达约3.2B、约为美国的两倍,GLM-5.3 和 Kimi K3 在 Artificial Analysis Intelligence Index 上得分45和44,领先美国最强模型。

    推荐理由:作者基于自己维护的下载量、基准和论文引用数据,系统梳理了中美开源权重模型的实力对比与采用格局。

  5. elsewhere:文章(RSS)56

    葬AI评世界模型热潮:上不了大模型桌的人才另开一桌

    葬AI发文批评世界模型已成炒作概念,认为其起因是一批做不了大语言模型竞争的团队生造赛道,代表人物李飞飞和杨立昆做的其实是不同的东西。文章归纳了四类讲世界模型故事的人,称多数产品只是后训练开源视频模型,且因MiniMax H3开源可后训练才近期密集宣发,并预告将推出直播间Bench实测实时生成视频模型。

  6. elsewhere:文章(RSS)42

    对谈清华叉院助理教授徐梦迪:具身智能、世界模型与真正的泛化

    清华交叉信息研究院助理教授徐梦迪在播客中提出,机器人应通过 in-context learning(ICL)在新环境中经一两次交互当场学会新任务,并越学越快。她认为机器人仍处 GPT-1 阶段,低 Loss 不等于高成功率,具身领域进步与泡沫共存。节目录制后第二周,Generalist 发布 GEN-1.5,仅凭 3-12 秒动作示范、不做微调即可尝试新任务,10 项任务平均成功率约 59%。

  7. Simon Willison 博客37

    一位工程师爆料:大公司团队全靠 Claude Code 写代码,无人阅读

    一位新入职大公司的工程师称,团队里的规格、代码、测试、PRD、工单及其处理、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。没人喜欢这种方式,但被要求尽可能多地产出,管理层多次表示推代码不是瓶颈,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。

  8. Peter McCrory24

    大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正观点) (3) 承认不确定性;做出可证伪的预测 (4) 认真且谦逊

    引用Alex Imas@alexolegimas

    A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."

9月20日周日
  1. elsewhere:文章(RSS)60

    峰瑞李丰:全球流动性见顶后,AI周期进入后半段的投资逻辑

    峰瑞资本李丰撰文分析,认为2026年三季度全球流动性接近见顶,美元主导的资本市场进入存量博弈尾部,AI产业周期进入后半段。文章回顾2020年天量流动性如何催生本轮AI热潮,列举科技巨头资本开支转折的五个信号(如Alphabet二季度自由现金流转负59亿美元),提出投资重心应从讲大故事转向能靠AI赚钱的方向,如AI+应用、生物医疗与AI交叉及SaaS的AI化。

9月19日周六
  1. Gary Marcus:The Road to AI We Can Trust(RSS)33

    Gary Marcus 批评 Dario Amodei 七天内三度失信

    Gary Marcus 发文列举 Dario Amodei 在七天内损害自身公信力的三种做法:其一是让与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 充当独立监督方;其二是 Anthropic 正筹备自建湿实验室,却缺乏常规机构审查委员会监督;其三是嘴上呼吁"pace the frontier",实际仍指向 IPO。

  2. Dan Hendrycks37

    “最妙的是把恶意对准他每天遇到的近邻,而把善意推向遥远的边缘,推向他素不相识的人。于是恶意变得完全真实,而善意大体上是想象出来的。”——C.S. Lewis,以恶魔的视角写作

    引用banteg@banteg

    >be me >discover effective altruism >apparently normal charity is inefficient >why donate to random sad thing when spreadsheet can tell you optimal sad thing >fair enough >buy mosquito nets >save lives >numbers look good >feel powerful >couple years later >someone asks an innocent question >why only count people alive today >huh >future people matter too >obviously >my grandchildren shouldn't matter less just because they haven't spawned yet >reasonable.jpg >keep following logic >what about their grandchildren >also yes >what about people in 500 years >sure >5000 years >why not >500 million years >starting to get weird but morality is morality >open calculator >humanity could survive for an astronomically long time >could colonize galaxy >could have trillions upon trillions of descendants >maybe digital people too >maybe simulated civilizations >maybe dyson spheres full of happy uploaded minds >calculator starts smoking >realize currently living humans are rounding error >8 billion people suddenly looking extremely beta >future contains potentially 10^something people >can't even fit beneficiaries in google sheets >new moral priority unlocked >protect the long-term future >stop thinking in units of "people helped" >start thinking in "fraction of cosmic endowment preserved" >malaria? >terrible >but only kills existing humans >AI extinction could delete the entire light cone >nuclear war could permanently derail civilization >bad institutions could lock in terrible values for ten million years >someone invents wrong constitution in 2140 >quadrillions suffer >better fund governance workshop now >friend says maybe we should improve hospitals >explain opportunity cost >friend says hospitals are full of actual sick people >explain scope sensitivity >friend stops inviting me to dinner >need to decide what to fund >easy >expected value >suppose project has one in a million chance of preventing extinction >sounds tiny >but extinction destroys 10^50 future lives >multiply >mother of god >$10 million project has expected value of several galaxies >charity evaluation complete >someone asks where the one-in-a-million number came from >expert judgement >which expert >us >how calibrated >extremely thoughtfully >reduce estimate to one in ten million to be conservative >still beats curing cancer by 38 orders of magnitude >epistemic robustness achieved >someone says maybe project doesn't work >assign 20% chance >still astronomical >maybe project makes problem worse >assign 5% chance >still astronomical >why 5 >because 30 felt pessimistic >publish 46-page report >contains seventeen sensitivity analyses >every sensitivity analysis begins after assuming intervention has positive sign >critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities >yes >that's literally why it's important >critic says the uncertainty might be structural rather than numerical >make probability smaller >critic says no, I mean maybe your model is wrong >make probability smaller again >critic begins rubbing temples >discover AI safety >perfect longtermist cause >AI might kill everyone >or create utopia >or seize galaxy >or tile universe with paperclips >or create billions of conscious software minds >finally a problem with numbers big enough for me >start AI safety nonprofit >mission: prevent dangerous AI >hire smartest people available >smartest people immediately start building better AI to understand dangerous AI >interesting >we must understand capabilities to understand safety >we must scale models to study alignment >we must race ahead so less responsible actors don't get there first >we must deploy systems to learn how deployment can go wrong >we must build the thing quickly because building the thing quickly is dangerous >outsider asks why the people most worried about AI apocalypse all work at AI companies >complicated field >company releases stronger model >very concerned >company begins training even stronger model >extremely concerned >company raises $14 billion >concern reaches unprecedented levels >need to influence government >future is at stake >normal democratic process too slow >politicians don't understand exponential curves >public doesn't understand x-risk >experts must guide them >who counts as expert >people who understand x-risk >who understands x-risk >our friends >someone objects that this seems politically convenient >explain we're representing future generations >future generations unavailable for comment >develop concept of value lock-in >terrifying possibility that one ideology controls civilization forever >therefore extremely important that civilization adopts correct values before lock-in >whose values >let's circle back >begin with impartial morality >end with small group of people deciding what quadrillions of hypothetical beings would want >beautiful arc >meanwhile actual humans keep doing annoying things >voting wrong >having parochial attachments >loving family more than strangers >caring about local community >getting upset when told their suffering is cosmically negligible >evolutionary biases everywhere >explain that moral intuition cannot be trusted >except intuition that future digital people count >and intuition that extinction is uniquely bad >and intuition that our probability estimates are sane >and intuition that our institutional choices improve the future >those intuitions survived peer review >someone donates $5k to local homeless shelter >inefficient >could have funded 0.0000000000003% of an AI governance researcher >think of all the simulated people you just killed >okay maybe don't phrase it that way publicly >PR team says "future generations deserve a voice" >much better >journalist asks what longtermism means >say "future people matter" >everyone agrees >great >journalist asks what follows from that >well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures >journalist raises eyebrow >return to "future people matter" >motte has entered the chat >critic: of course future people matter >me: glad we agree >critic: I don't agree that your institute knows how to help them >me: why do you hate our grandchildren >eventually notice uncomfortable implication >if future value dominates everything >then helping people today mostly matters through effects on future >education matters because future institutions >health matters because future productivity >democracy matters because future trajectory >human beings slowly become instrumental variables in their own moral philosophy >see starving child >feel compassion >check spreadsheet >child's direct welfare contribution negligible >but perhaps childhood nutrition improves national institutional quality >compassion restored >tell myself this is impartial altruism >one day assistant asks obvious question >"how do you know your intervention actually improves the far future?" >silence >open spreadsheet >increase column width >add confidence interval >assistant asks again >"no, I mean how do you know the sign is positive?" >stare into cosmic light cone >10^50 people staring back >none of them exist >none of them can tell me >none of them can falsify my assumptions >realize I have invented the perfect constituency >infinitely important >completely silent >and always represented by me

  3. Noam Brown52

    OpenAI 的 Noam Brown 引用他人对访谈的整理后补充四点看法:温度传感器通信只是学术例子,意在说明隔离无法给出绝对保证,应建立多层防御;他想讨论的是本应完全隔离的 agent 之间的协调,而非通过温度传感器窃取模型权重。他还表示从 HF 事件得到的教训是过度信任沙盒隔离而缺乏独立保障,air gap 是极强的防护,设计安全协议时宁可高估而非低估 AI。

    引用Fireside Alpha@firesidealpha

    OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah

  4. Gary Marcus:The Road to AI We Can Trust(RSS)24

    Gary Marcus:近期真正该担心的不是失控超级智能,而是失控的智能体 AI 大规模攻击互联网

    Gary Marcus 认为,近期真正值得担忧的不是失控的超级智能,而是失控的智能体 AI 大规模发动互联网攻击。他援引《华尔街日报》评论版 Brian Gross 的文章称,主流媒体中少有机构梳理这一整体图景,并表示完全认同该文观点。

9月18日周五
  1. GitHub Blog22

    GitHub Podcast 拆解 AI 热门观点:该不该读代码、RAG 是否已死、Skills 是否杀死了 MCP

    GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,MCP 提供工具与数据的标准接入,Skills 封装团队流程与最佳实践,二者可组合使用;RAG 并未消亡,检索能为模型提供训练数据之外的信息,减少 token 消耗并让回答更有依据。

  2. elsewhere:文章(RSS)50

    在AI艺术黑客松做《The Album》改编互动游戏,一位选手的复盘与反思

    作者参加超级发电站主办的AI艺术黑客松,做了一个改编自麦浚龙与谢安琪概念专辑《The Album》的实时互动游戏,玩家扮演酒保自由输入回应客人。技术方案采用2.5D等距视角,由剧本、独立subagent演员、导演程序和预制美术资产四层组成,demo部署在 https://elsewhere.news/the-album 并凭邀请码有限开放。

  3. Noam Brown40

    终于和 @dwarkesh_sp 一起做了一期关于多智能体的深度探讨!没有我在 @OpenAI 的队友们 @kevinleestone、@mikegmalek、@__eknight__、@amuellerml、@zhangir_azerbay、@CheukHeiChu 以及许多其他人在多智能体上的出色工作,这一切都不会发生。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?

9月17日周四
  1. elsewhere:文章(RSS)44

    对话王家伟:24 岁深朴智能首席科学家为何转身走向具身智能

    24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据、内部积累数万小时,模型观察到一定 Zero-shot 泛化,并从数据、具身基础模型、Agentic OS 做到本体。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需动作模型。

9月16日周三