跳到正文

#大佬观点

今日 38 条
9月30日周三
  1. Nathan Lambert25

    像这样的人的问题不是他们笨什么的,Timnit 拥有顶尖科学家的全部技能,问题在于他们所处的信息生态和同侪群体不鼓励对思想进行拷问。回音室是清晰思考的慢性死亡。

    引用Alec Stapp@AlecStapp

    Timnit Gebru doubles down on the "stochastic parrots" framing, saying you "cannot expect LLMs to be factual." As evidence to support this, she cites errors in... Google AI Overviews. We need to start a GoFundMe to pay for these people to have access to Opus 5.5 and Astra.

9月29日周二
  1. Aravind Srinivas38

    AI 安全是一个工程问题

    引用Perplexity@perplexity_ai

    Agent governance is an engineering problem. We’ve built safeguards into Perplexity’s infrastructure, harnesses, and tools, and put them to work across our products. Today we're sharing how we engineer safer agents: https://www.perplexity.ai/hub/blog/how-we-engineer-safer-agents

  2. Ars Technica:AI(RSS)22

    Mozilla Firefox 157 重新设计界面,负责人谈如何从 Chrome 争夺用户

    Mozilla 随 Firefox 157 在桌面和移动端推出界面重新设计,希望借此吸引隐私意识极客和开源倡导者之外的更广泛用户。Firefox 负责人 Ajit Varma 表示团队正借助 AI 工具提升开发速度,并恢复紧凑模式、增加自定义选项,让浏览器在体验上区别于基于 Chromium 的竞品。

  3. Ethan Mollick26

    嘿,Claude:“我让早期的 Claude 把《西线无战事》里的鱿鱼移除掉。现在你都能做电影之类的事了。我需要你展示一下你进步了多少” & 我把下面这条推文粘贴了进去 有些巧妙的东西,模型现在真的有幽默感了

    引用Ethan Mollick@emollick

    👀Claude handles an insane request: “Remove the squid” “The document appears to be the full text of the novel "All Quiet on the Western Front" by Erich Maria Remarque. It doesn't contain any mention of squid that I can see.” “Figure out a way to remove the 🦑​​​​​​​​​​​​​​​​“

  4. Diogo Almeida25

    再强调一下,我极度反对基准测试(: 引用推文例外处理——引用推文核心要点:别再搞“Jevbench”了,别再要求公开基准测试,它们完全没抓住Jevons悖论的重点,而且你不敢相信,你珍视的每一个基准测试有多容易被刷分。这是@CompleteSkeptic最惨痛的教训:选对任务胜过一切。

    引用Latent.Space@latentspacepod

    STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev and you won't believe how easy it is to game every benchmark you hold dear This is @CompleteSkeptic's bitterest lesson of all: picking the right task beats everything

  5. Thomas Wolf45

    OpenAI 的"地狱之夏"——@joedaroo 的好文 "准备的时间是现在,不是意外之后" "只给模型它需要的访问权限" "测试边界是否真的守得住" "把证据保留在模型控制范围之外" 安全团队与基础设施安全团队"应该是最好的朋友"

    引用Joe@joedaroo

    Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://x.com/i/article/2104258872957636608

  6. Simon Willison 博客42

    OpenAI 智能体安全负责人谈 AI 能力突跳带来的安全挑战

    OpenAI 智能体安全(Agent Security)负责人 @joedaroo 表示,模型在“cyber”“swarming”“message boards”等相关事件上能力跃升之快、之突然,远超团队预期。他指出安全态势需要时间积累,不只是加固系统,还要把安全融入公司文化,让人员随之演进。他呼吁各组织自问:人员、系统与流程能否承受 AI 能力的突然跃升,是否具备正确的事件响应与沟通机制。

  7. Arthur Mensch40

    只有开放生态才能保障 AI 的安全

    引用Jensen Huang@JensenHuang

    Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m

  8. TypeSafe AI46

    我们的座右铭背后有很多含义:Building Prod, Not God. 这项技术将改变世界,但这要靠勤勉的努力和创造力来实现,而不是靠故弄玄虚的诉求。 @a16z 与 @CompleteSkeptic 深入探讨了这一理念以及更多内容

    引用a16z@a16z

    TypeSafe AI's Diogo Almeida with a16z's Ben Horowitz and Martin Casado on Jev, the model built to live inside software: Diogo's elevator pitch for Jev is a simple question - where is all the automation? AI is unbelievably smart, but outside of chatbots and coding agents, it hardly touches any real work. His diagnosis is the industry built models that generate text for humans to read, and software can't consume that output. Jev reads natural language and returns a choice from a set of options with a confidence level assigned to each, so developers can build programs that reason about intent and make probabilistic decisions rather than relying on human interpretation. TypeSafe's philosophy is "We build prod, not God." 0:50 "Where the f**k is all the automation?" 2:50 Jev vs. Claude Code and Codex 6:55 Jev is a classifier and classifiers are sick 7:40 Chat vs. code: is Jev a slider? 9:00 Diogo: From mathlete to Kaggle to OpenAI 12:20 "We build prod, not God" 15:55 Reliability over demos 16:55 2021 thoughts: RLHF is AGI? 20:45 Optimizing for the wrong use case 21:50 Is the real world too messy to automate? 25:00 Nobody expected the Jev launch 26:35 Three kinds of reliability 28:05 Good at syntax, bad at architecture 30:00 The inverse SaaSpocalypse 33:40 Why coding agents automate so little 36:05 Probabilistic programming returns 38:45 Jev as the UDP-to-TCP layer for AI 40:20 The 5 stages of grief for embedding AI 41:30 Utopia: AI that actually does what you mean YouTube: https://youtu.be/Ut3LOjKNJaE @CompleteSkeptic @typesafeai @bhorowitz @martin_casado

9月28日周一
  1. AI as Normal Technology(RSS)48

    AI 存在性风险概率仍不可靠,不足以支撑政策制定

    针对当前 p(doom) 讨论推动政策关注的现象,该文重申 AI 存在性风险概率估计与 2024 年一样缺乏严谨性,不足以用于公共政策。作者指出,归纳法因不存在合适的参考类别而失效,概率本身不具权威性,政策制定者应认识到这些数字并非来自经过验证的模型或方法。

  2. Thomas Wolf46

    “如今,获取关于 AI 公司内部真实情况的经过验证的信息,显得尤为紧迫”——@RyanGreenblatt

    引用Ryan Greenblatt@RyanGreenblatt

    I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.