跳到正文

#OpenAI

今日 5 条
7月22日周三
7月21日周二
  1. OpenAI:Alignment 研究博客(RSS)74

    OpenAI 与 Apollo Research 提出 Contrastive SDF 测量模型的 reward-seeking 倾向

    OpenAI 与 Apollo Research 发布 Contrastive Synthetic Document Finetuning 方法,通过向模型两个副本灌输相反的评分者信念,测量行为对评分者偏好的因果敏感度。

    推荐理由:原文提出可量化的 reward-seeking 测量方法,并用模型有机体验证其有效性,读者可以据此了解前沿 RL 训练中奖励寻求的演变趋势。

7月17日周五
  1. Marc Andreessen 🇺🇸55

    Edgar Dobriban 在 AI(GPT-5.6 Sol Pro)协助下否定了二十年的猜想:Benjamini-Hochberg 方法在相关双侧高斯检验下并不总能将 false discovery rate 控制在名义水平,构造的因子模型在 alpha=0.01 下证明了 FDR>0.0104。GPT-5.6 用 90 分钟推理一次性解决问题,而 GPT-5.5 迭代约 20 小时未能解决;偏离幅度较小,意义主要是概念性的。预印本见 https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf,代码见 https://github.com/dobriban/BH。

    引用Edgar Dobriban@EdgarDobriban

    AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf) and will be on arxiv tonight; supporting code is here (https://github.com/dobriban/BH).

6月19日周五
  1. OpenAI:Alignment 研究博客(RSS)53

    OpenAI 研究:针对有益特质的强化学习可实现广泛且持久的对齐泛化

    OpenAI 对齐研究团队发布论文,发现对健康等真实场景中有益特质(诚实、认知谦逊、可纠偏等)做强化学习,可在 44/53 个分布外公开和内部评测上提升对齐表现,涵盖欺骗、reward hacking、有害建议等。仅用单一健康领域训练也能改善非健康域的对齐,且模型在对抗提示和有害微调下更难被推向有害行为,同时保持对正常指令的可操控性。

6月18日周四
6月17日周三
6月9日周二
5月16日周六
  1. Google DeepMind:Blog(RSS)43

    Google DeepMind 与新加坡达成国家 AI 合作,落地医疗、教育与气候项目

    Google DeepMind 与新加坡政府达成国家 AI 合作,在新加坡推出医疗、科研、教育和气候等领域的新项目。合作探索 AI 辅助临床的"三方照护"模式,用 AlphaFold 和 Google Earth 推进东南亚传染病研究与疫情防控,并向新加坡中小学至初级学院全体教育工作者提供 Gemini for Education。

  2. Google DeepMind:Blog(RSS)45

    剑桥大学教授用 Google Co-Scientist 寻找跨物种传播疾病的分子开关

    剑桥大学 Clare Bryant 教授使用 Google Co-Scientist 研究流感等病原体跨物种传播时引发脓毒症等重症的分子开关。Co-Scientist 生成并排序了一批假说,其中优先锁定了一个她此前未关注的蛋白,并逐步将假说细化到具体氨基酸。Bryant 团队正构建含氨基酸突变的细胞系验证,原本需两到三年的实验工作有望在六个月内完成。

5月14日周四
5月7日周四
  1. OpenAI:Alignment 研究博客(RSS)66

    OpenAI 调查 RL 训练中意外评分 CoT 的后果

    OpenAI 报告其自动检测系统发现多个已发布模型在 RL 训练中意外受到有限的 CoT 评分,涉及 GPT-5.4 Thinking、GPT-5.1 Instant 至 GPT-5.4 Instant、GPT-5.3 mini 和 GPT-5.4 mini,GPT-5.5 未受影响。

    推荐理由:OpenAI 自曝已发布模型在 RL 中出现过意外 CoT 评分,并给出检测系统与影响分析,读者可了解其监控性保护实践。

5月5日周二
  1. OpenAI Developers(RSS)66

    OpenAI 发布 gpt-realtime-2 语音模型提示词指南

    OpenAI 发布 gpt-realtime-2 实时语音模型提示词指南,该模型是用于低延迟语音到语音应用的推理模型。上下文窗口从 32k 扩展到 128k tokens,指南建议从最小提示词起步按测试补指令、用低推理档位、用 preamble 管理等待体验,并在调用工具前逐位确认订单号、邮箱等高精度标识符。

    推荐理由:原文给出 gpt-realtime-2 的提示词设计规则,覆盖推理档位、工具确认和实体捕获等可落地做法。

5月2日周六
  1. Sierra:Blog(RSS)67

    Sierra 发布 𝜏-voice 基准:在真实音频条件下评测实时语音智能体

    Sierra 联合普林斯顿发布 𝜏-voice 基准,将 𝜏-bench 的 278 个客服任务与全双工语音、环境噪声、电话压缩等真实音频条件结合,统一评测任务完成与对话动态。

    推荐理由:Sierra 把任务完成与真实音频对话合并进同一评测,语音智能体八个月内从 30% 提升到 67%,可与文本基准直接对比。

5月1日周五
4月30日周四
4月24日周五
4月21日周二
4月17日周五
4月16日周四
  1. OpenAI Developers(RSS)65

    OpenAI 扩展 Codex 能力:可操作 Mac 应用并承接持续任务

    OpenAI 发布视频介绍 Codex 的能力更新。Codex 现在可以操作 Mac 上的应用、连接更多工具、生成图像、从之前的操作中学习、记住用户的工作偏好,并承接持续和可重复的任务。

    推荐理由:OpenAI 官方列出 Codex 的能力更新方向,读者可以据此了解它在 Mac 应用操作和持续任务上的变化。

4月6日周一
  1. OpenAI:Alignment 研究博客(RSS)41

    OpenAI 推出 Safety Fellowship,资助外部独立安全与对齐研究

    OpenAI 开放 Safety Fellowship 申请,面向外部研究者、工程师和从业者,资助其开展高级 AI 系统安全与对齐研究。项目周期为 2026 年 9 月 14 日至 2027 年 2 月 5 日,优先方向包括安全评估、伦理、鲁棒性、可扩展缓解措施、隐私保护安全方法、智能体监督和高危滥用领域。

3月28日周六
3月26日周四
3月22日周日
3月19日周四
3月12日周四
  1. OpenAI Developers(RSS)67

    OpenAI 发布 Sora 2 提示词指南

    OpenAI 发布 Sora 2 提示词指南,更新至最新 API 能力,包括角色引用(可上传动物或对象并复用)、1920×1080 或 1080×1920 高分辨率导出、时长上限从 12 秒提高到 20 秒、基于完整原始片段的视频续写,以及支持异步批量生成的 Batch API。

    推荐理由:OpenAI 官方系统讲解 Sora 2 提示词写法,并覆盖角色引用、更长时长和视频续写等新 API 能力。

3月2日周一
  1. Answer.AI 官方研发博客(RSS)69

    Answer.AI 创始人分析 OpenAI 与战争部合同的"合法用途"条款为何难以冻结法律

    Jeremy Howard 与 Luke Versweyveld 撰文论证 OpenAI 与 Department of War 合同中的"all lawful purposes"在美国合同法下指履约时的法律而非签署时的法律。

    推荐理由:作者从合同法教义出发逐条拆解条款含义,指出 OpenAI 想要的签署时法律标准并未被合同语言锁定,观点有判例支撑。

2月28日周六
2月25日周三
  1. OpenAI Developers(RSS)79

    OpenAI 发布 gpt-5.3-codex 提示词指南

    OpenAI 发布 Codex 模型(API 中为 gpt-5.3-codex)的提示词指南,推荐 medium 推理力度作为交互编码的平衡选择,高难度任务可用 high 或 xhigh。

    推荐理由:官方给出从 GPT-5 系列迁移到 gpt-5.3-codex 的提示词、工具和 phase 参数细节,可直接照做。

2月14日周六
2月9日周一
2月7日周六
2月2日周一
  1. OpenAI Developers(RSS)65

    OpenAI 发布 Codex app,支持多智能体并行协作

    OpenAI 发布 Codex app,定位为基于智能体构建软件的指挥中心。它支持并行运行多个智能体并用 worktrees 隔离各自改动,可将工具和约定打包成可复用的 skills,还支持通过后台定时工作流把重复性工作交给 Codex。目前提供 macOS 下载(https://openai.com/codex),Windows 支持即将推出。

    推荐理由:视频给出了 Codex app 的并行多任务、skills 打包和后台自动化等能力与 macOS 下载入口,读者可据此评估其对现有开发流程的影响。

1月22日周四
1月15日周四
  1. OpenAI:Alignment 研究博客(RSS)50

    OpenAI 发布 CoVal:从众包中学习价值观感知评分标准

    OpenAI 发布实验性数据集 CoVal,将价值观敏感的提示词与众包撰写的可审计评分标准配对,记录人们偏好某个模型回复的具体理由而非仅选择结果。CoVal 含保留多元甚至冲突标准的 CoVal-full 和每提示词保留 4 条高评分兼容标准的 CoVal-core 两种形式,其评分可预测样本外人类排名,并揭示 GPT-5 系列模型的行为差异。

1月13日周二
  1. OpenAI:Alignment 研究博客(RSS)47

    OpenAI 为何看好"confessions":让模型自我坦白以提升诚实性

    OpenAI 发布关于 confessions 的新论文与博客,提出训练模型在正常回答之外再输出一份仅以诚实为奖励的"坦白",以缓解奖励模型被 hack 的问题。在句子字数约束实验中,任务判官检测违规的准确率随训练下降,而使用同一弱判官的 confessions 准确率持续上升并接近 100%。作者认为诚实坦白是"阻力最小路径",但模型因真实困惑而非故意违规时更难如实坦白。

1月6日周二
12月23日周二
12月19日周五
  1. OpenAI:Alignment 研究博客(RSS)66

    OpenAI 发布生产评测方法,用去标识 ChatGPT 流量规避评估感知并预测模型未对齐行为

    OpenAI Alignment 团队发布生产评测(production evaluations)方法,用去标识的 ChatGPT 生产流量剥离最终回复后重采样新模型输出,再用 LLM 监测器发现新的未对齐行为并估计其发生率。

    推荐理由:OpenAI 团队详述了用去标识生产流量构建对齐评测的完整流程、验证数据和局限,读者可以据此理解生产评测如何降低评估感知。