跳到正文

全部动态

今日 585 条
今天10月1日周四
  1. OpenAI:官网动态(RSS · 排除企业/客户案例)77

    OpenAI 披露并阻断一起有组织的对抗性模型蒸馏攻击

    OpenAI 披露已识别并阻断一起试图提取其模型 protected reasoning 的有组织攻击,最早活动出现在 7 月第一周。攻击者通过跨会话复制加密推理内容并请求解密转写等方式操纵模型交互,未入侵加密或数据库。

    推荐理由:OpenAI 官方复盘了一次模型蒸馏攻击的细节、归因和应对措施,读者可以借此了解针对 protected reasoning 的新型攻击手法。

  2. Claude Code:GitHub Releases(RSS)40

    Claude Code v2.1.286 发布

    Claude Code 发布 v2.1.286,为堆叠的权限请求加上"2 of 5"计数,并支持在全屏列表中点击"N more"行跳转。该版本修复了 claude --resume 和 --continue 在并行工具调用后丢失全部轮次、工具或 hook 返回对象/数字/布尔值时触发 API 400 错误,以及云会话因容器在 transcript 加载中被停止而无法唤醒等问题。

  3. Google AI:DEV 作者专属(RSS)59

    Ornith-1.0-9B vs Qwen3.5 vs Gemma4:CPU 本地实测对比

    作者在纯 CPU、32GB 内存的普通笔记本上,用 Ollama 以相同 Q4_K_M 量化对比 Ornith-1.0-9B、其基座模型 Qwen3.5-9B 和 Gemma4-12B,五个任务显示 Ornith 在 JSON 输出上最紧凑(16 token),但 bug 修复在未见过用例上出错,shell 命令与基座同样在含空格文件名上失败。

  4. Google AI:DEV 作者专属(RSS)39

    Sentinel-IR:AI 智能体运营省下数百万成本的非技术指南

    Sentinel-IR 是一种确定性数据压缩层,把代码、API 载荷和文档压缩成超紧凑的中间表示再喂给 LLM,充当上下文窗口的 ZIP 压缩。在 1,366 行的 12-billing-platform 测试文件上,原始 11,635 tokens 被压到 1,332 tokens,节省 88.6%;但 303 tokens(约 34 行)以下的微文件因压缩开销反而更贵。

  5. Google AI:DEV 作者专属(RSS)40

    我无法为代码辩护:i.c.stars 第四周的四份文档与产品规则设计

    i.c.stars 第四周产出四份非代码文档:缺陷日志、同行评审、三份 runbook 和一份 run of show,核心都是让工作能经受交接。作者在同行评审中复现对方八项发现中的六项,指出"擅长找 bug、弱于记录 bug"是常态。其产品贡献是主张检索系统只允许两种行为:附文档、章节和生效日期作答,或停止并转交人工,拒绝不是兜底而是同等功能。

  6. Google AI:DEV 作者专属(RSS)37

    图像质量评测:内容感知裁剪与居中裁剪的三种配方场景对比

    针对菜谱站点头像与菜品卡片,内容感知裁剪应作为默认方案,它能保留居中裁剪经常切掉的主体,但需存储裁剪框并提供人工调整路径。居中裁剪仅适用于拍摄时强制居中构图的流程,其确定性几何在异构上传中会稳定地切掉偏离中心的人脸或餐盘。Cloudinary、Imgix、ImageKit 等托管方案在运营边界上各有取舍,智能裁剪输出仍需审核与覆盖。

  7. Databricks:Blog(RSS)40

    Lakebase Postgres 成本优化实用指南

    Databricks 的 Lakebase Postgres 通过存储与计算分离架构实现成本优化:分支共享底层存储、自动扩缩容支持 scale to zero(数百毫秒恢复),关闭后按基线容量 25% 折扣计费。同步 Lakehouse 数据时建议只同步应用所需的活跃子集(如 60 天滚动窗口),并按数据新鲜度选择 Snapshot、Triggered 或 Continuous 三种同步模式。

  8. Google AI:DEV 作者专属(RSS)75

    一次 Agent 重构事故复盘:二十个正确改动如何掩盖了错误的假设

    作者复盘充电站地图去重任务的事故:一个由 Agent 编写、重构后测试全部通过的去重任务,因测试数据自造而未取自真实数据(9269 对重复记录中运营商名称仅 1 对匹配),导致约三分之一注册表站点在 100 米内存在重复显示,重构 78 分钟后被无审阅合并。

    推荐理由:作者以真实去重事故为底,给出从审代码转向审概念与真实数据的可迁移复核清单。

  9. Google AI:DEV 作者专属(RSS)23

    Portal 跨链流动性协议智能合约漏洞面分析:TVL 18.1 亿美元下的 10 类攻击向量

    针对 TVL 约 18.1 亿美元的跨链流动性路由协议 Portal 的智能合约漏洞面分析列出 10 类攻击向量,整体风险评级 7/10(高)。其中代理管理员可无时间锁升级、跨链桥回调重入、预言机操纵被列为 Critical 或 High 严重度,报告建议引入至少 48 小时时间锁升级机制并为桥接回调添加重入锁。

  10. Google AI:DEV 作者专属(RSS)47

    仅用 61 条数据构建气旋影响预测器:Cyclone Impact Forecaster 如何预测印度东海岸受灾人口

    Cyclone Impact Forecaster 用 61 条 EM-DAT 历史气旋记录训练出一个 2 参数对数线性模型,为北印度洋气旋影响半径内的每个地区输出受灾人口排名及 10-90% 预测区间,并叠加 XGBoost 的 6 小时风速强度预测,在 MapLibre GL 3D 卫星地球上可视化。

  11. MiniMax (official)48

    @HeyGen 发布 HeyGen Video,令人印象深刻!✨ 基于 MiniMax H3 构建,由 HeyGen 后训练,以更可及的成本为企业带来制作级视频。 很自豪能为这项工作提供基础,期待看到 HeyGen 团队将它带向多远。

    引用HeyGen@HeyGen

    We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog

  12. eric zakariasson53

    作者邀请试用 xAI 的 Grok Bot 市场工程类机器人,链接为 https://x.ai/bot/marketplace/engineering。其引用的 Grok Bot 官方内容称,Grok Bot 现在更适合软件开发,机器人可以把编码任务交给 Cursor、通过 GitHub 和 Origin 插件管理 PR,并分享构建内容的视频演示。截图显示工程分类下有多个可选机器人,包括 SWE by Cursor 等。

    引用Grok Bot@bot

    Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.

  13. Aravind Srinivas69

    Perplexity CEO Aravind Srinivas 宣布向所有人开放 Computer 的邮箱任务功能,无需 Perplexity 账号,限时免费执行所有通过邮箱委派的任务。用户将邮件发送、转发或抄送到 computer@perplexity.com,智能体会在后台完成任务并保留邮件上下文;每个任务以正常会话形式运行,可在网页和移动端查看,审计记录与应用内任务一致。

    引用Perplexity@perplexity_ai

    Computer now works in email. Send, forward, or cc computer@perplexity.com on any thread. Every email task runs as a normal session in Computer, viewable on web and mobile, with the same audit trail as any task in the app.

    推荐理由:Perplexity 开放无需账号的邮箱入口并限时免费,读者可以据此评估把任务转给 Computer 智能体执行的可行性。

  14. MiniMax (official)43

    基于 MiniMax H3 构建,@Creatify_Labs 的 Boreal-H3 是一款专为广告优化的视频模型,在更精准遵循创意简报的同时,保持产品和角色的一致性。 很高兴看到 MiniMax H3 成为更多面向特定行业的前沿模型的基础!✨

    引用Creatify Labs@Creatify_Labs

    Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.

  15. Ethan Mollick39

    这是当今时代最重要的问题之一:谁来决定AI的发展方向? Daron 主张在AI决策过程中引入更多民主参与,尽管这伴随着诸多挑战。

    引用Daron Acemoglu@DAcemogluMIT

    Second question on AI. We are told repeatedly that AI is going to transform every aspect of our lives – jobs, productivity, inequality, science, communication, daily activities, social order, and politics, among others. But this promise (or threat) is coupled with the rhetoric that such an important technology, with all of the risks and competitive pressures that it entails, should be left to experts or to “technocracy” (perhaps construed broadly to include some regulators). These two statements are hard to reconcile in a democratic society. If anything is half as important as AI is said to be (and I agree, AI is potentially very important and transformative), then involving democratic voice is essential. If something will shape our future in a democratic society, then its direction is for democratic institutions to decide. My instinct is that democratic voice is essential, and relying too much on technocracy could be both dangerous and counterproductive. The counterargument that AI’s direction can and should be entrusted to technocracy would go something along the following lines. First, democratic decision-making has become imperiled in our age of polarization. Second, AI is sufficiently complex that most citizens won’t have a deep enough understanding to meaningfully contribute to the debate (and even to the question of what we want from AI). Third, competition between different labs, and perhaps competition between the US and China, creates enough discipline for a socially beneficial direction of AI to be adopted. Fourth, today’s AI leaders are enlightened and ethical enough that within the framework created by competition, they can be broadly trusted. There are many aspects of this counterargument that I do not find convincing. Taking them in order: polarization can be overcome, and big decisions and challenges sometimes bring societies together; in fact, delegating key decisions to technocracy without democratic input may diminish trust in institutions and experts, and may worsen polarization. Second, democratic voice does not require citizens to write code or design new models; the debate should be informative enough that citizens can weigh in about what type of future they want and how they trade off the costs and benefits of different options. Third, competition doesn’t seem to be a good disciplining framework; on the contrary, competition sometimes brings the worst out of both organizations and people. Fourth, if three decades of work on political economy and institutions has taught me anything, it is that we should not bank on the ethical grounding of unconstrained leaders. But, still, I do not mean to immediately dismiss the technocracy option if there are more compelling arguments for it. The question is, then, whether there are any circumstances under which such important decisions can be delegated to AI experts and technocracy. One final secondary question: even if we managed to get democratic input in the United States or even in Europe, AI will shape the lives of everyone on this planet. How do we ensure that the voice of nearly 6 billion people who don’t live in the US, Europe and China also contributes to the debates on AI?

  16. HuggingFace Daily Papers(社区热门论文)56

    研究揭示分块 KV-cache 压缩导致模型出现周期性相位敏感弱位

    论文发现采用分块 KV-cache 压缩的模型存在相位敏感,检索准确率随压缩 stride 周期性波动,弱位置足以翻转答案。在 DeepSeek-V4 系列上,128K tokens 下最优与最差相位组差距达 15 到 40 个点;从零预训练的 0.6B 对照实验显示周期始终跟随 stride,full attention 各位置差距仅 6.1 个点而压缩模型可达 78 个点。

  17. HuggingFace Daily Papers(社区热门论文)3

    HuggingFace Daily Papers 论文 86 登上社区热门榜

    HuggingFace Daily Papers 收录的论文 86 登上社区热门榜。该条目来自 HuggingFace 每日论文板块,页面同时展示 Models、Datasets、Spaces 等入口及登录注册界面。原文未披露论文的具体模型、参数或评测信息。

  18. The Decoder:AI News(RSS)68

    Google 在 Gemini 中全球推出 Skills,取代 Gems

    Google 在 Gemini 聊天中全球推出 Skills,一种可由 AI 协助编写和完善的任务详细提示词格式,取代 Gems。用户将常用指令保存为 Skill,输入 "/" 加名称调用,Gemini 还能从历史对话生成 Skills、自动运行匹配的 Skill、链式组合多个 Skills,并支持文本文档、PDF 和图片作参考材料。

  19. Nature:Machine Learning 主题(RSS)49

    语音“衰老时钟”:AI 通过声音特征评估衰老速度

    科学家开发出一款“语音时钟”,通过音高、语速等数百项声音特征预测人的年龄,并计算“语音年龄差”。研究基于阿根廷、智利、哥伦比亚、墨西哥和秘鲁 2928 名西班牙语使用者的录音,用机器学习提取 700 多项随衰老和痴呆变化的语音特征。较大的语音年龄差与痴呆等认知问题强烈相关,成果已发表于 Science Advances。

  20. TechCrunch:AI(RSS)75

    Reddit 因 AI 爬虫将于 11 月 13 日关闭 RSS 并在 2027 年 3 月终止公共 API 访问

    Reddit 宣布因 RSS 被大规模爬取和自动化滥用,将于 11 月 13 日停止 RSS 支持,公共 API 访问也将于 2027 年 3 月终止。第三方应用和机器人开发者需在 2027 年 1 月 12 日前完成注册,否则会被移除 API 访问;版主被建议迁移到 Discord Relay Devvit 应用,Old Reddit 也将对最近 90 天未使用的登录用户限制访问。