OpenAI 披露并阻断一起有组织的对抗性模型蒸馏攻击
OpenAI 披露已识别并阻断一起试图提取其模型 protected reasoning 的有组织攻击,最早活动出现在 7 月第一周。攻击者通过跨会话复制加密推理内容并请求解密转写等方式操纵模型交互,未入侵加密或数据库。
推荐理由:OpenAI 官方复盘了一次模型蒸馏攻击的细节、归因和应对措施,读者可以借此了解针对 protected reasoning 的新型攻击手法。
OpenAI 披露已识别并阻断一起试图提取其模型 protected reasoning 的有组织攻击,最早活动出现在 7 月第一周。攻击者通过跨会话复制加密推理内容并请求解密转写等方式操纵模型交互,未入侵加密或数据库。
推荐理由:OpenAI 官方复盘了一次模型蒸馏攻击的细节、归因和应对措施,读者可以借此了解针对 protected reasoning 的新型攻击手法。
Claude Code 发布 v2.1.286,为堆叠的权限请求加上"2 of 5"计数,并支持在全屏列表中点击"N more"行跳转。该版本修复了 claude --resume 和 --continue 在并行工具调用后丢失全部轮次、工具或 hook 返回对象/数字/布尔值时触发 API 400 错误,以及云会话因容器在 transcript 加载中被停止而无法唤醒等问题。
作者在纯 CPU、32GB 内存的普通笔记本上,用 Ollama 以相同 Q4_K_M 量化对比 Ornith-1.0-9B、其基座模型 Qwen3.5-9B 和 Gemma4-12B,五个任务显示 Ornith 在 JSON 输出上最紧凑(16 token),但 bug 修复在未见过用例上出错,shell 命令与基座同样在含空格文件名上失败。
Sentinel-IR 是一种确定性数据压缩层,把代码、API 载荷和文档压缩成超紧凑的中间表示再喂给 LLM,充当上下文窗口的 ZIP 压缩。在 1,366 行的 12-billing-platform 测试文件上,原始 11,635 tokens 被压到 1,332 tokens,节省 88.6%;但 303 tokens(约 34 行)以下的微文件因压缩开销反而更贵。
i.c.stars 第四周产出四份非代码文档:缺陷日志、同行评审、三份 runbook 和一份 run of show,核心都是让工作能经受交接。作者在同行评审中复现对方八项发现中的六项,指出"擅长找 bug、弱于记录 bug"是常态。其产品贡献是主张检索系统只允许两种行为:附文档、章节和生效日期作答,或停止并转交人工,拒绝不是兜底而是同等功能。
针对菜谱站点头像与菜品卡片,内容感知裁剪应作为默认方案,它能保留居中裁剪经常切掉的主体,但需存储裁剪框并提供人工调整路径。居中裁剪仅适用于拍摄时强制居中构图的流程,其确定性几何在异构上传中会稳定地切掉偏离中心的人脸或餐盘。Cloudinary、Imgix、ImageKit 等托管方案在运营边界上各有取舍,智能裁剪输出仍需审核与覆盖。
作者用 PyTorch 从零实现并训练了一个 124M 参数的 GPT-2 风格 decoder-only Transformer,配置为 768 维嵌入、12 头注意力、12 层、词表 50,257、最大序列长度 512。
Databricks 的 Lakebase Postgres 通过存储与计算分离架构实现成本优化:分支共享底层存储、自动扩缩容支持 scale to zero(数百毫秒恢复),关闭后按基线容量 25% 折扣计费。同步 Lakehouse 数据时建议只同步应用所需的活跃子集(如 60 天滚动窗口),并按数据新鲜度选择 Snapshot、Triggered 或 Continuous 三种同步模式。
作者复盘充电站地图去重任务的事故:一个由 Agent 编写、重构后测试全部通过的去重任务,因测试数据自造而未取自真实数据(9269 对重复记录中运营商名称仅 1 对匹配),导致约三分之一注册表站点在 100 米内存在重复显示,重构 78 分钟后被无审阅合并。
推荐理由:作者以真实去重事故为底,给出从审代码转向审概念与真实数据的可迁移复核清单。
TypeSafe AI 推出 JEV,一款不生成文本、只对预定义选项打分的决策模型,输入百万 token 约 $0.042、输出 token 不收费,毫秒级返回结果。
印度独立开发者推出 ApexHierarchy,用 8 个层级(Town 到 Global 加两个特殊层级)和 Low/Mid/High/Top 档位,把个人、公司、机构和国家放进同一套权力标尺,目前收录 777 个档案、3127 条连接和 1648 个来源。
针对 TVL 约 18.1 亿美元的跨链流动性路由协议 Portal 的智能合约漏洞面分析列出 10 类攻击向量,整体风险评级 7/10(高)。其中代理管理员可无时间锁升级、跨链桥回调重入、预言机操纵被列为 Critical 或 High 严重度,报告建议引入至少 48 小时时间锁升级机制并为桥接回调添加重入锁。
Cyclone Impact Forecaster 用 61 条 EM-DAT 历史气旋记录训练出一个 2 参数对数线性模型,为北印度洋气旋影响半径内的每个地区输出受灾人口排名及 10-90% 预测区间,并叠加 XGBoost 的 6 小时风速强度预测,在 MapLibre GL 3D 卫星地球上可视化。
每周 Midjourney 办公时间 - 9/30 https://x.com/i/spaces/1nxeLMvZvQrJX
We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog
每家公司的客服人员即将被Dots & Muses等(AI智能体)用语音/聊天来谈判争取更优惠交易而淹没。人们把这类事委托给智能体并因此省钱的报道正在不断涌现,而且只会越来越火。
Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.
Computer now works in email. Send, forward, or cc computer@perplexity.com on any thread. Every email task runs as a normal session in Computer, viewable on web and mobile, with the same audit trail as any task in the app.
推荐理由:Perplexity 开放无需账号的邮箱入口并限时免费,读者可以据此评估把任务转给 Computer 智能体执行的可行性。
The ClawCast — Jev & OpenClaw Enterprise(第12期)https://x.com/i/broadcasts/1mxPaZewLAmKN
可分享的个人资料将你的 Sites 和插件汇集在 ChatGPT 中,让他人能够找到并复用你的作品。 做好了?就秀出来。
如今机器人能完成哪些工作?这对未来数年的经济结构又意味着什么? @rclegateyang 和 Maxim Massenkoff 的新研究今日发布,探讨了这些问题。
Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.
用 Manus Game Dev 把一个想法变成好玩的 3D 赛车游戏。🏎️ 赛车呢?由我们的合作伙伴 @tripoai 生成的 3D 模型。
这是当今时代最重要的问题之一:谁来决定AI的发展方向? Daron 主张在AI决策过程中引入更多民主参与,尽管这伴随着诸多挑战。
Second question on AI. We are told repeatedly that AI is going to transform every aspect of our lives – jobs, productivity, inequality, science, communication, daily activities, social order, and politics, among others. But this promise (or threat) is coupled with the rhetoric that such an important technology, with all of the risks and competitive pressures that it entails, should be left to experts or to “technocracy” (perhaps construed broadly to include some regulators). These two statements are hard to reconcile in a democratic society. If anything is half as important as AI is said to be (and I agree, AI is potentially very important and transformative), then involving democratic voice is essential. If something will shape our future in a democratic society, then its direction is for democratic institutions to decide. My instinct is that democratic voice is essential, and relying too much on technocracy could be both dangerous and counterproductive. The counterargument that AI’s direction can and should be entrusted to technocracy would go something along the following lines. First, democratic decision-making has become imperiled in our age of polarization. Second, AI is sufficiently complex that most citizens won’t have a deep enough understanding to meaningfully contribute to the debate (and even to the question of what we want from AI). Third, competition between different labs, and perhaps competition between the US and China, creates enough discipline for a socially beneficial direction of AI to be adopted. Fourth, today’s AI leaders are enlightened and ethical enough that within the framework created by competition, they can be broadly trusted. There are many aspects of this counterargument that I do not find convincing. Taking them in order: polarization can be overcome, and big decisions and challenges sometimes bring societies together; in fact, delegating key decisions to technocracy without democratic input may diminish trust in institutions and experts, and may worsen polarization. Second, democratic voice does not require citizens to write code or design new models; the debate should be informative enough that citizens can weigh in about what type of future they want and how they trade off the costs and benefits of different options. Third, competition doesn’t seem to be a good disciplining framework; on the contrary, competition sometimes brings the worst out of both organizations and people. Fourth, if three decades of work on political economy and institutions has taught me anything, it is that we should not bank on the ethical grounding of unconstrained leaders. But, still, I do not mean to immediately dismiss the technocracy option if there are more compelling arguments for it. The question is, then, whether there are any circumstances under which such important decisions can be delegated to AI experts and technocracy. One final secondary question: even if we managed to get democratic input in the United States or even in Europe, AI will shape the lives of everyone on this planet. How do we ensure that the voice of nearly 6 billion people who don’t live in the US, Europe and China also contributes to the debates on AI?
Arena 宣布 OpenAI 的 GPT-6.1 Sol (Max) 在 Code Arena: WebDev 榜以 1759 分排第 3,混合价格 $8/MToken。
GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today.
推荐理由:Arena 用自家榜单数据对比 GPT-6.1 Sol (Max) 与前代及竞品的价格性能位置,读者可据此评估其成本效率变化。
论文发现采用分块 KV-cache 压缩的模型存在相位敏感,检索准确率随压缩 stride 周期性波动,弱位置足以翻转答案。在 DeepSeek-V4 系列上,128K tokens 下最优与最差相位组差距达 15 到 40 个点;从零预训练的 0.6B 对照实验显示周期始终跟随 stride,full attention 各位置差距仅 6.1 个点而压缩模型可达 78 个点。
HuggingFace Daily Papers 收录的论文 86 登上社区热门榜。该条目来自 HuggingFace 每日论文板块,页面同时展示 Models、Datasets、Spaces 等入口及登录注册界面。原文未披露论文的具体模型、参数或评测信息。
Google 在 Gemini 聊天中全球推出 Skills,一种可由 AI 协助编写和完善的任务详细提示词格式,取代 Gems。用户将常用指令保存为 Skill,输入 "/" 加名称调用,Gemini 还能从历史对话生成 Skills、自动运行匹配的 Skill、链式组合多个 Skills,并支持文本文档、PDF 和图片作参考材料。
Reddit 为打击爬虫和自动化流量进一步限制 Old Reddit,未来几个月内用户需已登录且在过去六个月内使用过 Old Reddit 才能继续使用,版主可豁免近期使用要求。
科学家开发出一款“语音时钟”,通过音高、语速等数百项声音特征预测人的年龄,并计算“语音年龄差”。研究基于阿根廷、智利、哥伦比亚、墨西哥和秘鲁 2928 名西班牙语使用者的录音,用机器学习提取 700 多项随衰老和痴呆变化的语音特征。较大的语音年龄差与痴呆等认知问题强烈相关,成果已发表于 Science Advances。
The Verge 分析称,OpenAI 和 Meta 都在用可爱的软件智能体为实体 AI 硬件试水。OpenAI 与 Jony Ive 合作开发硬件,据公开文件计划最早 2027 年 2 月出货,并在 DevDay 发布了带彩色眼睛团子形象的智能体平台 Dots,Altman 称未来把它做进硬件是合理假设。
Reddit 宣布因 RSS 被大规模爬取和自动化滥用,将于 11 月 13 日停止 RSS 支持,公共 API 访问也将于 2027 年 3 月终止。第三方应用和机器人开发者需在 2027 年 1 月 12 日前完成注册,否则会被移除 API 访问;版主被建议迁移到 Discord Relay Devvit 应用,Old Reddit 也将对最近 90 天未使用的登录用户限制访问。
语音 AI 创业公司 ElevenLabs 宣布以 $22B 估值进行 $300 million 的 tender offer,允许员工出售部分已归属股份,估值为此前 $11B 的两倍。