跳到正文

#开源生态

今日 3 条
今天10月1日周四
  1. HuggingFace Daily Papers(社区热门论文)36

    DuoOPD:基于师生联合结果的多任务在线策略蒸馏

    DuoOPD 提出一种多任务在线策略蒸馏方法,由学生结果决定反馈方向、师生联合结果决定教师如何支持:仅教师答对时用其答案作为评分上下文,仅学生答对时按任务内共享权重强化整个回答。在 Qwen3 和 Llama 上,其平均宏准确率超越全部五个基线,较 OPD 分别提升 2.58 和 5.98 个百分点,并在科学计算、指令遵循与代码生成等任务混合上领先。消融显示,联合结果设计贡献了主要增益。

  2. Transluce(网页)55

    Transluce 将 Activation Oracles 扩展到 1.1T 参数模型,性能随模型与数据规模提升

    Transluce 训练激活预言机,用于分析模型内部激活以检测其行为,并扩展到最高 1.1T 参数的模型(包括 Kimi-K2.6 INT4),发现性能随模型规模、数据规模和数据质量提升。预言机在预测语言切换、检测编程智能体 reward hacking 等任务上取得成功,并采用全层激活读取架构,结合 LoRA 微调在 8xB200 上完成训练。

9月30日周三
  1. Prime Intellect(网页)50

    Prime Intellect 开源 General Agent:可自我进化的合成智能体环境

    Prime Intellect 开源首个 general-agent 环境,用 Synthesizer 与 Solver 双智能体博弈合成任务,目前包含 4,504 个任务、覆盖 1,040 个领域和 8,000 多个工具。每个任务由数据库、工具集、指令和带验证函数的 gold solution 构成,任务按难度分层进化,仅通过率落在校准难度区间的任务被采纳。

  2. Prime Intellect(网页)75

    Prime Intellect 用 nanoGPT speedrun 评测 18 个前沿模型的自主研究能力

    Prime Intellect 在 nanoGPT optimizer speedrun 上开展 153 次自主运行、覆盖 18 个前沿模型,每次运行使用 8xH200s、最长持续八天,baseline 为 3,290 steps,人类纪录为 2,600。

    推荐理由:实验规模和全部 trace 公开,模型差距主要来自实验执行而非想法本身,这个发现值得一看。

9月29日周二
9月16日周三
  1. Hugging Face:Blog(RSS)62

    Hugging Face 发布 ALTK-Evolve 一致性指南与 Consistency Analyzer,将智能体一致性差距减半

    IBM Research 与 Hugging Face 在 ALTK-Evolve 中推出一致性指南和 Consistency Analyzer,用于诊断和改善智能体重复运行的不稳定性。

    推荐理由:原文给出一致性差距的量化诊断方法与开源实现,读者可据此评估和改进智能体在重复运行下的可靠性。

8月28日周五
8月22日周六
8月21日周五
  1. Microsoft Research 博客(RSS)41

    微软 Skala 1.1 发布:训练数据增 2.5 倍,已接入 CP2K 并推进 Psi4、FHI-aims、ORCA、VASP 集成

    微软研究院发布深度学习交换关联泛函 Skala 1.1,训练数据比首个公开版本多 2.5 倍,在 GMTKN55 的 55 个类别中 32 项排名第一,加权平均误差 2.8 kcal/mol,精度超过当前领先的全局(范围分离)杂化泛函而保持半局域泛函的计算成本。

8月17日周一
8月13日周四
8月10日周一
7月29日周三
  1. BAIR:Berkeley AI Research Blog62

    从 CUDA 到 MLX:K-Search 如何把数十年内核经验带到 Apple Silicon

    IBM Research 基于 UC Berkeley Sky Lab 的 K-Search 框架扩展出 MLX 后端,并设计结构化 CUDA-to-MLX 翻译层,让进化式内核搜索把已有 CUDA 内核当作知识库适配到 Apple Silicon。

    推荐理由:K-Search 把 CUDA 内核优化经验迁移到 Apple Silicon,读者可了解跨平台内核搜索的方法与实测数据。

7月1日周三
  1. Jim Fan49

    ENPIRE -> ASPIRE,我们 Physical AutoResearch 系列的第二项工作。我们正在构建机器人自我改进的组件,一次一个 /skill。

    引用Jim Fan@DrJimFan

    Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an evolutionary search over control programs, and distill the best know-how into an ever-expanding library. ASPIRE is a new type of continual learning: "training" is skill refinement instead of gradient descent. "Trained model" is a repo of sensorimotor skills instead of floating weights. “Distributed training” is a panel of agents each practicing a different skill instead of sharded minibatches. Here's the beauty: ASPIRE gives the tired terms "sim2real transfer" and "cross-embodiment transfer" a whole new meaning. Bridging the sim-to-real gap is notoriously brutal. An end-to-end policy has to swallow both the visual shift (sim looks toyish next to a real camera) and the subtle contact physics it never quite gets right. ASPIRE sidesteps the mess, because it doesn't ship pixels or weights across the gap, but ships the know-how. The robot still has to practice in the real world, not zero-shot, but it gets there way faster because it isn't rediscovering the strategy from scratch. Same for going single-arm to bimanual hardware, which usually requires new data and retraining from zero. ASPIRE achieves up to ~10x cut in "transfer learning” tokens (yes, tokens are the new unit of *training* compute ;) Check out our gallery of 150+ tasks and 90+ skills the robots taught themselves, all on the website! Kind of wild that we can ship the "learned weights" as an HTML page rather than a GGUF. We'll open-source the full stack so your own robot library starts compounding from ours! Deep dive in thread:

  2. Jim Fan47

    Jim Fan 团队发布 ASPIRE,让编码智能体观察仿真与真实机器人的多模态感知轨迹,通过进化搜索控制程序并将最优经验蒸馏进持续扩张的技能库。ASPIRE 把"训练"重新定义为技能精炼而非梯度下降,跨仿真到真实与跨本体迁移时只传技能知识而非像素或权重,迁移学习 token 消耗最多降低约 10x。团队展示了 150+ 任务、90+ 技能,并将开源全栈。

6月18日周四
6月6日周六
  1. Ahead of AI(RSS)49

    LLM 研究论文:2026 年清单(1 月至 5 月)

    Ahead of AI 发布 2026 年 1 月至 5 月的 LLM 研究论文精选清单,按架构与模型设计、高效训练与扩展、推理效率与 KV Cache、稀疏注意力与长上下文、推理与测试时计算、强化学习与 RLVR、智能体系统与工具使用、编码智能体与软件工程、扩散语言模型、模型评估与基准等类别整理。

5月27日周三
4月16日周四
4月7日周二
3月17日周二
9月23日周二