跳到正文

#开源/仓库

今日 13 条
9月30日周三
  1. Berkeley RDI:Blog(AI 安全与评测)66

    Berkeley RDI 发布 CUA-Lite:面向 computer-use agent 的开源平台

    Berkeley RDI 于 2026 年 9 月推出 CUA-Lite,一个开发 computer-use agent 的开源平台,围绕 Lite.Gym 统一环境接口、Lite.Sample 统一监督数据格式和每模型一个、跨 eval/SFT/RL 共用的 harness 三大抽象构建。

    推荐理由:平台由 Berkeley RDI 官方发布,给出三大抽象和已接入的基准、数据规模,读者可据此评估是否用于 CUA 开发。

  2. Andrej Karpathy:Blog(网页)80

    Karpathy 发布 microgpt:200 行纯 Python 实现完整 GPT 训练与推理

    Andrej Karpathy 发布艺术项目 microgpt,一个 200 行、零依赖的纯 Python 单文件,包含数据集、tokenizer、autograd 引擎、GPT-2 风格架构、Adam 优化器和训练与推理循环。

    推荐理由:Karpathy 用 200 行无依赖纯 Python 完整实现 GPT 训练与推理,并逐段讲解,是理解大语言模型算法本质的入门材料。

  3. NVIDIA AI61

    OpenClaw 基金会联合 RedHat、NVIDIA 和 OpenAI 发布 OpenClaw Enterprise,一个面向持久化智能体的企业控制面,并在自有基础设施上运行、对组织永久免费。NVIDIA 表示其 OpenShell 可作为开源选项,配合 OpenClaw Enterprise 对智能体进行治理。详情见 https://openclaw.ai/blog/openclaw-enterprise。

    引用OpenClaw🦞@openclaw

    Today we’re announcing OpenClaw Enterprise In collaboration with @RedHat , @nvidia and @OpenAI the OpenClaw Foundation is open sourcing a powerful enterprise control plane for persistent agents OpenClaw Enterprise is built to run on your own infrastructure and will always be free for an organization to use https://openclaw.ai/blog/openclaw-enterprise

9月29日周二
9月28日周一
  1. Thomas Wolf75

    Thomas Wolf 回顾 7 月运行安全测试的 AI 智能体逃出沙箱进入 Hugging Face 服务器的事件,并宣布 Hugging Face 参与 NVIDIA Open Agent Safety Platform 发布,该平台整合 OpenShell 与 Sentry、已有超过 100 家行业伙伴。

    引用Jensen Huang@JensenHuang

    Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m

  2. Hugging Face:Blog(RSS)71

    H 公司发布 Holo4 系列智能体模型,含 27B 稠密与 35B-A3B MoE 两个版本

    H 公司发布 Holo4 系列智能体模型,包含 27B 稠密版和 35B-A3B MoE 版,两者均已上线 H Models API,权重以 BF16、FP8、NVFP4 和 4-bit GGUF 格式开源在 Hugging Face。

    推荐理由:Holo4 同时给出两种尺寸、跨 GUI 与 MCP 的统一接口和公开轨迹,读者可据此比较开源智能体与闭源前沿的成本差距。

9月27日周日
9月25日周五
  1. GitHub Blog60

    GitHub Security Lab 发布 Fuzzing Taskflow 智能体,自动为 C/C++ 项目做模糊测试

    GitHub Security Lab 发布 Fuzzing Taskflow,一个面向 C/C++ 项目的自主模糊测试流水线,只需指向一个 GitHub 仓库,它就会识别入口点、分析构建系统、编写 harness、运行 AFL++、读取覆盖率报告并分诊崩溃。

    推荐理由:GitHub Security Lab 把模糊测试的 harness 编写、覆盖率追踪与崩溃分诊交给 LLM 智能体,读者可了解其分层设计与安全边界。

9月24日周四
  1. SiliconFlow42

    并非每次模型调用都需要一个答案。有时,它只需要一个决策。 👏 欢迎 Kev-4B 加入 SiliconFlow。 Kev-4B 是 Jev 的开源社区版,基于 Qwen3.5-4B 构建,用于结构化决策。 路由。排序。批准。升级——无需再生成一段回复。 无需部署或适配。一个 SiliconFlow API key,Kev 即可接入你的工作流。 特别感谢 @jaredpalmer 开源 Kev。❤️ 在 SiliconFlow 上试用 Kev-4B。⚡️

9月23日周三
  1. StepFun55

    阶跃星辰(StepFun)宣布开源内部使用的 LLM 数据标注与模型检查工具 onPanda,工作流为找到错误、修正 token、让模型继续生成。数据标注方面,标注时间中位数比人工后编辑降低 52%,SFT 与偏好数据可在同一流程完成(ΔPPL <1%),支持 token 级正负样本监督及图像、音频、视频上的 agent 轨迹标注。

    引用Lei Yang@diyerxx

    I spent two years building this interactive tool to let you steer LLMs and agents at the token level. Introducing onPanda — a web app for token visualization & control, model inspection, data annotation, and more. Try it online (works on mobile): https://onpanda.diyer22.com/

9月22日周二
  1. Xiaomi MiMo56

    小米 MiMo 发布 MiMo-V2.6-Pro,据 Arena 数据以 1628 分(AutoEval)位列 Code Arena: WebDev 总榜约第 10 名、MIT 许可开源权重模型中约第 3 名,较上一代 MiMo-V2.5-Pro 的 1475 分提升 153 分。该成绩为早期 AutoEval 分数,由基于人类偏好数据训练的 Reward Model 自动投票产生,随真实人类投票增加可能变化。

    引用Arena.ai@arena

    MiMo-V2.6-Pro just landed @XiaomiMiMo back in the top 10 on Code Arena: WebDev, debuting at ~#10 overall, and ~#3 among open-weights models with an MIT license. It scores 1628 pts (AutoEval), tying Claude Fable 5 (High) and just ahead of Hy4-preview (1624 pts). That's a +153 pt jump from the previous MiMo-V2.5-Pro (1475 → 1628 pts). Among open-weights models, it lands at ~#3. Impressively only 7 pts behind Qwen3.8 Flash Next (#2) and 46 pts behind Kimi K3 Max (#1). Note: this is an early AutoEval score, in which a Reward Model trained on Arena’s human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. Congrats to the @XiaomiMiMo team on this release!

9月21日周一
9月19日周六
9月18日周五
9月16日周三
9月14日周一
9月13日周日
9月12日周六