跳到正文

全部动态

今日 49 条
9月21日周一
  1. MiniMax Design (H3)36

    🔥社区从不停下折腾的脚步。 不只是基于 H3 做开发,还在不断深入内部,寻找让它更聪明的新方法。

    引用Kamimoto(かみもと)@sep_is_heim

    流行のJevをMiniMax H3に組み込んで、動画生成を高速化してみた!Attention処理のスパース化にJevを使用。 ・層ごとにJevが重要度を判定(4step 49層が対象) ・Jevがスパース率1%, 3%, 5%, 10%を選択 RTX4070で6分7秒→3分34秒で41.7%短縮!動画生成中にJevクラウドに問合せしているのに速い!

  2. 小米 MiMo:GitHub 新仓库(模型发布)56

    小米 MiMo 开源 mimoagent:百行代码智能体在 SWE-bench Verified 得分超 74%

    小米 MiMo 在 GitHub 开源 mimoagent,一个仅约 100 行代码的 AI 智能体,可解决 GitHub issue 或在命令行中辅助用户。项目主打极简设计,无需庞大配置和大型 monorepo,并在 SWE-bench Verified 上取得超过 74% 的分数。仓库地址:https://github.com/XiaomiMiMo/mimoagent

  3. Peter McCrory24

    大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正观点) (3) 承认不确定性;做出可证伪的预测 (4) 认真且谦逊

    引用Alex Imas@alexolegimas

    A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."

9月20日周日
9月19日周六
  1. Gary Marcus:The Road to AI We Can Trust(RSS)33

    Gary Marcus 批评 Dario Amodei 七天内三度失信

    Gary Marcus 发文列举 Dario Amodei 在七天内损害自身公信力的三种做法:其一是让与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 充当独立监督方;其二是 Anthropic 正筹备自建湿实验室,却缺乏常规机构审查委员会监督;其三是嘴上呼吁"pace the frontier",实际仍指向 IPO。

  2. MIT News(RSS)27

    MIT Reads 十周年转型:转向虚构与回忆录,应对 AI 时代

    MIT Libraries 的 MIT Reads 项目在十周年之际转型,将重点转向虚构与回忆录,以在 AI 时代促进社交连接与共同人性。MIT 校长 Sally Kornbluth 选定 Ted Chiang 的《Exhalation》为 2026 年秋季书目,该书探讨人类与 AI 的关系等议题。MIT 教学与科研中 AI 使用特设委员会报告引用该项目,称其有助于推动校园关于共同规范的对话。

  3. Gary Marcus:The Road to AI We Can Trust(RSS)24

    Gary Marcus:近期真正该担心的不是失控超级智能,而是失控的智能体 AI 大规模攻击互联网

    Gary Marcus 认为,近期真正值得担忧的不是失控的超级智能,而是失控的智能体 AI 大规模发动互联网攻击。他援引《华尔街日报》评论版 Brian Gross 的文章称,主流媒体中少有机构梳理这一整体图景,并表示完全认同该文观点。

9月18日周五
  1. GitHub Blog22

    GitHub Podcast 拆解 AI 热门观点:该不该读代码、RAG 是否已死、Skills 是否杀死了 MCP

    GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,MCP 提供工具与数据的标准接入,Skills 封装团队流程与最佳实践,二者可组合使用;RAG 并未消亡,检索能为模型提供训练数据之外的信息,减少 token 消耗并让回答更有依据。

  2. MiniMax Design (H3)54

    MiniMax Design 转发用户演示:用 MiniMax H3 Max 的 r2v 功能将一张 3x3 分镜图一次性生成为视频,设置 480p、15 秒,提示词为「左上から右下のパネルにカットが切り替わる2Dアニメーション、複数パネル禁止、BGM禁止」(按左上到右下面板切换镜头的 2D 动画、禁止多面板、禁止 BGM),Quality 调整、参照强度标准,作者称生成画面与分镜设定基本一致。

    引用852話(hakoniwa)@8co28

    この3x3画像を1枚 r2v でAI動画化すると以下の設定でおおよそそのまま映像になる Minimax H3 Max r2v 480p 15秒 Prompt調整:Quality 参照強度:標準 「左上から右下のパネルにカットが切り替わる2Dアニメーション、 複数パネル禁止、BGM禁止」

  3. MiniMax (official)40

    Nunchux AI 与多校研究者推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上比 FlashAttention-4 快 1.6×、B300 上快 1.5×,保真度优于 SageAttention2。

    引用Nunchux AI@NunchuxAI

    Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.

  4. Karina56

    Epoch AI 推出 Benchmark Reviews 计划,对 AI 基准进行审计,首批覆盖 15 个基准,其中 4 个 Verified、9 个 Flawed、2 个信息不足。Karina Nguyen 转发并称赞这一举措能激励行业构建真正高质量的基准,并感谢其审计了 PostTrainBench、挖掘出旧的 SimpleQA。

    引用Epoch AI@EpochAIResearch

    Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

  5. Claude Code:GitHub Releases(RSS)39

    Claude Code v2.1.275 发布

    Claude Code 发布 v2.1.275,新增登录账号显示、ctrl+enter 立即发送排队消息,以及将 claude.ai 账号启用的技能和插件同步到终端会话。该版本还修复了恢复会话时提示缓存失效、全屏模式滚动卡顿、插件市场更新误删本地副本等问题,并改进 --system-prompt 中 __SYSTEM_PROMPT_DYNAMIC_BOUNDARY__ 行的提示缓存。