跳到正文

#Agent

今日 187 条
9月22日周二
  1. MeshyAI23

    Meshy 🤝 chrona 让你的 3D 创意尽情释放 🤩

    引用Chrona@chronaworld

    What if Tech Week wasn’t a calendar, but a city you could actually enter and play? We made a SF TECH WEEK CITY — rebuilding San Francisco as a 3D world with Tech Week events placed at their real venues. RSVP activity, trending events, and busy areas become part of the world itself. The birds flying above the city represent real community members, so you can see where people are heading and explore alongside them. Built with AI coding and published on Chrona. A small experiment in turning real-world data into a place worth entering. @MeshyAI @ChronaWorld #MeshyxChrona #TechWeek

  2. OpenBMB23

    非常感谢分享!看到 MiniCPM5-2B 被用在这样实用的多智能体工作流中真的很酷——尤其是 workers 通过工具调用实际处理匹配、短付款、重复引用和争议。感谢你在测试和记录上付出的所有努力。对社区来说是个很好的案例 🙌

    引用Joey@aijoey

    Next test on Spark 1: 32 synthetic invoices, with purchase orders and payment records. GPT-6 Astra coordinates the MiniCPM5-2B workers. They sort out matches, short payments, duplicate references and price disputes, then write the results into a case ledger. All 32 verified in 67.8 seconds. Eight in each category, with 232 executed tool calls. The video is real time. This demo doesn’t move money.

  3. elsewhere:文章(RSS)68

    阶跃 Step 5 Preview 实测评测:数据可视化与金融分析亮眼,泛化和审美仍有短板

    阶跃发布 Step 5 Preview,总参数量 600B、激活参数 27B,有视觉输入,官方称在 Artificial Analysis 上涨 44 分、单任务成本仅为 Claude Opus 5 的 1/8。作者与友人实测发现其在数据可视化、金融分析上表现不错,但泛化性、领域知识和审美偏弱,思考过程过长导致长任务耗时且易中断,且长上下文下安全指令遵循可被绕过。

  4. Xiaomi MiMo66

    小米 MiMo 发布 MiMo-V2.6 Pro 与 Flash 两款全模态模型,通过规模化强化学习训练。官方称 Pro 在多数 agent 基准上与 Claude Opus 5 和 GPT-5.6 Sol 相当,Artificial Analysis Intelligence Index 得分 46,为开源模型中最高;能力覆盖编码、computer use、3D 推理和创作。

    推荐理由:原文给出 Pro 与 Claude Opus 5、GPT-5.6 Sol 的 agent 基准对比和开源范围,可据此评估其相对位置。

  5. Grok29

    在 Box 中试用 Grok 4.7

    引用Box@Box

    We put Grok 4.7 from @SpaceXAI to work on a $2 million insurance claim where one deductible error alone changes the calculation by $143,000. In this Box Agent preview, @grok reconciles the claim against the policy and supporting records. It catches an $82,000 duplicate invoice and a missing $64,000 supplier credit. It also explains why the Business Income waiting period doesn’t apply to Extra Expense. The result is a cited claims review for the adjuster, with final coverage and payment decisions left to the insurer. Explore Box AI Studio to build custom agents for your own document-heavy workflows.

9月21日周一
  1. 小米 MiMo:GitHub 新仓库(模型发布)56

    小米 MiMo 开源 mimoagent:百行代码智能体在 SWE-bench Verified 得分超 74%

    小米 MiMo 在 GitHub 开源 mimoagent,一个仅约 100 行代码的 AI 智能体,可解决 GitHub issue 或在命令行中辅助用户。项目主打极简设计,无需庞大配置和大型 monorepo,并在 SWE-bench Verified 上取得超过 74% 的分数。仓库地址:https://github.com/XiaomiMiMo/mimoagent

9月20日周日
9月19日周六
  1. Andrew Milich35

    用 @bot 省钱! @Baconbrix: GROK BOT 刚在我邮箱里发现了 $986 🤯 我让 Grok 帮它给我买的车找个停车位,结果它顺路发现了我公寓近一千美元的额外收费,还发邮件要求退款。

    引用Evan Bacon@Baconbrix

    GROK BOT JUST FOUND $986 IN MY EMAIL 🤯 I asked Grok to get a parking spot for the car it bought me, and on the way it discovered nearly a thousand dollars of extraneous charges from my apartment and emailed them asking for a refund.

  2. Gary Marcus:The Road to AI We Can Trust(RSS)24

    Gary Marcus:近期真正该担心的不是失控超级智能,而是失控的智能体 AI 大规模攻击互联网

    Gary Marcus 认为,近期真正值得担忧的不是失控的超级智能,而是失控的智能体 AI 大规模发动互联网攻击。他援引《华尔街日报》评论版 Brian Gross 的文章称,主流媒体中少有机构梳理这一整体图景,并表示完全认同该文观点。

9月18日周五
  1. GitHub Blog22

    GitHub Podcast 拆解 AI 热门观点:该不该读代码、RAG 是否已死、Skills 是否杀死了 MCP

    GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,MCP 提供工具与数据的标准接入,Skills 封装团队流程与最佳实践,二者可组合使用;RAG 并未消亡,检索能为模型提供训练数据之外的信息,减少 token 消耗并让回答更有依据。