跳到正文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分62
AI 导读

Stanford 论文提出 DeLM 框架,用共享上下文和任务队列取代中央协调智能体,智能体异步认领任务、实时共享发现。

正文

A new Stanford paper just proved that decentralized multi-agent systems can beat Claude Code and Codex on complex coding tasks, running up to 2.49× faster!

DeLM tackles a key bottleneck in multi-agent systems: time wasted repeating work and waiting on other agents.

DeLM replaces the central coordinating agent with a shared context and task queue. Agents pick up tasks independently, share discoveries as they happen, and build on each other’s progress. When one agent finds a solution or hits a dead end, the others can use that information immediately.

DeLM beats state-of-the-art in both speed and accuracy on long-horizon tasks from Terminal-Bench 4.0, DeepSWE v1.1, and ProgramBench:

  • Up to 2.49× faster execution than the vanilla Claude Code and Codex baselines
  • Up to +19.2 percentage points in accuracy over the vanilla Claude Code and Codex baselines
  • Up to +19.9 points in ProgramBench test pass rate within the same 120-minute budget

Built on Codex and Claude Code. Code and 720 trajectories are available so you can explore how the agents collaborate! An open-source plugin lets you try DeLM directly in Codex and Claude Code.

来源:Rohan Paul · x.com