arXiv:cs.AI· Sumanyu Muku·· 6 小时前AI 评分49
并行编码智能体协调验证:NP-Bench 与调度规划器
Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner
AI 导读
研究者提出将并行编码智能体的协调问题重构为调度问题,通过预先划分工作范围并按生产者→消费者图排序合并,内置到 Nerveplane 中。配套 NP-Bench 三臂基准测试显示,规划器将干净集成场景从 1/9 提升到 9/9,合并冲突从 13 降至 0,且在破坏性契约变更中把干净集成率从 0 提升至前沿模型的 1.0 和小模型的 0.6,跨会话记忆则将重复错误率从 1.00 降至 0.00。
正文
Abstract:A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel on one codebase, the agents collide: two rewrite the same function, one codes against a contract a teammate just changed, and integration fails after the work is done. Most coordination tools react (watch for a conflict, then warn), but at agent speed the warning arrives after the wasted edit. We recast the problem as scheduling: take each work item's declared scope, partition the work into disjoint scopes, and order merges along the producer->consumer graph, all up front. We build this planner into Nerveplane and evaluate it with NP-Bench, an environment-grounded three-arm benchmark (no coordination; reactive detection; proactive planning) that verifies integration off a real git merge, both in a deterministic simulation and with live agents. The planner lifts clean-integration from 1/9 to 9/9 scenarios and cuts merge conflicts from 13 to 0, with a gap that grows in the number of agents. On a live breaking contract change it rescues an outcome both baselines miss on every seed: the clean-integration rate rises from 0 (no coordination and reactive detection) to 1.0 on a frontier model and 0.6 on a small one, while agents respect assigned scopes (0/5 leakage). A cross-session memory drops the repeated-mistake rate from 1.00 to 0.00 on strong and weak models alike. We also report a negative result: routing facts to agents does not rescue long-context accuracy at window-fitting scales; its value is cost and capacity, not attention. Across two capability tiers and two vendors, the benefit did not shrink as models got stronger, because it comes from how work is allocated, not model reasoning.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07261 [cs.AI] |
| (or arXiv:2610.07261v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07261 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sumanyu Muku [view email]
[v1]
Mon, 5 Oct 2026 19:02:27 UTC (25 KB)
来源:arXiv:cs.AI · arxiv.org