arXiv:cs.AI· Hans Schabert, Christoph Peters·· 6 小时前AI 评分52
arXiv 论文:通过 MCP 逐步下发流程步骤,用 LLM 智能体自主性换取流程可预测性
One Step at a Time: Trading LLM Autonomy for Process Predictability
AI 导读
Hans Schabert 与 Christoph Peters 提出用 MCP 服务器逐条下发 SOP 步骤、由智能体执行并返回结构化 step_output,使执行路径在运行前可预测并形成可审计的机器可读执行日志。
正文
Abstract:Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.
| Comments: | 14 pages, 12 tables |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7; I.2.8; H.4.1 |
| Cite as: | arXiv:2610.07817 [cs.CL] |
| (or arXiv:2610.07817v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07817 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hans Schabert [view email]
[v1]
Tue, 6 Oct 2026 06:10:42 UTC (42 KB)
来源:arXiv:cs.AI · arxiv.org