跳到正文
arXiv:cs.LG· Abolfazl Younesi·· 6 小时前AI 评分38

STEPGATE:面向小语言模型智能体的不确定性感知步骤级云端接管

Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents

AI 导读

STEPGATE 是一种不确定性感知的步骤级接管框架,对本地 SLM 每一步动作打分并选择性升级到更强模型。在 52 项 BFCL 单步测试中,Qwen2.5-1.5B/7B 组合以 30.8% 升级率取得 82.7% 任务成功率,高于纯本地的 67.3% 和随机升级的 75.4%。多轮评测中仅用 30.0% 云端动作即达 69.0% 轨迹成功率,但研究仅限单一模型族与脚本化任务。

正文

View PDF HTML (experimental)

Abstract:Small language models (SLMs) are attractive as local agent controllers because they reduce remote inference, latency, and deployment footprint, yet structured tool errors can cause an agent step to fail. Existing routers typically select a model once per query. However, agents expose sequential decision points whose difficulty dynamically changes based on intermediate observations. We propose STEPGATE, an uncertainty-aware handoff framework that scores each local SLM action and selectively escalates challenging steps to a stronger model. On a 52-task held-out single-step BFCL-derived test split, the Qwen2.5-1.5B/7B pair attains 82.7% task success with 30.8% escalation, versus 67.3% local-only and 75.4% random escalation (which uses 33.8% escalation). In a separate multi-turn evaluation, STEPGATE achieves 69.0% trajectory success and 84.0% action success using only 30.0% cloud actions, compared with 48.0%/70.5% local-only, 60.0%/78.2% random escalation, and 57.0%/77.1% query-level routing (strong-only achieves 82.0% trajectory success at 100% cloud actions). These results suggest that step-level escalation recovers a large share of the performance gap to the stronger Qwen2.5-7B backend at a matched cloud-action rate while transmitting fewer tokens remotely. However, our evaluation is limited to one model family, a single stronger backend, and scripted tasks. Furthermore, the test sets are small, multi-turn comparisons rely on paired intervals and statistical tests, and our risk tiers serve as research annotations rather than formal safety guarantees.
Comments: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: SLMs for Agentic Systems, Paris, France, 2026
Subjects: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Emerging Technologies (cs.ET); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
ACM classes: I.2; I.2.11
Cite as: arXiv:2610.07816 [cs.AI]
  (or arXiv:2610.07816v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07816

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Abolfazl Younesi [view email]
[v1] Tue, 6 Oct 2026 06:09:31 UTC (94 KB)

来源:arXiv:cs.LG · arxiv.org