arXiv:cs.LG· Haoran Zhang, Dongjun Kim, Seohyeon Cha, Kevin S Chan, Ananthram Swami, Gustavo De Veciana, Haris Vikalo·· 3 小时前AI 评分36
JOVE:面向资源感知 LLM 任务图的联合执行与验证框架
JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs
AI 导读
JOVE 是一个在线框架,将复杂推理查询分解为有向无环任务图并分配给异构 LLM,同时联合选择中间输出进行付费验证,验证异步运行并反馈用于优化后续分配。它通过求解逐查询混合整数线性规划做决策,在长期预算和单查询延迟约束下平衡执行开销与学习收益,并证明了次线性的质量学习 regret。在四个推理基准上,JOVE 精度与标准推理基线相当,平均成本与延迟至少降低 3.17 倍。
正文
Abstract:Complex reasoning queries can be decomposed into directed acyclic task graphs and distributed across heterogeneous LLMs, reducing latency through parallelism and enabling smaller models to solve complex tasks. In practice, however, the suitability of an LLM for a given subtask may be a priori unknown, and execution alone does not reveal output correctness. We propose JOVE, an online framework that jointly assigns executor LLMs and selects intermediate outputs for paid verification. Verification runs asynchronously and is used to improve future allocations, so the system must balance spending on execution now against learning for later. We study how to optimize this trade-off under a long-term budget and a per-query latency constraint, with stochastic, initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs. Online learning updates task-dependent estimates of LLM quality based on verification feedback, while an information-gain bonus incorporates the value of learning into allocation decisions. Under a natural set of assumptions, we establish sublinear quality-learning regret for JOVE. Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times.
| Comments: | preprint |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| ACM classes: | I.2.6; I.2.8; G.1.6 |
| Cite as: | arXiv:2610.03296 [cs.AI] |
| (or arXiv:2610.03296v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03296 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haoran Zhang [view email]
[v1]
Fri, 2 Oct 2026 13:37:59 UTC (388 KB)
来源:arXiv:cs.LG · arxiv.org