arXiv:cs.LG· Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager·· 7 小时前AI 评分36
DySCo:面向边缘-云协同 LLM 推理的动态分片与深度同步批处理
DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized Batching
AI 导读
研究者提出协同运行时 DySCo,通过 dyForward 支持在本地模型分片上执行可配置的连续层区间而无需重载权重,并用 depth-synchronized batching(DSB)将异构请求推进到最深切分点、批量复用公共后缀计算。
正文
Abstract:Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: this https URL
| Comments: | article under submission |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY) |
| Cite as: | arXiv:2610.08268 [cs.DC] |
| (or arXiv:2610.08268v1 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08268 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shashikant Ilager Mr [view email]
[v1]
Tue, 6 Oct 2026 12:42:11 UTC (145 KB)
来源:arXiv:cs.LG · arxiv.org