arXiv:cs.AI· Keyvan Dadashzadeh, Yuehong Zhou, Minyu Cui, Miquel Pericas·· 6 小时前AI 评分44
T-CCL:基于 Tensor Memory Accelerator 的资源高效集合通信库
T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator
AI 导读
T-CCL 是一个基于 Tensor Memory Accelerator(TMA)的节点内集合通信库,将数据搬运与归约操作卸载至 TMA,并以异步流水线方式执行集合通信。
正文
Abstract:Large transformer-based models increasingly depend on multi-GPU execution, which requires frequent collective communication among GPUs. Existing communication libraries often rely on many GPU threads to achieve high bandwidth or low latency, resulting in a large streaming multiprocessor (SM)-side resource footprint. This footprint can limit the resources available to other GPU work, particularly when communication and computation execute concurrently. Thus, efficient collective communication should not only achieve high collective performance but also reduce its SM-side resource usage. This paper presents T-CCL, a resource-efficient collective communication library based on the Tensor Memory Accelerator (TMA) for intra-node communication. T-CCL offloads both data movement and reduction operations to TMA and executes each collective as a pipelined series of asynchronous TMA operations, reducing the SM resources required for collective communication while maintaining high bandwidth. Evaluated across AllReduce, AllGather, and ReduceScatter collectives, T-CCL outperforms NCCL by up to 2.4x with unrestricted communication resources and up to 3.42x under restricted resource budgets, remains competitive with NCCL's recent symmetric-memory kernels, and occupies the same or fewer SMs in profiled cases. In a GEMM-collective overlap case study, switching the communication backend from NCCL to T-CCL raises the average operator-level speedup over a sequential baseline from 1.12x to 1.25x on two GPUs and from 1.04x to 1.14x on four GPUs, as T-CCL uses fewer SMs for communication, leaving more SMs available to the overlapped GEMM. Integrated into vLLM as a communication backend, T-CCL improves end-to-end inference throughput over vLLM's automatic backend dispatch by up to 1.31x, outperforming it at every evaluated batch size on both the conversation and decode-heavy workloads.
| Comments: | Workshops on Supercomputing (SC'26) |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07098 [cs.DC] |
| (or arXiv:2610.07098v1 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07098 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Minyu Cui [view email]
[v1]
Mon, 5 Oct 2026 14:19:11 UTC (1,030 KB)
来源:arXiv:cs.AI · arxiv.org