跳到正文
原文
NVIDIA Technical Blog(开发者技术博客 · RSS)· Tanya Lenz·· 9 天前AI 评分36

NVIDIA Dynamo-Triton 集成 TensorRT 多设备推理,支持跨多 GPU 模型服务

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

AI 导读

NVIDIA 在 Dynamo-Triton 中集成 TensorRT 多设备推理,让单个 TensorRT 网络借助 NCCL 分布式集合通信跨多 GPU 执行,同时保留 TensorRT 推理优化。该能力自 TensorRT 11.0 起获得完整支持,用于应对生成式 AI 超出单 GPU 的算力与显存需求。

来源:NVIDIA Technical Blog(开发者技术博客 · RSS) · developer.nvidia.com