跳到正文
arXiv:cs.AI· Rohit Verma, Anand V Bodas·· 3 小时前

BEVPIPE:用可移植 GPU 计算部署 BEV 感知,解决算子不匹配难题

The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute

AI 导读

针对 BEV 感知模型中稀疏 3D 卷积与几何 scatter 操作无法被标准推理运行时表达、且现有稀疏卷积库仅支持 CUDA 并绑定 PyTorch 的算子不匹配问题,研究者提出部署框架 BEVPIPE。

正文

View PDF HTML (experimental)

Abstract:Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space. These models are accurate, but they cannot be deployed through standard inference runtimes. The reason is an operator mismatch between dense convolutions (which runtimes handle well), sparse 3D convolutions (which runtimes cannot represent), and geometric scatter operations (which runtimes have no vocabulary for). Today, every sparse convolution library is CUDA-only and PyTorch-coupled, locking BEV deployment to a single vendor's hardware and a single execution framework.
We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes. BEVPIPE partitions the model into runtime-managed dense subgraphs and three external operator extensions (voxelizer, sparse encoder, BEV projector), connected through a shared GPU memory space. BEVPIPE achieves a 19.5x end-to-end speedup over conventional deployments while retaining 98.5% of reference mAP. We also showcase that BEVPIPE is portable across different GPU backends.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11504 [cs.AI]
  (or arXiv:2610.11504v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11504

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Rohit Verma [view email]
[v1] Thu, 8 Oct 2026 08:42:45 UTC (57 KB)

来源:arXiv:cs.AI · arxiv.org