跳到正文
arXiv:cs.LG· Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar·· 4 小时前AI 评分41

SpliTEE:用 GPU 辅助 TEE 结合差分隐私实现快速私密 LLM 推理

SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy

AI 导读

SpliTEE 将 LLM 推理拆分到 Intel TDX TEE 与不可信 GPU 之间,并用差分隐私替代加密保护中间表示,使推理速度接近纯 TDX 方案的两倍。该工作对 LLM 推理关键函数做了全局敏感度分析以界定 DP 噪声规模,并推导出噪声消除的浮点误差上界。基于 Llama-3.2-3B 和 Qwen3-4B 的实验中,SpliTEE 比 Slalom 最多快 43% 且准确率更高。

正文

View PDF HTML (experimental)

Abstract:User prompts provided to large language models (LLMs) may contain private information. One way to protect them is to execute the LLM inside a trusted execution environment (TEE). However, this results in slow inference times as current TEEs are significantly slower than GPUs for LLM inference. To circumvent this, Tramèr and Boneh (2019) proposed Slalom which splits neural network inference between a TEE and an untrusted GPU. They encrypt inputs to computations outsourced to the GPU. In this paper, we extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy (DP). We first demonstrate that masking intermediate representations is necessary by showing an 80% accuracy on a prompt-reconstruction attack from these representations. Our main contribution is a global sensitivity analysis of key functions in LLM inference, which bounds the required scale of DP noise. Unlike encryption, DP avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on the floating-point error from masking and subsequent noise cancellation as a function of the privacy parameter epsilon, keeping the same quality of the LLM response. We implement our architecture using the Intel TDX TEE and two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully TDX-based inference. Moreover, it is at most 43% faster than Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, even with knowledge of the DP mechanism, cannot recover more information than is contained in an unrelated prompt.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2609.15039 [cs.CR]
  (or arXiv:2609.15039v3 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2609.15039

arXiv-issued DOI via DataCite

Submission history

From: Robin Carpentier [view email]
[v1] Mon, 14 Sep 2026 04:55:17 UTC (325 KB)
[v2] Sun, 27 Sep 2026 03:23:20 UTC (325 KB)
[v3] Tue, 6 Oct 2026 10:45:55 UTC (344 KB)

来源:arXiv:cs.LG · arxiv.org