Google 发布基于 TEE 的下一代联邦学习系统
October 2, 2026 Toward provably private learning from federated data Mobile Systems · Security, Privacy and Abuse Prevention · Software Systems & Engineering
Google 宣布下一代联邦学习系统,利用可信执行环境(TEE)提供可远程验证和审计的数据匿名化保障,加密训练数据只能在符合访问策略的 TEE 内解密处理。访问策略发布到公开透明日志 Rekor,KMS 和数据处理二进制可从 Confidential Federated Compute GitHub 仓库复现构建。
原文给出新系统的隐私机制设计和 Gboard 实际部署后的训练提速,读者可对照旧方案理解差异。
In 2017, Google introduced Federated Learning (FL) a machine learning technique that trains models across decentralized, private data. It has been used to power everyday helpful features, including next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages, and Smart Text Selection in Android.
Our FL systems development is guided by four essential privacy principles: (1) data minimization, (2) data anonymization, (3) transparency and control, and (4) verifiability and auditability. Years of research development on anonymization have led to strong differential privacy (DP) guarantees for production models through algorithms like matrix factorization DP-FTRL (MF-DP-FTRL) and distributed differential privacy coupled with Secure Aggregation. In 2025, we introduced an evolved definition of FL centered on these four principles:
Federated learning (FL) is a machine learning setting where multiple entities (clients) collaborate in solving a machine learning problem, under the coordination of a service provider. A complete FL system should enable clients to maintain full control over their data, the set of workloads allowed to access their data, and the anonymization properties of those workloads. FL systems should provide appropriate transparency and control to the users whose data is managed by FL clients.
In “Toward provably private learning from federated data”, we announce the next generation of our FL system, which leverages Trusted Execution Environments (TEEs) to provide fully verifiable and auditable data anonymization guarantees. Logic that runs in TEEs is remotely attestable (third parties can verify the logic that is being executed), and it also gains confidentiality (its internal state cannot be observed) and integrity (the logic cannot be disrupted), subject to current-generation TEE limitations. Our new TEE-based FL system builds on these properties which TEEs offer at the level of a single machine to form a fully verifiable end-to-end FL system that achieves stronger privacy guarantees and improved accuracy. Gboard has already adopted the new system and is benefiting from substantially faster compute times than our previous FL system.
Building a verifiably private FL system out of individual TEEs
Our TEE-based FL system builds on techniques developed in our earlier work on confidential federated analytics and provably private insights.
The system coordinates four core operational concepts:
- Data upload: Client devices locally encrypt training examples and upload them. The devices pre-authorize an access policy, which is the set of TEE computations that will be allowed to process this data, and these computations only release anonymized results. Clients require access policies to be published to a public transparency log.
- KMS and policy verification: The Key Management System (KMS), which consists of a cluster of TEEs implementing the RAFT consensus protocol, only gives decryption keys to server-side TEE workloads that match the computations in the access policy.
- Workload execution: A data processing TEE executes a Python program that implements a training loop. This "root" TEE delegates parallelizable subtasks to a cluster of worker TEEs. Distributed logic is expressed using Federated Language, an open-source, framework-agnostic orchestration language derived from TensorFlow Federated, which powered our earlier FL system. The training loop periodically releases anonymized model weights to the data analyst.
- Fault-tolerant recovery: The Python program saves a KMS-encrypted recovery state at the end of the training round. This can be used to recover from intermittent root or worker failures, without leaking any additional privacy-sensitive information.
For more details on the TEE-based FL system design please see our whitepaper, Toward provably private learning from federated data.
How TEE-based FL strengthens privacy
In earlier FL systems, device data was uploaded for the purpose of immediate aggregation, but there was no way for external observers to verify that the data was never logged or inspected. Later, Secure Aggregation allowed uploads to be protected cryptographically, but was not compatible with state-of-the-art central DP guarantees. Our new TEE-based system represents the next milestone in our ongoing effort to completely remove the need to trust the server operator.
In our new TEE-based FL system, only metrics and differentially private model weights are visible to workload operators. Encrypted training data collected from devices can only be decrypted and processed within TEEs running Python training programs represented in the access policies, and only for a limited amount of time after upload.
Public transparency log and reproducible builds
Devices participating in our new TEE-based FL system know the full set of server workloads that may access data they’re uploading. The access policies representing these potential future server workloads are published to Rekor, a public transparency log, and external auditors are able to observe these logs to track the full set of server-side workloads that devices could potentially be participating in.
The KMS and data processing binaries used in our FL system can be reproducibly built from open source code published in the Confidential Federated Compute Github repository.
Verifiable execution with dynamic sideloading
In our earlier FL systems, the logic running on the server could be verified neither by devices nor auditors, and thus we needed to be trusted to correctly add random noise to gradient sums to provide differential privacy.
In our new TEE-based FL system, the access policies that are published to Rekor directly describe the Python program that expresses the FL training logic. To protect proprietary model architectures and data preprocessing logic while preserving auditability, our data processing TEEs support sideloading serialized information into the Python program at runtime. As long as all privacy-relevant logic remains hardcoded in the Python program, this sideloading functionality allows logic that must remain proprietary to run in the TEE while still providing strong externally verifiable privacy guarantees (see below for a discussion on side-channel observations).
How Gboard is using TEE-based FL
Gboard has deployed this TEE-based FL system to launch English and Japanese next word prediction models with stronger privacy guarantees and improved accuracy. These improvements can be attributed to several aspects of the new system design.
Mitigating diurnal availability constraints
By collecting all device uploads before running the server-side training workload, we no longer have to worry about diurnal variations in device availability impacting training progress. At server-side execution time, we can dynamically calculate the optimal device participation schedule within the program, and can use it to tune other DP parameters.
Training speedups
In the past, training these FL models could take 1-2 months each, with progress limited by device availability, on-device compute, and competition across multiple training workloads for the same set of device resources. With the new TEE-based system, bottlenecks have been moved to the server, and computation parallelization across many machines allows us to achieve significant speedups in training time, currently only limited by TEE resource availability.
What’s next?
In our new TEE-based FL system, computation of client gradients is shifted to the server, lifting limitations related to on-device compute resources that were present in earlier systems. This paves the way for training increasingly larger models using FL techniques. Integrating TEEs with accelerators will play an important role in such use cases.
The system we have described above is capable of executing not just FL training workloads in a verifiable manner, but also arbitrary workloads that can be expressed using Python. We are experimenting with running other types of workloads on this infrastructure, such as synthetic data generation workloads. Another area of exploration: using the data processing TEEs that execute arbitrary Python in combination with our other data processing TEEs that specialize in functionality such as LLM inference.
This work is a step toward rigorous proof that server side processing preserves individual privacy. With external verifiers able to inspect exactly what code we run at Google, we are able to offer strong assurances that data is processed server-side exactly as described. We expect future TEE hardware, along with ongoing research into mitigating side-channel observations, to offer deeper protections for dynamically loaded workloads against malicious server-side attacks. We anticipate that systems like ours may one day come with full proofs of correctness of the software implementations of the DP algorithms and system components.
Acknowledgements
The authors would like to thank the collaborators who contributed to the infrastructure design and implementation: Arun Ganesh, Brendan McMahan, Brett McLarnon, Chunxiang (Jake) Zheng, Emily Glanz, Maya Spivak, Michael Reneer, Nova Fallen, Stefan Dierauf, Suxin Guo, Timon Van Overveldt, Yu Xiao, Zachary Charles, and Zachary Garrett. We also thank close partners who supported the Gboard integration: Haicheng Sun, Heng Su, Jianpeng Hou, Liyang Jiang, Noriyuki Takahashi, Wenzhi Mao, Xiaojuan Fang, Yanxiang Zhang, Yingjie Liu, Yuanbo Zhang, and Yun Wang. This work was supported by Corinna Cortes, Shumin Zhai, and Yossi Matias. We additionally thank Maysam Moussalem for feedback on the writing of this post.
来源:Google Research:Blog(网页) · research.google