AI News Feed
Market watch
Companies

Google Research moves federated learning into trusted execution environments; Gboard trains with externally verifiable differential privacy

Google Research says TEE-based federated learning offers externally verifiable central DP and now trains Gboard models.

The report says earlier federated learning systems had a trust gap. Google introduced federated learning in 2017, and it powers next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages and Smart Text Selection in Android. Devices uploaded data for immediate aggregation, but outsiders could not verify that data was never logged or inspected. Secure Aggregation added cryptographic protection, but it was not compatible with state-of-the-art central differential privacy algorithms such as matrix factorization DP-FTRL. Google also had to be trusted to add differential privacy noise correctly.

The new design moves client gradient computation to the server and makes that server logic attestable, so the operator no longer needs to be trusted. It builds on Google's earlier confidential federated analytics work and coordinates four core components: data upload, key management and policy verification, workload execution, and fault-tolerant recovery. Devices encrypt training examples locally and pre-authorize an access policy that lists which TEE computations may process the data. Policies must appear in a public transparency log. A Key Management System built from TEEs running the RAFT consensus protocol releases keys only to workloads matching the policy. A root TEE runs a Python training loop and delegates subtasks to worker TEEs. Orchestration uses Federated Language, derived from TensorFlow Federated. Only differentially private model weights are released. Each round saves a KMS-encrypted recovery state to handle root or worker failures.

The privacy guarantee is verifiable because access policies are published to Rekor, Sigstore's public transparency log, according to the report. External auditors can track every server workload that a device could feed. The KMS and data processing binaries are reproducibly buildable from open source code, and the policies directly describe the Python training program. To protect proprietary model architectures, TEEs support sideloading serialized logic at runtime, but all privacy-relevant logic must stay hardcoded in the attested program. Workload operators see only metrics and differentially private model weights. Encrypted data can be decrypted only for a limited time after upload.

Gboard used the system to launch English and Japanese next-word prediction models, the report says. Two design choices drive the result. First, all uploads are collected before server-side training runs, so diurnal swings in device availability no longer slow training; the program can compute an optimal participation schedule and tune differential privacy parameters. Google's privacy-utility curves come from training an English model for 5,000 rounds with cohorts of 6,500 devices on both systems. Second, the bottleneck moved to the server. Previous federated learning models took one to two months each to train. Training now parallelizes across machines, limited only by TEE resource availability. Google reports substantially faster compute times but does not publish a single speedup figure.

The report compares Google's system with other federated learning frameworks. It says Google's approach is for production cross-device training and computes client updates in server-side TEEs. NVIDIA FLARE is described as a production federated learning SDK with Docker, Kubernetes and cloud tooling that computes updates at each participating site and supports hardware TEEs including AMD SEV-SNP, Intel TDX and NVIDIA GPU confidential computing. Flower is a framework for building federated AI systems that computes updates on clients and supports central and local differential privacy. Apple pfl-research is simulation only, not intended for third-party deployments, and does not include hardware TEE support in its core framework.