A Confidential GPU is a GPU whose memory, computation state, and data transfers are hardware-encrypted and isolated from the host system — extending the Trusted Execution Environment (TEE) boundary that technologies like TDX and SEV-SNP provide at the CPU level to encompass the GPU accelerator as well. Without that extension, a GPU sits outside the CPU TEE: it cannot read TEE memory, and any data offloaded to it would leave the confidential boundary. The primary implementation today is NVIDIA Confidential Computing, introduced on Hopper (H100), continued on Blackwell, and planned for Rubin as the third generation. In Confidential Computing mode, the GPU encrypts data resident in High Bandwidth Memory (HBM) under keys managed by its on-die security processor, and three hardware properties close the trust gap to the CPU TEE: an on-die Root of Trust that verifies GPU firmware authenticity before the OS can talk to the device; device attestation via the NVIDIA Remote Attestation Service (NRAS), which produces signed evidence that the GPU is genuine NVIDIA hardware in Confidential Computing mode with unmodified firmware (analogous to Intel DCAP or AMD KDS for CPU TEEs); and encrypted PCIe transfers between the CPU TEE and GPU TEE at line rate using a hardware AES-256-GCM implementation, so a host administrator, hypervisor, or co-tenant with DMA access to the bus sees only ciphertext. GPU attestation is verified alongside CPU attestation before secrets (model decryption keys, dataset credentials) are released to the combined CPU+GPU TEE. The technology requires no application code changes — existing TensorFlow, PyTorch, and CUDA workloads run unmodified inside the confidential boundary. The primary threat model is the same as CPU-level confidential computing (protecting data-in-use from the infrastructure operator) but applied to AI workloads: model intellectual property theft, training data exfiltration, and inference input/output interception during GPU computation.
Red Hat integrated Confidential GPU support in OpenShift sandboxed containers 1.12 (April 2026) as a Technology Preview, in collaboration with NVIDIA. The implementation extends the existing CoCo (Confidential Containers) architecture: a confidential container running inside a CPU TEE (TDX or SEV-SNP) is connected to an NVIDIA GPU operating in Confidential Computing mode, creating a unified TEE spanning both CPU memory and GPU memory. Red Hat build of Trustee 1.1 adds NRAS integration for composite attestation: CPU TEE evidence is verified by Trustee’s built-in TDX or SEV-SNP verifier, GPU TEE evidence is delegated to NRAS, and secrets are released only when both produce affirming trustworthiness claims (EAR / AR4SI). On OpenShift, Node Feature Discovery labels TEE- and CC-capable nodes; the NVIDIA GPU Operator enables CC mode (nvidia-cc-manager), binds devices for VFIO passthrough, and advertises GPUs via a Kata sandbox device plugin; the sandboxed containers operator creates the RuntimeClass kata-cc-nvidia-gpu (with kata-nvidia-gpu for non-confidential GPU passthrough) and supplies an initramfs that embeds the NVIDIA driver and the Trustee agent for secret injection (image signature verification, encrypted image decryption, network credentials). The key use case is model IP protection for distributed inference: a proprietary model vendor encrypts weights in a standard registry; an untrusted third-party OpenShift operator pulls them; Trustee releases the decryption key only after both CPU and GPU TEEs attest via NRAS — so the model never appears in plaintext outside verified hardware. Current operational limits matter for planning: a cluster’s worker nodes must be a single CPU TEE type (TDX or SEV-SNP, not mixed); OpenShift 4.21.9+ is required; tested confidential SKUs include H100 and RTX PRO 6000 Blackwell; only one confidential GPU can be cold-plugged into a confidential pod (regular non-CC pods can attach multiple GPUs); and all NVIDIA GPUs on a worker are statically bound to one runtime class until reconfigured. Roadmap items called out by Red Hat include dm-verity for the Kata VM image, broader GPU SKU coverage, and air-gapped NVIDIA attestation.
Additional Information#
- NVIDIA Confidential Computing
- AI meets security: POC to run workloads in confidential containers using NVIDIA accelerated computing (Nov 12, 2024)
- Secure AI inferencing: POC with NVIDIA NIM on CoCo with OpenShift AI (Mar 18, 2025)
- Red Hat OpenShift sandboxed containers 1.12 and Red Hat build of Trustee 1.1 bring confidential computing to bare metal and AI workloads (Apr 13, 2026)
- Protect data offloaded to GPU-accelerated environments with OpenShift sandboxed containers (May 22, 2026)
- Configuring confidential containers for NVIDIA GPUs (OpenShift sandboxed containers docs)
- Confidential Containers with NVIDIA Confidential GPU - Interactive Demo