The deployment of proprietary frontier models on shared cloud infrastructure has created a massive trust dilemma. We have successfully secured Data-at-Rest with AES-256 and Data-in-Transit with TLS 1.3. However, the moment an AI model begins processing prompts, generating embeddings, or performing inference, the data is decrypted in the server's RAM and GPU VRAM.
This creates the ultimate vulnerability: Data-in-Use. If a compromised hypervisor, a rogue cloud infrastructure administrator, or a hacker performs a memory-dump attack on the server's RAM, your multi-million-dollar model weights and your users' plaintext PII are exposed.
The only engineering solution to this is Confidential Computing (CC) on Bare Metal. By establishing a Zero Trust AI Infrastructure utilizing hardware-level Trusted Execution Environments (TEEs), you can mathematically lock out the host operating system. In this SRE playbook, we will dissect the architecture of Confidential VMs (CVMs), expose the PCIe interception threats, and reveal why public cloud confidential instances are a FinOps trap.
Phase 1: The CPU Hardware Root of Trust & Virtualization
Traditional virtualization (KVM) relies on a hypervisor to isolate tenants. But in a Zero Trust model, the hypervisor itself is treated as hostile. To fix this, modern CPUs create hardware-isolated boundaries.
- KVM vs CVM: KVM is like renting a room in a building where the building owner (Host OS) holds a master key. A CVM (Confidential Virtual Machine) is the "armor-plated" version. Once you lock the room, the master key no longer works. Not even the host can peek inside.
- Intel TDX (Trust Domain Extensions): Available only on 4th Gen (Sapphire Rapids) and newer Xeon processors. It creates a "Trust Domain" (TD)—think of it as a "Private VIP Room" inside the CPU. The encryption key is generated by the silicon itself and never leaves the CPU.
- AMD SEV-SNP (Secure Nested Paging): Available on modern EPYC processors. Earlier versions of SEV encrypted memory but failed to prevent the hypervisor from remapping pages (a thief breaking the wall instead of the door). SEV-SNP adds strict "Integrity" protection (Secure Nested Paging), immediately blocking any malicious memory remapping attempts by the hypervisor.
Phase 2: Securing the GPU (The Secret SPDM Handshake)
Securing the CPU RAM is useless if the data is sent to the GPU in plaintext. While physical hardware sniffers are a threat in bare-metal colocation setups, the real enterprise threat in cloud environments is a Rogue PCIe DMA (Direct Memory Access) Attack, where a compromised hypervisor or malicious neighboring VM attempts to scrape the PCIe bus memory space.

Architecture Flow: Data remains encrypted from the CPU Trust Domain across the PCIe bus into the GPU's isolated VRAM.
When Intel TDX / AMD SEV-SNP pairs with NVIDIA CC, magic happens:
- The Secret Handshake (SPDM): On boot, the CPU and GPU authenticate each other. "Are you a real NVIDIA GPU? Am I a real Intel CPU?" They negotiate a secure AES encryption key.
- CPU Boundary Encryption: Data leaving the processor is encrypted before it hits the PCIe slot.
- End-to-End Hardware Tunnel: An encrypted tunnel is formed across the motherboard. A compromised hypervisor attempting a DMA scrape on the PCIe bus sees only "Garbage Data".
- Decryption inside GPU Silicon: The data is decrypted only after it securely enters the NVIDIA Confidential Compute Engine (CCE) inside the GPU.
Phase 3: The 4-Step SRE Deployment Blueprint
Enabling this "Iron Shield" is not a simple
apt install. It is a full-stack hardware-to-software
architecture:
- Hardware Level (BIOS): SREs must enable TDX or SEV-SNP, along
with PCIe security (IOMMU and VT-d), directly in the Bare Metal server's
BIOS.
Cloud Deployment Note: In hosted environments like ServerMO, you cannot access the physical BIOS. You must request a "CC-Ready Reserved Bare Metal" instance, where the data center engineers will enable the CC flags at the BIOS/VBIOS level before delivering the server. - Hypervisor Level: The Linux Kernel must be configured with
CVM-supported KVM/QEMU flags to launch the secure VM. Specifically, you
must inject kernel parameters like
intel_iommu=on(oramd_iommu=on) andkvm_intel.tdx=1into the Host OS GRUB configuration to enable strict IOMMU passthrough. - NVIDIA Driver Level: The GPU firmware must be explicitly booted
into Confidential Computing (CC) Mode via NVIDIA management tools. For a
single GPU, CC mode is set to
on. For multi-GPU HGX clusters using NVSwitches, it must be set toppcie(Protected PCIe). Once enabled, the SRE must executenvidia-smi conf-compute -srs 1inside the guest VM to transition the GPU into the 'Ready State' before workloads can run. - Remote Attestation (The Climax): Before uploading model weights,
your client app sends a cryptographic challenge to the server. The Intel
CPU and NVIDIA H100 generate a hardware certificate stating: "I am
genuine, operating in CC mode, and untampered." Only after verifying
this PASS status do you release the data. In practice, you install the
nv-attestation-sdkand query the NVIDIA Remote Attestation Service (NRAS). NRAS returns a cryptographically signed JWT (JSON Web Token). You must parse this JWT and ensure thex-nvidia-overall-att-resultclaim strictly reads SUCCESS. Do not worry about inference latency. This remote attestation flow is a one-time startup cost. It typically takes 1 to 3 seconds during instance provisioning. Steady-state inference speed is completely unaffected.
Phase 4: The vLLM Crash Trap
Deploying an LLM on a Confidential GPU requires specific runtime configurations.
Phase 5: The Public Cloud FinOps Death Trap
Why not just use AWS, Azure, or GCP for Confidential Computing?
- The 10% Security Tax: Major public clouds charge an explicit premium. For example, enabling AMD SEV-SNP on AWS EC2 adds a 10% surcharge on top of the already expensive hourly rate, and permanently disables essential features like hibernation and Nitro Enclaves.
- Egress & Lock-in: Running a sovereign AI infrastructure requires terabytes of RAG embeddings and continuous data pipelines. Public cloud egress fees will bankrupt an AI startup.
- KMS Lock-in: Azure forces you into Azure Key Vault for
attestation. By deploying on ServerMO's Bare Metal AI Security servers,
you retain the freedom to bring your own KMS (like HashiCorp Vault),
bypass the 10% hourly cloud tax, and achieve 100% data sovereignty. AWS
KMS cannot natively parse NVIDIA GPU attestation JWTs (its
kms:RecipientAttestationonly works for Nitro Enclaves). You are forced to build an intermediate Lambda broker just to verify the GPU. On a Bare Metal setup with ServerMO, you can natively pipe the NVIDIA JWT directly into HashiCorp Vault using the JWT Auth method, completely bypassing cloud vendor lock-in.
Phase 6: Compliance Mandates & Sovereign AI
Enterprise & SRE Technical FAQ
The performance overhead is remarkably low—typically under 3% to 5% for large language models. The AES-256-GCM encryption is handled by dedicated hardware circuits parallel to the HBM controller, meaning it does not consume Tensor Core compute cycles.
In a standard cloud KVM, yes. But in a Confidential VM (CVM) running Intel TDX or AMD SEV-SNP combined with NVIDIA CC, the hypervisor is mathematically locked out of the memory space.
AWS Nitro Enclaves carve out a small, sub-VM boundary with no network access and no persistent storage, making it suitable only for tiny tasks like key signing. Bare Metal CVMs encrypt the entire operating system, enabling you to run massive, network-connected AI models natively.
Yes. Your host kernel must support confidential computing (like
Ubuntu 24.04). The guest OS must also have drivers (like
sev-guest or tdx-guest) to handle the
secure attestation handshakes.
Standard KVM relies on software-level isolation managed by the
host OS. A root admin or compromised hypervisor on standard KVM
can perform a memory dump (gcore or
/dev/mem) to read VM RAM in plaintext. A CVM (built
on Intel TDX or AMD SEV-SNP) encrypts the VM's memory pages in
hardware using ephemeral keys generated inside the CPU silicon.
The hypervisor can manage CPU cycles and I/O scheduling, but any
attempt to read or inspect CVM memory returns cryptographic
garbage.
Because physical BIOS/UEFI settings require out-of-band hardware management, ServerMO delivers "CC-Ready Reserved Bare Metal" instances. Our data center engineers pre-configure the motherboard BIOS (enabling IOMMU, VT-d, AMD SEV-SNP, or Intel TDX) and flash the matching CC-capable GPU VBIOS before handing over the server. You receive a fully hardware-attestable bare-metal host ready for immediate CVM orchestration.
Intel SGX (Software Guard Extensions) was process-based, requiring developers to completely rewrite application code into small "enclaves" with severe memory limits (unsuitable for multi-gigabyte LLMs). Intel TDX (introduced in 4th Gen Sapphire Rapids and 5th Gen Emerald Rapids) operates at the VM level (CVM). It allows a full "lift-and-shift" of entire operating systems, PyTorch, vLLM, and massive 70B+ model weights without modifying a single line of application code.
Early memory-encryption tech (plain SEV) encrypted RAM but allowed a malicious hypervisor to swap or remap encrypted memory pages without the guest knowing. AMD SEV-SNP introduces the Reverse Map Table (RMP) enforced in CPU hardware. Every physical memory access checks the RMP to verify page ownership, GPA-to-HPA mapping, and validation state. If a hypervisor attempts to replay or remap a page, the hardware instantly triggers a CPU fault.
On NVIDIA Hopper (H100/H200), Confidential Computing protects CPU-to-GPU transfers over the PCIe bus via SPDM/IDE AES-256 encryption. However, inter-GPU traffic crossing NVLink bridges is not encrypted at the hardware layer on Hopper. For multi-GPU Tensor Parallelism requiring 100% encrypted inter-GPU communication, you must deploy on NVIDIA Blackwell (B200/B300) architectures, which introduce native, line-rate NVLink Hardware Encryption.
No. Remote Attestation (via NVIDIA NRAS or Trustee) is a one-time
startup check performed during instance provisioning before
sensitive model weights are decrypted. The client challenges the
server, verifies the signed hardware JWT token (verifying
x-nvidia-overall-att-result: SUCCESS), and releases
the decryption keys. Once the model is loaded into encrypted
VRAM, steady-state inference runs at near-native speed with an
overhead of less than 3% to 5%.
Absolutely. Public clouds like AWS force you into complex Lambda
brokers because AWS KMS kms:RecipientAttestation
natively only parses AWS Nitro documents, not NVIDIA GPU JWTs.
On ServerMO Bare Metal, you have full control over the
attestation pipeline. You can pipe the signed NVIDIA NRAS JWT
token directly into HashiCorp Vault’s JWT Auth method to
conditionally release model keys with zero cloud vendor lock-in.
Inside your running CVM, execute
nvidia-smi conf-compute -q. You must verify that
CC State reports ON and
CPU CC Capabilities detects your hardware TEE
(AMD SEV-SNP or INTEL TDX). Next, run
nvattest attest --device gpu --verifier remote to
perform an end-to-end cryptographic check against NVIDIA's Root
of Trust. Once validated, execute
nvidia-smi conf-compute -srs 1 to transition the
GPU to Ready State.

























































