AI infrastructure is putting new pressure on the traditional boundaries between cloud and bare metal. For many years, organizations running demanding AI and HPC workloads have tended to favor bare-metal deployments, mainly to avoid the performance overhead historically associated with virtualization. At the same time, giving up virtualization also means giving up many of the capabilities that make cloud infrastructure easier to operate: isolation, automation, multi-tenancy, workload mobility, and consistent lifecycle management.
This is exactly the gap OpenNebula is designed to address. With today’s launch of the NVIDIA-Certified Hypervisors program, we are proud to announce that OpenNebula has successfully completed NVIDIA-Certified Hypervisor validation for NVIDIA GB200 NVL72. The validation demonstrates that OpenNebula delivers the performance needed for demanding AI workloads while enabling the operational benefits of virtualization through technologies such as PCI passthrough.
Virtualization Without Giving Up Accelerator Performance
The key to this architecture is the combination of virtualization with PCI passthrough, which exposes GPUs and other high-performance devices directly to a virtual machine. Instead of inserting an additional software layer in the accelerator data path, the guest operating system interacts directly with the assigned hardware.
That distinction matters a lot for AI and HPC. OpenNebula can therefore provide the operational advantages of a virtualized cloud platform while keeping the accelerator path close to a bare-metal deployment. The VM remains the unit of isolation and lifecycle management, but the workload gets direct access to the GPU resources it needs. This makes it possible to build AI infrastructure that does not force operators to choose between cloud flexibility and hardware performance.
Validated on NVIDIA GB200 NVL72
For the ARM certification track, OpenNebula 7.2 was validated on NVIDIA Grace Blackwell infrastructure using a virtual machine with direct GPU passthrough. Kubernetes was deployed inside the virtualized environment using OpenNebula’s Elastic Kubernetes Service, OneKS. This means the validation covered not only direct access to Grace Blackwell GPUs through PCI passthrough, but also the kind of Kubernetes-based environment in which organizations can realistically deploy and operate accelerated AI workloads.
The certification exercised a comprehensive set of NVIDIA validation suites covering the most performance-critical aspects of accelerated computing.
More importantly, the NVIDIA-Certified Hypervisor validation is not limited to checking whether a VM can see a GPU. The validation focuses on performance-critical behavior across compute, memory, data-path efficiency, and LLM inference. It also evaluates whether the virtualized environment correctly represents the hardware topology and implements the optimizations required to enable near bare-metal performance. That is particularly relevant for systems such as Grace Blackwell, where performance depends not only on individual GPUs but also on preserving the relationship between CPUs, GPUs, memory, and high-speed interconnects. For more information on the certification test, see the NVIDIA-Certified Hypervisors white paper.
Why Hardware Topology Matters
For conventional virtual machines, abstracting the underlying hardware is often an advantage. For tightly coupled AI workloads, though, topology matters. A workload may depend on the locality between CPU and GPU resources, the way accelerators communicate with one another, or the bandwidth available through the underlying data paths. If the virtualized environment does not expose that topology correctly, performance can suffer even when the individual devices themselves are fast.
The NVIDIA-Certified Hypervisor program specifically looks at this aspect of the platform. NVIDIA’s guidance highlights accurate hardware topology, key performance optimizations, and bare-metal performance as central characteristics of a certified hypervisor. For OpenNebula, this aligns closely with our approach to accelerated infrastructure: virtualize the workload environment, but avoid unnecessarily virtualizing the performance-critical hardware path.
A Cloud Operating Model for AI Infrastructure
The result is an architecture that combines two worlds that have often been treated separately. On one side, organizations get direct access to NVIDIA accelerators and the performance characteristics required by modern AI workloads. On the other hand, those resources remain part of a cloud environment managed through OpenNebula.
That means infrastructure teams can provision and manage GPU-enabled virtual machines using the same mechanisms they already use for the rest of their cloud. They retain strong workload isolation, controlled resource allocation, automation, and support for multi-tenant environments.
For enterprises building private AI infrastructure, research organizations operating shared GPU systems, and cloud providers delivering accelerated infrastructure as a service, this becomes increasingly important. As an infrastructure-first platform, this validation matters to OpenNebula for a broader reason. AI Factories increasingly need to manage accelerated resources as shared infrastructure, with dynamic provisioning, workload isolation, Kubernetes integration, and access for multiple users and teams. Successfully completing NVIDIA-Certified Hypervisor validation on Grace Blackwell shows that virtualization can support this model while preserving the performance required by demanding AI workloads through technologies such as PCI passthrough.




0 Comments