Blog Article:

Inside OpenNebula’s NVIDIA AI Cloud Ready Validated Architecture

OpenNebula has achieved NVIDIA AI Cloud Ready validation on NVIDIA GB200 NVL72.

Read the official announcement here.  

But what does that validation actually mean from an infrastructure perspective?

For us, the important part is that this was not simply a test of whether OpenNebula could provision a GPU-enabled virtual machine. The validated environment brings together the different infrastructure and orchestration layers required to operate a production AI cloud: physical infrastructure lifecycle management, virtualization, Kubernetes, HPC scheduling, accelerated computing and multi-tenancy.

The result is an architecture in which OpenNebula acts as the cloud management and orchestration layer across the AI factory, integrating NVIDIA infrastructure technologies with established cloud-native and HPC environments.

From NVIDIA GB200 NVL72 to a Multi-Tenant AI Cloud

The NVIDIA AI Cloud Ready validation initiative provides a technical validation framework for infrastructure and platform software deployed by NVIDIA Cloud Partners. Solutions are assessed against the NVIDIA Cloud Partner Software Reference Guide, covering requirements across the networking, compute, orchestration and AI platform layers. The objective is not simply to determine whether individual software components operate correctly, but whether the resulting stack can provide the capabilities required for production-scale AI infrastructure.

OpenNebula completed the validation as an integrated infrastructure platform combining OpenNebula with NVIDIA DSX OS™ components running on NVIDIA GB200 NVL72. The validated software architecture combines:

  • OpenNebula as the common cloud management, multi-tenancy and orchestration layer across the environment, including capabilities for workload scheduling, bare-metal resource management, virtualization, Kubernetes lifecycle management, service provisioning and infrastructure automation.
  • NVIDIA Infra Controller for automated bare-metal infrastructure lifecycle management.
  • Ubuntu as the operating system on the compute nodes.
  • KVM virtualization, managed through OpenNebula, for VM-based cloud and AI services.
  • RKE2 Kubernetes clusters, provisioned and lifecycle-managed through OpenNebula OneKS.
  • Slurm for HPC and distributed AI workloads.
  • NVIDIA NIM™ Microservices for GPU-accelerated LLM inference and model serving.
  • Additional open source components for observability, security, identity and access management, monitoring and supporting infrastructure services.

The important point is how these components work together.

One Control Plane, Multiple Ways to Consume Accelerated Infrastructure

AI factories increasingly need to support very different consumers. One team may need direct access to bare-metal accelerated infrastructure. Another may require virtual machines with GPU resources. AI development teams increasingly expect Kubernetes clusters on demand, while research and HPC environments continue to depend on Slurm for distributed workloads.

Building an independent infrastructure stack for each of these models creates operational silos. OpenNebula takes a different approach. The physical accelerated infrastructure becomes part of a common cloud environment from which operators can expose different consumption models according to the requirements of each tenant or workload:

NVIDIA AI ready validation 1

OpenNebula provides the management and orchestration layer across these environments, while NVIDIA infrastructure technologies provide the accelerated computing, networking and infrastructure capabilities underneath them.

This means an AI cloud operator can manage infrastructure centrally while allowing users to consume it through the model best suited to their workload.

Automating the Physical AI Infrastructure with NVIDIA Infra Controller

At rack scale, cloud automation has to extend below the hypervisor and Kubernetes layers. NVIDIA Infra Controller provides bare-metal infrastructure lifecycle management and exposes an API-first model that cloud platforms can integrate with their own control planes.

Within the validated architecture, this provides the bridge between the physical NVIDIA accelerated infrastructure and OpenNebula’s cloud management layer. Infrastructure can therefore move through its lifecycle from hardware discovery and provisioning to becoming tenant-ready capacity without requiring operators to manage the physical and cloud layers as independent environments.

This becomes particularly important in multi-tenant AI clouds, where infrastructure has to be allocated, isolated, reclaimed and prepared for the next workload in a repeatable way.

Virtualization with KVM

Virtual machines remain an important part of AI infrastructure. Not every AI workload belongs in Kubernetes, and cloud operators frequently need VMs for development environments, inference services, supporting infrastructure, legacy applications, and workloads requiring stronger isolation boundaries.

The validated architecture uses Ubuntu and KVM on the compute nodes, with OpenNebula managing virtual machine lifecycle, placement, resource allocation, and multi-tenancy. For accelerated workloads, the architecture supports direct device passthrough of GPUs and other high-performance devices to virtual machines, allowing workloads to access the underlying hardware directly and avoid the performance penalties associated with additional software abstraction layers.

This makes it possible to combine the isolation and operational flexibility of virtualization with native access to NVIDIA accelerated infrastructure, while keeping VM provisioning, scheduling, lifecycle management, and policy enforcement under the OpenNebula control plane.

In practice, this means the same OpenNebula environment can support conventional cloud workloads alongside performance-sensitive AI and HPC workloads, without forcing operators to choose between virtualization and direct hardware access.

Kubernetes-as-a-Service with OneKS

Kubernetes has become one of the primary execution environments for AI workloads. In the validated architecture, OpenNebula uses OneKS to provision and lifecycle-manage elastic RKE2 Kubernetes clusters on top of the underlying infrastructure.

Instead of requiring the AI cloud operator to maintain a separate Kubernetes management platform, Kubernetes becomes another service exposed through the OpenNebula cloud. Tenants can manage isolated Kubernetes environments while the provider retains centralized control over infrastructure allocation, lifecycle and governance.

This makes it possible to combine the elasticity expected from cloud infrastructure with the Kubernetes environments required by modern AI frameworks and applications.

Slurm and HPC Workloads

AI infrastructure also increasingly intersects with HPC. Large-scale training, scientific AI and other distributed computing workloads frequently depend on Slurm, and many organizations investing in accelerated computing already have substantial Slurm expertise and workloads.

The validated architecture therefore includes Slurm alongside virtualization and Kubernetes rather than treating HPC as a separate infrastructure domain. This allows the same accelerated infrastructure environment to support both cloud-native and HPC consumption models.

For organizations operating combined AI and HPC environments, this is particularly important: infrastructure does not have to be permanently divided into independent Kubernetes, virtualization and HPC islands.

From Infrastructure Silos to an AI Factory

This is ultimately what we believe the NVIDIA AI Cloud Ready validation demonstrates. An AI factory is more than a collection of GPUs. Production environments require physical infrastructure management, networking, isolation, virtualization, container orchestration, HPC scheduling, multi-tenancy, monitoring and automation. More importantly, these capabilities have to work together as an integrated system.

OpenNebula provides the common cloud layer connecting these domains.

NVIDIA AI CLOUD VALIDATION 850

The architecture gives infrastructure operators a common operational model while preserving flexibility at the workload layer.

Built for NVIDIA’s Current and Next Generation of AI Infrastructure

The NVIDIA AI Cloud Ready validation was performed on NVIDIA GB200 NVL72 infrastructure, providing a concrete rack-scale validation of the architecture. At the same time, the architecture is designed around infrastructure lifecycle management and orchestration rather than being tied to a single GPU generation. That becomes increasingly important as operators move from NVIDIA Blackwell infrastructure based on B200 GPUs toward NVIDIA Blackwell Ultra B300 and GB300 NVL72 systems, and eventually toward future generations of NVIDIA accelerated computing.

The objective is to give cloud operators a management architecture that can evolve alongside the underlying accelerated infrastructure rather than requiring the cloud stack to be redesigned for every hardware generation.

An Open European Platform for AI Factories

OpenNebula is developed in Europe as a fully open source cloud platform based on a single Apache-licensed codebase, a model that aligns closely with emerging cloud and AI sovereignty requirements.

The European Commission’s proposed Cloud and AI Development Act (CADA) places increasing emphasis on EU-based infrastructure, independence from third countries, software supply-chain transparency, European control, and open source technologies.

For cloud providers, enterprises, research organizations and public-sector operators, OpenNebula provides a way to combine NVIDIA accelerated computing with greater control over the cloud management layer that governs infrastructure, tenancy, workload placement and access to accelerated resources.

With NVIDIA AI Cloud Ready validation, OpenNebula combines these sovereignty characteristics with an integrated software stack technically validated for production NVIDIA Cloud Partner environments.

Production AI infrastructure also requires an operational support model that matches the criticality of the platform. OpenNebula Systems provides 24/7 support for the integrated OpenNebula-based environment, including lifecycle support, troubleshooting and assistance across the cloud management layer and its validated integrations with NVIDIA infrastructure technologies. This helps operators reduce the complexity of managing a multi-component AI factory and gives them a clear support path for production deployments.

What’s Next

The NVIDIA AI Cloud-Ready validation builds on our broader technical collaboration with NVIDIA and represents another milestone in a wider validation and integration roadmap.

OpenNebula has been integrating NVIDIA DSX OS™ technologies into its AI Factory architecture and has already completed additional NVIDIA validation initiatives, including NVIDIA Spectrum-X™ Ethernet networking and NVIDIA GPU virtualization for high-performance, multi-tenant AI environments. Together, these capabilities allow operators to combine full GPU passthrough for performance-sensitive workloads with virtualized GPU consumption models that improve flexibility, isolation and infrastructure utilization.

The next step is to extend this work further across NVIDIA AI infrastructure and workload management stack. This includes deeper integration with technologies such as NVIDIA Run:ai and NVIDIA Dynamo, together with the relevant interoperability and validation processes as those integrations mature.

At the same time, we will continue expanding support as NVIDIA accelerated computing infrastructure evolves, while keeping the same principle at the center of the architecture: one open cloud management layer for bare metal, virtual machines, Kubernetes, HPC and AI infrastructure.

For NVIDIA Cloud Partners and Neocloud providers, this offers a path from accelerated hardware to flexible, multi-tenant cloud services, with different ways of consuming GPU capacity according to workload requirements.

For enterprises, HPC centers and public-sector organizations, it provides an architecture for building private and sovereign AI factories without giving up control of the cloud layer.

For OpenNebula, NVIDIA AI Cloud-Ready validation confirms that an open cloud architecture and state-of-the-art NVIDIA accelerated computing can operate together as part of the same production AI factory, while providing a foundation for further integration and validation across NVIDIA’s evolving AI infrastructure stack.

Read the official press release here.

Meet the OpenNebula team at NVIDIA GTC™ Berlin and learn more about our latest work with NVIDIA on AI infrastructure.

 

Ruben S. Montero

Chief Technology Advisor at OpenNebula Systems

Sep 29, 2026

0 Comments

Submit a Comment