Inside the k0rdent AI + NICo Integration: Bare-metal Lifecycle and DPU-enforced Tenant Isolation
)
Summarize with AI
At AI factory scale, provisioning stops being a setup step and becomes a continuous operation. Nodes cycle through tenant assignments. Software stacks move forward a generation while the racks underneath them stay in service. Every one of those transitions has to be auditable, secure, and reversible — and none of it survives being done by hand. Manual provisioning and human-managed tenant boundaries may hold up fine in a lab, but they will fall apart the first time a multi-tenant GPU cloud hits real operational tempo.
The NVIDIA Infra Controller (NICo) solves this problem, and it's where we've focused our integration work in Mirantis k0rdent AI. Here's an overview of what NICo does, what it deliberately doesn't do, and how we've wired it into a control plane that takes an operator from racked hardware to a consumable AI service.
What NICo actually handles
NICo makes two things programmable that used to be manual and fragile.
The first is bare-metal lifecycle management through an API. A node enters the fleet — discovered, firmware-validated, and provisioned — gets assigned to a tenant, serves its workloads, and eventually gets reclaimed and reassigned, all as API-driven state transitions rather than runbook steps. That matters because at scale the interesting events aren't the initial build; they're the thousands of subsequent reassignments and replacements, each of which has to leave an audit trail.
The second, and the one worth understanding properly, is hardware-enforced tenant isolation through BlueField DPUs and the DOCA Platform Framework (DPF). The distinction here is hardware-enforced. In a conventional host, the isolation boundary lives in software running on the same CPU as the tenant workload. Push enough tenants through that model and the boundary becomes the thing you worry about. With a BlueField DPU, the infrastructure control functions (i.e.,networking, storage, the isolation boundary itself) are offloaded onto the DPU, below and outside the host the tenant runs on. The tenant doesn't share a trust domain with the control plane that isolates it. A compromised or noisy host has nothing to cross into, because the boundary isn't running where the host can reach it.
Why a provisioning controller isn't a platform
NICo provisions metal and isolates tenants. It does not, on its own, give you a sellable AI cloud. DSX OS is a set of open, modular components meant for incremental adoption, not a finished platform. Look at what sits around NICo in DSX OS, and the gap is obvious:
NVIDIA AI Cluster Runtime (AICR) captures validated runtime configurations as version-locked recipes, so the software state on a freshly provisioned node matches the rest of the fleet instead of drifting into the silent failures that configuration skew produces.
NVIDIA Run:ai provides AI workload and GPU orchestration, maximizing utilization across clusters.
NVSentinel does Kubernetes-native GPU health monitoring and fault remediation.
Networking, storage, and multi-cluster governance still have to come from somewhere.
Each of these is a separate component with its own lifecycle. The operational reality is that someone has to compose them into one coherent control plane, keep that composition consistent as nodes churn, and carry it forward when the next hardware generation lands without rebuilding from scratch. That composition problem, not provisioning in isolation, is the actual job of building and running a commercial AI cloud.
How k0rdent AI composes it
We've embedded NICo into the core of k0rdent AI, not as a separate add-on. This enables bare-metal lifecycle and tenant isolation to become native operations of the control plane, not a separate system an operator has to reconcile against everything else.
Here is a node’s lifecycle in the integrated stack:
A node joins the fleet. k0rdent drives NICo's API to provision the bare metal, without requiring a per-node runbook.
BlueField and the DOCA Platform Framework (DPF) establish the hardware-enforced tenant boundary for whoever that node is assigned to.
AICR applies the version-locked runtime recipe, so the node's software state is identical to the rest of the tenant's fleet on day one.
k0rdent composes the platform layer on top, including operators, networking, storage, and observability, across however many clusters the tenant spans.
NVIDIA Run:ai schedules the tenant's training or inference workloads onto the GPUs, and the operator is selling capacity.
When the workload completes, the environment tears down, the node returns to the pool, and the isolation boundary is re-established cleanly for the next tenant.The whole transition recorded.
The result is one control plane across the full NVIDIA NCP reference architecture portfolio. An operator plugs in NVIDIA-certified hardware and k0rdent AI manages it end to end, Metal-to-Model. We've validated the integration against NVIDIA Hopper and Grace Blackwell, and because the stack is built on open, cloud-native foundations, it carries forward to the next reference architecture as it arrives, rather than forcing a replatform every hardware cycle.
That is what turns AI infrastructure into a multi-generation investment instead of a per-cycle rebuild. An operator standing up Grace Blackwell today has day-zero readiness for NVIDIA Vera Rubin and the generations beyond it, using the same operational model, and the same provisioning workflows, and no need for infrastructure redesign. Aligning the control plane directly to NVIDIA's accelerated computing roadmap lets a long-horizon hardware purchase keep paying off across generations.
Open building blocks, open control plane
There's a strategic reason this is built the way it is, and it's the same pattern that played out with operating systems, with mobile, and with cloud, where Kubernetes became the open standard against proprietary hyperscaler APIs. AI infrastructure is the next wave, and the standard for it hasn't been set yet.
NVIDIA released DSX OS as open building blocks precisely so the ecosystem could pick them up incrementally and build differentiated platforms on a validated, open foundation. Mirantis isn't just consuming those building blocks downstream: as a primary open source infrastructure contributor to NICo, we work with NVIDIA upstream on how AI datacenter provisioning evolves across future accelerated computing platforms. Mirantis is also an inaugural partner for NVIDIA AI Cloud Ready, the validation initiative that qualifies and validates AI infrastructure software for deployment on AI clouds.
We think the platform layer should be open in the same spirit: an operator should be able to adopt NICo, compose it with the rest of DSX OS and with a certified ecosystem of hardware, networking, and storage — Dell, Supermicro, BE Networks, Netris, Saturn Cloud, VAST Data, and others — and not get locked into a closed stack to do it. Open building blocks deserve an open control plane on top.
Warren Barkley, vice president of product management at NVIDIA, stated: "DSX OS gives the ecosystem open, modular building blocks for operating AI factories at scale. By embedding the NVIDIA Infra Controller into k0rdent AI, Mirantis is giving operators an open control plane to turn validated NVIDIA infrastructure into production AI services.".
That's the integration in a sentence: NICo gives you programmable metal and a hardware-enforced tenant boundary; k0rdent AI turns that, plus the rest of DSX OS and a certified open ecosystem, into a modern multi-tenant AI factory.

)
)
)

)
)
