NVIDIA has published AI Cluster Runtime (AICR) v1.0, a release that formalizes version-locked recipes for configuring GPU-accelerated Kubernetes clusters and establishes a stable compatibility contract across the project’s public interfaces. Each AICR recipe pins component combinations that are known to work together, renders deployment artifacts for multiple deployers (Helm, Argo CD, Flux, Helmfile), and includes signed validation evidence from the hardware used for testing.
Why a reproducible contract is needed
GPU-accelerated Kubernetes clusters depend on compatible versions and settings across many independently released components: host kernels, GPU drivers, container runtimes, Kubernetes itself, networking, storage, device plugins, operators, schedulers, and workload frameworks. Upgrading or changing any one component can break a previously working configuration. A setup validated for a specific service, GPU generation, fabric type, machine shape, and Kubernetes release can fail under other conditions, and small version differences are often hard to trace after deployment.
Moreover, successful installation of components does not guarantee that the cluster meets the recipe’s intended configuration: components may not be healthy, required capabilities like gang scheduling or accelerator discovery might not function, and measured performance may fall short of thresholds.
Before AICR, knowledge about which combinations work and how they were validated lived in separate validation systems, deployment scripts, and runbooks, making discovery, reproduction, and updates difficult for teams.
What AICR provides
AICR implements four core, deliberately independent capabilities:
- Snapshot — records the observed cluster state, including Kubernetes, operating system, kernel, GPU and topology details.
- Recipe — describes the desired, version-locked component configuration and the constraints and validation phases that apply.
- Bundle — renders the recipe into artifacts for the operator’s preferred deployment tooling.
- Validation — compares the recipe with observed state and, where declared, runs deployment, conformance, and performance checks against the cluster.
These pieces can be combined in multiple sequences: snapshot data or explicit target criteria can produce a recipe; a recipe can be rendered into a bundle for an existing deployer; the recipe and observed state feed validation; and operators can explicitly verify bundles and evidence via the corresponding commands. Common open-source CD tools (Helm, Argo CD, Flux, Helmfile) are responsible for applying or reconciling bundles, while AICR validates and records signed evidence of the result.
What’s new in v1.0
AICR v1.0 defines compatibility rules and committed baselines for its public surfaces so integrators and operators can build against stable interfaces. The release covers:
- the aicr CLI public commands, flags, exit semantics, and structured output;
- the aicrd REST API and OpenAPI contract;
- the exported API of the github.com/NVIDIA/aicr/pkg/client/v1 package;
- the generated bundle layout and AICR artifact schemas.
Each public interface has a committed baseline that is checked before changes are merged. The release policy specifies semantic breaking changes: after v1.0, removing or incompatibly changing a stable public interface requires a new major release.
For Go integrators, pkg/client/v1 exposes the supported workflow without requiring imports from AICR’s internal packages. The CLI and REST server use the same facade to reduce the risk of differing behavior between public entry points.
Validation dashboard and ecosystem integrations
AICR’s validation dashboard lets operators find recipes by service, GPU, operating system, workload intent, and optional platform, and inspect each recipe’s status and any published evidence for the hardware configuration tested. Contributors can propose recipes for environments the maintainers cannot test, validate them on their own clusters, and submit signed evidence for maintainer review.
The recipe model is already being used in the ecosystem: Pulumi Labs exposes AICR through an infrastructure-as-code provider, and Mirantis’s k0rdent integration packages it for multi-cluster management. These integrations demonstrate the value of defining GPU-accelerated Kubernetes configuration once and consuming it through different tools.
AICR has grown in the last six months from a handful of recipes to a library that spans major Kubernetes services and the current NVIDIA accelerator portfolio, rendered as deployer-neutral bundles. The project now has over 100 distinct contributors, with almost half coming from outside NVIDIA.
Try it and contribute
Operators can try a recipe for their environment, inspect its status and any published evidence, and run the dashboard’s aicr evidence verify command where evidence is available. Contributions are welcomed for hardware and cluster combinations not currently covered. Possible contributions include adding features, integrations, or documentation; proposing and validating new recipes and submitting signed evidence; and reporting bugs or requesting features via GitHub Issues. Start with the project repository, contributing guide, and issue tracker to get involved.
Example workflow
An operator can select criteria such as EKS, GB300, Ubuntu, training intent, and Kubeflow; resolve those criteria to a pinned recipe; render it for Argo CD; deploy via an existing GitOps workflow; and validate the running cluster against the same recipe. The same recipe can be rendered for Helm, Flux, or Helmfile without changing the intended configuration.
In short, AICR v1.0 aims to provide a reproducible, verifiable contract for GPU-accelerated Kubernetes cluster configuration and a stable platform for operators, integrators and contributors to build on.



