Tools

AI-generated text

DIN Deploy: C++ examples using ONNX Runtime and NVIDIA TensorRT RTX for local AI acceleration

DIN Deploy is an open-source collection of C++ samples that show how to convert models to ONNX and run them in native, hardware-accelerated applications on Windows and Linux using ONNX Runtime together with the NVIDIA TensorRT RTX execution provider.

DIN Deploy: C++ examples using ONNX Runtime and NVIDIA TensorRT RTX for local AI acceleration

Do Inference Now (DIN) Deploy is an open-source set of practical C++ examples that demonstrates how to take AI models from checkpoints to native, hardware-accelerated applications on Windows and Linux. It pairs ONNX Runtime with the NVIDIA TensorRT RTX execution provider to enable accelerated inference. The same ONNX Runtime API can also be accessed via WinML 2.0.

Export versus deployment

Each DIN Deploy sample separates model export from deployment. A Python exporter downloads a model checkpoint from Hugging Face and converts it to an ONNX artifact. The application side is a native C++ command-line client built on ONNX Runtime (ORT). This separation lets teams convert models independently of the deployment logic, so an exported ONNX model can be integrated into a local application without requiring a model-specific runtime.

Most sample code uses ONNX Runtime session and tensor APIs in C++. Vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths. Any execution provider that supports the required ONNX Runtime tensor APIs can run the shared code. ORT’s copy tensor API helps manage data locality without embedding vendor-specific APIs into the shared code.

For preprocessing and postprocessing around exported-model inference, the FLUX.2 sample leverages ONNX Runtime’s graphics interop capability (introduced in ORT 1.25), using Vulkan and DirectX for sampling.

Platforms and build

The repository includes CMake presets for Windows and Linux, with Arm64 variants. DirectX support is available only on Windows. By default, CMake downloads ONNX Runtime and TensorRT RTX.

Tasks and example workloads

  • Automatic speech recognition (ASR): DIN Deploy provides both offline and streaming pipelines. OpenAI Whisper covers offline transcription at multiple model sizes, while NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming provide streaming pipelines. Samples demonstrate moving audio through a native application and returning transcriptions while using GPU acceleration where available.
  • Image segmentation: Meta SAM 2.1 samples support interactive masking for images and video, converting model outputs into segmentation masks for selection, tracking, and other computer-vision workflows.
  • Prompt-driven image generation: The FLUX.2-klein-4B sample demonstrates prompt-based image generation and includes graphics-API interops with Vulkan and DirectX so applications can integrate GPU-resident resources through a cross-vendor shader interface.

The sample also shows how post-training quantization (PTQ) with NVIDIA Model Optimizer can produce a quantized ONNX model. Quantization is hardware-dependent, but because the ONNX interfaces remain the same, the quantized model can be used as a drop-in replacement without changes to application code.

Performance examples (selected measurements)

The repository reports GPU and CPU performance measured on DGX Spark for several DIN Deploy workloads. For audio tasks the results are presented as multiples of real time (higher is faster):

  • openai/whisper-large-v3-turbo: GPU 58.5x, CPU 3.8x
  • nvidia/nemotron-3.5-asr-streaming-0.6b: GPU 39.01x, CPU 3.24x
  • nvidia/parakeet-tdt-0.6b-v3: GPU 206.41x, CPU 14.44x
  • facebook/sam2.1-hiera-base-plus: GPU 38.3 FPS, CPU 0.5 FPS

These measurements illustrate substantial GPU speedups for the selected models when run on supported hardware.

Getting started

Use the repository’s CMake presets for Windows, Linux, x86-64, and Arm64. After configuring and building the project, export a model to ONNX and run the CLI with TensorRT RTX, or integrate the sample code into your own application and use the provided pipeline implementations. CMake will download ONNX Runtime and TensorRT RTX by default.

Further reading

The project references additional resources for deeper information: TensorRT for RTX, NVIDIA Local AI, the DIN Deploy repository documentation, and an NVIDIA blog series on model quantization for guidance on quantization and local AI integration.