Tools

AI-generated text

Deploying HSTU Generative Recommenders with NVIDIA Dynamo‑Triton, PyTorch AOTI and FlexKV

NVIDIA demonstrates a production-ready workflow for serving Hierarchical Sequential Transduction Unit (HSTU) generative recommender models using Dynamo‑Triton, PyTorch Ahead‑of‑Time Inductor (AOTI) and FlexKV-backed GPU KV caching.

Deploying HSTU Generative Recommenders with NVIDIA Dynamo‑Triton, PyTorch AOTI and FlexKV

NVIDIA presents a complete production workflow for serving Hierarchical Sequential Transduction Unit (HSTU) generative recommender models. The workflow integrates HSTU model structure with PyTorch Ahead‑of‑Time Inductor (AOTI) export/compilation, NV Embedding Cache, a FlexKV‑backed GPU KV cache service, native C++ validation, and Dynamo‑Triton deployment. The objective is low‑latency inference for long user histories and large embedding tables typical of modern personalization systems.

Why use HSTU for generative recommendation?

Generative recommender (GR) systems frame recommendation as sequence modeling over user behavior instead of disjoint retrieval, ranking, and prediction stages. HSTUs accept token sequences composed of contextual tokens, item tokens and optional action tokens. Preprocessing retrieves embeddings, interleaves item and action embeddings when present, appends context and positional encoding, HSTU blocks process the sequence, and a prediction head emits multitask ranking outputs.

This architecture models recency, order and repeated interaction patterns well, but inference can be expensive if long histories are recomputed on every request.

Serving challenges for large sequential recommenders

Serving large sequential recommenders differs from serving compact dense ranking models: the stack must handle jagged sequence inputs, very large categorical embedding state, long histories, and request patterns where the same user returns with only a small amount of new data. Recomputing the full key‑value state for unchanged prefixes wastes computation and increases latency.

KV caching stores previously computed key‑value data so the model can avoid recomputing cached portions of a user’s history — particularly valuable when long‑term history remains mostly stable while new candidates or recent actions arrive.

The NVIDIA HSTU inference workflow

NVIDIA’s workflow includes a KVCacheManager that uses GPU memory and host storage for a paged KV table supporting lookup, allocation, append and eviction (LRU‑style when GPU space is constrained). Host storage provides an additional tier and FlexKV serves as the KV‑cache backend runtime.

The HSTU attention kernel can consume KV data from the paged cache; the exported inference path includes cache‑aware custom ops for lookup, allocation, onboarding, append and offload. This preserves the model’s sequence semantics while reducing redundant computation.

PyTorch AOTI for native inference

PyTorch AOTI exports a PyTorch model with torch.export and compiles it ahead of time into a package loadable by a native C++ runtime, lowering Python runtime overhead and producing a deployment‑friendly artifact for the Dynamo‑Triton torch_aoti backend. The export contains the AOTI archive (.pt2) plus metadata and embedding table files.

In the NVIDIA example, the embedding implementation uses DynamicEmb inference embedding tables combined with NV Embedding Cache so only hot embeddings are resident in GPU memory while the full table lives in CPU memory. The export writes layer metadata and embedding table data alongside the compiled archive to avoid duplicate embedding copies.

The workflow validates exported artifacts via Python export scripts that replay tensors and native C++ executables that load and replay the exported model for correctness and performance. Dynamo‑Triton then uses the same AOTI package and replay path to keep development validation and production serving aligned.

Dynamo‑Triton deployment path

Dynamo‑Triton provides the production serving layer: model repository management, request handling, backend integration, metrics and deployment structure. The AOTI deployment uses Dynamo‑Triton’s PyTorch backend with platform: "torch_aoti" so the server can load and serve the ahead‑of‑time compiled package.

The full workflow includes five stages:

  • Build required custom operators and runtime libraries
  • Export the HSTU ranking model with PyTorch AOTI
  • Start the FlexKV‑backed KV‑cache service
  • Validate the exported artifacts with native C++ replay
  • Serve the exported KV‑cache AOTI model with Dynamo‑Triton

NV Embedding Cache reduces GPU memory requirements by keeping only hot parts of embedding tables in GPU memory. FlexKV reduces recomputation for long user histories by caching attention blocks. Together, they form a serving stack aimed at realistic generative recommender inference.

Benchmarking HSTU serving latency

Benchmarks in NVIDIA/recsys‑examples measure latency on a single NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU using the KuaiRand‑1K ranking configuration. Key benchmark parameters:

  • Models: 3‑layer and 8‑layer HSTU variants
  • Hidden size: 512
  • Attention heads: 4
  • Weights: BF16
  • KV cache: BF16
  • Maximum history sequence length: 8,192 tokens (4,096 item+action pairs)
  • Maximum candidate sequence length: 100
  • Contextual features: 6
  • Effective aligned sequence length exported: 8,320 tokens

Latency is reported per logical request; Dynamo‑Triton calls contain one logical batch. Measured time excludes dataset loading, validation, rebatching, user‑ID generation, server startup, warmup and post‑warmup sleep.

Main results:

  • At Dynamo‑Triton batch size 2, PyTorch AOTI reduces latency compared with the Dynamo‑Triton Python backend even without KV cache hits; with a 20 GB GPU KV cache and hits, latency improvement is significantly larger.
  • With batch size 8 and 100% GPU KV‑cache hit rate, Dynamo‑Triton + PyTorch AOTI + FlexKV achieved up to:
    • 4.47× speedup for the three‑layer HSTU model (versus same AOTI config without KV caching)
    • 5.93× speedup for the eight‑layer HSTU model
  • Absolute latency at batch size 8 with GPU KV‑cache hits:
    • 3‑layer HSTU: ~0.423 ms per logical request
    • 8‑layer HSTU: ~0.678 ms per logical request

These results indicate that combining compiled model execution with cache‑aware serving yields strong latency reductions, especially for deeper models and larger logical batch sizes where avoiding recomputation across many layers matters most.

Benefits of accelerating HSTU GR inference

Recommendation systems operate under tight latency budgets; additional ranking latency affects page load, feed responsiveness and ad deadlines. As sequence‑aware, personalized models grow more complex, inference compute increases.

The presented stack—HSTU architecture, PyTorch AOTI compiled artifacts, FlexKV KV caching and NV Embedding Cache, served via Dynamo‑Triton—addresses this by lowering runtime overhead, reducing GPU memory usage and reusing previously computed attention state so only newly appended tokens require computation.

The savings scale with sequence length and model depth, because longer sequences and deeper HSTU layers otherwise multiply redundant attention computation.

How to get started

The NVIDIA/recsys‑examples GitHub repository contains the HSTU overview, AOTI inference guide, scripts to build images and libraries, prepare KuaiRand‑1K data, train a checkpoint, export the KV‑cache AOTI model, validate with C++ replay, package the Dynamo‑Triton runtime image and replay requests through the server.

Acknowledgments

This work is a cross‑functional effort across several NVIDIA teams. NVIDIA thanks J, Runchu Zhao, Yulu Liu, Lin Hu, Zhuofan Li, Jacob Subag and Tomer Bar‑On for their contributions.