Agentic AI extends inference into distributed workflows: a single request can invoke multiple model calls, tool executions, memory lookups, policy checks, storage accesses and network transfers before producing a final answer. That shift requires AI factory infrastructure to move, protect, retrieve and reuse data quickly enough to keep GPUs and CPUs productive.
Role of BlueField and DOCA
The NVIDIA BlueField platform provides dedicated, programmable infrastructure processing in the AI factory data path. BlueField offloads infrastructure work from host CPUs, accelerates data movement, enforces policy inline, and enables context reuse. Those capabilities translate into practical outcomes such as higher GPU utilization, more predictable latency, stronger tenant isolation, lower cost per token and improved tokens per watt.
The solution comprises BlueField‑4 DPUs, Vera BlueField‑4 STX storage processors, and the DOCA software stack. BlueField‑4 DPUs offload, accelerate and isolate networking, storage, security, telemetry and control‑plane services from host CPUs while speeding data movement across GPU and CPU systems. Vera BlueField‑4 STX storage processors power new data platforms for context memory, high‑performance storage infrastructure and secure data services.
DOCA provides the software foundation and programmability to build and operate these services across networking, storage, security, telemetry and lifecycle management. In the NVIDIA Vera Rubin platform and the broader NVIDIA DSX architecture, BlueField supplies accelerated infrastructure while DOCA supplies the programmable model for deploying those services across the data center.
Why infrastructure becomes part of inference
Agentic AI makes inference a cross‑component workflow spanning GPUs, CPUs, memory, networking, storage and security. GPUs run the model inference and generate tokens; CPUs orchestrate the agent runtime by executing tools, processing retrieval results, preparing prompts, validating outputs and coordinating reasoning steps. Preserving and retrieving context across turns—especially KV cache—is essential.
During prefill, an LLM creates KV cache that stores intermediate attention state. As prompts, conversations and agent workflows grow, cache state must persist and be reused. When GPU memory is constrained, systems may evict and recompute KV cache, limit context length, or move state to another memory tier—trading off latency, throughput or cost. Thus KV cache becomes part of the infrastructure data path and must be placed, protected and retrieved without slowing inference.
Network, storage, security, telemetry, control‑plane services and context‑memory management must process agent traffic inline so they do not delay GPU inference or consume host CPU resources essential for agent execution.
BlueField‑4 as the AI factory operating system
Agentic inference requires a fast, secure, programmable infrastructure data path. BlueField supplies a dedicated infrastructure processor that connects, secures, isolates and accelerates services across the AI factory.
BlueField‑4 operates as the DPU across Rubin GPUs and NVIDIA Vera CPUs. Key BlueField‑4 characteristics include:
- Up to 800 Gb/s Ethernet or InfiniBand connectivity,
- Integration of a 64‑core NVIDIA Grace CPU,
- High‑bandwidth LPDDR5X memory,
- PCIe Gen6 connectivity,
- Inline acceleration for networking, storage, security and data movement,
- And support for the DOCA software platform.
Compared with BlueField‑3, BlueField‑4 doubles networking bandwidth, delivers up to 6× more compute performance, 4× memory capacity and more than 3× memory bandwidth.
The BlueField‑4 STX storage processor combines the NVIDIA Vera CPU, NVIDIA ConnectX‑9 SuperNIC, up to 1.6 Tb/s Spectrum‑X Ethernet connectivity, high‑performance NVMe storage access, accelerated data movement and in‑silicon security together with DOCA programmability.
DOCA: turning hardware into programmable services
DOCA provides libraries and microservices to create and deploy accelerated infrastructure services. Selected DOCA capabilities described include:
- DOCA Host‑Based Networking (HBN): supports server‑side Layer‑3 routing with BlueField acting as a BGP router for scalable multi‑tenant designs.
- BlueField ASTRA: enables Spectrum‑X zero‑trust, multi‑tenant bare‑metal deployments with BlueField as the managed control point alongside ConnectX‑9.
- DOCA Memos: helps manage and share KV cache across compute and storage nodes so context can be reused for long‑context and agentic inference.
- DOCA security services: provide zero‑trust access, policy enforcement, runtime visibility and isolation to protect data, inference and agents across the AI factory.
Together these features keep networking, storage, security, context management and control services close to the data path instead of competing for host CPU resources, helping improve GPU utilization, reduce inference latency, strengthen multi‑tenant isolation and lower cost and energy per token.
System‑level capabilities and the Vera Rubin AI factory
The Vera Rubin platform is engineered for agentic AI workloads that require high‑throughput inference, dense CPU execution, large‑scale context memory and secure data movement. Rubin GPUs provide accelerated compute; Vera CPUs handle tool calls, orchestration and data movement. NVLink supports scale‑up communication while ConnectX‑9 and Spectrum‑X support scale‑out networking. BlueField operates across compute, networking, storage and security in this system.
GPU compute: BlueField‑4 acts as the infrastructure processor in GPU compute trays, supporting frontend (north‑south) networking, host CPU offload, secure data access, management, observability and infrastructure isolation. By offloading and isolating these services, BlueField helps prevent host CPU overhead, variable access controls or management traffic from gating GPU compute.
CPU execution: agentic workloads increase CPU demand because agents execute tools, run code, query data, transform inputs and orchestrate workflows. Vera CPUs provide high per‑core performance, concurrency and power‑efficient memory bandwidth for this layer. BlueField‑4 handles front‑end networking, storage access, security and control‑plane services in the DPU domain so Vera CPUs can focus on agent execution rather than infrastructure tasks.
AI‑native storage and context memory: agentic AI makes context active infrastructure data. The NVIDIA CMX context memory storage platform creates a shareable, scalable, energy‑efficient AI‑native storage tier between GPU memory and shared storage, optimized for KV cache using the Vera BlueField‑4 STX storage processor.
The storage processor runs KV I/O, metadata management, data placement, security and control operations close to the storage and network path. Rather than presenting flash as block storage, it tracks KV metadata, manages queues, supports cache recall and pre‑staging, enforces tenant policy and handles data protection. Placing these latency‑sensitive, control‑heavy operations in the storage processor helps preserve interactivity as context grows.
With DOCA Memos, CMX can manage and share KV cache across AI compute and CMX data nodes, making cache reuse an active part of the data path. Reusing KV cache reduces repeated prefill, context recomputation, GPU idle time and unnecessary data movement, improving tokens per second and power efficiency for long‑context and agentic workloads.
In‑silicon protection for data, inference and agents
Agentic AI changes the security surface because agents repeatedly touch data, models, tools, context memory and inference services. Host‑resident security alone shares resources and trust boundaries with the workloads it protects and can be vulnerable if the host is compromised. BlueField moves security processing in‑silicon and outside the host workload domain, strengthening the control boundary while preserving CPU and GPU resources for AI work.
For multi‑tenant AI factories this approach enables inline protection without slowing the data path. BlueField provides a trusted infrastructure control point for tenant isolation, network policy enforcement, secure access, runtime detection, encryption and telemetry across compute, storage and inference infrastructure. DOCA extends these capabilities through programmable services for zero‑trust data access, inference protection, agent behavior visibility and network‑level isolation.
Conclusion
Agentic AI makes infrastructure an integral part of the inference pipeline, requiring AI factories to move, protect, retrieve and reuse data without impeding GPUs, CPUs or context memory. BlueField‑4, Vera BlueField‑4 STX and DOCA together provide accelerated infrastructure processing and a programmable software layer to deploy those services across the data center, improving interactivity, isolation, GPU utilization and token cost and energy efficiency.
Getting started
NVIDIA recommends downloading DOCA, installing the latest release and building the samples to begin programming BlueField for accelerated networking, storage, security, telemetry and lifecycle management.



