NVIDIA is introducing Scale‑In, a network infrastructure model that combines the BlueField‑4 DPU, the DOCA software framework, and Spectrum‑X Ethernet to accelerate, secure and operate agentic AI factories. Scale‑In creates a host‑independent infrastructure processing domain so infrastructure services — networking, storage protocol processing, and security — run off host CPUs and can be enforced at line rate as AI compute scales.
Why Scale‑In is needed
Traditional cloud infrastructure was built for general‑purpose workloads and standard interfaces. Agentic AI factories, however, connect many users, agents, applications, enterprise data sources and storage systems to massively accelerated compute with multi‑terabit per‑server bandwidth. That mix raises demands on north‑south access networks: software‑only approaches on host CPUs are no longer sufficient for consistent access, security and operations at scale. Scale‑In aims to evolve north‑south networks into a coordinated, accelerated infrastructure domain where operators have consistent control over access, data movement, security, provisioning and observability.
Scale‑In among NVIDIA’s five infrastructure pillars
NVIDIA describes five complementary pillars that serve AI factories:
- Scale‑Up: NVIDIA NVLink for unifying GPUs as coherent accelerators.
- Scale‑Out: NVIDIA Spectrum‑X Ethernet and NVIDIA Quantum InfiniBand for server interconnect.
- Scale‑Across: NVIDIA Spectrum‑XGS Ethernet for distributed AI factories.
- Context Memory: NVIDIA CMX (built on NVIDIA STX) for pod‑level shared KV‑cache storage.
- Scale‑In: NVIDIA BlueField‑4, NVIDIA DOCA and NVIDIA Spectrum‑X Ethernet to accelerate access, security, data movement and infrastructure operations surrounding AI compute.
Scale‑In therefore complements compute and storage pillars by focusing on the access path into and out of the AI factory and providing accelerated infrastructure services outside tenant hosts.
BlueField‑4: the infrastructure processor
BlueField‑4 accelerates Scale‑In services across GPU servers, agentic CPU systems, AI factory storage systems and cloud services. Key components and roles:
- 64‑core NVIDIA Grace CPU: runs policies, provisioning, telemetry and infrastructure orchestration; NVIDIA states it provides 6x more compute than its predecessor.
- Inline acceleration engines: handle packet processing, RDMA, storage protocols, encryption, firewall rules and policy enforcement at up to 800 Gb/s, reducing load on the Grace CPU and host CPUs.
- LPDDR5X memory subsystem: supplies high memory bandwidth for service state.
- PCIe Gen6 host connection: high‑bandwidth path between host and Scale‑In processing domain.
- 800 Gb/s network interface: connects the server to the Scale‑In fabric.
Compared with BlueField‑3, BlueField‑4 offers 4x more memory bandwidth and 2x more network bandwidth, enabling more concurrent services, larger policy and telemetry datasets, and higher traffic and security throughput.
DOCA: programming and operating Scale‑In services
NVIDIA DOCA converts BlueField‑4 hardware accelerators into programmable, deployable infrastructure services:
- Production‑ready, containerized DOCA microservices can run directly on BlueField‑4.
- DOCA SDKs and libraries expose accelerated networking, security, storage and telemetry capabilities to developers.
- DOCA Flow programs hardware packet‑processing pipelines; DOCA PCC supports programmable congestion control; DOCA Telemetry exposes device and service health; DOCA Platform Framework (DPF) manages provisioning, deployment and updates.
BlueField‑4’s multiservice architecture and native service function chaining direct traffic flows through required service sequences, providing a unified software model to operate services across BlueField devices instead of managing separate server pipelines.
Spectrum‑X Ethernet: the Scale‑In fabric
NVIDIA Spectrum‑X Ethernet provides the high‑performance fabric across the Scale‑In access path, including external storage. While BlueField‑4 processes infrastructure services on each system, Spectrum‑X carries traffic between the AI factory and users, applications, data sources, services and external storage. Spectrum‑X addresses load balancing and congestion at scale, improves resource utilization, helps maintain high effective bandwidth and isolates concurrent traffic so access and storage flows achieve more predictable performance.
BlueField‑4 DPUs, DOCA microservices and Spectrum‑X networking are co‑designed with the NVIDIA Vera Rubin platform so processor performance, memory bandwidth, PCIe I/O, network bandwidth, hardware acceleration and software capabilities remain balanced across the Scale‑In path.
Representative Scale‑In use cases for agentic AI factories
The article highlights production‑oriented scenarios where Scale‑In capabilities are applied:
-
Build isolated AI factory virtual private clouds (VPCs): DOCA Host‑Based Networking (HBN) accelerates north‑south Layer‑3 routing and multi‑tenant isolation on BlueField‑4. DOCA Flow programs traffic classification and access controls; OVS‑DOCA applies policies on east‑west interfaces. BlueField Astra extends VPC policy across east‑west Scale‑Out interfaces. In NVIDIA Vera Rubin this coordinated model spans 7.2 Tb/s aggregate interface bandwidth: 800 Gb/s on the north‑south BlueField‑4 path and four 1.6 Tb/s east‑west paths per compute tray.
-
Enforce security in silicon: BlueField‑4 places enforcement in hardware outside the host OS so tenant software cannot disable or bypass controls. DOCA Argus provides runtime threat detection, DOCA Vault enforces file‑access policy, and DOCA Flow programs line‑rate network enforcement. BlueField Astra synchronizes policy, telemetry, keys and enforcement across north‑south and east‑west traffic, keeping privileged security control outside tenant hosts and preserving host CPU resources for AI workloads.
-
Accelerate storage access: BlueField‑4 accelerates NVMe‑oF, file and object protocols over RDMA and TCP, storage virtualization and data movement on dedicated infrastructure. Spectrum‑X provides the AI‑optimized Ethernet fabric along the storage path; NVIDIA reports up to 1.45x storage throughput versus off‑the‑shelf Ethernet. The result is reduced host CPU overhead and more consistent access for training, retrieval and inference workloads.
-
Run the AI factory control plane: BlueField‑4 onboards nodes, provisions network and storage resources, starts servers and loads host OS images over the network. Because the control plane is independent of tenants’ hosts, policies and resources can be established before tenant software runs. DOCA Platform Framework is a Kubernetes‑native orchestration framework for DPU discovery, provisioning, service deployment and updates across the AI factory.
-
Observe and optimize AI operations: DOCA Telemetry collects and exports network, storage and service health data to monitoring platforms, and DOCA libraries let observability vendors integrate the same signals into their tools. Infrastructure telemetry gathered via BlueField‑4 provides operators independent visibility of network and storage usage, aiding anomaly detection, bottleneck localization and incident resolution.
Conclusion
Scale‑In transforms north‑south access into a coordinated, accelerated infrastructure domain for data access, security and operations. BlueField‑4 provides host‑independent networking and storage acceleration, DOCA programs and operates infrastructure services, and Spectrum‑X Ethernet supplies high effective bandwidth and performance isolation. Together they aim to enable secure, observable and efficient AI factory platforms where scaled compute can reliably access data and infrastructure services.
Next steps
NVIDIA points readers to further resources on the BlueField platform, NVIDIA Spectrum‑X Ethernet and the NVIDIA DOCA software framework, plus DOCA downloads and a DOCA getting‑started guide for developers.



