Demand for AI compute continues to accelerate: workloads are larger, models more complex, and operators face increasing pressure to deploy infrastructure quickly. Data center–scale systems—often described as “AI factories,” which continuously convert data and energy into intelligence—are being built to meet this demand.
This AI factory approach has changed system design and operation. Peak accelerator FLOPS alone are insufficient. Today’s workloads (including trillion-plus-parameter models, mixture-of-experts (MoE) architectures, long-context reasoning, and disaggregated serving) require many accelerators operating together as a single unit of compute. That need drives the requirement for high-bandwidth, low-latency GPU-to-GPU communication, fast in-network compute for collectives, and software-aware scheduling. With so many components, resiliency must be designed into the entire data center, and the supply chain must move at industry pace.
Why scale-up networking matters
Scale-out networks connect servers across a data center; scale-up networks let the GPUs within a domain behave as a single compute engine. Both are important, but they solve different problems. Scale-out fabrics (for example, NVIDIA Quantum InfiniBand or NVIDIA Spectrum-X Ethernet) support very large clusters spanning thousands to hundreds of thousands of GPUs and are critical for large data-parallel training. Scale-up fabrics connect accelerators inside a domain with predictable low latency, shared high-bandwidth memory semantics, and the high all-to-all bandwidth that modern AI training and inference often require.
MoE inference illustrates the difference: expert parallelism distributes experts across GPUs and large batches maximize factory throughput, but that creates intensive all-to-all GPU communication. Tokens must be dispatched to selected experts, processed, gathered, reordered, and forwarded in parallel. If the underlying fabric has low bandwidth or high latency, communication overhead can erase the benefits of expert parallelism—this applies in training as well.
The takeaway: all-to-all bandwidth and latency are critical to AI factory performance. Purpose-built scale-up networking co-designed with the rest of the system increases ROI by improving tokens per watt, per dollar, and per square foot of factory space, shortening training runs, raising utilization, and lowering cost-per-token.
NVLink: the purpose-built scale-up fabric
NVIDIA NVLink is designed as the scale-up networking fabric for AI factories. The sixth generation NVLink interconnect with NVLink 6 Switch targets accelerated inference, training, and parallel workloads that require large, fast GPU-to-GPU communications. It supports in-network compute (SHARP-style) for offloading collective operations and includes rack-level resiliency features intended for production uptime. NVLink has been developed with extreme co-design—chips, systems, fabric, and software optimized together—and is part of NVIDIA’s annual AI infrastructure cadence.
Concrete performance numbers
On the Vera Rubin NVL72 platform, sixth-generation NVLink provides:
- 3.6 TB/s bidirectional GPU-to-GPU bandwidth per GPU in a 72-GPU scale-up domain
- 260 TB/s aggregate rack-level GPU bandwidth for a 72-GPU domain
- End-to-end GPU-to-GPU latency that is 3X lower than off-the-shelf Ethernet alternatives
- 10X higher packet rate compared to those alternatives
- 130 TFLOPS of in-network compute across the rack for accelerating reductions and other collectives
Each NVLink switch tray contains four NVLink 6 switch chips, 28.8 TB/s of tray bandwidth and 14.4 TFLOPS of FP8 in-network compute. NVIDIA’s roadmap mentions support for domain sizes up to 1,152 GPUs and co-packaged optics connectivity.
NVIDIA also reports that, in simulations on large MoE models (for example, DeepSeek-R1, Qwen 235B, and a simulated 2T-parameter LLM), NVLink can deliver up to 2.3x higher decode throughput compared with leading off-the-shelf Ethernet in a 72-accelerator scale-up domain.
Evaluating scale-up technologies requires a system view
Spec-sheet bandwidth figures (link rate, aggregate switch capacity, per-device headline bandwidth) are useful but insufficient. Evaluating a scale-up fabric for today’s AI workloads requires a factory-level perspective: delivered, full-system performance determines how many tokens can be processed per unit time, power, and facility footprint. That depends on all-to-all fabric bandwidth, end-to-end latency (from every GPU’s HBM through the fabric to every other GPU’s HBM), in-network compute for collective operations, and fully integrated software that optimizes routing, exposes collectives, balances link traffic, pipelines transfers, and supports current AI libraries and frameworks.
Delivered performance matters only if the factory is operational. Long uptime, continuous health monitoring and telemetry, and component-level serviceability while the rest of the factory runs convert delivered performance into goodput—the capacity the factory produces over its lifetime. Achieving optimal delivered performance and goodput across many workloads is difficult and benefits from years of experience, mature scale-up infrastructure, and a reliable supply chain. Unproven technology stacks or solutions not purpose-built for scale-up networking risk suboptimal factory performance, disruptive downtime, and supply unreliability.
Key metrics for scale-up networking
Three metrics are critical when evaluating a scale-up networking solution:
- Delivered performance (bandwidth, latency, end-to-end network performance, in-network compute, full-stack software)
- Factory resiliency (long uptime, telemetry, component-level maintenance support)
- Platform maturity and proven supply chain (broad deployments and realized ROI)
Resiliency and operability in NVLink 6
NVLink 6 adds management and resiliency features intended for AI factory operations, designed to make rack maintenance less disruptive and to provide better insight and isolation for faults. Key features include:
- Control plane resilience
- Support for partially populated racks
- Hot-swappable switch trays
- Software-defined routing with management controller fallback
- Dynamic traffic rerouting
- In-service software updates
- Fine-grained link telemetry for monitoring and fault attribution
These capabilities are intended so that a failed link, switch, tray, or controller does not force an entire rack offline and to provide the operational discipline matching the economic value of the infrastructure.
Software stack and co-design
NVLink’s hardware is supported by a software ecosystem that includes NVIDIA Dynamo, NVIDIA TensorRT-LLM, NVIDIA Collective Communications Library (NCCL), NIXL, and CUDA. Dynamo is an open framework for scaling generative AI and reasoning models in multi-node GPU environments (with disaggregated serving and dynamic GPU allocation); TensorRT-LLM is an open NVIDIA library for high-performance LLM inference; NIXL speeds data transfers across GPU memory, CPU memory, NVMe and remote storage; NCCL provides high-speed GPU communication with topology-aware collectives and SHARP support.
This extreme co-design enables performance improvements that siloed stack elements cannot achieve. For example, in the transition from NVIDIA Hopper to NVIDIA Blackwell (doubling NVLink bandwidth and expanding the scale-up domain from 8 to 72 GPUs, and incorporating Dynamo for disaggregated inference), NVIDIA reported a 50x improvement in MoE inference performance per watt. The Vera Rubin platform further extends gains by doubling NVLink bandwidth and in-network compute.
CPU–GPU coherence with NVLink-C2C
AI factories require more than GPU-to-GPU bandwidth: they also need high-bandwidth coherent CPU–GPU connectivity for orchestration, data movement, memory management, storage services, and agentic workloads. NVIDIA NVLink-C2C provides that path: with Vera CPUs in the Vera Rubin NVL72 platform, NVLink-C2C delivers 1.8 TB/s of coherent bandwidth between CPUs and GPUs, about seven times the bandwidth of PCIe Gen6. Vera CPUs include 88 NVIDIA custom Olympus cores, high memory bandwidth and energy-efficient operation, and the platform co-designs CPU, GPU, NVLink, HBM, system memory, networking and software to optimize performance.
NVLink Fusion for semi-custom XPU integration
Hyperscalers and AI native companies building specialized silicon face integration challenges: connecting custom chips into a scale-up networking stack, designing rack-scale architecture, and operating heterogeneous infrastructure. NVIDIA NVLink Fusion offers high-bandwidth, low-latency interconnect IP to attach custom silicon to the NVIDIA AI infrastructure platform, enabling adopters to leverage the NVLink scale-up stack and ecosystem, reduce development complexity, increase performance, and accelerate time-to-market.
Conclusion
The next-generation AI factories will be won by infrastructure that delivers the most intelligence at the lowest cost per token, on the fastest cadence and with the least deployment risk—not solely by the fastest standalone accelerator. NVLink’s sixth generation delivers lower latency (3X), higher packet rates (10X), 130 TFLOPS of in-network compute, factory resiliency features, and a mature, proven platform backed by more than a decade of investment and deployment. NVLink aims to provide the performance, operability, maturity and continuous innovation required to power production AI factories.



