Nvidia said on March 16, 2026 that its next-generation Vera Rubin platform had entered full production, with seven chips built as one coordinated system rather than a single GPU launch. The lineup includes the Rubin GPU, the Vera CPU, an NVLink 6 switch, a ConnectX-9 SuperNIC, a BlueField-4 DPU and a Spectrum-6 Ethernet switch, all designed together to move data between chips with less overhead than prior generations.

The company had previewed the architecture in January 2026, describing Rubin as the successor to the Blackwell generation that shipped through 2025, and followed up in July with a technical deep-dive detailing the GPU's die layout, its use of HBM4 memory, and a disaggregated approach to inference in which different phases of a model's response are handled by different pools of hardware.

Why it matters

Rubin is Nvidia's answer to a shift in what customers are actually buying compute for. Through 2023 and 2024 the dominant workload was training large models from scratch; by 2026 hyperscalers and AI labs are spending more on inference, particularly the multi-step 'agentic' inference where a model calls tools, checks its own work and iterates before answering. Nvidia has built its pitch around packaging chips and networking as a single 'AI factory' unit rather than selling GPUs as standalone parts.

For enterprise buyers and cloud providers, the practical question is whether Rubin racks deliver enough of a cost-per-token improvement to justify replacing still-new Blackwell fleets, since the capital committed to AI data centers is now measured in the hundreds of billions of dollars industry-wide.

How it works

Nvidia's technical blog describes the Rubin GPU as built to sustain throughput across many reasoning steps rather than to maximize a single forward pass, which is the profile agentic workloads need since a single user query can trigger dozens of internal model calls. The GPU uses HBM4 memory and pairs with the Vera CPU over NVLink, while NVLink 6 switches are meant to let racks of GPUs behave more like a single accelerator when handling large models, and Spectrum-6 Ethernet and BlueField-4 DPUs extend that coordination across a data center.

TechPowerUp's review of Nvidia's die annotation confirms the emphasis on HBM4 and describes 'disaggregated inference' as a core design goal, meaning the prefill and decode stages of generating a response can run on separate, purpose-tuned hardware pools instead of the same chip handling both.

Evidence

Nvidia's investor relations site confirmed on March 16, 2026 that the Vera Rubin NVL72 GPU racks and Vera CPU racks were in full production, describing the platform as 'configurable AI infrastructure optimized for every phase of AI, from pretraining, post-training and test-time scaling to agentic inference'. The newsroom release from January 5, 2026 first laid out the six-then-seven chip lineup and said the codesign was intended to slash both training time and inference token-generation cost.

The competing read

Nvidia's own materials are the primary source for Rubin's stated performance gains, and the company has an obvious interest in presenting each generation as transformative; independent, workload-specific benchmarks from customers running Rubin at scale were not yet widely available at the time of full production. Skeptics of the 'AI factory' framing note that Nvidia's marketing folds together GPU, networking and software gains into single headline multipliers that are hard to attribute to any one component.

What happens next

The near-term test is deployment pace: whether hyperscalers that have already committed capital to Blackwell-generation racks move quickly to Rubin, and whether Rubin supply is constrained by the same HBM and advanced-packaging bottlenecks that have limited every recent Nvidia generation. Nvidia's own July technical blog frames Rubin explicitly around 'agentic AI,' suggesting the company expects its next earnings cycles to be judged on inference economics rather than raw training throughput.