The NVIDIA Vera Rubin platform is a seven-chip, rack-scale AI factory that delivers 3,600 PFLOPS of NVFP4 inference per rack, 5x the throughput of Blackwell, at 10x lower cost per token. It entered full production in January 2026 and is now shipping to hyperscalers and neoclouds in H2 2026.
First announced at Computex 2024, Vera Rubin is not a single chip: it co-designs GPUs, CPUs, networking, storage, and security silicon into one coherent system. This guide covers every confirmed product, spec, and technology behind it.
Key takeaways:
- Rubin GPU delivers 50 PFLOPS FP4 and 22 TB/s HBM4 bandwidth — 5x and 2.8x over Blackwell B200/B300 respectively.
- The Vera Rubin NVL72 rack houses 72 Rubin GPUs and 36 Vera CPUs, totalling 3,600 PFLOPS and 20.7 TB HBM4.
- VR200 and R100 are unofficial names. NVIDIA calls it the "Rubin GPU." H300 does not exist.
- The Groq 3 LPX companion rack handles low-latency decode, delivering 35x higher throughput/megawatt than Blackwell for trillion-parameter models.
- Rubin Ultra (H2 2027) will double the die count to 4 per package, reaching 100 PFLOPS FP4 and 1 TB HBM4e at ~600 kW.
Everything We Know So Far
NVIDIA confirmed Rubin entered full production at CES 2026, ahead of the original H2 2026 target. Jensen Huang used the keynote to detail all six co-designed chips and announce H2 2026 partner availability. GTC 2026 in March finalized system architecture details and the cloud partner list.
Rubin succeeds Blackwell and targets two core demands of agentic AI:
- Long-context inference and low latency at scale
- Trillion-parameter mixture-of-experts (MoE) models
Compared to Blackwell, NVIDIA claims:
- 5x rack-level inference performance
- 10x lower inference token cost
- 4x fewer GPUs for MoE training
Cloud availability is rolling out in H2 2026 through hyperscalers (AWS, Google Cloud, Microsoft Azure) and select neoclouds (CoreWeave, Lambda, Nebius).
Read AI GPU rental market trends to understand how supply is shifting heading into H2 2026.
The NVIDIA Rubin Announced Product Line
The Rubin platform ships in several form factors, from single-GPU server nodes to full supercomputer configurations. Here is a breakdown of the three headline products.
| Specification | Vera Rubin NVL72 | Vera Rubin Superchip | Rubin GPU |
|---|---|---|---|
| Configuration | 72 Rubin GPUs + 36 Vera CPUs | 2 Rubin GPUs + 1 Vera CPU | 1 Rubin GPU |
| Inference (NVFP4) | 3,600 PFLOPS | 100 PFLOPS | 50 PFLOPS |
| Training (NVFP4) | 2,520 PFLOPS | 70 PFLOPS | 35 PFLOPS |
| GPU Memory | 20.7 TB HBM4 | 576 GB HBM4 | 288 GB HBM4 |
| Memory Bandwidth | 1,580 TB/s | 44 TB/s | 22 TB/s |
| NVLink Bandwidth | 260 TB/s | 7.2 TB/s | 3.6 TB/s |
| NVLink-C2C Bandwidth | 65 TB/s | 1.8 TB/s | - |
| CPU Cores | 3,168 Olympus cores | 88 Olympus cores | - |
| CPU Memory | 54 TB LPDDR5X | 1.5 TB LPDDR5X | - |
| Cooling | Full liquid | ||
NVIDIA Vera Rubin NVL72: The Rack-Scale AI Powerhouse
The Vera Rubin NVL72 is NVIDIA's flagship rack-scale system and the direct successor to the GB200 NVL72. It houses 72 Rubin GPUs and 36 Vera CPUs in a single liquid-cooled rack, connected by NVLink 6.

Vera Rubin Superchip: One Package to Run the AI Era
The Vera Rubin Superchip is the base compute unit of the platform. It pairs one Vera CPU and two Rubin GPUs in a single package. Applications can treat LPDDR5X and HBM4 as a unified memory pool, reducing data movement overhead.
The Rubin GPU uses TSMC's 3nm process with a dual-die design and 336B transistors, a 1.6x increase over Blackwell. The Vera CPU adds 227B transistors and 88 Arm Olympus cores; Spatial Multi-Threading brings the effective thread count to 176.
Rubin GPU: The Accelerator at the Heart of Rubin
The Rubin GPU is the inference and training accelerator at the center of every Vera Rubin system. NVIDIA's own product material uses the name "Rubin GPU" and has not published an official alphanumeric SKU; press and supply-chain coverage widely uses R100 or R200. Each Rubin GPU carries 288 GB of HBM4 at up to 22 TB/s, 2.8x the bandwidth of Blackwell's 8 TB/s.
The Rubin GPU is optimized for FP4 and FP6 workloads, long-context inference, and multi-modal generative tasks. For teams running models above 70B parameters, it is a meaningful step change from Blackwell.
VR200, R100, R200, H300: What the Names Mean
Given the lack of an official NVIDIA SKU, the Rubin GPU has several unofficial names:
- VR200 is the system-level designation for the Vera Rubin Superchip, following the GB200 naming convention from the Blackwell generation.
- R100 was early supply-chain shorthand before dual-die packaging was confirmed.
- R200 is the more accurate press shorthand for the two-die GPU package.
To date, NVIDIA's own product pages use "Rubin GPU" exclusively.
The NVIDIA Vera Rubin Platform: A Full-Stack AI Revolution
The Vera Rubin platform is a complete AI factory blueprint, not a single chip. NVIDIA describes it as a seven-chip, five-rack architecture where every component is co-designed to eliminate the bottlenecks of off-the-shelf parts.
The seven chips cover the full lifecycle of agentic AI workloads:
- Rubin GPU and Vera CPU for compute
- NVLink 6 Switch ASIC for rack-scale scale-up networking
- ConnectX-9 SuperNIC and Spectrum-6 Ethernet for scale-out connectivity
- BlueField-4 DPU for storage and infrastructure security
- Groq 3 LPU for high-throughput decode inference
A full Vera Rubin POD reaches 40 racks, 1,152 GPUs, and 60 exaflops.

Rubin GPU vs Blackwell vs Hopper: Full Spec Comparison
Memory bandwidth and FP4 compute are where the Rubin GPU pulls furthest ahead of Blackwell. Raw HBM capacity stays flat from B300 to Rubin at 288 GB per GPU; the gains are in bandwidth (2.8x) and FP4 throughput (3.3-5.6x), not capacity.
| Specification | Rubin GPU (VR200) | Blackwell B200 | Hopper H100 |
|---|---|---|---|
| Architecture | Rubin | Blackwell | Hopper |
| Process node | TSMC 3nm (N3P) | TSMC 4nm | TSMC 4nm |
| Transistors | 336B | 208B | 80B |
| GPU memory | 288 GB HBM4 | 180 GB HBM3e | 80 GB HBM3 |
| Memory bandwidth | 22 TB/s | 8 TB/s | 3.35 TB/s |
| FP4 inference | 50 PFLOPS | 9 PFLOPS | N/A |
| FP8 throughput | 17.5 PFLOPS | 4.5 PFLOPS | ~2 PFLOPS |
| NVLink bandwidth | 3.6 TB/s (NVLink 6) | 1.8 TB/s (NVLink 5) | 900 GB/s (NVLink 4) |
| TDP | ~2,300 W (Max P) | ~1,200 W | 700 W |
| Cloud availability | H2 2026 (enterprise-first) | Available now | Available now |
The Engineering Breakthroughs Inside the Rubin Architecture
NVIDIA Vera CPU: The Brain of the Rubin Platform
The Vera CPU replaces Grace (NVIDIA's previous CPU generation) and delivers 2x the data processing performance. It is built on TSMC's 3nm process with 227B transistors and 88 custom Arm Olympus cores. Spatial Multi-Threading doubles the effective thread count to 176, with each thread maintaining full single-core throughput.
The Vera CPU is purpose-built for agentic AI workloads: managing context memory, routing tokens, and keeping inference pipelines saturated. Unlike Grace, which used licensed Arm Neoverse cores, Vera uses NVIDIA-designed Olympus cores with up to 1.2 TB/s of LPDDR5X bandwidth.
NVIDIA Groq 3 LPU: Redefining Inference Speed
The Groq 3 LPU delivers deterministic, low-latency token generation (decode) in a companion LPX rack alongside the Vera Rubin NVL72. Following NVIDIA's $20B asset acquisition of Groq, the LPU was integrated into the Rubin platform as the decode-optimized engine for trillion-parameter and high-interactivity agentic models.
The LPX rack houses 256 LPUs, each carrying 500 MB of on-chip SRAM, for a rack-total of 128 GB SRAM, 40 PB/s of aggregate memory bandwidth, and 640 TB/s of scale-up bandwidth. Paired with the Vera Rubin NVL72, NVIDIA claims 35x inference throughput per megawatt for trillion-parameter models. The LPX rack ships liquid-cooled in H2 2026 and requires no CUDA code changes.
HBM4: The Memory Leap That Makes Rubin Possible
HBM4 gives each Rubin GPU 22 TB/s of memory bandwidth, a 2.8x improvement over Blackwell's HBM3e at 8 TB/s. The gain comes from doubling the interface bus width per stack to 2,048 bits, running at 10.8 GT/s per pin.
At the NVL72 rack level, total HBM4 capacity reaches 20.7 TB at 1.6 PB/s of aggregate bandwidth. This headroom allows Rubin to serve large-scale models from single-GPU memory, with the KV-cache tiering agentic applications require. SK Hynix, Samsung, and Micron all confirmed volume HBM4 production for Vera Rubin at GTC 2026.
NVLink 6: The Fastest GPU Interconnect
NVLink 6 doubles GPU-to-GPU bandwidth over Blackwell's NVLink 5. It delivers 3.6 TB/s of scale-up bandwidth per GPU, versus 1.8 TB/s on NVLink 5. At the NVL72 rack level, total fabric bandwidth reaches 260 TB/s.
The doubled bandwidth cuts gradient sync time for communication-bound training and reduces latency for tensor-parallel inference. NVLink 6 enables all 72 GPUs in the NVL72 rack to act as a single coherent compute unit.
NVIDIA GPU Roadmap: Rubin Ultra and Feynman
Rubin Ultra (referred to in press and supply-chain coverage as VR300) is NVIDIA's next planned architecture, targeting H2 2027. Jensen Huang confirmed the roadmap through 2028 at GTC 2025, and the cadence has held.
Rubin Ultra doubles the compute die count from two to four per package, raising FP4 inference to 100 PFLOPS per GPU package and moving to HBM4e with 1 TB per package. The rack-scale system is the NVL576, housing 576 GPU dies and drawing approximately 600 kW, roughly 3x the power envelope of the Vera Rubin NVL72.
A next-generation NVLink interconnect is expected to handle the expanded fabric, reported in press coverage as NVLink 7 but not yet officially named by NVIDIA.
Feynman follows in 2028. NVIDIA has disclosed only that it uses the Vera CPU architecture and succeeds Rubin Ultra; no compute or memory specs have been published.
| Generation | Product name | Target date | FP4 per GPU package | HBM per GPU package | Rack power (NVL system) |
|---|---|---|---|---|---|
| Blackwell Ultra | GB300 NVL72 | 2025 (shipping) | ~15 PFLOPS | 288 GB HBM3e | ~140 kW |
| Rubin | VR200 NVL72 | H2 2026 | 50 PFLOPS | 288 GB HBM4 | ~190-230 kW |
| Rubin Ultra | VR300 NVL576 | H2 2027 | 100 PFLOPS | 1 TB HBM4e | ~600 kW |
| Feynman | TBD | 2028 | TBD | TBD | TBD |
Last Thoughts on the NVIDIA Vera Rubin Architecture
The Vera Rubin platform is a generational reset, not an iteration: the Rubin GPU, Vera CPU, Groq 3 LPU, HBM4, and NVLink 6 are co-designed as one system. Cloud availability will be ramping through H2 2026, with broad neocloud access expected in 2027.
FAQ
What is the NVIDIA Rubin architecture?
NVIDIA Rubin is an AI factory platform combining GPUs, CPUs, networking, storage, and security silicon. It succeeds Blackwell and targets agentic AI workloads including long-context inference and trillion-parameter MoE models. It entered full production in January 2026.
What role does the NVIDIA Groq 3 LPU play in the Rubin platform?
The Groq 3 LPU handles low-latency token generation (decode) in a companion LPX rack alongside the Vera Rubin NVL72. The LPX rack houses 256 LPUs with 128 GB aggregate SRAM and 40 PB/s bandwidth, delivering 35x higher throughput/megawatt than Blackwell for trillion-parameter models.
How does HBM4 improve performance in the Rubin GPU?
HBM4 gives each Rubin GPU 22 TB/s of memory bandwidth, 2.8x faster than Blackwell's HBM3e at 8 TB/s. The gain comes from doubling the interface bus width per stack to 2,048 bits, running at 10.8 GT/s per pin.
What is VR200?
VR200 is press and supply-chain shorthand for the Vera Rubin Superchip (one Vera CPU plus two Rubin GPU dies).
How does the Rubin GPU compare to Blackwell B200 and B300?
The Rubin GPU delivers 22 TB/s memory bandwidth (2.8x B200/B300), 50 PFLOPS FP4 (5.6x B200, 3.3x B300), and NVLink 6 at 3.6 TB/s (2x B200). HBM capacity is identical to B300 at 288 GB.