The NVIDIA H200 targets large-model inference and memory-heavy training with 141GB of HBM3e memory, but that extra capacity carries a steep premium.
This guide covers the NVIDIA H200 pricing landscape for August 2026, including on-demand rates across hyperscalers like AWS and Azure and specialist GPU clouds like Hyperbolic and Runpod. It also covers the full H200 spec sheet, SXM vs NVL variants, and compares it to other relevant GPUs.
Key Takeaways
- H200 premiums remain steep. Even after AWS's June 2025 price cut, the cheapest hyperscaler H200 hour costs 3 times more than Thunder Compute's H100.
- Specialist clouds narrow the gap. Hyperbolic, Runpod, and Vast.ai all sit in the $2-5 range.
- Choose H200 only when you must. Running massive models that overflow 80GB VRAM or long-context inference. For development, fine-tuning, and most training, H100 80GB wins on ROI.
- Thunder Compute roadmap. We don't offer H200 nodes yet. You can launch an H100 80GB at $2.19/hr with one-click VS Code, per-minute billing, persistent volumes, and live hardware swaps.
| Provider | SKU / Instance | On-Demand $/GPU-hr* | Notes | |
|---|---|---|---|---|
| Hyperbolic | H200 | $3.49 | Marketplace pricing. | |
| Vast.ai | Marketplace | $4.08 | Marketplace current price per H200 GPU. | |
| Crusoe Cloud | H200 | $4.29 | Public listed hourly pricing. | |
| Runpod | H200 | $4.39 | Secure Cloud pricing from the public pricing page. | |
| Nebius | H200 | $4.50 | Public listed single-GPU equivalent. | |
| Modal | H200 | $4.54 | Public listed hourly pricing. | |
| CoreWeave | 8x H200 | $6.16 | $49.28/hr node price, normalized per GPU. | |
| AWS (p5e.48xlarge) | 8x H200 141GB | $7.91 | $63.28/hr node price, normalized per GPU. | |
| Oracle Cloud (BM.GPU.H200.8) | 8x H200 | $10.00 | Bare-metal node, $80/hr total. | |
| Azure (ND96isr H200 v5) | 8x H200 | $10.60 | Calculator price, normalized per GPU. | |
| Google Cloud (a3-ultragpu-8g) | 8x H200 | $10.60 | Calculator price, normalized per GPU. | |
Methodology – why you can trust these numbers
- On-demand only. We excluded capacity reservations longer than 14 days, reserved instances, and spot/preemptible offers.
- Same silicon. Every row is a 141GB NVIDIA H200 (SXM or NVL/PCIe).
- Public price lists only. Figures come straight from each provider's pricing page in August 2026.
- US regions, USD. Regional variation can add 5-20%; those are excluded for an apples-to-apples comparison.
H200 Cost Benchmark
| Provider | 10 hrs runtime | Effective cost |
|---|---|---|
| Thunder Compute – A100 80GB | 10 x $1.09 | $10.90 |
| Vast.ai | 10 x $4.08 | $40.80 |
| Crusoe Cloud | 10 x $4.29 | $42.90 |
| Runpod | 10 x $4.39 | $43.90 |
| CoreWeave | 10 x $6.16 | $61.60 |
| AWS | 10 x $7.91 | $79.10 |
| Azure | 10 x $10.60 | $106.00 |
Bottom line: an hour on Thunder Compute's A100 costs less than 15 minutes for an H200 for any provider, and buys roughly 8x more runtime than hyperscalers.
H200 Hardware Price: Buy vs Rent
After comparing cloud rates, it helps to understand what the hardware actually costs. NVIDIA does not publish list prices for data-center GPUs. The figures below come from OEM reseller quotes and secondary market listings tracked through 2026.
| Configuration | Price Range | Notes |
|---|---|---|
| Single H200 NVL (PCIe) | $28,000–$34,000 | Through authorized OEM resellers |
| Single H200 SXM5 | $32,000–$40,000 | SXM5 requires HGX server platform |
| 8-GPU HGX H200 server | $320,000–$420,000 | Includes NVLink fabric, NICs, integration |
| DGX H200 system | $350,000–$500,000 | Full NVIDIA-integrated system |
At $32,000–$40,000 per GPU, a single H200 rented at $4/hr on a specialist cloud would take roughly 8,000–10,000 GPU-hours to match in equivalent compute cost. At 8hrs/day, that's around 3 years.
For most teams, cloud rental is the right choice: no capital tied up, no cooling or power costs, no depreciation, and the flexibility to switch hardware as workload requirements change.
H200 Specifications
The H200 uses the same Hopper GH100 die as the H100. Every performance difference traces back to one change: the HBM3 memory subsystem from the H100 SXM5 was replaced with HBM3e memory, increasing capacity by 76% and bandwidth by 43%.
| Spec | H200 SXM | H200 NVL |
|---|---|---|
| Architecture | Hopper (GH100 die) | |
| CUDA Cores | 16,896 | |
| Tensor Cores | 528 (4th Gen, FP8 Transformer Engine) | |
| VRAM | 141 GB HBM3e | |
| Memory Bandwidth | 4.8 TB/s | |
| FP8 Throughput | 3,958 TFLOPS (with sparsity) | |
| FP16 / BF16 Throughput | 1,979 TFLOPS (with sparsity) | |
| TF32 Throughput | 989 TFLOPS (with sparsity) | |
| FP64 Throughput | 34 TFLOPS | |
| NVLink Bandwidth | 900 GB/s (NVLink 4.0, point-to-point) | 900 GB/s (2–4 way bridge) |
| Form Factor | SXM5 (HGX board) | PCIe add-in card |
| TDP | 700 W | 600 W |
| MIG Support | Up to 7 instances @ 18GB each | |
| Confidential Computing | Supported | |
| Source: NVIDIA H200 Tensor Core GPU Datasheet. | ||
The CUDA core count, Tensor Core generation, and FP8/FP16 compute ceiling are identical to the H100 SXM5. For any compute-bound workload that fits in 80GB, there is no meaningful performance difference between the two. The H200's advantage materializes when memory is the bottleneck like for most LLM inference, long-context processing, and large-batch serving fall into this category.
H200 SXM vs H200 NVL
Both variants carry 141GB HBM3e at 4.8 TB/s, but form factor and interconnect differ in ways that affect deployment planning.
| H200 SXM | H200 NVL | |
|---|---|---|
| Form factor | SXM5 socket on HGX server board | PCIe add-in card |
| NVLink | Full NVLink 4.0, 900 GB/s per GPU | 2-way or 4-way bridges, 900 GB/s per GPU |
| Multi-GPU scaling | Best: full NVSwitch fabric | Good for 2–4x GPU configs |
| Multi-GPU throughput delta | Baseline | 18% lower |
| TDP | 700 W | 600 W |
| Cooling requirement | Liquid preferred | Air-cooling compatible |
| Typical cloud availability | AWS p5e, Oracle BM.GPU.H200.8, CoreWeave | Vast.ai, Runpod NVL tier, Hyperbolic |
Choose SXM for distributed training and large-scale inference requiring all-reduce across 8 GPUs at full NVLink bandwidth. The full NVSwitch fabric lets every GPU communicate with every other at 900 GB/s without routing through a CPU. Choose NVL for single-GPU or small-cluster H200 access in a standard PCIe server, or when air-cooling is a constraint.
For most individual developers and small teams evaluating H200 access, NVL single-GPU instances from Runpod, Vast.ai, or Hyperbolic are the practical entry point.
H200 vs H100: Performance and Price
The H200 and H100 share the same compute die. The performance delta is entirely a function of memory bandwidth and capacity.
Why LLM inference is memory-bandwidth-bound
During the decode phase (autoregressive token generation), the GPU reads the full model weight tensor from VRAM for each new token. For a 70B parameter model in FP16, that is roughly 140GB of data movement per token — the Tensor Cores are nearly idle. The only thing that matters is how fast VRAM can deliver those weights, which is where the H200's 43% bandwidth advantage shows up.
Benchmark comparison (MLPerf Inference v4.0, Llama 2 70B)
| Scenario | H100 SXM | H200 SXM | Improvement |
|---|---|---|---|
| Offline (tokens/sec) | 22,290 | 31,712 | +42% |
| Server (tokens/sec) | 21,504 | 29,526 | +37% |
| Source: MLCommons MLPerf Inference v4.0, March 2024. Llama 2 70B benchmark, 8-GPU data center configuration. | |||
When H200 justifies the premium
The H200 earns its cost in three scenarios:
- A model exceeds 80GB VRAM and would otherwise require multi-GPU tensor parallelism on H100.
- When serving with long context windows (128K+ tokens) where KV cache consumes 30–50GB per request on a 70B FP8 model, exhausting the H100's headroom after weights.
- For large-batch production inference where the 37–42% throughput gain reduces GPU count and total cost per token.
When to stick with the H100
For development, fine-tuning, and any inference workload that fits inside 80GB, the H100 wins on cost-efficiency. The compute is identical. At Thunder Compute's $2.19/hr for an H100 80GB versus $3.49–$10.60/hr for H200 depending on provider, a workload that doesn't need 141GB pays a steep premium for capacity it isn't using.
H200 vs B200
The B200 is NVIDIA's Blackwell successor to the H200. It is not a memory upgrade within the same architecture but rather a full generational rebuild: a new die, native FP4 support, and nearly double the memory of the H200.
| Feature | H200 SXM | B200 SXM |
|---|---|---|
| Architecture | Hopper (GH100) | Blackwell (dual GB100 die) |
| VRAM | 141 GB HBM3e | 192 GB HBM3e |
| Memory Bandwidth | 4.8 TB/s | 8 TB/s |
| FP8 Throughput | 3,958 TFLOPS | ~9,000 TFLOPS |
| FP4 Support | No | Yes (5th-gen Tensor Cores) |
| TDP | 700 W (air-cooled viable) | ~1,000 W (liquid cooling required) |
| NVLink | NVLink 4.0, 900 GB/s | NVLink 5.0, 1.8 TB/s |
| Cloud price (on-demand) | $3.76–$10.60/GPU-hr | $6.03+/GPU-hr |
| Software maturity | Mature CUDA/Hopper stack | FP4 ecosystem still maturing |
When to choose H200
Choose the H200 when your model fits in 141GB. Llama 3 70B in FP16 is the canonical example.
The H200 also wins when you need a proven CUDA/Hopper software stack with no migration overhead. This includes most cloud frameworks, inference servers, and fine-tuning tooling are fully optimized for Hopper. For teams already running H100 workloads, the H200 is a drop-in upgrade with no kernel retuning.
When to choose B200
The B200 delivers roughly 2.7–3.5× lower cost-per-token than H200 for large models depending on workload, per SemiAnalysis InferenceX benchmarks, at optimized FP4 serving. A 405B parameter model fits on 3 B200s instead of 4 H200s, reducing NVLink overhead by one GPU's worth. The B200 is also the right choice for new deployments in 2026 with the current-generation software stack.
What LLMs Fit on an H200?
The H200's 141GB of VRAM determines which models run on a single card versus requiring multi-GPU parallelism. The estimates below use published weight sizes and assume 20–30% overhead for KV cache, activations, and framework buffers at a moderate context window.
| Model | Parameters | Precision | VRAM (weights only)1 | Single H200? |
|---|---|---|---|---|
| Llama 3 70B | 70B | FP8 | ~70 GB | ✅ Yes, with large KV cache headroom |
| Llama 3 70B | 70B | FP16 | ~140 GB | ✅ Fits (tight — short context only) |
| Qwen 3 72B | 72B | FP8 | ~72 GB | ✅ Yes, with headroom |
| Qwen 3 72B | 72B | FP16 | ~144 GB | ⚠️ Marginal — use FP8 instead |
| Llama 4 Scout | 109B MoE | Q4_K_M | ~55 GB | ✅ Yes, with headroom |
| Llama 4 Maverick | 400B MoE2 | Q4_K_M | ~200 GB | ✅ Fits on 8-GPU H200 node |
| Llama 4 Maverick | 400B MoE2 | FP16 | ~800 GB | ✅ Fits on 8-GPU H200 node with limited context headroom |
| Kimi K2.7 | 1T MoE2 | INT4 | ~594 GB | ✅ Fits on 8-GPU H200 node |
| Kimi K2.7 | 1T MoE2 | FP8 | ~1,000 GB | ❌ Requires multiple nodes |
| GLM-5.2 | 744B MoE2 | INT4 | ~372 GB | ✅ Fits on 8-GPU H200 node |
| GLM-5.2 | 744B MoE | FP8 | ~744 GB | ✅ Fits on 8-GPU H200 node |
| DeepSeek R1 | 671B | FP8 | ~670 GB | ✅ Fits on 8-GPU H200 node |
| 1 VRAM estimates include weights only. Add 20–30% for KV cache, activations, and framework overhead at moderate context. 2 MoE models must keep all expert weights in VRAM even though only a fraction activate per token. |
||||
A model at FP16 requires approximately 2 bytes per parameter. The H200's usable capacity after overhead fits roughly 60–70B parameters in FP16, or approximately 130–140B in FP8. At 128K context length on a 70B FP8 model, KV cache consumes 30–50GB depending on batch configuration making the H200's headroom over the H100 becomes operationally decisive.
An 8-GPU H200 node provides 1,128GB of pooled VRAM (8 × 141GB), fitting Llama 4 Maverick at FP8 and DeepSeek R1 at FP8.
Explore the full Cloud GPU market landscape in AI GPU rental market trends analysis.
Last Thoughts on NVIDIA H200 Pricing
The H200 is the right GPU when your workload genuinely needs more than 80GB of VRAM. For everything else, the H100 delivers better cost-efficiency. With specialist cloud prices now as low as $3.49/hr and Blackwell supply continuing to grow, H200 rates are likely to keep softening through the rest of 2026. Thunder Compute offers H100 80GB at $2.19/hr with no long-term commitment and instant access.
FAQ
How much does an NVIDIA H200 cost to buy outright?
A single H200 NVL (PCIe) costs $28,000–$34,000 through OEM resellers; the SXM5 variant runs $32,000–$40,000. A full 8-GPU HGX H200 server runs $320,000–$420,000.
What is the cheapest NVIDIA H200 GPU cloud provider?
Hyperbolic offers the lowest public on-demand H200 price at $3.49/GPU-hr as of August 2026. Among hyperscalers, AWS starts at $7.91/GPU-hr on the p5e.48xlarge.
Can I rent a single H200 GPU, or do I need a full 8-GPU node?
Hyperscalers (AWS, Azure, Oracle) require full 8-GPU nodes. Hyperbolic, Vast.ai, Runpod, Modal, and Nebius offer single-GPU H200 access for smaller workloads.
Is the H200 worth it over the H100?
Only for memory-bound workloads. The H200 delivers 37–42% higher throughput on 70B models. For workloads under 80GB, the H100 at $2.19/hr delivers the same FP8 compute at lower cost.
What AI models can fit on a single H200?
A single H200 (141GB) fits Llama 3 70B in FP16 (short context) and Qwen 3 72B in FP8. An 8-GPU node (1,128GB) fits Llama 4 Maverick at FP8 and DeepSeek R1 at FP8.
What is the difference between the H200 SXM and H200 NVL?
Both carry 141GB HBM3e at 4.8 TB/s. SXM runs at 700W with full NVLink 4.0 at 900GB/s per GPU. NVL is PCIe at 600W with 2-4 way bridges and roughly 18% lower multi-GPU throughput.
What is the AWS H200 price?
AWS prices the H200 at $7.91/GPU-hr on the p5e.48xlarge ($63.28/hr for the full 8-GPU node). Egress fees of $0.09/GB apply after the first 100GB/month.
What is Azure's H200 pricing?
Azure charges approximately $10.60/GPU-hr on ND96isr H200 v5 instances, normalized from the 8-GPU node price, plus egress and Azure ML surcharges.
Is Azure or Oracle Cloud cheaper for NVIDIA H200 GPUs?
Oracle Cloud is cheaper at $10.00/GPU-hr on BM.GPU.H200.8 bare-metal instances ($80/hr total), versus Azure at $10.60/GPU-hr on ND96isr H200 v5.
What is the price difference between Runpod and CoreWeave for H200 GPUs?
Runpod lists H200 at $4.39/GPU-hr with single-GPU access. CoreWeave lists $6.16/GPU-hr normalized from an 8-GPU node ($49.28/hr total).
What is Google Cloud's H200 pricing?
Google Cloud offers H200s at $10.60/GPU-hr the a3-ultragpu family.
When should I choose the H200 over the H100?
Choose H200 when your model exceeds 80GB VRAM or you need long-context inference with large KV caches. For most workloads under 80GB, the H100 delivers better cost-efficiency.