The NVIDIA H200 targets large-model inference and memory-heavy training with 141GB of HBM3e memory. But this extra capacity carries a steep premium.
This guide covers the NVIDIA H200 pricing landscape for September 2026, including on-demand rates across hyperscalers like AWS and Azure and specialist GPU clouds like Hyperbolic and Runpod. It also breaks down the full H200 spec sheet, SXM vs NVL variants, and comparisons to other relevant GPUs.
Key Takeaways
- H200 premiums remain steep. The cheapest hyperscaler H200 hour costs nearly 3 times more than Thunder Compute's H100.
- Specialist clouds narrow the gap. Vast.ai, Hyperbolic, Crusoe Cloud, and Runpod all sit in the $3.96-$4.59 range.
- Choose H200 only to run massive models that overflow 80GB VRAM or for long-context inference. For development, fine-tuning, and most training, H100 80GB wins on ROI.
- Thunder Compute roadmap. We don't offer H200 nodes yet. You can launch an H100 80GB at $3.20/hr with one-click VS Code, per-minute billing, persistent volumes, and live hardware swaps.
The pricing information in this guide is reviewed weekly.
| Provider | GPU / Instance | On-Demand $/GPU-hr* | Notes |
|---|---|---|---|
| Vast.ai | H200 | $3.96 | Median of 6 verified US/CA hosts, priced as compute plus 100 GB storage; bandwidth billed separately. |
| Hyperbolic | H200 | $3.99 | |
| Crusoe Cloud | H200 | $4.29 | |
| DigitalOcean | H200 | $4.47 | |
| Nebius | H200 | $4.50 | |
| Modal | H200 | $4.54 | |
| Runpod | H200 | $4.59 | |
| CoreWeave | 8 x H200 | $6.31 | Normalized from an 8-GPU node. |
| AWS | p5en.48xlarge | $7.91 | Normalized from an 8-GPU node. Cheapest US region (us-east-1); ranges $7.91-$9.89/GPU-hr across 4 US options. |
| Oracle Cloud | BM.GPU.H200.8 | $10.00 | |
| Azure | Standard_ND96isr_H200_v5 | $10.60 | Normalized from an 8-GPU node. Cheapest US region (eastus2); ranges $10.60-$13.78/GPU-hr across 7 US options. |
| Google Cloud | a3-ultragpu-8g | $10.85 | On-demand-equivalent rate for reservation, Spot, or Flex-start capacity. Normalized from an 8-GPU node. |
| Last reviewed on September 18, 2026. | |||
Methodology: Why You Can Trust These Numbers
- Public rates only. The table uses on-demand rates where available. Google Cloud H200 is shown as an on-demand-equivalent rate for reservation, Spot, or Flex-start capacity.
- Same silicon. Every row is a 141GB NVIDIA H200 (SXM or NVL/PCIe).
- Public price lists only. Figures come straight from each provider's pricing page or API in September 2026.
- US regions, USD.
H200 Cost Benchmark
| Provider | On-Demand $/GPU-hr | 10-Hour Cost |
|---|---|---|
| Vast.ai | $3.96 | $39.56 |
| Hyperbolic | $3.99 | $39.90 |
| Crusoe Cloud | $4.29 | $42.90 |
| DigitalOcean | $4.47 | $44.70 |
| Nebius | $4.50 | $45.00 |
| Modal | $4.54 | $45.40 |
| Runpod | $4.59 | $45.90 |
| CoreWeave | $6.31 | $63.05 |
| AWS | $7.91 | $79.12 |
| Oracle Cloud | $10.00 | $100.00 |
| Azure | $10.60 | $106.00 |
| Google Cloud | $10.85 | $108.45 |
Bottom line: an hour on Thunder Compute's A100 costs less than 15 minutes for an H200 for any provider, and buys roughly 8x more runtime than hyperscalers.
H200 Hardware Price: Buy vs Rent
After comparing cloud rates, it helps to understand what an NVIDIA H200 actually costs. The company doesn't publish list prices for data-center GPUs. The figures below come from OEM reseller quotes and secondary market listings tracked through 2026.
| Configuration | Price Range | Notes |
|---|---|---|
| Single H200 (SXM), effective street price | $31,000–$40,000 | Through OEM resellers. |
| 8-GPU HGX H200 system | $350,000–$450,0001 | Integrated platform with NVLink fabric, NICs, and CPUs. |
| 1 An integrated 8-GPU system is roughly eight GPUs plus $120,000–$180,000 of CPU, networking, chassis, and integration. NVIDIA does not publish list prices; ranges reflect reseller and system-implied estimates. | ||
At $31,000–$40,000 per GPU, a single H200 rented at $4/hr on a specialist cloud would take roughly 8,000–10,000 GPU-hours to match in equivalent compute cost. At 8hrs/day, that's around 3 years.
Cloud rental is the right choice for most teams: no capital tied up, no cooling or power costs, no depreciation, and the flexibility to switch hardware to match workload requirements.
H200 Specifications
The H200 and H100 GPUs use the exact same die (GH100). However, the H200 has increased memory capacity by 76% and bandwidth by 43% due to the HBM3 memory subsystem from the H100 SXM5 being replaced with HBM3e memory.
| Spec | H200 SXM | H200 NVL |
|---|---|---|
| Architecture | Hopper (GH100 die) | |
| CUDA Cores | 16,896 | |
| Tensor Cores | 528 (4th Gen, FP8 Transformer Engine) | |
| VRAM | 141 GB HBM3e | |
| Memory Bandwidth | 4.8 TB/s | |
| FP8 Throughput | 3,958 TFLOPS1 | |
| FP16 / BF16 Throughput | 1,979 TFLOPS1 | |
| TF32 Throughput | 989 TFLOPS1 | |
| FP64 Throughput | 34 TFLOPS | |
| NVLink Bandwidth | 900 GB/s (NVLink 4.0, point-to-point) | 900 GB/s (2–4 way bridge) |
| Form Factor | SXM5 (HGX board) | PCIe add-in card |
| TDP | 700 W | 600 W |
| MIG Support | Up to 7 instances @ 18GB each | |
| Confidential Computing | Supported | |
| Source: NVIDIA H200 Tensor Core GPU Datasheet. 1 With sparsity | ||
The CUDA core count, Tensor Core generation, and FP8/FP16 compute ceiling are identical to the H100 SXM5. For any compute-bound workload that fits in 80GB, there is no meaningful performance difference between the two. The H200's advantage materializes when memory is the bottleneck; most LLM inference, long-context processing, and large-batch serving.
H200 SXM vs H200 NVL
Both variants carry 141GB HBM3e at 4.8 TB/s, but form factor and interconnect differ in ways that affect deployment planning.
| H200 SXM | H200 NVL | |
|---|---|---|
| Form factor | SXM5 socket on HGX server board | PCIe add-in card |
| NVLink | Full NVLink 4.0, 900 GB/s per GPU | 2-way or 4-way bridges, 900 GB/s per GPU |
| Multi-GPU scaling | Best: full NVSwitch fabric | Good for 2–4x GPU configs |
| Multi-GPU throughput delta | Baseline | 18% lower |
| TDP | 700 W | 600 W |
| Cooling requirement | Liquid preferred | Air-cooling compatible |
| Typical cloud availability | AWS p5en, Oracle BM.GPU.H200.8, CoreWeave | Vast.ai, Runpod NVL tier, Hyperbolic |
Choose SXM for distributed training and large-scale inference requiring all-reduce across 8 GPUs at full NVLink bandwidth. The full NVSwitch fabric lets all GPUs in a node communicate directly at 900 GB/s.
Choose NVL for single-GPU or small-cluster H200 access in a standard PCIe server, or when air-cooling is a constraint.
H200 vs H100: Performance
The H200 and H100 share the same compute die; performance differences are a result of memory bandwidth and capacity.
Why LLM inference is memory-bandwidth-bound
During the decode phase (autoregressive token generation), the GPU reads the full model weight tensor from VRAM for each new token. For a 70B parameter model in FP16, that is roughly 140GB of data movement per token. Since this is memory-bound, both latency and throughput depend on how fast VRAM can deliver those weights, and the H200's 43% more bandwidth is a clear advantage.
Benchmark comparison (MLPerf Inference v4.0, Llama 2 70B)
| Scenario | H100 SXM | H200 SXM | Improvement |
|---|---|---|---|
| Offline (tokens/sec) | 22,290 | 31,712 | +42% |
| Server (tokens/sec) | 21,504 | 29,526 | +37% |
| Source: MLCommons MLPerf Inference v4.0, March 2024. Llama 2 70B benchmark, 8-GPU data center configuration. | |||
When H200 justifies the premium
The H200 earns its cost in three scenarios:
- A model exceeds 80GB VRAM and would otherwise require multi-GPU tensor parallelism on H100.
- When serving with long context windows (128K+ tokens) where KV cache consumes 30–50GB per request on a 70B FP8 model, exhausting the H100's headroom after weights.
- For large-batch production inference where the 37–42% throughput gain reduces GPU count and total cost per token.
When to stick with the H100
For development, fine-tuning, and any inference workload that fits inside 80GB, the H100 can be the more practical choice. The compute is comparable, and Thunder Compute's $3.20/hr H100 80GB availability avoids the H200 premium when a workload does not need 141GB.
H200 vs B200
The B200 is NVIDIA's Blackwell successor to the H200. It is not a memory upgrade within the same architecture but rather a full generational rebuild: a new die, native FP4 support, and more memory than the H200.
| Feature | H200 SXM | B200 SXM |
|---|---|---|
| Architecture | Hopper (GH100) | Blackwell (dual GB100 die) |
| Die | GH100 | 2x GB100 |
| VRAM | 141 GB HBM3e | 180 GB HBM3e |
| Memory Bandwidth | 4.8 TB/s | 8 TB/s |
| FP8 Performance | 3,958 TFLOPS | 9,000 TFLOPS |
| FP4 Performance | Not supported | 18,000 TFLOPS |
| TDP | 700 W (air-cooled viable) | ~1,000 W (liquid cooling required) |
| NVLink | NVLink 4.0, 900 GB/s | NVLink 5.0, 1.8 TB/s |
| Cloud price (on-demand) | $3.96-$10.85/GPU-hr | $6.03-$16.11/GPU-hr |
| Software maturity | Mature CUDA/Hopper stack | FP4 ecosystem still maturing |
When to choose H200
Choose the H200 when your model fits in 141GB. Llama 3 70B in FP16 is a great example.
The H200 also wins when you need a proven CUDA/Hopper software stack with no migration overhead. This includes most cloud frameworks, inference servers, and fine-tuning tooling are fully optimized for Hopper. For teams already running H100 workloads, the H200 is a drop-in upgrade with no kernel retuning.
When to choose B200
The B200 delivers roughly 2.7–3.5× lower cost-per-token than H200 for large models depending on workload, per SemiAnalysis InferenceX benchmarks, at optimized FP4 serving. A 405B parameter model fits on 3 B200s instead of 4 H200s, reducing NVLink overhead by one GPU's worth.
What LLMs Fit on H200 GPUs?
With 141GB of VRAM per card, an H200 runs mid-sized models on a single GPU, while the largest models need multi-GPU nodes or multi-node clusters. The estimates below are published weight sizes only; add 10–30% for KV cache, activations, and framework buffers at a moderate context window to gauge the real deployment footprint.
| Model | Parameters | Precision | VRAM (weights only)1 | H200 setup |
|---|---|---|---|---|
| Llama 3 70B | 70B | FP8 | ~70 GB | 1× H200, large KV headroom |
| Llama 3 70B | 70B | FP16 | ~140 GB | 2× H200 node |
| Qwen 2.5 72B | 72B | FP8 | ~72 GB | 1× H200, with headroom |
| Qwen 2.5 72B | 72B | FP16 | ~144 GB | 2× H200 node |
| Llama 4 Scout | 109B MoE | Q4_K_M | ~55 GB | 1× H200, with headroom |
| Llama 4 Maverick | 400B MoE2 | Q4_K_M | ~200 GB | 2× H200 node |
| Llama 4 Maverick | 400B MoE2 | FP16 | ~800 GB | 8× H200 node, limited context headroom |
| Kimi K2.7 | 1T MoE2 | INT4 | ~594 GB | 8× H200 node |
| Kimi K2.7 | 1T MoE2 | FP8 | ~1,000 GB | 16× H200 node |
| GLM-5.2 | 744B MoE2 | INT4 | ~372 GB | 4× H200 node |
| GLM-5.2 | 744B MoE | FP8 | ~744 GB | 8× H200 node |
| DeepSeek R1 | 671B | FP8 | ~670 GB | 8× H200 node |
| 1 VRAM estimates include weights only. Add 10–30% for KV cache, activations, and framework overhead at moderate context. | ||||
A model at FP16 requires approximately 2 bytes per parameter. The H200's usable capacity after overhead fits roughly 60–70B parameters in FP16, or approximately 130–140B in FP8. At 128K context length on a 70B FP8 model, KV cache consumes 30–50GB depending on batch configuration making the H200's headroom over the H100 becomes operationally decisive.
An 8-GPU H200 node provides 1,128GB of pooled VRAM (8 × 141GB), fitting Llama 4 Maverick at FP8 and DeepSeek R1 at FP8.
Explore the full Cloud GPU market landscape in AI GPU rental market trends analysis.
Last Thoughts on NVIDIA H200 Pricing
The H200 is the right GPU when your workload genuinely needs more than 80GB of VRAM. For everything else, H100 availability may be the more practical constraint.
With specialist cloud prices now as low as $3.96/hr and Blackwell supply continuing to grow, H200 rates are likely to keep softening through the rest of 2026.
FAQ
What is the cheapest NVIDIA H200 GPU cloud provider?
Vast.ai lists the lowest tracked H200 marketplace median at $3.96/GPU-hr as of September 2026. Among hyperscalers, AWS starts at $7.91/GPU-hr on the p5en.48xlarge.
What AI models can fit on a single H200?
A single H200 (141GB) fits Llama 3 70B in FP16 (short context) and Qwen 3 72B in FP8. An 8-GPU node (1,128GB) fits Llama 4 Maverick at FP8 and DeepSeek R1 at FP8.
What is the difference between the H200 SXM and H200 NVL?
Both carry 141GB HBM3e at 4.8 TB/s. SXM runs at 700W with full NVLink 4.0 at 900GB/s per GPU. NVL is PCIe at 600W with 2-4 way bridges and roughly 18% lower multi-GPU throughput.
What is the AWS H200 price?
AWS prices the H200 at $7.91/GPU-hr on the p5en.48xlarge ($63.30/hr for the full 8-GPU node). Egress fees of $0.09/GB apply after the first 100GB/month.
What is Azure's H200 pricing?
Azure charges approximately $10.60/GPU-hr on ND96isr H200 v5 instances, normalized from the 8-GPU node price, plus egress and Azure ML surcharges.