Go back

NVIDIA H200 Price Comparison (September 2026)

The NVIDIA H200 targets large-model inference and memory-heavy training with 141GB of HBM3e memory. But this extra capacity carries a steep premium.

This guide covers the NVIDIA H200 pricing landscape for September 2026, including on-demand rates across hyperscalers like AWS and Azure and specialist GPU clouds like Hyperbolic and Runpod. It also breaks down the full H200 spec sheet, SXM vs NVL variants, and comparisons to other relevant GPUs.

Key Takeaways

  • H200 premiums remain steep. The cheapest hyperscaler H200 hour costs nearly 3 times more than Thunder Compute's H100.
  • Specialist clouds narrow the gap. Vast.ai, Hyperbolic, Crusoe Cloud, and Runpod all sit in the $3.96-$4.59 range.
  • Choose H200 only to run massive models that overflow 80GB VRAM or for long-context inference. For development, fine-tuning, and most training, H100 80GB wins on ROI.
  • Thunder Compute roadmap. We don't offer H200 nodes yet. You can launch an H100 80GB at $3.20/hr with one-click VS Code, per-minute billing, persistent volumes, and live hardware swaps.

The pricing information in this guide is reviewed weekly.

Provider GPU / Instance On-Demand $/GPU-hr* Notes
Vast.ai H200 $3.96 Median of 6 verified US/CA hosts, priced as compute plus 100 GB storage; bandwidth billed separately.
Hyperbolic H200 $3.99
Crusoe Cloud H200 $4.29
DigitalOcean H200 $4.47
Nebius H200 $4.50
Modal H200 $4.54
Runpod H200 $4.59
CoreWeave 8 x H200 $6.31 Normalized from an 8-GPU node.
AWS p5en.48xlarge $7.91 Normalized from an 8-GPU node. Cheapest US region (us-east-1); ranges $7.91-$9.89/GPU-hr across 4 US options.
Oracle Cloud BM.GPU.H200.8 $10.00
Azure Standard_ND96isr_H200_v5 $10.60 Normalized from an 8-GPU node. Cheapest US region (eastus2); ranges $10.60-$13.78/GPU-hr across 7 US options.
Google Cloud a3-ultragpu-8g $10.85 On-demand-equivalent rate for reservation, Spot, or Flex-start capacity. Normalized from an 8-GPU node.

Methodology: Why You Can Trust These Numbers

  • Public rates only. The table uses on-demand rates where available. Google Cloud H200 is shown as an on-demand-equivalent rate for reservation, Spot, or Flex-start capacity.
  • Same silicon. Every row is a 141GB NVIDIA H200 (SXM or NVL/PCIe).
  • Public price lists only. Figures come straight from each provider's pricing page or API in September 2026.
  • US regions, USD.

H200 Cost Benchmark

Provider On-Demand $/GPU-hr 10-Hour Cost
Vast.ai $3.96 $39.56
Hyperbolic $3.99 $39.90
Crusoe Cloud $4.29 $42.90
DigitalOcean $4.47 $44.70
Nebius $4.50 $45.00
Modal $4.54 $45.40
Runpod $4.59 $45.90
CoreWeave $6.31 $63.05
AWS $7.91 $79.12
Oracle Cloud $10.00 $100.00
Azure $10.60 $106.00
Google Cloud $10.85 $108.45

Bottom line: an hour on Thunder Compute's A100 costs less than 15 minutes for an H200 for any provider, and buys roughly 8x more runtime than hyperscalers.

H200 Hardware Price: Buy vs Rent

After comparing cloud rates, it helps to understand what an NVIDIA H200 actually costs. The company doesn't publish list prices for data-center GPUs. The figures below come from OEM reseller quotes and secondary market listings tracked through 2026.

Configuration Price Range Notes
Single H200 (SXM), effective street price $31,000–$40,000 Through OEM resellers.
8-GPU HGX H200 system $350,000–$450,0001 Integrated platform with NVLink fabric, NICs, and CPUs.

At $31,000–$40,000 per GPU, a single H200 rented at $4/hr on a specialist cloud would take roughly 8,000–10,000 GPU-hours to match in equivalent compute cost. At 8hrs/day, that's around 3 years.

Cloud rental is the right choice for most teams: no capital tied up, no cooling or power costs, no depreciation, and the flexibility to switch hardware to match workload requirements.

H200 Specifications

The H200 and H100 GPUs use the exact same die (GH100). However, the H200 has increased memory capacity by 76% and bandwidth by 43% due to the HBM3 memory subsystem from the H100 SXM5 being replaced with HBM3e memory.

Spec H200 SXM H200 NVL
Architecture Hopper (GH100 die)
CUDA Cores 16,896
Tensor Cores 528 (4th Gen, FP8 Transformer Engine)
VRAM 141 GB HBM3e
Memory Bandwidth 4.8 TB/s
FP8 Throughput 3,958 TFLOPS1
FP16 / BF16 Throughput 1,979 TFLOPS1
TF32 Throughput 989 TFLOPS1
FP64 Throughput 34 TFLOPS
NVLink Bandwidth 900 GB/s (NVLink 4.0, point-to-point) 900 GB/s (2–4 way bridge)
Form Factor SXM5 (HGX board) PCIe add-in card
TDP 700 W 600 W
MIG Support Up to 7 instances @ 18GB each
Confidential Computing Supported

The CUDA core count, Tensor Core generation, and FP8/FP16 compute ceiling are identical to the H100 SXM5. For any compute-bound workload that fits in 80GB, there is no meaningful performance difference between the two. The H200's advantage materializes when memory is the bottleneck; most LLM inference, long-context processing, and large-batch serving.

H200 SXM vs H200 NVL

Both variants carry 141GB HBM3e at 4.8 TB/s, but form factor and interconnect differ in ways that affect deployment planning.

H200 SXM H200 NVL
Form factor SXM5 socket on HGX server board PCIe add-in card
NVLink Full NVLink 4.0, 900 GB/s per GPU 2-way or 4-way bridges, 900 GB/s per GPU
Multi-GPU scaling Best: full NVSwitch fabric Good for 2–4x GPU configs
Multi-GPU throughput delta Baseline 18% lower
TDP 700 W 600 W
Cooling requirement Liquid preferred Air-cooling compatible
Typical cloud availability AWS p5en, Oracle BM.GPU.H200.8, CoreWeave Vast.ai, Runpod NVL tier, Hyperbolic

Choose SXM for distributed training and large-scale inference requiring all-reduce across 8 GPUs at full NVLink bandwidth. The full NVSwitch fabric lets all GPUs in a node communicate directly at 900 GB/s.

Choose NVL for single-GPU or small-cluster H200 access in a standard PCIe server, or when air-cooling is a constraint.

H200 vs H100: Performance

The H200 and H100 share the same compute die; performance differences are a result of memory bandwidth and capacity.

Why LLM inference is memory-bandwidth-bound

During the decode phase (autoregressive token generation), the GPU reads the full model weight tensor from VRAM for each new token. For a 70B parameter model in FP16, that is roughly 140GB of data movement per token. Since this is memory-bound, both latency and throughput depend on how fast VRAM can deliver those weights, and the H200's 43% more bandwidth is a clear advantage.

Benchmark comparison (MLPerf Inference v4.0, Llama 2 70B)

Scenario H100 SXM H200 SXM Improvement
Offline (tokens/sec) 22,290 31,712 +42%
Server (tokens/sec) 21,504 29,526 +37%

When H200 justifies the premium

The H200 earns its cost in three scenarios:

  • A model exceeds 80GB VRAM and would otherwise require multi-GPU tensor parallelism on H100.
  • When serving with long context windows (128K+ tokens) where KV cache consumes 30–50GB per request on a 70B FP8 model, exhausting the H100's headroom after weights.
  • For large-batch production inference where the 37–42% throughput gain reduces GPU count and total cost per token.

When to stick with the H100

For development, fine-tuning, and any inference workload that fits inside 80GB, the H100 can be the more practical choice. The compute is comparable, and Thunder Compute's $3.20/hr H100 80GB availability avoids the H200 premium when a workload does not need 141GB.

H200 vs B200

The B200 is NVIDIA's Blackwell successor to the H200. It is not a memory upgrade within the same architecture but rather a full generational rebuild: a new die, native FP4 support, and more memory than the H200.

Feature H200 SXM B200 SXM
Architecture Hopper (GH100) Blackwell (dual GB100 die)
Die GH100 2x GB100
VRAM 141 GB HBM3e 180 GB HBM3e
Memory Bandwidth 4.8 TB/s 8 TB/s
FP8 Performance 3,958 TFLOPS 9,000 TFLOPS
FP4 Performance Not supported 18,000 TFLOPS
TDP 700 W (air-cooled viable) ~1,000 W (liquid cooling required)
NVLink NVLink 4.0, 900 GB/s NVLink 5.0, 1.8 TB/s
Cloud price (on-demand) $3.96-$10.85/GPU-hr $6.03-$16.11/GPU-hr
Software maturity Mature CUDA/Hopper stack FP4 ecosystem still maturing

When to choose H200

Choose the H200 when your model fits in 141GB. Llama 3 70B in FP16 is a great example.

The H200 also wins when you need a proven CUDA/Hopper software stack with no migration overhead. This includes most cloud frameworks, inference servers, and fine-tuning tooling are fully optimized for Hopper. For teams already running H100 workloads, the H200 is a drop-in upgrade with no kernel retuning.

When to choose B200

The B200 delivers roughly 2.7–3.5× lower cost-per-token than H200 for large models depending on workload, per SemiAnalysis InferenceX benchmarks, at optimized FP4 serving. A 405B parameter model fits on 3 B200s instead of 4 H200s, reducing NVLink overhead by one GPU's worth.

What LLMs Fit on H200 GPUs?

With 141GB of VRAM per card, an H200 runs mid-sized models on a single GPU, while the largest models need multi-GPU nodes or multi-node clusters. The estimates below are published weight sizes only; add 10–30% for KV cache, activations, and framework buffers at a moderate context window to gauge the real deployment footprint.

Model Parameters Precision VRAM (weights only)1 H200 setup
Llama 3 70B 70B FP8 ~70 GB 1× H200, large KV headroom
Llama 3 70B 70B FP16 ~140 GB 2× H200 node
Qwen 2.5 72B 72B FP8 ~72 GB 1× H200, with headroom
Qwen 2.5 72B 72B FP16 ~144 GB 2× H200 node
Llama 4 Scout 109B MoE Q4_K_M ~55 GB 1× H200, with headroom
Llama 4 Maverick 400B MoE2 Q4_K_M ~200 GB 2× H200 node
Llama 4 Maverick 400B MoE2 FP16 ~800 GB 8× H200 node, limited context headroom
Kimi K2.7 1T MoE2 INT4 ~594 GB 8× H200 node
Kimi K2.7 1T MoE2 FP8 ~1,000 GB 16× H200 node
GLM-5.2 744B MoE2 INT4 ~372 GB 4× H200 node
GLM-5.2 744B MoE FP8 ~744 GB 8× H200 node
DeepSeek R1 671B FP8 ~670 GB 8× H200 node

A model at FP16 requires approximately 2 bytes per parameter. The H200's usable capacity after overhead fits roughly 60–70B parameters in FP16, or approximately 130–140B in FP8. At 128K context length on a 70B FP8 model, KV cache consumes 30–50GB depending on batch configuration making the H200's headroom over the H100 becomes operationally decisive.

An 8-GPU H200 node provides 1,128GB of pooled VRAM (8 × 141GB), fitting Llama 4 Maverick at FP8 and DeepSeek R1 at FP8.

Explore the full Cloud GPU market landscape in AI GPU rental market trends analysis.

Last Thoughts on NVIDIA H200 Pricing

The H200 is the right GPU when your workload genuinely needs more than 80GB of VRAM. For everything else, H100 availability may be the more practical constraint.

With specialist cloud prices now as low as $3.96/hr and Blackwell supply continuing to grow, H200 rates are likely to keep softening through the rest of 2026.

FAQ

What is the cheapest NVIDIA H200 GPU cloud provider?

Vast.ai lists the lowest tracked H200 marketplace median at $3.96/GPU-hr as of September 2026. Among hyperscalers, AWS starts at $7.91/GPU-hr on the p5en.48xlarge.

What AI models can fit on a single H200?

A single H200 (141GB) fits Llama 3 70B in FP16 (short context) and Qwen 3 72B in FP8. An 8-GPU node (1,128GB) fits Llama 4 Maverick at FP8 and DeepSeek R1 at FP8.

What is the difference between the H200 SXM and H200 NVL?

Both carry 141GB HBM3e at 4.8 TB/s. SXM runs at 700W with full NVLink 4.0 at 900GB/s per GPU. NVL is PCIe at 600W with 2-4 way bridges and roughly 18% lower multi-GPU throughput.

What is the AWS H200 price?

AWS prices the H200 at $7.91/GPU-hr on the p5en.48xlarge ($63.30/hr for the full 8-GPU node). Egress fees of $0.09/GB apply after the first 100GB/month.

What is Azure's H200 pricing?

Azure charges approximately $10.60/GPU-hr on ND96isr H200 v5 instances, normalized from the 8-GPU node price, plus egress and Azure ML surcharges.