Go back

Cheapest GPU Clouds: Find Affordable Compute (August 2026)

The cheapest GPU cloud is rarely the one with the lowest advertised rate. Hidden fees, billing rounding, and reliability gaps routinely flip a seemingly cheap option into an expensive one. This guide ranks the most affordable providers and shows you how to evaluate total cost before you sign up.

  • Thunder Compute offers the lowest stable on-demand rates: A100 80 GB at $1.09/hr, RTX A6000 at $0.35/hr, H100 PCIe at $2.19/hr.
  • Spot pricing cuts GPU costs by 60-90%, but instances can be reclaimed with as little as 30 seconds of notice.
  • Per-minute billing saves ~40% vs hourly for short or bursty jobs. Per-second billing exists but makes a negligible difference in practice.
  • Consumer GPUs (RTX 4090, RTX 5090, L4) start below $0.40/hr on marketplaces .
  • Hyperscalers (AWS, GCP, Azure) are 3-10x more expensive on-demand but deep ecosystem integrations.

Pricing overview across GPU cloud platforms

Provider NVIDIA RTX A6000 /GPU-hr NVIDIA A100 80 GB /GPU-hr NVIDIA L40 48 GB /GPU-hr NVIDIA H100 80 GB /GPU-hr Free Credits
Thunder Compute $0.35 $1.09 $0.79 $2.19 Free $20 for students
Vast.ai - $1.32 $0.53 $2.21 -
Hyperstack $0.50 $1.35 $1.00 $2.50 -
Runpod $0.53 $1.39 $0.99 $2.89 -
Crusoe Cloud - $2.00 - $3.90 -
AWS - $3.43 - $6.88 750hr t2.micro + startup credits
Lambda N/A $2.79 - $3.99 -
Hyperbolic - - - $2.89 -
CoreWeave - $2.70 $1.25 $6.16 -
Nebius - - - $3.85 -
Google Cloud - $5.07 - $11.06 90-day $300 credit
*Prices from each vendor's public pricing page. On-demand rates. Spot and reserved options may be lower.

How To Find The Cheapest Cloud GPU (Without Overpaying)

The advertised GPU hourly rate is the most-compared number. Two instances at the same headline price can produce very different bills once you account for bundled vs. unbundled resources, billing granularity, and reliability.

Look Past The Advertised Rate

Some platforms quote a low GPU rate and bill CPU, RAM, and storage separately. A GPU at $1.50/hr that needs $0.50/hr of CPU and storage costs more in practice than an all-in $1.80/hr instance. Price the full instance you would actually run, not just the GPU line item.

Watch For Egress And Storage Fees

Egress fees are where hyperscaler bills quietly grow. AWS, GCP, and Azure charge for outbound data egress, so moving datasets or model checkpoints off-platform adds cost that never appears in the GPU rate. Some providers also charge for attached storage volumes while instances are stopped, so an "idle" GPU is not always free.

Factor In Reliability Cost

An interrupted training job costs more than a slightly higher hourly rate. If a marketplace instance drops your run twice and you lose hours of progress, the cheaper sticker price ends up more expensive than a stable instance that completes the job once. For workloads longer than a quick experiment, stable on-demand access is usually cheaper when you count wasted compute.

Check Billing Granularity

Hourly billing charges a full hour for any partial use. Per-minute billing cuts costs by roughly 40% on spiky or short workloads. Per-second billing exists but offers neglegible savings over its minute-based counterpart.

Spot Pricing vs. On-Demand: Where The Cheap Cloud GPU Savings Actually Come From

Spot instances are unused GPU capacity sold at a steep discount, often 60-90% below on-demand rates.

**The tradeoff is interruption risk. **Spot capacity can disappear with as little as 30 seconds of notice. Workloads must checkpoint and resume to use it safely.

Pricing model Typical discount Interruption risk Best for
On-demand Baseline None Development, inference APIs, training you cannot afford to restart
Spot / interruptible ~60% to 90% off High (reclaimed on short notice) Checkpointed training, batch jobs, fault-tolerant workloads
Reserved / committed ~30% to 60% off None Steady, predictable long-running workloads

The Cheapest GPU Cloud Options: Consumer-Tier Cards

Consumer-grade GPUs are the cheapest cloud compute when your model fits in 24-32 GB of VRAM. The RTX 4090, RTX 5090, NVIDIA L4, and RTX A5000 handle single-GPU inference, image generation, LoRA fine-tuning, and local LLM experiments at a fraction of data-center GPU rates. On marketplace platforms, these cards routinely start below $0.50/hr.

GPU VRAM Typical starting rate1 Good for
NVIDIA RTX A5000 24 GB ~$0.27/hr Inference, testing, dev environments
NVIDIA L4 24 GB ~$0.39/hr Efficient inference, light fine-tuning
NVIDIA RTX 4090 24 GB ~$0.31/hr Image generation, 7B to 13B LLMs
NVIDIA RTX 5090 32 GB ~$0.36/hr Longer context, 30B-class quantized LLMs
1 Representative on-demand marketplace rates. Consumer-card pricing and availability vary by host, region, and demand.

Thunder Compute's RTX A6000 is the better floor for developers who want consumer-tier pricing with datacenter reliability. At $0.35/hr it matches marketplace consumer cards on price, but delivers 48 GB of VRAM (double the RTX 4090 or RTX 5090).

1. AWS, GCP, Azure, Oracle (the big guys)

Pros Cons Best for
- Huge product catalog
- (Kubernetes, object storage, managed AI)
- Most expensive GPU hours.
- Data-egress lock-in.
- Steep learning curve
- Enterprises already married to a hyperscaler.
- VC-funded startups burning credits

AWS, GCP, Azure, and Oracle are the most expensive GPU cloud providers on an on-demand basis, but the most feature-complete. They offer built-in Kubernetes, managed AI services, object storage, and compliance tooling that specialized GPU clouds do not match. Teams already embedded in a hyperscaler's ecosystem often find the total cost defensible once egress, tooling, and integrations are factored in.

Startups should check each provider's credit programs, which can total hundreds of thousands of dollars. For teams without an existing cloud presence who want to get started quickly, the per-GPU rate and setup complexity make hyperscalers a poor first choice.

Recommended reads:

2. Thunder Compute

Thunder Compute homepage with low-cost GPU pricing.

Pros Cons Best for
- On-demand GPUs 80% cheaper than GCP.
- One-click VSCode integration.
- "Production mode" with maximum reliability and multi-GPU configs.
- Does not support autoscaling.
- No clusters available for large-scale training
- Researchers.
- Startups.
- ML Engineers

Thunder Compute features cost-efficient rates and a simple user experience.

  • On-demand instances for startups, research, and development.
  • Dedicated A100 hosts in U.S. data-centers.
  • Billed by the minute making it great for bursty workloads.

This solution is best suited for developers looking for the best bang for their buck.

GPU Model VRAM Hourly On-Demand Rate
NVIDIA RTX A6000 48 GB $0.35/hr
NVIDIA A100 80 GB $1.09/hr
NVIDIA L40 48 GB $0.79/hr
NVIDIA H100 PCIe 80 GB $2.19/hr

Use Thunder Compute's VS Code extension to launch an A100 80 GB in one click.

3. HyperStack

Hyperstack homepage with managed cloud GPU instances.

Pros Cons Best for
- High-speed networking and low-latency interconnects.
- 1-click deployment, hibernation, and on-demand Kubernetes.
- Integrated AI Studio for training and evaluation.
- Wide range of NVIDIA GPUs (A100, H100, H200)
- Not focused on low-cost GPUs.
- Requires familiarity with AI/ML workloads.
- High demand can occasionally impact stock
- Companies building production AI/ML pipelines.
- Organizations scaling GenAI inference.
- Teams needing premium hardware with a simplified software stack

Hyperstack is a high-performance cloud GPU platform built for AI, ML, generative AI and HPC workloads.

With reservation discounts, spot VMs and hibernation, Hyperstack lowers costs without compromising performance. Its NVMe storage enables fast data access, though some GPUs may be temporarily unavailable during peak demand.

4. Lambda

Lambda homepage with managed GPU infrastructure.

Pros Cons Best for
- GPU clusters with InfiniBand.
- Colocation options for custom hardware setups.
- Higher on-demand rates (starting from $2.79/A100).
- Limited self-service regions compared to hyperscalers.
- Research organizations looking for multi-GPU clusters.

Lambda specializes in GPU clusters for large-scale AI training, offering InfiniBand-connected clusters and colocation for teams that need a mix of cloud and on-premises hardware. Its on-demand rates are higher than most alternatives in this guide. Lambda charges for persistent storage even when instances are stopped, which adds ongoing cost for projects that pause frequently.

Read more about Lambda vs Thunder Compute and Lambda Alternatives.

5. Runpod

Runpod homepage with on-demand GPU pods and pricing.

Pros Cons Best for
- Container auto-scale with sub-second cold starts.
- A100 pricing starting from $1.39/hr.
- Community-tier GPUs can be less reliable.
- Documentation skews heavily toward inference use-cases
- Deploying production inference at the lowest possible cost.

RunPod is designed for container-based deployment with sub-second cold starts and auto-scaling, positioned between Modal and bare-metal providers on price and developer experience. Community-tier GPUs are cheaper but less reliable. Users report reliability concerns at scale, which limits its viability for production inference that requires consistent uptime.

Read more about Runpod vs Thunder Compute and Runpod vs. CoreWeave.

6. Modal

Modal homepage with serverless GPU platform features.

Pros Cons Best for
- Slick Python-native serverless API.
- Zero cold-start headaches.
- Container annotations add DX tax.
- Highest $/GPU-hr on this list.
- Teams who value development speed over price.

Modal is a Python-native GPU platform with strong developer experience and zero cold-start overhead. Deploying to Modal requires annotating Python code to containerize and scale functions.

It runs on Oracle Cloud infrastructure, with support for AWS, GCP, and Azure. Its main drawback is cost: Modal tends to have higher per-GPU-hr rates than most alternatives on this list.

7. Vast.ai

Vast.ai homepage with marketplace GPU listings.

Pros Cons Best for
- Lowest median marketplace price (~$0.15/hr). - UI friction.
- Poor reliability.
- Batch rendering / one-off experiments on a shoestring budget.

Vast.ai is a peer-to-peer GPU marketplace where individual hosts list hardware. It is primarily container-based, with limited Virtual Machine support now rolling out. Prices are the lowest available for interruptible workloads, but reliability varies significantly by host.

Read more about Vast.ai vs Thunder Compute.

8. CoreWeave

CoreWeave homepage with enterprise GPU cloud services.

Pros Cons Best for
- Kubernetes-native platform with strong multi-GPU and InfiniBand options.
- Fast storage and enterprise cluster tooling.
- Higher on-demand pricing than most alternatives.
- Hourly billing.
- Requires Kubernetes knowledge to use well.
- Ops-heavy teams running large training clusters or reserved enterprise workloads.

CoreWeave is built for teams that operate at cluster scale with Kubernetes, Slurm, or custom schedulers. It offers high-performance InfiniBand networking, fast storage, and reserved GPU environments suited to large training runs.

CoreWeave's pricing is significantly higher than Thunder Compute on equivalent GPUs, and the platform assumes Kubernetes fluency. It is not the right fit for budget-conscious developers or small teams.

Read more about CoreWeave vs Thunder Compute and CoreWeave Pricing Review.

9. Crusoe Cloud

Crusoe Cloud homepage with cloud GPU infrastructure.

Pros Cons Best for
- Solid enterprise GPU inventory.
- Minute-level billing.
- Strong fit for longer-running AI infrastructure projects.
- A100 and H100 pricing is higher than Thunder Compute.
- Full-node and allocation limits can reduce flexibility.
- Better discounts usually require longer commitments.
- Teams that want enterprise GPU capacity and can plan around reservations or larger deployments.

Crusoe Cloud sits between hyperscalers and developer-first GPU clouds. It offers enterprise GPU capacity, minute-level billing, and a fit for teams that plan workloads in advance and can commit to reservations.

For smaller teams and indie developers, the value case is harder: Crusoe's A100 and H100 pricing is higher than Thunder Compute's, and its larger-allocation model reduces flexibility for teams that resize workloads frequently.

Read more about Crusoe Cloud vs Thunder Compute.

10. Nebius

Nebius homepage with enterprise AI cloud infrastructure.

Pros Cons Best for
- Strong enterprise AI positioning.
- High-end GPU infrastructure.
- Focus on large-scale training environments.
- No public RTX A6000 or A100 price in the pricing dataset.
- Higher H100 pricing than Thunder Compute.
- Enterprise-oriented setup and contracts make it less friendly for small teams.
- Companies that need large distributed training infrastructure and can handle a more enterprise procurement process.

Nebius targets organizations training and serving models at larger scale. It positions itself closer to enterprise AI cloud than lightweight GPU rental, with an emphasis on infrastructure capacity over rapid self-serve workflows. For development, fine-tuning, or small-team experimentation, simpler pay-as-you-go options are usually a better fit.

Read more about Thunder Compute vs Nebius.

11. Hyperbolic

Hyperbolic homepage with decentralized GPU marketplace listings.

Pros Cons Best for
- Marketplace-style GPU access can surface competitive H100 pricing.
- Broad decentralized supply.
- Pricing can change with marketplace dynamics.
- Reliability depends on third-party suppliers.
- Harder to budget for production workloads.
- Experimental jobs, inference experiments, and cost-sensitive users comfortable with marketplace variability.

Hyperbolic uses a marketplace-driven model rather than managed GPU infrastructure. This can produce attractive pricing on specific GPU SKUs for users who tolerate variable availability. Marketplace pricing and supplier-dependent reliability make Hyperbolic better suited for opportunistic workloads than for production operations that require stable cost planning.

Read more about Thunder Compute vs Hyperbolic.

Last Thoughts on the Cheapest GPU Cloud

Match cost, reliability, and ecosystem fit to your workload before committing to a provider. For most developers and small teams, Thunder Compute's on-demand per-minute billing is the most cost-efficient starting point. You can always scale up or move to reserved capacity later.

If you work at a startup, check out our analysis of Startup-Friendly GPU Cloud Providers for tailored recommendations.

Frequently Asked Questions

Who is the cheapest GPU cloud provider in 2026?

Thunder Compute is among the cheapest for reliable on-demand compute: A100 80 GB at $1.09/hr, H100 PCIe at $2.19/hr. Spot marketplaces like Vast.ai reach lower prices for interruptible workloads; AWS and GCP sit at the top of the range.

How much does it cost to rent an A100 GPU per hour?

On-demand A100 80 GB runs from $1.09/hr to over $5/hr across providers. Thunder Compute's A100 80 GB is $1.09/hr, roughly 80% cheaper than GCP's $5.07/hr rate.

What is the cheapest cloud GPU for development?

Thunder Compute's RTX A6000 at $0.35/hr and A100 80 GB at $1.09/hr are the lowest stable on-demand rates for development. Consumer cards like the RTX 4090 rent cheaper on marketplaces but with less predictable availability.

How much does a cloud GPU cost per hour?

On-demand rates in this comparison run $0.35-$11.06/GPU-hr. Thunder Compute's A100 80 GB is $1.09/hr. Consumer marketplace cards start below $0.40/hr; hyperscaler H100 instances exceed $6/hr.

What hidden fees should I watch for when renting cloud GPUs?

Watch for outbound data egress fees (common on AWS, GCP, and Azure), storage charges on stopped instances, and unbundled CPU and RAM costs that raise the effective hourly rate above the advertised GPU price.