The NVIDIA L40 rents from $0.79/hr and the L40S from $1.09/hr in this October 2026 comparison of on-demand cloud pricing. Both are Ada Lovelace data center GPUs with 48 GB of memory, built for AI inference, fine-tuning, and rendering.
This guide covers current L40 and L40S cloud rates across providers, the specs that separate the two cards, which models fit in 48 GB, and how they compare to the A100.
Takeaways
- The L40 rents from $0.79/hr on Thunder Compute, the lowest rate tracked.
- The L40S ranges from $1.09/hr to $3.50/hr, and most providers list it rather than the base L40.
- The two cards share 48 GB GDDR6 and 864 GB/s, but the L40S adds the Transformer Engine and 2x FP8 throughput.
- Neither supports NVLink or MIG, so they suit single-GPU inference and LoRA fine-tuning, not large distributed training.
- For memory-bound or long-context work, an A100 or H100 is the better fit despite the higher rate.
NVIDIA L40 and L40S Cloud Pricing
The pricing information in this guide is reviewed weekly.
| Provider | GPU / Instance | On-Demand $/GPU-hr* | Notes |
|---|---|---|---|
| Thunder Compute | L40 | $0.79 | |
| Runpod | L40 | $0.82 | |
| Runpod | L40S | $1.09 | |
| CoreWeave | 8 x L40 | $1.25 | Normalized from an 8-GPU node. |
| Crusoe Cloud | L40S | $1.50 | |
| Nebius | L40S | $1.55 | |
| DigitalOcean | L40S | $1.57 | |
| Vultr | L40S | $1.67 | |
| AWS | g6e.xlarge | $1.86 | Cheapest US region (us-east-1); ranges $1.86-$7.58/GPU-hr across 24 US options. |
| Modal | L40S | $1.95 | vCPUs and RAM are billed separately. |
| CoreWeave | 8 x L40S | $2.25 | Normalized from an 8-GPU node. |
| Oracle Cloud | BM.GPU.L40S-NC.4 | $3.50 | Normalized from a 4-GPU node. |
| Last reviewed on October 1, 2026. | |||
Thunder Compute lists the lowest L40 rate at $0.79/hr, cheaper than every hyperscaler option. If your workload doesn't need the L40S's Transformer Engine, that is real savings.
Methodology
- On-demand only. The table excludes reserved, spot, and committed-use discounts.
- Same silicon per row. Each row is a 48 GB L40 or L40S; the GPU column marks which.
- Public price lists. Figures come from each provider's pricing page or API in October 2026.
- US regions, USD. Node prices are normalized to a per-GPU figure where a provider only sells multi-GPU nodes.
L40 and L40S Cost Benchmark
The 10-hour column shows what a single L40 or L40S costs for a typical inference or rendering session across the providers that list it.
| Provider | GPU | On-Demand $/GPU-hr | 10-Hour Cost |
|---|---|---|---|
| Thunder Compute | L40 | $0.79 | $7.90 |
| Runpod | L40S | $1.09 | $10.90 |
| Nebius | L40S | $1.55 | $15.50 |
| AWS | L40S | $1.86 | $18.61 |
| Oracle Cloud | L40S | $3.50 | $35.00 |
L40 vs L40S: What's the Difference?
The L40 and L40S share the same AD102 die, 48 GB of GDDR6 ECC, 864 GB/s of bandwidth, 18,176 CUDA cores, and 568 fourth-generation Tensor Cores. The difference is that the L40S has a Transformer Engine and a higher power budget, which doubles its FP8 throughput to 733 TFLOPS against the L40's 362.
For FP8-quantized LLM inference and generative AI, the L40S is the card to pick. The base L40 is aimed more at professional visualization and mixed compute, and it remains an option for image generation and 7B-34B inference that is not FP8-bound.
Read nextFor the full breakdown, see our NVIDIA L40 vs L40S comparison.
L40 and L40S Specs Compared
Both L40 cards sit a tier below the A100 on memory bandwidth, but their FP8 support and 48 GB of VRAM make them cost-effective inference GPUs. The table below compares the key specs that drive workload fit.
| Feature | NVIDIA L40 | NVIDIA L40S | NVIDIA A100 80GB |
|---|---|---|---|
| Architecture | Ada Lovelace | Ada Lovelace | Ampere |
| GPU Memory | 48 GB GDDR6 | 48 GB GDDR6 | 80 GB HBM2e |
| Memory Bandwidth | 864 GB/s | 864 GB/s | 1,935 GB/s (PCIe); 2,039 GB/s (SXM) |
| FP16 Tensor (dense) | 181 TFLOPS | 362 TFLOPS | 312 TFLOPS |
| Transformer Engine (FP8) | No | Yes | No |
| NVLink | No | No | 600 GB/s |
| MIG Support | No | No | Up to 7 |
| On-Demand Cloud Price | $0.79 - $1.25/hr | $1.09 - $3.50/hr | $1.09 - $5.03/hr |
Read nextFor the full spec sheet and architecture detail, see our NVIDIA L40 specs guide.
What Models Fit on a 48 GB L40 or L40S?
48 GB comfortably runs 7B-34B LLMs and popular image models on a single card, and fine-tunes models up to about 65B with QLoRA, or about 13B with 16-bit LoRA. The estimates below use published weight sizes and assume 20-30% overhead for KV cache, activations, and framework buffers.
| Model | Type | Setup | Approx VRAM1 | Fits 48 GB? |
|---|---|---|---|---|
| Llama 3 8B | LLM | FP16 | ~16 GB | Yes |
| Qwen 2.5 32B | LLM | 4-bit | ~20 GB | Yes |
| FLUX.1 [dev] | Image | bf16, full pipeline | ~30 GB | Yes |
| Gemma 2 27B | LLM | FP8 | ~30 GB | Yes |
| Qwen 2.5 32B | LLM | FP8 | ~33 GB | Yes |
| Llama 3 70B | LLM | 4-bit | ~40 GB | Tight |
| Qwen-Image | Image | bf16 | ~45 GB | Tight |
| Gemma 2 27B | LLM | FP16 | ~54 GB | No; use 4-bit |
| 1 Approximate VRAM for a single inference at typical settings. LLM figures are weights plus 20-30% for KV cache and activations; image figures include text encoders and the VAE and scale with resolution and batch size. | ||||
L40S vs A100: Which to Rent
The L40S wins on cost for FP8 batch inference; the A100 wins on memory bandwidth and capacity. At 864 GB/s, the L40S has roughly a third of the A100 80GB's 2,039 GB/s, so the A100 pulls ahead on memory-bound and long-context inference where weights stream from VRAM every token.
For 7B-13B models served in FP8 at moderate concurrency, the L40S often delivers comparable output at a lower hourly rate, which lowers cost per token. Choose the A100 when you need 80 GB of memory, higher bandwidth, or NVLink for multi-GPU scaling.
Read nextFor current A100 rates and the full comparison, see our NVIDIA A100 pricing guide.
NVIDIA L40S on the Major Clouds
Two hyperscalers list the L40S on-demand: AWS and Oracle. Both run well above the specialist-cloud rate.
NVIDIA L40S on AWS EC2
AWS offers the L40S through the EC2 G6e instance family, starting at $1.86/GPU-hr for g6e.xlarge in us-east-1. Rates climb to $7.58/GPU-hr across larger multi-GPU sizes and pricier regions. Egress fees apply on top of the hourly rate.
| SKU | GPUs | vCPUs | RAM | Region | Price Per-GPU |
|---|---|---|---|---|---|
| g6e.xlarge | 1 x L40S | 4 | 32 GiB | us-east-1 | $1.86 |
NVIDIA L40S on Oracle Cloud
Oracle offers the L40S through the BM.GPU.L40S-NC.4 bare-metal shape, a 4-GPU node at $3.50/GPU-hr. Bare-metal access means no hypervisor overhead.
| SKU | GPUs | Hourly Price | Price Per-GPU |
|---|---|---|---|
| BM.GPU.L40S-NC.4 | 4 x L40S | $14.00 | $3.50 |
Run L40 GPUs on Thunder Compute
Thunder Compute offers the L40 at $0.79/hr on-demand, the lowest rate tracked, with per-minute billing and no minimum commitment. The VS Code and Cursor extensions connect your editor directly to a running instance, with one-click templates for image generation and local LLMs.
Last Thoughts on NVIDIA L40 and L40S Pricing
The L40 and L40S are cost-effective inference and fine-tuning GPUs, with the L40 renting from $0.79/hr and the L40S from $1.09/hr. The L40S supports FP8 transformer workloads, while the base L40 suits image generation and general inference.
For workloads that need more than 48 GB, higher memory bandwidth, or multi-GPU NVLink scaling, step up to the A100 or H100. For everything a 48 GB Ada card handles well, the L40 is among the cheapest ways to run it.
FAQ
How much does it cost to rent an NVIDIA L40?
The NVIDIA L40 rents from $0.79/hr on Thunder Compute, the lowest tracked rate. Runpod lists it at $0.82/hr and CoreWeave at $1.25/hr, normalized from an 8-GPU node.
How much does it cost to rent an NVIDIA L40S?
The NVIDIA L40S rents from $1.09/hr to $3.50/hr in October 2026, with Runpod being the cheapest and Oracle's 4-GPU bare-metal shape at the high end.
What models fit on a 48 GB L40 or L40S?
A 48 GB L40 or L40S runs LLMs up to about 13B in FP16, 34B in FP8, or 70B in 4-bit, and image models like FLUX.1. It also handles QLoRA fine-tuning up to about 65B, or 16-bit LoRA up to about 13B.
Is the L40S better than the A100 for inference?
For FP8 batch inference on 7B-13B models, the L40S is often cheaper per token. The A100's 2,039 GB/s bandwidth wins on memory-bound and long-context workloads.
Can you fine-tune LLMs on an L40 or L40S?
Yes. With QLoRA they fine-tune models up to about 65B within 48 GB; with 16-bit LoRA, up to about 13B. Full fine-tuning fits only much smaller models, since optimizer states multiply the memory needed.