The NVIDIA RTX A6000 is a professional workstation GPU built on the Ampere architecture. It remains one of the most cost-efficient ways to access 48GB of ECC VRAM in a single card. Released in late 2020, it has aged into the affordable tier of cloud GPU pricing without losing relevance for AI inference, fine-tuning, and professional visualization.
This guide covers the full RTX A6000 specifications, real-world performance benchmarks, a practical LLM model size guide, and how the A6000 compares to its closest alternatives. If you already know the specs and want pricing, jump to the RTX A6000 cloud pricing comparison.
NVIDIA RTX A6000 at a Glance
The RTX A6000 is built on NVIDIA's GA102 chip, the same silicon family as the consumer RTX 3090. The A6000 uses the full die with all 84 Streaming Multiprocessors enabled; the RTX 3090 ships with two SMs disabled at 10,496 cores. It is a dual-slot, full-height card with 48GB of ECC GDDR6 memory, a 300W TDP, and professional drivers certified for ISV applications.
| Specification | Value |
|---|---|
| Architecture | NVIDIA Ampere (GA102) |
| CUDA Cores | 10,752 |
| Tensor Cores | 336 (3rd generation) |
| RT Cores | 84 (2nd generation) |
| GPU Memory | 48 GB GDDR6 ECC |
| Memory Bandwidth | 768 GB/s |
| Memory Interface | 384-bit |
| FP32 Performance | 38.7 TFLOPS |
| RT Core Performance | 75.6 TFLOPS |
| Tensor Performance | 309.7 TFLOPS (with sparsity) |
| NVLink Bandwidth | 112.5 GB/s (bidirectional) |
| System Interface | PCIe 4.0 x16 |
| Power Consumption | 300W TDP |
| Form Factor | Dual slot, full height (4.4" H x 10.5" L) |
| Display Connectors | 4x DisplayPort 1.4a |
| Encode / Decode | 1x NVENC, 2x NVDEC (+ AV1 decode) |
| Source: NVIDIA RTX A6000 Official Datasheet | |
Full NVIDIA RTX A6000 Specifications
GPU Architecture: Ampere (GA102)
The GA102 chip is fabricated on Samsung's 8nm process with 28.3B transistors. Ampere doubled the FP32 throughput per CUDA core versus Turing by enabling two FP32 data paths per core, delivering 38.7 TFLOPS of single-precision performance. That is roughly double the Quadro RTX 6000 it replaced.
The A6000 runs on standard PCIe 4.0, requires a single 8-pin power connector, and uses an active air cooler. There is no SXM baseboard or specialized power delivery required, making it straightforward to deploy in workstations, 1U/2U servers, and cloud instances.
Compute: CUDA Cores, Tensor Cores, and RT Cores
The A6000 has 10,752 CUDA cores across 84 Streaming Multiprocessors. Each SM contains 128 CUDA cores, one third-generation Tensor Core unit, and one second-generation RT Core. The Tensor Cores support TF32, BF16, FP16, INT8, and INT4, covering every common training and inference precision.
Third-generation Tensor Cores introduced TF32 precision, accelerating matrix operations by up to 5x without code changes when compared to the previous generation.
For PyTorch users on Ampere, TF32 behavior varies by version: enabled by default in PyTorch 1.7-1.11, then disabled for matmul from 1.12 onward. Enable it explicitly with torch.backends.cuda.matmul.allow_tf32 = True.
The 84 second-generation RT Cores handle ray-triangle intersection, delivering 75.6 TFLOPS of RT Core performance. They support concurrent ray tracing with shading or denoising, which suits architectural visualization, VFX, and virtual prototyping in applications like Blender, V-Ray, and NVIDIA Omniverse.
Memory: VRAM, Bandwidth, and Interface
The RTX A6000 carries 48GB of GDDR6 ECC memory on a 384-bit bus, delivering 768 GB/s of bandwidth. ECC detects and corrects single-bit errors in hardware, which matters for long training runs and simulations where a silent memory error could corrupt hours of compute.
Consumer Ampere cards use GDDR6X at higher clock speeds; the A6000 uses standard GDDR6 at 2000 MHz (16 Gbps effective) to support ECC without the power overhead of GDDR6X. For autoregressive LLM decoding, the 768 GB/s bandwidth governs tokens/second more than raw FLOPS.
Power, Form Factor, and Display Outputs
At 300W TDP, the A6000 draws less power than the A100 80GB SXM4 (400W) or RTX 4090 (450W). This makes it easier to fit into dense multi-GPU server configurations without specialized cooling infrastructure.
The card takes two PCIe slots and a single 8-pin connector. Four DisplayPort 1.4a outputs support up to four 4K displays at 120Hz or one 8K display at 60Hz. The A6000 is VR-ready and supports NVIDIA Quadro Sync II for multi-display synchronization.
NVLink: Dual-GPU Configuration and Bandwidth
Two RTX A6000 GPUs connected via NVIDIA's third-generation NVLink bridge create a single 96GB memory pool. The bridge delivers 112.5 GB/s of bidirectional bandwidth, giving both GPUs access to the full combined address space.
This configuration suits models too large for a single 48GB card at FP16, and batch-parallel inference pipelines. The A100 SXM4's NVLink runs at 600 GB/s, which is why dual A100s scale better for large-scale distributed training. Dual A6000 via NVLink is best for inference and fine-tuning of models at lower precision.
vGPU and Virtualization Support
The RTX A6000 supports NVIDIA's vGPU software stack, which partitions a single GPU into multiple virtual instances. Three tiers are available: NVIDIA Virtual PC (vPC), RTX Virtual Workstation (vWS), and Virtual Compute Server (vCS).
Ten vGPU profile sizes are supported, from 1GB up to 48GB. This suits multi-tenant environments where one physical card needs to serve multiple users with isolated GPU resources.
Video Encode and Decode: NVENC, NVDEC, and AV1
The A6000 has one NVENC encoder and two NVDEC decoders, plus a dedicated AV1 decode engine. NVENC offloads H.264 and HEVC encoding from CUDA cores, freeing compute for AI and rendering workloads in applications like DaVinci Resolve and FFmpeg.
The dual NVDEC configuration decodes two independent video streams simultaneously, useful for multi-camera video AI pipelines. AV1 decode is supported; AV1 encode is not. Workflows requiring AV1 encoding need Ada Lovelace or newer.
NVIDIA RTX A6000 Performance Benchmarks
AI Training Throughput
A single A6000 processes approximately 1,145 images/second on ResNet50 at batch size 1,024. Dual-GPU NVLink configurations reach around 2,400 images per second. On ResNet152, a single card delivers around 605 images/second, scaling to 1,128 with two GPUs.
A 20-epoch DenseNet121 run that takes 13hrs on CPU completes in 2hrs on the A6000 at batch size 64, and 1hr 15min at batch size 128. This gives a practical baseline for computer vision training timelines.
LLM Inference: Tokens/Second by Model Size
For autoregressive decoding, memory bandwidth governs tokens/second at batch size 1. The A6000 delivers approximately 102 tokens/second for Llama 2-7B at FP16 and around 40 tokens/second for Llama 2-13B at FP16. P95 latency typically stays under 100ms, within acceptable range for production inference APIs.
What Can You Run on 48GB of VRAM? (LLM Model Size Guide)
48GB is the threshold where a wide range of LLMs become runnable on a single GPU. The table below maps common models to their VRAM requirements on the A6000.
| Model | VRAM Required | Fits on A6000? | Precision |
|---|---|---|---|
| Llama 3 8B | ~16 GB | Yes | FP16 |
| Llama 3 13B | ~26 GB | Yes | FP16 |
| Mistral 7B | ~14 GB | Yes | FP16 |
| Llama 3 30B | ~60 GB | No (FP16) / Yes (4-bit) | 4-bit (~17 GB with GPTQ) |
| Stable Diffusion XL | ~8 GB | Yes, multiple concurrent | FP16 |
| Flux.1 Dev | ~24 GB | Yes | FP16 / BF16 |
| DeepSeek-R1 7B | ~14 GB | Yes | FP16 |
| DeepSeek-R1 32B | ~64 GB | No (FP16) / Yes (4-bit) | 4-bit (~20 GB with GPTQ) |
| VRAM estimates assume FP16 weights plus KV cache overhead at batch size 1. 4-bit values use GPTQ quantization. | |||
The 48GB capacity leaves headroom for KV cache at longer context lengths, which matters for multi-turn conversations and long document inputs. For 7B-13B models, the A6000 is rarely the bottleneck at typical production batch sizes.
For guidance on choosing between the A6000 and the A100, see NVIDIA RTX A6000 vs A100.
NVIDIA RTX A6000 Use Cases
AI Training and Fine-Tuning
The A6000 handles LoRA and QLoRA fine-tuning for 7B-13B models efficiently. Full-parameter fine-tuning of models up to 7B is feasible at FP16; larger models require gradient checkpointing or parameter-efficient methods. The ResNet50 throughput of 1,145 images per second per GPU gives a practical baseline for computer vision training pipelines.
LLM Inference
The A6000 is a strong single-GPU inference platform for 7B-30B class models. Its 768 GB/s bandwidth supports low-latency decoding for Llama, Mistral, Phi, and similar architectures without tensor parallelism. For 70B models at GPTQ-4bit, a model consumes approximately 40GB, leaving limited KV cache headroom at longer sequences; the A100 80GB is the better fit there.
Scientific Computing and Simulation
ECC memory protection makes the A6000 suitable for long-running simulations where data integrity is critical. Physics, CFD, and molecular dynamics workloads run via CUDA, DirectCompute, or OpenCL. The 300W PCIe form factor is also easier to deploy in dense HPC clusters than SXM-based accelerators.
NVIDIA RTX A6000 vs. Other GPUs
RTX A6000 vs. A100: When to Upgrade
| Specification | RTX A6000 | A100 80GB |
|---|---|---|
| Architecture | Ampere (GA102) | Ampere (GA100) |
| VRAM | 48 GB GDDR6 ECC | 80 GB HBM2e ECC |
| Memory Bandwidth | 768 GB/s | 2,039 GB/s |
| FP32 TFLOPS | 38.7 | 19.5 |
| NVLink Bandwidth | 112.5 GB/s (bidirectional) | 600 GB/s (bidirectional) |
| MIG Support | No | Yes |
| Power Consumption | 300W | 400W (SXM4) |
| On-Demand Pricing | from $0.35/hr | from $1.09/hr |
The A100's key advantage is memory bandwidth: 2,039 GB/s versus 768 GB/s. It also supports MIG partitioning for multi-tenant environments, and its NVLink enables distributed training at scale that dual A6000 cannot match.
Choose the A6000 when 48GB VRAM is enough and per-hour cost matters. Choose the A100 for 30B+ model training, high-concurrency 70B inference, or multi-GPU workloads where inter-GPU bandwidth is the bottleneck.
RTX A6000 vs. RTX 6000 Ada: The 2026 Decision
The RTX 6000 Ada is the direct successor to the A6000, with the same 48GB VRAM. The Ada variant adds FP8 Tensor Core support via the Transformer Engine and increases memory bandwidth to 960 GB/s. FP8 inference can significantly increase throughput over FP16 for compatible workloads.
For FP16 inference or fine-tuning at moderate batch sizes, the A6000 remains the better cost-per-result option. Teams that need maximum throughput from a single 48GB card, particularly for FP8-optimized models, should consider the Ada generation.
RTX A6000 vs. RTX 4090: Professional vs. Consumer
The RTX 4090 is a consumer Ada Lovelace card with 24GB of GDDR6X and 16,384 CUDA cores, delivering 82.6 TFLOPS of FP32. Workloads that fit in 24GB often run faster on a 4090; anything above that requires the A6000's 48GB.
The A6000 also offers ECC memory, professional ISV certifications, and NVLink support, none of which the RTX 4090 provides. For workloads that exceed 24GB VRAM or require certified professional drivers, the A6000 is the right choice.
NVIDIA RTX A6000 Cloud Pricing (July 2026)
Most major hyperscalers have dropped the RTX A6000, concentrating availability at specialist GPU providers. All prices below are on-demand rates with no reserved commitments.
| Provider | RTX A6000 ($/hr) | Notes |
|---|---|---|
| Thunder Compute | $0.35 | Lowest on-demand price. Dedicated, secure instances. |
| Vast.ai | $0.41 | Crowdsourced infrastructure. Prices vary by host. |
| RunPod | $0.49 | Secure Cloud rate. |
| Paperspace | $1.89 | Public cloud platform. |
| Prices as of July 2026. On-demand rates only. AWS, Google Cloud, and Oracle Cloud no longer list the RTX A6000. | ||
At $0.35/hr, breaking even against buying new ($4,650 on NVIDIA's marketplace) would take over 13,000 hours of continuous use, before power, cooling, and depreciation. For most teams, renting is more economical.
Getting Started with the RTX A6000 on Thunder Compute
Thunder Compute's VS Code and Cursor extensions connect your editor directly to an A6000 instance. You write and run code locally; compute happens remotely. There is no SSH setup, no driver installation, and no environment configuration beyond installing the extension.
One-click templates for Stable Diffusion, Flux, and LLM inference environments are available in the Thunder Compute console. If your workload does not match an existing template, you can bring your own Docker container or configure a custom environment from a base image.
Last Thoughts on NVIDIA RTX A6000 Specs
The RTX A6000 delivers 48GB ECC GDDR6, 10,752 CUDA cores, and 768 GB/s of memory bandwidth at a price point well below the A100 or H100. For teams that do not need FP8, distributed multi-GPU training, or more than 48GB on a single card, it is difficult to beat on a cost-per-result basis.
FAQ
How Much VRAM Does the NVIDIA RTX A6000 Have?
48GB of GDDR6 ECC memory. ECC detects and corrects single-bit memory errors in hardware, unlike consumer GPUs that omit ECC to cut costs.
What Is the RTX A6000's Memory Bandwidth?
768 GB/s over a 384-bit interface. For autoregressive LLM decoding, memory bandwidth governs tokens/second more than raw FLOPS.
How Many CUDA Cores Does the RTX A6000 Have?
10,752 CUDA cores across 84 Streaming Multiprocessors, plus 336 third-generation Tensor Cores and 84 second-generation RT Cores.
Does the RTX A6000 Support NVLink?
Yes. Two A6000s linked via third-generation NVLink share a 96GB memory pool at 112.5 GB/s bidirectional bandwidth. The A100's 600 GB/s NVLink is significantly faster for distributed training at scale.
What Is the Difference Between the RTX A6000 and the RTX 6000 Ada?
Both have 48GB VRAM, but the RTX 6000 Ada adds FP8 Tensor Core support and 960 GB/s memory bandwidth. The A6000 is the better value for FP16 inference and fine-tuning; the Ada generation suits FP8-optimized workloads.
What Is the Cheapest Way to Rent an NVIDIA RTX A6000?
Thunder Compute offers the RTX A6000 at $0.35/hr on-demand, the lowest rate for a dedicated, secure instance. See the full RTX A6000 pricing comparison for a breakdown across providers.
What LLM Models Can I Run on the RTX A6000?
7B and 13B models run comfortably at FP16. 30B models require 4-bit quantization (around 17GB with GPTQ). 70B models are marginal at 4-bit (around 40GB), leaving limited KV cache headroom at longer sequences.
How Does the RTX A6000 Compare to the RTX 4090?
The 4090 delivers higher FP32 throughput (82.6 vs 38.7 TFLOPS) but is limited to 24GB VRAM. The A6000 offers 48GB ECC memory, NVLink support, and professional ISV certifications the 4090 lacks. For workloads that fit in 24GB, the 4090 is often faster; beyond that, the A6000 is the correct choice.
Does the RTX A6000 Support AV1 Encoding?
No. The A6000 supports AV1 decode but not AV1 encode. AV1 encoding requires Ada Lovelace architecture or newer, such as the RTX 6000 Ada or L40S.