L
localai.computer
ModelsGPUsSystemsBuildsOpenClawMethodology

Resources

  • Methodology
  • Submit Benchmark
  • About

Browse

  • AI Models
  • GPUs
  • PC Builds
  • AI News

Guides

  • OpenClaw Guide
  • How-To Guides

Legal

  • Privacy
  • Terms
  • Contact

© 2026 localai.computer. Hardware recommendations for running AI models locally.

ℹ️We earn from qualifying purchases through affiliate links at no extra cost to you. This supports our free content and research.

  1. Home
  2. GPUs
  3. NVIDIA A6000

NVIDIA A6000

By NVIDIAReleased 2020-10Launch MSRP $4,699.00

This GPU offers reliable throughput for local AI workloads. Pair it with the right model quantization to hit your desired tokens/sec, and monitor prices below to catch the best deal.

Check Price on AmazonView Benchmarks
Specs snapshot
Key hardware metrics for AI workloads.
VRAM48GB
Cores10,752
TDP300W
ArchitectureAmpere

Quick Answer: NVIDIA A6000 has 48GB VRAM, enough for models up to roughly 120B parameters at 4-bit quantization. It draws 300W under load.

Key Takeaways
  • 48GB VRAM - runs models up to ~120B parameters
  • High-end compute for demanding workloads
  • Moderate power draw (300W) - 750W PSU typically sufficient
  • Strong price-to-VRAM value

What this means for you

With 48GB VRAM, NVIDIA A6000 can run models up to approximately 120B parameters using 4-bit quantization. That covers most popular models, including 70B-class ones at 4-bit.

Who should buy

  • Professional AI workloads requiring maximum VRAM
  • Running 100B+ parameter models with full precision

Looking to upgrade?

Consider H100 or MI300X — Maximum VRAM for enterprise workloads.

Where to Buy

Buy directly on Amazon with fast shipping and reliable customer service.

Amazon
See price on Amazon
Buy on Amazon

Prime shipping available • 30-day returns

Complete Your Build

Essential accessories to pair with NVIDIA A6000

Corsair RM750x ATX 3.1 750W
Minimum 750W recommended for RTX 40 series
~$119
View on Amazon
Corsair Vengeance RGB 32GB DDR5-6000
32GB ideal for AI workloads
~$129
View on Amazon
Noctua NF-A12x25 PWM
Quiet and efficient cooling
~$35
View on Amazon
Thermal Grizzly Kryonaut
Premium thermal paste for optimal cooling
~$15
View on Amazon

Accessories Total

Typical prices — check Amazon for current pricing

~$298
See Complete BuildsMore GPUs

💡 Not ready to buy? Try cloud GPUs first

Test NVIDIA A6000 performance in the cloud before investing in hardware. Pay by the hour with no commitment.

Vast.aifrom $0.20/hrRunPodfrom $0.30/hrLambda Labsenterprise-grade

GPU FAQs

Data-backed answers pulled from community benchmarks, manufacturer specs, and live pricing.

What throughput does an RTX A6000 deliver on 70B Q4?

Operators running dual RTX A6000/RTX 8000 cards inside oobabooga report roughly 6–7 tokens/sec on 70B IQ4 MiQu workloads—adequate for shared inference queues.

Source: Reddit – /r/LocalLLaMA (lnv0ww3)

Why do PCIe lanes limit RTX A6000 performance?

Enthusiasts caution that consumer boards seldom provide x16/x16 for two A6000s; dropping to x8/x4 starves llama.cpp workloads and erodes throughput.

Source: Reddit – /r/LocalLLaMA (mqpg0wp)

Is 48 GB still worth the workstation premium?

Even 2020-era RTX A6000 cards still list near $5,000, and the community expects scalpers to follow new workstation launches—showing how demand stays high.

Source: Reddit – /r/LocalLLaMA (movlqi2)

Should you swap to 48 GB RTX 4090 variants?

Some builders consider 48 GB 4090s, which keep full VRAM for inference but drop to 24 GB for PCIe peer-to-peer training—making the trade-off workload dependent.

Source: Reddit – /r/LocalLLaMA (mqoerg0)

What are the specs and latest prices?

RTX A6000 ships with 48 GB GDDR6 ECC and a 300 W TDP. As of Nov 2025 pricing on Amazon was around $4,899.

Source: TechPowerUp – NVIDIA RTX A6000 Specs

AI benchmarks

Showing 12 of 80 rows. Speeds are calculated estimates, not measurements — search for your model to jump straight to it.

ModelSizeQuantizationTokens/secVRAM used
Deepseek AI Deepseek Coder 1.3B Instruct1.3BQ4
~200 tok/sEstimated
1GB
Deepseek AI Deepseek R1 Distill Qwen 1.5B1.5BQ4
~200 tok/sEstimated
1GB
Deepseek AI Deepseek Ocr 2UnknownQ4
~165 tok/sEstimated
2GB
Deepseek AI Deepseek OcrUnknownQ4
~165 tok/sEstimated
2GB
Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 8bit8BQ4
~165 tok/sEstimated
4GB
Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 4bit8BQ4
~165 tok/sEstimated
4GB
Deepseek AI Deepseek R1 Distill Qwen 7B7BQ4
~165 tok/sEstimated
4GB
Nineninesix Kani Tts 2 EnUnknownQ4
~160 tok/sEstimated
1GB
Qwen Qwen3 Tts 12hz 1 7B Customvoice7BQ4
~160 tok/sEstimated
1GB
Zai Org Glm OcrUnknownQ4
~160 tok/sEstimated
1GB
Qwen Qwen3 Asr 1 7B7BQ4
~160 tok/sEstimated
2GB
Nari Labs Dia2 2B2BQ4
~160 tok/sEstimated
1GB
Deepseek AI Deepseek Coder 1.3B Instruct
Q4 · 1.3B
1GB
~200 tok/sEstimated
Deepseek AI Deepseek R1 Distill Qwen 1.5B
Q4 · 1.5B
1GB
~200 tok/sEstimated
Deepseek AI Deepseek Ocr 2
Q4 · Unknown
2GB
~165 tok/sEstimated
Deepseek AI Deepseek Ocr
Q4 · Unknown
2GB
~165 tok/sEstimated
Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 8bit
Q4 · 8B
4GB
~165 tok/sEstimated
Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 4bit
Q4 · 8B
4GB
~165 tok/sEstimated
Deepseek AI Deepseek R1 Distill Qwen 7B
Q4 · 7B
4GB
~165 tok/sEstimated
Nineninesix Kani Tts 2 En
Q4 · Unknown
1GB
~160 tok/sEstimated
Qwen Qwen3 Tts 12hz 1 7B Customvoice
Q4 · 7B
1GB
~160 tok/sEstimated
Zai Org Glm Ocr
Q4 · Unknown
1GB
~160 tok/sEstimated
Qwen Qwen3 Asr 1 7B
Q4 · 7B
2GB
~160 tok/sEstimated
Nari Labs Dia2 2B
Q4 · 2B
1GB
~160 tok/sEstimated

Model compatibility

Showing 12 of 240 rows.

ModelSizeQuantizationVerdictEstimated speedVRAM needed
01 AI Yi 1 5 34B Chat34BQ4Fits comfortably
~46 tok/sEstimated
18GB (have 48GB)
01 AI Yi 1 5 34B Chat34BQ8Fits comfortably
~32 tok/sEstimated
35GB (have 48GB)
01 AI Yi 1 5 34B Chat34BFP16Not supported
~18 tok/sEstimated
69GB (have 48GB)
AI Forever Rugpt 3.5 13B13BQ4Fits comfortably
~99 tok/sEstimated
7GB (have 48GB)
AI Forever Rugpt 3.5 13B13BQ8Fits comfortably
~70 tok/sEstimated
13GB (have 48GB)
AI Forever Rugpt 3.5 13B13BFP16Fits comfortably
~38 tok/sEstimated
26GB (have 48GB)
AI Mo Kimina Prover 72B72BQ4Fits comfortably
~26 tok/sEstimated
37GB (have 48GB)
AI Mo Kimina Prover 72B72BQ8Not supported
~19 tok/sEstimated
73GB (have 48GB)
AI Mo Kimina Prover 72B72BFP16Not supported
~10 tok/sEstimated
146GB (have 48GB)
Alibaba Nlp Gte Qwen2 1.5B Instruct1.5BQ4Fits comfortably
~160 tok/sEstimated
1GB (have 48GB)
Alibaba Nlp Gte Qwen2 1.5B Instruct1.5BQ8Fits comfortably
~110 tok/sEstimated
2GB (have 48GB)
Alibaba Nlp Gte Qwen2 1.5B Instruct1.5BFP16Fits comfortably
~60 tok/sEstimated
4GB (have 48GB)
01 AI Yi 1 5 34B ChatQ4
Size: 34B
Fits comfortably18GB required · 48GB available
~46 tok/sEstimated
01 AI Yi 1 5 34B ChatQ8
Size: 34B
Fits comfortably35GB required · 48GB available
~32 tok/sEstimated
01 AI Yi 1 5 34B ChatFP16
Size: 34B
Not supported69GB required · 48GB available
~18 tok/sEstimated
AI Forever Rugpt 3.5 13BQ4
Size: 13B
Fits comfortably7GB required · 48GB available
~99 tok/sEstimated
AI Forever Rugpt 3.5 13BQ8
Size: 13B
Fits comfortably13GB required · 48GB available
~70 tok/sEstimated
AI Forever Rugpt 3.5 13BFP16
Size: 13B
Fits comfortably26GB required · 48GB available
~38 tok/sEstimated
AI Mo Kimina Prover 72BQ4
Size: 72B
Fits comfortably37GB required · 48GB available
~26 tok/sEstimated
AI Mo Kimina Prover 72BQ8
Size: 72B
Not supported73GB required · 48GB available
~19 tok/sEstimated
AI Mo Kimina Prover 72BFP16
Size: 72B
Not supported146GB required · 48GB available
~10 tok/sEstimated
Alibaba Nlp Gte Qwen2 1.5B InstructQ4
Size: 1.5B
Fits comfortably1GB required · 48GB available
~160 tok/sEstimated
Alibaba Nlp Gte Qwen2 1.5B InstructQ8
Size: 1.5B
Fits comfortably2GB required · 48GB available
~110 tok/sEstimated
Alibaba Nlp Gte Qwen2 1.5B InstructFP16
Size: 1.5B
Fits comfortably4GB required · 48GB available
~60 tok/sEstimated

Note: Performance estimates are calculated. Real results may vary. Methodology · Submit real data

Alternative GPUs

RTX 4090
24GB

Explore how RTX 4090 stacks up for local inference workloads.

RTX 4080
16GB

Explore how RTX 4080 stacks up for local inference workloads.

RTX 4070 Ti
12GB

Explore how RTX 4070 Ti stacks up for local inference workloads.

RTX 3090
24GB

Explore how RTX 3090 stacks up for local inference workloads.

RX 7900 XTX
24GB

Explore how RX 7900 XTX stacks up for local inference workloads.