This GPU offers reliable throughput for local AI workloads. Pair it with the right model quantization to hit your desired tokens/sec, and monitor prices below to catch the best deal.
Quick Answer: RTX 3080 has 10GB VRAM, enough for models up to roughly 25B parameters at 4-bit quantization. It draws 320W under load.
With 10GB VRAM, RTX 3080 can run models up to approximately 25B parameters using 4-bit quantization. This is suitable for 7B-13B models like Llama 3 8B, Mistral 7B, and Qwen 7B.
Consider RTX 4070 or RTX 4080 — Significant performance increase for AI workloads.
Buy directly on Amazon with fast shipping and reliable customer service.
Essential accessories to pair with RTX 3080
Accessories Total
Typical prices — check Amazon for current pricing
💡 Not ready to buy? Try cloud GPUs first
Test RTX 3080 performance in the cloud before investing in hardware. Pay by the hour with no commitment.
Data-backed answers pulled from community benchmarks, manufacturer specs, and live pricing.
Owners running Qwen3-30B-A3B on a 10 GB RTX 3080 report roughly 15 tokens/sec after tuning, keeping interactive coding prompts responsive.
Source: Reddit – /r/LocalLLaMA (mquvxwc)
Some spec sheets assume higher ceilings, but real-world users note they already achieve ~10 tok/sec on a 10 GB 3080—showing how tuning beats blanket requirements.
Source: Reddit – /r/LocalLLaMA (mj408ke)
With larger context windows, Ollama reports 40% of layers moving to system RAM even on 12B models—illustrating the need to tune gpu_layers on 10 GB cards.
Source: Reddit – /r/LocalLLaMA (mnspe0d)
The RTX 3080 Founders Edition includes 10 GB GDDR6X, a 320 W board power, triple 8-pin power connectors, and NVIDIA recommends a 750 W PSU.
Source: TechPowerUp – RTX 3080 Specs
Showing 12 of 80 rows. Speeds are calculated estimates, not measurements — search for your model to jump straight to it.
| Model | Size | Quantization | Tokens/sec | VRAM used |
|---|---|---|---|---|
| Deepseek AI Deepseek Coder 1.3B Instruct | 1.3B | Q4 | ~190 tok/sEstimated | 1GB |
| Deepseek AI Deepseek R1 Distill Qwen 1.5B | 1.5B | Q4 | ~190 tok/sEstimated | 1GB |
| Deepseek AI Deepseek Ocr 2 | Unknown | Q4 | ~155 tok/sEstimated | 2GB |
| Deepseek AI Deepseek Ocr | Unknown | Q4 | ~155 tok/sEstimated | 2GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 8bit | 8B | Q4 | ~155 tok/sEstimated | 4GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 4bit | 8B | Q4 | ~155 tok/sEstimated | 4GB |
| Deepseek AI Deepseek R1 Distill Qwen 7B | 7B | Q4 | ~155 tok/sEstimated | 4GB |
| Nineninesix Kani Tts 2 En | Unknown | Q4 | ~150 tok/sEstimated | 1GB |
| Qwen Qwen3 Tts 12hz 1 7B Customvoice | 7B | Q4 | ~150 tok/sEstimated | 1GB |
| Zai Org Glm Ocr | Unknown | Q4 | ~150 tok/sEstimated | 1GB |
| Qwen Qwen3 Asr 1 7B | 7B | Q4 | ~150 tok/sEstimated | 2GB |
| Nari Labs Dia2 2B | 2B | Q4 | ~150 tok/sEstimated | 1GB |
Showing 12 of 240 rows.
| Model | Size | Quantization | Verdict | Estimated speed | VRAM needed |
|---|---|---|---|---|---|
| 01 AI Yi 1 5 34B Chat | 34B | Q4 | Not supported | ~44 tok/sEstimated | 18GB (have 10GB) |
| 01 AI Yi 1 5 34B Chat | 34B | Q8 | Not supported | ~31 tok/sEstimated | 35GB (have 10GB) |
| 01 AI Yi 1 5 34B Chat | 34B | FP16 | Not supported | ~17 tok/sEstimated | 69GB (have 10GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q4 | Fits comfortably | ~94 tok/sEstimated | 7GB (have 10GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q8 | Not supported | ~66 tok/sEstimated | 13GB (have 10GB) |
| AI Forever Rugpt 3.5 13B | 13B | FP16 | Not supported | ~36 tok/sEstimated | 26GB (have 10GB) |
| AI Mo Kimina Prover 72B | 72B | Q4 | Not supported | ~25 tok/sEstimated | 37GB (have 10GB) |
| AI Mo Kimina Prover 72B | 72B | Q8 | Not supported | ~18 tok/sEstimated | 73GB (have 10GB) |
| AI Mo Kimina Prover 72B | 72B | FP16 | Not supported | ~9.6 tok/sEstimated | 146GB (have 10GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q4 | Fits comfortably | ~150 tok/sEstimated | 1GB (have 10GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q8 | Fits comfortably | ~105 tok/sEstimated | 2GB (have 10GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | FP16 | Fits comfortably | ~57 tok/sEstimated | 4GB (have 10GB) |
Note: Performance estimates are calculated. Real results may vary. Methodology · Submit real data
Explore how RTX 4090 stacks up for local inference workloads.
Explore how RTX 4080 stacks up for local inference workloads.
Explore how RTX 4070 Ti stacks up for local inference workloads.
Explore how RTX 3090 stacks up for local inference workloads.
Explore how RX 7900 XTX stacks up for local inference workloads.
RPG • 2020
RPG • 2023
RPG • 2023
Action RPG • 2022
Action Adventure • 2019
Action Adventure • 2015
RPG • 2015
Racing • 2021
Action Adventure • 2023
Action RPG • 2020
Action RPG • 2024
Action RPG • 2023