RTX 3090 still delivers strong results for large language models thanks to its 24GB VRAM. It is ideal for enthusiasts leveraging the Ampere generation for budget workstation builds.
Quick Answer: RTX 3090 has 24GB VRAM, enough for models up to roughly 60B parameters at 4-bit quantization. It draws 350W under load.
With 24GB VRAM, RTX 3090 can run models up to approximately 60B parameters using 4-bit quantization. That covers most popular models, including 70B-class ones at 4-bit.
Consider RTX 4090 or RTX 6000 Ada — 24GB Ada offers better efficiency than Ampere.
Buy directly on Amazon with fast shipping and reliable customer service.
Essential accessories to pair with RTX 3090
Accessories Total
Typical prices — check Amazon for current pricing
💡 Not ready to buy? Try cloud GPUs first
Test RTX 3090 performance in the cloud before investing in hardware. Pay by the hour with no commitment.
Data-backed answers pulled from community benchmarks, manufacturer specs, and live pricing.
An Ampere builder reports ~18 tok/s on Llama 3 70B Q4 with a single 3090, and ~36 tok/s after adding tensor parallelism across four cards.
Source: Reddit – /r/LocalLLaMA (mqzh3yo)
Enthusiasts routinely see ~100 tokens/sec on Qwen 30B-A3B when tuned on a single RTX 3090, making it a budget-friendly coding workhorse.
Source: Reddit – /r/LocalLLaMA (mqs2r45)
Builders using x1 risers for dual 3080/3090 rigs measured no meaningful tokens/sec loss—the main penalty is slower model swaps, not slower inference.
Source: Reddit – /r/LocalLLaMA (mr10ib4)
RTX 3090 provides 24 GB GDDR6X, draws 350 W, and uses triple 8-pin PCIe power connectors. NVIDIA recommends a 750 W PSU.
Source: TechPowerUp – RTX 3090 Specs
Showing 12 of 80 rows. Speeds are calculated estimates, not measurements — search for your model to jump straight to it.
| Model | Size | Quantization | Tokens/sec | VRAM used |
|---|---|---|---|---|
| Deepseek AI Deepseek Coder 1.3B Instruct | 1.3B | Q4 | ~230 tok/sEstimated | 1GB |
| Deepseek AI Deepseek R1 Distill Qwen 1.5B | 1.5B | Q4 | ~230 tok/sEstimated | 1GB |
| Deepseek AI Deepseek Ocr 2 | Unknown | Q4 | ~195 tok/sEstimated | 2GB |
| Deepseek AI Deepseek Ocr | Unknown | Q4 | ~195 tok/sEstimated | 2GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 8bit | 8B | Q4 | ~195 tok/sEstimated | 4GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 4bit | 8B | Q4 | ~195 tok/sEstimated | 4GB |
| Deepseek AI Deepseek R1 Distill Qwen 7B | 7B | Q4 | ~195 tok/sEstimated | 4GB |
| Nineninesix Kani Tts 2 En | Unknown | Q4 | ~185 tok/sEstimated | 1GB |
| Qwen Qwen3 Tts 12hz 1 7B Customvoice | 7B | Q4 | ~185 tok/sEstimated | 1GB |
| Zai Org Glm Ocr | Unknown | Q4 | ~185 tok/sEstimated | 1GB |
| Qwen Qwen3 Asr 1 7B | 7B | Q4 | ~185 tok/sEstimated | 2GB |
| Nari Labs Dia2 2B | 2B | Q4 | ~185 tok/sEstimated | 1GB |
Showing 12 of 240 rows.
| Model | Size | Quantization | Verdict | Estimated speed | VRAM needed |
|---|---|---|---|---|---|
| 01 AI Yi 1 5 34B Chat | 34B | Q4 | Fits comfortably | ~54 tok/sEstimated | 18GB (have 24GB) |
| 01 AI Yi 1 5 34B Chat | 34B | Q8 | Not supported | ~38 tok/sEstimated | 35GB (have 24GB) |
| 01 AI Yi 1 5 34B Chat | 34B | FP16 | Not supported | ~21 tok/sEstimated | 69GB (have 24GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q4 | Fits comfortably | ~115 tok/sEstimated | 7GB (have 24GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q8 | Fits comfortably | ~81 tok/sEstimated | 13GB (have 24GB) |
| AI Forever Rugpt 3.5 13B | 13B | FP16 | Not supported | ~44 tok/sEstimated | 26GB (have 24GB) |
| AI Mo Kimina Prover 72B | 72B | Q4 | Not supported | ~31 tok/sEstimated | 37GB (have 24GB) |
| AI Mo Kimina Prover 72B | 72B | Q8 | Not supported | ~22 tok/sEstimated | 73GB (have 24GB) |
| AI Mo Kimina Prover 72B | 72B | FP16 | Not supported | ~12 tok/sEstimated | 146GB (have 24GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q4 | Fits comfortably | ~185 tok/sEstimated | 1GB (have 24GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q8 | Fits comfortably | ~130 tok/sEstimated | 2GB (have 24GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | FP16 | Fits comfortably | ~70 tok/sEstimated | 4GB (have 24GB) |
Note: Performance estimates are calculated. Real results may vary. Methodology · Submit real data
Explore how RTX 4090 stacks up for local inference workloads.
Explore how RTX 4080 stacks up for local inference workloads.
Explore how RTX 4070 Ti stacks up for local inference workloads.
Explore how RX 7900 XTX stacks up for local inference workloads.
Explore how RTX 4070 stacks up for local inference workloads.
RPG • 2020
RPG • 2023
Action RPG • 2023
RPG • 2023
Survival Horror • 2023
Action RPG • 2022
Action RPG • 2024
Action Adventure • 2025
Survival Horror • 2023
Action • 2022
Action Adventure • 2023
Action Adventure • 2019