RTX 4090 remains the go-to GPU for local AI workloads. It runs every mainstream 70B model, sustains the fastest consumer inference speeds, and anchors premium builds that scale to production deployments.
Quick Answer: RTX 4090 has 24GB VRAM, enough for models up to roughly 60B parameters at 4-bit quantization. It draws 450W under load.
With 24GB VRAM, RTX 4090 can run models up to approximately 60B parameters using 4-bit quantization. That covers most popular models, including 70B-class ones at 4-bit.
Consider RTX 4090 or RTX 6000 Ada — 24GB Ada offers better efficiency than Ampere.
Open direct compatibility pages for this GPU with VRAM fit and estimated speed.
Static benchmark coverage for nineninesix-kani-tts-2-en on rtx-4090.
Static benchmark coverage for fireredteam-firered-image-edit-1-0 on rtx-4090.
Static benchmark coverage for nanbeige-nanbeige4-1-3b on rtx-4090.
Static benchmark coverage for qwen-qwen3-tts-12hz-1-7b-customvoice on rtx-4090.
Static benchmark coverage for microsoft-vibevoice-asr on rtx-4090.
Static benchmark coverage for zai-org-glm-ocr on rtx-4090.
Static benchmark coverage for zai-org-glm-4-7-flash on rtx-4090.
Static benchmark coverage for deepseek-ai-deepseek-ocr-2 on rtx-4090.
Static benchmark coverage for qwen-qwen3-asr-1-7b on rtx-4090.
Static benchmark coverage for nvidia-personaplex-7b-v1 on rtx-4090.
Static benchmark coverage for essentialai-rnj-1 on rtx-4090.
Static benchmark coverage for mistralai-ministral-3-14b-instruct-2512 on rtx-4090.
Buy directly on Amazon with fast shipping and reliable customer service.
Essential accessories to pair with RTX 4090
Accessories Total
Typical prices — check Amazon for current pricing
💡 Not ready to buy? Try cloud GPUs first
Test RTX 4090 performance in the cloud before investing in hardware. Pay by the hour with no commitment.
Data-backed answers pulled from community benchmarks, manufacturer specs, and live pricing.
Community llama.cpp benchmarks of the ubergarm/Qwen3-30B-A3B-GGUF build show the RTX 4090 sustaining roughly 150–160 tokens/sec with CUDA kernels, keeping decode latency under 7 ms per token.
Source: Reddit – /r/LocalLLaMA (mq59v1k)
No. Builders loading Llama 3.1 70B Q4_K_M report roughly half the tensor pages spilling to system RAM on a 24 GB 4090, which drags throughput because PCIe becomes the bottleneck. Multi-GPU setups or 48 GB cards avoid the spill.
Source: Reddit – /r/LocalLLaMA (mqcouez)
Power users running multi-4090 racks note that a single 4090 comfortably hosts one 32B-class model; parallel agents or MoE workloads need tensor parallelism across multiple GPUs to keep speeds high.
Source: Reddit – /r/LocalLLaMA (mqwkgv3)
NVIDIA rates the RTX 4090 at 450 W board power and recommends at least an 850 W PSU with the 16-pin 12VHPWR connector to maintain headroom for AI workloads.
Source: TechPowerUp – RTX 4090 Specs
Showing 12 of 80 rows. Speeds are calculated estimates, not measurements — search for your model to jump straight to it.
| Model | Size | Quantization | Tokens/sec | VRAM used |
|---|---|---|---|---|
| Deepseek AI Deepseek Coder 1.3B Instruct | 1.3B | Q4 | ~270 tok/sEstimated | 1GB |
| Deepseek AI Deepseek R1 Distill Qwen 1.5B | 1.5B | Q4 | ~270 tok/sEstimated | 1GB |
| Deepseek AI Deepseek Ocr 2 | Unknown | Q4 | ~225 tok/sEstimated | 2GB |
| Deepseek AI Deepseek Ocr | Unknown | Q4 | ~225 tok/sEstimated | 2GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 8bit | 8B | Q4 | ~225 tok/sEstimated | 4GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 4bit | 8B | Q4 | ~225 tok/sEstimated | 4GB |
| Deepseek AI Deepseek R1 Distill Qwen 7B | 7B | Q4 | ~225 tok/sEstimated | 4GB |
| Nineninesix Kani Tts 2 En | Unknown | Q4 | ~215 tok/sEstimated | 1GB |
| Qwen Qwen3 Tts 12hz 1 7B Customvoice | 7B | Q4 | ~215 tok/sEstimated | 1GB |
| Zai Org Glm Ocr | Unknown | Q4 | ~215 tok/sEstimated | 1GB |
| Qwen Qwen3 Asr 1 7B | 7B | Q4 | ~215 tok/sEstimated | 2GB |
| Nari Labs Dia2 2B | 2B | Q4 | ~215 tok/sEstimated | 1GB |
Showing 12 of 240 rows.
| Model | Size | Quantization | Verdict | Estimated speed | VRAM needed |
|---|---|---|---|---|---|
| 01 AI Yi 1 5 34B Chat | 34B | Q4 | Fits comfortably | ~63 tok/sEstimated | 18GB (have 24GB) |
| 01 AI Yi 1 5 34B Chat | 34B | Q8 | Not supported | ~44 tok/sEstimated | 35GB (have 24GB) |
| 01 AI Yi 1 5 34B Chat | 34B | FP16 | Not supported | ~24 tok/sEstimated | 69GB (have 24GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q4 | Fits comfortably | ~135 tok/sEstimated | 7GB (have 24GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q8 | Fits comfortably | ~95 tok/sEstimated | 13GB (have 24GB) |
| AI Forever Rugpt 3.5 13B | 13B | FP16 | Not supported | ~51 tok/sEstimated | 26GB (have 24GB) |
| AI Mo Kimina Prover 72B | 72B | Q4 | Not supported | ~36 tok/sEstimated | 37GB (have 24GB) |
| AI Mo Kimina Prover 72B | 72B | Q8 | Not supported | ~25 tok/sEstimated | 73GB (have 24GB) |
| AI Mo Kimina Prover 72B | 72B | FP16 | Not supported | ~14 tok/sEstimated | 146GB (have 24GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q4 | Fits comfortably | ~215 tok/sEstimated | 1GB (have 24GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q8 | Fits comfortably | ~150 tok/sEstimated | 2GB (have 24GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | FP16 | Fits comfortably | ~82 tok/sEstimated | 4GB (have 24GB) |
Note: Performance estimates are calculated. Real results may vary. Methodology · Submit real data
Explore how RTX 4080 stacks up for local inference workloads.
Explore how RTX 4070 Ti stacks up for local inference workloads.
Explore how RTX 3090 stacks up for local inference workloads.
Explore how RX 7900 XTX stacks up for local inference workloads.
Explore how RTX 4070 stacks up for local inference workloads.
RPG • 2020
RPG • 2023
Action RPG • 2023
RPG • 2023
Survival Horror • 2023
Action RPG • 2022
Action RPG • 2024
Action Adventure • 2025
Survival Horror • 2023
Action • 2022
Action Adventure • 2023
Action Adventure • 2019