This GPU offers reliable throughput for local AI workloads. Pair it with the right model quantization to hit your desired tokens/sec, and monitor prices below to catch the best deal.
Quick Answer: AMD Instinct MI210 has 64GB VRAM, enough for models up to roughly 160B parameters at 4-bit quantization. It draws 300W under load.
With 64GB VRAM, AMD Instinct MI210 can run models up to approximately 160B parameters using 4-bit quantization. That covers most popular models, including 70B-class ones at 4-bit.
Consider H100 or MI300X — Maximum VRAM for enterprise workloads.
Buy directly on Amazon with fast shipping and reliable customer service.
Essential accessories to pair with AMD Instinct MI210
Accessories Total
Typical prices — check Amazon for current pricing
💡 Not ready to buy? Try cloud GPUs first
Test AMD Instinct MI210 performance in the cloud before investing in hardware. Pay by the hour with no commitment.
Showing 12 of 80 rows. Speeds are calculated estimates, not measurements — search for your model to jump straight to it.
| Model | Size | Quantization | Tokens/sec | VRAM used |
|---|---|---|---|---|
| Deepseek AI Deepseek Coder 1.3B Instruct | 1.3B | Q4 | ~355 tok/sEstimated | 1GB |
| Deepseek AI Deepseek R1 Distill Qwen 1.5B | 1.5B | Q4 | ~355 tok/sEstimated | 1GB |
| Deepseek AI Deepseek Ocr | Unknown | Q4 | ~295 tok/sEstimated | 2GB |
| Deepseek AI Deepseek Ocr 2 | Unknown | Q4 | ~295 tok/sEstimated | 2GB |
| Deepseek AI Deepseek R1 Distill Qwen 7B | 7B | Q4 | ~295 tok/sEstimated | 4GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 4bit | 8B | Q4 | ~295 tok/sEstimated | 4GB |
| Lmstudio Community Deepseek R1 0528 Qwen3 8B Mlx 8bit | 8B | Q4 | ~295 tok/sEstimated | 4GB |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q4 | ~285 tok/sEstimated | 1GB |
| Allenai Olmo 2 0425 1B | 1B | Q4 | ~285 tok/sEstimated | 1GB |
| Apple Openelm 1 1B Instruct | 1B | Q4 | ~285 tok/sEstimated | 1GB |
| Bigscience Bloomz 560M | Unknown | Q4 | ~285 tok/sEstimated | 1GB |
| Distilbert Distilgpt2 | Unknown | Q4 | ~285 tok/sEstimated | 1GB |
Showing 12 of 240 rows.
| Model | Size | Quantization | Verdict | Estimated speed | VRAM needed |
|---|---|---|---|---|---|
| 01 AI Yi 1 5 34B Chat | 34B | Q4 | Fits comfortably | ~83 tok/sEstimated | 18GB (have 64GB) |
| 01 AI Yi 1 5 34B Chat | 34B | Q8 | Fits comfortably | ~58 tok/sEstimated | 35GB (have 64GB) |
| 01 AI Yi 1 5 34B Chat | 34B | FP16 | Not supported | ~32 tok/sEstimated | 69GB (have 64GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q4 | Fits comfortably | ~180 tok/sEstimated | 7GB (have 64GB) |
| AI Forever Rugpt 3.5 13B | 13B | Q8 | Fits comfortably | ~125 tok/sEstimated | 13GB (have 64GB) |
| AI Forever Rugpt 3.5 13B | 13B | FP16 | Fits comfortably | ~68 tok/sEstimated | 26GB (have 64GB) |
| AI Mo Kimina Prover 72B | 72B | Q4 | Fits comfortably | ~48 tok/sEstimated | 37GB (have 64GB) |
| AI Mo Kimina Prover 72B | 72B | Q8 | Not supported | ~33 tok/sEstimated | 73GB (have 64GB) |
| AI Mo Kimina Prover 72B | 72B | FP16 | Not supported | ~18 tok/sEstimated | 146GB (have 64GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q4 | Fits comfortably | ~285 tok/sEstimated | 1GB (have 64GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | Q8 | Fits comfortably | ~200 tok/sEstimated | 2GB (have 64GB) |
| Alibaba Nlp Gte Qwen2 1.5B Instruct | 1.5B | FP16 | Fits comfortably | ~110 tok/sEstimated | 4GB (have 64GB) |
Note: Performance estimates are calculated. Real results may vary. Methodology · Submit real data
Explore how RTX 4090 stacks up for local inference workloads.
Explore how RTX 4080 stacks up for local inference workloads.
Explore how RTX 4070 Ti stacks up for local inference workloads.
Explore how RTX 3090 stacks up for local inference workloads.
Explore how RX 7900 XTX stacks up for local inference workloads.