Consumer · NVIDIA · 8 GB VRAM
What can you run on the RTX 4060 Ti 8GB?
The RTX 4060 Ti 8GB has 8 GB of VRAM. Of the 38 models we track, 23 run on this card — 9 at full FP16 precision and 14 once quantized. Every figure below counts model weights plus the KV cache at a 4,096-token context.
9
Run in FP16
14
Need quantization
15
Do not fit
Runs at full FP16 precision
These models fit without quantization, so you keep full output quality.
| Model | Params | Precision | Total VRAM | % of 8 GB | |
|---|---|---|---|---|---|
| Qwen2.5 3B Instruct | 3.09B | FP16 / BF16 | 7.2 GB | 90% | Details |
| Qwen3 1.7B | 2.03B | FP16 / BF16 | 5.1 GB | 64% | Details |
| DeepSeek R1 Distill Qwen 1.5B | 1.78B | FP16 / BF16 | 4.2 GB | 53% | Details |
| SmolLM2 1.7B Instruct | 1.71B | FP16 / BF16 | 4.7 GB | 59% | Details |
| Qwen2.5 1.5B Instruct | 1.54B | FP16 / BF16 | 3.6 GB | 45% | Details |
| Qwen2.5 Coder 1.5B Instruct | 1.54B | FP16 / BF16 | 3.6 GB | 45% | Details |
| Qwen3 0.6B | 0.75B | FP16 / BF16 | 2.1 GB | 26% | Details |
| Qwen2.5 0.5B Instruct | 0.49B | FP16 / BF16 | 1.1 GB | 14% | Details |
| SmolLM2 360M Instruct | 0.36B | FP16 / BF16 | 1 GB | 13% | Details |
Runs with quantization
These need an INT8 or 4-bit build (GPTQ, AWQ, or GGUF). Quality stays close to full precision for most chat work, but re-test structured output and tool calling before shipping.
| Model | Params | Precision | Total VRAM | % of 8 GB | |
|---|---|---|---|---|---|
| Mistral Nemo Instruct | 12.25B | INT4 (4-bit) | 7.6 GB | 95% | Details |
| Gemma 2 9B Instruct | 9.24B | INT4 (4-bit) | 6.6 GB | 83% | Details |
| Qwen3 8B | 8.19B | INT4 (4-bit) | 5.3 GB | 66% | Details |
| Granite 3.1 8B Instruct | 8.17B | INT4 (4-bit) | 5.3 GB | 66% | Details |
| Llama 3.1 8B Instruct | 8.03B | INT4 (4-bit) | 5.1 GB | 64% | Details |
| DeepSeek R1 Distill Llama 8B | 8.03B | INT4 (4-bit) | 5.1 GB | 64% | Details |
| Qwen2.5 7B Instruct | 7.62B | INT4 (4-bit) | 4.6 GB | 57% | Details |
| Qwen2.5 Coder 7B Instruct | 7.62B | INT4 (4-bit) | 4.6 GB | 57% | Details |
| DeepSeek R1 Distill Qwen 7B | 7.62B | INT4 (4-bit) | 4.6 GB | 57% | Details |
| OLMo 2 7B Instruct | 7.3B | INT4 (4-bit) | 6.2 GB | 78% | Details |
| Mistral 7B Instruct v0.3 | 7.25B | INT4 (4-bit) | 4.7 GB | 59% | Details |
| Qwen3 4B | 4.02B | INT8 (8-bit) | 5.2 GB | 65% | Details |
| Phi-3.5 Mini Instruct | 3.82B | INT8 (8-bit) | 5.9 GB | 74% | Details |
| Llama 3.2 3B Instruct | 3.21B | INT8 (8-bit) | 4.1 GB | 51% | Details |
Too large for this card
These exceed 8 GB even at 4-bit. You would need multiple cards with tensor parallelism, a larger GPU, or a smaller model.
| Model | Params | Precision | Total VRAM | % of 8 GB | |
|---|---|---|---|---|---|
| Qwen2.5 72B Instruct | 72.71B | INT4 (4-bit) | 43 GB | 538% | Details |
| Llama 3.1 70B Instruct | 70.6B | INT4 (4-bit) | 41.9 GB | 524% | Details |
| DeepSeek R1 Distill Llama 70B | 70.55B | INT4 (4-bit) | 41.9 GB | 524% | Details |
| Qwen2.5 32B Instruct | 32.8B | INT4 (4-bit) | 19.9 GB | 249% | Details |
| Qwen2.5 Coder 32B Instruct | 32.76B | INT4 (4-bit) | 19.8 GB | 248% | Details |
| Qwen3 32B | 32.76B | INT4 (4-bit) | 19.8 GB | 248% | Details |
| DeepSeek R1 Distill Qwen 32B | 32.76B | INT4 (4-bit) | 19.8 GB | 248% | Details |
| QwQ 32B | 32.76B | INT4 (4-bit) | 19.8 GB | 248% | Details |
| Gemma 2 27B Instruct | 27.2B | INT4 (4-bit) | 17 GB | 213% | Details |
| Mistral Small 24B Instruct | 23.57B | INT4 (4-bit) | 14.2 GB | 178% | Details |
| Qwen2.5 14B Instruct | 14.77B | INT4 (4-bit) | 9.3 GB | 116% | Details |
| Qwen3 14B | 14.77B | INT4 (4-bit) | 9.1 GB | 114% | Details |
| DeepSeek R1 Distill Qwen 14B | 14.77B | INT4 (4-bit) | 9.3 GB | 116% | Details |
| Qwen2.5 Coder 14B Instruct | 14.77B | INT4 (4-bit) | 9.3 GB | 116% | Details |
| Phi-4 | 14.66B | INT4 (4-bit) | 9.2 GB | 115% | Details |
How to read these numbers
Whether a model runs on the RTX 4060 Ti 8GB comes down to three memory costs measured against its 8 GB. First the weights: parameter count times bytes per parameter — 2 bytes in FP16, 1 in INT8, roughly 0.5 at 4-bit. Second the KV cache, which holds attention keys and values for every token in the context window and grows linearly with prompt length; it stays in FP16 even when the weights are quantized. Third, an allowance for activations, CUDA context, and fragmentation. A fit is only called comfortable when the total leaves about 10% headroom.
The practical consequence is that the table above is a starting point, not a guarantee. A model listed as fitting at 4,096 tokens can still run out of memory once conversations get long or several requests run at once, because the KV cache term grows with both. If you plan to use long context or serve concurrent users, size with the VRAM Calculator at your real context length before committing.
Compare other consumer GPUs
Frequently asked questions
What AI models can the RTX 4060 Ti 8GB run?
The RTX 4060 Ti 8GB has 8 GB of VRAM and runs 23 of the 38 models we track: 9 at full FP16 precision and 14 more once quantized to 8-bit or 4-bit.
What is the largest LLM the RTX 4060 Ti 8GB can run?
Qwen2.5 3B Instruct (3.09B parameters) is the largest model that fits, using about 7.2 GB at FP16 / BF16.
Do I need quantization on the RTX 4060 Ti 8GB?
For larger models, yes. 9 models run in full FP16, but 14 only fit once you drop to INT8 or 4-bit using a GPTQ, AWQ, or GGUF build.
Why does the VRAM number here differ from the model size?
Model weights are only part of the cost. Every estimate here also adds the KV cache, which stores attention keys and values for each token in the context window and grows as your prompt gets longer. These figures use a 4,096-token context and leave about 10% headroom for activations and fragmentation.
Can the RTX 4060 Ti 8GB run Qwen2.5 72B Instruct?
No. Even at 4-bit, Qwen2.5 72B Instruct needs about 43 GB, which is more than the 8 GB available. You would need roughly 6× RTX 4060 Ti 8GB with tensor parallelism, or a single larger card.