All compatibility checks

Consumer · NVIDIA · 8 GB VRAM

What can you run on the RTX 4060?

The RTX 4060 has 8 GB of VRAM. Of the 38 models we track, 23 run on this card — 9 at full FP16 precision and 14 once quantized. Every figure below counts model weights plus the KV cache at a 4,096-token context.

9

Run in FP16

14

Need quantization

15

Do not fit

Every model on the RTX 4060, at every precision

Total VRAM for each model at FP16, INT8, and 4-bit, including the KV cache at 4,096tokens. Green fits with headroom, amber fits but leaves under 10% spare, red exceeds the card's 8 GB. Models are ordered largest first, so the row where the colours change is the capacity limit of this card.

ModelParamsFP16INT84-bitKV @ 4KSpeedVerdict
Qwen2.5 72B Instruct72.71B168.4 GB84.8 GB43 GB1.3 GBNeeds 6× cards
Llama 3.1 70B Instruct70.6B163.7 GB82.5 GB41.9 GB1.3 GBNeeds 6× cards
DeepSeek R1 Distill Llama 70B70.55B163.6 GB82.3 GB41.9 GB1.3 GBNeeds 6× cards
Qwen2.5 32B Instruct32.8B76.4 GB38.7 GB19.9 GB1 GBNeeds 3× cards
Qwen2.5 Coder 32B Instruct32.76B76.3 GB38.7 GB19.8 GB1 GBNeeds 3× cards
Qwen3 32B32.76B76.3 GB38.7 GB19.8 GB1 GBNeeds 3× cards
DeepSeek R1 Distill Qwen 32B32.76B76.3 GB38.7 GB19.8 GB1 GBNeeds 3× cards
QwQ 32B32.76B76.3 GB38.7 GB19.8 GB1 GBNeeds 3× cards
Gemma 2 27B Instruct27.2B64 GB32.7 GB17 GB1.4 GBNeeds 3× cards
Mistral Small 24B Instruct23.57B54.8 GB27.7 GB14.2 GB0.6 GBNeeds 2× cards
Qwen2.5 14B Instruct14.77B34.8 GB17.8 GB9.3 GB0.8 GBNeeds 2× cards
Qwen3 14B14.77B34.6 GB17.6 GB9.1 GB0.6 GBNeeds 2× cards
DeepSeek R1 Distill Qwen 14B14.77B34.8 GB17.8 GB9.3 GB0.8 GBNeeds 2× cards
Qwen2.5 Coder 14B Instruct14.77B34.8 GB17.8 GB9.3 GB0.8 GBNeeds 2× cards
Phi-414.66B34.5 GB17.7 GB9.2 GB0.8 GBNeeds 2× cards
Mistral Nemo Instruct12.25B28.8 GB14.7 GB7.6 GB0.6 GB~25 tok/sfine for chatNeeds INT4 (4-bit)
Gemma 2 9B Instruct9.24B22.6 GB11.9 GB6.6 GB1.3 GB~33 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen3 8B8.19B19.4 GB10 GB5.3 GB0.6 GB~37 tok/sfaster than readingNeeds INT4 (4-bit)
Granite 3.1 8B Instruct8.17B19.4 GB10 GB5.3 GB0.6 GB~37 tok/sfaster than readingNeeds INT4 (4-bit)
Llama 3.1 8B Instruct8.03B19 GB9.7 GB5.1 GB0.5 GB~38 tok/sfaster than readingNeeds INT4 (4-bit)
DeepSeek R1 Distill Llama 8B8.03B19 GB9.7 GB5.1 GB0.5 GB~38 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen2.5 7B Instruct7.62B17.7 GB9 GB4.6 GB0.2 GB~39 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen2.5 Coder 7B Instruct7.62B17.7 GB9 GB4.6 GB0.2 GB~39 tok/sfaster than readingNeeds INT4 (4-bit)
DeepSeek R1 Distill Qwen 7B7.62B17.7 GB9 GB4.6 GB0.2 GB~39 tok/sfaster than readingNeeds INT4 (4-bit)
OLMo 2 7B Instruct7.3B18.8 GB10.4 GB6.2 GB2 GB~41 tok/sfaster than readingNeeds INT4 (4-bit)
Mistral 7B Instruct v0.37.25B17.2 GB8.8 GB4.7 GB0.5 GB~41 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen3 4B4.02B9.8 GB5.2 GB2.9 GB0.6 GB~38 tok/sfaster than readingNeeds INT8 (8-bit)
Phi-3.5 Mini Instruct3.82B10.3 GB5.9 GB3.7 GB1.5 GB~39 tok/sfaster than readingNeeds INT8 (8-bit)
Llama 3.2 3B Instruct3.21B7.8 GB4.1 GB2.2 GB0.4 GB~46 tok/sfaster than readingNeeds INT8 (8-bit)
Qwen2.5 3B Instruct3.09B7.2 GB3.7 GB1.9 GB0.1 GB~25 tok/sfine for chatRuns in FP16
Qwen3 1.7B2.03B5.1 GB2.7 GB1.6 GB0.4 GB~37 tok/sfaster than readingRuns in FP16
DeepSeek R1 Distill Qwen 1.5B1.78B4.2 GB2.1 GB1.1 GB0.1 GB~42 tok/sfaster than readingRuns in FP16
SmolLM2 1.7B Instruct1.71B4.7 GB2.8 GB1.8 GB0.8 GB~44 tok/sfaster than readingRuns in FP16
Qwen2.5 1.5B Instruct1.54B3.6 GB1.9 GB1 GB0.1 GB~48 tok/sfaster than readingRuns in FP16
Qwen2.5 Coder 1.5B Instruct1.54B3.6 GB1.9 GB1 GB0.1 GB~48 tok/sfaster than readingRuns in FP16
Qwen3 0.6B0.75B2.1 GB1.3 GB0.8 GB0.4 GB~89 tok/sfaster than readingRuns in FP16
Qwen2.5 0.5B Instruct0.49B1.1 GB0.6 GB0.3 GB0 GB~125 tok/sfaster than readingRuns in FP16
SmolLM2 360M Instruct0.36B1 GB0.6 GB0.4 GB0.2 GB~156 tok/sfaster than readingRuns in FP16

Weights scale with precision — roughly 2 bytes per parameter at FP16, 1 at INT8, 0.5 at 4-bit — but the KV cache column does not. It stays in FP16 regardless of how the weights are quantized, which is why 4-bit stops helping once the cache alone approaches the card's capacity. Speed is quoted at the best precision each model actually runs at on this card, for a single request.

Fitting is not the same as usable

A model can fit the RTX 4060and still be too slow to use, which is the gap the speed column closes. Generating text is memory-bandwidth-bound rather than compute-bound: producing each token requires reading every weight once, so throughput is roughly the card's bandwidth divided by the bytes those weights occupy. The RTX 4060 moves about 272 GB/s at peak, and a well-optimized runtime realizes roughly 60% of that on a single stream once attention, sampling, and kernel launch overhead are counted.

Two consequences follow. Quantization buys speed as well as memory — dropping from FP16 to 4-bit quarters the bytes read per token and so roughly quadruples generation speed, which is often the stronger argument for it. And bandwidth, not TFLOPS, is the spec that predicts how a card feels: a GPU with more compute but slower memory will generate text more slowly on the same model. As a rough guide, above 30 tokens per second output arrives faster than most people read, 10 to 30 is comfortable for chat but sluggish for long answers, and below about 3 the model is only practical for background or batch work.

These are single-request estimates and deliberately conservative. Serving several users at once raises total throughput well above these figures, because one weight read is shared across the whole batch, while each individual response gets no faster. Speculative decoding and heavily tuned runtimes can also beat them. Treat the numbers as a floor for interactive use rather than a benchmark result.

Where this card hits its limit

Full precision on the RTX 4060 runs out after Qwen2.5 3B Instruct at 3.09B parameters, which uses about 7.2 GB of the available 8 GB. That is the first cliff, and it is the one that costs you output quality rather than the ability to run at all. The second cliff is harder. Phi-4 needs roughly 9.2 GB even at 4-bit, past the point where quantization can rescue it. Below that size you are choosing a precision; above it you are choosing a different card, or roughly 2 of these with tensor parallelism.

The practical reading is that VRAM is a step function, not a slider. Between the two cliffs, dropping from FP16 to INT8 costs very little measurable quality on most chat and summarization work, and 4-bit is usually acceptable too — but structured output, tool calling, and code generation degrade first and degrade quietly, so re-run your own evaluations after quantizing rather than trusting the benchmark deltas.

What happens at longer context

The table above assumes a 4,096-token prompt. Weights do not change with context but the KV cache grows linearly with it, so a model that fits comfortably on a short prompt can run out of memory in a long conversation. These are the largest models that fit the RTX 4060, recomputed at their best precision as context grows.

Model4K tokens16K tokens32K tokens
Mistral Nemo Instruct7.6 GB · tight9.5 GB · over12 GB · over
Gemma 2 9B Instruct6.6 GB · fits10.6 GB · over15.8 GB · over
Qwen3 8B5.3 GB · fits7 GB · fits9.2 GB · over
Granite 3.1 8B Instruct5.3 GB · fits7.2 GB · fits9.7 GB · over
Llama 3.1 8B Instruct5.1 GB · fits6.6 GB · fits8.6 GB · over
DeepSeek R1 Distill Llama 8B5.1 GB · fits6.6 GB · fits8.6 GB · over

Any row that turns amber or red before the last column is a model you can demo but cannot ship on this card at that context. Two levers help before you buy more hardware: switch to a model with grouped-query attention, which cuts the cache by the ratio of attention heads to key-value heads, or cap the served context below the model's maximum. Batching works against you here — every concurrent request carries its own cache, so serving four users at 8K costs roughly the same as one user at 32K.

How to read these numbers

Whether a model runs on the RTX 4060 comes down to three memory costs measured against its 8 GB. First the weights: parameter count times bytes per parameter — 2 bytes in FP16, 1 in INT8, roughly 0.5 at 4-bit. Second the KV cache, which holds attention keys and values for every token in the context window and grows linearly with prompt length; it stays in FP16 even when the weights are quantized. Third, an allowance for activations, CUDA context, and fragmentation. A fit is only called comfortable when the total leaves about 10% headroom.

These are architecture-aware estimates, not benchmarks. Each model's layer count, key-value head count, and head dimension are read from its published config rather than assumed from parameter count, which matters because two 7B models with different attention layouts can differ by several gigabytes of cache. What the estimates cannot capture is runtime overhead: vLLM preallocates a large block of VRAM by design, llama.cpp can offload part of the model to system RAM, and driver and framework versions each take their own cut. Treat a comfortable fit as a green light and a tight fit as something to verify on the actual card before committing.

For numbers at your own context length, batch size, and quantization scheme, use the VRAM Calculator. To work the problem the other way — starting from a model and finding the cheapest card that runs it — use the GPU Picker.

Compare other consumer GPUs

Frequently asked questions

What AI models can the RTX 4060 run?

The RTX 4060 has 8 GB of VRAM and runs 23 of the 38 models we track: 9 at full FP16 precision and 14 more once quantized to 8-bit or 4-bit.

What is the largest LLM the RTX 4060 can run?

Qwen2.5 3B Instruct (3.09B parameters) is the largest model that fits, using about 7.2 GB at FP16 / BF16.

Do I need quantization on the RTX 4060?

For larger models, yes. 9 models run in full FP16, but 14 only fit once you drop to INT8 or 4-bit using a GPTQ, AWQ, or GGUF build.

Where does the RTX 4060 stop keeping up?

Full precision runs out after Qwen2.5 3B Instruct at 3.09B parameters. The card stops working entirely at Phi-4 (14.66B), which needs about 9.2 GB even at 4-bit.

How fast do these models actually generate on the RTX 4060?

Decoding is limited by memory bandwidth, not compute. The RTX 4060 peaks at about 272 GB/s, and each token requires reading the whole model once, so speed is roughly that bandwidth divided by the model's size in bytes — at about 60% realized efficiency for a single request. SmolLM2 360M Instruct is the quickest tracked model here at roughly 156 tokens per second. Quantizing helps twice over: 4-bit weights are a quarter the bytes of FP16, so they generate roughly four times faster.

Is more VRAM or more bandwidth better for LLM inference?

VRAM decides whether a model runs at all; bandwidth decides how fast it runs once it does. They are not interchangeable, and cards can be strong in one and weak in the other — an L4 has the same 24 GB as an RTX 4090 but roughly a third of the bandwidth, so it fits the same models and generates far more slowly. Size for VRAM first, then check the speed column to see whether the result is usable.

Why does the VRAM number here differ from the model size?

Model weights are only part of the cost. Every estimate here also adds the KV cache, which stores attention keys and values for each token in the context window and grows as your prompt gets longer. These figures use a 4,096-token context and leave about 10% headroom for activations and fragmentation.

Will these models still fit at 32K context?

Not always. The weights stay constant but the KV cache scales linearly with context, so a model sitting near the limit at 4,096 tokens can exceed 8 GB well before 32K. The long-context table above recomputes the largest fitting models at 4K, 16K, and 32K so you can see which ones lose their headroom first.

Can the RTX 4060 run Phi-4?

No. Even at 4-bit, Phi-4 needs about 9.2 GB, which is more than the 8 GB available. You would need roughly 2× RTX 4060 with tensor parallelism, or a single larger card.