Consumer · NVIDIA · 8 GB VRAM
What can you run on the RTX 3050?
The RTX 3050 has 8 GB of VRAM. Of the 38 models we track, 23 run on this card — 9 at full FP16 precision and 14 once quantized. Every figure below counts model weights plus the KV cache at a 4,096-token context.
9
Run in FP16
14
Need quantization
15
Do not fit
Every model on the RTX 3050, at every precision
Total VRAM for each model at FP16, INT8, and 4-bit, including the KV cache at 4,096tokens. Green fits with headroom, amber fits but leaves under 10% spare, red exceeds the card's 8 GB. Models are ordered largest first, so the row where the colours change is the capacity limit of this card.
| Model | Params | FP16 | INT8 | 4-bit | KV @ 4K | Speed | Verdict |
|---|---|---|---|---|---|---|---|
| Qwen2.5 72B Instruct | 72.71B | 168.4 GB | 84.8 GB | 43 GB | 1.3 GB | — | Needs 6× cards |
| Llama 3.1 70B Instruct | 70.6B | 163.7 GB | 82.5 GB | 41.9 GB | 1.3 GB | — | Needs 6× cards |
| DeepSeek R1 Distill Llama 70B | 70.55B | 163.6 GB | 82.3 GB | 41.9 GB | 1.3 GB | — | Needs 6× cards |
| Qwen2.5 32B Instruct | 32.8B | 76.4 GB | 38.7 GB | 19.9 GB | 1 GB | — | Needs 3× cards |
| Qwen2.5 Coder 32B Instruct | 32.76B | 76.3 GB | 38.7 GB | 19.8 GB | 1 GB | — | Needs 3× cards |
| Qwen3 32B | 32.76B | 76.3 GB | 38.7 GB | 19.8 GB | 1 GB | — | Needs 3× cards |
| DeepSeek R1 Distill Qwen 32B | 32.76B | 76.3 GB | 38.7 GB | 19.8 GB | 1 GB | — | Needs 3× cards |
| QwQ 32B | 32.76B | 76.3 GB | 38.7 GB | 19.8 GB | 1 GB | — | Needs 3× cards |
| Gemma 2 27B Instruct | 27.2B | 64 GB | 32.7 GB | 17 GB | 1.4 GB | — | Needs 3× cards |
| Mistral Small 24B Instruct | 23.57B | 54.8 GB | 27.7 GB | 14.2 GB | 0.6 GB | — | Needs 2× cards |
| Qwen2.5 14B Instruct | 14.77B | 34.8 GB | 17.8 GB | 9.3 GB | 0.8 GB | — | Needs 2× cards |
| Qwen3 14B | 14.77B | 34.6 GB | 17.6 GB | 9.1 GB | 0.6 GB | — | Needs 2× cards |
| DeepSeek R1 Distill Qwen 14B | 14.77B | 34.8 GB | 17.8 GB | 9.3 GB | 0.8 GB | — | Needs 2× cards |
| Qwen2.5 Coder 14B Instruct | 14.77B | 34.8 GB | 17.8 GB | 9.3 GB | 0.8 GB | — | Needs 2× cards |
| Phi-4 | 14.66B | 34.5 GB | 17.7 GB | 9.2 GB | 0.8 GB | — | Needs 2× cards |
| Mistral Nemo Instruct | 12.25B | 28.8 GB | 14.7 GB | 7.6 GB | 0.6 GB | ~21 tok/sfine for chat | Needs INT4 (4-bit) |
| Gemma 2 9B Instruct | 9.24B | 22.6 GB | 11.9 GB | 6.6 GB | 1.3 GB | ~27 tok/sfine for chat | Needs INT4 (4-bit) |
| Qwen3 8B | 8.19B | 19.4 GB | 10 GB | 5.3 GB | 0.6 GB | ~31 tok/sfaster than reading | Needs INT4 (4-bit) |
| Granite 3.1 8B Instruct | 8.17B | 19.4 GB | 10 GB | 5.3 GB | 0.6 GB | ~31 tok/sfaster than reading | Needs INT4 (4-bit) |
| Llama 3.1 8B Instruct | 8.03B | 19 GB | 9.7 GB | 5.1 GB | 0.5 GB | ~31 tok/sfaster than reading | Needs INT4 (4-bit) |
| DeepSeek R1 Distill Llama 8B | 8.03B | 19 GB | 9.7 GB | 5.1 GB | 0.5 GB | ~31 tok/sfaster than reading | Needs INT4 (4-bit) |
| Qwen2.5 7B Instruct | 7.62B | 17.7 GB | 9 GB | 4.6 GB | 0.2 GB | ~33 tok/sfaster than reading | Needs INT4 (4-bit) |
| Qwen2.5 Coder 7B Instruct | 7.62B | 17.7 GB | 9 GB | 4.6 GB | 0.2 GB | ~33 tok/sfaster than reading | Needs INT4 (4-bit) |
| DeepSeek R1 Distill Qwen 7B | 7.62B | 17.7 GB | 9 GB | 4.6 GB | 0.2 GB | ~33 tok/sfaster than reading | Needs INT4 (4-bit) |
| OLMo 2 7B Instruct | 7.3B | 18.8 GB | 10.4 GB | 6.2 GB | 2 GB | ~34 tok/sfaster than reading | Needs INT4 (4-bit) |
| Mistral 7B Instruct v0.3 | 7.25B | 17.2 GB | 8.8 GB | 4.7 GB | 0.5 GB | ~35 tok/sfaster than reading | Needs INT4 (4-bit) |
| Qwen3 4B | 4.02B | 9.8 GB | 5.2 GB | 2.9 GB | 0.6 GB | ~31 tok/sfaster than reading | Needs INT8 (8-bit) |
| Phi-3.5 Mini Instruct | 3.82B | 10.3 GB | 5.9 GB | 3.7 GB | 1.5 GB | ~33 tok/sfaster than reading | Needs INT8 (8-bit) |
| Llama 3.2 3B Instruct | 3.21B | 7.8 GB | 4.1 GB | 2.2 GB | 0.4 GB | ~39 tok/sfaster than reading | Needs INT8 (8-bit) |
| Qwen2.5 3B Instruct | 3.09B | 7.2 GB | 3.7 GB | 1.9 GB | 0.1 GB | ~21 tok/sfine for chat | Runs in FP16 |
| Qwen3 1.7B | 2.03B | 5.1 GB | 2.7 GB | 1.6 GB | 0.4 GB | ~31 tok/sfaster than reading | Runs in FP16 |
| DeepSeek R1 Distill Qwen 1.5B | 1.78B | 4.2 GB | 2.1 GB | 1.1 GB | 0.1 GB | ~35 tok/sfaster than reading | Runs in FP16 |
| SmolLM2 1.7B Instruct | 1.71B | 4.7 GB | 2.8 GB | 1.8 GB | 0.8 GB | ~36 tok/sfaster than reading | Runs in FP16 |
| Qwen2.5 1.5B Instruct | 1.54B | 3.6 GB | 1.9 GB | 1 GB | 0.1 GB | ~40 tok/sfaster than reading | Runs in FP16 |
| Qwen2.5 Coder 1.5B Instruct | 1.54B | 3.6 GB | 1.9 GB | 1 GB | 0.1 GB | ~40 tok/sfaster than reading | Runs in FP16 |
| Qwen3 0.6B | 0.75B | 2.1 GB | 1.3 GB | 0.8 GB | 0.4 GB | ~76 tok/sfaster than reading | Runs in FP16 |
| Qwen2.5 0.5B Instruct | 0.49B | 1.1 GB | 0.6 GB | 0.3 GB | 0 GB | ~108 tok/sfaster than reading | Runs in FP16 |
| SmolLM2 360M Instruct | 0.36B | 1 GB | 0.6 GB | 0.4 GB | 0.2 GB | ~136 tok/sfaster than reading | Runs in FP16 |
Weights scale with precision — roughly 2 bytes per parameter at FP16, 1 at INT8, 0.5 at 4-bit — but the KV cache column does not. It stays in FP16 regardless of how the weights are quantized, which is why 4-bit stops helping once the cache alone approaches the card's capacity. Speed is quoted at the best precision each model actually runs at on this card, for a single request.
Fitting is not the same as usable
A model can fit the RTX 3050and still be too slow to use, which is the gap the speed column closes. Generating text is memory-bandwidth-bound rather than compute-bound: producing each token requires reading every weight once, so throughput is roughly the card's bandwidth divided by the bytes those weights occupy. The RTX 3050 moves about 224 GB/s at peak, and a well-optimized runtime realizes roughly 60% of that on a single stream once attention, sampling, and kernel launch overhead are counted.
Two consequences follow. Quantization buys speed as well as memory — dropping from FP16 to 4-bit quarters the bytes read per token and so roughly quadruples generation speed, which is often the stronger argument for it. And bandwidth, not TFLOPS, is the spec that predicts how a card feels: a GPU with more compute but slower memory will generate text more slowly on the same model. As a rough guide, above 30 tokens per second output arrives faster than most people read, 10 to 30 is comfortable for chat but sluggish for long answers, and below about 3 the model is only practical for background or batch work.
These are single-request estimates and deliberately conservative. Serving several users at once raises total throughput well above these figures, because one weight read is shared across the whole batch, while each individual response gets no faster. Speculative decoding and heavily tuned runtimes can also beat them. Treat the numbers as a floor for interactive use rather than a benchmark result.
Where this card hits its limit
Full precision on the RTX 3050 runs out after Qwen2.5 3B Instruct at 3.09B parameters, which uses about 7.2 GB of the available 8 GB. That is the first cliff, and it is the one that costs you output quality rather than the ability to run at all. The second cliff is harder. Phi-4 needs roughly 9.2 GB even at 4-bit, past the point where quantization can rescue it. Below that size you are choosing a precision; above it you are choosing a different card, or roughly 2 of these with tensor parallelism.
The practical reading is that VRAM is a step function, not a slider. Between the two cliffs, dropping from FP16 to INT8 costs very little measurable quality on most chat and summarization work, and 4-bit is usually acceptable too — but structured output, tool calling, and code generation degrade first and degrade quietly, so re-run your own evaluations after quantizing rather than trusting the benchmark deltas.
What happens at longer context
The table above assumes a 4,096-token prompt. Weights do not change with context but the KV cache grows linearly with it, so a model that fits comfortably on a short prompt can run out of memory in a long conversation. These are the largest models that fit the RTX 3050, recomputed at their best precision as context grows.
| Model | 4K tokens | 16K tokens | 32K tokens |
|---|---|---|---|
| Mistral Nemo Instruct | 7.6 GB · tight | 9.5 GB · over | 12 GB · over |
| Gemma 2 9B Instruct | 6.6 GB · fits | 10.6 GB · over | 15.8 GB · over |
| Qwen3 8B | 5.3 GB · fits | 7 GB · fits | 9.2 GB · over |
| Granite 3.1 8B Instruct | 5.3 GB · fits | 7.2 GB · fits | 9.7 GB · over |
| Llama 3.1 8B Instruct | 5.1 GB · fits | 6.6 GB · fits | 8.6 GB · over |
| DeepSeek R1 Distill Llama 8B | 5.1 GB · fits | 6.6 GB · fits | 8.6 GB · over |
Any row that turns amber or red before the last column is a model you can demo but cannot ship on this card at that context. Two levers help before you buy more hardware: switch to a model with grouped-query attention, which cuts the cache by the ratio of attention heads to key-value heads, or cap the served context below the model's maximum. Batching works against you here — every concurrent request carries its own cache, so serving four users at 8K costs roughly the same as one user at 32K.
How to read these numbers
Whether a model runs on the RTX 3050 comes down to three memory costs measured against its 8 GB. First the weights: parameter count times bytes per parameter — 2 bytes in FP16, 1 in INT8, roughly 0.5 at 4-bit. Second the KV cache, which holds attention keys and values for every token in the context window and grows linearly with prompt length; it stays in FP16 even when the weights are quantized. Third, an allowance for activations, CUDA context, and fragmentation. A fit is only called comfortable when the total leaves about 10% headroom.
These are architecture-aware estimates, not benchmarks. Each model's layer count, key-value head count, and head dimension are read from its published config rather than assumed from parameter count, which matters because two 7B models with different attention layouts can differ by several gigabytes of cache. What the estimates cannot capture is runtime overhead: vLLM preallocates a large block of VRAM by design, llama.cpp can offload part of the model to system RAM, and driver and framework versions each take their own cut. Treat a comfortable fit as a green light and a tight fit as something to verify on the actual card before committing.
For numbers at your own context length, batch size, and quantization scheme, use the VRAM Calculator. To work the problem the other way — starting from a model and finding the cheapest card that runs it — use the GPU Picker.
Compare other consumer GPUs
Frequently asked questions
What AI models can the RTX 3050 run?
The RTX 3050 has 8 GB of VRAM and runs 23 of the 38 models we track: 9 at full FP16 precision and 14 more once quantized to 8-bit or 4-bit.
What is the largest LLM the RTX 3050 can run?
Qwen2.5 3B Instruct (3.09B parameters) is the largest model that fits, using about 7.2 GB at FP16 / BF16.
Do I need quantization on the RTX 3050?
For larger models, yes. 9 models run in full FP16, but 14 only fit once you drop to INT8 or 4-bit using a GPTQ, AWQ, or GGUF build.
Where does the RTX 3050 stop keeping up?
Full precision runs out after Qwen2.5 3B Instruct at 3.09B parameters. The card stops working entirely at Phi-4 (14.66B), which needs about 9.2 GB even at 4-bit.
How fast do these models actually generate on the RTX 3050?
Decoding is limited by memory bandwidth, not compute. The RTX 3050 peaks at about 224 GB/s, and each token requires reading the whole model once, so speed is roughly that bandwidth divided by the model's size in bytes — at about 60% realized efficiency for a single request. SmolLM2 360M Instruct is the quickest tracked model here at roughly 136 tokens per second. Quantizing helps twice over: 4-bit weights are a quarter the bytes of FP16, so they generate roughly four times faster.
Is more VRAM or more bandwidth better for LLM inference?
VRAM decides whether a model runs at all; bandwidth decides how fast it runs once it does. They are not interchangeable, and cards can be strong in one and weak in the other — an L4 has the same 24 GB as an RTX 4090 but roughly a third of the bandwidth, so it fits the same models and generates far more slowly. Size for VRAM first, then check the speed column to see whether the result is usable.
Why does the VRAM number here differ from the model size?
Model weights are only part of the cost. Every estimate here also adds the KV cache, which stores attention keys and values for each token in the context window and grows as your prompt gets longer. These figures use a 4,096-token context and leave about 10% headroom for activations and fragmentation.
Will these models still fit at 32K context?
Not always. The weights stay constant but the KV cache scales linearly with context, so a model sitting near the limit at 4,096 tokens can exceed 8 GB well before 32K. The long-context table above recomputes the largest fitting models at 4K, 16K, and 32K so you can see which ones lose their headroom first.
Can the RTX 3050 run Phi-4?
No. Even at 4-bit, Phi-4 needs about 9.2 GB, which is more than the 8 GB available. You would need roughly 2× RTX 3050 with tensor parallelism, or a single larger card.