All compatibility checks

Consumer · NVIDIA · 24 GB VRAM

What can you run on the RTX 3090?

The RTX 3090 has 24 GB of VRAM. Of the 38 models we track, 35 run on this card — 21 at full FP16 precision and 14 once quantized. Every figure below counts model weights plus the KV cache at a 4,096-token context.

21

Run in FP16

14

Need quantization

3

Do not fit

Every model on the RTX 3090, at every precision

Total VRAM for each model at FP16, INT8, and 4-bit, including the KV cache at 4,096tokens. Green fits with headroom, amber fits but leaves under 10% spare, red exceeds the card's 24 GB. Models are ordered largest first, so the row where the colours change is the capacity limit of this card.

ModelParamsFP16INT84-bitKV @ 4KSpeedVerdict
Qwen2.5 72B Instruct72.71B168.4 GB84.8 GB43 GB1.3 GBNeeds 2× cards
Llama 3.1 70B Instruct70.6B163.7 GB82.5 GB41.9 GB1.3 GBNeeds 2× cards
DeepSeek R1 Distill Llama 70B70.55B163.6 GB82.3 GB41.9 GB1.3 GBNeeds 2× cards
Qwen2.5 32B Instruct32.8B76.4 GB38.7 GB19.9 GB1 GB~32 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen2.5 Coder 32B Instruct32.76B76.3 GB38.7 GB19.8 GB1 GB~32 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen3 32B32.76B76.3 GB38.7 GB19.8 GB1 GB~32 tok/sfaster than readingNeeds INT4 (4-bit)
DeepSeek R1 Distill Qwen 32B32.76B76.3 GB38.7 GB19.8 GB1 GB~32 tok/sfaster than readingNeeds INT4 (4-bit)
QwQ 32B32.76B76.3 GB38.7 GB19.8 GB1 GB~32 tok/sfaster than readingNeeds INT4 (4-bit)
Gemma 2 27B Instruct27.2B64 GB32.7 GB17 GB1.4 GB~38 tok/sfaster than readingNeeds INT4 (4-bit)
Mistral Small 24B Instruct23.57B54.8 GB27.7 GB14.2 GB0.6 GB~44 tok/sfaster than readingNeeds INT4 (4-bit)
Qwen2.5 14B Instruct14.77B34.8 GB17.8 GB9.3 GB0.8 GB~35 tok/sfaster than readingNeeds INT8 (8-bit)
Qwen3 14B14.77B34.6 GB17.6 GB9.1 GB0.6 GB~35 tok/sfaster than readingNeeds INT8 (8-bit)
DeepSeek R1 Distill Qwen 14B14.77B34.8 GB17.8 GB9.3 GB0.8 GB~35 tok/sfaster than readingNeeds INT8 (8-bit)
Qwen2.5 Coder 14B Instruct14.77B34.8 GB17.8 GB9.3 GB0.8 GB~35 tok/sfaster than readingNeeds INT8 (8-bit)
Phi-414.66B34.5 GB17.7 GB9.2 GB0.8 GB~36 tok/sfaster than readingNeeds INT8 (8-bit)
Mistral Nemo Instruct12.25B28.8 GB14.7 GB7.6 GB0.6 GB~42 tok/sfaster than readingNeeds INT8 (8-bit)
Gemma 2 9B Instruct9.24B22.6 GB11.9 GB6.6 GB1.3 GB~54 tok/sfaster than readingNeeds INT8 (8-bit)
Qwen3 8B8.19B19.4 GB10 GB5.3 GB0.6 GB~32 tok/sfaster than readingRuns in FP16
Granite 3.1 8B Instruct8.17B19.4 GB10 GB5.3 GB0.6 GB~32 tok/sfaster than readingRuns in FP16
Llama 3.1 8B Instruct8.03B19 GB9.7 GB5.1 GB0.5 GB~33 tok/sfaster than readingRuns in FP16
DeepSeek R1 Distill Llama 8B8.03B19 GB9.7 GB5.1 GB0.5 GB~33 tok/sfaster than readingRuns in FP16
Qwen2.5 7B Instruct7.62B17.7 GB9 GB4.6 GB0.2 GB~34 tok/sfaster than readingRuns in FP16
Qwen2.5 Coder 7B Instruct7.62B17.7 GB9 GB4.6 GB0.2 GB~34 tok/sfaster than readingRuns in FP16
DeepSeek R1 Distill Qwen 7B7.62B17.7 GB9 GB4.6 GB0.2 GB~34 tok/sfaster than readingRuns in FP16
OLMo 2 7B Instruct7.3B18.8 GB10.4 GB6.2 GB2 GB~36 tok/sfaster than readingRuns in FP16
Mistral 7B Instruct v0.37.25B17.2 GB8.8 GB4.7 GB0.5 GB~36 tok/sfaster than readingRuns in FP16
Qwen3 4B4.02B9.8 GB5.2 GB2.9 GB0.6 GB~61 tok/sfaster than readingRuns in FP16
Phi-3.5 Mini Instruct3.82B10.3 GB5.9 GB3.7 GB1.5 GB~64 tok/sfaster than readingRuns in FP16
Llama 3.2 3B Instruct3.21B7.8 GB4.1 GB2.2 GB0.4 GB~74 tok/sfaster than readingRuns in FP16
Qwen2.5 3B Instruct3.09B7.2 GB3.7 GB1.9 GB0.1 GB~77 tok/sfaster than readingRuns in FP16
Qwen3 1.7B2.03B5.1 GB2.7 GB1.6 GB0.4 GB~108 tok/sfaster than readingRuns in FP16
DeepSeek R1 Distill Qwen 1.5B1.78B4.2 GB2.1 GB1.1 GB0.1 GB~120 tok/sfaster than readingRuns in FP16
SmolLM2 1.7B Instruct1.71B4.7 GB2.8 GB1.8 GB0.8 GB~124 tok/sfaster than readingRuns in FP16
Qwen2.5 1.5B Instruct1.54B3.6 GB1.9 GB1 GB0.1 GB~134 tok/sfaster than readingRuns in FP16
Qwen2.5 Coder 1.5B Instruct1.54B3.6 GB1.9 GB1 GB0.1 GB~134 tok/sfaster than readingRuns in FP16
Qwen3 0.6B0.75B2.1 GB1.3 GB0.8 GB0.4 GB~214 tok/sfaster than readingRuns in FP16
Qwen2.5 0.5B Instruct0.49B1.1 GB0.6 GB0.3 GB0 GB~267 tok/sfaster than readingRuns in FP16
SmolLM2 360M Instruct0.36B1 GB0.6 GB0.4 GB0.2 GB~305 tok/sfaster than readingRuns in FP16

Weights scale with precision — roughly 2 bytes per parameter at FP16, 1 at INT8, 0.5 at 4-bit — but the KV cache column does not. It stays in FP16 regardless of how the weights are quantized, which is why 4-bit stops helping once the cache alone approaches the card's capacity. Speed is quoted at the best precision each model actually runs at on this card, for a single request.

Fitting is not the same as usable

A model can fit the RTX 3090and still be too slow to use, which is the gap the speed column closes. Generating text is memory-bandwidth-bound rather than compute-bound: producing each token requires reading every weight once, so throughput is roughly the card's bandwidth divided by the bytes those weights occupy. The RTX 3090 moves about 936 GB/s at peak, and a well-optimized runtime realizes roughly 60% of that on a single stream once attention, sampling, and kernel launch overhead are counted.

Two consequences follow. Quantization buys speed as well as memory — dropping from FP16 to 4-bit quarters the bytes read per token and so roughly quadruples generation speed, which is often the stronger argument for it. And bandwidth, not TFLOPS, is the spec that predicts how a card feels: a GPU with more compute but slower memory will generate text more slowly on the same model. As a rough guide, above 30 tokens per second output arrives faster than most people read, 10 to 30 is comfortable for chat but sluggish for long answers, and below about 3 the model is only practical for background or batch work.

These are single-request estimates and deliberately conservative. Serving several users at once raises total throughput well above these figures, because one weight read is shared across the whole batch, while each individual response gets no faster. Speculative decoding and heavily tuned runtimes can also beat them. Treat the numbers as a floor for interactive use rather than a benchmark result.

Where this card hits its limit

Full precision on the RTX 3090 runs out after Qwen3 8B at 8.19B parameters, which uses about 19.4 GB of the available 24 GB. That is the first cliff, and it is the one that costs you output quality rather than the ability to run at all. The second cliff is harder. DeepSeek R1 Distill Llama 70B needs roughly 41.9 GB even at 4-bit, past the point where quantization can rescue it. Below that size you are choosing a precision; above it you are choosing a different card, or roughly 2 of these with tensor parallelism.

The practical reading is that VRAM is a step function, not a slider. Between the two cliffs, dropping from FP16 to INT8 costs very little measurable quality on most chat and summarization work, and 4-bit is usually acceptable too — but structured output, tool calling, and code generation degrade first and degrade quietly, so re-run your own evaluations after quantizing rather than trusting the benchmark deltas.

What happens at longer context

The table above assumes a 4,096-token prompt. Weights do not change with context but the KV cache grows linearly with it, so a model that fits comfortably on a short prompt can run out of memory in a long conversation. These are the largest models that fit the RTX 3090, recomputed at their best precision as context grows.

Model4K tokens16K tokens32K tokens
Qwen2.5 32B Instruct19.9 GB · fits22.9 GB · tight26.9 GB · over
Qwen2.5 Coder 32B Instruct19.8 GB · fits22.8 GB · tight26.8 GB · over
Qwen3 32B19.8 GB · fits22.8 GB · tight26.8 GB · over
DeepSeek R1 Distill Qwen 32B19.8 GB · fits22.8 GB · tight26.8 GB · over
QwQ 32B19.8 GB · fits22.8 GB · tight26.8 GB · over
Gemma 2 27B Instruct17 GB · fits21.4 GB · fits27.1 GB · over

Any row that turns amber or red before the last column is a model you can demo but cannot ship on this card at that context. Two levers help before you buy more hardware: switch to a model with grouped-query attention, which cuts the cache by the ratio of attention heads to key-value heads, or cap the served context below the model's maximum. Batching works against you here — every concurrent request carries its own cache, so serving four users at 8K costs roughly the same as one user at 32K.

How to read these numbers

Whether a model runs on the RTX 3090 comes down to three memory costs measured against its 24 GB. First the weights: parameter count times bytes per parameter — 2 bytes in FP16, 1 in INT8, roughly 0.5 at 4-bit. Second the KV cache, which holds attention keys and values for every token in the context window and grows linearly with prompt length; it stays in FP16 even when the weights are quantized. Third, an allowance for activations, CUDA context, and fragmentation. A fit is only called comfortable when the total leaves about 10% headroom.

These are architecture-aware estimates, not benchmarks. Each model's layer count, key-value head count, and head dimension are read from its published config rather than assumed from parameter count, which matters because two 7B models with different attention layouts can differ by several gigabytes of cache. What the estimates cannot capture is runtime overhead: vLLM preallocates a large block of VRAM by design, llama.cpp can offload part of the model to system RAM, and driver and framework versions each take their own cut. Treat a comfortable fit as a green light and a tight fit as something to verify on the actual card before committing.

For numbers at your own context length, batch size, and quantization scheme, use the VRAM Calculator. To work the problem the other way — starting from a model and finding the cheapest card that runs it — use the GPU Picker.

Compare other consumer GPUs

Frequently asked questions

What AI models can the RTX 3090 run?

The RTX 3090 has 24 GB of VRAM and runs 35 of the 38 models we track: 21 at full FP16 precision and 14 more once quantized to 8-bit or 4-bit.

What is the largest LLM the RTX 3090 can run?

Qwen3 8B (8.19B parameters) is the largest model that fits, using about 19.4 GB at FP16 / BF16.

Do I need quantization on the RTX 3090?

For larger models, yes. 21 models run in full FP16, but 14 only fit once you drop to INT8 or 4-bit using a GPTQ, AWQ, or GGUF build.

Where does the RTX 3090 stop keeping up?

Full precision runs out after Qwen3 8B at 8.19B parameters. The card stops working entirely at DeepSeek R1 Distill Llama 70B (70.55B), which needs about 41.9 GB even at 4-bit.

How fast do these models actually generate on the RTX 3090?

Decoding is limited by memory bandwidth, not compute. The RTX 3090 peaks at about 936 GB/s, and each token requires reading the whole model once, so speed is roughly that bandwidth divided by the model's size in bytes — at about 60% realized efficiency for a single request. SmolLM2 360M Instruct is the quickest tracked model here at roughly 305 tokens per second. Quantizing helps twice over: 4-bit weights are a quarter the bytes of FP16, so they generate roughly four times faster.

Is more VRAM or more bandwidth better for LLM inference?

VRAM decides whether a model runs at all; bandwidth decides how fast it runs once it does. They are not interchangeable, and cards can be strong in one and weak in the other — an L4 has the same 24 GB as an RTX 4090 but roughly a third of the bandwidth, so it fits the same models and generates far more slowly. Size for VRAM first, then check the speed column to see whether the result is usable.

Why does the VRAM number here differ from the model size?

Model weights are only part of the cost. Every estimate here also adds the KV cache, which stores attention keys and values for each token in the context window and grows as your prompt gets longer. These figures use a 4,096-token context and leave about 10% headroom for activations and fragmentation.

Will these models still fit at 32K context?

Not always. The weights stay constant but the KV cache scales linearly with context, so a model sitting near the limit at 4,096 tokens can exceed 24 GB well before 32K. The long-context table above recomputes the largest fitting models at 4K, 16K, and 32K so you can see which ones lose their headroom first.

Can the RTX 3090 run DeepSeek R1 Distill Llama 70B?

No. Even at 4-bit, DeepSeek R1 Distill Llama 70B needs about 41.9 GB, which is more than the 24 GB available. You would need roughly 2× RTX 3090 with tensor parallelism, or a single larger card.