All compatibility checks

Consumer · AMD · 16 GB VRAM

What can you run on the Radeon RX 9070 XT?

The Radeon RX 9070 XT has 16 GB of VRAM. Of the 38 models we track, 29 run on this card — 12 at full FP16 precision and 17 once quantized. Every figure below counts model weights plus the KV cache at a 4,096-token context.

12

Run in FP16

17

Need quantization

9

Do not fit

Runs at full FP16 precision

These models fit without quantization, so you keep full output quality.

ModelParamsPrecisionTotal VRAM% of 16 GB
Qwen3 4B4.02BFP16 / BF169.8 GB61%Details
Phi-3.5 Mini Instruct3.82BFP16 / BF1610.3 GB64%Details
Llama 3.2 3B Instruct3.21BFP16 / BF167.8 GB49%Details
Qwen2.5 3B Instruct3.09BFP16 / BF167.2 GB45%Details
Qwen3 1.7B2.03BFP16 / BF165.1 GB32%Details
DeepSeek R1 Distill Qwen 1.5B1.78BFP16 / BF164.2 GB26%Details
SmolLM2 1.7B Instruct1.71BFP16 / BF164.7 GB29%Details
Qwen2.5 1.5B Instruct1.54BFP16 / BF163.6 GB23%Details
Qwen2.5 Coder 1.5B Instruct1.54BFP16 / BF163.6 GB23%Details
Qwen3 0.6B0.75BFP16 / BF162.1 GB13%Details
Qwen2.5 0.5B Instruct0.49BFP16 / BF161.1 GB7%Details
SmolLM2 360M Instruct0.36BFP16 / BF161 GB6%Details

Runs with quantization

These need an INT8 or 4-bit build (GPTQ, AWQ, or GGUF). Quality stays close to full precision for most chat work, but re-test structured output and tool calling before shipping.

ModelParamsPrecisionTotal VRAM% of 16 GB
Mistral Small 24B Instruct23.57BINT4 (4-bit)14.2 GB89%Details
Qwen2.5 14B Instruct14.77BINT4 (4-bit)9.3 GB58%Details
Qwen3 14B14.77BINT4 (4-bit)9.1 GB57%Details
DeepSeek R1 Distill Qwen 14B14.77BINT4 (4-bit)9.3 GB58%Details
Qwen2.5 Coder 14B Instruct14.77BINT4 (4-bit)9.3 GB58%Details
Phi-414.66BINT4 (4-bit)9.2 GB57%Details
Mistral Nemo Instruct12.25BINT4 (4-bit)7.6 GB48%Details
Gemma 2 9B Instruct9.24BINT8 (8-bit)11.9 GB74%Details
Qwen3 8B8.19BINT8 (8-bit)10 GB63%Details
Granite 3.1 8B Instruct8.17BINT8 (8-bit)10 GB63%Details
Llama 3.1 8B Instruct8.03BINT8 (8-bit)9.7 GB61%Details
DeepSeek R1 Distill Llama 8B8.03BINT8 (8-bit)9.7 GB61%Details
Qwen2.5 7B Instruct7.62BINT8 (8-bit)9 GB56%Details
Qwen2.5 Coder 7B Instruct7.62BINT8 (8-bit)9 GB56%Details
DeepSeek R1 Distill Qwen 7B7.62BINT8 (8-bit)9 GB56%Details
OLMo 2 7B Instruct7.3BINT8 (8-bit)10.4 GB65%Details
Mistral 7B Instruct v0.37.25BINT8 (8-bit)8.8 GB55%Details

Too large for this card

These exceed 16 GB even at 4-bit. You would need multiple cards with tensor parallelism, a larger GPU, or a smaller model.

ModelParamsPrecisionTotal VRAM% of 16 GB
Qwen2.5 72B Instruct72.71BINT4 (4-bit)43 GB269%Details
Llama 3.1 70B Instruct70.6BINT4 (4-bit)41.9 GB262%Details
DeepSeek R1 Distill Llama 70B70.55BINT4 (4-bit)41.9 GB262%Details
Qwen2.5 32B Instruct32.8BINT4 (4-bit)19.9 GB124%Details
Qwen2.5 Coder 32B Instruct32.76BINT4 (4-bit)19.8 GB124%Details
Qwen3 32B32.76BINT4 (4-bit)19.8 GB124%Details
DeepSeek R1 Distill Qwen 32B32.76BINT4 (4-bit)19.8 GB124%Details
QwQ 32B32.76BINT4 (4-bit)19.8 GB124%Details
Gemma 2 27B Instruct27.2BINT4 (4-bit)17 GB106%Details

How to read these numbers

Whether a model runs on the Radeon RX 9070 XT comes down to three memory costs measured against its 16 GB. First the weights: parameter count times bytes per parameter — 2 bytes in FP16, 1 in INT8, roughly 0.5 at 4-bit. Second the KV cache, which holds attention keys and values for every token in the context window and grows linearly with prompt length; it stays in FP16 even when the weights are quantized. Third, an allowance for activations, CUDA context, and fragmentation. A fit is only called comfortable when the total leaves about 10% headroom.

The practical consequence is that the table above is a starting point, not a guarantee. A model listed as fitting at 4,096 tokens can still run out of memory once conversations get long or several requests run at once, because the KV cache term grows with both. If you plan to use long context or serve concurrent users, size with the VRAM Calculator at your real context length before committing.

Compare other consumer GPUs

Frequently asked questions

What AI models can the Radeon RX 9070 XT run?

The Radeon RX 9070 XT has 16 GB of VRAM and runs 29 of the 38 models we track: 12 at full FP16 precision and 17 more once quantized to 8-bit or 4-bit.

What is the largest LLM the Radeon RX 9070 XT can run?

Qwen3 4B (4.02B parameters) is the largest model that fits, using about 9.8 GB at FP16 / BF16.

Do I need quantization on the Radeon RX 9070 XT?

For larger models, yes. 12 models run in full FP16, but 17 only fit once you drop to INT8 or 4-bit using a GPTQ, AWQ, or GGUF build.

Why does the VRAM number here differ from the model size?

Model weights are only part of the cost. Every estimate here also adds the KV cache, which stores attention keys and values for each token in the context window and grows as your prompt gets longer. These figures use a 4,096-token context and leave about 10% headroom for activations and fragmentation.

Can the Radeon RX 9070 XT run Qwen2.5 72B Instruct?

No. Even at 4-bit, Qwen2.5 72B Instruct needs about 43 GB, which is more than the 16 GB available. You would need roughly 3× Radeon RX 9070 XT with tensor parallelism, or a single larger card.