All compatibility checks

Compatibility check

Can you run Mistral Small 24B Instruct on the RTX 6000 Ada?

Yes — with quantization

Yes, with quantization. Mistral Small 24B Instruct does not fit the RTX 6000 Ada (48 GB) in FP16, but it runs at INT8 (8-bit) using about 27.7 GB (58% of VRAM). Use a GPTQ, AWQ, or GGUF build to get there, and keep prompts moderate to leave room for the KV cache.

Memory breakdown

Weights plus a 0.6 GB KV cache at 4,096tokens, against the card's 48 GB. Verdicts leave ~10% headroom for activations and fragmentation.

PrecisionWeightsKV cacheTotal% of 48 GBFit
FP16 / BF16full quality54.2 GB0.6 GB54.8 GB114%No
INT8 (8-bit)near-full quality27.1 GB0.6 GB27.7 GB58%Fits
INT4 (4-bit)GPTQ / AWQ / GGUF Q413.6 GB0.6 GB14.2 GB30%Fits

Planning estimates, not a substitute for profiling. Real usage varies with the inference runtime, batch size, and how much context you actually use — the KV cache grows linearly with prompt length.

GPUs that run Mistral Small 24B Instruct

Cards where this model fits (at its best precision):

Go deeper

Frequently asked questions

Can the RTX 6000 Ada run Mistral Small 24B Instruct?

Not in FP16, but yes at INT8 (8-bit), where it uses about 27.7 GB versus the card's 48 GB.

How much VRAM does Mistral Small 24B Instruct need?

Approximately 54.2 GB in FP16, 27.1 GB in INT8, and 13.6 GB in 4-bit for the weights, plus a KV cache of about 0.6 GB at 4,096 tokens.

Does quantization let Mistral Small 24B Instruct fit on the RTX 6000 Ada?

Yes. Dropping to INT8 (8-bit) brings total usage to about 27.7 GB, which fits the 48 GB card with headroom for the KV cache.

What happens to memory with longer context?

The KV cache grows linearly with prompt length. At 4,096 tokens it is about 0.6 GB here; doubling the context roughly doubles that term, so long-context use can push a tight fit over the edge.