All compatibility checks

Compatibility check

Can you run DeepSeek R1 Distill Llama 70B on the H200 141GB?

Yes — with quantization

Yes, with quantization. DeepSeek R1 Distill Llama 70B does not fit the H200 141GB (141 GB) in FP16, but it runs at INT8 (8-bit) using about 82.3 GB (58% of VRAM). Use a GPTQ, AWQ, or GGUF build to get there, and keep prompts moderate to leave room for the KV cache.

Memory breakdown

Weights plus a 1.3 GB KV cache at 4,096tokens, against the card's 141 GB. Verdicts leave ~10% headroom for activations and fragmentation.

PrecisionWeightsKV cacheTotal% of 141 GBFit
FP16 / BF16full quality162.3 GB1.3 GB163.6 GB116%No
INT8 (8-bit)near-full quality81.1 GB1.3 GB82.3 GB58%Fits
INT4 (4-bit)GPTQ / AWQ / GGUF Q440.6 GB1.3 GB41.9 GB30%Fits

Planning estimates, not a substitute for profiling. Real usage varies with the inference runtime, batch size, and how much context you actually use — the KV cache grows linearly with prompt length.

GPUs that run DeepSeek R1 Distill Llama 70B

Cards where this model fits (at its best precision):

Go deeper

Frequently asked questions

Can the H200 141GB run DeepSeek R1 Distill Llama 70B?

Not in FP16, but yes at INT8 (8-bit), where it uses about 82.3 GB versus the card's 141 GB.

How much VRAM does DeepSeek R1 Distill Llama 70B need?

Approximately 162.3 GB in FP16, 81.1 GB in INT8, and 40.6 GB in 4-bit for the weights, plus a KV cache of about 1.3 GB at 4,096 tokens.

Does quantization let DeepSeek R1 Distill Llama 70B fit on the H200 141GB?

Yes. Dropping to INT8 (8-bit) brings total usage to about 82.3 GB, which fits the 141 GB card with headroom for the KV cache.

What happens to memory with longer context?

The KV cache grows linearly with prompt length. At 4,096 tokens it is about 1.3 GB here; doubling the context roughly doubles that term, so long-context use can push a tight fit over the edge.