RTX 4090 24GB for Local LLMs: What Fits, What Doesn't, and How Fast
The RTX 4090's 24GB of VRAM roughly doubles what a 12GB card can hold — and the payoff isn't "bigger everything," it's a specific jump: the 32B tier becomes usable at Q4, and 14B models move up to near-lossless Q8_0 with room to spare. Here's exactly what fits, at what quantization and context length, and where the honest ceiling actually is.
The hardware baseline
The RTX 4090 has 24GB of GDDR6X on a 384-bit bus and 1008 GB/s of memory bandwidth — the number that actually determines decode speed for local LLMs. Weight size determines what fits; bandwidth determines how fast tokens generate once it's loaded. They're independent, and the 4090 is strong on both.
At roughly 70% real-world utilization of peak bandwidth (accounting for KV cache reads, framework overhead, and memory controller efficiency), effective bandwidth is around 705 GB/s for decode. That's fast enough to run a 32B model at interactive speed — the range where a 12GB card can't follow.
What fits — model by model
Weight sizes below are calculated from real GGUF file sizes rather than theoretical bit counts. K-quants carry embedding tables, output heads, and per-block scale factors that naive bit-math misses — the effective bytes per parameter at Q4_K_M is closer to 0.61 than the 0.5 bytes (4 bits ÷ 8) a flat calculation gives.
| Model | Quant | Weight size | Fits 24GB | Est. speed |
|---|---|---|---|---|
| Llama 3.1 8B | Q8_0 | 8.51 GB | ✓ | ~83 tok/s |
| Qwen3 14B | Q8_0 | 15.69 GB | ✓ | ~45 tok/s |
| Phi-4 14B | Q8_0 | 15.58 GB | ✓ | ~45 tok/s |
| Mistral Small 24B | Q4_K_M | 14.64 GB | ✓ | ~48 tok/s |
| Mistral Small 24B | Q6_K | 19.68 GB | ✓ | ~36 tok/s |
| Qwen3 32B | Q4_K_M | 20.01 GB | ✓ | ~35 tok/s |
| DeepSeek-R1-Distill-Qwen 32B | Q4_K_M | 20.01 GB | ✓ | ~35 tok/s |
| Qwen3 32B | Q5_K_M | 23.29 GB | ✗ | — |
| Llama 3.1 70B | Q4_K_M | 43.07 GB | ✗ | — (offload) |
Weight sizes calculated from calibrated bytes-per-parameter values sourced from real GGUF file sizes. Speed estimates assume 1008 GB/s peak bandwidth at 70% real-world efficiency. Actual speed varies by backend and context length. Treat speeds as ±20%. Qwen3-32B Q5_K_M is listed to show the ceiling: weights alone leave no usable room for KV cache on 24GB.
Context length: the hidden constraint
On a 24GB card the 32B tier is close enough to the limit that context length decides your headroom. KV cache grows with context and comes out of the same pool as the weights — a 32B that loads fine at short context can OOM once the conversation gets long.
For Qwen3 32B at Q4_K_M (20.01GB weights), here's how context length changes the total VRAM requirement at F16 KV cache:
| Context length | KV cache | Total (weights + KV + overhead) | Fits 24GB |
|---|---|---|---|
| 4K (4,096) | 1.07 GB | 21.72 GB | ✓ |
| 8K (8,192) | 2.15 GB | 22.80 GB | ✓ |
| 16K (16,384) | 4.29 GB | 24.94 GB | ✗ |
| 32K (32,768) | 8.59 GB | 29.24 GB | ✗ |
The practical limit for Qwen3-32B Q4_K_M on a 4090 is around 12K context at F16 KV cache. If you need more, switching to Q8 KV cache roughly halves the KV cost and gets you past 24K — a good trade at this scale, since 8-bit KV cache has minimal quality impact. Alternatively, a 14B at Q8_0 leaves enough room for very long context with quality to spare.
Recommended configurations
Best overall: Qwen3 32B Q4_K_M
20.01GB weights, fits to ~12K context at F16 KV (further with Q8 KV), ~35 tok/s. This is what the 4090 buys you over a 12GB card — a genuinely capable 32B at interactive speed. Qwen3-32B holds up well at Q4_K_M for general use and reasoning.
Reasoning / math: DeepSeek-R1-Distill-Qwen 32B Q4_K_M
Same 20.01GB footprint and ~35 tok/s, tuned for chain-of-thought. If your workload is multi-step reasoning or math rather than general chat, this is the better use of the 32B tier — just watch context, since long reasoning traces eat KV cache quickly.
Quality + long context: Qwen3 14B Q8_0
15.69GB weights, ~8GB left for KV cache, ~45 tok/s. The config a 12GB card can't run at all — a 14B at near-lossless Q8_0 with room for very long context. Choose this over a 32B at Q4 when output quality per token and context length matter more than raw parameter count. Phi-4 14B Q8_0 (15.58GB) is the equivalent pick for reasoning and coding.
The 70B question
A 70B model needs about 43GB at Q4_K_M — it does not fit in 24GB. Running one on a single 4090 means RAM offload, which drops decode to roughly 1–5 tok/s. For 70B at usable speed you need a second 24GB card or a larger GPU. On a single 4090, a 32B at Q4/Q5 is the realistic quality ceiling.
Check any model or quant combination
The table above covers the most common configurations, but you can check any model, quantization, and context length combination — including KV cache quantization and estimated decode speed for your specific GPU — using the calculator.