llmfit

Calculators built and verified against real local inference setups — not just theoretical numbers. Figure out what fits your GPU before you download a 15GB model and find out the hard way.

VRAM Calculator

Calculate exact VRAM needed for any model, quantization, and context length.

What Can I Run?

Pick your GPU and see the best local LLM it can run — the highest-quality quant that fits, with estimated speed.

Context Length Calculator

See the max context length that fits your GPU — and how much more Q8/Q4 KV cache buys you.

RAM Calculator

Can you run a model on CPU / system RAM, and how fast? Built for big MoE models on RAM-heavy rigs.