llmfit
Calculators built and verified against real local inference setups — not just theoretical numbers. Figure out what fits your GPU before you download a 15GB model and find out the hard way.
VRAM Calculator
Calculate exact VRAM needed for any model, quantization, and context length.
What Can I Run?
Pick your GPU and see the best local LLM it can run — the highest-quality quant that fits, with estimated speed.
Context Length Calculator
See the max context length that fits your GPU — and how much more Q8/Q4 KV cache buys you.
RAM Calculator
Can you run a model on CPU / system RAM, and how fast? Built for big MoE models on RAM-heavy rigs.