Guides
Why Different VRAM Calculators Give Different Numbers
A methodology comparison — why the same model and quant can show different VRAM estimates across tools.
RTX 3060 12GB for Local LLMs: What Fits, What Doesn't, and How Fast
Exact VRAM requirements and estimated speeds for Llama 3.1, Qwen3, Phi-4, and more on a 3060 12GB.
RTX 4090 24GB for Local LLMs: What Fits, and How Fast
The 32B tier a 24GB card unlocks, what actually fits at each quant, and honest expectations for 70B.
RTX 4060 Ti 16GB for Local LLMs: What Fits, and How Fast
The 16GB tier that fits a high-quant 14B or a 24B at Q4 — and the bandwidth tradeoff that makes it slower than a 3060.