LLM VRAM Calculator — Will This Model Fit My GPU? (Spreadsheet + CLI + Cheat Sheet)
A downloadable tool
Buy Now$5.99 USD or more
Stop downloading 20 GB models that don't fit.
Inside
- Spreadsheet calculator: pick one of 15 presets (Llama 3.x, Qwen2.5/Qwen3 incl. 30B-A3B MoE, Mistral Nemo/Small, Phi-4, Gemma) or enter any model's config numbers. Choose the quant (Q2_K to F16), context length, KV-cache type (f16/q8_0/q4_0) and offload share. You get weights, KV cache, total, fits / doesn't fit, and the maximum context that fits your card.
- Command-line version (
vram_calc.py, Python, no dependencies). - Cheat sheet (PDF): the formula explained, a quant quality guide, fit tables for 8/12/16/24 GB GPUs at 8k and 32k context, and 10 ways to fit more (KV quantisation, MoE expert offload, slots and context, and more).
Honest notes: these are estimates, typically within about 10% for llama.cpp/GGUF on NVIDIA. Other backends allocate differently, and speed isn't modelled. Presets were checked against model configs; verify the exact variant you download. Created with AI assistance; the formulas were checked by hand.
| Published | 22 hours ago |
| Status | Released |
| Category | Tool |
| Author | redlinecortex |
| Tags | ai, calculator, gpu, llm, local-llm, vram |
| AI Disclosure | AI Assisted, Text |
Purchase
Buy Now$5.99 USD or more
In order to download this tool you must purchase it at or above the minimum price of $5.99 USD. You will get access to the following files:
LLM-VRAM-Calculator-v1.0.zip 277 kB

Leave a comment
Log in with itch.io to leave a comment.