A downloadable tool

Buy Now$5.99 USD or more

Stop downloading 20 GB models that don't fit.

Inside

  • Spreadsheet calculator: pick one of 15 presets (Llama 3.x, Qwen2.5/Qwen3 incl. 30B-A3B MoE, Mistral Nemo/Small, Phi-4, Gemma) or enter any model's config numbers. Choose the quant (Q2_K to F16), context length, KV-cache type (f16/q8_0/q4_0) and offload share. You get weights, KV cache, total, fits / doesn't fit, and the maximum context that fits your card.
  • Command-line version (vram_calc.py, Python, no dependencies).
  • Cheat sheet (PDF): the formula explained, a quant quality guide, fit tables for 8/12/16/24 GB GPUs at 8k and 32k context, and 10 ways to fit more (KV quantisation, MoE expert offload, slots and context, and more).

Honest notes: these are estimates, typically within about 10% for llama.cpp/GGUF on NVIDIA. Other backends allocate differently, and speed isn't modelled. Presets were checked against model configs; verify the exact variant you download. Created with AI assistance; the formulas were checked by hand.

Published 22 hours ago
StatusReleased
CategoryTool
Authorredlinecortex
Tagsai, calculator, gpu, llm, local-llm, vram
AI DisclosureAI Assisted, Text

Purchase

Buy Now$5.99 USD or more

In order to download this tool you must purchase it at or above the minimum price of $5.99 USD. You will get access to the following files:

LLM-VRAM-Calculator-v1.0.zip 277 kB

Leave a comment

Log in with itch.io to leave a comment.