A downloadable book

Buy Now$13.99 USD or more

You don't need a 24 GB card to run a genuinely capable local model.

This field guide shows exactly how a 35B-parameter Mixture-of-Experts model (Qwen3.6-35B-A3B) runs on a single RTX 3060 12 GB with a 131,072-token context, using only ~4 GB of VRAM — leaving the rest of the GPU free for games, image models or a 3D engine.

Measured on real hardware (RTX 3060 12 GB, i5-14600KF, 64 GB RAM):

  • Context: 8k → 128k tokens
  • Prompt processing: ~220 → ~540–670 tokens/s
  • Generation: ~27–29 tokens/s
  • GPU memory used by the model: ~4.0 GB

What you get

  • 7-page PDF guide (+ Markdown version): why MoE fits, the full launch command explained flag by flag, a VRAM budget planner for 8/12/16/24 GB cards, troubleshooting and a security checklist.
  • Ready-to-use systemd user service + 3-strike health-check timer — the server starts at boot and restarts itself if it hangs.
  • One-file config (llm.env) for all tunables.
  • A reproducible benchmark script that records VRAM and tokens/s for your own settings.
  • A dependency-free Python client for the OpenAI-compatible API.

Who it's for: Linux users with an 8–16 GB NVIDIA GPU and 32 GB+ RAM who want a fast, private, always-on local LLM for coding tools, agents or chat UIs.

Not included: model weights (download links/instructions are in the guide), llama.cpp binaries, Windows-specific service setup, or personal support. Your speed will depend on your CPU, RAM and build — figures are from one machine and are not guaranteed.

Made by a solo builder who runs this setup 24/7. Every purchase goes straight into more VRAM for the next experiments.

Published 1 day ago
StatusReleased
CategoryBook
Authorredlinecortex
Tagsai, gpu, guide, llama-cpp, llm, local-llm
AI DisclosureAI Assisted, Text

Purchase

Buy Now$13.99 USD or more

In order to download this book you must purchase it at or above the minimum price of $13.99 USD. You will get access to the following files:

12gb-35b-field-guide-v1.zip 442 kB

Leave a comment

Log in with itch.io to leave a comment.