How big an LLM fits in your server RAM
Size an LLM for a CPU server without guessing: weights in bytes, KV cache growth with context, runtime overhead, and why fitting is not the same as usable.
Filtering by topic #cpu-inference · clear
Size an LLM for a CPU server without guessing: weights in bytes, KV cache growth with context, runtime overhead, and why fitting is not the same as usable.
There is no Qwen 3.8 on Ollama yet. Here is the arithmetic for running the 27B tag that does exist on a CPU-only VPS, and what fits in 8 to 64 GB.