Ollama KV cache quantization: q8_0 vs f16
Raised num_ctx and Ollama no longer fits? The KV cache grew, not the weights. Set OLLAMA_KV_CACHE_TYPE with flash attention and watch the change in ollama ps.
Filtering by topic #quantization · clear
Raised num_ctx and Ollama no longer fits? The KV cache grew, not the weights. Set OLLAMA_KV_CACHE_TYPE with flash attention and watch the change in ollama ps.
Size an LLM for a CPU server without guessing: weights in bytes, KV cache growth with context, runtime overhead, and why fitting is not the same as usable.
Which Gemma 4 tag fits your VPS RAM, what each one costs in tokens per second on CPU, and the sizing arithmetic to check it before you pull 20GB.
GLM 5.2 is cloud only in Ollama's library. Here is the GLM model that actually fits a VPS, and the RAM each quantisation needs on your own box.
DeepSeek V4 Flash on Ollama is cloud only. Get the real disk and RAM numbers for every GGUF build, and see which VPS plan can actually load it.
Pick an Ollama quantization with arithmetic instead of guesswork: what q4_K_M, q8_0 and fp16 cost in RAM, and where the quality actually drops.
There is no Qwen 3.8 on Ollama yet. Here is the arithmetic for running the 27B tag that does exist on a CPU-only VPS, and what fits in 8 to 64 GB.