Ollama KV cache quantization: q8_0 vs f16
Raised num_ctx and Ollama no longer fits? The KV cache grew, not the weights. Set OLLAMA_KV_CACHE_TYPE with flash attention and watch the change in ollama ps.
Filtering by topic #kv-cache · clear
Raised num_ctx and Ollama no longer fits? The KV cache grew, not the weights. Set OLLAMA_KV_CACHE_TYPE with flash attention and watch the change in ollama ps.
The KV cache is RAM your server pays for on every request. Prompt caching is a bill someone else discounts. Only one can break a model load.