SSD Nodes Learn Hosting plans →
How to do am Matt ConnorBy Matt Connor · Updated 2026-09-08

Ollama q4_K_M vs q8_0 vs fp16: which one better?

Use simple arithmetic choose Ollama quantization: see RAM cost for q4_K_M, q8_0 and fp16, plus where quality really drops before you download.

Wetin Ollama quantization dey change

Ollama quantization dey store every weight for model with fewer bits than the file wey dem train am with. Tag wey end with q4_K_M dey keep about four bits per weight, while fp16 dey keep sixteen. So download go be roughly one quarter of the size, and machine go read one quarter as many bytes to produce each token. Dem dey round the weights onto a coarse grid; dem no throw dem away. For most models, four bits still make answers dey close to how dem behave for full precision.

Na the whole trade-off be this: memory footprint go much smaller and tokens per second go increase, but accuracy go reduce small. The next part explain how you fit predict both sides for one specific model on one specific box, before you spend twenty minutes downloading file wey no go fit.

If Ollama never dey run, start with installing Ollama for VPS. This page assume say ollama ls don already work.

How to read an Ollama quantization tag like q4_K_M

Local models dey ship as GGUF files, wey be the format llama.cpp dey use to store weights for disk. Ollama build on llama.cpp, so Ollama tags carry llama.cpp quantization names unchanged.

The number na the target width. q4 mean say most weight tensors dey packed at four bits each. q8 mean eight. fp16 no be quantized at all: na the model for sixteen bit floating point, wey be the precision most models dey publish with.

K mark a K-quant. Dem group weights into small blocks, and each block store e own scale beside the packed values. Block wey all e weights dey near 0.01 go get fine scale. Block wey hold one large outlier go get coarse one. Na these per-block scales dey make four bit file usable, and dem also be why four bit file no ever be exactly four bits per weight.

The last letter na the mixture. S, M and L decide how many tensors dem promote pass the target width. For q4_K_M, dem store tensors wey rounding dey affect most with wider precision, while most of the rest stay at four bits. Na why q4_K_M dey produce better output than the older q4_0 at almost the same file size.

Ask Ollama wetin e get for disk instead of guessing from the name wey you type:

ollama pull qwen3:8b-q4_K_M
ollama show qwen3:8b-q4_K_M

ollama show dey print architecture, parameters, quantization, context length and embedding length. The quantization line na the ground truth for model wey you pull months ago and no longer remember say you choose.

Bits per weight na where file size dey come from

Every size estimate dey start from one number: how many bits the format dey spend per weight, averaged across the whole file. llama.cpp publish measured figures for Llama 3.1 8B for its quantize documentation, and dem work well for any dense model wey get similar shape.

ChartGGUF quantization types measured on Llama 3.1 8B (published llama.cpp figures)
The data behind this chart
[
  {
    "label": "F16",
    "bits_per_weight": 16,
    "file_gib": 14.96
  },
  {
    "label": "Q8_0",
    "bits_per_weight": 8.5,
    "file_gib": 7.95
  },
  {
    "label": "Q6_K",
    "bits_per_weight": 6.56,
    "file_gib": 6.14
  },
  {
    "label": "Q5_K_M",
    "bits_per_weight": 5.7,
    "file_gib": 5.33
  },
  {
    "label": "Q4_K_M",
    "bits_per_weight": 4.89,
    "file_gib": 4.58
  },
  {
    "label": "Q3_K_M",
    "bits_per_weight": 3.99,
    "file_gib": 3.74
  }
]

The surprise for that table na the second column. Q4_K_M no be four bits per weight. E measure 4.89 bits, because block scales and promoted tensors dey cost real space too. Q8_0 measure 8.5 bits instead of eight, for the same reason. Use the measured number, and the arithmetic go land within a few percent of the real file:

weight bytes = parameter count x bits per weight / 8
8.03e9 params x 4.89 bits / 8 = 4.91e9 bytes = 4.57 GiB

Na the 4.58 GiB Q4_K_M file be this, recovered from two numbers. E also dey close enough to the memory wey the weights occupy after loading. Ollama no unpack anything when e load: the quantized weights dey sit for memory in the same packed form, and each block dey convert as dem use am.

Wetin Ollama actually dey release for each model size

The library dey publish q4_K_M, q8_0 and fp16 tags for most families. Some newer families no follow this pattern. Dem dey show for the library as cloud-only tags, with nothing wey you fit pull for any size. Na this wall you go hit when you dey try run GLM 5.2 for VPS. These na the Qwen3 sizes as of August 2026, based on the tag list for the model page. Every figure below na disk space before e ever enter RAM. Two or three of dem fit fill small VPS root volume, so e make sense to know where Ollama dey store the models wey e download before you start collecting tags.

ChartDownload size of Qwen3 tags in the Ollama library, GB
The data behind this chart
[
  {
    "label": "Qwen3 4B",
    "q4_K_M_gb": 2.6,
    "q8_0_gb": 4.4,
    "fp16_gb": 8.1
  },
  {
    "label": "Qwen3 8B",
    "q4_K_M_gb": 5.2,
    "q8_0_gb": 8.9,
    "fp16_gb": 16
  },
  {
    "label": "Qwen3 14B",
    "q4_K_M_gb": 9.3,
    "q8_0_gb": 16,
    "fp16_gb": 30
  },
  {
    "label": "Qwen3 32B",
    "q4_K_M_gb": 20,
    "q8_0_gb": 35,
    "fp16_gb": 66
  }
]

The default tag important for here. ollama pull qwen3:8b downloads exactly the same 5.2 GB as ollama pull qwen3:8b-q4_K_M, because the tag without suffix na the q4_K_M build. Q4_K_M no be compromise wey the library reluctantly dey offer. Na the default wey upstream choose, so matching am na the sensible first step for any model wey you never test by yourself. Na the same reasoning dey guide the tag choices for running Qwen 3 for VPS.

The ratios dey hold for every row. If you move from q4_K_M to q8_0, e go cost about seventy percent more, instead of exactly double, because the embedding and output tensors no scale the same way as the rest. fp16 na roughly three times q4_K_M. A 32B model for q4_K_M get 20 GB of weights. This don already pass wetin a 16 GB box fit hold with any context window at all. For broader view of which model fit which machine, see which models you fit self-host.

Why KV cache na second, context-dependent cost

Weights na the fixed cost. KV cache (key and value cache) na the variable cost. Every token for the context window keeps e own key and value vectors for every layer, so the cache dey grow in straight line with the window size wey you allow. The system allocate am for the whole window when model load, no be as conversation dey fill am. Na why long window go use memory even when prompt get only one word.

KV bytes per token = 2 (key and value) x layers x kv_heads x head_dim x bytes_per_element
Qwen3 8B at f16:  2 x 36 x 8 x 128 x 2 = 147456 bytes = 144 KiB per token

Those model numbers come from the model own configuration: 36 layers, 8 key/value heads, and head dimension of 128. ollama show gives you the architecture and parameter count, while the model own config.json for Hugging Face gives you the remaining details. Multiply the per-token cost by the window size, and the cache no longer be small rounding error.

Chartqwen3:8b-q4_K_M: KV cache and floor RAM by context length, GB
The data behind this chart
[
  {
    "label": "4k",
    "kv_cache_gb": 0.6,
    "floor_ram_gb": 5.8
  },
  {
    "label": "8k",
    "kv_cache_gb": 1.21,
    "floor_ram_gb": 6.4
  },
  {
    "label": "16k",
    "kv_cache_gb": 2.42,
    "floor_ram_gb": 7.6
  },
  {
    "label": "32k",
    "kv_cache_gb": 4.83,
    "floor_ram_gb": 10
  }
]

For Ollama default window of 4096 tokens, the cache adds 0.6 GB on top of the weights. If you raise the window to 32k, the cache alone reaches 4.83 GB. This one almost match the memory wey quantized weights use, and the minimum memory for the whole model becomes 10 GB. We call am a floor because compute buffers and the operating system still need memory on top. Read the actual figure from the SIZE column of ollama ps after model load.

You set the window on the server, no be per request, when you run Ollama as a service:

OLLAMA_CONTEXT_LENGTH=8192 ollama serve

For systemd install, put am inside a drop-in instead:

sudo systemctl edit ollama
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"

Restart with sudo systemctl restart ollama, then check the CONTEXT column of ollama ps to confirm the window wey the running model actually load with. OLLAMA_KV_CACHE_TYPE fit quantize the cache itself: f16 na the default, q8_0 use about half the memory of f16, while q4_0 use about one-quarter. Na global option this one be, so every model for that server go get the same setting. For small box wey get long window, cutting the cache by half frees more memory than any other single change. Set num_ctx and wetin e cost explains the window itself in detail. The cache too dey size once for every concurrent request slot, no be once for the server. So if you allow Ollama answer two prompts at once, the figure wey you just budget go double. Na this arithmetic dey support how to choose parallel slot count and queue limit.

Wetin fit run for 8, 16, or 32 GB VPS

Budget weights, plus KV cache, plus space wey operating system and anything else wey dey run need. Two GB extra space dey comfortable for small VPS.

8 GB. 4B model for q4_K_M na 2.6 GB, and e leave space for long window. 8B model for q4_K_M fit with default 4k window, but extra space go dey very small. No plan to use 8B with 32k window here, because floor of 10 GB don already pass wetin the machine get.

16 GB. 8B for q4_K_M with 16k or 32k window dey comfortable. 14B for q4_K_M na 9.3 GB of weights, and e fit with modest window. 8B for q8_0 na 8.9 GB, so e fit too. Compare both with your own prompts; na the most useful one hour you fit spend on this matter.

32 GB. 14B for q8_0 (16 GB) and 32B for q4_K_M (20 GB) both fit load. 32B build with large window go push near the memory limit, so monitor ollama ps instead of assuming say e go fit.

Wetin quantization dey spoil first

Quantization error no dey spread equally for all the things wey model dey do. Fluency dey last pass everything, and na why e easy to miss the damage: badly quantized model still fit write clean sentences. Precision dey go first. Exact recall of version number, API signature, or date. Long reasoning chains, where small error for step two fit turn to wrong answer for step eight. Strict output formats, where one wrong bracket fit make tool call fail.

That last one na the practical test. When model must return JSON wey your code go parse, quantization damage go show as parse error instead of prose wey just vaguely worse, so you go see am that same day. Coding agent na the harshest version of this test, because e dey drive model through tool call after tool call. So, pointing agent at your Ollama server go expose over-aggressive quantization within one afternoon.

Below four bits, the loss dey become sharp. The q3 and two-bit types dey for people wey dey squeeze large model onto small hardware. Dem be real option when the alternative na to no run the model at all. But dem no be good default. Between q4_K_M and q8_0, the gap small enough that published perplexity table no go settle which one better for your workload, so no try settle am that way. Run both against thirty of your own prompts and read the output.

When q8_0 or fp16 dey worth the RAM

Pull q8_0 when memory really dey spare and the task no dey tolerate small mistakes: structured extraction, tool calling, and code wey must compile. You dey buy insurance for that case, no be model wey go noticeably smarter.

Pull fp16 for only two reasons. Either you dey quantize the model by yourself and you need the source file, or you dey measure a baseline so you fit know how much your four bit build give up. Serving from fp16 dey use three times the memory of q4_K_M for difference wey most people no fit identify blindly. For CPU only box, e also reduce your token rate to one third.

The stronger rule, with fixed memory budget, be say: bigger model for q4_K_M usually pass smaller model for q8_0. 9.3 GB of 14B weights against 8.9 GB of 8B weights na almost the same RAM (random access memory), and the bigger model sabi more. Test am with your own prompts instead of just trusting am.

CPU-only inference dey limited by memory bandwidth

Most VPS plans no get GPU, so model dey run for system memory on host CPU. Generation come dey bound by memory bandwidth, no be arithmetic, because to produce one token, e need read every weight once. This give one ceiling wey no depend on how many cores you buy.

tokens per second ceiling = memory bandwidth / bytes read per token
50 GB/s / 5.2 GB  =  9.6 tokens/s    qwen3 8B q4_K_M
50 GB/s / 8.9 GB  =  5.6 tokens/s    qwen3 8B q8_0
50 GB/s / 16 GB   =  3.1 tokens/s    qwen3 8B fp16

Fifty GB/s na roughly the theoretical figure for dual-channel DDR4-3200 host. Your own share go smaller, because VPS dey share that bus with every other tenant for the machine. So treat those numbers as ceiling wey nobody dey reach. The useful pattern be say: for CPU, if you halve the bits per weight, token rate roughly go double. Quantization na the biggest speed lever wey you get for machine wey no get GPU. Whether the remaining rate dey tolerable depend on the model, and Nemotron 3.5 Lightning for VPS work through that calculation for one specific build, tag, and RAM figure. The other part of the waiting time na how much text the model decide to write. At ten tokens per second, six hundred-token answer go take full minute. So cap the reply with num_predict often fit save more waiting time than another step down for precision.

Prompt processing dey behave different. Reading long prompt dey bound by compute instead of bandwidth, so extra cores help for there but do almost nothing for generation speed. If machine fit ingest 4k prompt quickly and then generate slowly, na normal behaviour.

No just trust any of the arithmetic. Measure tokens per second for your own machine with the same prompt for each quantization, and allow your own numbers override these ones.

Quantize model by yourself

Ollama fit build quantized model from fp16 or fp32 source. This dey useful when you don fine-tune something and no library tag dey available. Point a Modelfile to the unquantized weights:

FROM /path/to/my/model/f16

Then build am and confirm:

ollama create --quantize q4_K_M mymodel
ollama show mymodel

--quantize accepts q8_0, q4_K_S and q4_K_M. No q6_K or q5_K_M option dey here. So, for those ones, use llama.cpp own tool to quantize am, then import the completed GGUF file. This import method get one problem too: chat template mismatch fit make the model reply with nonsense. Importing GGUF file into Ollama explain how to handle am. The quantization line from ollama show na how you verify say the build do wetin you request.

Wetin you go see when e go wrong

Everything dey run for CPU when you expect GPU. Read the PROCESSOR column:

ollama ps

E go print 100% GPU, 100% CPU, or split like 48%/52% CPU/GPU. Split mean say the weights plus KV cache no fit inside VRAM (video RAM, the memory for the graphics card), so part of the model enter system memory. Speed go then fall near CPU-only rate, because every token dey wait for the slow half. Reduce the context window, quantize the cache, or use smaller build. Adding more cores no go help.

The model dey get killed while e dey load. Check the kernel and service log:

sudo dmesg -T | grep -i "out of memory"
sudo journalctl -u ollama -n 50

Line wey contain Out of memory: Killed process mean say total size of weights, KV cache, and buffers pass the memory wey dey for the box. For VPS wey no get swap configured, the whole machine fit hang for some seconds before that line show.

Answers don worse even though you change nothing. Two builds of the same model fit dey side by side for ollama ls under different tags, and script wey pull the unsuffixed name go follow whichever model the library now point to. Run ollama show against the exact tag wey your client request, then read the quantization line instead of trusting the name for your configuration file.

FAQ

Which Ollama quantization I suppose pull?

Start with q4_K_M. Na wetin Ollama library dey ship as the default tag for most models, so ollama pull qwen3:8b and ollama pull qwen3:8b-q4_K_M go fetch the same file. Move go q8_0 only when memory still plenty and the task no tolerate small errors, like tool calling or structured JSON output. When memory budget fixed, bigger model for q4_K_M usually dey perform better than smaller model for q8_0, so test that pairing before you use RAM for more precision.

q4_K_M really mean four bits per weight?

No. When dem measure am for Llama 3.1 8B, e be 4.89 bits per weight, because every block of weights store its own scale, and the tensors wey sensitive pass dey use wider type. Q8_0 measure 8.5 bits instead of eight for the same reason. Use the measured value when you dey estimate: parameter count times bits per weight, divided by eight, go give the file size for bytes.

How much RAM 8B model need for CPU-only VPS?

Add weights, KV cache, and headroom. Qwen3 8B for q4_K_M get 5.2 GB of weights. For the default 4096 token window, the cache add 0.6 GB, so the minimum dey near 5.8 GB before compute buffers and the operating system. For 32k window, the cache alone na 4.83 GB. Plan for 8 GB if na short window, and 16 GB if you want long one.

Why my model dey use 100% CPU when the box get GPU?

Run ollama ps and read the PROCESSOR column. 100% CPU, or split like 48%/52% CPU/GPU, mean say the weights plus KV cache no fit inside VRAM, so Ollama put part or all of the model for system memory. The usual cause na context window wey pass wetin the card fit hold, because the cache dey allocate for the whole window when the model load. Reduce the window with OLLAMA_CONTEXT_LENGTH, set OLLAMA_KV_CACHE_TYPE=q8_0 to halve the cache, or pull smaller quantization.