Ollama num_ctx: How to Stop Prompt Truncation
Ollama fit quietly cut long prompts because default context fit be 2,048 or 4,096 tokens. Set num_ctx per request or server-wide, then check KV cache RAM.
num_ctx dey do wetin, and why dem cut your long prompt
Ollama context length na number of tokens wey loaded model fit hold for memory at once, and num_ctx na the option wey set am. Ollama dey choose default wey far below the maximum wey model advertise, so dem go cut longer prompt before model even read am. Nothing for response go tell you say e happen.
Ollama model library list Llama 3.1 8B with 128k context window. Stock server no go give you that. Ollama own documentation get different defaults for different pages: FAQ talk say na 4096 tokens, Modelfile reference talk say num_ctx default na 2048, and the context length page talk say default dey come from available VRAM (video RAM): 4k below 24 GiB, 32k from 24 to 48 GiB, and 256k above that. Each value dey correct for some build. The useful lesson from this disagreement be say make you read the value from your own running server instead of trusting any page, including this one.
Truncation dey happen quietly because model still dey answer, and the answer still dey read well. E write am from the tail of your input. Summary wey miss the first half of document fit look like say model weak. Most times, na small context window cause am.
Check the Ollama context length wey your server really apply
The check wey work for any build na prompt_eval_count, wey be the number of prompt tokens wey server report say e process. Send prompt wey pass the context capacity, and that number go stop for the limit.
sudo apt update && sudo apt install -y jq
LONG=$(python3 -c "print('the quick brown fox jumps over the lazy dog. ' * 2000)")
jq -n --arg p "$LONG" '{model:"llama3.1:8b", prompt:$p, stream:false, options:{num_ctx:4096}}' |
curl -s http://localhost:11434/api/generate -d @- |
jq '{prompt_eval_count, prompt_eval_duration}'That prompt get about 18,000 words, so e far pass 4096 tokens. prompt_eval_count go return near 4096 instead of near the real token count, because server drop the remaining part. Run am again with "num_ctx":16384 and the count go increase. If your build return error instead of truncating, na the same finding be that, but the signal go louder.
ollama psThe CONTEXT column, for builds wey print am, get the context length wey the loaded model dey use now. The PROCESSOR column next to am show where the model dey run. 100% CPU normal for VPS wey no get GPU. A split like 30%/70% CPU/GPU for GPU box mean say the weights plus the cache no fit again inside VRAM, and raised num_ctx na the usual reason.
journalctl -u ollama --no-pager | grep -i n_ctx | tail -n 5The inference runner print its context size for one line wey contain n_ctx. The exact wording dey change between releases, so treat missing line as a rename, no be proof of anything.
Places four wey you fit set num_ctx
For the request. Send "options": {"num_ctx": 16384} to /api/generate or /api/chat. This one override every other setting, and e apply only to that call. If the value no match wetin the loaded model dey use, the server go reload the model first. You fit see this for load_duration inside the response: e go jump from almost zero to complete seconds. The same wait go show when the model don stay idle reach the point wey dem unload am. So, after you don choose context size, e make sense to keep the model resident with keep_alive.
For the interactive session. Inside ollama run, type /set parameter num_ctx 16384. E go last for that session.
For a Modelfile. This one bake the value inside a named model, so every client go get am without any change for the client side.
FROM llama3.1:8b
PARAMETER num_ctx 16384ollama create llama3.1-16k -f ./Modelfile
ollama run llama3.1-16kFor the server. OLLAMA_CONTEXT_LENGTH set the default for every request wey no carry its own num_ctx. Under systemd, add a drop-in instead of editing the unit file.
sudo systemctl edit ollama.service[Service]
Environment="OLLAMA_CONTEXT_LENGTH=16384"sudo systemctl daemon-reload
sudo systemctl restart ollama
ollama psThe order of priority matter most when you dey debug another person client. Request wey carry num_ctx go override the server default. So, chat front end or agent wey send small value by itself fit quietly cancel the systemd change wey you make. When you point a coding agent to your Ollama server, check wetin the client dey send before you blame the server.
Why you no fit just set num_ctx to the model maximum
Attention dey make every token look at every token wey come before am. E keep the keys and values wey e calculate for earlier tokens, so e no need calculate dem again for every new token. This storage na KV cache (key/value cache). E allocate am for the whole of num_ctx when model load, not as conversation dey grow. So large context still use plenty memory even for one-line prompt.
DigitalOcean's inference cost tutorial show the calculation for one line:
kv_bytes_per_token = 2 * layers * kv_heads * head_dim * bytes_per_valueThe 2 dey count keys and values separately. Read the other numbers from your own model.
curl -s http://localhost:11434/api/show -d '{"model":"llama3.1:8b"}' |
jq '.model_info | {ctx: ."llama.context_length", layers: ."llama.block_count", heads: ."llama.attention.head_count", kv_heads: ."llama.attention.head_count_kv", embed: ."llama.embedding_length"}'Llama 3.1 8B report 32 layers and 8 key/value heads. The head dimension na embed divide by heads, so 4096 / 32 = 128 for this case. Some models publish am directly as llama.attention.key_length. The default cache dey hold f16 values, so bytes_per_value na 2. Therefore, 2 32 8 128 2 = 131,072 bytes. That na 128 KiB cache for every single token of context. Multiply am by the context length, and the cost no go remain abstract.
The data behind this chart
[
{
"label": "4k",
"kv_cache_gib": 0.5,
"total_ram_gib": 5.1
},
{
"label": "8k",
"kv_cache_gib": 1,
"total_ram_gib": 5.6
},
{
"label": "16k",
"kv_cache_gib": 2,
"total_ram_gib": 6.6
},
{
"label": "32k",
"kv_cache_gib": 4,
"total_ram_gib": 8.6
},
{
"label": "64k",
"kv_cache_gib": 8,
"total_ram_gib": 12.6
},
{
"label": "128k",
"kv_cache_gib": 16,
"total_ram_gib": 20.6
}
]Those 6 rows na calculations from the formula above, not measurements. The total column add the 4.9 GB download wey the Ollama library list for llama3.1:8b for August 2026. That na 4.6 GiB. E no include the compute buffers and the server process itself. Treat am as the minimum.
The shape na the main point. For 8k, the cache cost 1 GiB. This small compared with the weights. For the model full 128k, e cost 16 GiB. That pass three times the weights, giving total near 20.6 GiB. So 4 GB VPS no fit load this model with any useful context. 8 GB VPS dey comfortable at 8k. 16 GB VPS fit reach 32k and still get space for the rest of the box. All these limits dey increase as the weights increase. So if you dey compare a larger model with this 8B, the same calculations for Qwen's 27B tag on CPU-only VPS show how small the space wey the weights leave for context between 8 and 64 GB.
Wetin go happen when KV cache no fit
For VPS wey na CPU-only, the process go simply grow. Watch am while model dey load and while long request dey run.
free -m
ps -eo rss,comm --sort=-rss | head -n 5RSS (resident set size) dey print for kilobytes. If used swap for free -m start to increase, reduce the context. KV cache wey dey live for swap go make generation pause for seconds for each token, because every new token dey read the whole cache.
If the box run out of memory completely, kernel go pick the biggest process and kill am.
sudo dmesg | grep -i "killed process"Line wey read Out of memory: Killed process 1234 (ollama) mean say the context wey you request no fit. Ollama often go refuse before e reach that point, then the request go fail with message wey name the memory wey e need compared with the memory wey dey free.
For GPU box, the failure no too obvious. Layers go spill enter system RAM, ollama ps go show the CPU and GPU split, and throughput go drop sharply. How sharp the drop go be depend on your hardware, so measure tokens per second for your own box for every context setting instead make you trust figure from another person machine.
Prefill time dey grow faster than the prompt
Prefill na the work wey happen for your input before the first output token show. Each prompt token dey attend to every token before am, so the total work dey grow with the square of the input length. If you double the prompt, the wait for the first token go more than double.
The response carry the measurement, so you no need just trust am.
jq -n --arg p "$LONG" '{model:"llama3.1:8b", prompt:$p, stream:false, options:{num_ctx:16384}}' |
curl -s http://localhost:11434/api/generate -d @- |
jq '{tokens: .prompt_eval_count, prefill_seconds: (.prompt_eval_duration/1000000000)}'Run am with short prompt, then run am again with long one. For each run, divide tokens by seconds. For CPU-only VPS, prefill usually na the slowest part of long-context request. So, tokens per second wey you measure with short prompt no go predict long prompt performance.
If prefill take pass the timeout wey dey in front of am, long prompt fit return context deadline exceeded instead of answer. Na this one usually cause the problem. Before you reduce context size, find out which layer stop waiting and give up.
Concurrency na where this matter pass. Every request wey service dey handle need its own cache, so the memory for the chart above na per request, no be per server. One long request fit hold the whole box while short ones queue behind am. Set OLLAMA_NUM_PARALLEL deliberately, and read how many concurrent users one self-hosted LLM fit serve before you increase both numbers together.
Buy cache wey smaller to free memory
bytes_per_value for the formula na setting wey you fit control. Ollama FAQ document OLLAMA_KV_CACHE_TYPE, with f16 as default for 2 bytes, plus q8_0 for 1 byte and q4_0 below that. If you move to q8_0, the cache go reduce by half, so the 32k row go cost 2 GiB instead of 4 GiB. Quantising the weights go free memory from the other side of the same budget, and the GLM tag wey actually fit VPS dey explained quantisation by quantisation if na that trade-off you prefer. The same FAQ document OLLAMA_FLASH_ATTENTION=1, wey some builds need before quantised cache go take effect.
[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"Confirm am instead of assuming: restart the service, load the model with the same num_ctx as before, then compare RSS. Support depend on the model and backend, so if setting no change anything, e mean say una combination no dey covered. The documentation list these options but no promise any quality result, so test q4_0 with your own prompts before you rely on am. If na these knobs bring you here, Ollama and llama.cpp expose dem differently.
Recipe for how to choose num_ctx
- Read the model maximum context, layer count, and key/value head count from
/api/show. - Calculate bytes per token with the formula, then multiply am with the context wey you want.
- Add the weight size, compare am with free RAM, and leave at least 1 GiB for the rest of the server.
- Set the value, load the model, then confirm wetin apply with
ollama psandprompt_eval_count. - Run your real workload while you dey monitor
free -m, and cut the context by half if swap start to move.
Most jobs no need as much context as people dey give dem. You fit summarise one long report with 16k. Retrieval front end wey dey paste five document chunks rarely pass 8k. Coding agent wey dey read complete files na the case wey really need 64k or more. Na also this case make you size the machine around the context, instead of doing am the other way round. If the server still new, start from working Ollama installation for VPS and tune the context after models dey load cleanly.
FAQ
Ollama default context length na wetin?
E depend on the build and the hardware, so check am instead of assuming. Ollama FAQ document 4096 tokens, the Modelfile reference document a num_ctx default of 2048, and the context length page document a default wey e pick from available VRAM: 4k below 24 GiB, 32k from 24 to 48 GiB, and 256k above. CPU-only VPS go fall for the small end. ollama ps go print the applied context for builds wey get the column, and prompt_eval_count for an API response prove am for every build.
Why Ollama dey ignore the beginning of my long prompt?
Because the prompt long pass the context window, so the server cut am before the model see am, and no error come back. Send the same prompt again with bigger num_ctx and monitor prompt_eval_count as e grow for the response. If that number no move, something between you and the server dey set num_ctx by itself. Chat front ends and agent frameworks commonly cause this.
How much extra RAM larger num_ctx need?
Multiply the context length by the cache cost per token, wey be 2 * layers * kv_heads * head_dim * bytes_per_value. For Llama 3.1 8B at f16, na 128 KiB per token. So 32k tokens cost 4 GiB, and the full 128k cost 16 GiB on top of the weights. The cache dey allocate when the model load, so large num_ctx go use that memory even when your prompts remain short.
Larger context window dey make Ollama slower?
Yes, for two ways. Prefill work dey grow with the square of the prompt length, so long input dey delay the first token pass wetin the length suggest. The larger cache also dey compete for memory. For GPU box, e fit push layers enter system RAM. For CPU box, e fit push the machine toward swap. Large num_ctx wey you never fill still dey use the memory, but e no dey add prefill time.
I fit set num_ctx permanently for one model?
Yes. Write a Modelfile wey contain FROM llama3.1:8b and PARAMETER num_ctx 16384, then run ollama create llama3.1-16k -f ./Modelfile. Every client wey ask for llama3.1-16k go get that context without sending any options. Request wey carry its own num_ctx still get priority, so this one set default, no be ceiling.