Ollama num_predict: How to Limit Token Output
Learn how Ollama num_predict limits output tokens, which of the three settings wins, and how done_reason shows when generation stops at the cap.
num_predict dey do wetin for Ollama
num_predict na Ollama option wey dey limit how many tokens model fit generate for one response. E dey count output tokens only, so prompt no dey count inside am. When model reach the limit, generation go stop for where e dey, sometimes for middle of word, and response go come back with done_reason set to length.
Na everything this feature dey do. The hard part be say Ollama give you three different places to set the value, and setting wey dey closest to the request go win. Almost every report say “num_predict no dey do anything” happen because one layer dey quietly override another one.
num_predict no be num_ctx
People dey confuse these two options pass any other pair for Ollama, and this confusion dey waste real debugging time.
num_ctx na how much the model fit read. Na the size of the context window, wey dey hold the prompt plus everything wey e don produce so far. If you increase am, e go use more memory, because the key/value cache wey model dey keep for those tokens go grow with the window. How to size num_ctx for your hardware na separate work, with its own ways to fail.
num_predict na how much the model go write. Na stopping rule, no be allocation. If you increase am, e go take more wall-clock time instead of RAM, and nothing dey reserve ahead of time.
Dem meet for one place. As model dey produce tokens, the generated tokens enter the context window, so reply fit stop because the window don full instead of because e reach your cap. Ollama dey report length for both cases, so the number wey separate dem na eval_count, and we cover am further down.
Set am one time with a Modelfile
A Modelfile dey bake the value inside the model wey you create. Write the file:
FROM qwen3:8b
PARAMETER num_ctx 8192
PARAMETER num_predict 512Then build am and read back wetin you build:
ollama create qwen3-capped -f Modelfile
ollama show --parameters qwen3-cappedollama show --parameters dey print one line for each stored parameter together with the value. If num_predict no dey for that output, the model no get baked-in cap, and Ollama own default go apply. ollama show --modelfile qwen3-capped dey print the complete definition. This one na also the fastest way to copy the parameters wey existing model already ship with. Creating capped model this way dey use almost no extra disk space, because the new entry dey reuse the weight blobs wey the base model don already download instead of copying dem. You suppose know where Ollama dey keep those blobs before VPS root disk go fill up.
This na the correct layer for value wey you want every caller to inherit. But e no be the correct layer if you expect am to be final, because e no be final.
Set am for each request inside the options object
Every generation endpoint dey take an options object, and num_predict dey inside am:
curl http://localhost:11434/api/generate -d '{
"model": "qwen3:8b",
"prompt": "Explain what a reverse proxy does.",
"stream": false,
"options": { "num_predict": 128 }
}'/api/chat dey use the same options key with the same meaning. Any value wey you put here go apply only to that call and nothing else. Na this layer your tools dey use: chat front end, script, SDK wrapper, coding agent. All of dem dey send an options object, whether dem show you a box for am or not.
Set am for one session with /set parameter
Inside ollama run, the interactive session dey set options for the remaining part of that session:
>>> /set parameter num_predict 256
>>> /show parameters/show parameters go print wetin the session go send with your next message. This one na the fastest way to confirm say the change don take effect. The value go remain until you type /bye. To keep am, /save qwen3-capped go write the current session, including the parameters, as new model. Nothing wey you /set here go reach any other client.
Which setting dey win, and why e look like say dem ignore your own
The order short. Options wey request send dey beat everything else. A PARAMETER num_predict line for the model's Modelfile na fallback wey e go use when request no carry value. If both no dey, Ollama built-in default go apply.
/set parameter no be third rule. Interactive session na API client, so wetin you set there go dey sent as that request's options. Na exactly why e dey override Modelfile for the session.
Now, na this explain the failure. You add PARAMETER num_predict 512, rebuild the model, but replies still dey run reach thousands of tokens. Your setting dey present, and ollama show --parameters prove am. Every request dey override am because client dey send im own options object wey carry im own number. Often, na number wey you type for settings screen months ago and forget. ollama show dey read the stored model. E no fit show wetin arrive through HTTP.
Prove the server side with one command. Send request wey go produce long answer, force the cap low, then read two fields:
curl -s http://localhost:11434/api/generate -d '{
"model": "qwen3-capped",
"prompt": "Describe the Linux boot process in detail.",
"stream": false,
"options": { "num_predict": 32 }
}' | jq '.done_reason, .eval_count'That one suppose print "length" and 32. Install jq first with sudo apt install -y jq if e no dey. Response of "length" and 32 mean say server dey honour the option, while your application dey send different value. To see the server own record of request, restart am with OLLAMA_DEBUG=1 for the environment, then monitor journalctl -u ollama -f while your application dey talk to am.
Negative values, and numbers wey you no suppose copy
num_predict still dey accept negative values, and dem dey serve as sentinel values instead of counts. One negative value mean “no cap am, continue to generate”. Another one don mean “fill the remaining context”. As of August 2026, Ollama Modelfile reference dey give the default as -1, meaning infinite generation. Older versions of the same table also list -2 for filling the context.
Treat all these as version-dependent, because dem don change before. The reference document the default as 128 for long time before dem correct the entry at the end of 2024. Na why plenty guides still dey repeat the old number. Read the Modelfile parameter reference for the version wey you dey run, then confirm the behaviour with the eval_count check above. Value wey you verify for your own system better pass value wey you read anywhere, including this post.
Why output length na the main cost for CPU-only VPS
Generation get two phases wey dey run for very different speeds. Prompt tokens dey evaluate in batches, many at once. Output tokens dey produce one by one, and each one need full pass through the model weights. For CPU-only VPS, memory bandwidth dey limit that pass, so generating one token cost far more than processing one prompt token. Because that pass must read every weight, the number of bytes wey each weight occupy set the upper limit for your token rate. Na why q4 build dey decode faster than the same model for q8 or fp16.
Ask for response without streaming, and you go see the numbers directly:
"prompt_eval_count": 26,
"prompt_eval_duration": 107345000,
"eval_count": 237,
"eval_duration": 4289432000Durations dey for nanoseconds. For that block, wey be sample response published for Ollama API documentation instead of measurement from any particular server, 26 prompt tokens take about 0.1 seconds, while 237 output tokens take about 4.3 seconds. Your own generation rate na eval_count divided by eval_duration, then convert am to seconds. Measuring tokens per second for your own hardware worth doing once before you tune anything else. That rate depend on the model as much as the machine. So if long answers na the real cost, model wey dem build for fast decoding, like Nemotron 3.5 Lightning for VPS, fit recover some of the time wey low cap dey otherwise save.
The arithmetic go show the rest. For 8 tokens per second, 2,000-token answer go hold the machine for more than four minutes. The model no know say you want just one paragraph. Reasoning model dey spend part of that budget to think before e write the word wey you request. That thinking dey generate one token at a time like everything else. So the reasoning effort wey you request na another control for the same bill. Some models fit also loop and repeat phrase until something stop dem. Without cap, that single request go keep one core busy until the context window finish. num_predict na the setting wey limit am. This matter most for small self-hosted Ollama VPS, where one long request fit occupy the whole machine.
Output wey truncate na usually cap, no be say model spoil
The signs fit look like model failure. Answer wey stop for middle of sentence. JSON wey no go parse because closing brace no show. Instinct na to blame the model or the quantisation. First read the response.
done_reason mean say answer address the question directly. stop mean say model finish by itself, either e emit end-of-sequence token or e match one of the strings for your stop option. length mean say generation stop because space finish. When you see length, compare eval_count with your cap: if dem match exactly, na num_predict stop am; if the number smaller, context window fill first.
When you stream, those fields go arrive for final chunk, the one wey carry "done": true. Plenty client libraries discard that chunk and give your code only the text. Na why the same truncation fit look unexplained inside application but obvious under curl. If library dey hide am, send one request with curl to know wetin server really talk.
One more point fit save wasted afternoon. Increasing num_predict no make model write more. E only remove ceiling. If reply end for 200 tokens with done_reason of stop, model don decide say e finish, and bigger cap no go change anything. Short answers with stop na prompting problem. Short answers with length na cap problem.
Wetin value to choose
- For interactive chat, leave am uncapped and press Ctrl+C to stop reply wey no dey end. You dey watch the screen anyway.
- For anything scripted, set am. If generation wey no get cap dey inside loop, batch job wey suppose take ten minutes fit still dey run the next morning.
- For structured output, set the cap pass the biggest valid document wey you expect. Then treat
done_reasonoflengthas hard error and retry, instead of parsing wetin come back. - For coding agent, the value belong for the agent own configuration, because the agent dey send im own options for every request. How to point coding agent to Ollama explain where those settings dey.
The cap dey count tokens, no be words and no be characters, so no estimate am. Generate one representative answer without cap, read eval_count, then set the limit comfortably above that. Model families dey tokenise differently, so value wey fit Llama model fit truncate the same answer from Qwen 3 model for the same VPS.
FAQ
Wetin be difference between num_ctx and num_predict for Ollama?
num_ctx na size of the context window, so e set how much model fit read: the prompt plus everything wey e don produce so far. E dey use memory, because key/value cache dey grow with am. num_predict set how many tokens model fit write for one response. E dey use time instead of memory, and nothing dey reserve ahead of time. Generated tokens count against both, so either one fit cut reply short.
Why num_predict setting dey look like say e no dey work?
Because value wey request send go override value wey dey stored for model. Put PARAMETER num_predict 512 for Modelfile, then use chat front end or coding agent take drive that model, and client go send im own options object. Na that number go win. ollama show --parameters still print your value, because e dey read stored model and e no fit see wetin arrive through HTTP. Send one request with curl using "options": {"num_predict": 32} and check say eval_count come back as 32. This confirm say server itself dey behave well and move the search go your application.
How I fit know whether num_predict cut my output short?
Send request with "stream": false and read done_reason. Value of stop mean say model finish by itself. Value of length mean say e run out of space. Then compare eval_count with your cap: if dem match exactly, num_predict stop am, and if eval_count small pass, context window fill first. When streaming, both fields go arrive for final chunk with "done": true, but many client libraries dey throw am away before your code see am.
Wetin be default value of num_predict?
Read am from your own install instead of article. As of August 2026, Ollama Modelfile reference give default as -1. This mean generation no get cap, and dem correct that entry at the end of 2024 after years of documenting 128. Negative values na sentinels, no be counts, and older versions of the same table also list -2 for filling the remaining context. Check the Modelfile parameter reference for your version, then confirm am with ollama show --parameters and one curl request.
Raising num_predict go make model write longer answers?
No. E only remove ceiling. If reply end with done_reason of stop, model decide say e don finish, and bigger cap no go change anything. Length for that case na prompting matter: ask for specific structure, section count, or clear level of detail. Raise num_predict only when done_reason come back as length.