SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor · Updated 2026-09-26

Ollama temperature and sampling settings

Temperature does not make a model smarter. Here is what it does to the token distribution in Ollama, plus top_p, top_k, min_p, seed, and which place wins.

What Ollama temperature actually does to the token distribution

Ollama temperature is a divisor applied to the model's raw scores before those scores become probabilities. At every step the model emits a logit, a raw unnormalised score, for each token in its vocabulary. Ollama divides every logit by the temperature, then runs softmax, the function that turns a list of scores into probabilities that add up to 1. A small temperature widens the gaps between logits, so the leading token takes almost all the probability. A large temperature shrinks the gaps, so weak candidates get a real chance. One token is then drawn at random from whatever distribution came out.

That is the entire mechanism. Temperature adds no knowledge and buys no extra reasoning. It decides how much of the model's own uncertainty is allowed to reach the page.

The arithmetic is small enough to show in full. Take five candidate tokens with logits 4.0, 3.2, 2.8, 1.5 and 0.4, and run the softmax at four temperatures. These are computed numbers from one made-up set of logits, not a measurement of any model, so you can reproduce every cell yourself with a few lines of Python.

ChartProbability (%) of each candidate token at four temperatures
The data behind this chart
[
  {
    "label": "logit 4.0",
    "temp_0_2": 97.96,
    "temp_0_7": 65.23,
    "temp_1_0": 53.77,
    "temp_1_5": 43.19
  },
  {
    "label": "logit 3.2",
    "temp_0_2": 1.79,
    "temp_0_7": 20.8,
    "temp_1_0": 24.16,
    "temp_1_5": 25.33
  },
  {
    "label": "logit 2.8",
    "temp_0_2": 0.24,
    "temp_0_7": 11.75,
    "temp_1_0": 16.19,
    "temp_1_5": 19.4
  },
  {
    "label": "logit 1.5",
    "temp_0_2": 0.0,
    "temp_0_7": 1.83,
    "temp_1_0": 4.41,
    "temp_1_5": 8.16
  },
  {
    "label": "logit 0.4",
    "temp_0_2": 0.0,
    "temp_0_7": 0.38,
    "temp_1_0": 1.47,
    "temp_1_5": 3.92
  }
]

At temperature 0.2 the leading token holds 97.96 percent of the probability mass and the runner-up gets 1.79 percent. The draw is still random, but the outcome is close to fixed. At temperature 1.5 that same leading token is down to 43.19 percent, and the weakest of the 5 candidates climbs from 0.0 percent, where it rounds to zero, up to 3.92 percent.

Now remember this happens once per token. A 1-in-25 chance of a strange word is nothing on a single draw. Across a 600 token answer it is close to certain, and one odd token steers everything after it, because the next step conditions on the text already written. That is why a high temperature does not read as slightly more creative. It reads as an answer that drifts.

Setting temperature to 0

Set temperature to 0 and the draw collapses onto the highest-scoring token at every step. The tail filters below stop mattering, because there is no tail left to cut. Check this on your own install rather than trusting the paragraph. Substitute a model you actually have; ollama ls lists them.

curl -s http://localhost:11434/api/generate -d '{
  "model": "qwen3",
  "prompt": "Write a Python function that reverses a string.",
  "stream": false,
  "options": {"temperature": 0}
}' | python3 -c 'import sys,json; print(json.load(sys.stdin)["response"])'

Run it twice and compare the two outputs. They should match character for character. If they do not, something else is moving underneath you, and the seed section below lists the usual causes.

top_k, top_p and min_p remove candidates before one is drawn

Temperature reshapes the distribution. These parameters cut it down instead. Ollama's own interactive help is the shortest accurate description of each, and you can print it yourself: start a session with ollama run qwen3 and type /set parameter with nothing after it.

/set parameter top_k <int>            Pick from top k num of tokens
/set parameter top_p <float>          Pick token based on sum of probabilities
/set parameter min_p <float>          Pick token based on top token probability * min_p

top_k keeps a fixed number of the highest-scoring tokens and discards the rest. It is blunt because the right number changes at every step. After def there are thousands of reasonable identifiers. After import numpy as there is really one.

top_p, usually called nucleus sampling, keeps the smallest group of tokens whose probabilities add up to p. It adapts to the step, which is why it is generally preferred over top_k. Look at the chart above to see its limit: with the leading token at 97.96 percent, a top_p of 0.9 is already satisfied by that one token, so the group is that token alone.

min_p sets a floor relative to the leading token instead of an absolute budget. With min_p at 0.05 and a leading token at 60 percent, every candidate under 3 percent is discarded. It cuts hard on a confident step and leaves the field open on an uncertain one, which is the behaviour you want for open-ended writing.

min_p is not present on every Ollama build. That /set parameter help text is the check: if min_p is not in the list your build prints, do not rely on it, and do not assume a PARAMETER min_p line in a Modelfile is doing anything.

Use one tail filter at a time. Stacking two of them makes the effect impossible to reason about, because you can no longer tell which one removed the candidate you were hoping for.

repeat_penalty and repeat_last_n

repeat_penalty reduces the score of tokens that have already appeared recently, so the sampler is less likely to pick them again. repeat_last_n sets how many recent tokens count as recently. Ollama documents two special values for the window: 0 disables it, and -1 makes the window the whole context. A repeat_penalty of 1.0 means no penalty at all, because the penalty is a multiplier.

This pair is dangerous for code, and the reason is mechanical. Correct code repeats tokens constantly: the same closing brace, the same indentation, the same variable name on ten consecutive lines, the same import prefix. A repeat penalty above 1.0 pushes against all of it. Set repeat_last_n to -1 as well and the window covers the entire file you pasted in, so every identifier already on screen is penalised. The symptoms are specific and easy to recognise. The model renames a variable halfway through a function, drops a closing bracket, or switches from four spaces to two. For code, tool calls and structured output, leave repeat_penalty at 1.0.

How much text "the whole context" covers depends on the context window you configured with num_ctx, so the same penalty setting behaves very differently on a 4k and a 32k model.

If a model loops the same sentence forever, raising repeat_penalty hides the symptom rather than fixing it. Looping is normally a sign that the model is too small or too heavily quantized for the prompt you gave it.

What seed does, and what it does not guarantee

seed fixes the random stream the sampler draws from. Same seed, same prompt, same model, same options, same build, same hardware path: same output.

curl -s http://localhost:11434/api/chat -d '{
  "model": "qwen3",
  "messages": [{"role": "user", "content": "Describe one use for a Raspberry Pi."}],
  "stream": false,
  "options": {"seed": 101, "temperature": 0.8}
}'

Change any one of the following and the guarantee is gone.

  • A different quantization of the same model. Different weights produce different logits, so the same seed lands somewhere else.
  • A different Ollama version, or a build using different compute kernels.
  • A different split of layers between GPU and CPU. Moving a layer changes the order floating point numbers are added in.
  • Another request in flight against the same model. The batch shape changes, and so can the last decimal of a logit.
  • A different num_ctx, or any change to the prompt prefix, including the system message.

The mechanism behind most of those is the same: floating point addition is not associative, so summing the same numbers in a different order can change the last decimal. That almost never matters, because the leading token is usually far enough ahead that a tiny difference cannot overtake it. It matters at the steps where two candidates are nearly tied. One different token at step 40 gives you a different paragraph from step 40 onward, which is why the failure looks dramatic when it does happen.

If you need repeatable comparisons, pin concurrency to one request per model while you test.

sudo systemctl edit ollama.service
[Service]
Environment="OLLAMA_NUM_PARALLEL=1"
sudo systemctl restart ollama
systemctl show ollama --property=Environment

The last command should list OLLAMA_NUM_PARALLEL=1. If it prints an empty Environment line, the drop-in did not save.

One thing seed never does: it does not make an answer correct. It makes a wrong answer repeatable, which is still the most useful property in this whole guide, because it is the only way to change one parameter and know that the difference you see came from that parameter.

The three places a value can be set, and which one wins

Most confusion about Ollama sampling is not a misunderstood parameter. It is a value set in one place and expected in another. Ollama builds the options for a request in layers, and each layer overwrites the one before it.

  1. The built-in defaults of your Ollama build.
  2. The generation defaults the model itself ships with.
  3. The PARAMETER lines in the Modelfile that created the model you are calling.
  4. The options object on the /api/generate or /api/chat request.

The request wins. A Modelfile sets a default, not a lock, and there is no setting that stops a caller from overriding it.

Layer 3 looks like this. Write a file called Modelfile:

FROM qwen3
PARAMETER temperature 0.2
PARAMETER top_p 0.9
PARAMETER repeat_penalty 1.0
PARAMETER num_ctx 8192
ollama create qwen3-precise -f ./Modelfile
ollama show --parameters qwen3-precise

ollama show --parameters should print the four lines you wrote. If one is missing, ollama create did not read the file you think it did.

Layer 4 overrides it without warning:

curl -s http://localhost:11434/api/chat -d '{
  "model": "qwen3-precise",
  "messages": [{"role": "user", "content": "Refactor this function."}],
  "stream": false,
  "options": {"temperature": 0.9}
}'

That request runs at 0.9, not at the 0.2 in the Modelfile, because the options object is applied last. Nothing logs a conflict.

The interactive session is layer 4 wearing a different hat. /set parameter fills in the options object that the CLI attaches to its own requests.

>>> /set parameter temperature 0.9
>>> /show parameters
>>> /save qwen3-creative
>>> /bye

/show parameters prints two sections: model defined parameters, and user defined parameters. The second section holds what you just typed, and it exists only inside this session. /bye throws it away. /save writes the current session's parameters and system message into a new model, which gets you the same result as editing a Modelfile without leaving the prompt.

There is a fourth entry point worth knowing about, because it is the one that surprises people. Ollama also serves an OpenAI-compatible endpoint at /v1/chat/completions, and that is what most agents and SDKs speak. As of September 2026 that endpoint accepts temperature, top_p, seed, frequency_penalty, presence_penalty, max_tokens and stop. It has no field for top_k, min_p or repeat_penalty. So a client on /v1 can raise your temperature past anything your Modelfile said, and cannot touch your min_p at all. If you are pointing a coding agent at your own Ollama server, put the parameters that endpoint cannot send into the model itself with a Modelfile, and let the client own only temperature and top_p.

Read back what your install is actually using

Do not trust a default printed in any guide, including this one. Defaults move between releases and models carry their own. Read yours instead.

ollama show --parameters qwen3
ollama show --modelfile qwen3

ollama show --parameters prints only what the model carries. A parameter that does not appear in that output is not set to some visible value: it falls through to your build's default, which this command will not show you. If a value matters for your workload, set it explicitly so that it appears there. Anything you cannot see, you cannot review six months from now.

Which settings for which workload

Start from the job, not from a universal best value. There is no best value, because reliability and variety pull in opposite directions and different work wants opposite ends.

Coding agents, refactoring, anything whose output gets executed. Temperature 0, repeat_penalty at 1.0, and leave the tail filters alone. Creative code is wrong code. If you want the model to think before it answers, that is a separate control, covered in the reasoning effort settings on a local model, and you do not get it by raising temperature.

Structured extraction, JSON, classification. Temperature 0 plus a JSON schema in the request's format field. The schema restricts which tokens are legal, which is a far harder guarantee than any sampling value. Temperature 0 on its own still lets the model invent a key name.

Summarising or rewriting a source document. Around 0.3. Low enough to stay close to the source, high enough that the sentences do not all start the same way. If the summary misses the beginning of the document, the cause is the context window, not the temperature.

Open drafting and brainstorming. 0.8 to 1.0, with min_p near 0.05 if your build exposes it, or top_p near 0.9 if it does not. This is the one workload where the tail of the distribution is doing useful work.

Every number there is a starting point for your model, not a setting to copy permanently. Change one parameter at a time, with a fixed seed, and read both outputs.

Temperature is not a quality dial

The most expensive mistake in this area is treating temperature as a slider running from bad output to good output. It is not that. It only trades reliability against variety within what the model already knows. A 7B model at a heavy quantization that cannot keep your project's conventions straight will not keep them straight at 0.2 or at 0.9. You get confidently wrong output at one end and erratically wrong output at the other.

Match the symptom to its real cause before you touch a sampling value.

Sampling settings are worth tuning once you have the right model, the right context length and a prompt that works. They are the last 10 percent, and they cannot buy the first 90.

FAQ

What is the best temperature for Ollama?

There is no single best value, because temperature trades reliability for variety and different jobs want opposite ends of that trade. Use 0 for code, tool calls and structured extraction, where the same input should give the same output. Use 0.7 to 1.0 for drafting, where a repeated phrase is the failure you are trying to avoid. Around 0.3 suits summarising, where you want to stay close to a source. Whatever you pick, change one parameter at a time with a fixed seed so you can see what your change actually did.

Why is my Ollama temperature setting being ignored?

Almost always because it was set in a lower layer than the one the request reads. Ollama applies its built-in defaults, then the model's own defaults, then the Modelfile PARAMETER lines, then the options object on the API request, and each layer overwrites the one before it. An options object from your client beats any Modelfile, and nothing logs the conflict. Run ollama show --parameters <model> to see what the model carries, then check what your client sends. Many agents and SDKs set a temperature themselves without mentioning it.

Does setting a seed in Ollama guarantee identical output?

Only while everything else is identical: the same model file, the same quantization, the same Ollama build, the same options, the same split of layers between GPU and CPU, and no other request running against that model at the same time. Those conditions change the order floating point numbers are added in, and floating point addition is not associative, so a logit can differ in its last decimal. That matters only at steps where two candidates are nearly tied, but one different token changes every token after it. For a repeatable comparison, set OLLAMA_NUM_PARALLEL=1 on the server and change nothing except the parameter you are testing.

Should I use top_p, top_k or min_p?

Pick one and leave the others at their neutral setting, because stacking them makes the combined effect impossible to reason about. min_p is the easiest to predict: it drops any candidate below a fixed fraction of the leading token's probability, so it cuts hard when the model is confident and stays open when it is not. top_p is the safe fallback when your build does not list min_p in the /set parameter help. top_k is the bluntest, because the right number of candidates is different at every step. At temperature 0 none of them change anything, since only the leading token can be chosen.

#ollama#local-llm#sampling#temperature#self-hosting