Triage support tickets with Ollama on a VPS
Classify support tickets on your own VPS with an Ollama decision model: one /v1/systemone call answers every question at once, no API key.
Triage support tickets on your own VPS
Ollama v0.35.0 can triage support tickets on a VPS you own, with no API key and nothing leaving the box. You send the ticket to POST /v1/systemone on 127.0.0.1:11434 together with the questions you want answered about it. The model answers every question in one request and returns the option it picked plus the probability it gave each one. There is no prompt to engineer into producing JSON, and no retry loop for the times a chat model answers with a sentence instead.
This guide was written against Ollama v0.35.0 on 30 September 2026, using the /v1/systemone OpenAPI specification published in Ollama's own documentation. Every field name below comes from that specification.
You need a working Ollama install to start. If you do not have one, begin with installing Ollama on a VPS and serving your first model, then come back.
Update Ollama to v0.35.0 or later
/v1/systemone arrived in v0.35.0. An older server has no such route, so it returns a 404 for the path itself, which looks the same on the wire as a missing model until you read the response body.
ollama -vIf that prints anything below 0.35.0, run the install script again. It upgrades in place and restarts the systemd unit.
curl -fsSL https://ollama.com/install.sh | sh
ollama -v
sudo systemctl status ollama --no-pagerollama -v should now report 0.35.0 or higher, and the unit should read active (running). Install jq too, because you will be reading these responses by hand for the first hour.
sudo apt update && sudo apt install -y jqWhich decision model fits your VPS RAM?
Two decision models are published as of 30 September 2026. Nimble comes from Bespoke Labs and is 9B parameters. Tev1 comes from Together AI in a 4B and a 0.8B size. The number that decides which one your plan can run is the download size, because the weights sit in memory the whole time the model is loaded.
The data behind this chart
[
{
"label": "nimble:9b",
"params_b": 9,
"download_gb": 9.5
},
{
"label": "tev1:4b",
"params_b": 4,
"download_gb": 4.5
},
{
"label": "tev1:0.8b",
"params_b": 0.8,
"download_gb": 0.81
}
]Those are the published download sizes from the library pages, not memory measurements taken on a server. Resident memory is the weights plus the key-value cache (KV cache) for the context you actually send, so budget above the download size rather than at it.
Read the 3 rows this way. nimble:9b pulls about 9.5 GB of weights, so it belongs on a 16 GB VPS. It is not a job for the smallest plans: on an 8 GB box it will either fail to load or push everything else into swap, and a swapping model is slower than no model. tev1:4b is about 4.5 GB, which fits on an 8 GB VPS with room for your application next to it. tev1:0.8b is about 0.81 GB at 0.8B parameters, small enough for a 2 GB plan.
Start with tev1:4b if your box has 8 GB or more. Measure it against your own resolved tickets before you decide whether it is good enough, because that measurement is the only one that describes your queue. The library pages list a 256K context window for all of these tags, which is far more than a support ticket needs.
Pull the model
ollama pull tev1:4b
ollama listollama list should now show tev1:4b with a size close to the download figure above. The pull is a one-time cost and the weights live under /usr/share/ollama/.ollama/models on a default Linux install, so check you have the disk space before you start rather than after.
One request, three named questions
The request body has three required fields. model is the local tag you just pulled. state is the thing you are asking about: a nonempty string, or an object or array that Ollama serializes to JSON text for the model. It is not chat messages, and images, tools and generation controls are not supported on this route. questions is an object of named questions, between 1 and 64 of them.
Each question is scored separately against the full state, and answers are not passed to later questions. That has a practical consequence: a question cannot refer to another question's answer. If your second question only makes sense once the first is decided, you need two requests.
There are three question types:
choicepicks one of 2 to 26 named options. Itscriteriaobject maps each option key to a description of when that option applies. Anulldescription tells the model to use the key itself as the description.scorereturns a position on an ordered scale. Itscriteriais an array of 2 to 26 descriptions, lowest first, which defines a scale from 0 to the number of criteria minus 1.noulreturns the probability that a condition is true. Itscriteriais optional and only renames the two outcomes, which default to No and Yes.
That third name reads like a typo and it is not. noul is the type string in Ollama's OpenAPI schema, where the object is called SystemOneNoulQuestion, and noul is also the key in the answer. The value is a number between 0 and 1, so if answer["noul"]: is true for 0.02 exactly as it is for 0.98. Always compare it against a threshold.
Save this as request.json. The state is one ticket, and the three questions are the ones a triage desk actually asks.
{
"model": "tev1:4b",
"keep_alive": "30m",
"state": {
"subject": "Charged twice for September",
"from": "priya@example.com",
"plan": "Business",
"body": "My card was billed on 2026-09-02 and billed again on 2026-09-03. I only have one subscription. The billing page shows a single invoice. Please refund the second charge."
},
"questions": {
"category": {
"type": "choice",
"instructions": "Which queue should handle this ticket?",
"criteria": {
"billing": "Payments, invoices, refunds, or card charges",
"bug": "The product returns an error or behaves incorrectly",
"account": "Login, password, seats, or account access",
"howto": "The customer asks how to use a feature that is working",
"abuse": "Spam, threats, or content that breaks the terms of service"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket for the customer?",
"criteria": [
"No time pressure. A question or a comment.",
"Mildly inconvenient. The customer can keep working.",
"Blocked on one task. A workaround exists.",
"Blocked completely, or money has left the customer account.",
"Production is down, or the customer reports a security problem."
]
},
"needs_human_review": {
"type": "noul",
"instructions": "Must a person read this ticket before any automated reply is sent?",
"criteria": {
"false": "An automated reply or an automatic routing decision is safe",
"true": "A person must read it first: money, legal risk, an angry customer, or anything unclear"
}
}
}
}Send it:
curl -s http://127.0.0.1:11434/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary @request.json | jqThe first call after a service restart includes the model load, so it is slow. The second is the one that tells you anything.
Read the response field by field
A response for that request looks like this. Treat the numbers as an example: probabilities and confidence move with the model, the quantization and the server configuration.
{
"model": "tev1:4b",
"answers": {
"category": {
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.9614,
"bug": 0.0221,
"account": 0.0104,
"howto": 0.0039,
"abuse": 0.0022
},
"confidence": 0.8728
},
"urgency": {
"type": "score",
"score": 2.7613,
"legend": {
"0": "No time pressure. A question or a comment.",
"1": "Mildly inconvenient. The customer can keep working.",
"2": "Blocked on one task. A workaround exists.",
"3": "Blocked completely, or money has left the customer account.",
"4": "Production is down, or the customer reports a security problem."
},
"probabilities": {
"0": 0.0102,
"1": 0.0361,
"2": 0.2204,
"3": 0.6488,
"4": 0.0845
},
"confidence": 0.3851
},
"needs_human_review": {
"type": "noul",
"noul": 0.9312
}
},
"usage": {
"input_tokens": 1042,
"output_tokens": 3
}
}answers is keyed by the question names you sent, so category, urgency and needs_human_review come back under exactly those keys. Nothing is renamed and nothing is reordered.
choice is the option key with the highest probability. A tie picks the first option in request order, which matters more than it sounds: if one of your options wins far too often, check whether two of your criteria descriptions overlap and are splitting the model's probability between them.
probabilities is normalized over the candidates you supplied and sums to 1. Those five numbers add to 1.0000 above. This is the field to log. The winning label alone throws away the information that tells you whether the decision was close.
confidence is defined in the specification as 1 - H(p) / ln(N), where H(p) is the entropy of the probabilities and N is the number of candidates. It is 0 when the probabilities are spread evenly, and it approaches 1 when one candidate dominates. It measures concentration and nothing else. The documentation says so in plain words: it is not calibrated correctness.
The score answer has one more field than choice. score is the probability-weighted average of the zero-based criterion indices, so 2.7613 here, not rounded to a level and not normalized to 0 and 1. Work it out from the probabilities above and you get the same number, because that is all it is. legend maps those indices back to the descriptions you sent, which means a log line can be read a year later without the original request. A score of 2.7613 sits between "blocked on one task" and "money has left the customer account", and its confidence of 0.3851 says the model was genuinely split between those two levels.
The noul answer carries only type and noul. There is no confidence field on this type, because with two candidates the probability already tells you the concentration. A value of 0.9312 means the model puts the true case at 93%.
usage.input_tokens is the sum of the full rendered prompt lengths across all questions, including the shared state repeated for each one, even when the server cached it. Three questions about one ticket means the ticket text is counted three times. That is why 1042 input tokens comes back for a short email. output_tokens counts tokens generated internally for scoring, including prefix preparation and retries, so it can exceed the number of questions and it is not the length of the JSON you received.
Route on the answer, with thresholds you choose
A classifier that returns a label is not yet a triage system. The routing logic is where the work is, and it comes down to three bands:
- Above a high confidence, route the ticket automatically.
- In the middle band, route it as a suggestion and let a person confirm.
- Below a floor, treat the answer as no answer and send the ticket to the unsorted queue.
That bottom band is the one people skip. When confidence is near 0 the probability was spread almost evenly across your options and the top option won by rounding. Acting on it is acting at random with a number attached, which is worse than admitting you do not know.
The middle band is an approval queue, and it is the same shape as gating an AI agent's actions behind a human approval step: the model proposes, a person commits. Start with the middle band wide. Narrow it only after you have numbers.
Split the question set out of request.json so the script and your curl tests cannot drift apart:
jq '.questions' request.json > questions.jsonThen save this as triage.py next to it:
#!/usr/bin/env python3
"""Ask the local Ollama decision model about one ticket, then route it."""
import json
import sys
import urllib.error
import urllib.request
from pathlib import Path
ENDPOINT = "http://127.0.0.1:11434/v1/systemone"
MODEL = "tev1:4b"
QUESTIONS = json.loads(Path(__file__).with_name("questions.json").read_text())
# Your numbers. Set them from your own resolved tickets, not from this page.
AUTO_ABOVE = 0.90
NO_ANSWER_BELOW = 0.60
REVIEW_ABOVE = 0.30
def ask(ticket):
payload = json.dumps({
"model": MODEL,
"keep_alive": "30m",
"state": ticket,
"questions": QUESTIONS,
}).encode()
request = urllib.request.Request(
ENDPOINT, data=payload, headers={"Content-Type": "application/json"}
)
with urllib.request.urlopen(request, timeout=120) as response:
return json.load(response)
def route(answers):
category = answers["category"]
confidence = category["confidence"]
review = answers["needs_human_review"]["noul"]
decision = {
"queue": category["choice"],
"confidence": round(confidence, 4),
"urgency": round(answers["urgency"]["score"], 2),
"review_probability": round(review, 4),
}
if confidence < NO_ANSWER_BELOW:
decision["queue"] = "unsorted"
decision["action"] = "no_answer"
elif confidence >= AUTO_ABOVE and review <= REVIEW_ABOVE:
decision["action"] = "auto_route"
else:
decision["action"] = "propose_to_agent"
return decision
if __name__ == "__main__":
try:
result = ask(json.load(sys.stdin))
except urllib.error.HTTPError as error:
sys.exit("systemone %d: %s" % (error.code, error.read().decode()))
print(json.dumps(route(result["answers"]), indent=2))Run it against the ticket you already have:
jq '.state' request.json | python3 triage.pyWith the response above it prints:
{
"queue": "billing",
"confidence": 0.8728,
"urgency": 2.76,
"review_probability": 0.9312,
"action": "propose_to_agent"
}The label was billing at 96%, and the ticket still goes to a person, because 0.8728 is under AUTO_ABOVE and the review probability is far over REVIEW_ABOVE. That is the system working. Money is moving and the model said a human should look.
Those three constants are placeholders, and they are yours to set. Nobody can pick them for you, because they depend on your queue, your options and how much a misroute costs you. Until you have checked, you do not know whether confidence: 0.9 means right nine times in ten or right six times in ten. Export a few hundred already-resolved tickets with the queue they truly ended up in, run them through the endpoint, and count, inside each confidence band, how often the model agreed. The calibration section of the decision-model guide works through that check and what to do when the bands come out uneven, so there is no need to repeat it here.
Make it a service: a small webhook receiver
Once the thresholds behave, put the script behind an HTTP endpoint that your helpdesk can call on every new ticket. This one is standard library only and binds to loopback. Save it as receiver.py next to triage.py.
#!/usr/bin/env python3
"""Accept one ticket per POST and return the routing decision as JSON."""
import json
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from triage import ask, route
MAX_BODY = 48 * 1024
class Receiver(BaseHTTPRequestHandler):
def do_POST(self):
length = int(self.headers.get("Content-Length", 0))
if length <= 0 or length > MAX_BODY:
self.send_error(413, "ticket must be between 1 byte and 48 KiB")
return
try:
ticket = json.loads(self.rfile.read(length))
decision = route(ask(ticket)["answers"])
except Exception as error:
self.send_error(400, "could not triage: %s" % error)
return
payload = json.dumps(decision).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
if __name__ == "__main__":
ThreadingHTTPServer(("127.0.0.1", 8099), Receiver).serve_forever()The 48 KiB cap is not arbitrary. A /v1/systemone request must fit within 64 KiB in total, and your question definitions are part of that total, so the ticket budget is 64 KiB minus the rendered schema. Rejecting an oversized ticket at your own door gives a clearer error than a 413 from Ollama.
Install it as a unit that starts after Ollama:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin triage
sudo install -d -m 755 -o root -g root /opt/triage
sudo install -m 644 -o root -g root triage.py receiver.py questions.json /opt/triage/The files stay owned by root, so the service account can read its own code and cannot rewrite it. Write /etc/systemd/system/triage-receiver.service:
[Unit]
Description=Support ticket triage receiver
After=network-online.target ollama.service
Wants=ollama.service
[Service]
User=triage
Group=triage
WorkingDirectory=/opt/triage
ExecStart=/usr/bin/python3 /opt/triage/receiver.py
Restart=on-failure
RestartSec=5
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
[Install]
WantedBy=multi-user.targetsudo systemctl daemon-reload
sudo systemctl enable --now triage-receiver
systemctl status triage-receiver --no-pager
curl -s -X POST http://127.0.0.1:8099/ \
-H 'Content-Type: application/json' \
--data-binary "$(jq -c '.state' request.json)" | jqYou should get the same routing decision the script printed. enable --now is the half that matters, because a receiver started by hand disappears at the next reboot and tickets then pile up unsorted with nothing logging an error. If the unit fails, journalctl -u triage-receiver -n 50 has the traceback.
This receiver is the minimum that works. A production version needs an idempotency key so a retried webhook does not route twice, and a place to store every response for the calibration check above. Log the whole answers object, not the label you acted on.
Keep the model loaded between tickets
Loading 4.5 GB of weights takes real seconds, and the default keep-alive is 5 minutes. If tickets arrive every ten minutes, you pay that load cost on every single one, and your latency measurement is mostly disk.
There are two controls. Per request, keep_alive accepts a duration string such as "30m", a number of seconds such as 3600, 0 to unload immediately after the request, or any negative number to keep the model loaded. The request.json above sets "30m", which is why the script passes it on every call.
Server-wide, set OLLAMA_KEEP_ALIVE, which takes the same values:
sudo systemctl edit ollama[Service]
Environment="OLLAMA_KEEP_ALIVE=-1"sudo systemctl restart ollama
ollama psollama ps lists the loaded model with an UNTIL column. With a negative keep-alive that column stops counting down toward an unload, so the weights stay resident until you restart the service. The cost is honest and permanent: that memory is gone from the rest of the box for as long as Ollama runs. On a single-purpose triage VPS it is the right trade. On a box that also serves your application, it may not be, and keeping an Ollama model loaded and what it costs in RAM covers how to decide.
How many tickets can it handle at once?
OLLAMA_NUM_PARALLEL is the maximum number of requests each model will process at the same time, and its default is 1. With the default, two tickets arriving in the same second are handled one after the other, and the second one's latency includes all of the first one's. OLLAMA_MAX_QUEUE is how many requests Ollama will queue while busy before it rejects the rest, and its default is 512.
[Service]
Environment="OLLAMA_KEEP_ALIVE=-1"
Environment="OLLAMA_NUM_PARALLEL=2"
Environment="OLLAMA_MAX_QUEUE=64"Raising OLLAMA_NUM_PARALLEL costs memory, because each parallel slot needs its own KV cache. Lowering OLLAMA_MAX_QUEUE is the kinder setting for a webhook: a request rejected in a millisecond lets your helpdesk retry it, while a request that sits in a 512-deep queue for two minutes usually times out at the sender and gets retried anyway, now with a duplicate still working its way through. Pick a queue depth close to what you can actually clear.
Remember that one request with three questions is scored as three separate prompts inside the server. Adding a fourth question adds another full pass over the ticket text. If you need throughput, cut questions before you add parallelism. Tuning OLLAMA_NUM_PARALLEL and OLLAMA_MAX_QUEUE goes into the memory arithmetic, and how many concurrent users a self-hosted model can really serve covers the same question from the load side.
Keep port 11434 private
Ollama binds 127.0.0.1 port 11434 by default and there is no authentication in front of it. Anyone who can reach that port can use your model, read your state text back out of their own requests, and pull or delete models. Never set OLLAMA_HOST=0.0.0.0 to make a call from another machine work. If the receiver runs on the same VPS, it does not need to, which is why everything above uses 127.0.0.1.
sudo ss -ltnp | grep 11434
sudo ufw statusThe first command should show 127.0.0.1:11434 and nothing bound to a public address. If it shows 0.0.0.0:11434, the port is open to anything your firewall allows through, and you should fix it before the next ticket arrives. For the cases where the caller genuinely lives on another host, putting authentication and TLS in front of your Ollama API endpoint is the right way to do it, and what actually listens on port 11434 explains which routes are exposed when you open it.
Measure latency and accuracy on your own tickets
Published latency figures for this endpoint are worth very little, because the answer depends on your CPU, your ticket length and how many questions you ask. Measure it yourself. It takes five minutes.
sudo systemctl restart ollama
curl -s -o /dev/null -w '%{time_total}\n' \
http://127.0.0.1:11434/v1/systemone \
-H 'Content-Type: application/json' --data-binary @request.jsonThat first number includes the model load, so throw it away. Then take twenty warm calls:
for i in $(seq 1 20); do
curl -s -o /dev/null -w '%{time_total}\n' \
http://127.0.0.1:11434/v1/systemone \
-H 'Content-Type: application/json' --data-binary @request.json
done | sort -n | awk '{a[NR]=$1} END {print "median", a[int(NR/2)+1], "max", a[NR]}'Run the same loop against tev1:0.8b by changing one line in request.json. Now you have a real comparison between the two models on your hardware and your ticket length, which is the only comparison that should decide anything.
For accuracy, the method is the same shape. Export resolved tickets that already carry the queue a human chose, feed each state through triage.py, and compare queue against the recorded answer. Group the results by confidence band and count the agreement rate in each band. If the 0.9-and-up band agrees far less often than 90% of the time, your AUTO_ABOVE is too low and the model is not calibrated on your data. Keep the raw responses, because you will want to re-run this after every model or criteria change. Generation speed is a separate measurement that does not apply to this route, though measuring tokens per second on a local model is worth reading if you also run a chat model on the same box. If you do not yet have labelled history to test against, running your own helpdesk and ticket store is where that history comes from.
Failure modes, and the response you will see
A 404 for the whole route. The server is older than v0.35.0. Check ollama -v before you debug anything else, because a missing route and a missing model both arrive as 404.
A 404 for the model. The documentation is explicit: the local model was not found, so download it before making a request. Confirm with ollama list and watch for a typo in the tag, since tev1:4B and tev1:4b are not the same string.
A 400 for an unsupported model. The route needs compatible GGUF weights and a scoring-capable runner. Cloud models are rejected, and MLX or Safetensors models are not supported. Pointing model at a chat model you already had installed gives this error, not a bad answer.
A 400 because the rendered prompt exceeds the loaded context. Each rendered prompt must fit the loaded context window with two token positions left for scoring, and input is never truncated. A very long ticket combined with long criteria descriptions hits this. Shorten the ticket body yourself, or load the model with a larger context.
A 413 reading request body must not exceed 64 KiB. The whole request body counts, questions included. Cap the ticket text on your side, as the receiver above does.
A 500. Model loading, rendering or scoring failed on the server. Read journalctl -u ollama -n 50. On small plans this is usually the model being killed for memory, which sends you back to the sizing section.
A label that is always the same. Not an error, and the most common real problem. Log probabilities for a hundred tickets and look at the spread. If the second-place option is consistently close behind, two of your criteria descriptions overlap, and the fix is in the descriptions rather than in the model.
FAQ
Do I need an API key or an internet connection to triage tickets with Ollama?
No. /v1/systemone runs entirely against a local model on 127.0.0.1:11434, so there is no key, no account and no per-request cost. The ticket text never leaves the VPS. You need the internet once, to install Ollama and to pull the model, and after that the box can triage with outbound access closed. Cloud models are explicitly rejected on this route, so there is no way to accidentally send a ticket off the machine through it.
What is the noul question type in Ollama, and is that name a typo?
It is the yes/no question type, and the spelling is real. noul is the type string in Ollama's published OpenAPI schema for /v1/systemone, where the request object is named SystemOneNoulQuestion, and noul is also the key that carries the answer. The value is the probability that the condition is true, between 0 and 1, and not a Boolean. Testing it for truthiness in code is wrong, because 0.02 is a truthy number in most languages. Compare it against a threshold you picked. Unlike choice and score, a noul answer has no confidence field, since with two candidates the probability already tells you how concentrated the result is.
Which Ollama decision model do I need for an 8 GB VPS?
tev1:4b, which downloads at about 4.5 GB. That leaves room for your application and for the key-value cache alongside the weights. nimble:9b is about 9.5 GB of weights and wants 16 GB of RAM, so on an 8 GB box it will either refuse to load or force swapping, and a swapping model is slower than no model at all. If your plan is smaller than 8 GB, use tev1:0.8b. Then measure both on your own resolved tickets, because download size tells you what fits and nothing about whether the answers are good enough for your queue.
Does a high confidence value mean the answer is correct?
No. confidence is defined as 1 - H(p) / ln(N), which measures only how concentrated the probability distribution is over the candidates you supplied. A value near 1 means one option dominated the others. A value of 0 means the probability was spread evenly. Ollama's own documentation states that it is not calibrated correctness, so a model can be concentrated and wrong. To find out what a given number means for your data, run a few hundred already-resolved tickets through the endpoint and count the agreement rate inside each confidence band. Set your automation threshold from that count, not from the number's appearance.
Why does my request return HTTP 413 or 400?
A 413 means the request body went over the 64 KiB limit, and that limit covers the whole body, so your question definitions eat into the budget before the ticket text does. Cap the ticket length in your own code. A 400 has three common causes on this route: the request is malformed, the model is not supported here, or the rendered prompt does not fit the loaded context window. The endpoint never truncates input and needs two token positions free for scoring, so a very long ticket plus long criteria descriptions produces a 400 rather than a shortened answer. Shorten the ticket, or reload the model with a larger context.