Classify images on a VPS with Ollama's Clef
Run Cloudflare's Clef vision decision models in Ollama on your VPS. Moderate uploads and route tickets with screenshots, and the image never leaves the box.
Classify images on your own VPS with Ollama and Clef
You can classify images on your own VPS (virtual private server) with Ollama and Cloudflare's Clef decision models. The image and the answer never leave the server. You send one request with an image, a short text context and a set of named questions. Clef returns a probability for each possible answer instead of written text, so your code branches on a number.
Everything here was checked against Ollama 0.35.1, its documentation and the ollama.com library pages, as of October 2026. The guide builds three jobs you probably already do by hand:
- Checking user uploads before they are published.
- Labelling the screenshot attached to a bug report.
- Routing a support ticket by its text and its screenshot together.
A decision model is different from a chat model. It reads the whole input and scores the answers you defined. It does not write a reply. In the image example in Ollama's documentation, the response reports "output_tokens": 0.
Update Ollama to 0.35.1 or later
The /v1/systemone endpoint arrived in Ollama 0.35.0. Image support for Clef arrived in 0.35.1. An older server has no such endpoint, so check the version first. If Ollama is not installed yet, follow the guide to running Ollama on a VPS and come back here.
ollama -v
curl -fsSL https://ollama.com/install.sh | sh
ollama -v
systemctl status ollama --no-pagerThe install script also upgrades an existing install. The second ollama -v should print 0.35.1 or later, and systemctl status should show the ollama service as active (running).
Which Clef model fits your VPS RAM?
The ollama.com library lists two Clef models. Both take text and images, and both have a 256K token context window.
clef-flash(tagclef-flash:9b): 9 billion parameters, an 11GB download. The library page says it is fine-tuned from Qwen3.5-9B.clef(tagclef:27b): 27 billion parameters, an 18GB download.
The download size is the floor for RAM (random access memory), because Ollama loads the whole weights file into memory. It also needs room for the context cache, and the operating system and your own application need memory too. That gives a plain rule:
- Clef Flash needs a VPS with at least 16GB of RAM.
- Clef needs a VPS with at least 32GB of RAM.
Neither model belongs on the smallest plans. A 4GB or 8GB VPS cannot hold even Clef Flash. Check what the box has free before you pull 11GB:
free -h
df -h ~The available column in free -h must be well above the model size. Start with Clef Flash. Move to Clef only if your own tests show Flash gets too many images wrong. If you are still choosing hardware for several models, the overview of which AI models you can self-host lists other sizes.
Pull Clef Flash and check it
ollama pull clef-flash
ollama list
ollama show clef-flashollama list should show clef-flash:latest at about 11GB. Under Capabilities, ollama show lists only decision. That is expected. The 0.35.1 release notes say decision models report only the decision capability in ollama show and in model listings, even though Clef reads images.
Send one image with three questions
A request to /v1/systemone has these fields, taken from Ollama's API reference:
model: a local decision model, hereclef-flash.state: the context for the decision. It is required even when you send an image. It can be a string, an object or an array.images: an array of base64 strings. PNG, JPEG and WebP are accepted. URLs are not.questions: between 1 and 64 named questions.keep_alive: optional. How long the model stays in memory after the request. The default is 5 minutes.
There are two size limits. A request without images must not exceed 64 KiB. A request with images may be up to 32 MiB, and that limit counts the base64 text and the JSON (JavaScript Object Notation) around it. Base64 makes a file about one third larger, so keep the image itself well under 24 MiB.
Do not paste the base64 string into a curl -d '...' argument. Linux limits a single command-line argument to 128 KiB, so a normal screenshot fails with Argument list too long. Build the body in a file with jq and send the file instead.
This first request is the bug report job. It asks three question types in one pass: a choice with a criteria map, a noul (yes or no) question, and an ordered score.
sudo apt update && sudo apt install -y jq
base64 -w0 screenshot.png > screenshot.b64
jq -n --rawfile img screenshot.b64 '{
model: "clef-flash",
state: "A user attached this screenshot to a bug report titled: Checkout button does nothing.",
images: [$img],
keep_alive: "1h",
questions: {
area: {
type: "choice",
instructions: "Which part of the web app is shown in this screenshot?",
criteria: {
login: "The sign-in or password reset page",
checkout: "The cart or payment pages",
dashboard: "The main dashboard after sign-in",
settings: "Account or profile settings",
other: "Something else, or not our app"
}
},
shows_error: {
type: "noul",
instructions: "Does the screenshot show an error message or an error page?",
criteria: {
"false": "No error is visible",
"true": "An error message, error code or crash page is visible"
}
},
severity: {
type: "score",
instructions: "How badly is the user blocked, based on the screenshot and the report title?",
criteria: [
"Cosmetic: a layout or wording problem only",
"Degraded: a feature misbehaves but the user can continue",
"Blocked: the user cannot finish the task"
]
}
}
}' > request.json
curl -s http://127.0.0.1:11434/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary @request.json | jq .--data-binary sends the file exactly as written. Plain -d @file strips newlines. That does no harm to this JSON, but it can break other files.
What "shared by all questions" means for images
The 0.35.1 release notes say images in a request are "shared by all questions and scored jointly" with the text state. The image does not belong to one question. Every question sees the same image and the same state, and the model scores them together. Two things follow from that.
First, the report title in state can move the answer to an image question, and the screenshot can move the answer to a text question. That is useful for tickets, where neither half is enough alone. Second, if you put three images in the array, the answers describe the set. They do not describe each image. For one verdict per image, send one request per image.
Read the response field by field
Here is the image example response from Ollama's own documentation. It asked a noul question called has_ollama and a choice question called app:
{
"model": "clef-flash",
"answers": {
"has_ollama": {"type": "noul", "noul": 0.959},
"app": {
"type": "choice",
"choice": "vscode",
"probabilities": {"vscode": 0.966, "other": 0.017, "browser": 0.010, "terminal": 0.007},
"confidence": 0.868
}
},
"usage": {"input_tokens": 677, "output_tokens": 0}
}Your response has the same shape, with your question names as keys.
modelechoes the model you asked for.answersholds one entry per question, keyed by the name you gave it.- For a
choicequestion,choiceis the option with the highest probability.probabilitiescovers every option you defined and sums to 1.confidenceruns from 0 to 1 and measures how strongly the model favours one option over the others. - For a
noulquestion,noulis the probability that the answer is true. There is nochoicefield, so you set the cut-off yourself. - For a
scorequestion,scoreis the probability-weighted average of the zero-based levels. With three levels the scale runs from 0 to 2. Aseverityof 1.8 means most of the weight sits on "Blocked". The answer also carriesprobabilities, alegendwith your level descriptions, andconfidence. usage.input_tokenstells you how much input the model read, image included.output_tokensstays near zero because nothing is written.
Pull out only what your code needs:
curl -s http://127.0.0.1:11434/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary @request.json \
| jq '{area: .answers.area.choice, error: .answers.shows_error.noul, severity: .answers.severity.score}'One warning about confidence. Ollama's API reference describes it as distribution concentration, "Not calibrated correctness". A choice at 0.97 is not right 97 times in 100. It only means the model put most of its weight on one option. Whether a given probability matches real accuracy on your images is a calibration question. The guide to running a Jev-style decision model covers how to measure it on your own labelled data.
How fast is Clef on a CPU-only VPS?
Plan for this from the start. On a VPS with no GPU (graphics processing unit), a 9B vision model is a background queue tool. It does not belong in the request path, where a user waits for the page to load. Accept the upload, tell the user it is being checked, and publish when the worker is done.
Measure your own box instead of trusting anyone's figure. The CPU (central processing unit) model and the image size both change the answer. Run the same request three times:
for i in 1 2 3; do
curl -s -o /dev/null -w '%{time_total}\n' \
-H 'Content-Type: application/json' \
--data-binary @request.json \
http://127.0.0.1:11434/v1/systemone
done
ollama psThe first time includes loading 11GB from disk, so it is the slowest. The second and third show the real cost of one image. ollama ps should list clef-flash with 100% CPU under PROCESSOR on a box without a GPU. Almost all the time goes into reading the input, so time it with images the same size your users really send. A phone photo is a much larger input than a cropped screenshot. The method in measuring tokens per second for a local LLM applies to the input side too.
Keep the model loaded between jobs
With the default keep_alive of 5 minutes, a quiet queue lets the model unload. The next upload then pays the full load time again. The request examples here send "keep_alive": "1h", which keeps Clef in memory for an hour after each job. ollama ps shows when it will unload in the UNTIL column. To set this once for the whole server instead of per request, follow keeping an Ollama model loaded in memory.
Turn it into an upload moderation worker
A worker is a small loop. It takes uploads from a queue, calls the local endpoint, and acts on thresholds. The queue here is a spool directory, because it needs no extra software:
- Your application writes each upload into
/srv/uploads/inboxunder a temporary name ending in.part, then renames it to its final name. A rename inside one filesystem is atomic, so the worker never reads half a file. - The worker moves each file into
published,revieworrejected, and writes the model's answers next to it as a.jsonfile.
Create a system user and the directory:
sudo useradd --system --no-create-home --home-dir /srv/uploads --shell /usr/sbin/nologin clefworker
sudo install -d -o clefworker -g clefworker -m 750 /srv/uploadsSave this as clef-worker.py. It uses only the Python standard library.
#!/usr/bin/env python3
"""Moderate queued uploads with a local Clef model, one file at a time."""
import base64
import json
import os
import pathlib
import time
import urllib.error
import urllib.request
SPOOL = pathlib.Path(os.environ.get("SPOOL", "/srv/uploads"))
MODEL = os.environ.get("CLEF_MODEL", "clef-flash")
URL = "http://127.0.0.1:11434/v1/systemone"
IMAGE_TYPES = {".png", ".jpg", ".jpeg", ".webp"}
MAX_BYTES = 20 * 1024 * 1024 # base64 adds a third; the request cap is 32 MiB
QUESTIONS = {
"kind": {
"type": "choice",
"instructions": "What kind of image is this?",
"criteria": {
"photo": "A photograph of people, places or things",
"screenshot": "A screenshot of an app or website",
"document": "A scan or photo of a document, card or letter",
"artwork": "A drawing, diagram or other artwork",
},
},
"personal_data": {
"type": "noul",
"instructions": "Does the image show personal data such as an ID card, a bank card, or a name with an address or phone number?",
},
"risk": {
"type": "score",
"instructions": "How safe is this image to publish in a public community gallery?",
"criteria": [
"Safe: fine for any audience",
"Sensitive: needs a content warning or a closer look",
"Unsafe: sexual or violent content that must not be published",
],
},
}
def classify(path):
body = {
"model": MODEL,
"state": "A user uploaded this image to a public community gallery.",
"images": [base64.b64encode(path.read_bytes()).decode("ascii")],
"questions": QUESTIONS,
"keep_alive": "1h",
}
req = urllib.request.Request(
URL,
data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=900) as resp:
return json.load(resp)["answers"]
def decide(answers):
risk = answers["risk"]["score"] # 0 safe, 1 sensitive, 2 unsafe
personal = answers["personal_data"]["noul"]
if risk >= 1.5:
return "rejected"
if risk <= 0.2 and personal <= 0.1 and answers["kind"]["choice"] != "document":
return "published"
return "review"
def main():
for name in ("inbox", "published", "review", "rejected"):
(SPOOL / name).mkdir(parents=True, exist_ok=True)
while True:
for path in sorted((SPOOL / "inbox").iterdir()):
if path.suffix.lower() not in IMAGE_TYPES:
continue
answers = None
if path.stat().st_size > MAX_BYTES:
verdict = "review"
else:
try:
answers = classify(path)
verdict = decide(answers)
except urllib.error.HTTPError as exc:
if exc.code >= 500:
print(f"{path.name}: HTTP {exc.code}, will retry", flush=True)
continue
verdict = "review"
except (urllib.error.URLError, TimeoutError) as exc:
print(f"{path.name}: {exc}, will retry", flush=True)
continue
dest = SPOOL / verdict / path.name
path.rename(dest)
dest.with_name(dest.name + ".json").write_text(json.dumps(answers, indent=2))
print(f"{path.name}: {verdict}", flush=True)
time.sleep(5)
if __name__ == "__main__":
main()The worker skips files that do not end in an image suffix, which is why the .part names are safe. A 4xx (client error) answer, such as a 413 for an oversized request, sends the file to review, because retrying the same file cannot fix it. A 5xx (server error) answer or a lost connection leaves the file in inbox for the next pass.
Install it as a systemd service in /etc/systemd/system/clef-worker.service:
[Unit]
Description=Clef upload moderation worker
After=ollama.service
Wants=ollama.service
[Service]
User=clefworker
Group=clefworker
Environment=SPOOL=/srv/uploads
Environment=CLEF_MODEL=clef-flash
ExecStart=/usr/bin/python3 /usr/local/bin/clef-worker.py
Restart=on-failure
[Install]
WantedBy=multi-user.targetsudo install -m 755 clef-worker.py /usr/local/bin/clef-worker.py
sudo systemctl daemon-reload
sudo systemctl enable --now clef-worker
systemctl status clef-worker --no-pagerTest it with one image, handed over the same way your application will do it:
sudo install -o clefworker -g clefworker -m 640 test.jpg /srv/uploads/inbox/test.jpg.part
sudo mv /srv/uploads/inbox/test.jpg.part /srv/uploads/inbox/test.jpg
journalctl -u clef-worker -fAfter the model has read the image, the journal prints a line such as test.jpg: review. The file is now in that folder, and test.jpg.json beside it holds the full answers. Read that JSON for the first few dozen uploads. It is the fastest way to see whether your questions mean what you think they mean.
Pick thresholds, and send the middle band to a person
decide() has three outcomes. A risk score of 1.5 or more means most of the weight sits on "Unsafe", so the file is rejected. A score of 0.2 or less, with low odds of personal data and an image that is not a document, is published. Everything else goes to review.
That middle band is the point of the design. The model handles the clear cases well and is unreliable on the borderline ones, so a person decides those. Gating AI agent actions behind human approval shows how to build the review step so nothing in review goes live by accident. Treat 1.5, 0.2 and 0.1 as starting values only. Keep the human decisions from the review folder, compare them with the scores, and move the cut-offs. The calibration method in the Jev-style guide linked above applies directly.
Route support tickets by text and screenshot together
Ticket routing uses the same worker shape with different questions. Here state is an object holding the ticket text, and the screenshot goes in images. Because the image and text are scored jointly, a vague ticket such as "it does not work" can still reach the right team when the screenshot shows a payment page.
{
"model": "clef-flash",
"state": {
"subject": "Payment failed",
"body": "I tried three times and it still does not work. Screenshot attached."
},
"images": ["<base64-encoded screenshot>"],
"keep_alive": "1h",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and integrations",
"account": "Sign-in and account access"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
"criteria": {
"false": "No refund is requested",
"true": "The customer requests a refund"
}
},
"urgency": {
"type": "score",
"instructions": "How urgently does this ticket need a response?",
"criteria": [
"Routine: no time pressure",
"Soon: a customer is inconvenienced",
"Immediate: a critical service is unavailable"
]
}
}
}A sensible action rule assigns team only when the probability of the chosen option is high, and leaves the rest in a general triage queue. Ticket text can contain anything, so treat the routing as a suggestion. Never let the model's answer trigger a refund by itself. If your tickets have no screenshots, you do not need a vision model at all. Triaging support tickets with a text-only decision model does the same job with a much smaller model.
How many requests can run at once?
The worker above sends one request at a time. That is deliberate. On a CPU-only box, parallel image requests share the same cores, so each one gets slower. The queue lives in the spool directory, where waiting costs nothing. If you later run more than one worker, or point a second application at the same model, read how OLLAMA_NUM_PARALLEL and the request queue work first. More parallel slots also need more memory.
Keep port 11434 private
The worker calls 127.0.0.1:11434, and nothing else needs to. A default Ollama install on Linux listens only on the loopback address. Check that it still does:
ss -ltn | grep 11434The output should show 127.0.0.1:11434. If it shows 0.0.0.0:11434 or *:11434, someone set OLLAMA_HOST to listen on every interface, and anyone who can reach the port can use your model. Uploads that arrive on another server should move to this box as files, not as calls to an open Ollama port. Securing your Ollama API endpoint covers the cases where remote access is really needed.
When a request fails
Ollama's API reference lists four error codes for this endpoint:
- 400: the request is invalid or the model is not supported. Images need Clef or Clef Flash, so sending
imagesto a text-only decision model is one way to get it. - 404: the model is not on this server. Run
ollama listand check the exact name. - 413: the request is over 64 KiB without images or 32 MiB with them. Resize the image or trim the state.
- 500: the model failed to load or to score. Read
journalctl -u ollama -n 50for the reason.
Argument list too long from your shell is not an Ollama error. It means the base64 text went into a command-line argument. Use the jq and --data-binary @request.json pattern above.
FAQ
Can I classify images with Ollama without sending them to a cloud API?
Yes. Pull clef-flash or clef with Ollama 0.35.1 or later and send the image as base64 to http://127.0.0.1:11434/v1/systemone. The model runs on your server, so the image and the answer stay there. Keep port 11434 bound to 127.0.0.1 so no one else can use the endpoint.
How much RAM do I need for Clef Flash and Clef?
Clef Flash is an 11GB download and needs a VPS with at least 16GB of RAM. Clef is an 18GB download and needs at least 32GB. Ollama loads the whole model into memory and needs extra room for the context cache, so neither model fits on a 4GB or 8GB plan.
Is the confidence value the chance that the answer is right?
No. Ollama documents confidence as distribution concentration, not calibrated correctness. A high value means the model favoured one option strongly. To know how often it is right, compare its answers with human decisions on your own images and set thresholds from that.
Can I send several images in one request?
Yes, images is an array. All images are shared by all questions and scored jointly with the text state, so the answers describe the whole set. For a separate verdict on each image, send one request per image.
Why does clef-flash fail with "Clef: non-finite logit"?
This error is reported in open Ollama issue 18769, against 0.35.1 on Windows with a CUDA GPU. On CPU the same reporter saw Clef: cannot open model, while clef:27b worked on the same hardware. If you see either string, check for a newer Ollama release first, then try clef if your VPS has the RAM for it.