Run a Jev-style decision model on a VPS
Jev has no public weights. Build the same typed decisions on your own server with llama.cpp grammars or GLiNER2.5-Decide, and check the confidence.
What you can run on a VPS instead of Jev
Jev has no published weights, so you cannot run a Jev-style decision model on a VPS by downloading Jev. What you can run is the same contract: hand a model some state plus typed questions, get back typed answers and a number that claims to be a confidence. Two paths do that today on a small server. One constrains an open LLM (large language model) with a grammar in llama.cpp and reads the token probabilities back. The other runs an open-weight decision model built for the job, fastino/GLiNER2.5-Decide, which is 340M parameters under Apache 2.0.
The typed answer is the easy half. Constrained decoding has been solid for two years and it cannot emit an invalid value. The confidence number is the hard half, because you are asking for calibration from a model that was never trained to be calibrated. Most of this guide is about that number and how you would check it.
What Jev's API contract actually is
TypeSafe AI put Jev into early access on 15 September 2026 and calls it a System One model. You POST one request to https://api.typesafe.ai/v1/systemone carrying a state, a model, and a map of questions. Every declared question is answered in the same parallel pass, so mixing question types costs no extra round trip. As measured on 29 September 2026, input is billed at $0.042 per million tokens and output is free. The published budget is 64k tokens for the state and all questions together, with a tighter 32k limit on the state plus the longest single question.
There are three question types, and they are the whole surface you are rebuilding.
Choicepicks one option from a set of up to 255, and returns a probability for each option plus an overall confidence.Scoreplaces the state on an ordered scale of 2 to 10 levels, and returns a probability-weighted float.Noulanswers a yes or no statement with a single float from 0 to 1. It has no separate confidence field, because the float is the answer and the confidence at the same time.
None of that needs Jev's weights. It needs a model that will only emit values from a set you declared, and a way to read how much probability mass sat on each one.
Path one: llama.cpp with a grammar, and the probabilities behind it
Start a server on the box. The model download needs outbound network reach, so run this yourself on a machine that has it.
llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF -c 8192 --host 127.0.0.1 --port 8080Bind to 127.0.0.1. llama-server has no authentication of its own, so a bind to 0.0.0.0 on a VPS with a public IP puts an open inference endpoint on the internet within seconds of the port opening. Put it behind a reverse proxy or reach it over a tunnel. If you have not stood one of these up before, the full llama-server setup on a VPS covers the build, the service unit and the proxy.
A Choice is a GBNF grammar with one alternation. Send it on the request, not at server start, so different questions can use different option sets.
curl -s http://127.0.0.1:8080/completion -H 'Content-Type: application/json' -d '{
"prompt": "Ticket: my card was charged twice this month. Category:",
"grammar": "root ::= \"billing\" | \"technical\" | \"sales\" | \"spam\"",
"n_predict": 1,
"n_probs": 4,
"post_sampling_probs": true
}'The response carries the chosen string and the numbers you came for.
{
"content": "billing",
"probs": [
{
"top_probs": [
{"tok_str": "billing", "prob": <float>},
{"tok_str": "technical", "prob": <float>},
{"tok_str": "sales", "prob": <float>},
{"tok_str": "spam", "prob": <float>}
]
}
]
}probs holds one entry per generated token, so n_predict: 1 gives exactly one entry. Its top_probs list is your option set, because the grammar removed every other token from the sampling chain before the probabilities were taken. Flip post_sampling_probs to false and the field is named top_logprobs with a logprob on each entry instead, taken before the sampling chain. That version covers the whole vocabulary and you renormalise over your own options. Both are useful, for different reasons, and the reason matters later.
Score is the same shape with numeric labels: root ::= "1" | "2" | "3" | "4" | "5". Multiply each level by its probability and sum, and you have the probability-weighted float that Jev returns. Noul is root ::= "yes" | "no", and the answer is the probability on yes.
One token position can only separate options that differ in their first token.billingandbilling_disputeshare a prefix, so at position 0 the model has not chosen between them yet and your two probabilities are really one. Pick option labels whose first characters differ, or read the whole constrained continuation and accumulate probability along it.
The OpenAI-compatible route works too. POST /v1/chat/completions accepts response_format with a schema-constrained JSON object, which is the friendlier developer experience and slots into any existing client. The llama.cpp server documentation covers logprobs on /completion and not on that route, so the schema path gives you a valid answer without the number. If the confidence is the point, use /completion. That split is a good example of what llama.cpp exposes that higher-level runners wrap away.
Path two: GLiNER2.5-Decide, weights you can actually download
Fastino released GLiNER2.5-Decide on 24 September 2026. It is a 340M-parameter encoder on DeBERTa-v3-large, Apache 2.0, English, built for exactly this class of decision: routing, intent, triage, moderation, severity. It is not a general chat model and its own card says so plainly: it does not reason, explain, or answer open questions.
pip install gliner2from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")
text = "My card was charged twice this month and nobody has replied."
# Choice
print(model.classify_text(text, {"category": ["billing", "technical", "sales", "spam"]}))
# Score, as an ordered scale
print(model.classify_text(text, {"urgency": ["0", "1", "2", "3", "4", "5"]}))
# Noul, as a yes/no prompt
print(model.classify_text(text, {
"answer": {"labels": ["yes", "no"], "prompt": "Does this ticket need a human?"}
}))The three primitives map straight across, which is the point. Two honest limits before you plan around it. The quickstart returns chosen labels, and the multi-label form takes a cls_threshold, so scores exist inside the library. Read the library's return shape yourself before you build a routing rule on a field name. Second, this is an encoder, and an encoder takes a short fixed input window, far below Jev's 64k. Check the window on the model card before you plan to feed it a whole agent conversation. A long state has to be summarised or selected down first.
Fastino publishes a p50 of 38.3 ms on an NVIDIA V100 and 167.3 ms on a 48-vCPU Xeon, and 60.1% average exact match across a 17-dataset internal benchmark. Those are the vendor's numbers on the vendor's hardware. Your VPS is neither, so measure your own before you promise anyone a latency.
The hard part: a confidence number nobody calibrated
A probability from a language model is a probability over tokens. It is not a probability over the world. Nothing in pretraining or instruction tuning pushes it toward meaning "correct 80% of the time when it says 0.8". TypeSafe trains Jev with reinforcement learning for calibrated decisions precisely because that property does not arrive for free, and when you skip that training step you do not get the property.
The local ecosystem quietly admits this. TypeSafe's own system-one-adapter, a drop-in TypeSafeClient replacement backed by ordinary LLM APIs, ships a normalize_probabilities option whose documented job is to rescale invalid LLM probability distributions so they sum to 1. A set of numbers that needs rescaling to sum to 1 was never a distribution. Treat whatever you read out as an unvalidated score until you have measured it.
Two mechanisms distort it further. The grammar constraint itself renormalises over the allowed set, so a model that wanted to answer something outside your options still hands you a confident-looking split across options it did not want. And the prompt wording moves the numbers, because the model is predicting text, and different text has different likely continuations.
How to check calibration without guessing
You need labels. There is no way around that, and no amount of prompt engineering substitutes for it.
Collect 300 to 500 real examples of the decision with the correct answer attached. Run them through your setup and record the predicted answer and the stated confidence. Sort the predictions into buckets by confidence, in steps of 0.1. For each bucket, compare the mean stated confidence against the fraction that were actually correct. That table is a reliability diagram in numbers. The size-weighted average gap between those two columns is expected calibration error (ECE). For the yes/no case, the Brier score, the mean squared error between your float and the 0 or 1 outcome, gives you one number to track over time.
What you will usually find is that the 0.9 bucket is right well under 90% of the time. The standard fix is temperature scaling: divide the logits by a single scalar T before the softmax, and fit T on the holdout set to minimise negative log likelihood. One free parameter means a few hundred examples is genuinely enough. A T above 1 flattens an overconfident model. This is the reason to care about post_sampling_probs: to fit T you want the pre-sampling top_logprobs, renormalised over your option set yourself, because you cannot undo a softmax that llama.cpp already applied.
Then ship the threshold, not the number. The useful output of all this is one rule: below confidence X, escalate to a slower model or a person. Read X off the reliability table at the accuracy your product needs. That escalation is the same shape as routing between a cheap model and an expensive one inside an agent, with the confidence taking the place of a heuristic.
The arithmetic: when a VPS beats $0.042 per Mtok
Assume 2,000 input tokens per decision, which is a realistic agent state plus a handful of questions. At $0.042 per million tokens with output free, one decision costs $0.000084. Over 30 days:
The data behind this chart
[
{
"label": "1k/day",
"hosted_usd_month": "2.52",
"vps_usd_month": 12
},
{
"label": "5k/day",
"hosted_usd_month": "12.60",
"vps_usd_month": 12
},
{
"label": "20k/day",
"hosted_usd_month": "50.40",
"vps_usd_month": 12
},
{
"label": "100k/day",
"hosted_usd_month": "252.00",
"vps_usd_month": 12
},
{
"label": "500k/day",
"hosted_usd_month": "1,260.00",
"vps_usd_month": 12
}
]At 1,000 decisions a day the hosted bill is $2.52 a month. No server beats that, and self-hosting at that volume is a hobby, not a saving. The crossover against an assumed $12 a month box lands near 5,000 decisions a day, where the hosted bill reaches $12.60. Past that the hosted line keeps climbing and the flat line does not: $252.00 at 100,000 a day, and $1,260.00 at half a million, across the 5 volumes above.
Two caveats on that flat line. It stays flat only while one box keeps up, so size against your peak rate rather than your daily total: 500,000 a day is 5.8 per second averaged, and real traffic is not averaged. And it prices hardware, not your time. Calibrating and maintaining this yourself is real work that the hosted price includes.
The RAM floor for a decision model on a VPS
Weights in memory is parameter count times bytes per weight. That is arithmetic, not a measurement, so you can plan against it before you rent anything.
The data behind this chart
[
{
"label": "GLiNER2.5-Decide 340M, F32",
"bytes_per_weight": 4,
"weights_gb": 1.36
},
{
"label": "GLiNER2.5-Decide-1B, F32",
"bytes_per_weight": 4,
"weights_gb": 4.0
},
{
"label": "0.8B LLM, Q8_0",
"bytes_per_weight": 1,
"weights_gb": 0.8
},
{
"label": "4B LLM, Q4_K_M",
"bytes_per_weight": 0.55,
"weights_gb": 2.2
},
{
"label": "8B LLM, Q4_K_M",
"bytes_per_weight": 0.55,
"weights_gb": 4.4
}
]GLiNER2.5-Decide at F32 is 1.36 GB of weights, so a 2 GB VPS is too tight once Python, the tokenizer and the OS are resident, and 4 GB is comfortable. The 1B variant needs 4.0 GB at F32, which puts you on an 8 GB plan. On the llama.cpp side an 0.8B model at Q8_0 is 0.8 GB and an 8B at Q4_K_M is 4.4 GB. Q4_K_M averages about 4.5 bits per weight, so roughly 0.55 bytes, which is why that column is not a round number.
Add the KV cache on top for the llama.cpp path. It grows with the context length you set with -c, it is separate from the weights, and it is the usual reason a model that "fits" gets killed by the out-of-memory killer under load. Working out the real memory ceiling for a given model is worth doing before you pick a plan, and the wider question of which models are self-hostable at all narrows the field further.
No GPU is required for either path at these sizes. A 340M encoder is a CPU workload, and an 0.8B model at Q8_0 on a few vCPUs handles a steady trickle of decisions. GPU matters when your peak rate does.
The ecosystem, and what state each piece is in
system-one-adapter is published under TypeSafe's own typesafe-ai GitHub organisation, MIT licensed, in Python and JavaScript. Install with pip install 'system-one-adapter[openai]' or npm install system-one-adapter. It reimplements the TypeSafeClient surface, including Choice, Score and Noul, on top of an ordinary LLM API, and custom endpoints set through OPENAI_BASE_URL default to Chat Completions. That is the flag that points it at your own llama-server. Community forks of it exist and track the upstream loosely.
PragmaTwice/jeva.cpp is a llama.cpp fork, MIT, that adds a JEV-compatible decision API to llama-server and derives Choice, Score and Noul evaluations directly from model logits while leaving normal autoregressive generation intact. Its stated requirement is a model that exposes next-token vocabulary logits. Several other forks aim at the same target under names like jev.cpp and openjev.cpp. All of them are single-maintainer forks of a fast-moving upstream, so running one means carrying the rebase cost yourself. The grammar plus n_probs approach above needs no fork, which is why it is the one to start with.
GLiNER2.5-Decide is the only piece here with open weights built for the task rather than adapted to it, and the 1B sibling on the Ettin encoder scores marginally lower on Fastino's own suite than the 340M model does. Smaller is not obviously worse in this class.
What none of them give you is calibration. That is the part Jev sells, it is the part you cannot download, and it is the part you have to measure on your own labelled data before the number in your routing rule means anything.
FAQ
Can I download Jev and run it on my own server?
No. TypeSafe has not published Jev's weights, and the model is only reachable through their hosted API at https://api.typesafe.ai/v1/systemone. Anything described as running Jev locally is running a different model behind a compatible interface. The closest open-weight substitute built for the same job is fastino/GLiNER2.5-Decide, 340M parameters under Apache 2.0, and the closest general substitute is any open LLM constrained by a GBNF grammar or a JSON schema in llama.cpp.
How do I get a confidence number out of llama.cpp?
Use the native /completion endpoint, not the OpenAI-compatible route. Send grammar with your allowed options, n_predict: 1, and n_probs set to the number of options. The response carries a probs array, and with post_sampling_probs: true each entry holds top_probs with a tok_str and a prob. Set post_sampling_probs to false and you get top_logprobs with raw logprob values across the vocabulary instead, which is what you want if you plan to fit a temperature. The llama.cpp server docs describe these fields on /completion and not on /v1/chat/completions, so the schema-constrained OpenAI route gives you a valid answer with no number attached.
Is the confidence from a constrained LLM calibrated?
Not by default, and you should assume it is not until you have checked. A token probability is a statement about likely text, not about how often the answer is right. The grammar makes it worse in one specific way: it renormalises the probability mass over your allowed set, so an option the model never wanted still receives a share. TypeSafe's own adapter library ships a normalize_probabilities setting whose documented purpose is rescaling invalid LLM probability distributions to sum to 1, which tells you what these numbers are before you touch them.
How do I actually measure calibration?
Label 300 to 500 real examples of the decision. Run them, record the predicted answer and the stated confidence, then bucket the predictions by confidence in steps of 0.1. In each bucket compare the mean stated confidence against the fraction that were correct. The size-weighted gap between those two columns is expected calibration error, and for yes/no questions the Brier score gives you a single tracking number. If the model is overconfident, fit one temperature scalar on that same holdout by minimising negative log likelihood, and divide the logits by it before the softmax.
At what request volume does a VPS beat Jev's per-token price?
At 2,000 input tokens per decision and $0.042 per million input tokens as measured on 29 September 2026, the hosted bill crosses a $12 a month server at roughly 5,000 decisions a day. Below that, paying per token is cheaper and far less work. Above it the gap widens fast, reaching $252.00 a month at 100,000 decisions a day. Size the server against your peak rate rather than the daily total, because the flat line only stays flat while one box keeps up.