SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Run Laya on a VPS: pip, Ollaya or Ollama?

Self-host the Laya decision model on a CPU VPS with laya-serve 0.3.26, keep it off the public internet, and learn why Ollama cannot load it but Ollaya can.

Which runtime actually runs the Laya decision model?

To run the Laya decision model on a VPS today, install the laya Python package in a virtual environment and start its own server, laya-serve. As of 2026-10-04 the current release on PyPI is 0.3.26, and this guide pins that exact version. Ollaya, an independent runtime, also lists Laya. Ollama does not run it.

Laya is a decision model from Convai Innovations, released under the Apache-2.0 license. The model card is convaiinnovations/laya on Hugging Face and the code lives at NandhaKishorM/laya on GitHub. A decision model does not write text. You send it a piece of text, called the state, plus a set of typed questions: yes or no, a choice between options, or a score. It returns a probability for every option in one forward pass. That is what you want for ticket routing, tagging and moderation, where the output is a label and a confidence, not a paragraph. Laya speaks the same HTTP contract as TypeSafe's hosted Jev API, POST /v1/systemone, so most of the guide to running a Jev-style decision model on a VPS applies here too.

These are the three runtimes people search for, as they stood on 2026-10-04:

  • laya-serve 0.3.26, installed from PyPI. This is the reference server, from the same author as the model. It routes each request to the right Laya checkpoint. The rest of this guide uses it.
  • Ollaya v0.9.0, released 2026-10-02 at ollaya-dev/ollaya by Mert Cobanov. It is an independent project that describes itself as Ollama for decision models. Its README lists laya:en and laya:multilingual, served on port 11435 behind a TypeSafe-compatible API.
  • Ollama 0.35, which added decision models in a blog post dated 2026-09-29. Its /v1/systemone endpoint serves nimble, tev1 and tev1:0.8b. Laya is not on that list.

The only result for "laya" in the Ollama model library is a community upload, s1gnature/laya. It is tagged as an embedding model and has no readme. An embedding model returns vectors from /api/embed. It does not answer /v1/systemone questions. So the plain answer to "ollama laya" is: Ollama does not load Laya as a decision model, as of this date. Use laya-serve, or Ollaya if you want the Ollama-style pull and serve workflow.

What Convai claims, and what you should measure yourself

Every number in this section is Convai's own published figure. None of them was measured for this guide. Treat them as vendor claims until your own box confirms them.

ChartLaya checkpoints, as published by Convai
The data behind this chart
[
  {
    "label": "laya (English)",
    "params_millions": 421,
    "context_tokens": 512,
    "download_mb": 808
  },
  {
    "label": "laya-multilingual",
    "params_millions": 322,
    "context_tokens": "1,024",
    "download_mb": 647
  }
]

According to Convai, the English checkpoint has 421 million parameters, a context of 512 tokens, and a download of about 808 MB. The multilingual checkpoint has 322 million parameters, a context of 1,024 tokens, and a download of about 647 MB. A third checkpoint, laya-typed-decisions, is built on the same base model as the English one and targets domain-specific decision tasks.

Convai also claims better calibration than Jev. ECE (expected calibration error) measures the gap between the confidence a model states and how often it is actually right. Lower is better.

ChartCalibration error (ECE), Convai's published figures, lower is better
The data behind this chart
[
  {
    "label": "Laya",
    "ece": 0.081
  },
  {
    "label": "Jev",
    "ece": 0.246
  }
]

The model card gives Laya an ECE of 0.081 against 0.246 for Jev. The README also says Laya is six to eight times faster than Jev. Those speed figures come from GPU runs, and a CPU VPS will be slower. Do not repeat the vendor's calibration numbers as fact in your own planning. The Jev guide above includes a calibration check you can run against your own labelled data, and it works unchanged against laya-serve because the API is the same.

Install laya-serve 0.3.26 in a virtual environment

Ubuntu 24.04 marks the system Python as externally managed, so a bare pip install fails with error: externally-managed-environment. A virtual environment avoids that and keeps Laya's dependencies away from the system packages. If you are deciding between a venv, pipx and uv for server tools, the comparison of venv, pipx and uv on a server covers the trade-offs. This guide uses a plain venv in /opt/laya, owned by root, and runs the server as its own system user.

sudo apt update && sudo apt install -y python3-venv
sudo useradd --system --create-home --home-dir /var/lib/laya --shell /usr/sbin/nologin laya
sudo python3 -m venv /opt/laya
sudo /opt/laya/bin/pip install --upgrade pip
sudo /opt/laya/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu
sudo /opt/laya/bin/pip install "laya[serve]==0.3.26"

The torch line comes first for a reason. Laya runs on PyTorch, and on Linux the default torch wheel from PyPI pulls in the NVIDIA CUDA libraries. That is several gigabytes of disk a CPU VPS never uses. Installing the CPU build from PyTorch's own index first means pip finds torch already present when it resolves Laya. The [serve] extra adds FastAPI and uvicorn, which laya-serve needs.

Check both versions:

/opt/laya/bin/pip show laya | head -2
/opt/laya/bin/python -c "import torch; print(torch.__version__)"

The first command should print Version: 0.3.26. The second should print a version ending in +cpu. If it does not end in +cpu, pip replaced your CPU build with the default wheel, usually because Laya asked for a torch version the CPU index did not have.

What does laya-serve bind to by default?

The PyPI page says laya-serve binds 0.0.0.0:8000. That means every network interface, including the public one. Do not take the docs on trust. Read the line in the release you installed:

grep -n 'LAYA_HOST' /opt/laya/lib/python3*/site-packages/laya/serve.py

In the published source the default is os.environ.get("LAYA_HOST", "0.0.0.0"). If your output shows that line, the docs are right for your version. On a VPS with a public IP and no firewall rule in the way, anyone who finds port 8000 can send requests and spend your CPU. Authentication is also off unless you set LAYA_API_KEY, so the default is an open endpoint.

Start the server in the foreground with loopback set explicitly:

sudo -u laya -H env LAYA_HOST=127.0.0.1 LAYA_PORT=8000 LAYA_DEVICE=cpu /opt/laya/bin/laya-serve

In a second SSH session, check what is actually listening:

ss -ltn 'sport = :8000'

The Local Address column should read 127.0.0.1:8000. If it reads 0.0.0.0:8000, the variable did not reach the process. The usual cause is that sudo dropped it, which is why the command above passes it through env. The -H flag sets HOME to /var/lib/laya, so the checkpoint download lands in that user's cache and not in yours.

Send the first /v1/systemone request

Save the README's example question to a file so you can send it again:

cat > q.json <<'EOF'
{"state": {"body": "billed twice, refund please or we cancel"},
 "questions": {"dept": {"type": "choice", "instructions": "which team?",
   "criteria": {"billing": "refunds", "tech": "bugs"}}}}
EOF
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d @q.json | python3 -m json.tool

The response holds an answers object keyed by your question name, dept, with a probability for each option. For this text, billing should get the higher probability. The first request is slow because the server downloads the checkpoint, about 808 MB for English, and then loads it into memory. Requests after that show the real speed.

The server also has a health route, which reports checkpoint status and the device in use:

curl -s localhost:8000/health | python3 -m json.tool

The device should say cpu. If a checkpoint shows as not loaded, it has not been requested yet.

Measure RAM and latency on your own VPS

Published latency comes from a GPU. Your number depends on your CPU model and your vCPU count, so take it yourself after the first warm-up request:

for i in 1 2 3 4 5; do
  curl -s -o /dev/null -w '%{time_total}\n' localhost:8000/v1/systemone \
    -H 'content-type: application/json' -d @q.json
done
ps -o rss=,cmd= -u laya

The loop prints the total time of each request in seconds. The ps line prints resident memory in kilobytes for each process run by the laya user. Note both figures with the plan size you tested on.

Two settings change memory use. The router picks the multilingual checkpoint for non-English text, and LAYA_MAX_LOADED defaults to 2. One Spanish ticket can therefore load a second checkpoint next to the English one, and memory grows to match. Set LAYA_MAX_LOADED=1 if you only need one at a time, or LAYA_IDLE_UNLOAD_SECONDS to unload a checkpoint that has been idle. If Laya shares the box with other services, LAYA_THREADS caps the CPU threads it uses, so it does not take every core.

Add bearer auth with LAYA_API_KEY and run it under systemd

When LAYA_API_KEY is set, every request needs an Authorization: Bearer <key> header. Put the settings in a root-only file, then write a systemd unit that reads it.

sudo install -m 600 /dev/null /etc/laya.env
printf 'LAYA_HOST=127.0.0.1\nLAYA_PORT=8000\nLAYA_DEVICE=cpu\nLAYA_API_KEY=%s\n' "$(openssl rand -hex 32)" | sudo tee /etc/laya.env > /dev/null

Create /etc/systemd/system/laya.service:

[Unit]
Description=Laya decision model server
After=network-online.target
Wants=network-online.target

[Service]
User=laya
Group=laya
EnvironmentFile=/etc/laya.env
Environment=HOME=/var/lib/laya
ExecStart=/opt/laya/bin/laya-serve
Restart=on-failure

[Install]
WantedBy=multi-user.target

Stop the foreground server with Ctrl+C, then start the service:

sudo systemctl daemon-reload
sudo systemctl enable --now laya
sudo systemctl status laya --no-pager

Now test both sides of the key:

curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d @q.json
KEY=$(sudo sed -n 's/^LAYA_API_KEY=//p' /etc/laya.env)
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' \
  -H "Authorization: Bearer $KEY" -d @q.json | python3 -m json.tool

The first call should return HTTP 401 with {"detail":"invalid or missing bearer token"}. The second should return the answers as before. Run ss -ltn 'sport = :8000' again and confirm the service still listens on 127.0.0.1.

A key is not encryption. Over plain HTTP the bearer token crosses the network in cleartext, so a remote client should reach Laya through a TLS reverse proxy, with Laya itself staying on loopback. The pattern is the same one used for Ollama, and securing a self-hosted model API endpoint walks through the proxy, TLS and firewall steps.

Does a Jev client work after changing baseUrl?

The Laya README says an existing Jev client only needs its baseUrl pointed at laya-serve, and that the response schema is identical. One real request settles it for your client. If your client is the TypeSafe SDK, it reads TYPESAFE_BASE_URL and TYPESAFE_API_KEY. Set the first to http://127.0.0.1:8000 and the second to your LAYA_API_KEY, then send one question your code already sends to Jev. A pass means your code parses the answer exactly as before. A 401 means the key or the header is wrong.

The README also lists the porting differences, and these are where a working Jev request can break:

  • A question can carry about 20 options, because the options share one token budget per checkpoint.
  • Every score level needs a description. A level with no description is not accepted.
  • Confidence is calibrated differently, so a threshold you tuned on Jev's confidence values will not mean the same thing on Laya.

The last point matters most in production. Re-tune any threshold before you move live traffic, using the calibration check linked above.

Running Laya through Ollaya instead

Ollaya ships as a single binary with Ollama-style commands: pull, run, list, ps and serve. Its README gives a pipe-to-shell installer. Download it and read it before you run it:

curl -fsSL https://ollaya.dev/install.sh -o install-ollaya.sh
less install-ollaya.sh
sh install-ollaya.sh
ollaya pull laya:en

If the installer did not start a background service, run ollaya serve. Per its documentation it listens on 127.0.0.1:11435 by default, and OLLAYA_HOST changes that. So, unlike laya-serve, it starts on loopback. Requests use the same /v1/systemone body, plus a model field set to laya:en or laya:multilingual. Port 11435 sits next to Ollama's 11434, so both can run on one box. How Ollama's API uses port 11434 explains the Ollama side if you run both. The procedure in this guide was written for laya-serve. The Ollaya commands here follow its README for v0.9.0.

Pick laya-serve if Laya is the only model you need and you want the reference server pinned through pip. Pick Ollaya if you want several decision model families behind one daemon. If you are already on Ollama for its own decision models, the ticket triage setup with Ollama on a VPS shows the same kind of routing job end to end.

FAQ

Can Ollama run the Laya decision model?

Not as of 2026-10-04. Ollama 0.35 added a /v1/systemone endpoint for decision models, but it serves nimble, tev1 and tev1:0.8b. The only "laya" in the Ollama library is a community upload tagged as an embedding model, which returns vectors and does not answer decision questions. To run Laya, use its own laya-serve from PyPI or the independent Ollaya runtime, which lists laya:en and laya:multilingual.

Is laya-serve safe to expose on port 8000?

Not with its defaults. The server reads LAYA_HOST and falls back to 0.0.0.0, which listens on every interface, and it has no authentication unless LAYA_API_KEY is set. Set LAYA_HOST=127.0.0.1 and an API key, confirm the bind with ss -ltn 'sport = :8000', and put a TLS reverse proxy in front if remote clients need it.

How much RAM does Laya need on a CPU VPS?

Measure it on your own plan. Send one warm-up request, then run ps -o rss=,cmd= -u laya to see resident memory in kilobytes. Memory grows when non-English text loads the multilingual checkpoint next to the English one, because LAYA_MAX_LOADED defaults to 2. Set it to 1, or set LAYA_IDLE_UNLOAD_SECONDS, to keep memory lower.

Why is the first Laya request so slow?

The server downloads the checkpoint on first use, about 808 MB for English according to Convai, and then loads it into memory. Later requests skip both steps. Time a few requests with curl -w '%{time_total}' after the first one to see your real latency.

Do I need a GPU to run Laya?

No. Set LAYA_DEVICE=cpu and Laya runs on an ordinary VPS. Convai's speed figures come from GPU runs, so expect higher latency on CPU, and measure it before you decide whether it is fast enough for your traffic.

#laya#decision-models#ollama#self-hosted-ai#text-classification