SSD Nodes Learn Hosting plans →
How to do am Matt ConnorBy Matt Connor

How to run Ternary Bonsai 2 27B for CPU VPS

Ternary Bonsai 2 na Qwen3.8 27B for 1.72 bits, and e fit run for 12 to 16 GB VPS. We build di llama.cpp fork from fixed tag, do di RAM sum, and serve am without per-token bill.

How to run Ternary Bonsai 2 27B for CPU VPS: di short answer

To run Ternary Bonsai 2 27B for CPU VPS, you need a VPS (virtual private server) wey get at least 12 GB RAM, plus di PrismML fork of llama.cpp, built from one fixed release tag. Normal llama.cpp and Ollama no fit load di file. Di file wey we go use na PQ2_0. E weigh 7.21 GB, and with 8,192 tokens of context di whole thing dey sit around 9 GiB of RAM.

If you don read our guide to run Qwen3.8 27B on a VPS, you go remember say e turn back anybody wey get only 8 to 16 GB RAM. Dis post na for dat person. Ternary Bonsai 2 27B na di same Qwen3.8-27B, but PrismML squeeze am down to 1.72 bits per weight. Na dat one make 27 billion parameters fit enter small server.

Di gain for you dey plain: model wey sabi work dey run for server wey you rent, and nobody go charge you dollar for every token on top your naira card. You pay for di VPS every month, and dat na all.

Wetin "ternary" mean for dis model

For normal model, every weight na 16-bit number. For ternary model, every weight fit be only one of three values: -1, 0 or +1. PrismML share di weights into groups of 128, and every group carry one FP16 (16-bit floating point) scale. Dem call am "ternary g128". Di scale tell di runtime how big di -1 and +1 for dat group really be.

Di base model na Qwen3.8-27B, with 27.36 billion parameters, and di license na Apache 2.0. Na hybrid attention model: about 75% of di layers use linear attention and about 25% use full attention. Dis one matter for RAM later, because only di full-attention layers keep a KV (key-value) cache wey grow as di conversation dey long.

PrismML ship two GGUF files for CPU and GPU, as di model card show am on 2026-10-03:

  • Ternary-Bonsai-2-27B-PTQ1_0.gguf, 5.95 GB. E pack di three-value weights tight, about 1.75 bits per weight inside di file.
  • Ternary-Bonsai-2-27B-PQ2_0.gguf, 7.21 GB. E put every weight inside 2-bit slot, about 2.13 bits per weight inside di file. E bigger, but di kernels wey read am dey simpler.

PrismML talk say di model keep 98.2% of di FP16 score across 14 thinking-mode benchmarks: average 84.78 against 86.32 for di full FP16 model. Na PrismML claim be dat. We never run dose benchmarks ourselves, so treat am as di maker own number.

Why Ollama and normal llama.cpp no fit load dis file

Every tensor inside GGUF (di file format wey llama.cpp dey use) carry one type number. Di runtime look dat number to know how to read di bytes. PQ2_0 na type 142 for di PrismML fork. Mainline llama.cpp no get any type 142, so e no fit read di weights at all.

Ollama carry mainline llama.cpp code inside am. Dat na why di usual way to import GGUF into Ollama no go help you here. You fit write di Modelfile well well, but di engine wey go read di file still no sabi type 142. PrismML own known-issues page talk am direct: "Stock llama.cpp, Ollama and LM Studio (GGUF) can't load the PQ2_0 or PTQ1_0 files."

No try di F16 file as escape road. E dey 53.8 GB, and PrismML talk say e go load for stock llama.cpp but e go bring rubbish output, because e depend on metadata wey only di PrismML build dey apply.

As of 2026-10-03, mainline llama.cpp no support PQ2_0 or PTQ1_0. Dat one fit change later. Di fork README already list one group-64 format, Q2_0_g64, wey mainline fit run, but di 27B repo no ship any file for dat format. So check di model page again before you build. If mainline or Ollama don add dis support, you fit skip di fork.

Which tag you go build, and why you no go use branch head

Di PrismML fork get three lines wey fit confuse you:

  • prism na di main branch. Dem dey develop am as prism-v7.
  • prism-v6 na old snapshot wey di README tell you make you avoid.
  • prism-v5 don freeze. Di last release for dat line na prism-b9601, and di README keep am for di older Q2_0 files. No use am for Bonsai 2.

Di tag wey we go use na prism-b10743-adfffbe, wey dem release on 2026-09-25. We pick am because PrismML own Bonsai-demo repo pin dis exact tag for Bonsai 2. E also carry two fixes wey matter for CPU. Di first one na pull request #206 (merged 2026-09-21), wey add AVX2 and AVX-VNNI kernels for PQ2_0. Di second one na pull request #245 (merged 2026-09-23), wey stop PQ2_0 from crashing for CPUs wey get AVX-512 (Advanced Vector Extensions, 512-bit). One newer tag, prism-b10754-2459f68 from 2026-10-02, dey there too, but di changes inside am na mostly GPU work.

Why fixed tag and no be branch head? Di head of prism dey move almost every day. If you build head today and rebuild next month, na two different programs you get, and you no go know which commit bring new problem. Fixed tag mean say di build wey work today go still be di same build tomorrow. Our guide to pin llama.cpp releases on a server explain di full habit.

Check your CPU before you build

Di fast PQ2_0 kernels need AVX2 at least. Run dis one:

grep -o -w -E 'avx2|avx_vnni|avx512f|avx512_vnni' /proc/cpuinfo | sort -u
nproc
free -h
df -h ~

You suppose see avx2 for di list. If you see avx512f too, your CPU get AVX-512. Before pull request #245, PQ2_0 dey segfault while e dey load for dis kind CPU, and PrismML list "some server and virtualised CPUs" among di ones wey e affect. Di tag for dis guide carry di fix, so na only older tags go crash like dat.

free -h show your total RAM, and df -h ~ show your free disk. You need about 15 GB free disk: 7.21 GB for di model, plus di source code and di build.

Build di PrismML fork for CPU

Install di build tools, clone di exact tag, and confirm say na di right commit:

sudo apt update
sudo apt install -y build-essential cmake git wget libssl-dev libcurl4-openssl-dev
git clone --depth 1 --branch prism-b10743-adfffbe https://github.com/PrismML-Eng/llama.cpp.git llama.cpp-prism
git -C llama.cpp-prism log -1 --format=%h

Di last command suppose print adfffbe. If e print another hash, you no dey on di tag wey dis guide use.

Now configure and build. GGML_NATIVE=ON tell di compiler make e use every instruction wey your own CPU get. BUILD_SHARED_LIBS=OFF put everything inside one binary, so you fit copy am comot from di build folder:

cd llama.cpp-prism
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON -DBUILD_SHARED_LIBS=OFF
cmake --build build -j"$(nproc)" --target llama-server llama-bench
sudo install -m 755 build/bin/llama-server /usr/local/bin/llama-server-prism
sudo install -m 755 build/bin/llama-bench /usr/local/bin/llama-bench-prism
llama-server-prism --version

We give di binaries new names, llama-server-prism and llama-bench-prism, so dem no go clash with mainline llama.cpp if you install am later. --version suppose print one line wey get adfffbe inside. Di build number for front fit look small, because shallow clone no carry di full git history wey di build dey use count commits.

PQ2_0 or PTQ1_0: which one for CPU?

Use PQ2_0. Di fork README describe am as "preferred on Metal, CUDA, HIP and CPU". Di reason na speed. PrismML known-issues page, wey dem last check on 2026-09-23, still list di fast x86 kernel for PTQ1_0 as "in review". So for x86 VPS, PTQ1_0 save you about 1.2 GB of RAM, but e fit run slower.

Download di model with wget -c. Di -c mean say if di network cut, you fit run di same command again and e go continue from where e stop:

sudo mkdir -p /opt/bonsai/models
cd /opt/bonsai/models
sudo wget -c https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PQ2_0.gguf
ls -l /opt/bonsai/models

ls -l suppose show di file at about 7,210,000,000 bytes. If e small pass dat one, di download never finish. Run di wget -c again.

How much RAM Ternary Bonsai 2 27B need? Di full sum

First, one trap. Hugging Face show file size for GB (1,000,000,000 bytes). free -h show RAM for GiB (1,073,741,824 bytes). 7.21 GB na only about 6.7 GiB. All di numbers below na GiB, so dem go match wetin free -h show you.

Di sum get four parts:

  • Di weights. Di whole file go enter RAM.
  • Di KV cache. PrismML Bonsai-demo README give about 6.3 GiB of FP16 KV cache for 100,000 tokens. Dat one na about 64 KiB for every token. Di cache grow only for di full-attention layers, and dat na why e small pass wetin you go expect for 27B model. If you set di cache type to q8_0, e go reduce to a little more than half.
  • Compute buffers. Dis na working memory wey llama-server keep for di maths. We budget 0.5 GiB. Na estimate be dis, so later we go show you di log line wey give di real number for your own box.
  • Di operating system. Ubuntu, SSH and small page cache. We budget 1 GiB.
ChartEstimated RAM for Ternary Bonsai 2 27B (GiB)
The data behind this chart
[
  {
    "config": "PTQ1_0, 8K context, f16 KV",
    "weights_gib": 5.5,
    "kv_cache_gib": 0.5,
    "buffers_gib": 0.5,
    "os_gib": 1.0,
    "total_gib": 7.5
  },
  {
    "config": "PQ2_0, 8K context, f16 KV",
    "weights_gib": 6.7,
    "kv_cache_gib": 0.5,
    "buffers_gib": 0.5,
    "os_gib": 1.0,
    "total_gib": 8.7
  },
  {
    "config": "PQ2_0, 32K context, q8_0 KV",
    "weights_gib": 6.7,
    "kv_cache_gib": 1.1,
    "buffers_gib": 0.5,
    "os_gib": 1.0,
    "total_gib": 9.3
  },
  {
    "config": "PQ2_0, 32K context, f16 KV",
    "weights_gib": 6.7,
    "kv_cache_gib": 2.0,
    "buffers_gib": 0.5,
    "os_gib": 1.0,
    "total_gib": 10.2
  }
]

Wetin dis one mean for di VPS wey you go rent:

  • 8 GB VPS. Even PTQ1_0 with small context need about 7.5 GiB, and 8 GB plan dey usually show you about 7.7 GiB or less inside free -h. E go swap, or di kernel go kill di process. No go dis road.
  • 12 GB VPS. PQ2_0 with 8K context need about 8.7 GiB. E go run, with small space left.
  • 16 GB VPS. Dis na di size wey we recommend. PQ2_0 with 32K context and q8_0 KV cache need about 9.3 GiB. Even FP16 cache for 32K, about 10.2 GiB, still leave space.

If you wan understand dis kind sum for any other model, our guide on how big LLM fit enter your RAM break am down step by step. And our post on KV cache quantization explain wetin q8_0 cache dey cost you for quality.

Run llama-server for 127.0.0.1 only

Start di server by hand first, so you fit see di errors with your own eye:

llama-server-prism -m /opt/bonsai/models/Ternary-Bonsai-2-27B-PQ2_0.gguf \
  --host 127.0.0.1 --port 8080 -np 1 -c 8192 -fa on \
  -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95

Wetin di flags dey do:

  • --host 127.0.0.1 mean say only programs inside di same server fit reach am. Nobody from internet fit talk to am.
  • -np 1 give one slot, so di whole context belong to one conversation.
  • -c 8192 set di context to 8,192 tokens. Raise am to 32768 if you get 16 GB.
  • -fa on switch on flash attention. Di q8_0 V cache need am.
  • -ctk q8_0 -ctv q8_0 make di KV cache smaller. No use q5_0, because PrismML known-issues page talk say e dey "several times slower" than di other cache types.
  • --temp 1.0 --top-p 0.95 na di sampling values wey di model card use.

For another terminal, check am:

curl -s http://127.0.0.1:8080/health
curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"Explain swap memory in two sentences."}]}'
ss -ltnp | grep 8080

/health suppose answer {"status":"ok"} once di model don load. Di second command go return JSON with di model answer inside choices. ss suppose show 127.0.0.1:8080. If e show 0.0.0.0:8080, di server dey open to di whole internet, so check your --host again.

To use am from your laptop, open SSH tunnel and point your app to http://127.0.0.1:8080 for your laptop:

ssh -N -L 8080:127.0.0.1:8080 youruser@your-vps-ip

Our full guide to run llama.cpp server on a VPS cover API keys and reverse proxy if you need more than tunnel.

Make di server start by itself with systemd

Create user wey no fit login, then write di unit file:

sudo useradd --system --home-dir /opt/bonsai --shell /usr/sbin/nologin bonsai
sudo chown -R bonsai:bonsai /opt/bonsai
sudo tee /etc/systemd/system/bonsai.service > /dev/null <<'EOF'
[Unit]
Description=Ternary Bonsai 2 27B on llama-server (PrismML fork)
After=network.target

[Service]
User=bonsai
ExecStart=/usr/local/bin/llama-server-prism -m /opt/bonsai/models/Ternary-Bonsai-2-27B-PQ2_0.gguf --host 127.0.0.1 --port 8080 -np 1 -c 8192 -fa on -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95
Restart=on-failure

[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now bonsai
journalctl -u bonsai -n 200 --no-pager | grep -i 'buffer size'

Di grep go show di lines wey llama-server print when e allocate memory, like di KV cache buffer and di compute buffer, all for MiB. If e show nothing, di model never finish to load. Wait some seconds and run di journalctl line again. Na di real numbers be dat for your box. Put dem inside di sum from di last section instead of our estimate.

How fast e go run for x86 CPU?

Di speed wey PrismML put for front of di model card na GPU and Apple numbers, like 47.0 tokens per second for Apple M5 Max with Metal. Dose numbers no tell you anything about x86 VPS. Apple chip and big GPU get memory bandwidth wey VPS no fit reach.

Di only x86 CPU number wey we find, as of 2026-10-03, dey inside PrismML pull request #206. Dem test PQ2_0 for one Intel i7-13620H laptop with single-channel DDR5-5200 memory, and dem get 3.9 to 4.2 tokens per second for decode and 16 tokens per second for prompt processing. Laptop no be VPS. Your vCPUs fit share di same physical cores with other customers, and di memory channels no be your own.

Decode speed follow memory bandwidth, because for every new token di CPU must read almost all 6.7 GiB of weights from RAM. So more cores no go help much once di memory don reach im limit. Measure your own box:

llama-bench-prism -m /opt/bonsai/models/Ternary-Bonsai-2-27B-PQ2_0.gguf -t 2,4,8 -p 512 -n 128

Stop di systemd service first with sudo systemctl stop bonsai, or di two programs go fight for RAM. Di output na one row for every thread count and every test. pp512 na prompt processing speed. tg128 na decode speed, di number wey you go feel when di answer dey come out. Look di t/s column. Put di thread count wey give di best tg128 inside your service with -t.

Wetin fit break, and di sign wey you go see

Di process just die. Run sudo dmesg | grep -i 'out of memory'. If you see Out of memory: Killed process with llama-server-prism inside di line, di kernel kill am because RAM finish. Reduce -c, use q8_0 cache, or move to bigger plan.

Illegal instruction after your provider move your VPS. GGML_NATIVE=ON build di binary for di exact CPU wey you get during di build. If di provider move your VM go host with older CPU, di binary go try instruction wey di new CPU no get, and di kernel go stop am. Rebuild for di new host.

Segfault while e dey load, for CPU wey get AVX-512. You dey run tag wey old pass pull request #245. Run llama-server-prism --version and confirm say adfffbe or newer commit dey inside.

Di first answer dey take long before e show. Qwen3.8 dey think before e answer, and every thinking token cost you time for CPU. Our guide to reasoning effort settings for local LLMs show how to control am. PrismML known-issues page also warn say --reasoning on for di server go override client wey send reasoning_effort: "none", so leave di server for --reasoning auto.

FAQ

Fit I run Ternary Bonsai 2 27B for 8 GB VPS?

E no go work well. Even di smaller PTQ1_0 file need about 5.5 GiB for weights, and when you add KV cache, compute buffers and di operating system, di total pass 7 GiB. 8 GB plan dey usually give you about 7.7 GiB or less, so di server go swap or di kernel go kill am. Use 12 GB at least, and 16 GB if you want 32K context.

Why Ollama no fit load di Ternary Bonsai 2 GGUF file?

Ollama carry mainline llama.cpp code, and mainline no know di PQ2_0 and PTQ1_0 tensor types wey PrismML create. PrismML own known-issues page talk say stock llama.cpp, Ollama and LM Studio no fit load dis files. You need di PrismML fork of llama.cpp until mainline add di types, and as of 2026-10-03 e never add dem.

Which llama.cpp tag go load Ternary Bonsai 2 27B?

PrismML Bonsai-demo repo pin prism-b10743-adfffbe from di PrismML-Eng/llama.cpp fork, and dat tag carry di AVX-512 crash fix and di AVX2 kernels for PQ2_0. No use prism-b9601, because na di frozen prism-v5 line for older files. No build from branch head, because di code dey change every day.

How many tokens per second Ternary Bonsai 2 27B go give me for CPU?

Nobody don publish VPS number as of 2026-10-03. Di only x86 figure na 3.9 to 4.2 tokens per second decode for one Intel i7-13620H laptop, inside PrismML pull request #206. Run llama-bench-prism with -p 512 -n 128 for your own server and read di tg128 row, because your memory bandwidth decide di answer.