Run Gemma 4 on a VPS: sizing and speed
Which Gemma 4 tag fits your VPS RAM, what each one costs in tokens per second on CPU, and the sizing arithmetic to check it before you pull 20GB.
Filtering by topic #sizing · clear
Which Gemma 4 tag fits your VPS RAM, what each one costs in tokens per second on CPU, and the sizing arithmetic to check it before you pull 20GB.
k3s puts the real Kubernetes API on one VPS. What it costs in RAM, why port 80 collides at install, and when Docker Compose is still the right answer.
Pick a model by the RAM you actually have. Sizing arithmetic for 4 GB, 16 GB and 64 GB VPS plans, honest CPU token rates, and the hidden cost of context.
One always-on coding agent runs on 4 GB and 2 vCPU. The builds and language servers it launches are what fill the box, and what makes it hang.
Immich documents 6 GB of RAM as its minimum. Here is what the server, Postgres, Redis and machine learning each need, and how to run it on 4 GB.
Docker on a VPS is the same engine with less room: RAM runs out, published ports skip UFW, containers stay down after a reboot, and disk fills up.