How I Fit Limit VPS RAM and CPU with systemd
Runaway process fit freeze your VPS. Set MemoryHigh, MemoryMax, CPUQuota and TasksMax, then confirm cgroup v2 values and check the OOM kill.
Limit process memory and CPU with a systemd drop-in
You fit limit process memory and CPU for Linux VPS by adding small number of lines to the unit wey dey run the process. MemoryMax= na the hard ceiling for memory. CPUQuota= na the ceiling for processor time. cgroup v2 (control groups, version 2) dey enforce both. Na kernel feature be this, and systemd already dey use am to account for every service for the server.
sudo systemctl edit myapp.serviceThat one go open drop-in file wey get instructions inside comments. Add this above dem:
[Service]
MemoryHigh=512M
MemoryMax=768M
MemorySwapMax=0
CPUQuota=80%
TasksMax=128sudo systemctl daemon-reload
sudo systemctl restart myapp.service
systemctl show myapp.service -p MemoryHigh -p MemoryMax -p CPUQuotaPerSecUSec -p TasksMaxsystemctl show suppose repeat your numbers back with the kernel own units: MemoryMax=805306368 and CPUQuotaPerSecUSec=800ms. If e print MemoryMax=infinity, drop-in no load at all. Check say file dey for /etc/systemd/system/myapp.service.d/override.conf, and say e start with [Service] header. If settings line no get section above am, systemd go log Assignment outside of section. Ignoring. and start the service without any limit.
The rest of this guide explain how to choose those numbers, and wetin fit still go wrong after you set dem.
Why runaway process fit freeze VPS even when e no full am
Process wey hit hard memory cap go die for about one second, then service go restart. Na the good case be that. The bad case na when nothing die: the box dey answer ping, SSH dey accept connection, but shell prompt no dey show. Machine dey alive and busy, but none of that work dey useful.
This na how e dey happen, because e no obvious. When free memory dey low, kernel dey reclaim pages instead of giving out new ones. File-backed pages dey cheapest to reclaim, and page cache dey hold executable code for everything wey dey run. So kernel go evict the text pages of sshd, and the next instruction wey sshd run go be page fault wey must read those bytes back from storage. Every process go end up waiting for disk instead of running. The same pages go leave and come back for loop. Dem dey call this thrashing.
Two things dey make this worse for VPS pass laptop. Storage often dey network attached or shared, so every fault dey cost more milliseconds than local NVMe device go cost. And kernel no dey measure time; e dey measure failure: as long as reclaim dey return one page, no matter how slow, kernel believe say e dey make progress and e no call out of memory (OOM) killer. Box fit remain for that state for many minutes before anything get killed.
You fit watch am happen. Linux 4.20 and newer dey export pressure stall information (PSI):
cat /proc/pressure/memory
cat /proc/pressure/iosome avg10=63.72 avg60=41.02 avg300=12.33 total=13729481
full avg10=48.15 avg60=30.44 avg300=8.90 total=9114233The full line na the important one. full avg10=48.15 mean say for the last ten seconds, 48% of the time every runnable task for the box dey stalled, waiting for memory work, so nothing run. Healthy server dey read close to zero for full. Above 10, human go feel say e slow. 40 or more na the state wey people dey describe as frozen.
Na this one still explain why limit by itself no be promise. Unit wey dem hold under MemoryHigh= go get throttled instead of getting killed, so e go remain alive and remain slow. Nothing go restart am because from systemd point of view, e never fail. Capped unit wey still get permission to swap go generate reads and writes wey dem charge to that unit, but one shared device go serve dem. So e fit push /proc/pressure/io up for every other service for the box. Limits decide who go pay for shortage; dem no fit create capacity.
Check say your VPS dey run cgroup v2
stat -fc %T /sys/fs/cgroupcgroup2fs na the unified hierarchy, and every setting below need am. tmpfs mean say the box boot with the older v1 layout, where MemoryHigh= and MemorySwapMax= no dey, and per-unit OOM behaviour dey different. Ubuntu 22.04 and later, plus Debian 11 and later, dey use v2 by default. Old image, or kernel wey boot with systemd.unified_cgroup_hierarchy=0, no dey use am.
For cgroup v2, systemd dey turn on memory accounting for every unit by default, so the numbers don already dey available:
systemd-cgtop -mThat one go list cgroups according to memory use, and na the fastest way to answer “wetin dey chop this box” while e still fit respond. If the server new, do the account and firewall work for the first ten minutes for a new VPS before this.
MemoryHigh dey throttle. MemoryMax dey kill.
Difference between the two memory settings na wetin go happen when failure occur.
MemoryHigh=na soft cap. When usage pass am, kernel dey reclaim aggressively from that cgroup and dey slow down its allocations on purpose. Usage still fit pass the number, and nothing go kill.MemoryMax=na hard cap. When allocation no fit succeed under the cap, OOM killer go run inside that cgroup and kill one process wey belong to that unit.
Na that second part make e important to set MemoryMax= for anything wey you no trust completely. Without cap, shortage become problem for the whole machine, and global OOM killer choose victim based on oom_score, wey mostly mean the biggest process. The biggest process normally na your database, no be the script wey leak memory. With cap, the kill go happen inside the unit wey cause am.
Set both, with MemoryHigh= about 20 to 30 percent below MemoryMax=. The gap na warning zone: slow leak go pass High and show as service wey don slow down, while sudden spike go pass straight through Max and die.
Percentage values dey calculate against installed physical memory, so MemoryMax=25% for 4 GB plan na 1 GB and e still remain one-quarter of the machine after you resize the plan. MemorySwapMax=0 keep that unit completely out of swap, wey turn long crawl into fast, obvious kill.
Some services let you decide their memory appetite ahead of time instead of measuring am: Ollama unit dey size its KV cache from the context window wey you give am, so read wetin raising num_ctx cost for RAM before you choose ceiling for one.
A cap need restart policy beside am, otherwise the kill go just leave you with stopped service.
[Unit]
StartLimitIntervalSec=300
StartLimitBurst=5
[Service]
Restart=on-failure
RestartSec=5sStartLimit* belong inside [Unit] and Restart= inside [Service]. If you put either one for wrong section, systemd go ignore am. Five restarts within five minutes mean say na leak, no be temporary blip, so after that systemd go give up and leave the unit failed. Na that state you want find later, instead of crash loop wey dey hide the problem.
CPUQuota use do cap CPU, or CPUWeight use do share am
CPUQuota= dey take percentage of time wey one CPU get available. CPUQuota=50% na half of one core. CPUQuota=200% na equivalent of two cores, and the unit fit spread am across as many threads as e want. For 2 vCPU plan, CPUQuota=200% na the whole machine.
CPUWeight= na the better default for most services. E dey represent relative share from 1 to 10000, and kernel default na 100. E only matter when something dey compete: backup job for CPUWeight=20 go yield to web server for 100 when load dey high, but e still fit use the whole machine when machine dey idle. Hard quota go throw away that idle capacity.
Make you honest about wetin CPU limit fit give you. CPU-bound process rarely go freeze Linux, because scheduler dey continue give everybody time. Na memory dey usually bring machine down. Use CPUQuota= when you want predictable ceiling, like for build or agent wey for otherwise run flat out for one hour. How to size that kind workload na separate matter, and how much RAM and CPU coding agent VPS needs cover am.
If CPU dey show as busy while none of your processes dey do plenty work, the cause fit dey on the other side of hypervisor. Na CPU steal time from a noisy neighbour be this, and no quota wey you set go change am.
TasksMax fork loop stop dey
TasksMax= na the number of processes and threads wey one unit fit hold. Threads dey count too, so Java or Go service need more allowance than wetin process list dey show. Na the cheapest protection against script wey dey fork for loop, because the fork go fail inside the unit instead of making the whole box run out of process IDs.
TasksMax=128When one unit reach the limit, kernel go log one line wey name the cgroup:
cgroup: fork rejected by pids controller in /system.slice/myapp.serviceThe program itself normally report fork: retry: Resource temporarily unavailable. Check wetin manager dey apply by default with systemctl show -p DefaultTasksMax.
Limit one one-off job with systemd-run
You no need unit file to use any of dis. systemd-run dey build one transient unit around one command.
sudo systemd-run --scope -p MemoryMax=1G -p MemorySwapMax=0 -p CPUQuota=50% -p TasksMax=64 ./import-data.sh--scope dey run the command for your terminal, after e print Running scope as unit: run-r7c1a....scope. Output dey remain for your screen, and the limits go disappear when command exit. Any property from systemd.resource-control go work after -p.
For long job, comot --scope and give am name. E go then run for background as transient service and log go journal:
sudo systemd-run --unit=nightly-import -p MemoryMax=1G -p CPUWeight=20 ./import-data.sh
journalctl -u nightly-import -fThe same options dey work with --user when you no be root, but your user manager only get controllers wey dem delegate give am, so e fit reject one property there. Run am with sudo if that one happen. When job get permanent place, move the settings unchanged enter real unit: see how to run script as systemd service and timer.
Swap question wey get honest answer
Swap dey change how failure go happen; e no dey stop failure.
If no swap dey, memory leak go reach limit, and something go die within seconds. Outage go loud, short, and easy to understand for journal afterwards. If swap dey, kernel go write cold anonymous pages go disk and buy more time. If process suppose level off, swap go save you. But if na runaway process, swap go turn five second outage to twenty minute stall. The stall worse, because dead process still leave working shell, but thrashing box no dey.
swapon --show
free -hFor small VPS, workable middle na to keep modest swap file for pages wey system allocate once and never touch again. Set MemorySwapMax=0 for the units wey you fit afford to lose. Important services keep their swap. Unpredictable ones go hit the limit quickly and restart.
Lowering vm.swappiness na weak lever, and you suppose know why. E only shift the balance between evicting page cache and swapping anonymous pages. Both options go cost disk read later. E change which pages dey thrash, not whether the box go thrash.
Early OOM daemon fit kill process before system hang
Kernel dey wait make memory reclaim fail finish, and for small VPS, na this same waiting time dey make you lose the machine. Two userspace daemons fit close this gap by monitoring memory by themselves and killing process earlier.
earlyoom dey monitor available memory and free swap, then e kill process wey get the highest score when either one drop below threshold.
sudo apt install earlyoom
systemctl status earlyoomDebian and Ubuntu package go start the service when installation finish. E options dey inside /etc/default/earlyoom:
EARLYOOM_ARGS="-m 5,2 -s 5,2 --avoid '^(sshd|systemd)$' --prefer '^(node|python3)$'"-m PERCENT set the minimum available memory, while -s PERCENT set the minimum free swap. Both dey 10 percent by default. The second number for each pair na the SIGKILL point. earlyoom go send SIGTERM once memory drop below the first value, then send SIGKILL below the second value. By default, the second value na half of the first one. Apply change with sudo systemctl restart earlyoom, and read journalctl -u earlyoom to see which process e kill and how much memory that process dey hold.
systemd-oomd na the other option. E manual page describe am as "a system service that uses cgroups-v2 and pressure stall information (PSI) to monitor and take corrective action before an OOM occurs in the kernel space". E dey act on complete cgroups instead of individual processes, so e go kill one unit, not stray child process. Units fit opt in with ManagedOOMMemoryPressure=kill or ManagedOOMSwap=kill, and thresholds dey inside /etc/systemd/oomd.conf.
systemctl status systemd-oomd
oomctloomctl go print wetin e dey monitor currently. For server image, this one often na nothing because setting dey opt in for each unit. Choose one daemon and stop there. If you run both, the two of dem go race to choose victim, and e go harder to find out why any process get killed.
Which unit cause am?
Start with the kernel, because e dey record every kill wey e make.
journalctl -k --grep "Killed process" --since "2 hours ago"A kill from the global OOM killer go look like this:
Out of memory: Killed process 4127 (node) total-vm:2731084kB, anon-rss:1874232kB, file-rss:0kB, shmem-rss:0kB, UID:1000 pgtables:4212kB oom_score_adj:0anon-rss na the memory wey that process hold for RAM when e die, about 1.8 GB for here. Read the name inside brackets with suspicion. Na the victim wey kernel choose be that, and kernel dey choose the biggest process. But na not always that process cause the shortage.
A kill from cgroup limit get different prefix, and the report wey print above am name the cgroup wey reach im own limit:
Memory cgroup out of memory: Killed process 8811 (python3) total-vm:1044320kB, anon-rss:769112kB, file-rss:0kB, shmem-rss:0kB, UID:998 pgtables:1720kB oom_score_adj:0That prefix na most of the diagnosis. Memory cgroup out of memory mean say one unit reach the MemoryMax= wey you give am, while the rest of the box dey fine. Plain Out of memory mean say the whole machine run out, so your caps either no dey or dem too generous to add up.
Then ask systemd wetin e see:
systemctl status myapp.service
journalctl -u myapp.service -n 50myapp.service: A process of this unit has been killed by the OOM killer.
myapp.service: Main process exited, code=killed, status=9/KILL
myapp.service: Failed with result 'oom-kill'.systemctl status talk the same thing for one line, as Active: failed (Result: oom-kill).
The cgroup counters na the third source, and na the only one wey record throttling, wey no dey produce log line at all:
cat /sys/fs/cgroup/system.slice/myapp.service/memory.events
cat /sys/fs/cgroup/system.slice/myapp.service/memory.peaklow 0
high 4213
max 118
oom 12
oom_kill 12high count how many times the unit pass MemoryHigh= and get throttled. max count how many times e reach the hard cap, and oom_kill count processes wey kernel actually kill. Big high together with oom_kill 0 na the silent case we talk earlier: the service dey run, e slow reach crawl, and e never report failure to anybody. memory.peak (Linux 5.19 and newer) hold the highest usage wey the cgroup reach. Na this number you use size MemoryMax= against. Both files reset when the unit restart, because systemd create the cgroup again.
One prerequisite dey under all this. If /var/log/journal no exist, the journal dey live for RAM, and every line go disappear after the reboot wey you need to recover the box.
sudo mkdir -p /var/log/journal
sudo systemd-tmpfiles --create --prefix /var/log/journal
sudo systemctl restart systemd-journald
journalctl --list-bootsjournalctl --list-boots showing more than the current boot mean say history don dey survive, so journalctl -k -b -1 fit show you the kernel messages from the boot wey die.
A starting point for one small VPS
For a 2 GB plan, leave 300 to 400 MB for the kernel and page cache. No let all the caps add reach the full 2 GB, because every unit fit reach peak at the same time. Give the service wey matter the biggest share. Then cap everything wey no sure around am.
[Service]
MemoryHigh=256M
MemoryMax=384M
MemorySwapMax=0
CPUWeight=20
TasksMax=64
Restart=on-failure
RestartSec=5sTo keep one way to enter the server, add one more setting. OOMScoreAdjust=-500 for a drop-in under ssh.service go make the global OOM killer much less likely to choose your SSH daemon as victim. This fit be the difference between fixing the server and rebooting am from the control panel. E only change the kernel choice of victim. E no reduce how long the stall go last.
Containers dey run inside their own cgroups. The container runtime create dem, not your unit files. So, limit for docker.service no become limit for one container. The per-container equivalents of MemoryMax= and CPUQuota= dey covered for how to set memory and CPU limits for Docker Compose.
FAQ
Why my VPS freeze instead of killing the runaway process?
Because kernel dey judge progress by whether reclaim return pages, not by how long e take. When memory scarce, e dey evict page cache, including executable pages of programs wey dey run, then e read dem back for the next instruction. Everything dey wait for storage, and no allocation don technically fail, so OOM killer no dey get called. Check /proc/pressure/memory while e dey happen: full avg10 wey pass 40 mean say almost no task run for the last ten seconds. Userspace daemon like earlyoom fit kill before the machine reach that state.
Wetin be the difference between MemoryHigh and MemoryMax?
MemoryHigh= na soft cap wey dey throttle. Kernel dey reclaim memory hard from the unit and slow down its allocations, but usage fit pass the number and nothing go get killed. MemoryMax= na hard cap: allocation wey no fit happen under am go invoke OOM killer inside that unit own cgroup, so the process wey cause the problem na im go die instead of the biggest process for the machine. Set MemoryHigh= below MemoryMax= and treat the gap between dem as warning zone.
How I fit find which service OOM killer hit?
Run journalctl -k --grep "Killed process" --since "2 hours ago". Line wey start with Memory cgroup out of memory mean say one unit hit its own MemoryMax=, while plain Out of memory mean say the whole machine run out of memory. Then run journalctl -u <unit> -n 50 and look for Failed with result 'oom-kill'. If /var/log/journal no dey your server, journal don dey keep for RAM and the evidence die when reboot happen, so create that directory before the next incident.
I suppose add swap to small VPS?
Small swap file fit help with cold pages wey dem allocate once and never touch again. E no help with runaway process: e dey delay the kill and replace short outage with long stall wey you no fit log in fix. Keep swap modest, and set MemorySwapMax=0 on the units wey you fit lose, so dem reach their ceiling and restart quickly while important services keep their swap.
I fit limit command without writing unit file?
Yes. sudo systemd-run --scope -p MemoryMax=1G -p CPUQuota=50% ./script.sh dey run the command for your terminal inside transient scope with those limits, and the limits go disappear when e exit. Every property from systemd.resource-control dey available after -p, so MemorySwapMax=, TasksMax= and CPUWeight= work there too. Remove --scope and add --unit=name to run the job for background with its output inside the journal.