Which AI Models Fit My RAM for Self-Hosting?
Use your real RAM take choose model: see sizing for 4 GB, 16 GB and 64 GB VPS, honest CPU token rates, plus the hidden context-window cost.
Wetin dey decide which AI models you fit self-host
Na one number dey decide which AI models you fit self-host: the RAM wey dey for the machine. Model family and framework matter far less than whether the weights fit inside memory with some space remaining. This post na the arithmetic wey you need to calculate am. Installing runtime na separate task, and the guide to running Ollama on a VPS cover am.
Two costs dey decide the answer. The weights na fixed cost, based on parameter count and quantisation. Context window na the ongoing cost, and na the one people dey forget until model wey load yesterday refuse load today.
The sizing arithmetic: bits per parameter
Model file nearly dey consist only of weights. Each weight dey store with some number of bits. Quantisation mean say dem dey store weights with fewer bits than the precision wey dem train with. This one reduce small accuracy, but e save plenty memory. The size follow directly from this:
weights in GB = (parameters in billions x bits per weight) / 8Dem dey release models with 16 bits, wey be 2 GB for every billion parameters. Na why almost nobody dey run release precision for VPS. Na these quantisations you go actually meet, with their real average bits per weight:
Q8_0dey store about 8.5 bits per weight, so e dey roughly 1.1 GB for every billion parameters.Q6_Kdey store about 6.6 bits, so e dey roughly 0.83 GB for every billion.Q5_K_Mdey store about 5.7 bits, so e dey roughly 0.71 GB for every billion.Q4_K_Mdey store about 4.8 bits, so e dey roughly 0.6 GB for every billion.
Use 0.6 GB for every billion parameters as your working number. Q4_K_M na sensible default for box wey memory bind. Quality loss compared with 8 bits small for most tasks, and the file nearly half the size. Below 4 bits, the loss dey increase quickly. So, 70B wey dem squeeze to 2 bits usually dey answer worse than 32B for 4 bits from the same generation. When memory tight, reduce one size class before you go below 4 bits.
The data behind this chart
[
{
"label": "3B",
"weights_gb": 1.8,
"kv_8k_gb": 0.9,
"kv_128k_gb": 14
},
{
"label": "8B",
"weights_gb": 4.8,
"kv_8k_gb": 1,
"kv_128k_gb": 16
},
{
"label": "14B",
"weights_gb": 8.4,
"kv_8k_gb": 1.5,
"kv_128k_gb": 24
},
{
"label": "32B",
"weights_gb": 19.2,
"kv_8k_gb": 2,
"kv_128k_gb": 32
},
{
"label": "70B",
"weights_gb": 42,
"kv_8k_gb": 2.5,
"kv_128k_gb": 40
}
]The weight column above na that 0.6 GB for every billion rule wey dem apply. Real GGUF files dey fall within a few percent of am because embedding and output layers dey keep for higher precision than the rest. 3B model for 4 bits na about 1.8 GB. 8B na 4.8 GB. 32B na 19.2 GB, and 70B na 42 GB.
Why context length dey cost more RAM than weights
The KV cache (key value cache, wey be the attention state wey model dey keep for every token wey currently dey conversation) na the second cost. Dem dey allocate am when model load, size am based on the context length wey you request, and e dey grow straight with that length.
The KV cache formula, and where to read the numbers
bytes per token = 2 x layers x kv_heads x head_dim x bytes per elementThe 2 dey count the key and the value. The values for layers, kv_heads (listed as num_key_value_heads) and head_dim dey inside config.json for the model card page. Bytes per element na 2 for a 16 bit cache. One normal 8B model get 32 layers, 8 key value heads and head dimension of 128, so 2 x 32 x 8 x 128 x 2 = 131072 bytes, wey be 128 KiB for each token.
For Ollama default context, that 8B model dey use half gigabyte for cache. For 8192 tokens e dey use 1 GB. For the 128k context wey e advertise for model card, e dey use 16 GB, wey pass three times the weights. The 70B na the opposite case: e cache for 128k na 40 GB, less than the weights, because grouped query attention stop the per-token cost from growing anywhere near as fast as parameter count.
Ollama default context length na 4096 tokens for server wey dey use CPU only. When GPU dey present, e dey choose default from VRAM instead: 32k between 24 and 48 GiB, and 256k for 48 GiB and above. Raise am with the OLLAMA_CONTEXT_LENGTH variable for server, then check wetin running model actually get for CONTEXT column inside ollama ps. The memory calculation behind that setting dey explained inside the post about num_ctx and context length.
You fit reduce the cache in two ways. Request the context wey you need instead of the context wey model card advertise, because most chat and coding work dey fit inside 8k to 32k. Or quantise the cache itself to 8 bits, wey go cut am by half, but e fit reduce recall for long context.
Model wey dey resident go hold RAM until something unload am
Ollama dey keep model for memory for 5 minutes after the last request, then e go unload am. This default fit laptop, but e no correct for server, because the first request after every idle period go pay the load time again.
ollama ps
ollama stop qwen3:4bollama ps dey list the models wey dey resident. E get SIZE column wey show how much memory each model dey hold, and UNTIL column wey show when e go expire. To pin model permanently, set OLLAMA_KEEP_ALIVE=-1 for the service. Value of 0 go unload am as soon as each response finish.
sudo systemctl edit ollama.service[Service]
Environment="OLLAMA_KEEP_ALIVE=-1"
Environment="OLLAMA_CONTEXT_LENGTH=8192"sudo systemctl daemon-reload
sudo systemctl restart ollamaSend one prompt, then run ollama ps again 10 minutes later. The model still dey listed. Na exactly this be the point: e dey hold that RAM whether anybody dey use am or not. Pinned model no be spare capacity. For 16 GB VPS, 8B with 8k context fit hold roughly 6 GB for as long as the service dey run. So size the box based on the model plus your application, no be the model alone. Pin model for memory explain the trade-off against cold start latency.
Wetin fit run for a 4 GB VPS
Keep about 1 GB for the operating system and the model server. That one leave roughly 3 GB. E fit carry 1B to 4B model for 4 bits, with the default 4096 token context. As of August 2026, this class include Llama 3.2 for 3B, Qwen 3 for 1.7B and 4B, plus the small Gemma and Phi releases. Take these as size examples, no be recommendations. The names dey change every few months, but the arithmetic no change.
Expect roughly 6 to 14 tokens per second. Models wey small like this dey do narrow tasks well: classification, tag extraction, short summaries, and rewriting one paragraph to match house style. Dem no strong for multi step reasoning or code wey span several files, and prompting no fit solve that.
The main failure for this tier na swap. If the model no fit inside memory, Linux no go refuse to load am. Instead, e go page memory out to disk. Because generating one token dey read every weight once, generation go slow reach seconds per token. Monitor free -h and the si and so columns for vmstat 1 while the model dey answer. If swap in and swap out no be zero during generation, e mean say the model too big for the plan.
Wetin fit run for 8 to 16 GB VPS
Na for here self-hosted model dey become generally useful. For 8 GB, you fit run 7B or 8B for 4 bits, about 4.8 GB of weights, with 8k context. For 16 GB, you fit run 13B or 14B for 4 bits, about 8.4 GB, or keep 8B for 8 bits if you prefer use the memory for precision instead of parameter count.
Speed na the main problem. 8B for CPU dey generate about 3 to 7 tokens per second, while 14B dey generate about 1.5 to 3.5. Person dey read around 5 to 10 tokens per second, so 8B for CPU VPS go feel like say you dey watch slow typist. E good for background job, but e fit tire person for interactive chat. Measured runs of Qwen 3 for 8B and bigger models for VPS show how e dey work for real life.
Wetin fit run for 32 to 64 GB VPS
A 32B for 4 bits na about 19.2 GB, so e fit enter 32 GB plan with short context, and e go get enough space for 48 GB or 64 GB. A 70B for 4 bits na about 42 GB, so e need 64 GB before you add any cache at all.
Then read the speed as e really be. A 32B for CPU dey run around 0.6 to 1.5 tokens per second, while 70B dey run around 0.2 to 0.5. One 500 token answer from that 70B fit take about twenty minutes. For that speed, request usually go fail before model finish, because one client or proxy timeout wey dey before Ollama go trigger first. Na from there context deadline exceeded error dey come. These tools dey work for batch processing. Feed dem queue of documents overnight, and speed no too matter. Put dem behind chat window, and speed matter well well.
Mixture of experts routing dey change this calculation, and na the only architecture detail wey worth learning. An MoE model dey send each token through only small part of its weights. Model wey get 30B total parameters and 3B active per token need memory for 30B, but e dey generate almost near the speed of dense 3B, because each token dey read only the active experts. For 32 GB machine, MoE with this kind setup dey much more usable than dense 30B. The rule wey you suppose remember be say: total parameters dey set the memory, while active parameters dey set the speed.
How fast CPU inference dey, honestly?
To generate one token, system need read every active weight from memory one time. Nothing fit avoid this, so generation speed for CPU depend on memory bandwidth, no be core count. The maximum na usable memory bandwidth divide by weight size for bytes. Small shared VPS normally fit deliver 10 to 25 GB per second across e vCPUs, so 4.8 GB model fit reach around 2 to 5 tokens per second.
The data behind this chart
[
{
"label": "3B",
"tokens_per_second_low": 6,
"tokens_per_second_high": 14
},
{
"label": "8B",
"tokens_per_second_low": 3,
"tokens_per_second_high": 7
},
{
"label": "14B",
"tokens_per_second_low": 1.5,
"tokens_per_second_high": 3.5
},
{
"label": "32B",
"tokens_per_second_low": 0.6,
"tokens_per_second_high": 1.5
},
{
"label": "70B",
"tokens_per_second_low": 0.2,
"tokens_per_second_high": 0.5
}
]Na ranges wey people commonly report for ordinary VPS hardware, no be benchmark for one particular machine. Your own result depend on memory generation, number of channels for the host, and how many neighbours dey compete for am. Measure your own result with any model tag wey you already get:
ollama run qwen3:4b --verbose "Write three sentences about disk latency."The summary wey e print after answer finish get one line wey dey read eval rate: ... tokens/s. Na your generation speed be that. Ignore the first run for one session, because load duration for the same summary include reading the weights from disk. How to measure tokens per second well explain how you fit get number wey worth compare.
Two results dey surprise people for here. Adding vCPUs stop to help quickly, because after roughly 8 cores, the extra cores dey wait for memory instead of doing arithmetic. For shared plan, the same command fit return different numbers from one hour to another. Na CPU steal time from noisy neighbour cause this, no be anything wey you configure wrongly.
Reading your prompt na different work from generating the answer. Prompt processing dey limited by compute, so e scale with core count, and na here GPU get the biggest advantage. Long document fit take CPU minutes to read, but GPU fit do am in seconds. Na the first wall you go hit when you point coding agent to model wey you host, because every turn resend the file context and tool definitions before even one answer token come back.
Wetin change when you add GPU
The arithmetic no change; na only the pool wey e apply to change. VRAM na hard limit, so calculate wetin go fit before you rent am:
- 8 GB VRAM fit hold 7B or 8B for 4 bits with short context.
- 16 GB fit hold 14B for 4 bits with real context, or 8B for 8 bits.
- 24 GB fit hold 32B for 4 bits when you keep the context short.
- 48 GB and above fit hold 70B for 4 bits with space for cache and concurrency.
When model no fit, Ollama go split am: some layers for GPU, the rest for CPU. ollama ps reports the split for e PROCESSOR column, with something like 78%/22% CPU/GPU. Take that as warning, no be feature. The CPU half dey set the speed, because every token still dey wait for those layers. So, if one-quarter of the model layers dey for CPU, the model go run much closer to CPU speed than GPU speed. If you see split wey you no plan, reduce the context length first. Na usually cache push am pass the limit.
Concurrency na another reason to choose bigger size. Requests wey dey run at the same time share the weights, but every active request need e own KV cache. So ten users wey dey use 8B with 8k context at the same time need ten times 1 GB cache on top of the weights. How to serve concurrent users from one self-hosted model explain where that limit dey land.
Whether GPU worth renting na arithmetic question too. E depend on how many tokens you really generate every month. The break-even point between GPU VPS and API tokens get those numbers.
Wetin you no fit self-host
Two different walls dey here, and e help to know which one you dey face.
The first one na closed weights. Frontier commercial models no dey distributed, so no file dey to download and any amount of RAM no go change that. You fit self-host everything around dem: the interface, the retrieval layer, the agent loop, and the logs. But the model itself go remain a remote API. Whether you fit self-host Claude explain this fully.
The second one na open weights wey just too big. The biggest open releases na mixture of experts designs with hundreds of billions of total parameters. The same rule apply to dem: a 400B total parameter model for 4 bits need about 240 GB for weights alone, before any cache. Na specialist hardware be this, and to rent am monthly cost far more than wetin most people dey spend on API tokens for one year. Wetin e take to self-host Kimi class model go through the real requirement. You go see the same split inside Ollama own library, where GLM 5.2 dey listed only as cloud model and na one much smaller sibling actually dey download onto a VPS.
The honest line between both: self-host when the load steady and the data no suppose leave your server. Buy tokens when the load dey bursty, or when frontier answer quality na the thing you really need.
Check wetin you get before you choose
free -h
nproc
lscpu | grep 'Model name'Plan against the available column for free -h, no be the total column, because total include memory wey system dey already use. Remove about 1 GB for the operating system and the model server. Divide wetin remain by 0.6 to get the biggest parameter count for billions wey you fit hold at 4 bits. Then subtract the KV cache for the context wey you really want. Wetin remain na your answer, and unlike list of model names, e no dey become outdated.
FAQ
How much RAM I need to run an 8B model?
About 4.8 GB for the weights with 4 bit quantisation, plus KV cache for your context length, plus roughly 1 GB for the operating system and the model server. For 8192 token context, the cache go add about 1 GB, so 8 GB plan go work but 4 GB plan no go work. If you want the full 128k context wey the model card advertise, the cache alone na 16 GB, so you dey look at 32 GB plan.
Why my model slow even though the VPS get plenty vCPUs?
Because generation dey limited by memory bandwidth, no be by cores. Every token need pull the whole active weight set from RAM, so once few cores saturate the memory channels, the remaining ones just dey wait. Swap na the other common cause. If vmstat 1 show non zero si and so while the model dey answer, the weights no fit inside RAM and part of every token dey served from disk. This one cost far more than e suppose look.
Longer context window really need more memory?
Yes, and the growth dey linear in tokens. Typical 8B model dey use about 128 KiB of KV cache per token, so 8192 tokens cost 1 GB and 131072 tokens cost 16 GB. The cache dey allocated when the model loads, instead of when the conversation dey grow. So if you ask for 128k context, the system reserve that memory immediately, even if every prompt wey you send na 200 tokens long.
I suppose run large model at 2 bits or smaller one at 4 bits?
Use the smaller model at 4 bits. Quality dey reduce slowly from 8 bits down to 4, but e dey reduce quickly below 4. So 70B squeezed to 2 bits usually give worse answers than 32B at 4 bits from the same model generation. Heavy quantisation dey show as repetition and missed instructions, instead of error message, so e easy to blame your prompt. Treat 4 bits as the floor and change the parameter count instead.
I fit self-host model wey get the same ability as the big commercial ones?
No be for ordinary VPS. The strongest open weight models get hundreds of billions of parameters. At 4 bits, this mean over 200 GB of RAM before any KV cache, and the strongest commercial models no dey distributed at all. Ordinary hardware dey do well when e run good 8B to 32B model for one specific job, where narrow and well prompted small model fit often match general one. If you need frontier quality, compare the API price with the hardware cost before you buy either one.