Why Your Self-Hosted LLM Dey Stall at 5 Users
One user dey okay, five dey crawl. See how batching, KV cache limits, prefill, queue depth, and Ollama default parallelism decide your real user limit.
Why self-hosted LLM dey slow down when more users show?
Self-hosted LLM dey stop for 5 concurrent users because server still dey generate one reply at a time, while the other four dey queue. Ollama documentation talk am clear about the default: OLLAMA_NUM_PARALLEL na "the maximum number of parallel requests each model will process at the same time, default 1." Nothing spoil. Four out of your five people dey wait for their turn.
The fix rarely be bigger server. Na serving engine wey fit push many requests through the model for the same forward pass, plus enough spare memory to hold everybody conversation while e dey work. Both parts matter, and na the second one really dey set your limit.
The two phase wey every request dey pass through
Prefill dey read the whole prompt at once and build attention cache for am. Every prompt token dey pass through the model together, so prefill na one big matrix multiply, and arithmetic throughput dey limit am. Decode then dey write the answer one token at a time. For every token, e need read the model full weights from memory again, while the arithmetic wey happen for that one token small well-well. Memory bandwidth dey limit decode.
Na this difference make batching work. Decode for one user fit read, for example, 5 GB of weights for every token, while e leave most arithmetic units idle. Add second request, and the engine read the same 5 GB once, then calculate two tokens from am. The second user almost no add extra time. If you serve requests strictly one after another, you lose this benefit.
Two numbers dey describe wetin user dey experience. TTFT (time to first token) na queue wait plus prefill. ITL (inter-token latency) na the gap between streamed tokens, and decode dey determine am. Slow server usually slow for one of these two, and the fixes no be the same. E make sense establish which one you dey fight before you change any setting, and time prefill and decode separately na how you go find out.
Static batching dey make everybody wait for the slowest reply
Static batching na the basic version, and na wetin you get if you group requests by yourself for application code. The engine go collect N requests, run dem together, and keep every slot busy until the longest generation for the group finish.
One user wey ask for 1,200-token summary fit keep four one-line answers locked inside the batch, because the batch no go free any slot until the slowest member finish.
Two costs dey follow. Sequences wey don finish still dey occupy slots wey no dey do any useful computation, so effective throughput go drop as output lengths dey vary, and chat output lengths dey vary well well. Request wey arrive one step after the batch form go wait for the whole batch to drain before e even start prefill, meaning say na another person essay go determine im TTFT.
Continuous batching dey admit and retire requests for every token
Continuous batching dey schedule work for the level of one decoding step. After every step, scheduler dey remove sequences wey just emit their stop token, then admit waiting requests enter the free slots. Reply wey end for step 40 go free its slot for step 40, no be when batch finish.
This one no be strange thing. llama-server document -cb, --cont-batching as "whether to enable continuous batching (a.k.a dynamic batching) (default: enabled)", and vLLM dey built around this idea. Ollama too fit serve parallel requests. The default just limit the number to one. Na why plenty people dey conclude say their hardware no fit handle concurrency, when na their configuration talk no.
People normally measure published continuous batching results for datacenter cards wey get spare compute and tens of gigabytes for cache. The general shape of those results still apply to your machine. But their size no apply, and the memory section below explain why.
Prefill dey compete with decode for the same compute
When new request land while four replies dey stream, e first need prefill the prompt, and prefill dey use plenty compute. If scheduler give that prefill one step by itself, the four users wey dey stream no go receive any token during that time. For long prompt, every open window go show that pause clearly. Na this stutter people mean when dem talk say server dey hiccup anytime somebody else press send.
Chunked prefill dey break long prompt into pieces, then mix each piece into the same step with the decodes wey dey run. vLLM's tuning guide state the tradeoff directly: smaller chunk budgets "achieve better ITL because there are fewer prefills slowing down decodes", while higher values "achieve better time to first token (TTFT) as you can process more prefill tokens in a batch". You dey choose whose experience to protect: the person wey dey wait for reply to start, or the people wey dey watch text stream.
Prompt length dey decide how much this go affect performance. A 6,000 token prompt with 200 token answer na 6,000 tokens of prefill work against 200 decode steps. Retrieval-augmented chat and long system prompts both fit push you enter this situation, so prefill no longer be small rounding error; na the thing users dey wait for. Prefix caching dey help when the long part dey repeat: vLLM exposes --enable-prefix-caching, wey dey reuse cache for shared prompt prefix instead of recomputing am for every request.
Memory wey go finish first na KV cache
Every token for every active conversation dey leave one key vector and one value vector for every layer of the model. Na KV cache (key/value cache) be this, and na wetin make decode avoid recompute the whole prompt for every new token. The size for each token dey fixed by the model shape: 2 (one key, one value) times the layer count, times the number of key/value heads, times the head dimension, times the bytes per value. Read those numbers from the model's config.json.
Calculate am once, and the limit no go remain mystery again. One normal 8B model wey get 36 layers, 8 key/value heads, and head dimension of 128, with cache for 16-bit, dey use 2 36 8 128 2 bytes for each token. That na 147,456 bytes, about 144 KiB. One 8,192-token conversation therefore need roughly 1.2 GB cache. Five of dem need roughly 6 GB, on top of the weights. Na this be the real answer to how many users fit run.
Concurrency dey multiply context, and the tools dey show am clearly. Ollama FAQ talk say: "Parallel request processing for a given model results in increasing the context size by the number of parallel requests. For example, a 2K context with 4 parallel requests will result in an 8K context and additional memory allocation." RAM wey you need dey scale with OLLAMA_NUM_PARALLEL multiplied by OLLAMA_CONTEXT_LENGTH. For llama-server, the context wey you request with -c dey share among the -np slots. So, if you increase the slot count by itself, the amount each request fit hold go reduce. Read the context for each slot from the startup log instead of assuming am.
vLLM dey preallocate memory instead. --gpu-memory-utilization (default 0.92) na "the fraction of GPU memory to be used for the model executor". The memory wey remain after the weights become the paged KV pool. When that pool no get enough space, the scheduler go evict one request instead of failing am:
WARNING 05-09 00:49:33 scheduler.py:1057] Sequence group 0 is preempted by PreemptionMode.RECOMPUTE mode because there is not enough KV cache space.For vLLM V1 engine, the default preemption mode na RECOMPUTE. So, when e evict one request, e discard the cache and prefill again when e admit the request back. That work go happen two times. The documentation warn say "preemption and recomputation can adversely affect end-to-end latency". This log line na the best single explanation for why one unlucky user wait much longer than everybody else, even when your average latency look healthy. Set disable_log_stats=False to log the cumulative count, or read the preemption counter from the Prometheus metrics wey vLLM expose.
Wetin dey change for 2, 5 and 20 concurrent users
Two users. E almost no show for GPU wey still get cache space, because the second decode stream dey follow the first one with very small extra time. For CPU-only VPS wey get 4 to 8 GB RAM, e no be free: both streams dey share the same few vCPUs and the same RAM bandwidth, so each user go see roughly half the tokens per second, and cache demand go double against much smaller budget.
Five users. Na here default settings stop to dey enough, and e first become queue problem. With OLLAMA_NUM_PARALLEL at 1, four people go wait for whoever request the long answer, and each person go see normal speed as soon as e reach their turn. If you increase the parallel count, the problem go change shape: five slots at 8K context each mean 40K token cache to find space for. If e no fit enter VRAM, the engine go offload layers to system RAM. If e no fit enter RAM, the machine go use swap and tokens per second go collapse.
Twenty users. Twenty people for chat UI usually no be twenty concurrent requests, and na the most useful thing to understand before you buy hardware. Person go read reply and think for 20 to 60 seconds between turns, so most of their session dey idle. Twenty agents, or twenty document summarisation jobs, na twenty real streams with no idle time at all. Na different kind machine be that. One developer wey don point coding agent to their own Ollama server dey closer to the second case than the first, because the agent dey keep sending requests as long as the task dey run and e no get any of the reading pauses wey person dey make.
Users dey make requests at the same time, or dem just dey logged in?
First calculate requests wey dey in flight before you size anything. The arithmetic simple: requests in flight equal users, times seconds wey generation take for each turn, divided by seconds between turns.
- First measure your own speed for one stream, both prefill and decode. No use number from another person card: measure tokens per second for your own box and use the result.
- Estimate the duty cycle. Twenty chat users, 12 seconds of generation for each turn, and one turn every 90 seconds gives 20 * 12 / 90, wey be about 2.7 requests in flight.
- Set the slot count a little above that, then check memory: slots times per-request context must fit inside the cache tokens wey you actually get.
- Keep the queue short so overflow go fail fast and show clearly.
Available cache tokens na free memory after the weights, divided by the per-token cost from the section above. A 24 GB card wey dey run an 8B model for 16-bit spends about 16 GB on weights and get roughly 6 GB usable cache for the default utilisation, wey be around five 8K conversations. To fit more, shorten the per-request context, or store the cache for 8-bit (llama-server takes --cache-type-k q8_0). Both options increase concurrency by giving up something, and you suppose read the honest explanation of this trade before you spend money on hardware: where GPU VPS break even against API tokens.
Where Ollama default no longer enough
Increase the parallel count through the service unit, because shell export no go reach daemon wey systemd dey manage.
sudo systemctl edit ollama.service[Service]
Environment="OLLAMA_NUM_PARALLEL=4"
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_MAX_QUEUE=64"sudo systemctl daemon-reload
sudo systemctl restart ollama
systemctl show ollama --property=Environment
ollama pssystemctl show suppose print the three variables wey you just set. If e no do am, the drop-in no save, and nothing else wey you do go matter. ollama ps then list the loaded model with size wey pass the weights alone, because four slots at 8,192 tokens reserve 32,768 tokens of cache beside dem. If PROCESSOR column show say part of the model dey for CPU when you expect all of am for GPU, e mean say you request more cache than the card get available. Reduce one of the two numbers. Reducing the context usually safer, but window wey too small go quietly truncate long prompts instead of showing error. So e make sense to size num_ctx deliberately instead of reducing am until the model fit.
The queue default need another look. Ollama queue up to OLLAMA_MAX_QUEUE requests, and "the default is 512". After that, e respond "with a 503 error indicating the server is overloaded". Queue wey deep reach 512 for a server wey fit serve four requests at once na promise wey you no fit keep, because client for position 300 go timeout long before na im turn. Short queue go return error wey your application fit retry or report, and that better than spinner wey no dey ever finish.
Test am properly. Send two requests at the same time from two terminals and monitor both. If the second one no produce anything until the first one finish, the parallel setting no take effect.
When real serving engine start to pay for itself
vLLM worth the extra setup when you get GPU wey still get spare capacity and more than around four requests genuinely dey in flight. Its scheduler dey work per token, its cache dey paged so e fit reuse free fragments, and e converts spare VRAM to concurrency instead of leaving am idle. As of August 2026, the documented install and launch na two commands:
uv pip install vllm --torch-backend=auto
vllm serve Qwen/Qwen2.5-1.5B-Instructcurl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Say hello."}]
}'Reply wey contain a choices array mean say server don come up and model don load. Under load, the two knobs wey matter na --max-num-seqs, the "maximum number of sequences to be processed in a single iteration", and --max-num-batched-tokens, the "maximum number of tokens that can be processed in a single iteration". The first one set the concurrency limit. The second one na the chunked prefill budget we describe earlier.
When fewer than around four requests dey in flight, or for any box wey no get supported GPU, vLLM add complexity but give small benefit. E expect a CUDA-class card and e claim most of the memory when e start, so this no be the right trade-off for a 4 to 8 GB VPS. For that kind setup, use smaller model with shorter context and queue wey you control. how Ollama and vLLM differ as serving engines explain the choice fully, and running Qwen 3 8B on a VPS show wetin a mid-sized model need before you add even one extra user.
Wetin folklore dey hide
Continuous batching dey increase total throughput, and e normally improve median latency too, because queued request fit start earlier. Tail latency dey move for the opposite direction, and people rarely dey mention that part.
Every extra sequence for one step dey add small work, so ITL dey rise for everybody as batch dey fill up. When new request arrive, e prefill dey take part of one step wey streaming users for otherwise get. When cache pressure happen, scheduler preempt, and this one send half-generated request back to the beginning of its prefill.
Chat UI dey show tail latency, no be average. Stream wey pause for two seconds for middle of sentence go look like say e break, even when total time to completion good. Measure p95 TTFT and p95 ITL under the load wey you expect, and treat mean tokens per second as capacity number, no be description of the user experience.
The practical setting follow from this. Cap concurrency small below wetin memory fit handle, so engine no go need preempt. Short queue wey predictable better pass deep batch wey dey thrash, because user wey wait four seconds then stream smoothly go happier than user wey start immediately but stall two times.
Wetin to check when e slow
Each user dey normal, but wait time long. Na queue be that, no be speed problem. Check the parallel setting first. The model dey serve correctly, one request at a time.
HTTP 503 from Ollama. Queue full. Either the box don truly reach capacity, or OLLAMA_MAX_QUEUE set low on purpose to shed load, and na wetin you want am to do be that.
Tokens per second dey collapse under load for CPU box. Run vmstat 1 while e dey happen. Nonzero si and so columns mean say machine dey swap, so e dey read weights from disk for every token. No configuration change fit save that. Reduce the model size or slot count.
One user for every ten dey wait far longer than the others. Search the vLLM log for preempted. Preemption and the recompute wey follow am na the usual cause. E mean say cache don get more demand than the context length wey you allow fit handle.
TTFT bad even when server idle. Na prefill be that, no be concurrency. Long prompts dey take real time before the first token show, so check prompt size and prefix caching before you check hardware. If na only the first person wey return after quiet period dey wait long, while everybody after am dey fine, then no be prefill at all. Na Ollama unload the model and dey read the weights from disk again. You fit rule this out by keeping the model resident between requests.
FAQ
Why my self-hosted LLM dey slow down when second person dey use am?
Most times e no really slow down. E dey queue. Ollama dey ship with OLLAMA_NUM_PARALLEL set to 1, so second request go wait until first one emit final token. To tell the two cases apart, time one user stream while another one dey wait: if tokens per second normal once e start, na queue you get, and increasing parallel count go fix am. If both streams dey run at half speed, una dey genuinely share memory bandwidth, and na hardware limit be that.
How many concurrent users fit one small GPU serve?
Count memory, no be users. First na weights, then KV cache. KV cache cost 2 times layers times key/value heads times head dimension times bytes, per token, for each active conversation. Typical 8B model wey get 36 layers, 8 key/value heads, and head dimension 128 cost about 144 KiB per token for 16-bit. So, 8,192 token conversation need roughly 1.2 GB. 24 GB card wey dey hold that model for 16-bit get about 6 GB left for cache. That one fit support about five conversations with full context, or more if you make the context shorter.
Continuous batching dey make each user's reply slower?
Median latency normally dey improve, because requests no longer wait for complete batch to finish. Tail latency dey worse. Every extra sequence add work to every decoding step. New arrival's prefill take part of one step from users wey dey stream. If request get preempted, e must run prefill twice. Measure p95 inter-token latency, no be average. Chat window dey show pauses clearly in a way wey average fit hide.
Make I increase OLLAMA_NUM_PARALLEL or move to vLLM?
Increase parallel count first. E free and e need only one drop-in file. E dey fix the common case where four people queue behind one long answer. Memory na the limit: parallel requests multiply the context wey you must hold, so monitor whether layers dey spill to CPU. Move to vLLM when your GPU get spare VRAM and more than about four requests dey genuinely run at the same time. Na for that point paged cache and per-token scheduling dey give more benefit than cost.
More CPU cores go fix slow LLM server?
No be for the part wey users notice pass. Decode dey read the whole model from memory for every token, so RAM bandwidth dey limit am. Extra cores stop to help once bandwidth don saturate. Prefill dey scale with cores, so more cores fit reduce time to first token for long prompts. For 4 to 8 GB VPS, memory capacity usually na the main constraint. The effective fix na smaller model or shorter context, no be more vCPUs.