Prefill vs decode: why token one lags
Prefill is compute bound and owns time to first token. Decode is memory bandwidth bound and owns tokens per second. Measure them separately on your own box.
Filtering by topic #batching · clear
Prefill is compute bound and owns time to first token. Decode is memory bandwidth bound and owns tokens per second. Measure them separately on your own box.
One user was fine, five crawl. How batching, KV cache limits, prefill and queue depth decide how many people your LLM server can serve at once.