Ollama concurrency: NUM_PARALLEL and MAX_QUEUE
The second request waits, or it is refused: how OLLAMA_NUM_PARALLEL and OLLAMA_MAX_QUEUE decide which, and why every parallel slot costs you VRAM.
Filtering by topic #vram · clear
The second request waits, or it is refused: how OLLAMA_NUM_PARALLEL and OLLAMA_MAX_QUEUE decide which, and why every parallel slot costs you VRAM.
Pick an Ollama quantization with arithmetic instead of guesswork: what q4_K_M, q8_0 and fp16 cost in RAM, and where the quality actually drops.
Kimi K3 is 2.8 trillion parameters. Here is the VRAM arithmetic, the KV cache math, and the three honest ways to run it without a 32 GPU cluster.
What open video models really produce in July 2026, the VRAM class each one needs, what a five second clip costs on a rented GPU, and when to use an API.