What does an AI server cost to build?
The accelerator sets most of the bill. Learn how VRAM, memory bandwidth, power headroom, cooling and your own tariff decide what the rest costs.
Filtering by topic #self-hosted-llm · clear
The accelerator sets most of the bill. Learn how VRAM, memory bandwidth, power headroom, cooling and your own tariff decide what the rest costs.
Zanus AI sells a no-code AI appliance priced by quote. What the same private stack takes to build on a rented GPU server, and when an API is the better buy.
Every Claude plan price in USD as of September 2026, why a card gets declined in an unsupported country, and the open model option you can still run.
Which Gemma 4 tag fits your VPS RAM, what each one costs in tokens per second on CPU, and the sizing arithmetic to check it before you pull 20GB.
Curl your Ollama server on port 11434, read what a refused connection means, tour the API endpoints, and set OLLAMA_HOST without exposing the box.
The second request waits, or it is refused: how OLLAMA_NUM_PARALLEL and OLLAMA_MAX_QUEUE decide which, and why every parallel slot costs you VRAM.
GLM 5.2 is cloud only in Ollama's library. Here is the GLM model that actually fits a VPS, and the RAM each quantisation needs on your own box.
Build llama-server from a pinned tag, serve GGUF models on the OpenAI-compatible API, bind it to localhost, and run it under systemd with memory limits.
DeepSeek V4 Flash on Ollama is cloud only. Get the real disk and RAM numbers for every GGUF build, and see which VPS plan can actually load it.
Kimi K3 is 2.8 trillion parameters. Here is the VRAM arithmetic, the KV cache math, and the three honest ways to run it without a 32 GPU cluster.