SSD Nodes Learn Hosting plans →

#self-hosted-llm

Filtering by topic #self-hosted-llm · clear

Guides

Run GLM 5.2 on a VPS with Ollama

GLM 5.2 is cloud only in Ollama's library. Here is the GLM model that actually fits a VPS, and the RAM each quantisation needs on your own box.

Guides

Run the llama.cpp server on a VPS

Build llama-server from a pinned tag, serve GGUF models on the OpenAI-compatible API, bind it to localhost, and run it under systemd with memory limits.

Guides

Run DeepSeek V4 Flash on a VPS

DeepSeek V4 Flash on Ollama is cloud only. Get the real disk and RAM numbers for every GGUF build, and see which VPS plan can actually load it.

Guides

What it takes to self-host Kimi K3

Kimi K3 is 2.8 trillion parameters. Here is the VRAM arithmetic, the KV cache math, and the three honest ways to run it without a 32 GPU cluster.