How to run Nemotron 3.5 Lightning for Ollama on VPS
Wanna run Nemotron 3.5 Lightning on your server? See how to pull the tag, check if 32GB RAM dey enough, and why CPU-only mode fit slow down your agent tasks too much.
Wetin Nemotron 3.5 Lightning dey do
Nemotron 3.5 Lightning na NVIDIA open 30B mixture-of-experts model, wey dem release for August 2026, and dem build am for agents wey dey run for hours instead of just one chat window. MoE (mixture of experts) mean say dem divide the weights into plenty expert sub-networks, and dem dey route each token through only small portion of dem. NVIDIA model card talk say e get 30 billion total parameters with 3 billion wey dey active per token. You dey pay for the big number for memory. You dey get the small number back as speed.
Dat trade-off na the reason why you suppose look this model for server wey you rent. Agent wey dey do real work dey send thousands of short requests throughout the day, so throughput per dollar na wetin go decide whether e fit live for your own box. Model wey dey take 40 seconds per reply na okay assistant but e no good as agent, because one task fit make twenty calls and you go wait for every one of dem.
NVIDIA describe the architecture as hybrid: e get interleaved Mamba-2 and MoE layers with some attention layers. The model card talk say maximum context length fit reach 1M tokens and e get OpenMDW-1.1 license, wey mean say e ready for commercial use. The main languages na English and code, but dem also list Spanish, French, German, Italian and Japanese.
Artificial Analysis publish launch measurements for August 2026 wey show say e fit reach nearly 670 output tokens per second, for one pre-release DeepInfra endpoint wey dey serve the NVFP4 weights. Dat one na hosted GPU endpoint. Make you read am as wetin the architecture fit do, no be wetin your VPS go fit do.
Which Ollama tag fit which VPS
Ollama library dey publish plenti builds of the same weights. Wetin dey change between dem na quantisation, wey be how many bits dem take store each weight, and dat one dey change the download size well-well.
The data behind this chart
[
{
"label": "30b-a3b-q4_K_M",
"size_gb": 25
},
{
"label": "30b-a3b-q8_0",
"size_gb": 35
},
{
"label": "30b-a3b-bf16",
"size_gb": 66
},
{
"label": "30b-a3b-mlx",
"size_gb": 23
}
]The tags wey dem name latest, 30b and 30b-a3b all point to the same digest as 30b-a3b-q4_K_M, so the default download na the 25 GB four-bit build wey get full 1M context. Q8_0 na 35 GB and bf16 na 66 GB, both still dey for 1M. The MLX builds wey dey 23 GB na for Apple silicon and e dey stop for 256K context, so dem no be the correct choice for Linux VPS.
Dem be download sizes, no be memory requirement. NVIDIA no dey publish minimum VRAM (video RAM) figure for Ollama builds, so make you treat the download size as the bare minimum and nothing else. The weights must stay somewhere, inside GPU memory if the card fit carry am, or inside system RAM if e no fit; plus the KV cache (key/value cache, the model memory for each token for the conversation) go join on top. The correct number for your hardware dey come from one command, no be from calculation, and e dey below. If you never decide on the quantisation level, wetin Q4, Q8 and FP16 each one dey cost you explain wetin you go lose for each step.
Pull the exact tag, no be latest
latest na pointer wey dey move. If the library update am, your agent go change behaviour for the next pull and you no go get any note to explain why. Use the correct tag name.
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
ollama pull nemotron-3.5-lightning:30b-a3b-q4_K_MThe install script go setup one systemd service wey dey run as the ollama user and e dey keep models for inside /usr/share/ollama/.ollama/models. For most VPS images, that path dey the root filesystem, so check space before you ask for 25 GB. If that filesystem space dey tight, where Ollama stores its models and how to move them go help you before you pull, instead of after the disk don full.
df -h /usr/share/ollamaIf pull stop for middle and e show no space left on device, e mean exactly wetin e talk, and the partial blobs go stay for disk until you delete dem. After that, confirm wetin actually land:
ollama show nemotron-3.5-lightning:30b-a3b-q4_K_Mollama show go print the architecture, the parameter count, the context length, and the quantisation wey the file carry. If any of those no match wetin dey the library page, e mean say you pull different tag wey you no intend.
Serve am, come check where e actually run
sudo systemctl enable --now ollama
ollama run nemotron-3.5-lightning:30b-a3b-q4_K_M "Reply with one word: ready"While the model still dey load, for one second shell:
ollama psDis na the command wey go answer the memory question for your machine. ollama ps go print the model wey load, the size wey e dey take for memory, plus one PROCESSOR column. 100% GPU mean say everything dey inside VRAM. 100% CPU mean say nothing dey inside, and every token na the processor dey compute am from system RAM. One split like 65%/35% CPU/GPU mean say the layers no fit enter everything, and the CPU share na im go determine your speed. No try guess the requirement. Load am make you read dis line.
If e no fit load at all, Ollama go refuse cleanly instead of make e crash:
Error: model requires more system memory (28.4 GiB) than is available (15.6 GiB)CPU-only VPS go fast reach?
General purpose VPS no get GPU, so na the CPU dey do all the work and e dey read every weight wey e need from system RAM. MoE dey help for this side, because na only about 3 billion out of the 30 billion parameters dem dey touch per token, so the calculation per token dey far smaller pass wetin dense 30B model go need. Memory no dey get any help at all. All 30 billion parameters must stay resident, because the router fit pick any expert for any token.
So, CPU-only inference for this model dey limited by memory bandwidth, no be by the number of cores. If you add vCPUs to a plan wey already get enough cores, e no go change anything. Wetin you need na enough RAM to hold the weights plus your KV cache, and the fastest memory wey the plan fit give you.
Measure am before you commit any agent to am, use the method wey dey measuring tokens per second for a local LLM:
ollama run --verbose nemotron-3.5-lightning:30b-a3b-q4_K_M "Write a 200 word summary of TCP slow start."The eval rate line wey dem print for the end na your generation speed for tokens per second. That single number na wetin go decide the matter, because na that one dey determine how long the agent go take finish work. Multiply am by the length of reply wey you expect, and if the answer take time pass wetin you fit wait, capping the output with num_predict na the one way to limit a single call without you changing your hardware.
The data behind this chart
[
{
"label": "Nemotron 3.5 Lightning",
"sec_per_task": 30
},
{
"label": "gpt-oss-120b",
"sec_per_task": 204
},
{
"label": "Qwen3.6 35B",
"sec_per_task": 210
}
]Those figures na third-party data wey dem publish, wey dem convert from the minutes per task wey Artificial Analysis report when dem launch, and dem measure am for hosted GPU endpoints, no be for VPS. Nemotron 3.5 Lightning average about 30 seconds per task, where gpt-oss-120b take roughly 204 and Qwen3.6 35B take roughly 210. Use dem to see the difference, no be as promise for your own hardware.
The correct advice depend on who dey wait. If person dey wait for the agent, or if the agent dey do long chains of calls one after another, rent GPU capacity. If the thing dey run for schedule for night and nobody dey watch, a large-RAM CPU plan na better place to put am. Anyhow, the setup na the same, and running Ollama on a VPS explain how to size your plan and how GPU instance take compare with paying API provider per token. The break-even point depend on how you dey use am: GPU instance dey charge every hour wey e dey run, while API tokens dey charge only when you use am, so if your agent dey busy most of the day, e better make you own the box, but if the agent only run two times per hour, e no usually worth am.
The 1M context window no dey free
1M tokens na the maximum wey the model fit take, and Ollama no dey give am to you by default. Ollama dey use small default window and e dey comot the oldest tokens once the conversation pass the limit. Nothing dey show for log when that one happen, so for agent, e go be like say the model forget the beginning of im own task.
Set the window by yourself. For the whole server, edit the service:
sudo systemctl edit ollamaAdd this one, then run sudo systemctl restart ollama:
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=32768"Per request, send num_ctx inside the options object instead:
curl http://localhost:11434/api/chat -d '{
"model": "nemotron-3.5-lightning:30b-a3b-q4_K_M",
"messages": [{"role": "user", "content": "Say ready"}],
"options": {"num_ctx": 32768},
"stream": false
}'Every increase dey cost memory, because the KV cache dey grow as you increase the number of tokens. Increase the value, restart, then run ollama ps again and watch the reported size as e dey climb. If the PROCESSOR column change from 100% GPU to a split after that change, the KV cache don push model layers comot from VRAM and your speed go drop well-well. Choosing num_ctx in Ollama explain that trade-off for detail. No set 1000000 just because the model card allow am, because the allocation dey happen for front and the load go just fail.
Make am join always-on agent
Ollama launch post for this model show one shortcut wey fit start agent wey don already point go the model:
ollama launch claude --model nemotron-3.5-lightningThe post show claude, opencode, openclaw and hermes for that position. The subcommand need current version of Ollama, so check ollama --version first, and if e no dey, point the agent go the API by yourself. Ollama get OpenAI-compatible endpoint, wey most agent harnesses dey accept:
export OPENAI_BASE_URL=http://localhost:11434/v1
export OPENAI_API_KEY=ollamaOllama no dey care about the key, but most clients no go gree start if you no set one. We don cover the harness side for how to point coding agent go Ollama and for how to build your own OpenClaw agent.
Two server settings dey important once the agent dey run by itself. OLLAMA_KEEP_ALIVE dey control how long model go stay for memory after the last request, and the default dey unload am after five minutes, so the next call go take time to load again. For 25 GB file wey no get GPU, that wait time fit make the connection timeout. Set OLLAMA_KEEP_ALIVE=-1 make the model stay for memory. OLLAMA_HOST=0.0.0.0:11434 dey make the API reachable from other machines, and e no get any kind authentication, so open am only if you get firewall rule or private network.
Failure modes, with the strings wey you go see
The pull fail immediately. Error: pull model manifest: file does not exist mean say that tag no dey exist. Tag names na exact strings, so copy one from the library page instead of make you dey guess the quantisation suffix.
The model no go load. Error: model requires more system memory (28.4 GiB) than is available (15.6 GiB) mean say the tag too big for this plan as you configure am. Drop go smaller quantisation, or lower OLLAMA_CONTEXT_LENGTH, because the KV cache dey count inside that requirement.
Nothing answer for port 11434. curl: (7) Failed to connect to localhost port 11434 mean say the service no dey run, or e no dey listen where you expect. Read systemctl status ollama and journalctl -u ollama -n 50. If you self start ollama serve by hand, the second copy go exit with Error: listen tcp 127.0.0.1:11434: bind: address already in use.
E answer, but e dey very slow. Check ollama ps before you change anything. Any CPU share for the PROCESSOR column for a GPU machine mean say part of the model don spill comot from VRAM, so lower the context or take the smaller quantisation. For machine wey no get GPU, slow na the expected outcome and no setting fit repair am.
The agent dey forget im instructions as e dey do task. The conversation don pass the context window and dem don silently throway the oldest tokens. Raise OLLAMA_CONTEXT_LENGTH, confirm with ollama ps say the model still fit, and if e no fit again, the fix na bigger machine instead of smaller window.
Where this model dey stand compare to other options
30B MoE na heavy load for small work. If dense 8B model fit do your task, e go cost you far less to run and e go load sharp-sharp. Qwen 3 for 8B and 27B for VPS na the direct comparison wey you fit use decide. For better survey of wetin your plan fit carry, start from which AI models you fit self-host. If you plan to serve many agents at once instead of one, read Ollama compare with vLLM first, because Ollama no dey batch concurrent requests like how production inference server dey do, and na there single-user setup dey reach im limit.
FAQ
Which Nemotron 3.5 Lightning tag I suppose pull for Linux VPS?
Use nemotron-3.5-lightning:30b-a3b-q4_K_M. E size na 25 GB, e get the full 1M maximum context, and na the same digest wey latest, 30b and 30b-a3b tags point to as of August 2026. Call the name directly instead of pulling latest, so say if dem republish that pointer later, e no go change how your agent dey behave without you notice. The mlx tags na for Apple silicon, e no go work for Linux.
How much RAM Nemotron 3.5 Lightning need?
NVIDIA no publish minimum memory figure for Ollama builds, so make you measure am instead of to dey guess. Pull the tag, run the model one time, then check ollama ps while the model dey load: e go show the size wey e dey occupy and whether e land for GPU or CPU. The download size, 25 GB for the default tag, na the base, because the KV cache dey add join and e dey grow as you set the context window. If your plan too small, Ollama go refuse with model requires more system memory and e go show you the two numbers.
I fit run Nemotron 3.5 Lightning for VPS wey no get GPU?
Yes, if your plan get enough RAM to hold the weights. The MoE design dey help because na only about 3 out of the 30 billion parameters dem dey compute per token. The wahala na speed. Without GPU, the model speed dey depend on memory bandwidth, so to add vCPUs no go really make am fast. Run ollama run --verbose with one fixed prompt, read the eval rate line, and check whether that number fit your agent deadline. If na batch job wey go run for night, e dey okay. If na something wey person dey wait for, e no go work well.
Why Ollama no dey give me the full 1M context window?
1M na the maximum wey the model fit take, no be the default for Ollama. Ollama dey use smaller window and e dey comot the oldest tokens once conversation pass the limit, e no go show any error, so the agent go just dey forget wetin you tell am before. Set OLLAMA_CONTEXT_LENGTH for your systemd service, or pass num_ctx per request. Increase am small-small and re-check ollama ps every time, because the KV cache memory dey increase as the window increase, and e fit push model layers comot from GPU.
Nemotron 3.5 Lightning free for commercial use?
NVIDIA model card put the model under OpenMDW-1.1 license and e mark am say e ready for commercial use. That one cover the weights wey you download and run by yourself. E no talk anything about the other software for your stack, so check the license of your agent harness and any tools wey you connect to am separately, and read the current model card before you use am for any contract work.