How to run Meta Muse Glimmer 30B on Linux VPS
Muse Glimmer tags range from 17GB to 59GB. Check the RAM and disk space wey your Linux VPS need before you pull the model and see wetin CPU only inference go cost you for server.
Wetin Muse Glimmer need for VPS
Muse Glimmer dey run for normal Linux VPS wey no get GPU, and the tag wey you pull go decide weda e fit enter your memory. Meta Superintelligence Labs release the model for 10 August 2026 under Apache 2.0: e get 30 billion parameters, 128K context window, and one special 1.8B parameter perception encoder so e fit read image join text. Meta design am for agents wey dey always on, no be just for chat, and you fit set the reasoning strength for every request.
The Ollama tags wey dem publish, wey dem read for 16 August 2026, dey range from 17 GB reach 59 GB. That range na the whole wahala of how you go size your server. The default tag dey around 18 GB, so the smallest server wey make sense must get pass 18 GB of free RAM. You still need extra disk space for the download and extra memory for the context window.
Which muse-glimmer tag you suppose pull?
The data behind this chart
[
{
"label": "30b-nvfp4",
"size_gb": 17
},
{
"label": "30b (default)",
"size_gb": 18
},
{
"label": "30b-q4_K_M",
"size_gb": 18
},
{
"label": "30b-q4_K_M-dflash",
"size_gb": 20
},
{
"label": "30b-nvfp4-dflash",
"size_gb": 21
},
{
"label": "30b-q8_0",
"size_gb": 31
},
{
"label": "30b-mxfp8",
"size_gb": 33
},
{
"label": "30b-q8_0-dflash",
"size_gb": 33
},
{
"label": "30b-mxfp8-dflash",
"size_gb": 35
},
{
"label": "30b-bf16",
"size_gb": 57
},
{
"label": "30b-bf16-dflash",
"size_gb": 59
}
]Ollama list 11 tags for this model wey no be Apple builds. Dem get the same 30 billion weights wey dem store for different numeric precisions. The size wey you see na wetin you go download, and e still be roughly wetin you must get for memory before you add any context.
The two 4-bit builds na the small ones: 30b-nvfp4 for 17 GB and 30b-q4_K_M for 18 GB. The default 30b tag get the same size as the q4_K_M build. The 8-bit builds, 30b-q8_0 and 30b-mxfp8, dey near 31 GB. 30b-bf16 na the unquantised 16-bit release for 57 GB, wey be more RAM pass wetin most rented servers dey offer for price wey person fit pay for side project.
The -dflash tags na the same builds with DFlash support, and each one dey larger pass ein plain twin. Ollama describe DFlash as speed feature and e show am for Apple Silicon and desktop GPUs. For CPU only VPS, you go dey pay that extra size for real memory sake of feature wey dem measure for other hardware, so start with the plain tag and change one thing at a time.
Start for 4-bit unless you get specific reason why you no go do am. To move from 4-bit go 8-bit roughly double the bytes wey the CPU must read for every token e generate, so throughput go drop while memory use go rise. That trade na the subject of wetin q4, q8 and fp16 quantisation actually cost you, and for CPU box, the short answer na say the 4-bit build be the only one wey make sense to start with.
Why the MLX tags no dey work for Linux server
MLX na Apple array framework, and Ollama MLX engine na di backend for Apple Silicon. Any tag wey get mlx for inside di name, na for dat engine and dat hardware dem build am. For x86 Linux VPS, dat one na tens of gigabytes of download wey you no fit run, and e go just full your disk without do anything. Di speed figures wey dem show for di announcement, na for Mac dem measure am, so e no concern your server. Wen you dey check di tag list for di model page, first remove every mlx name, den you come check di size from wetin remain.
How much RAM and disk e really need?
Two things dey chop memory, and na only one of dem be the tag size. The weights dey fixed based on the tag wey you pull. The KV cache, wey be the state wey the model dey keep for every token inside conversation, dey grow as the context length wey you configure dey grow. Ollama own documentation talk say if you dey serve parallel requests, e go multiply the context by the number of requests wey dey run at the same time, so one box wey dey answer two agents at once need more memory pass the same box wey dey answer one.
No just follow any RAM figure wey you see for guide, even this one. Pull the tag, send am one prompt, and while the model still dey resident for memory, run these two commands.
ollama ps
free -hollama ps go show wetin dey loaded right now and how the work dey split between CPU and GPU. free -h go show wetin remain. Those two outputs for your own box better pass any table wey dem publish, because dem don already include your context setting, your quantisation, and everything else wey the server dey run.
Disk na the easier part. Ollama dey store models inside /usr/share/ollama/.ollama/models for Linux, wey dey sit on the root filesystem for most VPS images. 40GB root volume no go fit hold the bf16 build at 57 GB, and e no go fit hold two 8-bit tags side by side too. If you never check wetin pull command dey write before, where Ollama stores models and how to move them go show you how that directory dey work. Move the store go one mounted volume before you pull anything.
sudo systemctl edit ollama[Service]
Environment="OLLAMA_MODELS=/mnt/models"sudo mkdir -p /mnt/models
sudo chown -R ollama:ollama /mnt/models
sudo systemctl daemon-reload
sudo systemctl restart ollamaThe ollama user must get that directory, because the service dey run as ollama and e dey write its blobs there as itself. If pull fail because of permissions, journalctl -u ollama -n 50 na where the reason go show.
Swap need one plain statement: swap no go let you run bigger tag. Generation dey touch the weights for every token wey e produce, so weights wey dey inside swap go dey read back from disk over and over, vmstat 1 go show say the si and so columns dey busy, and output go slow reach seconds per token. Keep small swap file as insurance against the out of memory killer. Size your RAM for the tag wey you actually want.
Install Ollama and pin one named tag
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
systemctl status ollamaDi install script go set up one systemd service, so di server go start back afta reboot. If you no want make e run as root-managed system service, running Ollama rootless under Podman don explain how you go do am. Afta dat, pull one specific tag.
ollama pull muse-glimmer:30b
ollama listCheck di size column for ollama list by yourself and compare am wit di current tag list wey dey di model page. Dem dey add, rename, and remove published tags, so di size wey you see for guide na just snapshot of one day.
Make you no ever write ollama pull muse-glimmer for server wey you dey depend on. One bare model name dey resolve go di latest tag, and latest na pointer wey di publisher fit move go different build. One routine pull fit come swap di model wey dey your agent backend, wey go come get different memory needs and different behaviour, and nothing for your logs go announce am. Write di tag inside your scripts, your unit files, and your agent config. Self-hosting an LLM with Ollama on a VPS don cover di rest of di server setup.
Fit I run Muse Glimmer without GPU?
Yes, and make we talk true about the limit. To generate one token mean say you must read model weights from memory, so the speed depend on memory bandwidth, no be how many vCPUs your plan get. Once you pass small number of cores, extra cores no go help you much. For shared VPS, that bandwidth dey shared with every other person wey dey the host, so 30B model wey be 4-bit go produce small number of tokens per second.
No believe anybody number for that, even my own. Measure tokens per second for your own machine and decide based on wetin you see.
The result show clear difference for wetin the model fit do. Interactive chat go dey painful, because you go read faster pass how the server dey write, and every reply go start with long pause. Background agent work dey fine, because task wey dey run for ten minutes no care whether e slow. That second workload na exactly wetin Meta talk say this model fit do.
If you need interactive speed, the two honest answer na GPU or hosted API. Calculate the break even point between GPU VPS and API tokens before you rent anything, and wetin GPU VPS actually give you explain wetin you dey buy. For the wider question of wetin one box fit carry, start from which models you fit self-host, and running similar size Qwen model on VPS na the closest comparison for this size class. If the numbers wey you measure come out too slow for you, Nemotron 3.5 Lightning on VPS ask the same RAM and tokens per second questions for model wey dem build for speed instead of size.
Why e dey forget tins long before 128K tokens?
Because Ollama default context window na 4096 tokens, no matter wetin the model support. Dat default dey inside Ollama own FAQ as of August 2026. The tag dey advertise 128K, but the server go give the model 4096 until you talk otherwise, so long agent transcript go lose im early turns and the model go look like say e get amnesia.
Increase am for the server for every request:
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=32768"Inside one interactive session, /set parameter num_ctx 32768 go change am for dat session only. Over the API, send num_ctx inside the request options.
Every extra token of context dey cost memory on top of the weights. If you ask for the full 128K on one box wey size just for the weights, the load go fail or e go fall back to sometin wey slow. Increase am small-small and run ollama ps after every step. How num_ctx and context length dey work for Ollama explain the calculation.
Reasoning strength: low, medium, high and xhigh
Meta show four reasoning strengths for Muse Glimmer, from low go reach xhigh, and dem recommend the two wey high pass for complex coding and agent tasks. For Ollama, this one dey work through the think parameter. Use --think= for command line, or send think inside the API body.
ollama run muse-glimmer:30b --think=high "Summarise the changes in /tmp/patch.diff"Inside one interactive session, /set think and /set nothink dey toggle am. Ollama documentation talk say most models fit take boolean or level like low, medium or high, and some fit take max for the highest one wey dey available. The correct strings wey this model dey accept dey for im model page, so make you read am instead of to dey guess, and try one by hand before you connect am go inside agent.
For box wey only get CPU, this setting get weight. Higher strength mean say more thinking tokens go generate before the first word of the answer show, and one thinking token dey cost the same time as one answer token. Make you leave routine work for low setting. The length of the answer sef need the same care, so cap the reply with num_predict instead of make you allow one long talk response hold your slow box for several minutes.
Make the model stay loaded for agent wey dey always on
Ollama dey unload model wey no dey do anything after five minutes by default. For agent wey dey run every ten minutes, this mean say you go dey load 18 GB from disk every time e run, and for VPS wey use network attached storage, that loading no dey fast. Make you pin am for memory instead.
[Service]
Environment="OLLAMA_KEEP_ALIVE=-1"If you use negative value, the model go stay for RAM until something else unload am, and keep_alive inside API request go override wetin the server set for that one call. The cost dey clear: the RAM go stay occupied even when nothing dey happen, so this setting na for server wey you dedicate only for the agent. Make Ollama model stay loaded explain the different ways wey you fit do am.
Point coding agent go di place
Ollama dey serve OpenAI compatible API for http://127.0.0.1:11434/v1, so most agent tools go connect with base URL and any API key wey no empty. Ollama Muse Glimmer page self show shortcut wey you fit use take link supported agent to local model with one command, and you suppose pin di tag for there too.
ollama launch claude --model muse-glimmer:30bAgents dey send big prompts. File contents, tool output, and transcript wey dey grow, all of dem dey come as input tokens. For CPU box, prompt processing na di part wey dey heavy before generation even start. Make you keep context setting small as di task allow. Point coding agent go Ollama explain di client side, run coding agent for VPS explain di box wey e dey live, and control agent cost for VPS explain wetin dey happen when e run all day.
Image input dey work di same way. Ollama API dey take images for images field of message, so client wey be text-only no go ever send one, no matter how di perception encoder strong reach.
No open port 11434
Ollama API no get authentication. If you set OLLAMA_HOST=0.0.0.0:11434 make you fit reach am from your laptop, you don put model runner wey no get security for public internet. Anybody wey find am fit load models enter your disk and read anything wey your agent send pass inside. Make you leave am bound to localhost and use tunnel instead.
ssh -N -L 11434:127.0.0.1:11434 user@your-vpsHow to secure Ollama API endpoint show you the correct options, including how to use reverse proxy wey go ask for password before e allow access.
Wetin fit break, and wetin you go see
The pull stop for middle. Disk. Run df -h against the model directory. One 57 GB bf16 build no fit enter 40GB root volume, and two 8-bit tags wey dey side-by-side no fit enter too.
The model load finish then the process die. Out of memory. dmesg -T dey record as the kernel out of memory killer choose one process, and journalctl -u ollama -n 100 show the service side of the same event. The fix na smaller tag or smaller num_ctx. No be more swap.
E dey run for seconds per token. Run vmstat 1 and watch the si and so columns. If swap activity dey steady, e mean say the weights no fit enter RAM and the box dey read dem back from disk as e dey work.
One tag wey work last week don disappear. Tag lists dey change. Read the model page again, pin wetin dey current, and write the tag name for place wey you go see am again.
Check the sizes yourself before you pull
The sizes wey dey the chart na wetin dem read from the model tag page for 16 August 2026, and the tag list wey dem publish no be promise. Read the current list for the model page, then confirm wetin actually land for your disk:
ollama pull muse-glimmer:30b
ollama list
sudo du -sh /usr/share/ollama/.ollama/modelsOllama dey store model layers as shared blobs, so two tags wey share one layer no go cost double disk space. Compare wetin du report against the published size and plan your disk space based on the one wey big pass.
FAQ
How much RAM Muse Glimmer need for VPS?
Start from the tag size come add the context window. Dem list the default tag as 18 GB for 16 August 2026, so 16GB box no fit hold am at all, and 24GB box go hold am but e no go get space for context. Make you see this one as starting point, no be final answer. Pull the tag, load am one time, then run ollama ps and free -h for your own box make you read your own numbers. Longer context and parallel requests dey add memory join the weights.
I fit run Muse Glimmer without GPU?
Yes. E dey load and answer for CPU-only VPS. Generation speed dey depend on memory bandwidth pass core count, and for shared host, that bandwidth na shared, so expect small number of tokens per second for 4-bit. This one dey okay for background agent work wey no need person, but e go slow well-well for interactive chat. Run ollama ps during request and check the processor column to confirm where the work dey run.
The MLX tags get use for Linux VPS?
No. Every tag wey get mlx for name na for Ollama MLX engine, wey be the backend for Apple Silicon. For x86 Linux server, those tags na big download wey you no fit run. Use the plain 30b tag, or any other non-MLX tag, and ignore the Apple hardware benchmarks wey dem put for the MLX builds.
Why the model dey forget things before e reach 128K tokens?
Because Ollama default context window na 4096 tokens, no matter wetin the model support, so the server dey cut long conversation before the model even see dem. Set OLLAMA_CONTEXT_LENGTH for the server, or /set parameter num_ctx for one session, or send num_ctx for the API request options. Memory use dey increase as you increase the window, so raise am small-small and check ollama ps every time.
I suppose pin the tag or I go just use latest?
Pin am. muse-glimmer wey no get tag dey resolve to latest, wey be pointer wey the publisher fit move to different build anytime, so routine pull fit change the model wey your agent dey use. Write muse-glimmer:30b for scripts, unit files, and agent config. Check the tag list for the model page before you pin, because published tags dey change.