Ollama pull vs run: wetin be the difference?
Ollama pull go download model stop, but ollama run go open chat. Learn where these files dey save for your VPS root disk and how to move them so your storage no go full sharp sharp.
ollama pull vs ollama run
ollama pull go download model come stop. ollama run go download the model only if e no dey, then e go load am enter memory come open interactive chat. The download process na the same and the files go land for the same place. Only run go continue after that.
Na that one difference decide which command you go use for script and which one you go use for keyboard.
ollama pull gemma4
ollama run gemma4
ollama run gemma4 "Reply with one word: ready"The first line go fetch the model come exit, so e safe to use for provisioning and inside systemd unit. The second one go open chat session; type /bye or press Ctrl+D to comot. The third one go send one prompt, print the answer, then exit; this na the format wey script need if you want answer instead of session. That third format still leave the length of the reply for the model to decide, so one-line question fit return three paragraphs, and capping the answer with num_predict na wetin go keep scripted run inside size wey the caller fit use. Model names dey change fast, so treat gemma4 for here as placeholder: na the example wey official Ollama documentation use as of August 2026, and any tag from the library go behave the same way. If you prefer to use model wey person don already size against real server, running Nemotron 3.5 Lightning on a VPS give the exact tag to pull and the memory wey e expect to see.
Why the first ollama run looks like it has frozen
If you run run for the first time for fresh VPS, e fit stay like say e freeze for some minutes. Nothing spoil. The chat prompt no go show until the model don finish download enter disk and load enter memory, so run dey do multi-gigabyte download before e fit show you anything.
Two things dey hide that work. Ollama dey draw progress bar only when the output na terminal, so if you run run inside shell script, cron job, CI step or plain ssh host ollama run ..., e no go print anything while e dey download. Then, after the bytes don land, the file still get to read from disk enter RAM before the first token show, and for small VPS, that read dey slow. If the box no get enough memory for the model, the kernel go start swapping and the wait go come long well well.
Watch am from second session instead of just dey guess:
df -h /
watch -n5 df -h /If free space dey reduce small-small, e mean say the download still dey run. If free space stop to reduce while the command still dey busy, e mean say the download don finish and the load enter memory don start.
This na the reason why you suppose pull am before time. The person wey type ollama run no suppose be the one wey dey wait for the download.
Pull the model before anybody ask for am
The same thing apply to anything wey no be person: coding agent wey point go your Ollama endpoint go usually give up for the first request instead of make e wait for multi-gigabyte download. For new box, pull the same script wey dey install the server:
curl -fsSL https://ollama.com/install.sh | sh
ollama pull gemma4If you dey set up the server for the first time, the full Ollama on a VPS install cover the service itself and who get permission to reach am. After that, the thing wey make sense to set up na pull wey go last pass your terminal, because if download cut for middle, na so people take dey get model store wey no complete.
Run am inside tmux, or give am to systemd as one-shot unit wey dey run for boot. Write /etc/systemd/system/ollama-pull.service:
[Unit]
Description=Pre-pull Ollama models
Wants=ollama.service network-online.target
After=ollama.service network-online.target
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/bin/sh -c 'until ollama list >/dev/null 2>&1; do sleep 2; done'
ExecStart=/bin/sh -c 'ollama pull gemma4'
[Install]
WantedBy=multi-user.targetBoth commands dey run through /bin/sh -c because of purpose. One bare ExecStart= need absolute path, and the installer no dey always put the binary for the same directory, so command -v ollama for your own box na the only reliable answer. To go through the shell dey use the service PATH instead of path wey you copy from guide. The first ExecStart matter too: After=ollama.service mean say the server unit don start, wey no be the same thing as say e ready, so the loop dey wait until ollama list answer before the pull begin.
sudo systemctl daemon-reload
sudo systemctl enable --now ollama-pull.service
journalctl -u ollama-pull.serviceThe journal suppose show say the pull finish without error, and ollama list suppose show the model after. To keep moving tag current, add systemd timer or weekly cron entry wey go run the same pull. If you re-pull tag wey don move, e go download the new layers and leave the old ones without anything point to dem, and the server go clean dem comot next time wey e start.
Wetin dey happen if pull stop halfway
Every layer of one model dey store under one hash wey represent wetin dey inside. If pull stop halfway, e no mean say you waste work: run the same ollama pull again, and the layers wey don finish go dey recognized and skipped, so the download go continue from the layer wey cut off.
One action fit destroy that progress. When the Ollama server start, e dey remove stored layers wey no model manifest point to, and the partial layer wey dead pull leave behind na exactly that. So if you restart the service before you retry, you don throw away the part wey you don download finish. Retry the pull first before you restart. If you really need make partial download survive restart, set OLLAMA_NOPRUNE=1 for the service environment, then remove am later, because that startup cleanup na wetin dey stop orphaned layers from full your disk.
If the pull die with no space left on device, free space before you retry. If df report say disk full and du for the model directory no show wetin take the space, the space go another place, and the reasons df and du disagree dey important make you read before you delete anything.
Where Ollama dey keep models for VPS?
Make you check your own server instead of follow wetin any guide talk, even this one. The location dey different if you install am as package or container, and e fit change again if person don set OLLAMA_MODELS.
systemctl cat ollama.service
getent passwd ollama
sudo find / -xdev -type d -name blobs 2>/dev/nullsystemctl cat go print the unit file join with every drop-in, so any OLLAMA_MODELS line wey you set or wey dey inside your image go show there. If that line no dey, the store dey under the home directory of the account wey the service dey run as, and getent passwd go print that home directory for the sixth field wey colon separate. The find go search the filesystem for the blobs directory, na there dem dey write the layers. Use -xdev if the models fit already dey on top separate mount.
Now, make you measure and read your own numbers:
ollama list
df -h /
sudo du -sh /the/directory/you/found
sudo du -h -d1 /the/directory/you/foundThe store get two parts. manifests dey keep one small file for every model tag, and that file list the layers wey the tag take build. blobs dey keep the layers demsef, each one get name base on the hash of wetin dey inside, and almost all the size dey there. Because layers dey shared between tags, two models wey dem build on top the same weights go report their own size for ollama list even though dem occupy that space only once for disk, so the sizes wey dem list fit add up pass wetin du report for the directory.
Model files dey fill small VPS root filesystem quick pass any other thing wey you fit install, and the biggest thing wey you fit do to control the size na the weight format. Choosing between q4, q8 and fp16 fit save you gigabytes for every model.
Move the models go data volume wit OLLAMA_MODELS
If your plan get second disk or bigger data volume, move the store before the root filesystem full. Stop the server first, so you no go copy file wey still dey write.
sudo systemctl stop ollama
sudo mkdir -p /mnt/data/ollama-models
sudo rsync -a /the/directory/you/found/ /mnt/data/ollama-models/
sudo chown -R ollama:ollama /mnt/data/ollama-models
sudo systemctl edit ollama.servicesystemctl edit go open editor for drop-in file, so the packaged unit no go change and package upgrade no go overwrite your work. Add these two lines:
[Service]
Environment="OLLAMA_MODELS=/mnt/data/ollama-models"sudo systemctl daemon-reload
sudo systemctl restart ollama
systemctl show ollama --property=Environment
ollama listsystemctl show suppose print your new path, and ollama list suppose show the same models wey e show before you move am. If the list empty, e mean say the server no fit read the new directory. The service dey run as ollama user, so that user need read and write access to the destination, na why we put chown line for up. Check journalctl -e -u ollama for permission error wey mention the new path. Delete the old copy only after the list correct, because if you move am finish and you delete the source before you check, you go need download everything again.
The other option na to keep the original path and mount the data volume on top:
echo '/mnt/data/ollama-models /the/directory/you/found none bind 0 0' | sudo tee -a /etc/fstab
sudo mount -a
findmnt /the/directory/you/found
df -h /findmnt wey print the mount mean say the bind dey active. Bind mount dey help if another thing for the box already expect the default location. E get one trap: the files wey you copy out still dey under the mount point for the root disk, the mount cover dem, so the space no go free until you unmount and remove dem. The environment variable na the one wey easy to explain to anybody wey login next.
Where the container dey keep dem instead
The official image dey store models inside whatever you mount, e no dey store am for any host directory wey belong to one ollama user. The run command wey dem document na:
docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollamaollama wey dey before the colon na named Docker volume, and /root/.ollama na where the server dey write inside the container. So, du against the paths from the previous section no go see anything, because nothing dey there. Print the real location and size:
docker volume inspect ollama
docker system df -v
docker exec -it ollama ollama listRead the Mountpoint field from docker volume inspect, then run sudo du -sh against am. To put the models for one data volume, replace the named volume with one host directory (-v /mnt/data/ollama:/root/.ollama) and recreate the container. The container dey write as root, so that host directory go come belong to root. Under rootless Podman, the ids dey mapped into your user subuid range instead, so host ownership go look different again: running Ollama under rootless Podman explain that mapping.
One warning about cleanup. docker volume prune dey remove every volume wey no container dey use. If you remove or recreate the ollama container without im volume, and you later run prune, e go delete every model wey you download, and you no go fit recover dem except you download dem again. Read how to prune Docker disk usage on a VPS before you run prune for one box wey dey host models.
Remove model wit ollama rm, no be wit rm
ollama list
ollama rm gemma4
ollama list
df -h /ollama rm dey delete the manifest for dat tag, den e go delete the layers wey no manifest dey point to again. The space go come back as soon as dem unlink dose files, so df go move sharp sharp. Because layers dey shared, if you remove one out of two tags wey resemble each oda, the space wey you go free fit small pass the size wey ollama list show for side. Dat one na correct behaviour, no be say the delete fail.
If you delete files wit your hand, you go break the pair. If you remove blob wit rm, the manifest go still list am, so ollama list go continue to show the model and any attempt to use am go fail wen the system try read the layer wey miss. If you remove manifest wit your hand, the layers go still dey for disk wit nothing point to dem, and dem go dey occupy space wey no Ollama command go fit report give you. If you don already do am, ollama rm on the tag go clear the leftover entry, and if you restart the server, e go clear the layers wey nothing dey refer to.
One last difference, because pipo dey always mix the two. ollama rm concern disk. ollama stop gemma4 dey unload model from memory and e no dey free any disk space at all. How long model go stay for RAM after the download finish na separate setting, and keeping a model loaded instead of reloading am for every request don explain am.
FAQ
Wetin be the difference between ollama pull and ollama run?
ollama pull go download model go your disk come close. ollama run go check whether the model dey for disk already, download am if e no dey, load am enter memory, then open interactive chat session. Both of dem dey write the same files for the same directory. Use pull for provisioning and inside scripts, and use run when person dey for keyboard. ollama run <model> "your prompt" go send one prompt come close, wey be the scriptable form of run.
Why my first ollama run dey look like say e hang?
E dey download. The chat prompt no fit show until the model don finish download enter disk and load enter memory, and model fit reach several gigabytes. Ollama dey show progress bar only when output na terminal, so run inside script, cron job or ssh host ollama run ... no go show anything while e dey work. Open second session and run watch -n5 df -h /: if free space dey reduce small-small, e mean say download dey go on. Pull the model before time and the wait go comot.
Where Ollama dey store im models?
The location depend on how you install am, so make you print am instead of to guess. Run systemctl cat ollama.service to see whether OLLAMA_MODELS dey set for the unit or drop-in. If e no dey, the store dey under the home directory of the account wey the service dey run as, wey getent passwd ollama go print. sudo find / -xdev -type d -name blobs 2>/dev/null go locate the layer directory directly. For container image, the store dey inside the mounted volume, and docker volume inspect ollama go print the host Mountpoint.
How I go move Ollama models go another disk?
Stop the service, copy the store go the new location with rsync -a, give the directory to the service account with sudo chown -R ollama:ollama <directory>, then run sudo systemctl edit ollama.service and add Environment="OLLAMA_MODELS=<directory>" under [Service] line. Reload with sudo systemctl daemon-reload and restart. Confirm with systemctl show ollama --property=Environment and ollama list. If the list empty, e almost always mean say the ollama user no fit read the new directory; journalctl -e -u ollama go show the path.
If I delete the model files, e go free the space?
If you delete files by hand, e go free the bytes but the store go come get error. If you remove blob, the manifest go still list that model, so e go still dey show for ollama list and e go fail when you try use am. If you remove manifest, the layers go still dey for disk but nothing go point to dem. Use ollama rm <model>, wey go delete the manifest and then the layers wey no other model need. If you don already delete files by hand, run ollama rm on the tag to clear the entry, then restart the server, wey go remove layers wey no manifest point to.