SSD Nodes Learn Hosting plans →
How to do am Matt ConnorBy Matt Connor

Download Hugging Face model enter your VPS

Pull model weights from Hugging Face enter your VPS with terminal alone: disk check, resolve link, download wey fit resume, gated repo token, and where the file land.

Wetin you go do here

To download Hugging Face model enter your VPS, wetin you really need na the correct file link plus one downloader wey fit continue after network cut. The page wey you dey look inside browser no be the file. That address carry blob inside am, and blob na the preview page, no be the weights.

Everything for this guide na terminal work on one Linux VPS wey no get desktop. Browser dey enter for only two reasons: to collect your access token, and to click the agree button for repo wey lock. After that, na SSH alone.

Check disk before you start, no be after

Model weights big. If your disk full when the download reach 90 percent, you don waste bandwidth wey you pay for, and the half file wey remain still dey chop space. So check first.

df -h ~
lsblk

df -h ~ dey answer the correct question, because the download dey land on the filesystem wey hold your home folder, no be on "the server" as one whole thing. If lsblk show one extra volume wey you attach, na there model suppose go, no be the small root disk.

Next thing, know how big the fetch be before you start am. The Files and versions tab on the model page dey show size for every file. If you don install the official CLI, one flag go count am for you without pulling anything:

hf download <user>/<repo> --dry-run
hf download <user>/<repo> --include "*.gguf" --dry-run

E go list the files plus the total bytes wey e go download. Leave headroom pass that number. If you download into cache and you later copy the file go another folder, you don hold two copy of the same weights on the same disk. Whether your RAM fit carry the model na different question, and how to know the model size wey your RAM fit hold cover that side.

Open the repo, click Files and versions, then click the file. Your address bar go carry something like this:

https://huggingface.co/openai-community/gpt2/blob/main/config.json

That one na HTML page wey dey show the file inside browser. Change blob to resolve and you get the file itself:

https://huggingface.co/openai-community/gpt2/resolve/main/config.json

The pattern na https://huggingface.co/<user>/<repo>/resolve/<branch>/<path>. main na the default branch. If the repo get tag or commit wey you want, e dey enter that same slot.

Test the idea with small file first, no be with 7GB:

curl -L -o gpt2-config.json \
  "https://huggingface.co/openai-community/gpt2/resolve/main/config.json"
head -c 200 gpt2-config.json

You suppose see JSON wey start with {. If wetin you see na angle bracket and HTML, the link no resolve to a file, and na that same check you go run on your big download.

Download the Hugging Face model with client wey fit resume

The resolve link no dey serve the bytes by itself. E dey redirect go a storage host. So your client must follow redirect, and e must sabi continue from where e stop.

curl -L -C - -o model.gguf \
  "https://huggingface.co/<user>/<repo>/resolve/main/<file>.gguf"

-L mean follow the redirect. Without am, curl go save the redirect response, and you go end up with tiny file wey no be weights at all. -C - mean continue: curl go look how many bytes don already land inside model.gguf, then ask the server for only the remaining part. Run the exact same command again after the line cut and e go continue from there. If the server refuse the range request, the progress bar go start from 0 percent, and you go see am one time.

wget dey do the same work:

wget -c -O model.gguf "https://huggingface.co/<user>/<repo>/resolve/main/<file>.gguf"

When e finish, check say the size make sense:

ls -lh model.gguf

File wey suppose be gigabytes but e land as few kilobytes mean your fetch catch an error page instead of weights. Run head -c 200 model.gguf and you go read the text wey dey inside.

Make light wahala no kill your download

If you run the download inside plain SSH session and your line cut, the shell die and everything wey e start die with am. Na SIGHUP. NEPA take light, your router off, and 5GB download wey don reach 80 percent just stop. Put the work inside tmux so the process dey live on the server side, no be on your side.

sudo apt update && sudo apt install -y tmux
tmux new -s dl

Start the download inside that window. Press Ctrl-b then d to comot without killing am. When light come back:

tmux attach -t dl

The progress bar go still dey run. tmux ls go show you which session still dey alive.

Use the official CLI when you want the whole repo

One file na curl work. Whole repo wey carry tokenizer plus plenty shard file, na the huggingface_hub CLI. On Ubuntu 24.04, if you try pip install straight into the system Python, pip go refuse and talk about externally-managed-environment. No fight am. Na venv:

sudo apt install -y python3-venv
python3 -m venv ~/.venvs/hf
~/.venvs/hf/bin/pip install -U "huggingface_hub"
echo 'export PATH="$HOME/.venvs/hf/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
hf download --help

If hf download --help print the usage line, the install set. The project also publish a standalone installer wey no need Python venv at all:

curl -LsSf https://hf.co/cli/install.sh | bash

Now the actual fetch. If you no name any file, e go pull the whole repo:

hf download <user>/<repo>
hf download <user>/<repo> <file>.gguf --local-dir /srv/models/mymodel
hf download <user>/<repo> --include "*.safetensors" --exclude "*.bin"

--include and --exclude dey take glob pattern, so you fit leave the format wey you no need. --local-dir go put the files inside the folder wey you name, and e go keep the repo folder structure. The most useful part: the command dey print the path where e save the file. That print na your real answer to "where the thing go", so read am instead of guessing.

Gated repo and the licence matter

Some repo lock. The author turn on access request, so you must login for browser, open the model page, read the terms, and click the agree button before anything go work. Some repo dey approve you automatic. Some need the author approve by hand, and that one fit take days. Na the author get the final say, and no terminal flag dey bypass am. Accept the terms, or find another repo wey carry the same weights.

Even after you get access, your VPS still need token, because curl and the CLI no dey share your browser login. Create a read token on https://huggingface.co/settings/tokens. If na fine grained token you pick, turn on the read access to contents of public gated repos option, otherwise the token go still bounce for gated repo.

hf auth login
hf auth whoami

hf auth login go ask for the token and save am under your Hugging Face home folder, so you no need paste am every time. hf auth whoami go print the account wey that machine dey use. If e no print your username, every download go behave like say you be stranger.

For curl, na header:

read -rs HF_TOKEN && export HF_TOKEN
curl -L -C - -H "Authorization: Bearer $HF_TOKEN" -o model.gguf \
  "https://huggingface.co/<user>/<repo>/resolve/main/<file>.gguf"

read -rs dey collect the token without showing am on screen and without dropping am inside your shell history. Before you burn bandwidth, ask the server whether you even get permission:

curl -sIL -o /dev/null -w '%{http_code}\n' \
  -H "Authorization: Bearer $HF_TOKEN" \
  "https://huggingface.co/<user>/<repo>/resolve/main/<file>.gguf"

If e print 200, you fit download. If e print 401, the server no see valid token, so check hf auth whoami or check say you export the variable inside the same shell. If e print 403, the token dey fine but that account never get access to the repo, so go back to the model page and see whether your request still dey pending.

Where the file land, and how to send am go bigger disk

By default the library keep everything inside one cache under your home folder, ~/.cache/huggingface/hub. Inside am, every repo get im own folder named like models--<user>--<repo>, and that folder carry blobs, snapshots, refs and trees. The real bytes dey inside blobs. Wetin dey inside snapshots na symlink wey dey point go the blob. Na why two revision wey share the same file no dey double your disk usage.

To see wetin dey inside:

hf cache ls
hf cache ls --revisions
du -sh ~/.cache/huggingface

That symlink layout get one trap. If you use rsync -a or plain tar to copy a snapshot folder go another box, wetin you carry na the link, no be the data, so the copy go be dead link for the other side. Use rsync -aL or tar -h so e go follow the link and carry the actual file. Or simpler: when you already know say the file go travel, download am with --local-dir from the start.

VPS root disk dey small and model dey chop space fast, so point the cache go the bigger volume:

mkdir -p /mnt/data
mv ~/.cache/huggingface /mnt/data/hf
echo 'export HF_HOME=/mnt/data/hf' >> ~/.bashrc
source ~/.bashrc
hf download openai-community/gpt2 config.json

The last command go print the full path where e save that small file. If the path start with /mnt/data/hf, your setting don work. HF_HOME dey move the cache together with the saved token. If na only the model cache you wan move, use HF_HUB_CACHE and leave the rest where e dey.

Two things dey catch person here. The library dey read the environment variable when the process dey start, so any shell wey you open before you edit .bashrc still dey use the old path. And systemd service or another user account no dey read your .bashrc at all, so if a daemon dey pull model, you must set the variable inside that service own environment.

Download wey break dey leave partial blob as .incomplete file inside the cache. hf cache ls go mention say e see incomplete download, but e no go remove them. hf cache prune na the command wey dey clear them, plus any revision wey no get branch or tag again.

Confirm say the file complete before you trust am

Half file dey behave like complete file until the loader try read am, and then the error message go point you go the model when na the download be the problem. Check am first.

hf cache verify <user>/<repo>

E dey compare the checksum of wetin dey your cache with wetin the Hub get, and e go tell you whether everything match. Note say e dey look the cache only, so file wey you pull with --local-dir or with curl no dey inside that check. For those one na manual checksum work, and how to verify your download with checksum show the full method.

Now hand the file over to something wey go run am

GGUF wey just dey sit inside folder no dey do anything. For Ollama, you point one Modelfile go the file:

echo "FROM /srv/models/mymodel/model.gguf" > Modelfile
ollama create mymodel -f Modelfile
ollama run mymodel

The full GGUF import steps for Ollama cover the template and parameter part wey dey decide whether the model go answer well or talk nonsense. If na just one community GGUF you want and you no send where the file land, Ollama fit fetch am straight from the Hub:

ollama run hf.co/<user>/<repo>
ollama run hf.co/<user>/<repo>:Q4_K_M

That one keep the weights inside Ollama own storage, no be inside the Hugging Face cache, so you go dey watch two different folder on the same disk. Where Ollama dey keep the model wey you pull explain that side. For image model, the checkpoint dey follow the same download rule, but the folder wey the interface dey expect na different thing, and setting up your own image generator show where the checkpoint suppose land.

FAQ

blob na the web page wey dey show the file inside browser, so curl go just save HTML. resolve na the file itself. Change the word inside the URL and you get https://huggingface.co/<user>/<repo>/resolve/main/<file>. The resolve link dey redirect go a storage host, so pass -L to curl. Without -L, curl go save the redirect response and you go get tiny file wey no be weights.

My download cut for middle. I go start from zero?

No, if your client fit resume. curl -L -C - -o model.gguf <url> go check how many bytes don land inside the local file and ask the server for only the remaining part. wget -c do the same job. For the hf CLI, any file wey don complete inside the cache no dey download again, so running the same command again go only fetch the ones wey never finish. Partial blob dey stay as .incomplete file, and hf cache prune na wetin dey clear them when you wan free space.

Where hf download dey put the file, and how I go change am?

Default na the cache under your home folder, ~/.cache/huggingface/hub. No guess am: the command dey print the full path after e finish, and hf cache ls dey list every repo wey dey inside with im size. To move the cache go a bigger volume, set HF_HOME (or HF_HUB_CACHE for the model cache alone) before you run the command, then pull one small file and read the path wey e print. To drop the file inside one exact folder instead of the cache, use --local-dir.

Why gated repo still dey reject me after I login?

Login and access na two different thing. You must open the model page inside browser while you don login, read the terms, and click the agree button. Some repo dey approve automatic, some need the author approve by hand, and that one fit stay days. Even after approval, your VPS need token wey carry read access, and fine grained token need the read access to contents of public gated repos option turned on. Check am with curl -sIL -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $HF_TOKEN" <url>. 401 mean token problem, 403 mean that account never get access to the repo.