Air-gapped AI server: what it takes
An air gap means no network path at all, and a rented VPS cannot give you one. Build egress-denied inference instead: default-deny outbound, staged weights.
What an air-gapped AI server is, and why a rented VPS is not one
An air-gapped AI server has no network path to anything: no cable in the socket and no route out of the room. Data arrives and leaves on physical media that a person carries. That is a property of a building and a locked rack, not of a configuration file.
A VPS is a guest on hardware you do not own. It has a virtual NIC (network interface card) wired into the provider's network, and a hypervisor on the other side of that NIC with access to the guest's memory and its disk. You can drop every packet in software. The path still exists. If a requirement uses the words "air-gapped" and means them, a rented server does not satisfy it, and no firewall rule changes that.
Most people who ask for an air gap want something narrower: this data must not leave my control. That you can build on a server you rent. The working name for it is egress-denied inference. The model weights sit on local disk. The box refuses to open outbound connections. The inference endpoint listens only where you can reach it, and nothing left running on the machine has a standing reason to call home. This guide builds that, names each thing it breaks, and closes on the requirements it genuinely cannot meet. If you have not yet decided between running this yourself and buying it as a service, the trade between a managed private AI server and self-hosting is the question to settle first, and what the hardware actually costs to build is the other half of it.
What a stock Ubuntu server talks to on the internet
Before you deny anything, write down what the box currently reaches. A blanket deny applied to an uninventoried server does not fail loudly. It fails weeks later, in one subsystem, for a reason nobody connects to the firewall change.
On a stock Ubuntu 24.04 server these are the usual outbound talkers:
apt-daily.timerandapt-daily-upgrade.timerfetch package lists and apply unattended upgrades over HTTPS to the archive mirrors.snapdrefreshes installed snaps from the Snap Store over HTTPS, on its own schedule, whether or not you use snaps.motd-news.servicefetches the login banner text from an Ubuntu URL. It is controlled byENABLEDin/etc/default/motd-news.ua-timer.timer, from the Ubuntu Pro client, checks entitlement and ESM (expanded security maintenance) package availability.systemd-timesyncdsends NTP (network time protocol) packets on UDP 123 to keep the clock correct.systemd-resolvedsends DNS (domain name system) queries on UDP and TCP 53. Almost everything above depends on this one.cloud-initand any provider guest agent talk to the metadata service on the link-local address 169.254.169.254, mostly at boot.- An ACME (automatic certificate management environment) client such as certbot reaches the certificate authority on TCP 443, roughly every 60 days.
apportcollects crash reports, and on images wherewhoopsieis installed, those reports are uploaded.
Enumerate yours rather than trusting that list. Run systemctl list-timers --all and read every unit name against what you actually installed. Run sudo ss -tunap a few times across an hour and note the remote addresses that appear. Then decide each entry deliberately: keep it and write an allow rule, or drop it and turn the service off so it stops retrying forever in the background.
A slower, more honest inventory method
Rather than guessing, let the firewall tell you. Raise the log level with sudo ufw logging medium before you switch the outgoing policy, leave the box running under its normal workload for a day, then read /var/log/ufw.log and group the blocked entries by destination port. Each distinct port is a decision you have not made yet. This finds the talkers that only wake on a weekly timer, which a one-hour ss sample will miss.
Stage the model weights before you close the door
Once egress is denied, ollama pull cannot reach the registry, so a model that is not already on disk is not going to arrive. Pull everything first, while the box can still fetch.
ollama pull qwen3:8b
ollama listCompare the ollama list output against the models your application names in its requests. A missing tag here becomes a failed request after the lockdown, and the failure will read as an inference error rather than a network one. If you are still choosing, which models fit on a server you can afford is worth settling before you stage anything, because re-opening egress to fetch a second model undoes the point of the exercise.
When the server should never have registry access at all, stage on a different machine and copy the model directory across. On Linux, the packaged Ollama service stores models under /usr/share/ollama/.ollama/models.
# on the staging machine
sudo tar -C /usr/share/ollama/.ollama -cf models.tar models
# on the locked server, after transferring models.tar
sudo tar -C /usr/share/ollama/.ollama -xf models.tar
sudo chown -R ollama:ollama /usr/share/ollama/.ollama/models
sudo systemctl restart ollamaThat directory holds blobs/, the content-addressed weight files, and manifests/, which maps a model name like qwen3:8b onto those blobs. Copy blobs/ without manifests/ and the disk fills up while ollama list shows nothing, because the name lookup has nothing to resolve against. Copy both. The chown matters because the service runs as the ollama user, so files arriving as root are unreadable to the process that needs them.
Set the default outbound policy to deny
ufw calls the outbound direction outgoing, and it maps to the kernel's OUTPUT path: packets that originate on this machine. Set the policy before you enable the firewall, and keep your allow rules narrow.
sudo ufw default deny incoming
sudo ufw default deny outgoing
sudo ufw allow in 22/tcp
sudo ufw enable
sudo ufw status verboseRead the Default: line that ufw status verbose reports on your image and confirm the outgoing policy is the one you just set. That single line is the lockdown. Everything else in this guide is exceptions to it.
Your current SSH session survives this, because /etc/ufw/before.rules accepts packets in the RELATED,ESTABLISHED connection state on the output path. Replies to an inbound connection are part of an established flow, so they are accepted before any user rule is consulted. New outbound connections are what stops. Keep a second SSH session open anyway while you work, which is the same discipline that applies to any ufw rule change on a remote VPS.
Outbound allow rules use allow out:
sudo ufw allow out to 10.10.0.5 port 53 proto udp
sudo ufw allow out to 10.10.0.5 port 53 proto tcpufw resolves a hostname at the moment you add the rule and stores the resulting address, so a rule written against a name that later points somewhere else keeps filtering for the old address with no warning. Write outbound rules against addresses you control, and treat any rule aimed at a CDN-backed hostname as temporary.
If IPv6 is enabled in /etc/default/ufw, ufw writes a parallel v6 ruleset, and an address-family you forgot about is a working route out. Confirm what is listening and what is permitted on both families, because an IPv6 port left open behind a v4-only ruleset is the standard way a lockdown turns out to be half a lockdown.
Read the same ruleset through nftables
ufw is a front end. On Ubuntu 24.04 the iptables command is the nft-backed implementation, so ufw's chains land in the kernel's nftables ruleset alongside anything else that writes rules. Checking through nft shows you the whole picture instead of ufw's view of its own work.
sudo update-alternatives --display iptables
sudo nft list rulesetThe first command tells you which backend is selected on your image, which decides whether the second command can see ufw's rules at all. In the ruleset, find the chain that carries ufw's user-defined output rules and check that its entries match what ufw status claims. A mismatch means something else has written to the kernel since ufw last reloaded.
This is also where you find rules you did not write. Docker, Kubernetes and some VPN installers insert their own chains, and they do not ask ufw first.
Does denying outgoing traffic stop Docker containers?
No, and this catches people who have done everything else correctly. ufw default deny outgoing sets the policy on the OUTPUT path, which handles packets the host itself originates. A container sits on a bridge network, so its packets are routed by the host and traverse the FORWARD path instead, where Docker has installed its own accept rules. Your model container can reach the internet while ufw status reports outgoing as denied.
Run your egress test from inside the container, not from the host shell, or you will test the wrong path. To filter container traffic, write rules into Docker's DOCKER-USER chain, which is consulted before Docker's own accept rules, and verify the result with nft list ruleset rather than ufw status. The inbound version of this same gap is well documented: published container ports bypass ufw entirely, for the matching reason on the DNAT path.
What breaks, and what to do about each
DNS stops. Deny UDP and TCP 53 outbound and no name on the box resolves, which takes down apt, ACME and any application making outbound API calls, all with errors that mention hostnames rather than firewalls. Decide between running a resolver on the box with a static zone for the few names you need, or allowing 53 out to one resolver address on your private network.
The clock drifts. systemd-timesyncd needs UDP 123 outbound. Time matters because TLS (transport layer security) certificate validation compares the current time against the certificate's validity window, and token expiry checks do the same, so a clock that has slipped produces authentication failures that look like credential problems. Either allow UDP 123 to one NTP address, or take time from the hypervisor: load the ptp_kvm module, check whether /dev/ptp0 appears, and if it does, point chrony at it with a refclock PHC /dev/ptp0 poll 2 line in /etc/chrony/chrony.conf. Run that check on your own host rather than assuming, because the device only appears when the hypervisor exposes it. Run one time daemon, and use timedatectl to confirm which one is active.
Certificate renewal fails silently. An ACME client needs outbound 443 to reach the certificate authority's API, and http-01 validation also needs inbound 80 from the CA. A blanket outbound deny does not break anything on the day you apply it. It breaks about 60 days later, when the certificate expires and browsers start refusing the site. Pick one of three answers deliberately: allow outbound 443 to the CA and accept that hole, move to a DNS-01 challenge with an allow rule aimed at your DNS provider's API, or stop using a public CA for an endpoint nobody on the internet reaches and issue an internal certificate instead. The last option needs the reader-run setup in adding your own CA to the Ubuntu trust store, otherwise every client on the box rejects your certificate.
Patching stops. With no path to the mirrors you cannot install security updates, and a private model server running an unpatched kernel is not more secure than one that can reach an archive. Pinning an allow rule to a mirror's address is fragile, because the archive hostname resolves to a rotating set of CDN addresses and your rule follows only one of them. Use a maintenance window instead: re-open outbound 443, run the upgrade, close it again, and record the date. In between, track what you are exposed to by checking the installed package versions against known CVEs rather than assuming quiet means safe.
Provider integration breaks. cloud-init and guest agents read configuration from the metadata service on 169.254.169.254. That address is link-local, so an allow rule for it does not grant internet access, but denying it can change what happens on your next reboot. Read your provider's documentation and decide rather than discovering it during a restart.
Telemetry and update checks. Turn off what you are not using rather than leaving it to retry: sudo systemctl disable --now apt-daily.timer apt-daily-upgrade.timer, set ENABLED=0 in /etc/default/motd-news, and hold snap refreshes with sudo snap refresh --hold. For the inference stack itself, treat update checks as something to test, not to assume. Start the service under the deny policy, exercise it, and read journalctl -u ollama -n 200 for connection failures. Whatever appears there is a call home you now know about.
Bind the inference endpoint where only you can reach it
Ollama's documented default bind address is 127.0.0.1 on port 11434, which is loopback only. Anything that changed OLLAMA_HOST to 0.0.0.0 undid that, and on a box whose firewall you are still building, that is a listening API with no authentication in front of it. Set it explicitly through a systemd drop-in so the value is in version control rather than in someone's shell history.
sudo install -d -m 755 /etc/systemd/system/ollama.service.d
printf '[Service]\nEnvironment="OLLAMA_HOST=127.0.0.1:11434"\n' \
| sudo tee /etc/systemd/system/ollama.service.d/bind.conf
sudo systemctl daemon-reload
sudo systemctl restart ollamaThe Ollama documentation describes the same change through systemctl edit ollama.service, which opens an editor and writes the same drop-in file. The file form is what you want in a build script.
Then list the listening sockets with sudo ss -tlnp and find the row for port 11434. Read the local address column and confirm it is the interface you intended. A loopback bind means the only way in is from the machine itself, so reach it over an SSH tunnel from your laptop:
ssh -N -L 11434:127.0.0.1:11434 you@your-serverNow the port exists nowhere on the public internet and there is no inbound rule to get wrong. When a tunnel per user is too awkward and you need the API reachable from a private network, do not simply widen the bind address: put authentication in front of the Ollama API first, because the API has none of its own and any client that can reach the port can load models and read every prompt.
How do you prove the box is not talking out?
Do not trust the configuration. Test it, and run each test yourself on your own image rather than taking a published result on faith.
Start with resolution and reach. Run getent hosts example.com and then curl -sS --max-time 5 https://example.com, and record what each one does and how long it takes. The --max-time matters: a denied packet that is dropped rather than rejected produces a hang, not an error, so without a timeout the test looks like a broken terminal. Repeat both commands inside your model container if you run one, because that traffic takes the FORWARD path and is governed by different rules.
Then check the listening side from somewhere else. From a second machine, scan the ports your server exposes and compare the result against the inbound rules you wrote, using the methods in checking whether a port is really open from outside. A port that answers from the outside and does not appear in ufw status is a rule some other tool installed.
Finally, re-read /var/log/ufw.log a week later. Blocked entries appearing at a weekly cadence are the timers you missed in the inventory step. If you are hardening a fresh machine rather than retrofitting one, the ordering in the first ten minutes on a new VPS puts the firewall and SSH work before anything is installed, which makes this inventory much shorter.
What a rented server cannot give you
Everything above controls the guest. None of it controls the host. The hypervisor operator can read guest memory, snapshot the disk, and observe traffic on the virtual NIC. Full-disk encryption on the guest protects the volume when it is detached and at rest, and encrypting a storage volume you rent is worth doing, but the key is in the running guest's memory, which the host can read. That is the honest boundary, and no ruleset moves it.
Confidential computing narrows the gap. AMD SEV-SNP and Intel TDX encrypt guest memory with a key the host cannot use, and remote attestation lets you verify what launched before you send it a secret. The practical catch is that attestation wants to fetch vendor certificates over the network, which is exactly the outbound reach you just removed, so you end up caching those certificates during the build step. What confidential computing on a VPS does and does not prove is worth reading before you rely on it, because it moves trust from the operator to the CPU vendor rather than removing trust.
So here is the boundary, plainly. If your requirement is that no third party can read the data you process, a rented VPS cannot satisfy it, and confidential computing changes who you must trust rather than eliminating the question. If your requirement is that your prompts and documents are never sent to a model API, that no vendor retains them, and that the server does not open connections you did not authorise, an egress-denied VPS satisfies all of it and you can demonstrate each part with the commands above. If a regulator, a contract or an auditor wrote the word "air-gapped", buy hardware, put it in a room you control, and do not connect it. Everything in between is this guide.
FAQ
Can a VPS ever be genuinely air-gapped?
No. An air gap means no network path exists, and a VPS has a virtual NIC attached to the provider's network by definition. A default-deny outbound firewall stops the guest from using that path, which is a real and useful control, but the path and the hypervisor on the far side of it are still there. If a compliance document specifies physical isolation, that requirement can only be met on hardware in a location you control.
Will ufw default deny outgoing disconnect my SSH session?
It should not, because /etc/ufw/before.rules accepts outbound packets in the RELATED,ESTABLISHED connection state before any user rule is evaluated, and the replies in your SSH session belong to an already established connection. What stops is new outbound connections the server initiates. Keep a second session open while you make the change anyway, so a mistake in the inbound rules does not leave you locked out.
Why did my Let's Encrypt certificate expire after I denied outbound traffic?
Because the ACME client needs outbound TCP 443 to reach the certificate authority's API, and the deny policy blocked the renewal attempt. The failure is delayed: nothing breaks on the day you apply the policy, and the site stops working about 60 days later when the existing certificate runs out. Either allow outbound 443 to the CA, switch to a DNS-01 challenge with an allow rule for your DNS provider's API, or issue certificates from an internal CA and add that CA to the server's trust store.
Does denying outgoing traffic stop my Docker containers from reaching the internet?
No. The ufw outgoing policy applies to the OUTPUT path, which covers packets the host itself originates. Container traffic is routed by the host and traverses the FORWARD path, where Docker's own accept rules apply. Test egress from inside the container rather than from the host shell, and filter container traffic by adding rules to the DOCKER-USER chain, then confirm the result with sudo nft list ruleset.
How do I update packages on a server with no outbound access?
Use a scheduled maintenance window. Re-open outbound 443, run the upgrade, close the policy again, and record the date you did it. Pinning a permanent allow rule to the archive mirror is unreliable, because the mirror hostname resolves to a rotating set of CDN addresses and a rule stores only the address it resolved when you added it. Between windows, track the CVEs that apply to your installed package versions so that a quiet server does not become an unpatched one.