Ansible reboot: only when the server needs it
Gate the reboot on evidence the host reports: the Ubuntu marker file, the Red Hat exit code, then reboot in slices so the whole fleet stays up.
The Ansible reboot that only fires when needed
An Ansible reboot should run only when the host says it needs one. Every Linux family already keeps that answer on disk. Debian and Ubuntu write a marker file. Rocky and Alma answer a command with an exit code. Read the answer, register it, and gate the reboot with when:. A run across thirty servers then restarts the four that need it and leaves the rest alone.
Here is the whole idea in one play. The sections after it explain each part, and the ways each part goes wrong.
- name: Reboot only the hosts that need it
hosts: all
become: true
tasks:
- name: Look for the Debian family reboot marker
ansible.builtin.stat:
path: /var/run/reboot-required
register: deb_marker
when: ansible_facts['os_family'] == 'Debian'
- name: Ask the Red Hat family whether a reboot is needed
ansible.builtin.command: needs-restarting -r
register: rh_marker
changed_when: false
failed_when: false
check_mode: false
when: ansible_facts['os_family'] == 'RedHat'
- name: Store one answer for both families
ansible.builtin.set_fact:
reboot_needed: '{{ (deb_marker.stat.exists | default(false)) or (rh_marker.rc | default(0) != 0) }}'
- name: Reboot the hosts that asked for it
ansible.builtin.reboot:
msg: Reboot requested by the patching playbook
reboot_timeout: 900
post_reboot_delay: 30
when: reboot_needed | boolTwo details in that play are easy to miss. set_fact stores the result of a template as text, so reboot_needed holds the string True or the string False, not a boolean. That is why the gate reads reboot_needed | bool. The second detail is the default() filters. A task skipped by its own when: still registers a result, and that result carries no stat key and no rc key. Remove the defaults and the play stops on the first Ubuntu host with 'dict object' has no attribute 'rc'.
If this is your first playbook, start with a plain playbook run against a single VPS and come back. A reboot is a poor place to learn the syntax.
How does Ubuntu say that a reboot is needed?
On Debian and Ubuntu the signal is the existence of the file /var/run/reboot-required. Its content is a sentence for humans. Your playbook should read nothing except whether the path is there, which is what ansible.builtin.stat reports in stat.exists. A companion file, /var/run/reboot-required.pkgs, lists the packages that asked, one name per line, which tells you whether the trigger was the kernel or something smaller.
/var/run is a symbolic link to /run, and /run is a tmpfs held in memory. The marker is erased by the reboot itself, so nothing has to clean it up and the file can never go stale.
The file is created by the script /usr/share/update-notifier/notify-reboot-required, which package maintainer scripts call after they install something that needs a restart. That script belongs to the update-notifier-common package. Ubuntu server images ship it. Minimal Debian images often do not, and this is the failure that quietly empties the whole gate: with no script present, no package can create the marker, so stat.exists stays false forever and your playbook decides that no Debian host has ever needed a reboot.
Check this once per image, on your own host:
ls -l /usr/share/update-notifier/notify-reboot-required
ls -l /var/run/reboot-required /var/run/reboot-required.pkgsDoes the first path exist? If it does not, install the package before you trust the gate.
sudo apt update && sudo apt install -y update-notifier-commonDebian 12 and later also ship needrestart, which answers a different question: which running processes still use library files that have already been replaced on disk. It prompts a human for service restarts, and it falls back to list only mode when it runs non-interactively, which is what happens under automatic security updates on Ubuntu. Service restarts and reboots are separate problems, so solve them separately.
How does Rocky or Alma say that a reboot is needed?
The Red Hat family has no marker file. It has a command, needs-restarting, from the dnf-plugins-core package.
sudo dnf install -y dnf-plugins-core
needs-restarting -r; echo $?The documented job of -r is to "only report whether a reboot is required (exit code 1) or not (exit code 0)". That non-zero exit is the whole signal, and it is also why the Ansible task needs two extra lines. A non-zero exit code makes ansible.builtin.command fail the task, so failed_when: false keeps the run alive and hands you the rc to test yourself. changed_when: false stops a question from being counted as a change, which keeps the run summary honest.
This command has moved between packages and between dnf versions over the years. Run it by hand once on every image you manage and watch what it does. Whatever text it prints, the exit code is the part your playbook reads. What needs-restarting reports after an update on Rocky and Alma walks through the output line by line, including needs-restarting -s, which lists the systemd services that would be enough on their own.
The reboot module, parameter by parameter
ansible.builtin.reboot does much more than run reboot. Before it sends anything it records a boot identifier using boot_time_command, which defaults to cat /proc/sys/kernel/random/boot_id. After the connection returns it reads that identifier again, and it continues only once the value has changed. This is what makes the module trustworthy. A host that ignored your reboot command still answers SSH, so a connection that works again proves nothing on its own. A changed boot id proves the kernel is new.
The parameters that matter, with the defaults given in the module documentation:
reboot_timeout, default 600 seconds. The entire budget for the host to go down and answer again. A server that runs a filesystem check at boot, or one on slow shared storage, can pass 600 seconds without being broken.post_reboot_delay, default 0. Seconds to wait after the connection returns, before the play continues. Set it when the next task talks to a service that starts late.pre_reboot_delay, default 0. Seconds between the module deciding to reboot and the shutdown command running.test_command, defaultwhoami. The command that has to succeed on the far side.msg, defaultReboot initiated by Ansible. This is the wall message that logged in users see.search_paths, default/sbin,/bin,/usr/sbin,/usr/binand/usr/local/sbin. PATH is ignored on the remote node while the module hunts for the shutdown binary, so an unusual location has to be added here.
test_command deserves a thought. whoami proves that a login shell works. On a systemd host, systemctl is-system-running --wait waits for the boot to settle and reports degraded with a non-zero exit when any unit failed to start, which turns a quietly broken boot into a failed play. That is usually what you want inside a patch window. It is not what you want on a host that has been sitting at degraded for a month for an unrelated reason, so look at the current state of your fleet before you switch.
The module supports check mode and returns a prediction without touching the host. It returns elapsed, the seconds it waited, and rebooted. It targets POSIX systems; Windows hosts use ansible.windows.win_reboot instead.
Why wait_for_connection is what makes the next task safe
When you use the reboot module, the waiting is already handled. You need ansible.builtin.wait_for_connection when the restart happens some other way: a handler that schedules shutdown -r +1, a reboot from your provider's control panel, or a fire and forget task like this one.
- name: Ask for a reboot without waiting for the reply
ansible.builtin.shell: sleep 2 && /sbin/shutdown -r now
async: 1
poll: 0
become: true
- name: Wait for the host to answer on the new boot
ansible.builtin.wait_for_connection:
delay: 30
sleep: 5
timeout: 900delay is the line people leave out, and leaving it out is the reason a play continues too early. The defaults are 0 seconds of delay, 1 second of sleep between attempts, a 5 second connect timeout and a 600 second total timeout. With no delay, the first attempt happens while the old boot is still up and still accepting SSH. The wait succeeds at once, and the next task then runs against a machine in the middle of shutting down. Thirty seconds of delay costs nothing and removes the race.
wait_for_connection is also stronger than the older habit of waiting on TCP port 22 from the control node with ansible.builtin.wait_for. A port check proves that something accepted a connection. That something can be the dying boot, a load balancer, or a bastion forwarding the port. wait_for_connection goes through the real connection plugin and runs the ping module, so it succeeds only once Ansible can actually execute a module on the host, remote Python interpreter included.
Roll the fleet with serial and a health check
Without serial, a play on hosts: all reboots every host at once, up to your fork count. For a pair of load balanced web servers that is an outage. serial splits the run into batches and finishes one batch before starting the next.
- name: Patch and reboot in slices
hosts: webservers
become: true
serial: '25%'
max_fail_percentage: 0
tasks:
- name: Apply updates on the Debian family
ansible.builtin.apt:
update_cache: true
upgrade: dist
when: ansible_facts['os_family'] == 'Debian'
- name: Apply updates on the Red Hat family
ansible.builtin.dnf:
name: '*'
state: latest
when: ansible_facts['os_family'] == 'RedHat'
- name: Reboot if the host asks for it
ansible.builtin.reboot:
reboot_timeout: 900
post_reboot_delay: 30
when: reboot_needed | bool
- name: Hold this slice until it serves traffic again
ansible.builtin.uri:
url: '{{ health_check_url }}'
status_code: 200
delegate_to: localhost
become: false
register: health
until: health is succeeded
retries: 20
delay: 15The detection tasks from the first play belong in this one as well, right before the reboot task. They are left out here to keep the example short.
max_fail_percentage: 0 is evaluated once per batch, so a single failed host in the first slice stops the play before the second slice goes anywhere. Define health_check_url per host or per group in your Ansible inventory file, and point it at something that touches the part you care about, such as a page that queries the database rather than the web server's welcome page. delegate_to: localhost runs the check from the control node, which is the direction your users arrive from. A check run on the host itself reports success happily while a firewall keeps the world out.
Where a kernel livepatch removes the reboot, and where it does not
A livepatch replaces functions inside the kernel that is running right now, so a fix lands with no restart. Canonical ships canonical-livepatch for Ubuntu through Ubuntu Pro, whose free personal tier covered five machines as of September 2026, and Enterprise Linux has kpatch. On a patched Ubuntu host, canonical-livepatch status reports what is currently applied.
Two limits decide how much this changes your playbook.
First, a livepatch applies to the running kernel only. It does not install a kernel, and it does not stop the package manager from installing one later. When a new kernel package lands, the maintainer script still writes /var/run/reboot-required, so your gate still says a reboot is needed. The gate is right. The reboot is simply far less urgent now, which is the real benefit: it moves to a maintenance window instead of happening tonight.
Second, a kernel livepatch does nothing for userspace. An update to a shared library such as glibc or OpenSSL takes effect only in processes started after the update. Every long running daemon keeps the old copy mapped until it restarts. That is a service restart, not a reboot, and it is exactly what needs-restarting -s and needrestart report. How live kernel patching works on a VPS covers which fixes are eligible, because the answer is a subset of them and never all of them.
Testing the play without breaking production
Run it in check mode first, with a dry run and diff output. One trap is waiting there. ansible.builtin.command does not execute in check mode, so the needs-restarting task is skipped, rh_marker carries no rc, and the fact falls back to the default of 0. Your dry run then announces that no Red Hat host needs a reboot, which is a confident and completely empty answer. check_mode: false on that one task is the fix. Asking the question changes nothing on the host, and asking it is the only way a dry run can tell you anything true.
After check mode, run it against one host with --limit, and watch what happens.
ansible-playbook reboot-if-needed.yml --check --diff
ansible-playbook reboot-if-needed.yml --limit web01 -vDoes the play report a change on the reboot task for the host you expected, and skip it on the others? Does uptime on that host now show a fresh boot? Answer both on your own servers before this play meets a whole group.
One thing you cannot do is rehearse a reboot inside a container. A container shares the kernel of its host and has no init system of its own to shut down, so systemctl reboot inside one either fails or kills the container. There is no kernel there to replace, so the thing you are testing does not exist. Use a throwaway virtual machine you can afford to lose. Build it, patch it, break the play against it, then point it at the fleet of Linux servers you actually manage.
FAQ
How do I check whether a Linux server needs a reboot?
On Debian and Ubuntu, check whether the file /var/run/reboot-required exists. It is written by package scripts through update-notifier-common, and /var/run/reboot-required.pkgs names the packages that asked for it. On Rocky, Alma and other Red Hat family systems, run needs-restarting -r from dnf-plugins-core and read the exit code: 1 means a reboot is required and 0 means it is not. Neither check contacts a repository, so both are cheap enough to run on every host at the start of a play.
Does the Ansible reboot module wait for the server to come back?
Yes. It records a boot identifier first, by default with cat /proc/sys/kernel/random/boot_id, sends the reboot, then waits for the connection to return, for that identifier to change, and for test_command to succeed. It gives up after reboot_timeout, which is 600 seconds by default. You need wait_for_connection only when something other than the module performs the restart.
Why does my playbook continue before the server has finished rebooting?
Because the host answering you is still the old boot. A shutdown command returns straight away and SSH keeps accepting connections for a few more seconds, so a wait_for_connection with its default delay of 0 connects, succeeds, and lets the next task run into a shutdown in progress. Set delay to about 30 seconds, or use the reboot module, which watches the boot identifier rather than the connection.
Does kernel livepatching mean I never have to reboot?
No. A livepatch fixes the kernel that is running, so it removes the urgency rather than the reboot. Installing a new kernel package still writes the reboot marker, and the machine keeps running the old kernel until it restarts. Livepatching also does nothing for userspace libraries such as glibc or OpenSSL, because a process that mapped the old copy keeps using it until the service restarts, which needs-restarting -s will show you.