Patched, and still three kernels behind
- linux
- almalinux
- security
- operations
- patching
On 1 September a monthly reminder I set up in August went off and told me to check the kernel on the VPS I run. It is the same box I counted 193,000 failed logins on. I expected the check to find nothing. It found the box running a kernel with three Important security advisories outstanding, all three of which had already been downloaded, installed and sat waiting on disk for days.
Nothing was broken. That is the part worth writing down.
What the check found
The server had been up two weeks and three days, running kernel 5.14.0-687.38.1. Installed and waiting since 29 August: 687.42.1. In between it had also silently acquired 687.39.1 and 687.41.1.
Three Important kernel advisories, then, none of them running:
| Advisory | Kernel |
|---|---|
| ALSA-2026:54443 | 687.39.1 |
| ALSA-2026:57252 | 687.41.1 |
| ALSA-2026:59723 | 687.42.1 |
I had predicted this in August, which is the only reason the reminder existed. I did not predict it would be three deep after a fortnight.
Why the box was both patched and vulnerable
dnf-automatic is set to apply_updates = yes and reboot = never. It runs, it fetches the security updates, it installs them, and it stops. That is exactly what I told it to do and it does it faultlessly.
The trouble is what “installed” means for a kernel. Every other package is live the moment it lands. A kernel is a file on disk until something boots it. So rpm -qa is honest, the package manager is content, dnf-automatic’s logs are clean, and the code actually executing in memory is whatever was current the last time the machine restarted.
There is no error state here. Nothing fails, nothing warns, nothing goes red. The gap only exists in the space between two questions that most tooling treats as the same question: what is installed, and what is running.
The command that says it out loud
Last time I counted advisories with dnf updateinfo list --security | wc -l. That came back with 15 and I had to read all fifteen, because the list mixes kernels I had already installed in with userland packages that were genuinely still pending. A number that needs interpreting is not much of a check.
dnf check-update --security is better, and it is better because it stops being a list and starts being a sentence:
kernel-core-...687.42.1 is an installed security update
kernel-core-...687.38.1 is the currently running version
Two lines, side by side, naming the exact thing that is wrong. I have swapped the monthly check over to it.
Two things I nearly got wrong on the way
Before rebooting I check /root/.pm2/dump.pm2, because pm2 brings back the apps it has saved and a stale dump quietly restores fewer than you had. Last month the dump predated an app going live. This month it was dated 19 August and did list all the apps by name, so on the face of it there was nothing to do.
I ran pm2 save anyway, and I would again. It costs nothing, it is safe to repeat, and it is the same instinct as making sure every script backs up what it touches. Names matching is not the whole test. The dump also stores each app’s working directory, script path and environment, and any of those can drift while the name sits there looking correct. The check I actually did was weaker than the check I thought I was doing.
The app count was wrong too. Every note I had, including the reminder’s own verification step, said six pm2 apps. It is seven. lrn-chat, the chatbot behind this site, joined in late August and none of the documentation caught up. A verification step that counts to six on a seven-app box passes while something is missing.
Baseline before, not after
The other correction is smaller and I will keep making it if I do not write it down.
I curl a list of URLs after a reboot to confirm the sites came back. The apex alphaitsolutions.uk, with no www, returns 301. That is correct and healthy: it redirects to the www host, as it has for months. But an expectation of “everything should be 200” turns that into a regression, and I would have spent twenty minutes chasing a redirect that was working perfectly.
So the check is the comparison, not the code. Capture what every URL returns before you touch anything, then look for what changed.
The result, and what I left alone
The reboot went as it should. Kernel 687.42.1, all seven pm2 apps back at restart_time=0, pm2-root clean, nginx and mongod and fail2ban all active, no failed units, and all ten URLs matching what they returned beforehand.
Three Moderate userland updates are still pending and I deliberately left them: dbus-broker 28-7 to 28-9, gzip 1.12-1 to 1.12-2, tar 1.34-11 to 1.34-13. The automatic timer last ran at 06:02 UTC on the 1st and those advisories landed after it, so the next run picks them up and none of them needs a reboot. I have made a note to look next month, because if those same three are still sitting there in October then dnf-automatic has stopped applying updates and that is a real fault rather than a scheduling artefact.
I am not changing the cause
reboot = never stays. A box running seven production applications should not restart itself at three in the morning because a package landed, and I would rather carry a known gap I check on a date I chose than hand that decision to a timer.
The same shape turned up a week later on something with no server in it at all: an Android app I could no longer update, because a Google Play deadline had passed while every screen I looked at said the app was fine. What the evidence has changed is my sense of the size of it. On this box the gap grows by roughly one kernel a fortnight. Miss a single month and you are three or four Important advisories behind, with every dashboard you own reporting the machine fully patched, because by its own definition it is.
The honest weakness in my fix is that the reminder runs on my desktop and only fires while I have the right application open. It ran on the 1st because I happened to be at the machine. A systemd timer on the server itself, with somewhere to send the result, does not depend on my Monday morning. That is the version I should build, and until I do, the control here is a habit rather than a system.
