Skip to content
RO
← All writing

What's actually running on your server?

  • linux
  • hardening
  • operations
  • nginx

The call is nearly always the same shape. Somebody set a server up two or three years ago, it has been running ever since, and that somebody has left or stopped answering. It works. Nobody knows what is on it.

People expect me to start with hardening, because that is the word on the service page. I don’t. The first day is spent finding out what the machine is actually doing, and on an inherited box that takes longer than the hardening does.

Here is what that looks like.

The first twenty minutes

Four questions, in this order, and none of them involves opening a config file.

What is listening, and which process owns it: ss -tulpn. What is set to come back after a reboot: systemctl list-unit-files --state=enabled. What runs on a timer, which means every user’s crontab and not just root’s, plus systemctl list-timers. And what has been logging in: lastlog, the auth log, and the authorized_keys file of every account that has one.

That last one produces the most awkward conversations. Keys belonging to people who left in 2023 are extremely common, and nobody ever removed them because nobody knew they were there.

What the machine is serving is not what is in sites-available

If the box runs nginx, nginx -T is the command that matters. It dumps the full configuration as the running process understands it, every include resolved, rather than the files somebody thinks are in play.

I have yet to inherit a server where those two things matched. There is a symlink in sites-enabled pointing at a file somebody edited a copy of. There is a server block for a site that was decommissioned, still holding the default for any hostname that doesn’t match anything else. There is an SSL certificate for a domain that moved to a different provider in March.

The same trap exists everywhere on a Linux box and it is the single most expensive assumption in this work. A file on disk is a statement of intent. The running process is the fact. My own server reported itself fully patched while booted into a kernel three versions behind, and my own header checker told me a site was missing four headers it had had for months, because nginx quietly discards inherited headers the moment a block adds one of its own. Both were mine. Both looked right in the file.

Three piles

Everything the inventory turns up goes into one of three piles.

Things genuinely in use, which are usually fewer than anyone expects. Things nobody remembers installing, which are harmless until the day one of them is the way in. And things listening to the internet with no reason to be, which is the pile that pays for the job.

That third pile is where I find the interesting stuff: a database bound to 0.0.0.0 because a tutorial said to, an admin interface on a high port that was meant to be temporary, a development copy of the site that was never taken down and is running whatever version of everything it was frozen at. I have found all three on my own kit, which is how I know to look.

Then, and only then, the hardening

The list is short and boring, which is the point. SSH cut back to key-only. A host firewall that denies by default rather than one with a long list of allow rules nobody can explain. fail2ban on anything still exposed. Everything in the third pile switched off or bound to localhost.

Then it gets audited with Lynis, because “I made the changes” and “the changes are in effect” are different claims, and I have been wrong about the second one often enough to stop assuming. The verification happens from outside the box. Reading the config back only tells you what you typed.

Two rules I don’t break while doing any of this. The working session stays open and every change gets tested from a second one, so a lockout is always recoverable. And every script backs up what it touches before it changes it. That one came out of breaking a working config with no way back, and it costs about three lines.

Backups are checked by restoring them, not by reading a status page

This is the part I am most insistent about, and I have two write-ups of my own systems that explain why better than an argument would.

One where a server told me its backups were failing and they were not: eight of them sat on disk, one a day, the most recent from the previous evening. What had broken was a manual export almost nobody presses. And one where every job succeeded every night, green across the board, while the largest database on the machine had never been in a backup of any kind.

Neither raised a warning. Warnings tell you what broke. Nothing warns you about the thing nobody thought to include, which is why the only test I trust is restoring the thing somewhere else and opening it.

The part that makes me uncomfortable

If I do all of the above and write nothing down, I have replaced one person who understands the server with a different person who understands the server. That is not an improvement. It is the same risk wearing my face.

So the documentation is part of the job rather than an optional extra, and the standard it has to meet is that somebody who is not me can pick it up: what runs where, what depends on what, where the backups go and how to get one back, and what to do at two in the morning if it stops. If you ever want to take it in-house or hand it to somebody cheaper, that document is the thing that makes it possible. I would rather you could.

What it costs, and what it doesn’t buy you

One-off setup and hardening is a fixed price. Ongoing cover is monthly and priced per server rather than per hour, because per-hour pricing quietly rewards me for things taking longer, and because you want to be able to ring me about something small without doing sums first.

What it does not buy is a team. I am one person. There is an alert on my phone, not a room with staff in it at three in the morning, and if that is the level of cover you need then you want a managed hosting provider and I will say so on the first call. It also doesn’t buy certification: I don’t hold Cyber Essentials and can’t certify anybody, though I can tell you where you’d stand.

If you are sitting on a server nobody has looked at properly, the inventory is worth having on its own, even if you then decide to do nothing. Sometimes it ends with the box being retired rather than hardened, which is a perfectly good outcome and the other half of whether you needed a server at all. Knowing what is listening is cheap. Finding out the hard way is not.