Skip to content
RO
← All writing

How do you test a server backup?

  • linux
  • backups
  • restore
  • vps

Most people asking this already have backups. They have had them for years, and the job reports success every night. What they want to know is whether that means anything.

On its own, no. A success message tells you a program finished. It says nothing about whether the thing it wrote contains your data, or whether you could turn it back into a working server on the day you need to. Those are separate questions, and only one of them gets checked by default.

Five things that get called testing

They are not equal. I’d put them in this order, weakest first.

The job exited 0. This is the green tick, and it is worth surprisingly little. When I built a nightly backup for every app on my own VPS, mongodump reported success for seven databases that did not exist, writing a small stub file each time. Exit code 0 throughout. Not one byte of anybody’s data.

The file exists and is a sensible size. Better, and cheap to automate. The same script now rejects any archive under 200 bytes. Size also caught the opposite problem that week: a tar exclude written with a leading slash matched nothing, and the archive came out at 421M instead of about 100M.

The archive passes an integrity check. gzip -t on a tarball. restic check on a restic repository, with one catch that is easy to miss. Restic’s own documentation says that by default the check “does not verify that the actual pack files on disk in the repository are unmodified”, because that means reading every one of them. You have to ask for it with --read-data, or --read-data-subset to work through a large repository a slice at a time.

The data comes back somewhere else and you look at it. This is the first one I would call a test.

The whole service comes back from nothing, and you time it. This is the one almost nobody does.

The first three are useful as alarms. They run unattended and they catch a job that has quietly stopped. None of them proves you can recover.

Restore it somewhere that isn’t production

Never test by restoring over the live server. If the backup is bad, you have just found out in the worst possible way.

Spin up a throwaway VPS billed by the hour, or a container on another machine, and restore into that. For a database, open it and count rows in the two or three biggest tables against the live copy. A dump that restores cleanly but is missing last Tuesday’s orders is still a failed backup. For files, pick a handful at random and compare checksums, and open at least one document a person would actually miss.

Then start the application against the restored data and log in. That last step is where most of the surprises are.

Expect the tools to lie to you a little on the way. mongorestore --dryRun reports “0 documents restored” for a perfectly good archive, because a dry run doesn’t restore anything; walking the index with -vv is what tells you what’s inside. And when I checked an AES-encrypted zip with Info-ZIP’s unzip, it failed, which looked like a broken backup and was really a tool that cannot decrypt AES. Three false signals on my own servers, each one confidently wrong. A test you don’t understand is just a different kind of green tick.

What a restore test actually measures

Not the data, mostly. By the time you’re restoring, the data is usually fine. What a restore finds is everything that was never in the backup because nobody thought of it as data.

The nginx config. The .env file with the database password in it. The cron entries, the systemd timers, the firewall rules. The TLS setup, which will want the domain pointing at the new machine before it’ll issue anything. A backup of /var/www and a database dump gets you the website’s contents and leaves you a day of reconstruction from memory.

The passphrase, above all. If your backups are encrypted, and off-site ones should be, the passphrase is the backup now. The restic docs print the warning every time you create a repository: “Losing your password means that your data is irrecoverably lost.” So where is it written down, and could anybody other than you find it?

And the clock. Write down how long the restore took, start to finish, including the bits where you had to go and look something up. That number is your real recovery time, and it is usually several times what anybody would have guessed. The UK GDPR asks for exactly this. Article 32(1)(c) wants “the ability to restore the availability and access to personal data in a timely manner” after a physical or technical incident, and (d) wants “a process for regularly testing, assessing and evaluating the effectiveness” of your security measures. A restore you have timed is evidence for both. A nightly success email is evidence for neither.

How often

The NCSC doesn’t give a number, and I think it’s right not to. Its August 2019 post on offline backups says backups “should also be regularly tested to check they work as expected”. Its principles for ransomware-resistant cloud backups, published November 2024, go further and say a restore test can be run on demand, in a way that is not destructive to your existing infrastructure, and that the owner should test “as part of a regular monitoring process”.

My own answer has two parts. The cheap checks (size, age, integrity) run automatically, every time. A real restore into a scratch machine happens on a schedule you will actually keep, and quarterly is a sensible floor for a small business server.

The second part matters more. Test again whenever something new lands on the box. A new app, a new database, a new directory somebody started saving uploads into. Backups go stale on the day the server changes, not on a calendar. The biggest database on my own VPS sat outside every backup for months, while every job carried on succeeding, because it arrived after the jobs were written.

One more failure that no restore test will catch: the backup that stops arriving. My off-site copies are pulled to a machine at home, and when that machine dropped off its own network for an hour, the pulls stopped and nothing on the server looked wrong. Something needs to alert when the newest copy is too old, and it needs to run somewhere other than the machine that might be down.

Host snapshots count, but not alone

If your VPS provider takes a daily image, keep it. It’s a good fast way back from a bad update. It’s also stored by the same company and reached through the same account login as the server, which I’ve written about separately. Restore from it once too. An image taken while a database was mid-write isn’t something to trust untested.

What I do, and what I don’t

Backups are part of my infrastructure service: encrypted, off-site, and tested by restoring them rather than by reading a status page. One-off setup, including a first full restore test, is a fixed price. Ongoing cover is monthly and priced per server, not per hour, and repeating the tests is part of what that pays for. The documentation records where the backups go, where the passphrase lives, and the steps to rebuild, written so somebody who isn’t me can follow them.

The limits, plainly. This is Linux servers. Laptops and Microsoft 365 mailboxes are a different job and I don’t sell it. I can’t restore something that was never in a backup, which is why the first job on any server I take over is working out what lives on it. And I’m one person, not a recovery team on shift, so if an hour of downtime costs you thousands, you want a provider with people awake at 4am as well as me.

If you’d rather do it yourself, do this one thing this week. Take last night’s backup of your biggest database, restore it onto a machine you can throw away, and count the rows. If the numbers match, you have a backup. If you can’t do it at all, you’ve found the problem in the cheapest way there is.