Skip to content
RO
← All writing

The backup that had not failed

  • linux
  • backups
  • operations
  • monitoring

On 14 August the MeshCentral server I run put a red line on its own status page:

Backup failed (10000)
handleBackupRequest: Backup error

That box holds the identity of every agent that connects to it. I read “backup failed”, drew the obvious conclusion, and wrote down that there was no backup of the most important machine I administer.

That was wrong. It took about twenty minutes to find out how wrong, and the correction is more useful than the bug.

10000 is not an error code

It looks like one. It is a bitmask, printed in binary, by this line:

backupStatus.toString(2).slice(-8)

BACKUPFAIL_ZIPMODULE is 0x0010, which is 10000 in binary. So the number is not a code you can look up, it is five bits with one of them set, and the digits happen to look like a round decimal number. Worth knowing the neighbours too: 0x0001 prints as 1, and 0x0100 prints as 00000000, because slicing the last eight characters of a nine-character string quietly drops the bit you care about.

A status line that renders one failure as 10000 and another as 00000000 is not really a status line.

The backups had never stopped

Here is the part I got wrong. I went and looked in /opt/meshcentral/meshcentral-backups expecting an empty directory or a stale file from months ago.

Eight zips. About 13MB each, one a day, every day at 18:43, the most recent one from the previous evening.

Only the manual “Download server backup” button was failing. That path asks for a password, so it takes the encrypted branch of the code. The scheduled nightly run has no password configured, takes the else branch, and produces a plain zip that had been working the whole time. One warning, two completely different code paths, and the banner does not tell you which one it is talking about.

Read the logs before you believe a red warning describes everything. I did not, at first, and I had already written the wrong version down.

What was actually broken

db.js line 3714 does require('archiver-zip-encrypted') when a backup password is in play. That module is not a dependency of MeshCentral. Not in dependencies, not in optionalDependencies, so npm install meshcentral never brings it and it was simply absent from the tree.

The require threw, the catch set the failure bit, the backup aborted.

One thing that deserves credit: the code refuses to fall back to an unencrypted zip when you asked for an encrypted one, and it skips the old-backup cleanup on failure. So nothing was silently downgraded to plaintext and nothing was deleted to make room for a backup that never arrived. Plenty of software would have done both and told you it was fine.

Why I did not just run npm install

The obvious fix is one command. It was also the wrong one.

package.json pins "meshcentral": "^1.2.3". Installed was 1.2.3. The registry had 1.2.5. A plain npm install in that directory would have quietly upgraded the live server that every agent connects to, on a box where "SelfUpdate": false is set deliberately, in the middle of fixing something unrelated.

So instead: install the module into a scratch directory under /tmp, list its seven direct dependencies, and compare them against what was already on the box. Six were already there at identical versions. I copied across the two that were missing, then confirmed meshcentral was still 1.2.3 afterwards.

No restart either. The require happens inside the backup function at call time, and Node does not cache failed resolutions, so the running process picks up the module on the next attempt. Restarting MeshCentral drops every agent connection, and there was no reason to pay that.

The test that failed for the wrong reason

Then a third false signal, which is why I am counting them.

I ran unzip -P against the new encrypted output to prove it opened. It failed. For a moment that looked like the fix had not worked.

Info-ZIP’s unzip cannot decrypt AES at all. The failure said nothing about my file. Verifying it properly meant reading the bytes: compression method 99, an 0x9901 extra field, strength byte 3 for AES-256, deflate underneath, and the plaintext genuinely absent from the archive. There is no 7z on that VPS, so opening one by hand needs a different machine.

Three things reported a problem that day and none of the three was describing what I thought it was.

The two things worth fixing were never in the warning

Nothing had alerted on either of these, because neither is an error.

The nightly zips were unencrypted, since no backup password is configured and that is what the plain branch produces. That one is a decision rather than a bug: a password you lose is a backup you lose, so it needs a documented home before it is worth turning on.

The second was the real one. The backups lived on the machine they were backing up, with a default retention of ten days. Lose the VPS and you lose the thing that was supposed to survive losing the VPS.

That got fixed the same day, and the design choice is the interesting part. My home server pulls the backups over WireGuard rather than the VPS pushing them. The VPS holds no credentials for the other end at all, so an attacker who takes the VPS cannot reach through it and destroy the off-box copies, which is the exact event those copies exist to survive. The rsync has no --delete for the same reason: mirroring deletions would faithfully reproduce “somebody wiped the backups”. Retention is handled separately at each end instead, ten days on one and ninety on the other.

That design got extended in September to cover every app and website on the box, which is when I found out that the largest database on the server had never been in a backup at all. Good architecture copying the wrong set of things is still the wrong set of things.

The whole security story is one line in authorized_keys:

command="/usr/local/bin/rrsync -ro /opt/meshcentral/meshcentral-backups",restrict,from="10.8.0.2"

Read-only, no pty, no forwarding, and only from one address on the WireGuard network. It is the same instinct as making every script back up what it touches: assume the thing you are protecting against is the one that will happen.

The cost of that design showed up in September, when the home server dropped off its own LAN for an hour and the pull quietly stopped happening while every backup on the VPS carried on succeeding. Nothing anywhere looked wrong. That one is written up separately, because the tools I used to find it were the problem.

What this is actually about

A monthly arrangement to look after somebody’s systems is mostly this. Not heroics, and not a dashboard full of green ticks. Somebody who opens the directory and counts the files when a banner claims something is broken, and who goes looking on the days nothing is claiming anything at all, which is how I found 193,000 failed logins on the same server.

Backups are the clearest case, because a backup you have not restored is a rumour. If you want somebody doing that work on your systems, that is what the care plan is.

One more thing, noted for whoever inherits this. Those two modules are not tracked in any package.json on that box. A future MeshCentral upgrade will not know they should be there, and the failure will look exactly like it did on 14 August: one button, one red number that is not a number, and eight perfectly good backups nobody thought to count.