The database nobody had backed up
- linux
- backups
- mongodb
- operations
On Tuesday 9 September I set up a nightly backup of every application and website on my VPS, pulled off the box to a machine at home, seven nights kept at each end. Ordinary work. A systemd timer, a shell script, an rsync over the tunnel.
The backup is not the interesting bit. What building it turned up is.
Nobody had ever asked what was on the box
There were already backups running, and they were fine. The client portal’s cron tarred its own SQLite file every night. MeshCentral wrote its archive of the agent identities at 18:43, which is the job I misread a red warning about in August. Both had been running for months. Both did exactly what they were written to do.
Neither of them went anywhere near MongoDB.
The largest thing on that server is a tenant database holding one real client’s data for a live app. 351MB gzipped. 468MB of BSON. It had never been inside a backup of any kind, on any machine, at any point. Roughly a thousand times more client data than everything that was actually being copied every night.
There was no bug. Each job backed up the thing it had been written to back up, on the day somebody wrote it, and every one of them succeeded every night for months. The question that had never been asked was what else lives here, and a script cannot answer a question nobody put to it. That is the whole failure, and it has no error message anywhere in it.
Three checks, all of which said it worked
Writing the job took an hour. Proving it was real took considerably longer, because the first three signals I got were wrong.
The exclude that matched nothing. The archive of /opt was meant to skip MeshCentral, which has its own encrypted backup and its own pull. I wrote the exclude with a leading slash, the way the path appears on the system. Under tar -C / opt the members are named opt/... with no leading slash, so the pattern silently matched nothing at all. No warning. The only symptom was the size: the archive came out at 421M instead of about 100M, carrying MeshCentral and two backup directories along with it. Exclude patterns are relative to the -C root, and tar will not tell you when one is dead.
The seven dumps that held nothing. The script found the databases to dump by listing them in mongosh and printing each id. In mongosh, printing an ObjectId gives you ObjectId('...') rather than the hex inside it, so the names it generated were things like company_ObjectId('6a35...'). Databases that do not exist.
Here is the part worth keeping. mongodump exits 0 on a database that does not exist. It writes a small stub file and returns success. Seven lines of output saying each database had been dumped, seven tiny files, exit code 0 throughout, and not one byte of anybody’s data. The script now rejects any archive under 200 bytes, which is a crude check that would have caught it instantly.
The dry run that proves nothing. Having got real archives, I wanted to know they would restore. mongorestore --dryRun reports “0 documents restored” for a perfectly good archive. It is not lying and it is not broken. It simply does not do what the name suggests, and if you read that line as a result you will conclude your backup is empty when it is fine. What genuinely walks the archive index is --dryRun -vv, grepped for found collection, which names every collection it can see.
Three tools. Two false failures and one false success, and the false success is the one that would have sat there for a year.
The credentials file that named a role
Before any of that, I spent a while convinced there was no way to administer MongoDB on the box at all.
The credentials file had a line labelled with the word root. Authenticating as a user called root failed. It kept failing. MongoDB returns Authentication failed for a wrong password and for a user that does not exist, with no way to tell the two apart from the outside, so a wrong username reads exactly like a lost password.
The label was the role, not the account. The admin user has the root role and is called something else entirely, and the password in that file had been correct the whole time. Nothing had ever been broken. I had written down that it was, which is the second time this month I have recorded a confident wrong conclusion before checking, and the fix both times was to go and read the actual state: admin.system.users lists who exists.
The key in that file is now named after the account rather than the role. Somewhere in the middle of misunderstanding it I also created a duplicate admin user, which was dropped the same day once the real one turned up.
What a copy inherits
One more, because it is the kind of thing that stays wrong quietly.
The backup directory on the server is mode 0750, owned by root with a dedicated group. The pull uses rsync -a, and the a includes preserving ownership. Ownership crosses the wire as a numeric id, not a name. The group that owns those files is GID 988 on the server. GID 988 on the machine at home belongs to something completely unrelated that ships with the desktop packages.
So the archives arrived and were readable by a group that had nothing to do with backups. It was latent rather than live, since that group has no members and is nobody’s primary group, but every nightly pull re-applied it and a package update could have added one. Now fixed, with --no-owner --no-group and an explicit --chmod, and the permissions read back correctly after a fresh pull.
The general rule is worth more than the incident: never let rsync -a set ownership across two machines that do not share a user database. The numbers match. The names do not. Nothing reports it.
There is a related point about what these archives contain. A full backup of a server is a copy of its secrets as well as its data, so the copy needs permissions at least as tight as the original, not the defaults the older backup directories on the same box are still sitting on.
What is still wrong
The pull is failure-silent. If the machine at home is off, the timer never fires, nothing runs, and no part of the system complains. I have added a warning when the newest night on disk is more than two days old, which helps, except that the warning lands in the journal of the box doing the checking. The machine that is down is the machine watching for the machine being down.
That needs to live somewhere that stays up, and it does not yet. Same open point as the last time I wrote about this, and I would rather say so than let a green tick on the server stand in for an answer.
Everything else verified. All eleven archives at home match the server’s own sha256 list byte for byte, every one passes gzip -t, and the pull moved 484MB in twelve seconds.
The reason any of this matters to somebody paying monthly is narrow and specific. Backups are not a product you buy once. They are a question somebody has to keep asking about a machine that keeps changing, and the answer was wrong here for months on a server I administer myself, with backups running the whole time. That is what the care plan is: somebody opening the archive and looking inside it, on a day when nothing is on fire. I have since written out every line of what that plan covers, including the four things it does not. The same instinct as making every script back up what it touches before it changes anything.
