193,000 failed logins, and the three things I'd left open
- linux
- security
- ssh
- operations
Let me start with the boring part, because it matters. Nothing got in. I went looking for a compromise in August and there wasn’t one: no unexpected accounts with root privileges, no listeners I couldn’t account for, no cron jobs anyone had added, no miner quietly eating the CPU, nothing odd going out. The server was fine.
What wasn’t fine was how hard it had been trying not to be.
What a month of logs actually looks like
The box is one I administer at Alpha IT Solutions. It carries production sites and a few internal services, and by any measure it is not an interesting target: no press coverage, no bug bounty, nothing that would put it on anybody’s list. (It has since given me a second, quieter problem: it was three kernels behind while reporting itself fully patched.)
Here is one month of SSH authentication logs on it.
| Week to | Failed passwords | Invalid usernames | Aimed at root |
|---|---|---|---|
| 26 Jul | 59,628 | 41,416 | 16,682 |
| 1 Aug | 60,928 | 24,720 | 34,347 |
| 9 Aug | 66,760 | 40,825 | 22,080 |
| 15 Aug | 6,165 | 5,934 | 3,565 |
| 19 Aug | 0 | 2,461 | 365 |
That’s roughly 193,000 failed password attempts in a month, against a server nobody has any particular reason to care about. It isn’t personal. It’s automated, it’s continuous, and it starts within hours of an IP address going live.
The usernames tell you who’s knocking. Top of the list across the whole period: admin 8,085 times, ubuntu 6,593, user 4,414, then wallet at 3,323, which is a crypto bot hunting for a machine with a hot wallet on it. Then debian, test, deploy, guest, oracle.
And 1,115 attempts on alphaitsolutions.
That one stopped me. Everything above it is a dictionary. That one is a username derived from a domain name, which means at least one of these things had looked at what the server hosts and made a guess. It’s still automated. It’s just less blind than I’d assumed.
The three things I’d left open
The logs were never the problem. The problem was what they’d have found if they’d got lucky.
Password authentication was on, and so was root login. Every one of those 193,000 attempts was a real attempt at a real door. The odds of guessing a strong password are nil, but “the odds are nil” is a probability argument, not a security control, and it only takes one reused credential to stop being an argument at all.
MongoDB had no authentication. It wasn’t reachable from the internet, so a scan wouldn’t have found it, and that is the entire reason it survived. But “not exposed” is one firewall rule away from “exposed”, and I’d been treating a network control as if it were an access control. Any process on that box, or anyone who got a foothold through any of the applications, could read every database on it.
One application port was published straight to the internet. The app also sat behind nginx, which is how it was actually meant to be reached, so the direct port was doing nothing except offering a second way in that bypassed everything nginx does. I’d opened it during setup and never closed it.
None of those is exotic. That’s rather the point. Nobody gets breached by something clever when there’s a default sitting there.
What changed, and what it did
9 August. Closed the application port, turned off rpcbind, tightened file permissions on the environment files, and put fail2ban on SSH.
Failed passwords dropped from 66,760 that week to 6,165 the next. That’s fail2ban doing exactly what it says: five wrong guesses and the address stops being able to talk to port 22 for an hour.
13 August. Disabled password authentication entirely. Keys only, root included.
Failed passwords went to zero, and have stayed there.
Not “fewer”. Zero, because there is no longer a password to fail at. The bots are still coming, and you can see them in the same logs: 2,461 invalid-username attempts in the last three days alone. They knock, sshd tells them there is no password authentication available, and they leave. The traffic didn’t stop. It stopped mattering.
Over the same month there were 2,336 successful logins. Every single one was a public key. Not one password login, because there is no longer any such thing.
16 August. Turned on MongoDB authentication, gave each application its own user with rights to only its own database, and verified from an anonymous shell that it now refuses to list anything.
That last one had been on my list twice before, and I’d talked myself out of it both times, on the grounds that changing a database’s auth model on a box running live applications was the sort of thing that breaks a Sunday. It took about forty minutes in the end.
The bit I’d rather skip
On 10 August, fail2ban banned me.
I was mid-session, port 22 stopped answering, and every website on the box carried on serving perfectly, which is the specific combination that makes you think the network is broken rather than that you have been locked out of a server you administer. It was my home IP, sitting in the jail alongside the actual attackers, because a client on my machine had been retrying a stale connection often enough to trip the same rule.
I waited out the hour rather than unbanning myself. Partly because it proved the thing worked. Mostly because the alternative was reaching for the fix that says “add my address to the ignore list”, and a permanent exception for a dynamic home IP is how you end up with a rule that protects nothing.
What I actually took from it
Nobody had to get in for this to have been worth doing. That’s the part I keep coming back to. If I’d read those logs as “no breach, therefore fine”, all three of those doors would still be open, and the evidence supporting that conclusion would have been exactly as strong as it was.
The fix that mattered took four minutes. One directive in sshd_config, one reload, and a third of a million knocks a year became noise. The work isn’t hard. Deciding to do it before something forces you is the hard part, and I didn’t manage that either.
The same sweep turned up a quieter class of problem a month later: headers that were set, loaded, and thrown away by nginx before they reached the page.
Anything I write that changes a server now backs up what it’s about to touch and can be run twice safely. That rule predates all of this and it’s the only reason I was willing to make these changes on a live box on a Wednesday afternoon rather than scheduling them for a weekend that never comes.
If you look after a server, go and count the failed logins on it. Not because you’ll find an intruder. Because the number is going to be larger than you think, and then you’ll go and look at what’s behind the door.
