Skip to content
RO
← All writing

A handshake from a machine that was gone

  • linux
  • networking
  • wireguard
  • operations

On Friday 5 September I unplugged a server from one switch and plugged it into another. Ten seconds of work. It took the machine off the network for about an hour, and in that hour three separate tools told me convincing lies about why.

The box is server2, an old Acer laptop running headless Debian 13 that acts as my Docker host at home. It had been sitting behind a mesh node on 10.110.1.36/24, with 10.110.1.19 as its gateway. I moved it onto the same switch as my desktop, which is the ISP router’s own segment, 192.168.1.0/24.

Its address is static. So it came back up holding an address for a subnet that no longer existed, pointed at a gateway that was no longer reachable. No route off the LAN. No internet. And because it dials out to a WireGuard hub on a VPS rather than accepting anything inbound, the tunnel died with the route.

That is what made it awkward. The LAN address was wrong, the tunnel was down, and SSH failed from every direction at once. There is no meaningful difference between that and a machine somebody has switched off.

The handshake that was stuck

Here is the part worth keeping.

I checked from the hub, which is normally the fastest honest answer, because the VPS can see all three peers and does not care what my desktop thinks. wg show reported server2’s latest handshake as 6 minutes ago, against endpoint 149.34.140.80:33014. That is my real home public address and a perfectly plausible age for a peer with PersistentKeepalive = 25.

The machine was completely off the network at that moment. It had no default route and no path to the internet at all.

latest handshake is a stored timestamp, not a liveness check. WireGuard records the last successful handshake and then leaves the value where it is. Nothing decrements it, nothing ages it out into an obvious “gone” state. A peer that drops out three minutes into its handshake interval keeps reporting a handshake from three minutes ago for as long as you keep asking, and on a quiet link that can look fresher than a healthy peer that simply has not needed to talk recently.

The test that works is to make it move. Ping the peer from the hub, then read the raw integer back:

ssh vps "wg show wg0 latest-handshakes"

Use latest-handshakes rather than the pretty wg show output. It gives you epoch seconds, which you can compare between two readings. If the number does not advance after you have sent traffic at the peer, the peer is not there, whatever the human-readable line says.

There is a local wrinkle that made me slower than I should have been. server2 drops ICMP on its LAN side, so Test-NetConnection from Windows reports PingSucceeded: False and TcpTestSucceeded: True, and that is the normal state of things rather than a fault. Pinging a box is the reflex that would have caught this in five seconds, and I had trained myself out of it on this particular box for a perfectly good reason.

The gateway that failed silently

Fixing the LAN should have been one nmcli command. I typed the gateway as 198.168.1.1.

NetworkManager accepted it. It did not object that the gateway sits outside the interface’s own subnet. It installed an onlink route for it and carried on, and nmcli con show reported the profile exactly as I had asked for it. Everything on the local subnet kept working, DNS included, because 192.168.1.1 is directly reachable and needs no gateway to get to. Only off-subnet traffic broke.

So the symptom was: LAN fine, DNS fine, no internet. Which looks nothing like “you typed one digit wrong ninety seconds ago”, and I spent a few minutes suspecting the router before I read the route table properly.

Read ip route and look at the actual octets of the default. nmcli connection show tells you what is stored. Only the kernel tells you what is installed.

Correcting it was not clean either. After nmcli con mod and an ip route replace, the kernel held two defaults: the good one at metric 0 and the bad one still sitting behind it at metric 100. nmcli device reapply enp1s0 cleared it out.

Stored is not applied

Same family of mistake, one level up. nmcli connection modify writes the profile and does not touch the running interface. In between, nmcli connection show says manual while ip route still says proto dhcp. It looks finished and it is not, until an nmcli connection up or a reboot.

I knew that. I had written it down in August, on this same machine, after it bit me the first time. Knowing it did not stop me reaching for the wrong command again, which is roughly the argument for writing things down where somebody else can read them rather than trusting that the person who learned it will be available.

What the move bought, and what it broke

The gain is real. Both machines are on 192.168.1.0/24 now, so the LAN path works whenever the box is up, and it went from behind two NATs to one. Nothing is forwarded to it, so it is still dark from outside.

Not one line of WireGuard config changed. The peer dials out to 185.230.219.110:51820 and the 10.8.0.x addresses have nothing to do with which LAN the box sits on. Fix the LAN and the tunnel comes back on its own inside twenty-five seconds. That is the design working, and it is easy to miss while you are busy assuming the tunnel is the thing that is broken.

One casualty survives. The CUPS queue office-hp still points at ipp://10.110.1.96/ipp/print, and that printer is behind the mesh node on the old subnet. Verified unreachable from server2 the same afternoon. Either the printer moves or the queue stays broken, and I have left it broken, which tells you something about how much I print.

There is also a quieter consequence I only thought about afterwards. server2 is the machine that pulls my MeshCentral backups off the VPS over that tunnel. While it was down, the off-box copies silently stopped advancing, and the backups on the VPS carried on succeeding every night, so nothing anywhere looked wrong. That is the same shape of problem I wrote about in August, arriving from a completely different direction.

Two places now define one address

I asked the router to reserve 192.168.1.36 for the box’s MAC, and left the static profile on the box as well. Both on purpose: the reservation holds the address out of the DHCP pool, the static profile is what actually puts it on the interface.

That means two systems now define the same address, and changing one without the other makes the next reboot interesting. The router is a Technicolor gateway with a web interface and nothing else. No SSH, no SNMP, no API. So the reservation can only be verified by logging into it and looking at a page, by a human, which is the part of this I like least and the reason it is written down instead of remembered.

The pattern

Nothing here was difficult. Every step reported success.

nmcli said the profile was configured. wg show said the peer had handshaked six minutes ago. Both were accurately reporting stored state, and neither was reporting reality. ip route was the only tool in the hour that told me something true, and only because I eventually read the numbers instead of the shape.

If you have a Linux box somewhere that works, that nobody has looked at properly, and whose current state nobody could describe from memory, that is the job I do. The interesting part is almost never the fix. It is finding out which of your instruments is lying.

I have since written up what the first day on an inherited server actually involves, which starts from the same problem: the file on disk and the running process are two different things.