Everything I build works on the day I build it. That’s the day I’m looking at it. I care more about what it does on day 180, when I’m not.

I went back through three of my own systems with that question. The answer came out the same every time. The layers that fail loudly stayed correct. Every layer that fails silently drifted. None of them degraded from a design flaw. They degraded because nothing complained.

A rename that finished on day zero and never finished after that

I moved a runtime root from /opt/openclaw to /opt/hog on 2026-02-07. One commit, 109 files, a 65-line migration script. The service came up against the new path. By every check I had, the migration was done.

Seven months later the repo still references the old path in 6 files. One of them is the runbook I’d open during an incident. README.md:10 names /opt/openclaw as the runtime location. docs/runbook.md gets it wrong ten times, including the log paths I’d tail at 2am. Those paths don’t exist anymore.

That pattern is exact. The runtime path had a test, since the service either started or it didn’t. The repo tree had a test too, which was the build. The documentation had nothing, so the documentation is the layer that’s wrong. I wasn’t careless. There was just nothing in place that could notice.

Four lines of CI would have caught it. Grep for the retired path, fail on a hit. I wrote the 65-line migration script and skipped the 4-line check, which is backwards. The script ran once. The check would have run every day since.

A routing tree I couldn’t review by reading it

I have 117 Prometheus alert rules. They fan through Alertmanager into an automation ladder and a human channel. I put continue: true on the automation route. I read that as “also try the next route.” It doesn’t mean that. A parent route’s own receiver only fires when no child matched. A child that matches and continues still suppresses the fallback.

42 of 117 rules went to automation and to no channel a human reads. That held for five weeks.

An accident saved it. My daily digest script queries ALERTS{alertstate="firing"} straight against Prometheus. It never consults Alertmanager. All 42 still showed up in the 07:30 summary. The failure cost me 24 hours of latency instead of total silence. Only because one tool declined to trust the layer above it.

Two things I took from that.

A config whose behavior I verify by reading it is not verified. I read that routing tree several times. Reading it is how I got it wrong, repeatedly, with confidence. amtool config routes test takes label sets instead of intentions and disagrees out loud. Seven label sets against the old and new file showed exactly one difference.

And redundancy that shares an assumption isn’t redundancy. Two receivers both fed by Alertmanager’s routing tree fail together. My digest survived because it reads a different source. Not because it was a second copy.

The system whose maintenance outran its use

In June I built my own coding agent. 24,800 lines of Go, a real tool that builds clean and works. By August its own ledger said this:

SignalMeasurement
Lifetime turns413
Share on a single dogfood day87%
Turns in the final 7 weeks5
Commits in the last 30 days22
Sessions on the tool it was replacing, same window180

Twenty-two commits against five turns. Development had fully decoupled from use. Everything left on the roadmap was re-implementing features the tool it replaced already had. So I archived it.

That’s a maintenance outcome and not a failure. It never cost money. It ran on subscriptions and local models, $0 lifetime. It cost attention, and attention is the budget maintenance actually spends. A system I keep building and stop using is charging me rent.

Archiving it taught me one more thing. Shutting down a repo doesn’t remove what it installed. Its own OAuth credential file was still in my home directory. Refresh tokens for two paid accounts, with no consumer. Decommissioning is a checklist, not a git archive.

What I do differently now

  • Every silent layer gets one loud check. Docs get a grep in CI. Routing gets a test command. Backups get an alert on staleness, not a green checkmark on the last run.
  • Verify behavior with a tool that disagrees. Anything I can only confirm by re-reading my own config is unverified.
  • Prefer redundancy across sources, not across copies. The second path should make a different assumption than the first. Otherwise it dies in the same failure.
  • Measure use, not activity. Commits are activity. Turns, sessions, and requests are use. When the first number climbs and the second doesn’t, delete the thing.

I’m not arguing for building less. I’m arguing for picking which layers I’m willing to have wrong for seven months. At build time, on purpose. In practice I’m making that choice anyway, whether or not I write it down.

One more. The digest that saved my alert routing has no watcher above it. Single point of failure. I know what the fix costs and I haven’t spent it yet. Every system has one of these. The dangerous one is the one you haven’t named.