Skip to content

All writing

Infrastructure8 min read

Monitoring that lives somewhere else

Ten scrape jobs, twenty-nine alert rules, and a watchdog that can reach me without touching the stack it is watching. The point is the address, not the tooling.

The failure that teaches it

Every monitoring setup is fine until the failure is big enough to be interesting. A container dies, the alert fires, you get a message, everyone is happy. Then the whole machine goes away, and with it the thing that was supposed to tell you the machine went away.

The silence that follows looks exactly like a quiet night. That is the failure mode worth designing against: not a missing metric, but a missing alarm that nobody notices is missing.

A box that is not the box

So the observability contour sits on a machine of its own. Prometheus, Alertmanager, Grafana and Loki live there and nowhere else, which costs one more machine and buys the only property that matters: the watcher and the watched can fail independently.

If the alarm rides on the thing it is watching, the loudest failure is the one you never hear about.

Ten jobs, twenty-nine rules

Ten scrape jobs collect from the application, the database, the host and the edge. Twenty-nine alert rules turn that into statements a person can act on at four in the morning, and every one of them had to justify its own existence, because an alert that fires when nothing is wrong teaches the team to ignore alerts that fire when something is.

Logs go to Loki on the same machine. Not because logs are metrics, but because when you are woken up you want one address, not two.

The watchdog that skips the stack

A separate cross-machine watchdog reaches Telegram directly, without passing through the application or its queue. It exists so that the message that says everything is down does not need everything to be up.

It is the smallest piece of the contour and the one I would rebuild first. Deliberately dumb: it checks, it sends, it has no opinions and no dependencies worth failing.

What I would build first

Put the monitoring somewhere else, describe it with the same automation as everything else so it can be rebuilt rather than nursed, and give the last-resort alert its own path out of the building. Then test that path on a normal Tuesday, because the first time you use it should not be the first time you try it.