Green Is Not Working
A dashboard is a claim about a system. Claims drift. Here is how green has lied in every monitoring generation since ICMP.
There is a specific kind of phone call I have taken for twenty-five years.
The customer is down. Hard down. And somewhere, on a wall or a laptop, a dashboard is green.
Not broken. Not stale. Green. Confidently, accurately green – accurate about the thing it was measuring, which turned out not to be the thing that mattered.
Green has always meant something narrower than you think
Every monitoring generation picked a proxy for “working”, and every generation the proxy drifted away from reality.
ICMP. We pinged things. A reply meant the network stack was alive. It said nothing about whether the application had deadlocked ten minutes ago. The box was up. The business was down. Both statements true.
SNMP. We graphed CPU, memory, interface counters. Beautiful graphs. All within threshold, while the database sat behind a lock nobody was polling for. We measured what the device offered, not what the service needed.
Uptime checks. HTTP 200 from a robot in another country. The homepage returned 200 all night. The checkout returned 200 too – with a page that said “we are unable to process your order”. Two hundred is a status code, not a business outcome.
Synthetics. Better. Real transactions, real journeys. Then the script drifted from what users do, or ran from a network path no customer uses, and passed for a month while the mobile app failed for everyone.
Dashboards and SOCs. Now we aggregate all of it into one pane, which means a single green tile is standing in for hundreds of assumptions – and nobody alive can enumerate them.
See the shape? Each generation measured more, from further away, with more confidence. The proxy got prettier. The gap between “green” and “working” never closed.
The failure that has no dashboard at all
Now the part that should genuinely bother you, because it is worse than a bad proxy.
I audited a perimeter firewall this year. Serious box, serious network. Its local counters were extraordinary: over 1.4 billion network events retained, millions of VPN events, hundreds of thousands of user events. Every logging category ticked. Every screen agreed that logging was on.
It had zero syslog servers configured.
The box was logging. Diligently, correctly, for years – into a local ring buffer that overwrites itself. Nothing had ever left the device. Which means the perimeter history of that network, for any investigation, any breach, any dispute, was zero.
There was no dashboard for this. There was nothing to be green. The box was doing exactly what it had been told to do, and what it had been told to do was useless. The ticked category is not the export. The destination is.
Why this survives every technology cycle
Because green is a claim, and a claim is cheap to produce and expensive to verify.
The device asserts something about itself. The dashboard aggregates assertions. Nobody in the chain is incentivised to ask “and did that assertion ever correspond to anything?” – because asking is work, and the answer is usually yes, and the one time it is no, you find out from a customer.
Green also drifts silently. It does not decay into yellow. It stays green while the world moves underneath it: a collector is decommissioned in a migration, a threshold is copied from a box with different capacity, a check outlives the service it checks. Nothing alerts you that your alerting has stopped meaning anything.
What I actually trust
After twenty-five years, this is the short list. It has not changed much, which is the point.
- Measure the transaction, not the component. The customer does not buy your CPU. If the check does not do what a user does, it is decoration.
- Trust destinations, not sources. A source reporting on itself is not a witness. Go to where the output should have landed and look for it.
- Every check needs an expiry date. A check that has not failed in two years is not proof of stability – it is an untested claim. Break it on purpose and see if anyone hears.
- Alert on absence. The most dangerous failure is silence: no events, no heartbeat, no delivery. Most systems only alert on bad news arriving, never on good news stopping.
- Ask what the tile cannot see. For any green indicator: name three real failures that would leave it green. If you cannot, you do not understand it well enough to rely on it.
The uncomfortable question
Look at your monitoring. Now answer honestly: when did it last catch something before a human did?
If the answer is “I cannot remember”, you do not have monitoring. You have a mood ring that is stuck on calm – and it will stay calm right through the next outage, exactly as it was configured to.