Delivery and observability

Grafana

The dashboards and alerts that turn telemetry into something an on-call engineer can act on at four in the morning.

What we reach for first

When we reach for it

Any system with an on-call rota, and every performance engagement where a before and after has to be shown. A dashboard the client’s own team reads without us is part of what we hand over, not a nicety at the end of the engagement.

When we would argue against it

Dashboards nobody has agreed to look at. An unread dashboard is not observability, it is decoration with a maintenance cost. Every panel should answer a question somebody actually asks and every alert should be actionable; the rest trains people to ignore the screen.

What it looks like in delivery

Alerts tied to user-visible symptoms rather than to resource metrics, so a page means something is actually wrong. The dashboard is built around the questions asked during an incident rather than the metrics that happened to be easy to collect, and every alert links to what to do about it.

Where this appears on the site

Published work that names it

Derived from the stacks published on those pages, not written here — so this list cannot claim something the page it points at does not say.

Working in Grafana?

Tell us what it is running, what it costs you today, and what you need it to do next. A senior engineer will tell you what we would keep and what we would change.