Monitoring and observability
Monitoring tells you when a known condition occurs. Observability helps you investigate service behaviour, including situations you did not anticipate in advance.
Learn
A few useful ideas to help teams get more value from their signals and make reliability work easier to discuss.
Field notes
Use these as a starting point, then adapt the ideas to the services and teams you support.
Monitoring tells you when a known condition occurs. Observability helps you investigate service behaviour, including situations you did not anticipate in advance.
Metrics, logs, and traces answer different questions. Connecting them to a service and its owners makes it easier to move from an alert to useful investigation.
A useful alert points to a service condition that needs attention and gives the responder a reasonable next step. Regular review helps keep alerts relevant.
A service level objective makes a reliability expectation explicit. It can help teams balance changes, operational work, and the experience their service provides.
Put the ideas into practice