Observability
Observability
How to instrument, collect, and act on telemetry across distributed systems.
Topics
Instrumentation
- OpenTelemetry setup for .NET on Azure
- Structured logging conventions
Metrics and SLOs
- Defining SLIs for a web API
- Error budget policies that teams actually follow
Alerting
- Designing SLO-based alerting
- Alert routing and on-call hygiene
- Reducing alert noise with composite signals
Dashboards and triage
- Golden signals dashboard template
- Incident triage workflow with distributed traces
Cost of observability
- Sampling strategies for high-volume traces
- Log retention tiers that balance cost and compliance
Last updated on • Steve Rackham