Plain, vendor neutral writing from the practice. What goes wrong in large observability estates, and how a coverage first, context aware, cause driven approach fixes it.
Coverage decisions get made by default, not design. The fix is to start from the application and the business, and to place every host and service in a tier on purpose.
Read the article →Log strategy is a list of questions, not a list of sources. The three questions to ask of every log before it is ingested, kept and linked.
Read the article →An agent can be up, fast and wrong. What quality signals, tail latency, tool call loops and cost attribution mean for observing AI workloads.
Read the article →Selection gets the attention and the migration gets a decommission date. The failure pattern, and the handover by service that avoids it.
Read the article →An hour with the practice, no slides, no obligation. Bring your incident history and your tooling list. Leave with a plain view of where your observability is working, where it is not, and what to do first.
Book an observability sense check