Why tool consolidation fails at the migration, not the selection
Consolidating several monitoring tools onto one platform is the right instinct, and most large organisations are right to want it. Most consolidations fail anyway. They fail in the migration, and almost never in the selection.
The selection gets all the attention. There are demos, scorecards, a shortlist, a decision. It feels like the hard part, so it gets the effort. Then the migration is handed a decommission date and treated as a formality, and that is where it comes apart.
The failure pattern
It goes like this. The new platform is chosen. The old tools are given a date to be switched off. In between there is a period where the old tools are half switched off and the new one is half configured, and during that period the estate can see less than it could before either project started. Alerts that lived in an old tool go missing because nobody rebuilt them yet. A team that trusted its old dashboard loses it, does not trust the new one yet, and quietly keeps the old tool running so it is not flying blind. Multiply that across a dozen teams and, two years later, there are five tools where there were four. The consolidation added a tool.
None of that is a fault of the platform that was chosen. It is a fault of treating a migration as a switch to be flipped rather than a sequence to be run.
The alternative: handover by service
A consolidation that works is planned as a series of coverage handovers, one service at a time. For each service, the old tool and the new platform run side by side. The receiving team keeps using what it trusts while the new platform is configured to show the same thing and more. Only when that team confirms it can see what it used to see, and better, is the old tool switched off for that service. Then the next one.
It is slower to start, because the first handover has to establish the pattern, and much faster to finish, because nothing is ever running blind and no team is ever forced to choose between an unfamiliar new tool and no tool at all. The unofficial old tools never appear, because nobody had a reason to keep one.
Context does not migrate on its own
The part that surprises people is that the topology and the business meaning in the old tools rarely carry across by themselves. The dashboards, the alert thresholds tuned over years, the knowledge of which signal means what, all of that is context, and it has to be rebuilt deliberately on the new platform. A migration plan that assumes context comes for free ends up with a technically complete platform that no team can actually use, which is another way to end up keeping the old tool.
Make a past incident the acceptance test
The cleanest test of whether a handover is really done is a real incident from the past. Take one that the old estate diagnosed, and check that the new platform finds its root cause faster. If it does, that service is ready to hand over. If it does not, the migration for that service is not finished, no matter what the plan says. It is an honest test because it measures the thing the platform is for, getting from symptom to cause, rather than whether a box has been ticked.
A Time to Cause Review is often where this starts, because the past incidents are already the evidence, and the consolidation plan falls out of them.
Why do observability tool consolidations fail?
Because the migration is treated as a switch rather than a sequence. The new platform is chosen well, the old tools get a decommission date, and in the gap teams lose visibility they relied on and quietly keep the old tool running. The estate ends up with more tools than it started with, not fewer.
What is the right acceptance test for a consolidation?
A real past incident. Check that the new platform finds its root cause faster than the old estate did before the old estate is switched off. If it cannot, the migration is not finished, whatever the project plan says.