Every observability estate is a set of decisions about what to see and how deeply. Most of those decisions were never made. They were defaults. Instrumentation design is the practice of making them on purpose.
What goes wrong
Agents get deployed to everything at the deepest setting because it was easy, and the estate produces data nobody reads at a cost nobody planned. Or agents get deployed to a handful of servers a project cared about, and the service that failed last Tuesday was never on the list. Gateways are placed where the network team allowed rather than where the traffic is. Container platforms get the same treatment as virtual machines and neither works well.
How C4C approaches it · Coverage
We start with the application, not the server list. Which services make money or carry risk, which are supporting cast, which are inventory. Each host, container and service is placed in one of three tiers: full depth, infrastructure only, or discovered. Then the physical layout: where gateways sit, which zones agents report through, what the security team needs to approve, and in what order it rolls out so the important services are visible first. The output is a design the operations team can read and the platform team can execute.
What you instrument and how deep. Full depth, infrastructure only, or discovered but not monitored, decided on purpose for every part of the estate.
The topology and the business meaning attached to the telemetry, so a technical signal arrives already attached to its consequence.
Getting from symptom to root cause in minutes rather than hours. One problem, one cause, the owner obvious. It only works when Coverage and Context are right.
What you get
Productised as the Instrumentation Blueprint, a fixed scope engagement that produces the plan a platform rollout should start from.
Questions
No. Full depth belongs on the services that carry business value or risk. Most of an estate needs basic health and nothing more. Instrumenting everything at full depth produces cost and noise without adding visibility where it matters.
A gateway is a relay inside the organisation’s network. Agents report to the gateway and the gateway reports out to the platform. It means the servers themselves do not need outbound internet access, which most enterprise security teams require.
Agents are deployed per node rather than per container, workloads appear and disappear in minutes, and the useful unit of observation is the service rather than the machine. Designs built for virtual machines tend to produce gaps and noise when applied to containers.
An hour with the practice, no slides, no obligation. Bring your incident history and your tooling list. Leave with a plain view of where your observability is working, where it is not, and what to do first.
Book an observability sense check