Storage SLAs and Uptime Guarantees: What to Actually Contract For
Every enterprise storage vendor leads with a big availability number. Pure, NetApp and Dell EMC all promise a level of uptime that sounds like your data can never be unreachable. The headline is rarely the thing that protects you. What protects you is buried in the exclusions and the remedy, and that is exactly the part the datasheet does not show. Here is how storage availability guarantees really work, and what is worth negotiating before you sign.
If you are comparing uptime SLAs across Pure Storage, NetApp and Dell EMC for a mission critical workload, you will find the headline numbers are close enough to be useless as a decider. One says six nines, another says a data availability guarantee, a third wraps it in a subscription programme with its own name. They are marketing figures first and contractual commitments second, and the gap between those two things is where the risk lives. The right question is not whose number is bigger. It is what each vendor will actually stand behind in writing, what they exclude, and what they owe you when it slips.
C4C is an independent, vendor neutral technology consultancy that helps enterprises hold storage vendors to what they promise. We translate the marketing availability number into what is actually guaranteed in the contract, and close the gap before you sign, drawing on years spent on the vendor side writing exactly these commitments. We know which nines are real, which are conditional, and which questions force an honest answer. No array of our own to sell.
What the nines actually mean
Availability is expressed as a percentage of time the system is reachable, and each extra nine is an order of magnitude less downtime. The numbers are smaller than they feel. Five nines, 99.999 percent, allows about five minutes of downtime a year. Six nines, 99.9999 percent, allows about 32 seconds a year. Those are tiny windows, and a vendor promising them is making a serious claim. The catch is that the claim is almost always measured and bounded in ways the headline does not mention.
Three things quietly shape what the number is worth. First, what counts as downtime, because planned maintenance, degraded performance and partial outages are often excluded from the measurement. Second, what has to be true for the guarantee to apply, because it usually assumes a specific supported configuration, a maintenance contract in good standing, and a redundant deployment you have to buy and run correctly. Third, who measures it, because a figure the vendor calculates from its own telemetry is not the same as one you can independently verify. None of that makes the number dishonest. It makes it conditional, and the conditions are the actual product.
Availability, durability and data loss are three different promises
The most expensive confusion in this whole area is treating one guarantee as if it covers the others. It does not, and vendors are precise about the distinction even when buyers are not.
Availability is whether you can reach your data right now. Durability is whether the data still exists and is intact over time, usually expressed as a probability that an object survives a given year. Data loss is governed by your recovery point and recovery time objectives, the RPO and RTO, which describe how much data you could lose and how long recovery takes after a failure. A platform can be highly available and still lose data in a disaster. It can be extremely durable and still be unreachable during an outage. A high availability array does not remove the need for backup, replication and a tested recovery plan, and any read of an SLA that assumes it does is a read that will cost you when something breaks.
How storage availability guarantees are really structured
Under the branding, the major vendors offer variations on a few mechanisms, and it helps to see them for what they are rather than by their programme names.
- A headline availability figure, the nines, offered against a specific supported configuration. It is real, but it is a floor with conditions, not a blanket promise across whatever you happen to deploy.
- A data availability or anytime availability guarantee, often part of a subscription or evergreen style programme, which commits the vendor to keep the system reachable and sometimes to non disruptive upgrades. These are genuine and valuable, and they are also the part most tightly bound to keeping the maintenance and subscription current.
- A credit remedy, which is what you receive when the guarantee is missed. This is the part that matters most and gets read least.
The point of naming them plainly is that the three big platforms are more alike here than the marketing suggests. Once you strip the programme names away, you are comparing configurations, exclusions and remedies, not slogans.
Why the credit is the part that matters
When a storage SLA is missed, the standard remedy is a service credit, usually a percentage of the fees for the affected period. That sounds fair until you weigh it against what an outage actually costs. If a mission critical platform is down and the trading floor, the warehouse or the clinical system stops, the business impact runs to figures that a few percent of a monthly storage charge does not begin to touch. The credit is a refund on the storage, not compensation for the outage, and it is almost always capped.
This is the single most important thing to understand about storage guarantees. The headline nines tell you what the vendor is aiming for. The remedy tells you what they are actually willing to be accountable for, and for most standard contracts the answer is very little in proportion to your exposure. That is not a scandal, it is how the industry prices risk, but you should go in knowing it and decide deliberately whether it is enough.
The exclusions do most of the work
Read any storage SLA to the end and the exclusions are longer than the guarantee. Common ones are entirely reasonable and worth knowing anyway: downtime during scheduled maintenance windows, issues caused by your own environment such as power, cooling or network, configurations outside the supported list, and lapses in the maintenance contract. Others are worth pushing on, such as whether degraded performance counts as an outage, whether a single component failure in a redundant system is excluded, and how a partial outage affecting some volumes is treated. The guarantee you think you are buying and the guarantee after exclusions can be meaningfully different platforms, and the difference is only visible if you read for it.
What to actually negotiate for a mission critical workload
For a genuinely mission critical system, the availability percentage is table stakes and the terms around it are where the value is. The things worth spending your negotiating capital on:
- A meaningful remedy, not a token credit. Push for credits that scale with severity and duration, and understand the cap. If the cap is trivial next to your exposure, the guarantee is largely decorative and you should price the residual risk yourself.
- A clear, tested definition of downtime. Get degraded performance and partial outages addressed explicitly rather than left to interpretation after an incident.
- Recovery commitments, not just uptime. Availability says nothing about how fast you recover from a real failure. Tie the conversation to your RPO and RTO and the replication or backup that delivers them.
- Response and escalation times. For a critical platform, how fast a senior engineer is engaged during a severity one incident often matters more day to day than the last nine.
- Non disruptive upgrade commitments in writing. If the programme promises upgrades without downtime, make sure that promise is contractual and its exclusions are acceptable.
How to compare Pure, NetApp and Dell EMC fairly
All three are credible platforms with mature availability engineering, and none of them is the loser here. To compare them honestly, normalise the terms before you weigh them. Put each guarantee on the same page: the headline figure, the exact configuration it assumes, the full exclusion list, the remedy and its cap, and the recovery commitments. Compared like for like, the decision usually turns not on whose number is highest but on whose exclusions and remedy fit your risk, and on which vendor will move the boilerplate terms during negotiation. That last point is where an independent read pays for itself, because knowing what each vendor can concede, and has conceded before, is not something a datasheet will ever tell you.
How C4C helps
We read these guarantees for a living, from the side that used to write them. We normalise the SLAs across Pure, NetApp, Dell EMC and the rest onto a single comparison, expose the exclusions and the real value of the remedy, and separate the availability promise from the durability and recovery ones so you are not buying one and assuming the others. Then we help you negotiate the terms that actually protect the workload, with a clear view of what is genuinely movable. Independent, vendor neutral, and with no array of our own to sell. Where a renewal or refresh is driving the decision, our storage refresh guide and independent decision guide are the natural next reads.
Weighing a storage SLA for a critical workload?
Send us the guarantees you are comparing, or the terms on the table, and we will give you an independent read: what each one really promises once the exclusions are in, what the remedy is worth against your exposure, and what is worth negotiating before you sign. Independent, with no array of our own to push. We spent years on the vendor side writing exactly these commitments.
Prefer email? Reach us directly at hello@c4cgroup.co.uk.
Frequently asked questions
What uptime SLA do enterprise storage vendors offer?
The major platforms typically headline an availability guarantee in the range of five to six nines, meaning 99.999 to 99.9999 percent, which allows roughly five minutes down to about 32 seconds of downtime a year. Pure, NetApp and Dell EMC each wrap this in a named programme, often tied to a subscription or maintenance agreement. The headline figure is real but conditional: it assumes a specific supported and usually redundant configuration, a maintenance contract in good standing, and a measurement that excludes planned maintenance and various fault categories. Treat it as a starting floor with conditions, not a blanket promise.
What is the difference between availability and durability?
They are two different promises and one does not imply the other. Availability is whether you can reach your data right now, expressed as an uptime percentage. Durability is whether the data still exists and is intact over time, usually expressed as the probability that data survives a given year. A platform can be highly available and still lose data in a disaster, and it can be extremely durable and still be unreachable during an outage. Neither one replaces backup, replication and a tested recovery plan, which are governed separately by your recovery point and recovery time objectives.
Are storage availability guarantees actually worth anything?
They are worth something, but usually far less than the headline suggests, because the remedy is the real measure of accountability. When an SLA is missed the standard remedy is a service credit, typically a capped percentage of the fees for the affected period. Weighed against the business cost of a mission critical outage, that credit is a refund on the storage rather than compensation for the disruption. The guarantee signals what the vendor is aiming for, the remedy shows what they will actually be accountable for, and for most standard contracts the second number is small in proportion to your exposure. Read the remedy and the cap before you trust the nines.
Why do the exclusions matter more than the headline number?
Because the exclusions decide when the guarantee actually applies, and they are usually longer than the guarantee itself. Common exclusions cover scheduled maintenance, problems caused by your own power, cooling or network, unsupported configurations and lapsed maintenance contracts. Others are more consequential and worth challenging, such as whether degraded performance counts as an outage, how a partial outage affecting only some volumes is treated, and whether a single component failure in a redundant system is excluded. The guarantee you think you bought and the guarantee after exclusions can be meaningfully different, and the difference is only visible if you read for it.
What should I negotiate into a storage SLA?
For a mission critical workload, spend your negotiating capital on the terms rather than the last nine. Push for a remedy that scales with severity and duration and understand its cap. Get a clear, written definition of downtime that addresses degraded performance and partial outages. Tie the conversation to recovery, not just uptime, by anchoring on your recovery point and recovery time objectives. Pin down response and escalation times for a severity one incident, since how fast a senior engineer engages often matters more day to day than the headline figure. And if non disruptive upgrades are promised, make that promise contractual with acceptable exclusions.
How should I compare uptime SLAs across Pure, NetApp and Dell EMC?
Normalise the terms before you weigh them, because all three are credible and none is the obvious loser. Put each guarantee on the same page: the headline figure, the exact configuration it assumes, the full exclusion list, the remedy and its cap, and the recovery commitments. Compared like for like, the decision usually turns not on whose number is highest but on whose exclusions and remedy fit your risk, and on which vendor will actually move the boilerplate during negotiation. Knowing what each vendor can concede is where an independent view earns its keep, because a datasheet will never tell you.