Data Centre · Storage

Data Reduction and Effective Capacity: How to Compare and Contract It

Effective capacity is the biggest number on a storage quote and the least reliable. It depends entirely on how well your data reduces, and the assumption behind it is set by the vendor. Here is how to compare two platforms on the same honest terms, and how to make the vendor stand behind the ratio in the contract rather than on the datasheet, from people who built these claims from the vendor side.

Two all flash platforms quoted at the same effective capacity can be completely different deals, and the difference is hidden in an assumption neither quote spells out. Data reduction, the combined effect of deduplication and compression, is real and valuable, but the ratio a vendor applies to reach its headline capacity is an average drawn from the data that reduces best. Apply it to your data and the number can hold up beautifully or collapse by half. This guide is about doing that comparison honestly, and then closing the gap between the promise and the contract so you get the capacity you paid for.

Who we are

C4C is an independent, vendor neutral technology consultancy that helps enterprises compare storage on real terms and hold vendors to their claims. We cut through the data reduction and effective capacity marketing, model each platform on your actual data rather than the datasheet average, and get the ratio guaranteed in the contract so it is enforceable, drawing on years spent on the vendor side making exactly these claims. No array of our own to sell.

This is the deeper, comparison and contracting companion to our shorter primer. For the plain explanation of the three capacity numbers and how a quote slides between them, start with effective capacity claims: how to read a storage quote honestly and come back here for the vendor to vendor comparison and the contract mechanics.

The number to distrust, in one line

Every array has three capacity numbers. Raw is the physical total of the drives. Usable is what survives data protection, spares and overhead, and it is the honest floor you are guaranteed regardless of your data. Effective is usable multiplied by an assumed reduction ratio, and it is the figure printed largest because it is the biggest. The trap is simple: usable is a fact about the array, effective is a guess about your data. Compare platforms on the fact, then treat the guess as something to be tested and contracted, not believed.

Why data reduction ratios swing so wildly

Your ratio is a property of your data, not of the array. The same platform can deliver eight to one on one workload and barely better than one to one on another, and both are correct. What separates them is how much redundancy and compressibility the data actually contains.

  • Reduces poorly: anything already compressed or encrypted. Images, video, audio, and databases using transparent encryption arrive at the array with the redundancy already squeezed out, so they land close to one to one. If most of your estate is this kind of data, a high assumed ratio is fiction for you.
  • Reduces well: virtual machines and virtual desktops, where thousands of near identical operating system images sit side by side, general file shares, and verbose logs. These can meet or beat the assumed ratio comfortably.
  • Reduces variably: production databases without encryption, which depend heavily on the schema and how much repetition the tables carry.

This is why a quote built on a four to one assumption, applied to a workload that genuinely reduces at two to one, delivers half the capacity you were shown. Nobody lied. The average was simply not your average, and you find out when the array fills early and the top up quote arrives.

The special case of backup and archive

Backup and archive is where reduction claims get most interesting, and where the comparison between platforms genuinely diverges. Backup data is highly repetitive by nature, the same files captured again and again across successive backups, so deduplication has an enormous amount to work with and ratios can look spectacular. But two things complicate it. First, if your backup software already deduplicates and compresses before the data reaches the array, most of the reduction has happened upstream and the array has little left to find. Second, purpose built backup appliances, general purpose all flash platforms, and the newer unified platforms approach this workload differently, so a ratio that is true on one is not transferable to another.

This is exactly the ground on which platforms like Pure Storage FlashBlade and VAST Data are often compared for backup and archive. Both are capable scale out platforms with strong reduction on the right data. The honest answer is that neither has a universal ratio you can lift from a slide, because the outcome depends on whether your backup stream is already reduced, how much genuine duplication exists across your retention, and the mix of data types underneath. The right comparison is not which vendor claims the higher number, it is which platform delivers on your data, measured, with the result you were shown standing behind it in writing.

How to compare two platforms on the same terms

The mistake is comparing two effective figures against each other, because they rest on different assumptions and are not the same currency. Put both on a common basis instead.

  1. Normalise both quotes back to usable capacity, after protection and overhead. This is the fact you can compare directly.
  2. Apply one conservative reduction ratio based on your real data, the same ratio to both, drawn from what your workloads actually achieve today rather than from either vendor's assumption.
  3. Only now divide price by the genuinely usable terabytes each delivers. That is the real price per terabyte, and it is often nothing like the headline.
  4. Model the cost of being wrong. Ask what it costs to add capacity later if your ratio comes in low, at what unit price, because a cheap headline propped up by premium priced top ups is not a cheap deal, it is a deferred one.

Done this way, two arrays sold as identical on effective capacity routinely turn out to differ by a wide margin on the number that matters. The platform with the lower headline sometimes wins outright.

The guarantee programmes, and what they actually protect

Most vendors offer a data reduction guarantee, and on the surface it looks like it removes the risk. It rarely removes as much as it appears to, because it is written to defend the headline number rather than your budget. Read for these things before you rely on it.

  • The exclusions carry the risk. Pre compressed and encrypted data is almost always excluded from the guaranteed ratio, which carves out precisely the data most likely to underperform.
  • The remedy is usually more drives, not money. When the array misses the ratio, the typical fix is additional capacity shipped to you rather than a refund, so the vendor protects its story while you absorb the rack space and power.
  • You often have to enable every reduction feature to qualify, including ones you might otherwise turn off for performance.
  • Claiming is not automatic. It generally needs a support case and a validation exercise, on your time.

None of this makes the guarantee worthless. It makes it a thing to read carefully and negotiate, not a box to tick.

How to get the ratio into the contract

Here is the resolution, and it is the opposite of the fight most buyers pick. Do not try to force the vendor down onto raw or usable and refuse to discuss effective capacity. Effective capacity is a legitimate way to buy storage, and a vendor confident in its platform should be willing to stand behind its number. The honest move is to accept the effective figure and make it contractual and specific to you.

In practice that means a small number of things agreed in writing before you sign: the assumed ratio stated explicitly on the order, a guarantee that applies to your data profile rather than a generic exclusion list, a defined outcome if the array misses the ratio in production that actually compensates you rather than just topping up drives, and a locked unit price for expansion capacity so a low ratio cannot be turned into an expensive surprise later. Get those into the contract and the effective number stops being a marketing line and becomes a commitment. That shift, from datasheet to contract, is where the risk actually moves off your side of the table.

How C4C helps

This is squarely our ground. We spent years on the vendor side of the storage market, building exactly these effective capacity figures and knowing which assumptions they rest on, so we read a quote the way the person who wrote it does. We will normalise competing quotes back to the same honest basis, model the reduction you will realistically achieve on your own data rather than the vendor's average, and pin the ratio down in the contract so the number you were shown is the number you can rely on. Independent, with no array of our own to push, and no preferred logo. We work quietly behind your team or as your named advisor, whichever serves you better. If you want the wider commercial picture of how a storage deal is assembled, our buying enterprise storage guide covers it, and the all flash and NVMe economics guide goes deeper on where the real cost sits.

Comparing storage quotes on effective capacity?

Send us the quotes, or the platforms you are weighing, and tell us roughly what your data looks like. We will read them back to you on the same honest basis, model the reduction you will actually get, and show you where the contract needs to change before you commit. Independent, with no storage line of our own to push. We built these claims from the inside.

Prefer email? Reach us directly at hello@c4cgroup.co.uk.

Frequently asked questions

What is the difference between raw, usable and effective capacity?

Raw is the physical total of every drive at its rated size. Usable is what survives data protection, spare capacity and system overhead, and it is the honest floor you are guaranteed no matter what your data looks like. Effective is usable multiplied by an assumed data reduction ratio, the combined effect of deduplication and compression, and it is the number a quote leads with because it is the largest. Usable is a fact about the array, effective is a guess about your data, so compare platforms on usable and treat effective as something to test and contract.

Why do data reduction ratios vary so much between workloads?

Because the ratio is a property of your data, not of the array. Data that is already compressed or encrypted, such as images, video and databases using transparent encryption, arrives with the redundancy squeezed out and reduces close to one to one. Data with a lot of repetition, such as virtual machines, virtual desktops and verbose logs, reduces well and can beat the assumed ratio. The vendor applies an average drawn from the workloads that reduce best, so a four to one assumption on a workload that genuinely reduces at two to one delivers half the capacity you were shown.

How do I compare data reduction between vendors fairly?

Do not compare two effective figures against each other, because they rest on different assumptions and are not the same currency. Normalise both quotes back to usable capacity, apply one conservative reduction ratio based on your own real data to both, and only then compare the price per genuinely usable terabyte. Also model what it costs to add capacity later if your ratio comes in low, because a cheap headline propped up by premium priced top ups is a deferred cost, not a saving. Compared this way, two arrays sold as identical on effective capacity often differ by a wide margin.

How much does backup and archive data reduce?

Backup data is highly repetitive, the same files captured again and again across successive backups, so deduplication has a lot to work with and ratios can look very high. Two things complicate it. If your backup software already deduplicates and compresses before the data reaches the array, most of the reduction has happened upstream and the array has little left to find. And purpose built backup appliances, general purpose all flash platforms and newer unified platforms handle the workload differently, so a ratio that is true on one does not transfer to another. Measure it on your actual backup stream rather than trusting a headline.

Do Pure Storage FlashBlade and VAST Data reduce data differently?

Both are capable scale out platforms with strong data reduction on the right data, and they are often compared for backup and archive. Neither has a universal ratio you can lift from a slide, because the outcome depends on whether your backup stream is already reduced upstream, how much genuine duplication exists across your retention, and the mix of data types underneath. The right comparison is not which vendor claims the higher number, it is which platform delivers on your data, measured, with the result standing behind it in the contract. On your data, run both against the same sample before you decide.

Can I get a data reduction guarantee written into the contract?

Yes, and it is the right move, but read what most guarantees actually protect first. They tend to exclude pre compressed and encrypted data, which is the data most likely to underperform, and the usual remedy is additional drives rather than money, so the vendor defends its story while you absorb the rack space and power. The honest approach is to accept the effective figure and make it specific and contractual: the assumed ratio stated on the order, a guarantee that applies to your data profile, a defined outcome that genuinely compensates you if the array misses in production, and a locked unit price for expansion. That shift from datasheet to contract is where the risk moves off your side of the table.