Is N+1 Still the Right Question? Rethinking Cooling Redundancy for the Liquid Cooling Era

nVent's Nick Onyett explains why the liquid cooling era demands that the industry evaluate redundancy based on what is inside the unit, not just how many units are in the row.

The Assumption Nobody Questions

 N+1 is one of the most expensive unexamined assumptions in data center cooling design. The concept is simple: deploy enough cooling units to handle the full heat load, then add one more as a backup. For decades, that math made perfect sense. Traditional cooling units, the CRAC and CRAH air handlers that have cooled data centers for decades, were built as largely monolithic systems, with limited internal redundancies and few components that could be serviced while the unit ran. Servicing a part or recovering from an internal failure generally meant taking the whole unit offline, so operators kept a spare ready to pick up the load.

The data center industry's tier classification system, a widely used framework that rates facilities from Tier 1 through Tier 4 based on levels of redundancy and uptime, codified this approach into a standard. And the standard became invisible. Nobody questions N+1 for cooling. It's just how it's done. The problem is that liquid cooling CDUs don't have to work this way, and the gap between what modern cooling units can actually do and what the legacy redundancy model assumes they can't, is costing operators real money.

What the Spare Is Actually For

The backup unit exists to serve two very different scenarios, and the industry rarely separates them. The first is routine maintenance. Pumps need service, filters need replacement, firmware needs updating. Taking a unit offline for planned work is inevitable, and the spare keeps cooling uninterrupted while that happens. This is the everyday reality of data center operations, and it's the scenario most operators are actually thinking about when they specify N+1.

The second scenario is catastrophic failure. A coolant leak floods the unit's internals. An electrical fault takes out the control board. A piece of equipment physically damages the housing. These are the events that tier classifications are ultimately designed to protect against, particularly at Tier 4, where the standard demands fully independent backup systems with complete physical isolation.

But at Tier 3, where most enterprise and hyperscale facilities operate, the requirement is concurrent maintainability: the ability to take any single component offline for service without interrupting operations. The question worth asking is how much of your redundancy investment is protecting against the maintenance scenario, and how much is protecting against the catastrophic one. Because those two risks have very different probabilities, and very different solutions. A coolant pump that needs servicing every eighteen months is a certainty. A forklift driving through a live data hall with tens of millions of dollars in equipment is not a cooling problem. It's a planning and security problem. Yet the conventional N+1 model treats both scenarios the same way: add another unit.

The Architecture Already Changed

The reason the +1 unit became standard is that traditional cooling units gave operators no alternative. If anything inside the box failed, the box went down. But the architecture of liquid cooling CDUs has moved well past that limitation, and the implications for redundancy planning are significant.

Consider what a modern CDU can look like. Early CDUs carried over the same single-string design: one pump, one power feed, sealed components that forced a full shutdown for service. A modern CDU breaks from that. Instead of a single pump, multiple pumps operate simultaneously. If one fails, the remaining pumps are already running and ramp up to compensate. There is no switchover delay, no activation sequence, no thermal excursion while a backup spins up. The cooling simply continues. Instead of a single power feed, multiple independent feeds ensure that losing one input doesn't interrupt operation. Instead of sealed components that require a full shutdown to access, critical parts are designed for hot-swap replacement during live operation.

This changes the maintenance equation entirely. The routine scenario described earlier, taking a unit offline to service a pump or swap a filter, no longer requires a full shutdown. Much of that work now happens while the unit runs. That is concurrent maintainability delivered at the component level, inside the unit, rather than at the unit level across the row. It's the outcome that Tier 3 demands, achieved through better engineering rather than additional hardware.

None of this eliminates the catastrophic scenario. A CDU with redundant pumps and redundant power feeds is still a single physical object that can be damaged or destroyed. But as noted earlier, that scenario belongs to a different risk category with a very different probability, and arguably a very different set of solutions.

The Cost of the Idle Unit

If the +1 unit is no longer solving the everyday maintenance problem, then its cost needs to be evaluated against the narrow set of scenarios it is actually protecting against. And at scale, that cost is not trivial.

Start with power. In an active-standby design, a backup CDU draws electricity around the clock whether it is doing productive cooling work or not. For a single unit, the draw is a rounding error. Across a hyperscale facility with hundreds of CDUs, the cumulative power consumption of idle backup units becomes a line item. Load-sharing designs reduce that penalty, though the redundant capacity still carries a cost. Every watt powering a unit that isn't actively cooling a rack is a watt that could be driving compute. In an industry that increasingly measures efficiency in cost per token, that's quantifiable waste with no productive return.

Then there is the physical footprint. Each backup unit occupies floor space, requires its own power distribution infrastructure, consumes coolant volume, and needs connection to the facility water loop. That footprint represents capacity that could be serving productive racks. For operators competing on density and speed to deployment, an idle unit sitting in the row waiting for an event that may never come is a real opportunity cost.

Finally, there is operational overhead. The backup unit doesn't maintain itself. It requires the same firmware updates, the same commissioning verification, the same periodic testing and inspection as every productive unit on the floor. Maintenance teams are servicing equipment whose entire purpose is to wait. Multiply that across a large facility and the labor and scheduling burden adds up, quietly increasing the operational cost per rack without adding a single watt of cooling capacity.

Asking a Better Question

The tier classification system is not broken. It has provided the data center industry with a shared language for reliability that has stood for decades, and its core principles around uptime and maintainability remain sound. But it was written for a generation of cooling technology that no longer represents the state of the art, and applying its assumptions without reexamination leads to infrastructure decisions that prioritize convention over engineering.

The operators who will deploy cooling most efficiently in the coming years are the ones asking a different set of questions. Not just how many backup units do I need, but how many single points of failure exist inside each unit. Not just can I take a unit offline for service, but can I service individual components while the unit keeps running. Not just what tier am I designing to, but what am I actually protecting against and at what cost.

The liquid cooling era demands that the industry evaluate redundancy based on what is inside the unit, not just how many units are in the row. The technology has evolved. The evaluation criteria should too.

About the Author

Nick Onyett

Nick Onyett

Nick Onyett is Integrated Marketing Manager, Cooling and Power at nVent.

nVent is a leading global provider of electrical connection and protection solutions. We believe our inventive electrical solutions enable safer systems and ensure a more secure world. We design, manufacture, market, install and service high performance products and solutions that connect and protect some of the world's most sensitive equipment, buildings and critical processes. We offer a comprehensive range of systems protection and electrical connections solutions across industry-leading brands that are recognized globally for quality, reliability and innovation. Our principal office is in London and our management office in the United States is in Minneapolis.

Sign up for our eNewsletters
Get the latest news and updates
Pkaza – Data Center Recruiting
Source: Pkaza – Data Center Recruiting
Sponsored
Peter Kazella of Pkaza – Data Center Recruiting provides an insiders look at how AI is impacting the recruitment process.
Dmitry Demidovich/Shutterstock.com
Source: Dmitry Demidovich/Shutterstock.com
Sponsored
Chrissy Olsen, Vice President of Critical Power Solutions at MPI Energy, explains why resilience must be defined by redundancy and adaptability in the AI era.