Building smarter: How infrastructure is responding to the demands of AI workloads 

Building smarter: How infrastructure is responding to the demands of AI workloads 

Jon Abbott, Technologies Director, Global Strategic Clients at Vertiv discusses AI demand on data centres and how hardware must keep up 

Artificial Intelligence (AI) is no longer limited to R&D labs or big tech campuses. From fraud detection in financial services to predictive maintenance in manufacturing, AI has entered production, and it’s here to stay. But the shift isn’t just happening in software. Behind the scenes, the critical digital infrastructure needed to support AI is undergoing an evolution too. 

Established data centres – which are traditionally designed for predictable, low-variance workloads – are now reaching their limits. AI brings with it dense hardware, erratic power profiles and thermal outputs that strain even the most robust legacy systems. And the challenges are happening now. 

For those responsible for IT infrastructure, the pressure is mounting to adapt quickly and intelligently. 

What makes AI workloads different? 

AI hardware doesn’t behave like conventional compute. Graphics processing units (GPUs), which form the backbone of most AI systems, are designed for parallel processing. This ability to process vast amounts of information makes them process efficient, but also power-hungry and heat intensive. 

Where a typical enterprise server might draw 5 kilowatts to 10 kilowatts, a rack filled with GPU servers can easily demand 30, 50, or even beyond 100 kilowatts. And these loads aren’t always steady. AI tasks often run in bursts, with workloads ramping from idle to peak output and back in a matter of seconds. This volatility creates unique challenges for both electrical and cooling systems. 

In short, AI puts more stress on infrastructure and it does so in ways that older data centre designs simply weren’t built to handle. 

Cooling is now a strategic concern 

As densities increase, the limits of traditional air cooling are being tested. Fans, vents and aisle containment strategies still have a role to play, but on their own, they’re not enough. Heat builds quickly and when it isn’t removed efficiently, it can lead to serious incidents such as thermal throttling or even catastrophic system failure. 

That’s why many operators are turning to liquid cooling. Direct-to-chip systems and immersion cooling solutions are becoming more common, particularly in facilities that need to house dense clusters of AI servers. 

These cooling methods introduce new considerations: everything from leak prevention and fluid maintenance to regulating flow rates and temperatures in real time. But there are benefits: improved energy efficiency, higher density per square foot and greater long-term flexibility. 

Electrical infrastructure is under strain 

AI workloads don’t just generate heat. They demand a different kind of electrical backbone. 

The spike-driven nature of AI tasks means power systems must be responsive and resilient. Power distribution units (PDUs), uninterruptible power supplies (UPSs), switchgear and circuit protection systems all need to be sized and selected with this behaviour in mind. 

Failures can cascade quickly when AI nodes are involved. A momentary voltage drop or power quality issue may go unnoticed in a standard IT environment but could interrupt or corrupt an active AI training process, costing hours or even days of processing time. 

This has placed a renewed focus on electrical commissioning. Systems must be tested under stress, not just at average operating conditions. Load banks, simulated faults and failover exercises are becoming essential parts of the handover process. 

Commissioning for dynamic environments 

Data centre infrastructure used to be relatively static. Once a data hall was provisioned, it often ran close to capacity with consistent, stable loads. AI changes that norm. The infrastructure must now handle dynamic scaling by adding and removing servers, integrating new hardware generations, and adjusting cooling and power delivery on the fly. 

This requires a different commissioning mindset. Facilities are increasingly deploying digital twins to test new layouts and configurations before rolling them out. These virtual environments help predict airflow patterns, simulate energy loads and even model thermal hotspots – reducing surprises during live deployment. 

Moreover, commissioning no longer ends once a system is signed off. With workloads evolving rapidly, operators are introducing continuous commissioning models – re-testing infrastructure periodically as usage changes. 

The grid is becoming a bottleneck 

Power availability is one of the most pressing concerns for data centre expansion today. In parts of the UK, securing new grid connections can take years. In urban centres, the situation is especially acute. Some developers are being told that capacity won’t be available until the 2030s. 

For AI infrastructure, which requires significantly more power than general-purpose compute, this is a major constraint. To work around it, some operators are turning to on-site generation – either through combined heat and power (CHP) systems, natural gas turbines, or hydrogen-ready backup units. Others are investing in large-scale battery energy storage systems (BESS) to help manage peak loads and maintain power continuity. 

Cooling systems are tightly linked to these developments. Liquid cooling systems rely on consistent energy delivery. If the power goes, so does the cooling – and when racks are drawing tens of kilowatts, thermal runaway can happen fast. 

Reuse of heat is being explored seriously 

A broader benefit of AI’s thermal intensity is that it creates an opportunity to recapture waste heat. This isn’t a new idea, but the practicality has often been limited by low-grade heat and complex routing challenges. 

Liquid cooling changes the equation. It delivers higher-temperature thermal energy in a concentrated form, which can be more easily captured and redirected. In some cases, data centres are feeding heat into district heating networks or using it to warm adjacent offices or greenhouses. 

While heat reuse isn’t viable in every case, it’s gaining traction; especially where local authorities require sustainability planning as part of development approvals. 

Visibility is now essential 

With critical digital infrastructure becoming more complex and workloads more unpredictable, monitoring and analytics are taking centre stage. Power and thermal telemetry, fault detection, predictive maintenance and workload tracking are no longer optional. They’re core to operating high-performance environments. 

Modern data centres are investing in integrated monitoring platforms that can track power draw, cooling performance and infrastructure availability across multiple systems. AI itself is being used to optimise the environment that supports it – modelling failure risk, rebalancing load and even predicting when and where maintenance will be needed. 

For network and infrastructure teams, this means more collaboration between IT and facilities, and a sharper focus on real-time data. 

AI infrastructure isn’t just about bigger power and better cooling.  It’s also about designing systems that can adapt and scale. What works today may need to change tomorrow as models grow, workloads shift and new demands emerge. Flexibility is fast becoming the most valuable asset in any data centre. That applies to rack design, cooling choice, power provisioning and network architecture.

Browse our latest issue

Intelligent Data Centres

View Magazine Archive