Next-Gen AI Cluster Infrastructure_E-Guide.pdf

Next-Gen AI Cluster Infrastructure_E-Guide.pdf

Next-Gen AI Cluster Infrastructure_E-Guide.pdf

https://www.se.com/ww/en/

I N T R O D U C T I O N

A I - D R I V E N C H A L L E N G E S F O R D A T A C E N T E R S

O U R S O L U T I O N S

R A C K I N F R A S T R U C T U R E

P O D I N F R A S T R U C T U R E

S O F T W A R E , S E R V I C E S , A N D S U S T A I N A B I L I T Y

R E F E R E N C E D E S I G N S

Imagine a bustling metropolis of AI servers, each acting as

a powerhouse of computation, generating immense heat

and consuming vast amounts of energy. As the

computational power increases, so too does the complex

orchestration of power and cooling infrastructure that

supports it. This evolution impacts the entire infrastructure,

from the smallest chip inside a server to the expansive

power grids that support them.

In this new era, achieving higher rack density is essential.

More servers are packed into tighter spaces, generating

more heat and requiring more power and cooling. At the

core of AI operation, the white space is where the intricate

interaction of power and cooling systems takes place at the

rack and pod levels. Here, space is valuable. It presents a

challenging paradox: AI servers need to be densely

clustered to maximize the number of GPUs per rack, but

with the interior rack space pushed to its limits, organizing

power and cooling infrastructure becomes a complex

puzzle akin to a high-stakes game of Tetris.

Beyond these design fundamentals, there are advanced

challenges to address, such as continuous reliability,

speed and ease of scale, and energy efficiency.

This guide will cover design considerations, key

elements for success, and the solutions needed to

achieve AI-ready white space.

At the core of AI operation, the white

space is where the intricate interaction

of power and cooling systems takes

place at the rack and pod levels.

Here, space is valuable.

Rapidly evolving AI needs bring

up three clear challenges.

Space

• More equipment in the rack leads to space and air flow constraints

• Compute, storage, and networking applications have different needs

• Racks must be stronger, deeper, optimized for high-density computing,

and flexible for multiple applications

Complexity

• White space is just as important as the rest of the data center and can’t

be viewed in isolation

• Engineering silos (power, cooling, services, etc.) lead to inefficiencies

and failure risk

• Hybrid cooling and power solutions are often required (such as air and

liquid cooling)

Unpredictability

• AI innovation is happening at unprecedented speeds

• It’s difficult to plan for IT needs even one or two years in the

future — yet deployments must happen quickly and at scale

• Regulatory changes and sustainability goals are in flux

The best AI data centers will address space, complexity, and unpredictability with an approach that is…

• White space infrastructure must evolve to meet

the rapidly changing demands of accelerated

compute

• Deployments must be streamlined, modular,

and scalable — at the rack and pod levels

• Environmental sustainability must be integrated

from the ground up — not as an afterthought

• No more silos: Design teams implementing

compute, power, cooling, and management must

work together

• Integrators and design teams must be able to trust

that what they order, ship, and implement will

arrive safely without damage and function properly

• Leave nothing to chance: Data-driven

modeling, planning, digital remote

monitoring, and management

optimize data center operations

• Reference designs validated between

chip, server, and physical

infrastructure vendors take work off

your plate

Unlock peak compute density — and scale it — with modular, integrated solutions you know will work. EcoStruxure Data

Center Solutions are engineered to meet the power and GPU-intensive demands of AI clusters and replace the complexity of

onsite construction. Our pod makes deployments in the white space less complex, so you can focus on core business.

Factory integration and

testing offsite

Offsite integration and

testing Plug-and-play onsite

Adaptable

We are always innovating our industry leading portfolio of reference designs,

prefabricated pods, and rack solutions to support high-density compute and AI.

To put it simply, we have the infrastructure to support high-density computing

— so you can adapt to a rapidly changing environment.

Collaborative

We partner with chip manufacturers, server manufacturers, and integrators to

break design silos and optimize IT infrastructure. Bringing together all the

parties required for a successful deployment is complicated. We’ve already

done the legwork so you can rest easy.

Data-driven

Our unrivaled reference designs, services, and software capabilities ensure

stress-free deployments and operation. Schneider Electric really is a one-stop-

shop for building and optimizing IT infrastructure.

Scott Data embraces AI

transformation with collaborative

effort involving Schneider Electric

• Read case study

Compass Datacenters deploys IT

faster with Schneider partnership

• Watch video

• Read case study

Pictured: Pod infrastructure for small- to mid-size clusters

https://www.se.com/ww/en/ https://www.wwt.com/case-study/scott-data-embraces-ai-transformation-to-offer-gpu-as-a-service https://youtu.be/SFFH5lTGSsc?si=BmEwr4C8Gwab8psd https://www.se.com/us/en/work/campaign/life-is-on/case-study/compass-datacenters.jsp

GPU innovation continues to drive the need for

more compute and advanced cooling. The higher

the rack density, the more advanced the

technology and infrastructure must be to

support it.

Efficient rack and pod architectures are key for:

• Scaling quickly

• Keeping up with the latest cooling technologies

• Adapting to different design needs (ex. AC / DC)

• Reducing your overall IT footprint

Cluster Pod Rack

Scalable, modular, prefabricated pods Flexible, integration-ready rack

architectures

Power Cooling Rack

systems

Security &

environmental monitoring

Software

& control

Data centers optimized for AI must bring these groups into a cohesive whole

where collaboration is built into the solution from the ground up.

Design teams and vendors often work

in silos, with different groups focusing on:

• Chips, servers, storage, and

networking gear

• Racks, power, and cooling

• Consulting and design to build out

infrastructure

• Deployment and installation

• Daily operations and management

O U R

E C O S Y S T E M

O F P A R T N E R S

C H I P

M A N U F A C T U R E R S

S E R V E R

M A N U F A C T U R E R S

T E C H

A G G R E G A T O R S /

S Y S T E M

I N T E G R A T O R S

H Y P E R S C A L E R S

AI infrastructure needs a full end-to-

end view from planning to operation:

• Advanced digital tools model, analyze,

and simulate power and cooling

performance

• Little or no margin for error requires real-

time monitoring in every domain

• Monitoring rapid power shifts is key in

some applications

• Sustainability reporting is becoming

mandatory — and IT teams are struggling

to do it

Digitized asset tracking & monitoring

Sustainability & modernization services

Consulting services for design & build

optimization Collect asset data to perform health

monitoring via connected sensorsDesign system architecture and

optimal asset strategy

Maintenance & support execution

Optimize asset maintenance

strategy with analytic-based

predictions

Optimize asset life and minimize

impact with circularity &

modernization services

Proactive

Asset

Management

at system

level

4 3

21

Engineered for demanding AI workloads, our NetShelter racks offer flexible, extra-large enclosures

with reinforced support for high-power servers and advanced cooling, including integrated Motivair

liquid cooling options.

Backed by our data-driven design expertise and AI ecosystem validation, these rack solutions

maximize compute per rack while ensuring reliable operation and simplified deployment, giving you

certainty in your AI infrastructure investment.

• Maximize compute per rack without compromising

operational performance

• Adapt to a wide array of servers, design standards,

and power and cooling architectures

• Support rack and stack integration for plug-and-play

deployment

Tame the complexity and confidently execute

build-out backed by:

• Robust, partner-validated reference designs

• Data-driven software and services

• Reliable, efficient high current rack power for EIA

and ORV-inspired architecture

• Maximize compute per rack, reducing costs

• Extra large racks for EIA and ORV-inspired architecture

• Reinforced load ratings for power & liquid cooling

• Rack, stack, and ship-ready for faster deployments

• Complete liquid cooling systems for Standard EIA

and ORV-inspired architecture ensure extreme

heat is removed effectively and efficiently

• Improve cooling performance and protect your

investment (and your bottom line)

• AI training and high-density inference

• AI networking and storage

• High-performance computing (HPC)

• GenAI and machine learning

• AI factories

• Far edge inference and networking

• Remote applications and non-IT environments

• Commercial and industrial applications

• Geographically dispersed IT portfolios

• Standard EIA 19” architecture

• Open AI system architecture

• Superpod architecture

• New builds or retrofit

• Mixed density applications

• Liquid, air, or hybrid cooling

• Mixed EIA and OCP design

• Ideal for L10 and L11 rack-level integrations

• Liquid- or air-cooled servers

• Rack and stack deployments

• Custom engineering, configuration, and

design aesthetics

Flexible layouts

to accommodate

a wide variety of

server models

Efficient, high-

current rack power

in horizontal,

vertical, and open

rack models

Liquid, air, or hybrid

high-density

cooling

Taller, deeper,

stronger racks with

shock-packaged

options

EIA

MGX OCP-inspired,

ORV3

Edge accelerated compute Data center accelerated compute Data center accelerated compute

Capacity Up to 20 kW Up to 160 kW Up to 300 kW

Power Vertical or horizontal

rack PDU and modular rack UPS

Vertical or horizontal

standard rack PDU

50V DC ORV3 rack busbar and

high-density power shelves

Cooling Air cooled Liquid-to-liquid and air cooled Liquid-to-liquid cooled

Rack standard EIA EIA ORV3/MGX

Flexible Options

Integration Level 10 ready

Customizable for Level 11

Power Vertical or horizontal PDUs

DC voltage in-rack busbar / power shelves

Cooling Air or liquid cooling

In-rack/room-based architecture

Rack Standard EIA 19”, ORV3, MGX and OCP-inspired

Standard EIA OCP-inspired

Power

High amperage 3-phase power

High quantity of dedicated circuits

Compact vertical and horizontal models

High-density power shelves

High-voltage DC in-rack busbar

Cooling

Rear Door Heat Exchanger (RDHx)

In-rack Manifold

In-rack CDU

Rear Door Heat Exchanger (ORV3)

In-rack Manifold (ORV3)

In-rack CDU (ORV3)

Rack

systems

Standard, large, X-large, XX-large sizes

High weight load ratings

Secure shock packaging options

ORV3 compatible racks

MGX approved racks

Security and

environmental

monitoring

Video surveillance

Intelligent access control

Temperature and humidity thresholds

Spot and rope leak detection

Video surveillance

Temperature and humidity thresholds

Spot and rope leak detection

• Engineered for AI and accelerated computing (more

room, more power, more cooling)

• Flexible for a wide array of applications — to protect your

investment now and in the future

• Pre-tested and validated reference guides for leading

server and chip brands (including everything needed to

configure, integrate, and deploy at the rack level)

• Combine with pod architecture for end-to-end data center

solutions that scale as needed

• Built with sustainability in mind — to maximize IT footprint

and optimize energy use

Mid scale Mid scale Hyperscale

Max kW N/A 73kW per rack

896kW per pod

132kW per rack

1.2 MW per pod

Scale increments 6+ racks 8-12 racks 40+ racks

Application scale Less than 60 racks Less than 60 racks Unlimited

Time Saved

CapEX Saved

• Hyper and exascale cluster infrastructure for easy plug-and-play

• Pre-built in a Schneider facility (very little onsite assembly required)

• Drives easier, faster, and more reliable IT roll-outs

• Streamlined infrastructure for small- to mid-sized deployments

• Configured and tested for seamless integration on site

• Efficiently scales IT power and cooling at smaller increments

• AI training and high-density inference

• AI networking and storage

• High-performance computing (HPC)

• GenAI and machine learning

• AI factories

• Max 132kW per rack

• Max 1.2MW per pod

• Pod scale at 40+ rack increments

• Hyperscale rollouts

• Standard EIA 19” architecture

• Open AI system architecture

• Superpod architecture

• New builds or retrofit

• Mixed density applications

• Liquid, air, or hybrid cooling

• Mixed EIA and OCP design

• Max 73kW per rack

• Max 896kW per pod

• Pod scale at 8-12 rack increments

• Deployments up to 60 racks

• Large, but not hyperscale rollouts

Confidential Property of Schneider Electric |Page 26

Aisle configuration:

Hot or cold

Total capacity:

Up to 1.2 MW

Rack size:

42U–58U

Cooling:

Supports air- and liquid-

cooled architectures

(InRow, RDHx, and DTC)

Racks per side:

21 racks per side

(1 networking rack)

Power:

Compatible with high-current busbar

or remote power panel

Cabling:

Large networking cable trays

Assembly and delivery:

• Prefabricated in Schneider

Electric prefab facility

• Transported to site partially

assembled

Prefabricated modular pod Configurable pod

Power

Low industrial voltage busway — integrated

Low industrial voltage remote power panel

Power quality metering

Standard voltage busway — integrated

Standard voltage remote power panel

Power quality metering

Cooling

Liquid to liquid cooling — integrated

Rear door heat exchanger

In-row and room-based air cooling

Hot or cold aisle configuration

Liquid to liquid cooling — non-integrated

Rear door heat exchanger

In-row and room-based air cooling

Hot or cold aisle configuration

Rack

systems

EIA, ORV3, MGX racks

High-performance networking cable management

EIA, ORV3 racks

Standard cable management

Security and

environmental

monitoring

Video surveillance

Environmental monitoring hub

Video surveillance

Environmental monitoring hub

• Prefabricated, modular pods for hyper and exascale —

Arrive on site partially-built and factory integrated for

premium speed-to-market and reliability.

• Configured pods for small- to mid-scale — Pre-tested

and validated to streamline onsite integration

• Save time and CapEx — Reduce planning, modeling,

and testing time; speed up deployment; and gain

efficiency

• Validated reference guides with complete architecture

for power, cooling, and white space, plus software and

services

• Built with sustainability in mind— to optimize IT

footprint and energy use

Capabilities

Monitoring and

management

Cloud-based or on-premise

Preventative alarming and management

Proactive, AI-driven insights

Device monitoring and control

Security reinforcement

Proactive monitoring,

optimization, and onsite

services

Predictive analytics

Condition-based maintenance

Expert remote and onsite support

Planning, modeling, and

optimizing

Visualization and simulation

Capacity tracking and modeling

Optimized cooling and energy consumption

Colocation tenant portal

Consulting and

customization services

Audits and analysis

Optimization recommendations

Digital twin modeling and simulation

Automated reports

Custom dashboards and features

Third-party integrations

Setting a bold,

actionable strategy

Implementing sustainable

data center designs

Driving sustainability

in operations

Securing sustainable

power

Decarbonizing

supply chains

https://download.schneider-electric.com/files?p_Doc_Ref=SPD_WP212_EN&p_enDocType=White+Paper&p_File_Name=WP212_V2_EN.pdf

Most reference designs Schneider Electric

reference designs

Basic bill of materials

Full architecture and analysis since 2012

Include engineering schematics, floor or rack

layouts, complete component list, and more

Cover electrical, mechanical, and IT

space systems

Room Pod Rack

We continually partner with leading server and chip manufacturers to create

reference designs that support AI workloads at varying densities and different

power and cooling architectures.

Learn more about AI-ready reference designs.

Schneider Electric reference designs are available for data centers,

pods, and even down to the rack level. This is crucial to help data

center designers reduce planning time and deploy with confidence.

https://www.se.com/ww/en/work/solutions/for-business/data-centers-and-networks/reference-designs/

https://www.se.com/ww/en/ https://www.se.com/ww/en/ http://www.apc.com/

Intro Slide 1: Next-gen AI Clusters Modular pod and integrated rack infrastructure for AI and accelerated computing Slide 2 Slide 3

AI-driven challenges for data centers Slide 4 Slide 5 Slide 6

Our solutions Slide 7 Slide 8 Slide 9 Slide 10 Slide 11 Slide 12

Rack infrastructure Slide 13 Slide 14 Slide 15: EcoStruxure Rack Solutions Slide 16 Slide 17 Slide 18 Slide 19 Slide 20 Slide 21

Pod Infrastructure Slide 22 Slide 23 Slide 24 Slide 25 Slide 26 Slide 27 Slide 28

Software, services, and sustainability Slide 29 Slide 30 Slide 31

Reference designs Slide 32 Slide 33 Slide 34 Slide 35 Slide 36


Item Type: pdf