Next-Gen AI Cluster Infrastructure_E-Guide.pdf

https://www.se.com/ww/en/
I N T R O D U C T I O N
A I - D R I V E N C H A L L E N G E S F O R D A T A C E N T E R S
O U R S O L U T I O N S
R A C K I N F R A S T R U C T U R E
P O D I N F R A S T R U C T U R E
S O F T W A R E , S E R V I C E S , A N D S U S T A I N A B I L I T Y
R E F E R E N C E D E S I G N S
Imagine a bustling metropolis of AI servers, each acting as
a powerhouse of computation, generating immense heat
and consuming vast amounts of energy. As the
computational power increases, so too does the complex
orchestration of power and cooling infrastructure that
supports it. This evolution impacts the entire infrastructure,
from the smallest chip inside a server to the expansive
power grids that support them.
In this new era, achieving higher rack density is essential.
More servers are packed into tighter spaces, generating
more heat and requiring more power and cooling. At the
core of AI operation, the white space is where the intricate
interaction of power and cooling systems takes place at the
rack and pod levels. Here, space is valuable. It presents a
challenging paradox: AI servers need to be densely
clustered to maximize the number of GPUs per rack, but
with the interior rack space pushed to its limits, organizing
power and cooling infrastructure becomes a complex
puzzle akin to a high-stakes game of Tetris.
Beyond these design fundamentals, there are advanced
challenges to address, such as continuous reliability,
speed and ease of scale, and energy efficiency.
This guide will cover design considerations, key
elements for success, and the solutions needed to
achieve AI-ready white space.
At the core of AI operation, the white
space is where the intricate interaction
of power and cooling systems takes
place at the rack and pod levels.
Here, space is valuable.
Rapidly evolving AI needs bring
up three clear challenges.
Space
• More equipment in the rack leads to space and air flow constraints
• Compute, storage, and networking applications have different needs
• Racks must be stronger, deeper, optimized for high-density computing,
and flexible for multiple applications
Complexity
• White space is just as important as the rest of the data center and can’t
be viewed in isolation
• Engineering silos (power, cooling, services, etc.) lead to inefficiencies
and failure risk
• Hybrid cooling and power solutions are often required (such as air and
liquid cooling)
Unpredictability
• AI innovation is happening at unprecedented speeds
• It’s difficult to plan for IT needs even one or two years in the
future — yet deployments must happen quickly and at scale
• Regulatory changes and sustainability goals are in flux
The best AI data centers will address space, complexity, and unpredictability with an approach that is…
• White space infrastructure must evolve to meet
the rapidly changing demands of accelerated
compute
• Deployments must be streamlined, modular,
and scalable — at the rack and pod levels
• Environmental sustainability must be integrated
from the ground up — not as an afterthought
• No more silos: Design teams implementing
compute, power, cooling, and management must
work together
• Integrators and design teams must be able to trust
that what they order, ship, and implement will
arrive safely without damage and function properly
• Leave nothing to chance: Data-driven
modeling, planning, digital remote
monitoring, and management
optimize data center operations
• Reference designs validated between
chip, server, and physical
infrastructure vendors take work off
your plate
Unlock peak compute density — and scale it — with modular, integrated solutions you know will work. EcoStruxure Data
Center Solutions are engineered to meet the power and GPU-intensive demands of AI clusters and replace the complexity of
onsite construction. Our pod makes deployments in the white space less complex, so you can focus on core business.
Factory integration and
testing offsite
Offsite integration and
testing Plug-and-play onsite
Adaptable
We are always innovating our industry leading portfolio of reference designs,
prefabricated pods, and rack solutions to support high-density compute and AI.
To put it simply, we have the infrastructure to support high-density computing
— so you can adapt to a rapidly changing environment.
Collaborative
We partner with chip manufacturers, server manufacturers, and integrators to
break design silos and optimize IT infrastructure. Bringing together all the
parties required for a successful deployment is complicated. We’ve already
done the legwork so you can rest easy.
Data-driven
Our unrivaled reference designs, services, and software capabilities ensure
stress-free deployments and operation. Schneider Electric really is a one-stop-
shop for building and optimizing IT infrastructure.
Scott Data embraces AI
transformation with collaborative
effort involving Schneider Electric
• Read case study
Compass Datacenters deploys IT
faster with Schneider partnership
• Watch video
• Read case study
Pictured: Pod infrastructure for small- to mid-size clusters
https://www.se.com/ww/en/ https://www.wwt.com/case-study/scott-data-embraces-ai-transformation-to-offer-gpu-as-a-service https://youtu.be/SFFH5lTGSsc?si=BmEwr4C8Gwab8psd https://www.se.com/us/en/work/campaign/life-is-on/case-study/compass-datacenters.jsp
GPU innovation continues to drive the need for
more compute and advanced cooling. The higher
the rack density, the more advanced the
technology and infrastructure must be to
support it.
Efficient rack and pod architectures are key for:
• Scaling quickly
• Keeping up with the latest cooling technologies
• Adapting to different design needs (ex. AC / DC)
• Reducing your overall IT footprint
Cluster Pod Rack
Scalable, modular, prefabricated pods Flexible, integration-ready rack
architectures
Power Cooling Rack
systems
Security &
environmental monitoring
Software
& control
Data centers optimized for AI must bring these groups into a cohesive whole
where collaboration is built into the solution from the ground up.
Design teams and vendors often work
in silos, with different groups focusing on:
• Chips, servers, storage, and
networking gear
• Racks, power, and cooling
• Consulting and design to build out
infrastructure
• Deployment and installation
• Daily operations and management
O U R
E C O S Y S T E M
O F P A R T N E R S
C H I P
M A N U F A C T U R E R S
S E R V E R
M A N U F A C T U R E R S
T E C H
A G G R E G A T O R S /
S Y S T E M
I N T E G R A T O R S
H Y P E R S C A L E R S
AI infrastructure needs a full end-to-
end view from planning to operation:
• Advanced digital tools model, analyze,
and simulate power and cooling
performance
• Little or no margin for error requires real-
time monitoring in every domain
• Monitoring rapid power shifts is key in
some applications
• Sustainability reporting is becoming
mandatory — and IT teams are struggling
to do it
Digitized asset tracking & monitoring
Sustainability & modernization services
Consulting services for design & build
optimization Collect asset data to perform health
monitoring via connected sensorsDesign system architecture and
optimal asset strategy
Maintenance & support execution
Optimize asset maintenance
strategy with analytic-based
predictions
Optimize asset life and minimize
impact with circularity &
modernization services
Proactive
Asset
Management
at system
level
4 3
21
Engineered for demanding AI workloads, our NetShelter racks offer flexible, extra-large enclosures
with reinforced support for high-power servers and advanced cooling, including integrated Motivair
liquid cooling options.
Backed by our data-driven design expertise and AI ecosystem validation, these rack solutions
maximize compute per rack while ensuring reliable operation and simplified deployment, giving you
certainty in your AI infrastructure investment.
• Maximize compute per rack without compromising
operational performance
• Adapt to a wide array of servers, design standards,
and power and cooling architectures
• Support rack and stack integration for plug-and-play
deployment
Tame the complexity and confidently execute
build-out backed by:
• Robust, partner-validated reference designs
• Data-driven software and services
• Reliable, efficient high current rack power for EIA
and ORV-inspired architecture
• Maximize compute per rack, reducing costs
• Extra large racks for EIA and ORV-inspired architecture
• Reinforced load ratings for power & liquid cooling
• Rack, stack, and ship-ready for faster deployments
• Complete liquid cooling systems for Standard EIA
and ORV-inspired architecture ensure extreme
heat is removed effectively and efficiently
• Improve cooling performance and protect your
investment (and your bottom line)
• AI training and high-density inference
• AI networking and storage
• High-performance computing (HPC)
• GenAI and machine learning
• AI factories
• Far edge inference and networking
• Remote applications and non-IT environments
• Commercial and industrial applications
• Geographically dispersed IT portfolios
• Standard EIA 19” architecture
• Open AI system architecture
• Superpod architecture
• New builds or retrofit
• Mixed density applications
• Liquid, air, or hybrid cooling
• Mixed EIA and OCP design
• Ideal for L10 and L11 rack-level integrations
• Liquid- or air-cooled servers
• Rack and stack deployments
• Custom engineering, configuration, and
design aesthetics
Flexible layouts
to accommodate
a wide variety of
server models
Efficient, high-
current rack power
in horizontal,
vertical, and open
rack models
Liquid, air, or hybrid
high-density
cooling
Taller, deeper,
stronger racks with
shock-packaged
options
EIA
MGX OCP-inspired,
ORV3
Edge accelerated compute Data center accelerated compute Data center accelerated compute
Capacity Up to 20 kW Up to 160 kW Up to 300 kW
Power Vertical or horizontal
rack PDU and modular rack UPS
Vertical or horizontal
standard rack PDU
50V DC ORV3 rack busbar and
high-density power shelves
Cooling Air cooled Liquid-to-liquid and air cooled Liquid-to-liquid cooled
Rack standard EIA EIA ORV3/MGX
Flexible Options
Integration Level 10 ready
Customizable for Level 11
Power Vertical or horizontal PDUs
DC voltage in-rack busbar / power shelves
Cooling Air or liquid cooling
In-rack/room-based architecture
Rack Standard EIA 19”, ORV3, MGX and OCP-inspired
Standard EIA OCP-inspired
Power
High amperage 3-phase power
High quantity of dedicated circuits
Compact vertical and horizontal models
High-density power shelves
High-voltage DC in-rack busbar
Cooling
Rear Door Heat Exchanger (RDHx)
In-rack Manifold
In-rack CDU
Rear Door Heat Exchanger (ORV3)
In-rack Manifold (ORV3)
In-rack CDU (ORV3)
Rack
systems
Standard, large, X-large, XX-large sizes
High weight load ratings
Secure shock packaging options
ORV3 compatible racks
MGX approved racks
Security and
environmental
monitoring
Video surveillance
Intelligent access control
Temperature and humidity thresholds
Spot and rope leak detection
Video surveillance
Temperature and humidity thresholds
Spot and rope leak detection
• Engineered for AI and accelerated computing (more
room, more power, more cooling)
• Flexible for a wide array of applications — to protect your
investment now and in the future
• Pre-tested and validated reference guides for leading
server and chip brands (including everything needed to
configure, integrate, and deploy at the rack level)
• Combine with pod architecture for end-to-end data center
solutions that scale as needed
• Built with sustainability in mind — to maximize IT footprint
and optimize energy use
Mid scale Mid scale Hyperscale
Max kW N/A 73kW per rack
896kW per pod
132kW per rack
1.2 MW per pod
Scale increments 6+ racks 8-12 racks 40+ racks
Application scale Less than 60 racks Less than 60 racks Unlimited
Time Saved
CapEX Saved
• Hyper and exascale cluster infrastructure for easy plug-and-play
• Pre-built in a Schneider facility (very little onsite assembly required)
• Drives easier, faster, and more reliable IT roll-outs
• Streamlined infrastructure for small- to mid-sized deployments
• Configured and tested for seamless integration on site
• Efficiently scales IT power and cooling at smaller increments
• AI training and high-density inference
• AI networking and storage
• High-performance computing (HPC)
• GenAI and machine learning
• AI factories
• Max 132kW per rack
• Max 1.2MW per pod
• Pod scale at 40+ rack increments
• Hyperscale rollouts
• Standard EIA 19” architecture
• Open AI system architecture
• Superpod architecture
• New builds or retrofit
• Mixed density applications
• Liquid, air, or hybrid cooling
• Mixed EIA and OCP design
• Max 73kW per rack
• Max 896kW per pod
• Pod scale at 8-12 rack increments
• Deployments up to 60 racks
• Large, but not hyperscale rollouts
Confidential Property of Schneider Electric |Page 26
Aisle configuration:
Hot or cold
Total capacity:
Up to 1.2 MW
Rack size:
42U–58U
Cooling:
Supports air- and liquid-
cooled architectures
(InRow, RDHx, and DTC)
Racks per side:
21 racks per side
(1 networking rack)
Power:
Compatible with high-current busbar
or remote power panel
Cabling:
Large networking cable trays
Assembly and delivery:
• Prefabricated in Schneider
Electric prefab facility
• Transported to site partially
assembled
Prefabricated modular pod Configurable pod
Power
Low industrial voltage busway — integrated
Low industrial voltage remote power panel
Power quality metering
Standard voltage busway — integrated
Standard voltage remote power panel
Power quality metering
Cooling
Liquid to liquid cooling — integrated
Rear door heat exchanger
In-row and room-based air cooling
Hot or cold aisle configuration
Liquid to liquid cooling — non-integrated
Rear door heat exchanger
In-row and room-based air cooling
Hot or cold aisle configuration
Rack
systems
EIA, ORV3, MGX racks
High-performance networking cable management
EIA, ORV3 racks
Standard cable management
Security and
environmental
monitoring
Video surveillance
Environmental monitoring hub
Video surveillance
Environmental monitoring hub
• Prefabricated, modular pods for hyper and exascale —
Arrive on site partially-built and factory integrated for
premium speed-to-market and reliability.
• Configured pods for small- to mid-scale — Pre-tested
and validated to streamline onsite integration
• Save time and CapEx — Reduce planning, modeling,
and testing time; speed up deployment; and gain
efficiency
• Validated reference guides with complete architecture
for power, cooling, and white space, plus software and
services
• Built with sustainability in mind— to optimize IT
footprint and energy use
Capabilities
Monitoring and
management
Cloud-based or on-premise
Preventative alarming and management
Proactive, AI-driven insights
Device monitoring and control
Security reinforcement
Proactive monitoring,
optimization, and onsite
services
Predictive analytics
Condition-based maintenance
Expert remote and onsite support
Planning, modeling, and
optimizing
Visualization and simulation
Capacity tracking and modeling
Optimized cooling and energy consumption
Colocation tenant portal
Consulting and
customization services
Audits and analysis
Optimization recommendations
Digital twin modeling and simulation
Automated reports
Custom dashboards and features
Third-party integrations
Setting a bold,
actionable strategy
Implementing sustainable
data center designs
Driving sustainability
in operations
Securing sustainable
power
Decarbonizing
supply chains
https://download.schneider-electric.com/files?p_Doc_Ref=SPD_WP212_EN&p_enDocType=White+Paper&p_File_Name=WP212_V2_EN.pdf
Most reference designs Schneider Electric
reference designs
Basic bill of materials
Full architecture and analysis since 2012
Include engineering schematics, floor or rack
layouts, complete component list, and more
Cover electrical, mechanical, and IT
space systems
Room Pod Rack
We continually partner with leading server and chip manufacturers to create
reference designs that support AI workloads at varying densities and different
power and cooling architectures.
Learn more about AI-ready reference designs.
Schneider Electric reference designs are available for data centers,
pods, and even down to the rack level. This is crucial to help data
center designers reduce planning time and deploy with confidence.
https://www.se.com/ww/en/work/solutions/for-business/data-centers-and-networks/reference-designs/
https://www.se.com/ww/en/ https://www.se.com/ww/en/ http://www.apc.com/
Intro Slide 1: Next-gen AI Clusters Modular pod and integrated rack infrastructure for AI and accelerated computing Slide 2 Slide 3
AI-driven challenges for data centers Slide 4 Slide 5 Slide 6
Our solutions Slide 7 Slide 8 Slide 9 Slide 10 Slide 11 Slide 12
Rack infrastructure Slide 13 Slide 14 Slide 15: EcoStruxure Rack Solutions Slide 16 Slide 17 Slide 18 Slide 19 Slide 20 Slide 21
Pod Infrastructure Slide 22 Slide 23 Slide 24 Slide 25 Slide 26 Slide 27 Slide 28
Software, services, and sustainability Slide 29 Slide 30 Slide 31
Reference designs Slide 32 Slide 33 Slide 34 Slide 35 Slide 36