SK hynix AI Infrastructure Hits a Physical Limit: Power and Cooling
SK hynix has shifted its AI infrastructure analysis toward a constraint that faster processors cannot solve alone: dense computing systems need enough electricity and cooling capacity to run.
The company’s latest examination of AI data centers argues that performance now depends on an integrated system. Compute, memory, storage, networking, power delivery, and thermal management must operate together. One weak layer can leave expensive accelerators waiting, throttled, or unavailable.
That argument changes the center of the AI infrastructure race. Nvidia, hyperscale cloud providers, memory suppliers, utilities, and cooling vendors are no longer solving separate problems. They are building around one physical limit: every watt entering a server eventually becomes heat that the facility must remove.
The shift does not make processors or high-bandwidth memory less important. It makes their usable performance dependent on infrastructure outside the chip package. A faster accelerator has limited value if a site cannot energize the rack, move coolant through it, or connect the facility to the grid.
SK hynix AI Infrastructure Analysis Moves Beyond the Chip
SK hynix is treating power and cooling as performance components, not background facility services.
The company’s infrastructure analysis extends a broader argument about AI system design. Individual components matter, but their specifications do not guarantee application-level performance.
An AI training cluster repeatedly transfers model parameters, activations, checkpoints, and training data. Inference systems must retrieve context, move it through memory, coordinate accelerators, and return answers within acceptable latency limits.
Every stage consumes energy. Every transfer also produces heat, including transfers that do not directly contribute to the calculation requested by the user.
This creates a chain of dependencies. Accelerators need data from memory. Memory and storage need high-throughput connections. Network switches must coordinate thousands of devices. Power systems must supply all of them without interruption.
Cooling equipment then has to remove the resulting heat. If it cannot, hardware reduces its operating frequency to remain within safe temperatures. That process, called thermal throttling, protects components while lowering available performance.
The same issue applies at facility scale. A data center operator can purchase processors without possessing the electrical capacity needed to deploy them. The operator can also reserve utility power yet lack the internal distribution equipment required to deliver that power to dense racks.
Transformers, switchgear, uninterruptible power supplies, busways, and backup systems become part of the deployment schedule. Their specifications affect how much computing capacity a building can support.
The thermal path matters just as much. Heat must travel from the silicon through packaging materials, cold plates or heat sinks, coolant loops, heat exchangers, and the facility’s final heat-rejection system.
Any weak link raises temperatures or limits rack density. Operators may have to spread equipment across more floor space, reduce processor power, or delay deployment while upgrading the building.
SK hynix has a direct interest in this transition. Its high-bandwidth memory products sit beside accelerators in dense computing packages. HBM increases memory bandwidth by stacking memory dies and connecting them through short vertical pathways.
That architecture helps supply data to processors faster. However, it also concentrates components and heat within a smaller area.
In May 2026, SK hynix announced an integrated cooling approach for future HBM products. The company said its iHBM thermal design reduced package thermal resistance by 30% in its testing.
That figure remains a company claim, and commercial results will depend on final products and system designs. Still, the announcement shows why memory suppliers now discuss cooling alongside bandwidth and capacity.
Thermal performance is becoming part of product architecture. It can no longer be left entirely to the server maker or facility operator.
AI Data Center Power Demand Is Growing Faster Than the Grid
The central power problem is not simply total energy consumption. It is delivering large amounts of electricity at the required place and time.
The International Energy Agency estimated that data centers consumed about 415 terawatt-hours of electricity in 2024. That represented approximately 1.5% of global electricity consumption.
Its base case projects data center consumption reaching about 945 terawatt-hours in 2030. That would place the sector just below 3% of global electricity use.
The IEA expects consumption from accelerated servers, which are primarily associated with AI adoption, to rise by about 30% annually through 2030. Those servers account for almost half of the projected net increase in global data center electricity consumption.
The figures are significant, but their location creates the sharper operational challenge. Data centers concentrate loads within particular utility territories, substations, and transmission regions.
An electric vehicle fleet spreads demand across many homes and charging locations. An AI campus can request hundreds of megawatts behind a small number of grid connections.
The IEA’s energy demand outlook notes that a data center can become operational within two or three years. Major energy infrastructure often requires longer planning and construction cycles.
That timing mismatch places utilities and developers under pressure. Technology companies want computing capacity while demand remains strong. Utilities must study grid effects, procure transformers, build substations, and sometimes expand generation or transmission.
A planned campus is not usable AI capacity until power reaches the servers. Announced capacity, contracted capacity, and energized capacity are different measures.
The United States illustrates the scale of the uncertainty. A 2024 Department of Energy report estimated that data centers used about 4.4% of national electricity in 2023.
The report projected that their share could reach between 6.7% and 12% by 2028. That range is unusually wide because AI adoption, hardware efficiency, utilization, and construction schedules remain difficult to predict.
The U.S. energy estimate does not mean every proposed facility will receive a connection. It shows how different deployment scenarios can create very different grid outcomes.
Power availability therefore affects where AI systems are built. Training workloads can sometimes move to regions with available electricity because they do not always require immediate proximity to end users.
Production inference is less flexible. A service supporting factories, hospitals, financial systems, or interactive applications may have latency, resilience, and data-residency requirements.
Companies must balance those requirements against local grid capacity. A region with users, fiber connectivity, and suitable land can still be a poor deployment location if electricity will not arrive on schedule.
This pressure is changing procurement. Technology companies are signing long-term renewable contracts, supporting nuclear restarts, investing in advanced geothermal projects, and considering on-site generation.
Those actions can increase supply, but they do not eliminate the infrastructure between a generator and a server. Transmission lines, substations, transformers, and internal electrical systems remain necessary.
Clean generation also does not automatically provide constant power at one specific campus. Operators still need grid balancing, storage, firm generation, or other resources when wind and solar output changes.
The result is a more complicated definition of AI performance. A benchmark can describe how quickly a processor completes a workload. It cannot guarantee that an operator can power thousands of those processors in the desired location.
A 120-Kilowatt Rack Changes the Cooling Equation
Rack-level power density is forcing AI data centers to replace air-only assumptions with liquid-assisted thermal designs.
Traditional data centers distribute computing equipment across racks that can be cooled by moving conditioned air through server fans and room-level systems. That approach remains practical for many enterprise and cloud workloads.
Dense accelerator systems produce a different thermal profile. They place more computing equipment, memory, networking, and power hardware within each rack.
Nvidia lists the approximate power consumption of its GB200 NVL72 system at 120 kilowatts per rack. The rack combines 72 Blackwell GPUs with Grace CPUs and NVLink switching equipment.
For comparison, many conventional enterprise racks operate at a fraction of that density. The exact difference varies widely by facility and workload, so there is no single traditional-rack number that applies everywhere.
The important change is architectural. A 120-kilowatt rack concentrates heat into a small footprint and requires an electrical path designed for the same load.
Nvidia uses a hybrid system for the GB200 NVL72. Its documentation describes roughly 85% liquid cooling and 15% air cooling.
The company’s rack-scale cooling design sends liquid to cold plates attached to major heat-producing components. The coolant absorbs heat closer to its source than room air can.
Direct-to-chip liquid cooling does not submerge the entire server. It circulates coolant through sealed cold plates that make thermal contact with processors, memory-adjacent components, or other hot devices.
A coolant distribution unit transfers heat between the technology equipment loop and the facility water loop. Pumps, manifolds, sensors, controls, and heat exchangers maintain flow and temperature.
Liquid carries more heat per unit of volume than air. That makes it suitable for racks where fans alone would require excessive airflow, space, or electrical power.
However, liquid cooling is not a universal replacement for air. Servers still contain components that need airflow, and facilities must remove the collected heat from the building.
Operators also inherit new engineering requirements. They must monitor coolant chemistry, pressure, flow, connector integrity, condensation risk, and leak detection.
Maintenance procedures become more involved. Technicians need training for fluid systems near high-value electronics. Operators must decide how to isolate a rack, replace a tray, or service a pump without disrupting the cluster.
The cooling system must also match the computing architecture. A cold plate designed for one processor package cannot automatically support every future package.
Supply temperatures and flow rates can change between hardware generations. Facility planners must avoid building a thermal system that becomes obsolete when the next rack arrives.
Retrofitting existing data centers presents another challenge. An older building may have sufficient total power but lack the piping, floor loading, rack dimensions, or heat-rejection capacity required for dense liquid-cooled systems.
Operators can spread accelerator servers across more racks, but that increases cable lengths and floor-space requirements. It can also complicate high-speed networking, where physical distance affects latency and signal integrity.
They can lower processor power, but that reduces performance. They can install rear-door heat exchangers or contained air systems, though those approaches still face limits at rising densities.
New construction offers greater flexibility. Designers can coordinate the electrical and thermal systems around expected rack loads from the start.
That does not remove uncertainty. A building designed today must support hardware delivered years later, even though processor power, cooling interfaces, and cluster architecture continue to change.
Power Versus Performance Is the New Infrastructure Tradeoff
The competition is no longer maximum chip performance versus a slower chip. It is theoretical performance versus deployable performance.
Hardware vendors can improve performance by adding transistors, memory bandwidth, and faster interconnects. Those gains often increase package complexity and system power, even when performance per watt improves.
Performance per watt measures how much computing work a system completes for each unit of energy. It is essential, but it does not tell operators whether total demand will fall.
More efficient hardware can lower the energy required for one task. Lower costs can then encourage more tasks, larger models, longer reasoning processes, and wider deployment.
This rebound effect explains why efficiency gains can coexist with growing electricity consumption. Demand expands faster than the energy required for each unit of work declines.
AI agents add another variable. An agent may call a model repeatedly, retrieve documents, use software tools, and evaluate intermediate results before producing one answer.
A single user action can therefore trigger multiple inference operations. Video generation, scientific workloads, and extended reasoning can also require more computation than short text responses.
Operators need efficiency at several levels. The processor must perform calculations efficiently. Memory must feed the processor without excessive waiting or data movement.
Networks must coordinate devices with minimal delay. Software must schedule jobs and avoid leaving equipment idle while still consuming power.
The cooling plant must remove heat without adding excessive overhead. The facility must convert and distribute electricity efficiently.
This is why power usage effectiveness, or PUE, cannot describe the entire problem. PUE compares a facility’s total electricity consumption with the electricity used by its technology equipment.
A lower PUE indicates less facility overhead. It does not reveal whether the processors are well utilized, whether software schedules work efficiently, or whether a model uses more computation than necessary.
A facility can have an excellent PUE while running poorly utilized accelerators. Conversely, a slightly higher PUE might support more useful computing work if the overall system is better matched to its workload.
The relevant unit is useful work delivered within power, cost, latency, and reliability limits. Measuring that outcome remains difficult because companies disclose little standardized data about AI utilization or energy per workload.
SK hynix’s system-level framing is useful here. Memory bandwidth can keep accelerators working, but greater bandwidth also affects package power and heat.
Reducing data movement can save energy because information does not travel repeatedly across long electrical paths. Placing memory closer to compute can improve speed and efficiency, although denser packaging makes heat removal harder.
This is the core tradeoff. Tighter integration improves data movement while concentrating thermal load.
The same tension appears at rack scale. High-density systems shorten network paths and place more accelerators within one tightly connected domain.
They also require larger power feeds and more capable cooling loops. A facility that cannot support the rack may need a less dense design, losing some of the networking and space advantages.
Infrastructure vendors are responding with coordinated reference architectures. Chipmakers work with server manufacturers, power-equipment suppliers, and cooling specialists before a system reaches customers.
That cooperation reduces integration risk, but it can also narrow operator choice. A tightly specified rack may require particular coolant temperatures, power shelves, connectors, and management software.
Customers gain a validated design while becoming more dependent on a particular equipment chain. Changing one component can require validation across the whole system.
Cloud providers possess an advantage because they can coordinate custom servers, networking, software, and buildings at large scale. Smaller operators and enterprises usually work with existing facilities and multiple suppliers.
They face harder retrofit decisions. Moving an AI workload to the cloud can avoid a local facility upgrade, but it introduces questions about operating cost, data control, latency, and supplier dependence.
Building locally offers control but requires electrical capacity, cooling expertise, and longer planning. The correct choice depends on workload characteristics, not on processor specifications alone.
Liquid Cooling Solves Heat Transfer, Not Every Constraint
Cooling technology can unlock denser computing, but it does not create grid capacity or guarantee economical AI services.
The excitement around liquid cooling sometimes compresses several problems into one. Liquid moves heat away from chips more effectively than air, but electricity must still reach the rack.
The captured heat must still be rejected outdoors, reused, or transferred elsewhere. Pumps and heat-rejection equipment still consume energy.
Water use also depends on the complete system. A liquid-cooled rack can operate within a closed technology loop while the facility uses evaporative cooling towers elsewhere.
Another facility may use dry coolers that consume less water but require more energy or provide less cooling efficiency during hot weather. Climate and local water availability affect the best design.
Claims that liquid cooling either eliminates water use or necessarily consumes more water need context. The answer depends on coolant loops, heat-rejection equipment, operating temperatures, climate, and power generation.
Reliability requires similar caution. Liquid near electronics introduces leak concerns, yet conventional air-cooled facilities also contain water systems and mechanical equipment.
Engineering quality, monitoring, component design, and maintenance determine operational risk. A poorly managed cooling system can fail regardless of the heat-transfer medium.
There is also no guarantee that every rack will operate continuously at its nameplate power. Workload patterns, software scheduling, chip availability, and customer demand determine actual utilization.
Designers still need capacity for expected peaks. Overbuilding every part of the facility for a peak that rarely occurs can raise capital costs and reduce efficiency at lower loads.
Underbuilding creates the opposite problem. The facility might own expensive accelerators that cannot operate simultaneously at full performance.
These uncertainties complicate investment decisions. The IEA explicitly models multiple demand scenarios because AI adoption, efficiency, and infrastructure bottlenecks can move consumption in different directions.
Its high-efficiency case shows how hardware, software, and facility improvements can reduce electricity use while delivering the same digital services. Its higher-growth scenarios reflect faster adoption and fewer supply constraints.
Neither outcome is guaranteed. Companies are making large infrastructure commitments before they know the long-term revenue generated by each AI workload.
That creates a risk of stranded capacity. A facility optimized around one rack format or cooling standard may need costly changes if hardware designs shift.
The reverse risk is also real. Waiting for standards to stabilize can leave operators without capacity when customers need it.
SK hynix faces its own uncertainty. The company benefits if demand for HBM and dense AI systems continues growing. Its public analysis therefore comes from a participant in the supply chain, not a neutral observer.
Its conclusion should be evaluated against measured deployment data. Useful evidence would include real rack utilization, delivered computing work, cooling energy, component temperatures, downtime, and total facility costs.
The physical argument remains sound even if demand grows more slowly. Electrical and thermal limits determine what any installed system can sustain.
What remains uncertain is the amount of infrastructure the market will need, where it will be built, and whether the revenue from AI services will justify the investment.
Three Signals Will Show Whether the Bottleneck Is Easing
The next phase will be measured by energized capacity, operational cooling data, and efficiency under real workloads.
The first signal is the gap between announced data center projects and facilities that receive firm grid connections. Construction announcements show developer ambition, but energized megawatts reveal deployable capacity.
Watch utility interconnection timelines, transformer availability, substation construction, and project delays. Shorter timelines would support the view that power supply is catching up with AI investment.
Persistent delays would strengthen the opposite conclusion. They would show that processor production can expand faster than the infrastructure needed to use those processors.
The second signal is operating data from dense liquid-cooled racks. Vendors already publish design specifications, but operators need evidence from sustained production workloads.
Relevant measures include coolant temperatures, pumping energy, rack uptime, maintenance requirements, component failure rates, and computing output per unit of facility power.
Greater standardization would also matter. Common connectors, thermal interfaces, and operating ranges could reduce integration costs and make hardware upgrades easier.
Fragmented requirements would slow adoption, especially in colocation facilities that must support equipment from many customers and vendors.
The third signal is whether AI service demand grows faster than efficiency. The IEA expects substantial improvement in the energy required for individual tasks, but total consumption depends on usage volume and workload complexity.
More efficient inference will ease infrastructure pressure only if computing demand grows at a manageable rate. Agents, video generation, and longer reasoning workloads can absorb those gains.
Standardized disclosure would help customers and policymakers separate useful computing growth from avoidable waste. Companies currently report energy and emissions at broad organizational levels, which makes individual AI services difficult to assess.
The global electricity outlook projects that data center consumption will more than double by 2030. That forecast is not destiny, but it establishes the scale of the planning challenge.
SK hynix AI infrastructure analysis points to the right physical question. The next limit is not only how many accelerators manufacturers can produce. It is how many operators can power, cool, connect, and utilize effectively.
Developers and enterprise buyers should now ask where their AI capacity runs, how efficiently it uses accelerators, and what happens when demand increases. Procurement teams should evaluate electrical and cooling readiness alongside processor, memory, and model performance.
The most important benchmark may no longer fit on a chip specification sheet. It is the amount of reliable AI work that an entire facility can deliver from each constrained megawatt.



