← Back to all insights
AI Racks Hit a Thermal Ceiling: How 1kW-Plus TDP Reshapes Data-Center Part Structure in 2026

Published on: August 21, 2026

AI Racks Hit a Thermal Ceiling: How 1kW-Plus TDP Reshapes Data-Center Part Structure in 2026

Individual AI processors from NVIDIA, AMD and Google have crossed 1kW TDP, and rack-scale systems now draw hundreds of kilowatts. Researchers at KAIST and Seoul National University this week placed packaging thermals ahead of further performance gains as the primary bottleneck on AI infrastructure expansion. Liquid cooling penetration moves from roughly 33% in 2025 to 53% in 2026 and near 60% in 2027. For secondary-channel and industrial OEM sourcing, the consequence sits in cold plates, pumps, quick-disconnects and high-voltage DC power components rather than in compute silicon.

ai-demandsupply-chainmarketsupply-risk
Also available in:日本語·中文

For two years, build plans for AI server programmes have turned on a single question: whether the memory lands. HBM allocation, DDR5 lead times and enterprise SSD priority formed the critical path on most BOMs.

This week's public discussion pushed that assumption to its limit.

The binding constraint has moved. Individual AI processors from NVIDIA, AMD and Google now exceed 1kW TDP, and rack-scale systems draw hundreds of kilowatts. Seoul National University Professor Kim Sung-dong noted that the industry now prioritises thermal management over further performance gains. The weight of that statement lies in the ordering: heat has not merely become important, it has become the precondition for everything else shipping.

KAIST Professor Kim Joung-ho points the problem at 3D stacking. Higher stacks concentrate heat flux and lengthen the dissipation path. His lab uses an HBM Design AI Agent to search heat-dissipation structures, substituting AI-driven design automation and digital twins for manual iteration. The choice of method indicates the scale of the difficulty.

Three numerical anchors

  • Single-die TDP above 1kW: flagship AI silicon from NVIDIA, AMD and Google has all crossed the line.
  • Racks at hundreds of kilowatts: a single rack now approaches the total draw of a small legacy data centre.
  • Liquid cooling at roughly 33% in 2025, 53% in 2026 and near 60% in 2027: a doubling in two years, with the inflection already behind.

The third figure carries the most direct sourcing consequence. Moving from 33% to 53% means liquid cooling is the majority case in 2026 rack designs rather than an option. Quotation templates built on an air-cooled default will keep mispricing programmes through the second half of 2026.

Where the BOM widens

Once heat becomes the constraint, the critical path on an AI server BOM extends past compute and memory. Four categories warrant their own lead-time tracking:

  • Cold plates: dimensionally tied to package geometry, refreshed on the GPU cadence, with low reuse across generations.
  • Pumps and redundant pump assemblies: high reliability class, a short qualified supplier list.
  • Quick-disconnects: leak risk maps directly to rack downtime, and qualification cycles run long.
  • High-voltage DC power components: rack power in the hundreds of kilowatts is changing the distribution architecture itself, which introduces busbar and protection parts as a new category.

These parts share a profile: low unit cost, no shipment without them, and mostly outside the coverage of traditional semiconductor channels. Programme management that treats memory lead time as the sole critical path no longer matches 2026 rack projects.

Three mitigation paths

CPO (co-packaged optics) moves optical transceivers into the processor package and substitutes optical signalling for copper, which is a thermal measure in itself. SK hynix published a CPO roadmap in Nature Electronics on 20 August with quantified targets: over 100 Tb/s per node, sub-1 pJ/bit efficiency and chip-to-chip latency below 10ns, spanning 2D packaging, 2.5D interposers and 3D heterogeneous integration.

STCO (system-technology co-optimization) pulls the thermal constraint forward into system design rather than optimising any single component.

New memory tiers: HBF (stacked NAND) and HBS (stacked SRAM) are expected alongside HBM. Samsung's zHBM claims 10x density versus HBM5, a 3x efficiency gain and a thermal-resistance reduction above 50%.

All three share a timing problem. The CPO roadmap spans three packaging generations, and zHBM, HBF and HBS sit in a next-generation product sequence. None of them changes 2026 or 2027 build schedules.

The supply-side cross-check

One set of figures belongs alongside this. Semiconductor equipment lead times now run 12 to 24 months: 12 months for conventional etch and deposition, over 18 months for high-end back-end tools, 24 months for RF power supplies. Heat is partly an engineering problem, but advanced packaging capacity itself sits under the equipment constraint, which holds the mitigation paths to the same curve.

In the same week, Samsung HBM4 yield rose from below 60% at February mass-production start to approximately 80%, and a KRW 6 trillion Onyang HBM fab was scheduled to break ground in September. The yield gain is the portion deliverable within 2026; the fab output falls after 2029. Supply-side increments and thermal-side constraints sit on two different timelines.

The practical reading

The bottleneck list on AI server programmes changed structurally in 2026. Memory remains tight and is no longer the sole critical path. Cold plates, pumps, quick-disconnects and high-voltage DC parts without their own lead-time entry on the BOM are the most underestimated risk category at present.

A build plan treating compute-component lead time as the only bottleneck stopped holding this week.