Ziyue (Alvin) Liu
DAC 2023

DeepOHeat: Operator Learning-based Ultra-fast Thermal Simulation in 3D-IC Design

Ziyue Liu1, Yixing Li2, Jing Hu2, Xinling Yu1, Shinyu Shiau2, Xin Ai2, Zhiyu Zeng2, Zheng Zhang1

1University of California, Santa Barbara2Cadence Design Systems

In short

Thermal design of a 3D chip needs many heat simulations, and each finite-element run takes minutes to hours. DeepOHeat learns the operator that maps a design’s configuration, such as its power map and boundary conditions, to its full 3D temperature field, and it is trained on the heat equation alone, with no simulation data. For designs it has not seen, it matches the commercial solver Celsius 3D to within 0.16% mean error, 1000× to 300000× faster.

Design configuration
Top power map u1sampled on a 21 × 21 grid
Bottom heat transfer hbone value for a uniform surface
Other boundary conditionseach as one more input
A point y in the chip(y1, y2, y3)
Each configuration is a function, not a single parameter, so a new design means new input functions.
Multi-input DeepONet
Branch net 1u1 → b1, a q-vector
Branch net 2hb → b2
More branch netsone for each further input
Trunk netFourier features of y → t
Element-wise product, then sum: T(y) = Σj b1j b2j ··· tj
Physics-informed loss
Heat equationk∇2T + qV = 0 inside the chip
Convection surfaces−k ∂T/∂n = h(T − Tamb)
Top surface−k ∂T/∂n equals the power map
Other surfacestheir own boundary conditions
Loss: the sum of these residuals on randomly drawn designs, from automatic differentiation. No simulation data.
A trained model gives the temperature at any point of an unseen design in one forward pass. Other configurations, such as a 3D power map or a non-uniform surface, are encoded the same way.
300,000×
faster than Celsius 3D with one V100 GPU, and 3,000× on the same CPU
0.16%
largest mean error over ten unseen power maps
0
simulations needed for training: the loss is the physics itself

Abstract

Thermal issue is a major concern in 3D integrated circuit (IC) design. Thermal optimization of 3D IC often requires massive expensive PDE simulations. Neural network-based thermal prediction models can perform real-time prediction for many unseen new designs. However, existing works either solve 2D temperature fields only or do not generalize well to new designs with unseen design configurations (e.g., heat sources and boundary conditions). In this paper, for the first time, we propose DeepOHeat, a physics-aware operator learning framework to predict the temperature field of a family of heat equations with multiple parametric or non-parametric design configurations. This framework learns a functional map from the function space of multiple key PDE configurations (e.g., boundary conditions, power maps, heat transfer coefficients) to the function space of the corresponding solution (i.e., temperature fields), enabling fast thermal analysis and optimization by changing key design configurations (rather than just some parameters). We test DeepOHeat on some industrial design cases and compare it against Celsius 3D from Cadence Design Systems. Our results show that, for the unseen testing cases, a well-trained DeepOHeat can produce accurate results with 1000× to 300000× speedup.

Learn the solution operator, not one solution

In steady state, the temperature T of a chip with conductivity k and internal power qV follows the heat equation, and each exposed surface adds a boundary condition: a fixed temperature, a fixed heat flux (a 2D power map is one), an insulated surface, or convection with a heat transfer coefficient h.

k ∇2T + qV = 0−k ∂T/∂n = h (T − Tamb)  on a convection surface

Power maps and heat transfer coefficients are functions, and changing them changes the problem itself, so a model that takes a few design parameters cannot cover them. Earlier neural models either predict only 2D fields or do not carry over to new configurations. DeepOHeat learns the operator Gθ from the configuration functions to the temperature field, so one trained model serves every design drawn from the same family.

A multi-input DeepONet

Each configuration is sampled at fixed points and sent to its own branch net: a 21 × 21 power map becomes 441 numbers, and a uniform heat transfer coefficient a single one. A point y in the chip goes to a trunk net whose first layer maps it to Fourier features, which helps with sharp temperature changes. The temperature at y is the sum of the element-wise product of all the nets’ outputs.

Trained by the physics

A single finite-element run of a complex chip can take hours, so collecting enough simulations to train on is not practical. DeepOHeat instead minimizes the residual of the heat equation inside the chip plus the residual of every boundary condition and power map on its surface, computed with automatic differentiation for randomly drawn configurations. Training needs no solver output at all; the solver is only used to test.

Results

Two industrial test cases, each a single cuboid chip of about 1 mm × 1 mm × 0.5 mm, compared point by point with Celsius 3D, Cadence’s finite-element solver, on configurations the model never saw in training.

Unseen power maps

The top surface carries a 2D power map; the sides are insulated and the bottom cools by convection. DeepOHeat trains for 10,000 iterations, 10 hours on one V100, on random smooth power maps drawn from a Gaussian random field. It is then tested on ten block-style power maps from Celsius 3D, from a uniform map to an irregular one with many small heat sources. Its fields match the solver’s closely: the mean error stays between 0.02% and 0.16%, and the peak error between 0.10% and 1.00%. On the most irregular map, it slightly overestimates the temperature between the small heat sources.

Ten unseen top-surface power maps p1 to p10, from a uniform map to an irregular one with many small heat sources. For each map, the 3D temperature field from Celsius 3D and from DeepOHeat look the same, and the error maps below them stay under 0.01, with the largest errors for p9.

Scroll sideways to see all ten maps.

From top: the ten unseen power maps, p1 to p10 from left to right; the temperature field from Celsius 3D and from DeepOHeat, each with its range in K; and the error between them. Figure 3 of the paper.
Error against Celsius 3D (%)
hollow: mean absolute percentage error · filled: peak
p1
0.030.10
p2
0.030.20
p3
0.020.24
p4
0.050.38
p5
0.140.52
p6
0.040.49
p7
0.130.71
p8
0.070.66
p9
0.161.00
p10
0.080.40
00.250.50.751
Show the numbers
Power mapMean error (%)Peak error (%)
p10.030.10
p20.030.20
p30.020.24
p40.050.38
p50.140.52
p60.040.49
p70.130.71
p80.070.66
p90.161.00
p100.080.40

Heat transfer coefficients as inputs

A second model takes the heat transfer coefficients of the top and bottom surfaces as two inputs, each drawn from 333.33 to 1,000 W/m2K in training, with a thin layer of volumetric power inside the chip. After 5,000 iterations, about 2 hours, it matches Celsius 3D on unseen pairs to within 0.032% mean error. These coefficients change the field only slightly, and DeepOHeat still follows the change.

Top and bottom HTC (W/m2K)Mean errorPeak error
1,000 and 333.330.032%0.043%
500 and 5000.011%0.025%

Speed

One Celsius 3D run takes about 5 minutes for the power-map case and 2 minutes for the heat-transfer case. A DeepOHeat prediction takes 0.1 s on the same CPU and 0.001 s on a V100 GPU. For larger designs the solver’s cost grows, while DeepOHeat’s stays the same.

Test caseCelsius 3DXeon Gold 6148DeepOHeatsame CPUDeepOHeatTesla V100
Power map≈ 5 min0.1 s3,000×0.001 s300,000×
Heat transfer≈ 2 min0.1 s1,200×0.001 s120,000×
Time for one temperature field, and the speedup over Celsius 3D. Training takes 10 hours (power maps, one V100) and about 2 hours (heat transfer), once per family of designs.

Scope and follow-up work

Citation

BibTeX
@inproceedings{liu2023deepoheat,
  author    = {Liu, Ziyue and Li, Yixing and Hu, Jing and Yu, Xinling and Shiau, Shinyu and Ai, Xin and Zeng, Zhiyu and Zhang, Zheng},
  title     = {{DeepOHeat}: Operator Learning-based Ultra-fast Thermal Simulation in {3D-IC} Design},
  booktitle = {2023 60th ACM/IEEE Design Automation Conference (DAC)},
  pages     = {1--6},
  year      = {2023},
  address   = {San Francisco, CA, USA},
  publisher = {IEEE},
  doi       = {10.1109/DAC56929.2023.10247998},
  url       = {https://ieeexplore.ieee.org/document/10247998}
}