IC-ThermBench: An Open, Progressive Benchmark for Generalizable 2.5D/3D-IC Thermal Learning
This paper introduces IC-ThermBench, an open and progressive benchmark that unifies diverse 2.5D/3D-IC thermal tasks and evaluation protocols to enable fair, reproducible comparison of learning-based thermal models while revealing significant performance degradation in cross-package out-of-distribution scenarios that can be mitigated through target-domain adaptation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Heat is the silent enemy of modern computing. As computer chips become smaller and more powerful, they pack more energy into tiny spaces, generating intense heat that can warp delicate circuits or cause a system to shut down. To prevent this, engineers must predict exactly where hot spots will form before a chip is ever built. For decades, the only reliable way to do this was to run slow, complex physics simulations for every single design change. Recently, researchers have tried to speed this up by teaching artificial intelligence to guess the temperature patterns instead, hoping these digital models could learn the rules of heat flow and offer instant answers. However, a major problem has held this field back: without a common yardstick, it is impossible to know if one new AI model is truly better than another, because everyone has been testing their creations on different sets of data with different rules.
To solve this confusion, a team of researchers has built a new, open standard called IC-ThermBench, designed to test how well these AI models can handle the messy reality of chip design. They created a massive, shared library of fifty thousand simulated chip scenarios, covering everything from simple, fixed designs to complex, ever-changing layouts where materials, cooling conditions, and physical shapes all vary. The researchers then put eight different AI models through a series of increasingly difficult tests. They started with familiar designs the models had seen before, then asked the models to predict heat for new arrangements of the same basic parts, and finally, they challenged the models with entirely new types of chip packages they had never encountered. The results revealed a clear pattern: while the models could handle gradual changes in design and material with only a slight drop in accuracy, they struggled immensely when faced with a completely new type of chip structure. In these unfamiliar situations, the models' predictions became wildly inaccurate, with errors jumping from less than one degree to nearly sixteen degrees. However, the study also found that giving the models just a tiny amount of new information—specifically, temperature data from only ten examples of the new chip type—allowed them to recover most of their lost accuracy, bringing errors back down to a manageable level.
The core of this work is the realization that being good at predicting heat for one type of chip does not automatically make an AI good at predicting heat for all chips. The researchers organized their tests into five distinct levels of difficulty. The first level tested the models on designs they had already studied, serving as a baseline. The next three levels introduced changes that a model might see in the real world: shifting the position of tiny computing blocks, swapping out the materials that fill the gaps between them, and altering the external cooling conditions. In these scenarios, the models performed reasonably well, with their errors creeping up slowly as the variations became more complex. This suggested that learning to handle a wider range of known physical conditions is a manageable task for these systems.
The true test came at the final level, where the researchers presented the models with five completely new chip packages that were entirely absent from the training data. This is akin to asking a chef who has mastered French cuisine to instantly cook a perfect meal from a cuisine they have never seen, using ingredients they do not recognize. The results were stark. The models that had performed well on the previous levels suddenly failed, with their prediction errors skyrocketing. The best model, which had been off by less than one degree in the earlier tests, missed the mark by nearly sixteen degrees on these new structures. This sharp drop in performance proved that the ability to generalize within a known family of designs is fundamentally different from the ability to adapt to a completely new system. The study explicitly ruled out the idea that simply making the AI models larger or more complex would solve this problem; even the most massive models in the test did not outperform the smaller, more efficient ones when faced with these unseen challenges.
The researchers also explored whether these models could be quickly taught to handle the new chip types. They found that by showing the models just ten labeled examples of the new chip's temperature behavior, the models could adapt rapidly. With this small amount of new information, the error rates dropped dramatically, returning to levels close to those seen in the earlier, easier tests. This suggests that while these AI systems cannot magically guess the behavior of a brand-new chip from scratch, they can learn to do so with very little guidance. The study concludes that the path forward for thermal modeling is not just about building bigger AI, but about creating systems that can be efficiently adapted to new designs. By providing a shared, rigorous testing ground, this new benchmark allows engineers to see exactly where these models succeed and where they fail, ensuring that future tools for chip design are built on a foundation of reliable, reproducible science rather than isolated, uncomparable experiments.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.