Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wei Huang 0004

dblp:81/6685-4 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
1since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
15 papers
Energy-efficient computing · 45% High-performance computing · 12% GPUs and heterogeneous computing · 11%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
power management
1.162017
Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017
Ti-states: Processor power management in the temperature inversion region · MICRO 2016
Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015
Energy-efficient computing
thermal management
1.082019
Understanding the Impact of Socket Density in Density Optimized Servers · HPCA 2019
Ti-states: Processor power management in the temperature inversion region · MICRO 2016
Accurate, Pre-RTL Temperature-Aware Design Using a Parameterized, Geometric Thermal Model · IEEE Trans. Computers 2008
High-performance computing › supercomputing
exascale computing
0.922023
A Research Retrospective on AMD's Exascale Computing Journey · ISCA 2023
Design and Analysis of an APU for Exascale Computing · HPCA 2017
GPUs and heterogeneous computing
GPU power management
0.522017
Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017
Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.522017
Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017
PPEP: Online Performance, Power, and Energy Prediction Framework and DVFS Space Exploration · MICRO 2014
Cloud and datacenter computing › datacenter architecture
datacenter server architecture
0.412019
Understanding the Impact of Socket Density in Density Optimized Servers · HPCA 2019
Energy-efficient computing › thermal management
thermal-aware scheduling
0.412019
Understanding the Impact of Socket Density in Density Optimized Servers · HPCA 2019
Memory systems
3d-stacked memory
0.312017
Design and Analysis of an APU for Exascale Computing · HPCA 2017
GPUs and heterogeneous computing › heterogeneous architecture
accelerated processing unit
0.312017
Design and Analysis of an APU for Exascale Computing · HPCA 2017
Hardware accelerators and domain-specific architectures › accelerator integration
CPU-GPU integration
0.312017
Design and Analysis of an APU for Exascale Computing · HPCA 2017
Processor architecture and microarchitecture › multicore design
heterogeneous processor architecture
0.312017
Design and Analysis of an APU for Exascale Computing · HPCA 2017
Memory systems › DRAM › DRAM architecture
high bandwidth memory
0.312017
Design and Analysis of an APU for Exascale Computing · HPCA 2017
Integrated circuit design
low-power circuit design
0.212016
Ti-states: Processor power management in the temperature inversion region · MICRO 2016
Energy-efficient computing › voltage scaling
near-threshold voltage operation
0.212016
Ti-states: Processor power management in the temperature inversion region · MICRO 2016
Energy-efficient computing › power management › device power management
processor power management
0.212016
Ti-states: Processor power management in the temperature inversion region · MICRO 2016
GPUs and heterogeneous computing
GPU architecture
0.212015
Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015
Energy-efficient computing › thermal management
thermal-aware design
0.232008
Accurate, Pre-RTL Temperature-Aware Design Using a Parameterized, Geometric Thermal Model · IEEE Trans. Computers 2008
Many-core design from a thermal perspective · DAC 2008
Compact thermal modeling for temperature-aware design · DAC 2004
High-performance computing
supercomputing
0.212023
A Research Retrospective on AMD's Exascale Computing Journey · ISCA 2023
Performance modeling and evaluation
workload characterization
0.212014
PPEP: Online Performance, Power, and Energy Prediction Framework and DVFS Space Exploration · MICRO 2014
Memory systems
DRAM
0.112012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Memory systems
tiered memory
0.112012
Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall · IEEE Trans. Computers 2012
Energy-efficient computing
thermal modeling
0.132004
Temperature-aware microarchitecture: Modeling and implementation · ACM Trans. Archit. Code Optim. 2004
Compact thermal modeling for temperature-aware design · DAC 2004
Temperature-Aware Microarchitecture · ISCA 2003
Energy-efficient computing › power management › energy-efficient networking
network power management
0.112011
Power shifting in Thrifty Interconnection Network · HPCA 2011
Parallel and multicore computing
task scheduling
0.112019
Understanding the Impact of Socket Density in Density Optimized Servers · HPCA 2019
Energy-efficient computing › thermal management
dynamic thermal management
0.122004
Temperature-aware microarchitecture: Modeling and implementation · ACM Trans. Archit. Code Optim. 2004
Temperature-Aware Microarchitecture · ISCA 2003
Integrated circuit design › heterogeneous integration
chiplet-based design
0.112017
Design and Analysis of an APU for Exascale Computing · HPCA 2017
Performance modeling and evaluation
performance prediction
0.112017
Dynamic GPGPU Power Management Using Adaptive Model Predictive Control · HPCA 2017
Processor architecture and microarchitecture
many-core architecture
0.112008
Many-core design from a thermal perspective · DAC 2008
Processor architecture and microarchitecture
multicore design
0.112008
Many-core design from a thermal perspective · DAC 2008
Energy-efficient computing › power management
power capping
0.112014
PPEP: Online Performance, Power, and Energy Prediction Framework and DVFS Space Exploration · MICRO 2014

Methods — techniques the papers use, named apart from their topics

simulation · 0.5analytical modeling · 0.4performance prediction · 0.3model predictive control · 0.3architectural simulation · 0.3measurement-based analysis · 0.2per-core power model · 0.2cycles-per-instruction model · 0.2power model training · 0.1on-chip power sensing · 0.1
YearPublicationVenuePosition
2023 A Research Retrospective on AMD's Exascale Computing Journey
abstract
The pace of advancement of the top-end supercomputers historically followed an exponential curve similar to (and driven in part by) Moore's Law. Shortly after hitting the petaflop mark, the community started looking ahead to the next milestone: Exascale. However, many obstacles were already looming on the horizon, such as the slowing of Moore's Law, and others like the end of Dennard Scaling had already arrived. Anticipating significant challenges for the overall high-performance computing (HPC) community to achieve the next 1000x improvement, the U.S. Department of Energy (DOE) launched the Exascale Computing Program to enable and accelerate fundamental research across the many technologies needed to achieve exascale computing.
Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang 0004, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, René van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov
ISCA34
2019 Understanding the Impact of Socket Density in Density Optimized Servers
abstract
The increasing demand for computational power has led to the creation and deployment of large-scale data centers. During the last few years, data centers have seen improvements aimed at increasing computational density - the amount of throughput that can be achieved within the allocated physical footprint. This need to pack more compute in the same physical space has led to density optimized server designs. Density optimized servers push compute density significantly beyond what can be achieved by blade servers by using innovative modular chassis based designs. This paper presents a comprehensive analysis of the impact of socket density on intra-server thermals and demonstrates that increased socket density inside the server leads to large temperature variations among sockets due to inter-socket thermal coupling. The paper shows that traditional chip-level and data center-level temperature-aware scheduling techniques do not work well for thermally-coupled sockets. The paper proposes new scheduling techniques that account for the thermals of the socket a task is scheduled on, as well as thermally coupled nearby sockets. The proposed mechanisms provide 2.5% to 6.5% performance improvements across various workloads and as much as 17% over traditional temperature-aware schedulers for computation-heavy workloads.
Manish Arora, Matt Skach, Wei Huang 0004, Xudong An, Jason Mars, Lingjia Tang, Dean M. Tullsen
HPCA3
2017 Dynamic GPGPU Power Management Using Adaptive Model Predictive Control
abstract
Modern processors can greatly increase energy efficiency through techniques such as dynamic voltage and frequency scaling. Traditional predictive schemes are limited in their effectiveness by their inability to plan for the performance and energy characteristics of upcoming phases. To date, there has been little research exploring more proactive techniques that account for expected future behavior when making decisions. This paper proposes using Model Predictive Control (MPC) to attempt to maximize the energy efficiency of GPU kernels without compromising performance. We develop performance and power prediction models for a recent CPU-GPU heterogeneous processor. Our system then dynamically adjusts hardware states based on recent execution history, the pattern of upcoming kernels, and the predicted behavior of those kernels. We also dynamically trade off the performance overhead and the effectiveness of MPC in finding the best configuration by adapting the horizon length at runtime. Our MPC technique limits performance loss by proactively spending energy on the kernel iterations that will gain the most performance from that energy. This energy can then be recovered in future iterations that are less performance sensitive. Our scheme also avoids wasting energy on low-throughput phases when it foresees future high-throughput kernels that could better use that energy. Compared to state-of-the-practice schemes, our approach achieves 24.8% energy savings with a performance loss (including MPC overheads) of 1.8%. Compared to state-of-the-art history-based schemes, our approach achieves 6.6% chip-wide energy savings while simultaneously improving performance by 9.6%.
Abhinandan Majumdar, Leonardo Piga, Indrani Paul, Joseph L. Greathouse, Wei Huang 0004, David H. Albonesi
HPCA5
2017 Design and Analysis of an APU for Exascale Computing
abstract
The challenges to push computing to exaflop levels are difficult given desired targets for memory capacity, memory bandwidth, power efficiency, reliability, and cost. This paper presents a vision for an architecture that can be used to construct exascale systems. We describe a conceptual Exascale Node Architecture (ENA), which is the computational building block for an exascale supercomputer. The ENA consists of an Exascale Heterogeneous Processor (EHP) coupled with an advanced memory system. The EHP provides a high-performance accelerated processing unit (CPU+GPU), in-package high-bandwidth 3D memory, and aggressive use of die-stacking and chiplet technologies to meet the requirements for exascale computing in a balanced manner. We present initial experimental analysis to demonstrate the promise of our approach, and we discuss remaining open research challenges for the community.
Thiruvengadam Vijayaraghavan, Yasuko Eckert, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Bradford M. Beckmann, William C. Brantley, Joseph L. Greathouse, Wei Huang 0004, Arun Karunanithi, Onur Kayiran, Mitesh R. Meswani, Indrani Paul, Matthew Poremba, Steven E. Raasch, Steven K. Reinhardt, Greg Sadowski, Vilas Sridharan
HPCA9
2016 A Case for Criticality Models in Exascale Systems
abstract
Performance variation is a significant problem for large scale HPC systems and will increase on future exascale systems. In this work, we show that performance variation impacts the performance and energy efficiency of contemporary large-scale computing systems in highly temporally inconsistent ways. We thus present a case for criticality models, a learning based mechanism that allows a system to generate holistic models of performance variation as it occurs during application runtime. Criticality models are designed to provide a mechanism by which applications can detect performance variation at runtime and take action to mitigate its effects. We present a promising preliminary analysis of criticality models on a small scale cluster. Our results demonstrate that models based on logistic regression scan accurately model criticality at this scale.
Brian Kocoloski, Leonardo Piga, Wei Huang 0004, Indrani Paul, Jack Lange
CLUSTER3
2016 Performance Boosting Opportunities under Communication Imbalance in Power-Constrained HPC Clusters
abstract
This paper provides a detailed message-passing interface (MPI) communication characterization across representative HPC applications. It further evaluates performance and power efficiency improvement opportunities. Specifically, it shows that the traditional approach of active polling while waiting for MPI messages is extremely power inefficient, especially under a constrained cluster-level power budget, where processors can only operate at some percentage of their labeled thermal design power (TDP) due to data center infrastructure limits. To mitigate the communication imbalance among different nodes, one can choose to power gate waiting processes and shift remaining power budget to processes that are in the critical execution paths, a technique we call Gate&Shift. With considerations of overheads from power gating and control-loop, Gate&Shift leads to performance improvement without additional power overhead. Gate&Shift is a reactive scheme that does not require prediction mechanisms. With the aid of real MPI traces and hardware measured power data from an HPC cluster, we show that (1) 1 ms control period for power-shifting is sufficient to achieve most potential performance gains, and (2) for a cluster with processors running at 65% of their labeled TDP, Gate&Shift can achieve 7%, 8.5% and 9% performance improvement for AMR Boxlib, Fill Boundary and Big FFT, respectively.
Leonardo Piga, Indrani Paul, Wei Huang 0004
ICPP3
2016 Ti-states: Processor power management in the temperature inversion region
abstract
Temperature inversion is a transistor-level effect that can improve performance when temperature increases. It has largely been ignored in the past because it does not occur in the typical operating region of a processor, but temperature inversion is becoming increasing important in current and future technologies. In this paper, we study temperature inversion's implications on architecture design, and power and performance management. We present the first public comprehensive measurement-based analysis on the effects of temperature inversion on a real processor, using the AMD A10-8700P processor as our system under test. We show that the extra timing margin introduced by temperature inversion can provide more than 5% Vddreduction benefit, and this improvement increases to more than 8% when operating in the near-threshold, low-voltage region. To harness this opportunity, we present Ti-states, a power management technique that sets the processor's voltage based on real-time silicon temperature to improve power efficiency. Ti-states lead to 6% to 12% measured power saving across a range of different temperatures compared to a fixed margin. As technology scales to FD-SOI and FinFET, we show there is an ideal operating temperature for various workloads to maximize the benefits of temperature inversion. The key is to counterbalance leakage power increase at higher temperatures with dynamic power reduction by the Ti-states. The projected optimal temperature is typically around 60°C and yields 8% to 9% chip power saving. The optimal high-temperature can be exploited to reduce design cost and runtime operating power for overall cooling. Our findings are important for power and thermal management in future chips and process technologies.
Yazhou Zu, Wei Huang 0004, Indrani Paul, Vijay Janapa Reddi
MICRO2
2015 Harmonia: balancing compute and memory power in high-performance GPUs
abstract
In this paper, we address the problem of efficiently managing the relative power demands of a high-performance GPU and its memory subsystem. We develop a management approach that dynamically tunes the hardware operating configurations to maintain balance between the power dissipated in compute versus memory access across GPGPU application phases. Our goal is to reduce power with minimal performance degradation.
Indrani Paul, Wei Huang 0004, Manish Arora, Sudhakar Yalamanchili
ISCA2
2014 PPEP: Online Performance, Power, and Energy Prediction Framework and DVFS Space Exploration
abstract
Performance, power, and energy (PPE) are critical aspects of modern computing. It is challenging to accurately predict, in real time, the effect of dynamic voltage and frequency scaling (DVFS) on PPE across a wide range of voltages and frequencies. This results in the use of reactive, iterative, and inefficient algorithms for dynamically finding good DVFS states. We propose PPEP, an online PPE prediction framework that proactively and rapidly searches the DVFS space. PPEP uses hardware events to implement both a cycles-per-instruction (CPI) model as well as a per-core power model in order to predict PPE across all DVFS states. We verify on modern AMD CPUs that the PPEP power model achieves an average error of 4.6% (2.8% standard deviation) on 152 benchmark combinations at 5 distinct voltage-frequency states. Predicting average chip power across different DVFS states achieves an average error of 4.2% with a 3.6% standard deviation. Further, we demonstrate the usage of PPEP by creating and evaluating a highly responsive power capping mechanism that can meet power targets in a single step. PPEP also provides insights for future development of DVFS technologies. For example, we find that it is important to carefully consider background workloads for DVFS policies and that enabling north bridge DVFS can offer up to 20% additional energy saving or a 1.4x performance improvement.
Junli Gu, Li Shen 0007, Wei Huang 0004, Joseph L. Greathouse, Zhiying Wang 0003
MICRO4
2013 Architectural implications of spatial thermal filtering
Karthik Sankaranarayanan, Brett H. Meyer, Wei Huang 0004, Robert J. Ribando, Hossein Haj-Hariri, Mircea R. Stan, Kevin Skadron
Integr.3
2012 Power-efficient time-sensitive mapping in heterogeneous systems
abstract
Heterogeneous systems that contain multiple types of resources, such as CPUs and GPUs, are becoming increasingly popular thanks to the potential of achieving high performance and energy efficiency. In such systems, the problem of data mapping and communication for time-sensitive applications while reducing power and energy consumption is more challenging, since applications may have varied data management and computing patterns on different types of resources. In this paper, we propose power-aware mapping techniques for CPU/GPU heterogeneous system that are able to meet applications' timing requirements while reducing power and energy consumption by applying DVFS on both CPUs and GPUs. We have implemented the proposed techniques in a real CPU/GPU heterogeneous system. Experimental results with several data analytics workloads show that compared to performance-driven mapping, our power-efficient mapping techniques can often achieve a reduction of more than 20% in power and energy consumption.
Cong Liu 0017, Jian Li 0059, Wei Huang 0004, Juan C. Rubio, William Evan Speight, Xiaozhu Lin
PACT3
2012 Accurate Fine-Grained Processor Power Proxies
abstract
There are not yet practical and accurate ways to directly measure core power in a microprocessor. This limits the granularity of measurement and control for computer power management. We overcome this limitation by presenting an accurate runtime per-core power proxy which closely estimates true core power. This enables new fine-grained microprocessor power management techniques at the core level. For example, cloud environments could manage and bill virtual machines for energy consumption associated with the core. The power model underlying our power proxy also enables energy-efficiency controllers to perform what-if analysis, instead of merely reacting to current conditions. We develop and validate a methodology for accurate power proxy training at both chip and core levels. Our implementation of power proxies uses on-chip logic in a high-performance multi-core processor and associated platform firmware. The power proxies account for full voltage and frequency ranges, as well as chip-to-chip process variations. For fixed clock frequency operation, a mean unsigned error of 1.8% for fine-grained 32ms samples across all workloads was achieved. For an interval of an entire workload, we achieve an average error of-0.2%. Similar results were achieved for voltage-scaling scenarios, too. We also present two sample applications of the power proxy: (1) per-core power billing for cloud computing services, and (2) simultaneous runtime energy saving comparisons among different power management policies without running each policy separately.
Wei Huang 0004, Charles Lefurgy, William Kuk, Alper Buyuktosunoglu, Michael S. Floyd, Karthick Rajamani, Malcolm Allen-Ware, Bishop Brock
MICRO1
2012 Tiered Memory: An Iso-Power Memory Architecture to Address the Memory Power Wall
abstract
Moore's Law improvement in transistor density is driving a rapid increase in the number of cores per processor. DRAM device capacity and energy efficiency are increasing at a slower pace, so the importance of DRAM power is increasing. This problem presents system designers with two nominal options when designing future systems: 1) decrease off-chip memory capacity and bandwidth per core or 2) increase the fraction of system power allocated to main memory. Reducing capacity and bandwidth leads to imbalanced systems with poor processor utilization for noncache-resident applications, so designers have chosen to increase DRAM power budget. This choice has been viable to date, but is fast running into a memory power wall. To address the looming memory power wall problem, we propose a novel iso-power tiered memory architecture that supports 2-3X more memory capacity for the same power budget as traditional designs by aggressively exploiting low-power DRAM modes. We employ two "tiers” of DRAM, a "hot” tier with active DRAM and a "cold” tier in which DRAM is placed in self-refresh mode. The DRAM capacity of each tier is adjusted dynamically based on aggregate workload requirements and the most frequently accessed data are migrated to the "hot” tier. This design allows larger memory capacities at a fixed power budget while mitigating the performance impact of using low-power DRAM modes. We target our solution at server consolidation scenarios where physical memory capacity is typically the primary factor limiting the number of virtual machines a server can support. Using iso-power tiered memory, we can run 3× as many virtual machines, achieving a 250 percent improvement in average aggregate performance, compared to a conventional memory design with the same power budget.
Kshitij Sudan, Karthick Rajamani, Wei Huang 0004, John B. Carter
IEEE Trans. Computers3
2011 Power shifting in Thrifty Interconnection Network
abstract
This paper presents two complementary techniques to manage the power consumption of large-scale systems with a packet-switched interconnection network. First, we propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically with little or no overhead by using inherent system events to timely trigger link activation or de-activation. Second, we propose Network Power Shifting (NPS) that dynamically shifts the power budget between the compute nodes and their corresponding network components. TIN activates and trains the links in the interconnection network, just-in-time before the network communication is about to happen, and thriftily puts them into a low-power mode when communication is finished, hence reducing unnecessary network power consumption. Furthermore, the compute nodes can absorb the extra power budget shifted from its attached network components and increase their processor frequency for higher performance with NPS. Our simulation results on a set of real-world workload traces show that TIN can achieve on average 60% network power reduction, with the support of only one low-power mode. When NPS is enabled, the two together can achieve 12% application performance improvement and 13% overall system energy reduction. Further performance improvement is possible if the compute nodes can speed up more and fully utilize the extra power budget reinvested from the thrifty network with more aggressive cooling support.
Jian Li 0059, Wei Huang 0004, Charles Lefurgy, Lixin Zhang 0002, Wolfgang E. Denzel, Richard R. Treumann, Kun Wang 0005
HPCA2
2010 Temperature-to-power mapping
abstract
Accurate power maps are useful for power model validation, process variation characterization, leakage estimation, and power optimization, but are hard to measure directly. Deriving power maps from measured thermal maps is the inverse problem of the power-to-temperature mapping, extensively studied through thermal simulation. Until recently this inverse heat conduction problem has received little attention in the microarchitecture research community. This paper first identifies the source of difficulties for the problem. The inverse mapping is then performed by applying constraints from microarchitecture-level observations. The inherent large sensitivity of the resultant power map is minimized through thermal map-filtering and constrained least-squares optimization. Choices of filter parameters and optimization constraints are investigated and their effects are evaluated. Furthermore, the paper highlights the differences between the grid and block modeling in the inverse mapping which were often ignored by previous schemes. The proposed methods reduce the mapping error by more than 10× compared to unoptimized solutions. To our best knowledge this is the first work to quantitatively evaluate and minimize the noise effect in the temperature to power mapping problem at the microarchitecture level for both grid and block mode, and for the steady and transient case.
Zhenyu Qi 0001, Brett H. Meyer, Wei Huang 0004, Robert J. Ribando, Kevin Skadron, Mircea R. Stan
ICCD3
2009 Differentiating the roles of IR measurement and simulation for power and temperature-aware design
abstract
In temperature-aware design, the presence or absence of a heatsink fundamentally changes the thermal behavior with important design implications. In recent years, chip-level infrared (IR) thermal imaging has been gaining popularity in studying thermal phenomena and thermal management, as well as reverse-engineering chip power consumption. Unfortunately, IR thermal imaging needs a peculiar cooling solution, which removes the heatsink and applies an IR-transparent liquid flow over the exposed bare die to carry away the dissipated heat. Because this cooling solution is drastically different from a normal thermal package, its thermal characteristics need to be closely examined. In this paper, we characterize the differences between two cooling configurations-forced air flow over a copper heatsink (AIR-SINK) and laminar oil flow over bare silicon (OIL-SILICON). For the comparison, we modify the HotSpot thermal model by adding the IR-transparent oil flow and the secondary heat transfer path through the package pins, hence modeling what the IR camera actually sees at runtime. We show that OIL-SILICON and AIR-SINK are significantly different in both transient and steady-state thermal responses. OIL-SILICON has a much slower short-term transient response, which makes dynamic thermal management less efficient. In addition, for OIL-SILICON, the direction of oil flow plays an important role by changing hot spot location, thus impacting hot spot identification and thermal sensor placement. These results imply that the power- and temperature-aware design process cannot just rely on IR measurements. Simulation and IR measurement are both needed and are complementary techniques.
Wei Huang 0004, Kevin Skadron, Sudhanva Gurumurthi, Robert J. Ribando, Mircea R. Stan
ISPASS1
2008 Many-core design from a thermal perspective
abstract
Air cooling limits have been a major design challenge in recent years for integrated circuits. Multi-core exacerbates thermal challenges because power scales with the number of cores, but also creates new opportunities for temperature-aware design, because multi-core designs offer more design parameters than single-core designs. This paper investigates the relationship between core size and on-chip hot spot temperature and shows that with the same power density, smaller cores are cooler than larger cores due to a spatial low-pass filtering effect of temperature. This phenomenon suggests that designs exploiting low-pass filtering can dissipate more power within the same cooling budget than contemporary designs.
Wei Huang 0004, Mircea R. Stan, Karthik Sankaranarayanan, Robert J. Ribando, Kevin Skadron
DAC1
2008 Accurate, Pre-RTL Temperature-Aware Design Using a Parameterized, Geometric Thermal Model
abstract
Preventing silicon chips from negative, even disastrous thermal hazards has become increasingly challenging these days; considering thermal effects early in the design cycle is thus required. To achieve this, an accurate yet fast temperature model together with an early-stage, thermally optimized, design flow are needed. In this paper, we present an improved block-based compact thermal model (HotSpot 4.0) that automatically achieves good accuracy even under extreme conditions. The model has been extensively validated with detailed finite-element thermal simulation tools. We also show that properly modeling package components and applying the right boundary conditions are crucial to making full-chip thermal models like HotSpot accurately resemble what happens in the real world. Ignoring or over-simplifying package components can lead to inaccurate temperature estimations and potential thermal hazards that are costly to fix in later designs stages. Such a full-chip and package thermal model can then be incorporated into a thermally optimized design flow where it acts as an efficient communication medium among computer architects, circuit designers and package designers in early microprocessor design stages, to achieve early and accurate design decisions and also faster design convergence. For example, the temperature-leakage interaction can be readily analyzed within such a design flow to predict potential thermal hazards such as thermal runaway.
Wei Huang 0004, Karthik Sankaranarayanan, Kevin Skadron, Robert J. Ribando, Mircea R. Stan
IEEE Trans. Computers1
2007 Interconnect Lifetime Prediction for Reliability-Aware Systems
abstract
Thermal effects are becoming a limiting factor in high-performance circuit design due to the strong temperature dependence of leakage power, circuit performance, IC package cost, and reliability. While many interconnect reliability models assume a constant temperature, this paper analyzes the effects of temporal and spatial thermal gradients on interconnect lifetime in terms of electromigration, and presents a physics-based dynamic reliability model which returns reliability equivalent temperature and current density that can be used in traditional reliability analysis tools. The model is verified with numerical simulations and reveals that blindly using the maximum temperature leads to too pessimistic lifetime estimation. Therefore, the proposed model not only increases the accuracy of reliability estimates, but also enables designers to reclaim design margin in reliability-aware design. In addition, the model is useful for improving the performance of temperature-aware runtime management by modeling system lifetime as a resource to be consumed at a stress-dependent rate
Zhijian Lu, Wei Huang 0004, Mircea R. Stan, Kevin Skadron, John C. Lach
IEEE Trans. Very Large Scale Integr. Syst.2
2006 HotSpot: A Compact Thermal Modeling Methodology for Early-Stage VLSI Design
abstract
This paper presents HotSpot-a modeling methodology for developing compact thermal models based on the popular stacked-layer packaging scheme in modern very large-scale integration systems. In addition to modeling silicon and packaging layers, HotSpot includes a high-level on-chip interconnect self-heating power and thermal model such that the thermal impacts on interconnects can also be considered during early design stages. The HotSpot compact thermal modeling approach is especially well suited for preregister transfer level (RTL) and presynthesis thermal analysis and is able to provide detailed static and transient temperature information across the die and the package, as it is also computationally efficient.
Wei Huang 0004, Shougata Ghosh, Sivakumar Velusamy, Karthik Sankaranarayanan, Kevin Skadron, Mircea R. Stan
IEEE Trans. Very Large Scale Integr. Syst.1
2005 Analytical Model for Sensor Placement on Microprocessors
abstract
Thermal management in microprocessors has become a major design challenge in recent years. Thermal monitoring through hardware sensors is important, and these sensors must be carefully placed on the chip to account for thermal gradients. In this paper, we present an analytical model that describes the maximum temperature differential between a hot spot and a region of interest based on their distance and processor packaging information. We also use a run-time thermal model, as an illustration of virtual sensors, and examine two benchmarks that exhibit highly concentrated thermal stress. We then use our analytical model to demonstrate the safety margins of the chip. Ultimately, the mathematical expression allows designers to obtain worst-case behavior of thermal heatup and select the optimal location of additional sensors.
Kyeong-Jae Lee, Kevin Skadron, Wei Huang 0004
ICCD3
2005 Monitoring Temperature in FPGA based SoCs
abstract
FPGA logic densities continue to increase at a tremendous rate. This has had the undesired consequence of increased power density, which manifests itself as higher on-die temperatures and local hotspots. Sophisticated packaging techniques have become essential to maintain the health of the chip. In addition to static techniques to reduce the temperature, dynamic thermal management techniques are essential. Such techniques rely on accurate on-chip temperature information. In this paper, we present the design of a system that monitors the temperatures at various locations on the FPGA. This system is composed of a controller interfacing to an array of temperature sensors that are implemented on the FPGA fabric. Such a system can be used to implement dynamic thermal management techniques. We cross validate the sensor readings with values obtained from HotSpot, a pre-RTL architectural level thermal modeling tool.
Sivakumar Velusamy, Wei Huang 0004, John C. Lach, Mircea R. Stan, Kevin Skadron
ICCD2
2005 The need for a full-chip and package thermal model for thermally optimized IC designs
abstract
Modeling and analyzing detailed die temperature with a full-chip thermal model at early design stages is important to discover and avoid potential thermal hazards. However, omitting important aspects of package details in a thermal model can result in significant temperature estimation errors. In this paper, we discuss the applications of an existing compact thermal model that models both die and package temperature details. As an example, a thermally selfconsistent leakage power calculation of a POWER4-like microprocessor design is presented. We then demonstrate the importance of including detailed package information in the thermal model by several examples considering the impact of thermal interface material (TIM), which glues the die to the heat spreader. The fact that detailed package information is needed to build an accurate compact thermal model implies a design flow, in which the chip- and package-level compact thermal model acts as a convenient medium for more productive collaborations among circuit designers, computer architects and package designers, leading to early and efficient evaluations of different design tradeoffs for an optimal design from a thermal point of view. Categories and Subject Descriptors:
Wei Huang 0004, Eric Humenay, Kevin Skadron, Mircea R. Stan
ISLPED1
2004 Compact thermal modeling for temperature-aware design
abstract
Thermal design in sub-100nm technologies is one of the major challenges to the CAD community. In this paper, we first introduce the idea of temperature-aware design. We then propose a compact thermal model which can be integrated with modern CAD tools to achieve a temperature-aware design methodology. Finally, we use the compact thermal model in a case study of microprocessor design to show the importance of using temperature as a guideline for the design. Results from our thermal model show that a temperature-aware design approach can provide more accurate estimations, and therefore better decisions and faster design convergence.
Wei Huang 0004, Mircea R. Stan, Kevin Skadron, Karthik Sankaranarayanan, Shougata Ghosh, Sivakumar Velusamy
DAC1
2004 Interconnect lifetime prediction under dynamic stress for reliability-aware design
abstract
Thermal effects are becoming a limiting factor in high-performance circuit design due to the strong temperature-dependence of leakage power, circuit performance, IC package cost and reliability. While many interconnect reliability models assume a constant temperature, this paper presents a physics-based model for estimating interconnect lifetime for any time-varying temperature/current profile. This model is verified with numerical solutions. With this model, we show that designers may be more aggressive with the temperature profiles that are allowed on a chip. In fact, our model reveals that when the temperature magnitude variation is small, average temperature (instead of worst-case temperature) can be used to accurately predict interconnect lifetime, allowing for significant design margin reclamation in reliability-aware design. Even when the variation of temperature magnitude is large, our model shows that using the maximum temperature is still too conservative for interconnect lifetime prediction. Therefore, our model not only increases the accuracy of reliability estimates, but also enables designers to consider more aggressive designs. This model is similarly useful for temperature-aware dynamic runtime management.
Zhijian Lu, Wei Huang 0004, John C. Lach, Mircea R. Stan, Kevin Skadron
ICCAD2
2004 Temperature-aware microarchitecture: Modeling and implementation
abstract
With cooling costs rising exponentially, designing cooling solutions for worst-case power dissipation is prohibitively expensive. Chips that can autonomously modify their execution and power-dissipation characteristics permit the use of lower-cost cooling solutions while still guaranteeing safe temperature regulation. Evaluating techniques for thisdynamic thermal management(DTM), however, requires a thermal model that is practical for architectural studies.This paper describesHotSpot, an accurate yet fast and practical model based on an equivalent circuit of thermal resistances and capacitances that correspond to microarchitecture blocks and essential aspects of the thermal package. Validation was performed using finite-element simulation. The paper also introduces several effective methods for DTM: "temperature-tracking" frequency scaling, "migrating computation" to spare hardware units, and a "hybrid" policy that combines fetch gating with dynamic voltage scaling. The latter two achieve their performance advantage by exploiting instruction-level parallelism, showing the importance of microarchitecture research in helping control the growth of cooling costs.Modeling temperature at the microarchitecture level also shows that power metrics are poor predictors of temperature, that sensor imprecision has a substantial impact on the performance of DTM, and that the inclusion of lateral resistances for thermal diffusion is important for accuracy.
Kevin Skadron, Mircea R. Stan, Karthik Sankaranarayanan, Wei Huang 0004, Sivakumar Velusamy, David Tarjan
ACM Trans. Archit. Code Optim.4
2003 Temperature-Aware Microarchitecture
abstract
With power density and hence cooling costs rising exponentially, processor packaging can no longer be designed for the worst case, and there is an urgent need for runtime processor-level techniques that can regulate operating temperature when the package's capacity is exceeded. Evaluating such techniques, however, requires a thermal model that is practical for architectural studies.This paper describes HotSpot, an accurate yet fast model based on an equivalent circuit of thermal resistances and capacitances that correspond to microarchitecture blocks and essential aspects of the thermal package. Validation was performed using finite-element simulation. The paper also introduces several effective methods for dynamic thermal management (DTM): "temperature-tracking" frequency scaling, localized toggling, and migrating computation to spare hardware units. Modeling temperature at the microarchitecture level also shows that power metrics are poor predictors of temperature, and that sensor imprecision has a substantial impact on the performance of DTM.
Kevin Skadron, Mircea R. Stan, Wei Huang 0004, Sivakumar Velusamy, Karthik Sankaranarayanan, David Tarjan
ISCA3