Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hai Wang 0002

dblp:59/3767-2 · DBLP profile ↗
← Back
41ranked-venue papers
18as first author
4since 2021 · last 2026
0000-0002-4003-2758ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 16 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Energy-efficient computing · 49% Processor architecture and microarchitecture · 22% Integrated circuit design · 10%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
thermal management
1.342021
Leakage-Aware Predictive Thermal Management for Multicore Systems Using Echo State Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
STREAM: Stress and Thermal Aware Reliability Management for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
A Fast Leakage-Aware Full-Chip Transient Thermal Estimation Method · IEEE Trans. Computers 2018
Processor architecture and microarchitecture
multicore design
0.722022
DBP: Distributed Power Budgeting for Many-Core Systems in Dark Silicon · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
GDP: A Greedy Based Dynamic Power Budgeting Method for Multi/Many-Core Systems in Dark Silicon · IEEE Trans. Computers 2019
Energy-efficient computing › power management
power budgeting
0.612022
DBP: Distributed Power Budgeting for Many-Core Systems in Dark Silicon · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Energy-efficient computing
power management
0.612022
DBP: Distributed Power Budgeting for Many-Core Systems in Dark Silicon · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Processor architecture and microarchitecture › microprocessor
3d microprocessors
0.512021
Runtime Performance Optimization of 3-D Microprocessors in Dark Silicon · IEEE Trans. Computers 2021
Memory systems
cache management
0.512021
Runtime Performance Optimization of 3-D Microprocessors in Dark Silicon · IEEE Trans. Computers 2021
Processor architecture and microarchitecture
chip multiprocessor
0.512021
Runtime Performance Optimization of 3-D Microprocessors in Dark Silicon · IEEE Trans. Computers 2021
Energy-efficient computing › thermal management
dynamic thermal management
0.412020
Leakage-Aware Predictive Thermal Management for Multicore Systems Using Echo State Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Integrated circuit design
3d integration
0.412019
STREAM: Stress and Thermal Aware Reliability Management for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Energy-efficient computing › power management › power budgeting
dynamic power budgeting
0.412019
GDP: A Greedy Based Dynamic Power Budgeting Method for Multi/Many-Core Systems in Dark Silicon · IEEE Trans. Computers 2019
Hardware reliability and fault tolerance
reliability management
0.412019
STREAM: Stress and Thermal Aware Reliability Management for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Energy-efficient computing
thermal modeling
0.322020
Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Leakage-Aware Predictive Thermal Management for Multicore Systems Using Echo State Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation
thermal analysis
0.322018
Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
A Fast Leakage-Aware Full-Chip Transient Thermal Estimation Method · IEEE Trans. Computers 2018
Integrated circuit design › 3d integration
3-D stacked IC
0.212014
Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Integrated circuit design › 3d integration
through-silicon via
0.212014
Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing › energy-constrained computing
dark silicon
0.112019
GDP: A Greedy Based Dynamic Power Budgeting Method for Multi/Many-Core Systems in Dark Silicon · IEEE Trans. Computers 2019
Electronic design automation
physical design
0.112018
A Fast Leakage-Aware Full-Chip Transient Thermal Estimation Method · IEEE Trans. Computers 2018
Performance modeling and evaluation › numerical algorithms
finite difference method
0.112014
Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Performance modeling and evaluation
simulation
0.112014
Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014

Methods — techniques the papers use, named apart from their topics

greedy algorithm · 0.9model predictive control · 0.8distributed power budget computing · 0.6distributed active core locating · 0.6recurrent neural network · 0.4echo state network · 0.4lifetime banking · 0.4dynamic voltage and frequency scaling · 0.4artificial neural network · 0.4adaptive model order reduction · 0.3
YearPublicationVenuePosition
2026 fastiSSM: Fast inference of state space model with online model approximation in frequency-domain
Yancheng Xie, Hai Wang 0002
Neural Networks3
2023 fastESN: Fast Echo State Network
abstract
Echo state networks (ESNs) are reservoir computing-based recurrent neural networks widely used in pattern analysis and machine intelligence applications. In order to achieve high accuracy with large model capacity, ESNs usually contain a large-sized internal layer (reservoir), making the evaluation process too slow for some applications. In this work, we speed up the evaluation of ESN by building a reduced network called the fast ESN (fastESN) and achieve an ESN evaluation complexity independent of the original ESN size for the first time. FastESN is generated using three techniques. First, the high-dimensional state of the original ESN is approximated by a low-dimensional state through proper orthogonal decomposition (POD)-based projection. Second, the activation function evaluation number is reduced through the discrete empirical interpolation method (DEIM). Third, we show the directly generated fastESN has instability problems and provide a stabilization scheme as a solution. Through experiments on four popular benchmarks, we show that fastESN is able to accelerate the sparse storage-based ESN evaluation with a high parameter compression ratio and a fast evaluation speed.
Hai Wang 0002, Xingyi Long, Xue-Xin Liu
IEEE Trans. Neural Networks Learn. Syst.1
2022 DBP: Distributed Power Budgeting for Many-Core Systems in Dark Silicon
abstract
Power budget is an important power constraint provided to guarantee the thermal reliability of an integrated system. In this work, we present DBP, a distributed power budgeting method, for dark silicon many-core systems. In DBP, there are two new techniques proposed to bring accurate and optimized power budgets in a distributed way. First, a distributed active core locating technique is developed to find an active core distribution that leads to a high-power budget. Second, a distributed power budget computing technique is introduced which computes the power budget for each active core accurately. Experiments show DBP outperforms the state-of-the-art power budgeting methods’ thermal safe power (TSP) and greedy dynamic power (GDP) on many-core dark silicon systems by providing a high and accurate power budget with low overhead and good scalability.
Hai Wang 0002, Wenjun He, Qinhui Yang, Xizhu Peng, He Tang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Runtime Performance Optimization of 3-D Microprocessors in Dark Silicon
abstract
Because the increasing power density is limited by the thermal constraint, multi-core integrated systems have stepped into the dark silicon era recently, meaning not all parts of the system can be powered on at the same time. Dark silicon effects are, especially severe for 3-D microprocessors due to the even higher power density caused by the stacked structures, which greatly limit the system performances. In this article, we propose a greedy based core-cache co-optimization algorithm to optimize the performance of 3-D microprocessors in dark silicon at runtime. The new method determines many runtime settings of the 3-D system on the fly, including the active core and cache bank positions, active cache bank number, and the voltage/frequency (V/f) level of each active core, which optimizes the performance of the 3-D microprocessor under thermal constraint. Because the core-cache settings are co-optimized in the 3-D space and the power budgets are computed dynamically according to the running state of the 3-D microprocessor, the new method leads to a higher system performance compared with the existing methods. Experiments on two 3-D microprocessors show the greedy-based core-cache co-optimization algorithm outperforms the state-of-the-art 3-D dark silicon microprocessor performance optimization method by achieving a higher processing throughput with guaranteed thermal safety.
Hai Wang 0002, Wei Li 0216, Wenjie Qi, Diya Tang, Letian Huang, He Tang 0003
IEEE Trans. Computers1
2020 Leakage-Aware Predictive Thermal Management for Multicore Systems Using Echo State Network
abstract
Leakage power is becoming significant in new generation IC chips. As leakage power is nonlinearly related to temperature, it is challenging to manage the thermal behavior of today's multicore systems, since thermal management becomes a nonlinear control problem. In this paper, a new predictive dynamic thermal management (DTM) method with neural network thermal model is proposed to naturally consider the inherent nonlinearity between leakage and temperature. We start with analyzing the problems of using recurrent neural network (RNN) to build the nonlinear thermal model, and point out that there is exploding gradient induced long-term dependencies problem, leading to large model prediction errors. Based on this analysis, we further propose to use echo state network (ESN), which is a special type of RNN, as the leakage-aware nonlinear thermal model. We theoretically and experimentally show that ESN achieves much higher accuracy by completely avoiding the long-term dependencies problem. On top of this nonlinear ESN thermal model, we propose a novel model predictive control (MPC) scheme called ESN MPC, which uses iterative steps to find the optimal future power recommendations for thermal management. Being able to consider the leakage-temperature nonlinear effects and equipped with advanced control technique, the new method achieves an overall high quality temperature management with smooth and accurate temperature tracking. The experimental results show the new method outperforms the state-of-the-art leakage-aware multicore DTM method in both temperature management quality and computing overhead.
Hai Wang 0002, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Compact Piecewise Linear Model Based Temperature Control of Multicore Systems Considering Leakage Power
abstract
Temperature control of the new-generation integrated multicore system is challenging. This is because the leakage power, which is significant in modern systems, is nonlinearly related to temperature, resulting in a complex nonlinear control problem in thermal management. In this article, a new dynamic thermal management (DTM) method with compact piecewise linear (PWL) model based predictive control is proposed to solve the nonlinear control problem. First, a compact PWL thermal model, which takes dynamic power as input, is built by combining multiple local compact linear thermal models expanded at several Taylor expansion points. These local compact linear thermal models are obtained by sampling-based model order reduction with high accuracy. Their Taylor expansion points are selected by a systematic scheme, which exploits the thermal behavior property of the multicore chips. Based on the compact PWL thermal model, a new predictive control method is proposed to compute the future power recommendation for DTM. By approximating the nonlinearity accurately with the compact PWL thermal model and being equipped with predictive control technique, the new DTM achieves an overall high quality temperature management with smooth and accurate temperature tracking. Experimental results show that the new method outperforms the linear model predictive control based method and the echo state network based predictive thermal management method in temperature management quality with lower computing overhead.
Hai Wang 0002, Liwen Hu 0005, Yang Nie, He Tang 0003
IEEE Trans. Ind. Informatics1
2019 Leakage-aware thermal management for multi-core systems using piecewise linear model based predictive control
abstract
Performing thermal management on new generation IC chips is challenging. This is because the leakage power, which is significant in today's chips, is nonlinearly related to temperature, resulting in a complex nonlinear control problem in thermal management. In this paper, a new dynamic thermal management (DTM) method with piecewise linear (PWL) thermal model based predictive control is proposed to solve the nonlinear control problem. First, a PWL thermal model is built by combining multiple local linear thermal models expanded at several Taylor expansion points. These Taylor expansion points are carefully selected by a systematic scheme which exploits the thermal behavior property of the IC chips. Based on the PWL thermal model, a new predictive control method is proposed to compute the future power recommendation for DTM. By approximating the nonlinearity accurately with the PWL thermal model and being equipped with predictive control technique, the new DTM can achieve an overall high quality temperature management with smooth and accurate temperature tracking. Experimental results show the new method outperforms the linear model predictive control based method in temperature management quality with negligible computing overhead.
Hai Wang 0002, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
ASP-DAC2
2019 GDP: A Greedy Based Dynamic Power Budgeting Method for Multi/Many-Core Systems in Dark Silicon
abstract
Dark silicon phenomenon is significant in today's multi/many-core systems manufactured using new generation technology. In order to enhance performance of dark silicon systems, power budget constrained dynamic optimizations are performed in various ways including dynamic voltage and frequency scaling (DVFS) and task scheduling. However, power budgets given by existing methods are generally over pessimistic, which greatly limit the capability of dynamic performance optimization methods. In order to resolve this problem, we propose a dynamic power budgeting method, called Greedy based Dynamic Power (GDP). Different from existing methods, which are steady state based and ignore active core distributions, GDP formulates the power budgeting problem as a thermal-constrained combinational power optimization problem. To efficiently solve this problem, we propose two new ideas: first, we transform the original power-optimization problem to an easier solving temperature-optimization problem; second, we employ a more efficient greedy based algorithm that finds a sub-optimal active core distribution which maximizes power budget. The new method can consider current temperature states and transient thermal effects, which were ignored by existing methods. Both theoretical studies and experimental results show that GDP outperforms existing methods by providing a higher and less pessimistic power budget with low computing cost and guaranteed thermal safety.
Hai Wang 0002, Diya Tang, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
IEEE Trans. Computers1
2019 STREAM: Stress and Thermal Aware Reliability Management for 3-D ICs
abstract
Accurate and fast reliability management is important for 3-D integrated circuits (3-D ICs) because of the severe on-chip thermal and reliability problems. However, due to the lack of stress information and difficulties in implementing management method for reliability, existing full-chip reliability management methods suffer from low management accuracy and high system performance degradation. In this paper, we propose a new stress and thermal aware reliability management method for 3-D ICs called STREAM. Unlike traditional methods which do not perform explicit stress analysis due to the large computing cost, STREAM employs an artificial neural network-based stress model to estimate stress accurately at runtime. In order to further improve the reliability management accuracy and improve the system performance, a lifetime estimator with lifetime banking technology and a specially designed lifetime model predictive control are integrated into the reliability management framework. Our numerical results show that STREAM performs the stress and thermal aware full-chip reliability management with both high accuracy and speed. It is able to boost the performance of 3-D ICs and outperforms the state-of-the-art 3-D IC reliability management method.
Hai Wang 0002, Darong Huang 0003, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Runtime Stress Estimation for Three-dimensional IC Reliability Management Using Artificial Neural Network
abstract
Heat dissipation and the related thermal-mechanical stress problems are the major obstacles in the development of the three-dimensional integrated circuit (3D IC). Reliability management techniques can be used to alleviate such problems and enhance the reliability of 3D IC. However, it is difficult to obtain the time-varying stress information at runtime, which limits the effectiveness of the reliability management. In this article, we propose a fast stress estimation method for runtime reliability management using artificial neural network (ANN). The new method builds ANN-based stress model by training offline using temperature and stress data. The ANN stress model is then used to estimate the important stress information, such as the maximum stress around each TSV, for reliability management at runtime. Since there are a variety of potential ANN structures to choose from for the ANN stress model, we analyze and test three ANN-based stress models with three major types of ANNs in this work: the normal ANN-based stress model, the ANN stress model with hand-crafted feature extraction, and the convolutional neural network–(CNN) based stress model. The structures of each ANN stress model and the functions of these structures in 3D IC stress estimation are demonstrated and explained. The new runtime stress estimation method is tested using the three ANN stress models with different layer configurations. Experiments show that the new method is able to estimate important stress information at extremely fast speed with good accuracy for runtime 3D IC reliability enhancement. Although all three ANN stress models show acceptable capabilities in runtime stress estimation, the CNN-based stress model achieves the best performance considering both stress estimation accuracy and computing overhead. Comparison with traditional method reveals that the new ANN-based stress estimation method is much more accurate with a slightly larger but still very small computing overhead.
Hai Wang 0002, Darong Huang 0003, Lang Zhang, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
ACM Trans. Design Autom. Electr. Syst.1
2018 A Fast Leakage-Aware Full-Chip Transient Thermal Estimation Method
abstract
Accurate and fast thermal estimation is important for the runtime thermal regulation of modern microprocessors due to excessive on-chip temperatures. However, due to the nonlinear relationship between the leakage power and temperature, full-chip thermal estimation methods suffer slow speed and scalability issue when the increasing static leakage power is considered. In this work, we propose a new fast leakage-aware full-chip thermal estimation method. Unlike traditional methods, which use iteration to handle the leakage-temperature nonlinearity dependency issue, the new method applies a dynamic linearization algorithm, which adaptively transforms the original nonlinear thermal model into a number of local linear thermal models. In order to further improve the thermal estimation efficiency, a specially-designed adaptive model order reduction method is integrated into the thermal estimation framework to generate local compact thermal models. Our numerical results show that the new method is able to accurately estimate full-chip transient temperature distribution by fully considering the nonlinear leakage-temperature dependency with fast speed. On different chips with core number ranging from 9 to 36, it achieved 85x to 589x speedup in average against traditional iteration based method, with average thermal estimation error to be around 0.2°C.
Hai Wang 0002, Jiachun Wan, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030, Keheng Huang, Zhenghong Zhang
IEEE Trans. Computers1
2018 Thermal-Sensor-Based Occupancy Detection for Smart Buildings Using Machine-Learning Methods
abstract
In this article, we propose a novel approach to detect the occupancy behavior of a building through the temperature and/or possible heat source information. The new method can be used for energy reduction and security monitoring for emerging smart buildings. Our work is based on a building simulation program, EnergyPlus, from the Department of Energy. EnergyPlus can model various time-series inputs to a building such as ambient temperature; heating, ventilation, and air-conditioning (HVAC) inputs; power consumption of electronic equipment; lighting; and number of occupants in a room, sampled each hour, and produce resulting temperature traces of zones (rooms). Two machine-learning-based approaches for detecting human occupancy of a smart building are applied herein, namely support vector regression (SVR) and recurrent neural network (RNN). Experimental results with SVR show that the four-feature model provides accurate detection rates, giving a 0.638 average error and 5.32% error rate, and the five-feature model delivers a 0.317 average error and 2.64% error rate. This indicates that SVR is a viable option for occupancy detection. In the RNN method, Elman’s RNN can estimate occupancy information of each room of a building with high accuracy. It has local feedback in each layer and, for a five-zone building, it is very accurate for occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum, considering ambient, room temperatures, and HVAC powers as detectable information. Without knowing HVAC powers, the estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5. Our article further shows that both methods deliver similar accuracy in the occupancy detection. But the SVR model is more stable for adding or removing features of the system, while the RNN method can deliver more accuracy when the features used in the model do not change a lot.
Hengyang Zhao, Qi Hua, Haibao Chen, Yaoyao Ye, Hai Wang 0002, Sheldon X.-D. Tan, Esteban Tlelo-Cuautle
ACM Trans. Design Autom. Electr. Syst.5
2017 A quantitative design methodology for high-speed interpolation/averaging ADCs
He Tang 0003, Albert Wang 0001, Hai Wang 0002
Integr.5
2017 Energy and Lifetime Optimizations for Dark Silicon Manycore Microprocessor Considering Both Hard and Soft Errors
abstract
In this paper, we propose a new energy and lifetime optimization techniques for emerging dark silicon manycore microprocessors considering both hard long-term reliability effects (hard errors) and transient soft errors, which have been studied less in the past. We consider a recently proposed physics-based electromigration (EM) reliability model to predict the EM-induced reliability. We employ both dynamic voltage and frequency scaling (DVFS) and dark silicon core state using ON/OFF switching action as the two control knobs. We show that on-chip power consumption has different (even contradicting) impacts on soft and hard reliability effects. This paper also shows that soft error should be mitigated by other techniques if aggressive low power and high long-term reliability are pursued. We focus on two optimization techniques for improving lifetime and reducing energy. To optimize EM-induced lifetime, we first apply the adaptive Q-learning-based method, which is suitable for dynamic runtime operation as it can provide cost-effective yet good solutions. The second lifetime optimization approach is the mixed-integer linear programming (MILP) method, which typically yields better solutions but at higher computational costs. To optimize the energy of a dark silicon chip subject to the both hard and soft reliability effects, power budgets, and performance limits, the Q-learning method has been applied as well. A large class of multithreaded applications is used as our benchmarks to validate and compare the proposed dynamic reliability management methods. Experimental results on a 64-core dark silicon chip show that the proposed DRM algorithm can effectively manage and optimize the lifetime of a dark silicon microprocessor under the given power budget and performance limit. Also, the proposed energy optimization can effectively manage and optimize energy consumption subject to both hard and soft-error rates, power budget, and performance limits as constraints. We also show that the under tightened power and performance constraints, we cannot satisfy both hard and soft errors at the same time as there is no simple tradeoff between performance/power and reliability in this case. Some other soft-error mitigation techniques are required in this case.
Taeyoung Kim 0001, Zeyu Sun 0001, Haibao Chen, Hai Wang 0002, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Thermal modeling for energy-efficient smart building with advanced overfitting mitigation technique
abstract
Building energy accounts large amount of the total energy consumption, and smart building energy control leads to high energy efficiency and significant energy savings. A compact and accurate building thermal model is important for designing the efficient energy control system. In this paper, we propose an accurate thermal behavior modeling technique for general and complicated buildings. This new modeling technique builds compact thermal model by system identification using temperature and power data obtained from EnergyPlus software, which can provide realistic temperature, weather and power data for buildings. In order to make the best use of data from EnergyPlus and avoid the overfitting problem associated with the system identification method, a cross-validation technique is employed to generate multiple thermal models to find the optimal model order. The final model is then generated by performing a regular system identification using the previously selected order. Experimental results from a case study of a 5-zone building have shown that the proposed method is able to find the optimal model order, and the building models built by the proposed method can achieve 1-3% average errors and less than 10-18% maximum errors for the estimation of zone temperatures for about a one year period.
Wandi Liu, Hai Wang 0002, Hengyang Zhao, Shujuan Wang, Haibao Chen, Yuzhuo Fu, Jian Ma 0002, Xin Li 0001, Sheldon X.-D. Tan
ASP-DAC2
2016 Dynamic reliability management for near-threshold dark silicon processors
abstract
In this article, we propose a new dynamic reliability management (DRM) techniques at the system level for emerging low power dark silicon manycore microprocessors operating in near-threshold region. We mainly consider the electromigration (EM) failures. To leverage the EM recovery effects, which was ignored in the past, at the system-level, we propose a new equivalent DC current model to consider recovery effects for general time-varying current waveforms so that existing compact EM model can be applied. The new equivalent DC current is calculated in two steps: firstly, the equivalent square waveform is calculated so that peak and terminal stresses are matched, secondly, the parameterized equivalent DC current is derived in terms of the parameters of the periodic fitted square waveforms from the first step. The new recovery EM model can allow EM-induced lifetime to be better managed at the system level. The system level energy optimization problem considering EM lifetime subject to power and performance constraints is framed by seeking the best dark silicon cores' voltage and on/off status. The resulting problem is solved by the State-Action-Reward-State-Action (SARSA) reinforcement learning algorithm. Experimental results on a 64-core near-threshold dark silicon processor show that the new equivalent EM DC currents can fully exhibit the recovery effects at the system-level so that trade-off between EM lifetime and energy/performance can be easily made. We further show that the proposed learning-based energy optimization can effectively manage and optimize energy subject to reliability, given power budget and performance limits. When the recovery effects are considered, the new optimization method can achieve 8.6× longer lifetime at the costs of 2.0× more energy and 3.3× more performance degradation.
Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Jagadeesh Gaddipati, Hai Wang 0002, Haibao Chen, Sheldon X.-D. Tan
ICCAD5
2016 Learning-based occupancy behavior detection for smart buildings
abstract
In this article, we propose a novel method to detect the occupancy behavior of a building through the temperature and/or possible heat source information, which can be used for energy reduction, security monitoring for emerging smart buildings. Our work is based on a realistic building simulation program, EnergyPlus, from Department of Energy. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). The new approach is based on a learning based approach in which a recurrent neutral network (RNN) is trained to detect the number of people in a room based on the room temperature and other information such as ambient temperature, and other related heat sources. We applied the Elman's recurrent neural network (ELNN), which has local feedbacks in each layer. We use an empirical formula to calculate the RNN layer number and layer size to configure RNN architecture to avoid overfitting and under-fitting problems. Experimental results from a case study of a 5-zone building show that ELNN can lead to very accurate occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum when we consider ambient, room temperatures and HVAC powers as detectable information. Without knowing HVAC powers, estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5.
Hengyang Zhao, Zhongdong Qi, Shujuan Wang, Kambiz Vafai, Hai Wang 0002, Haibao Chen, Sheldon X.-D. Tan
ISCAS5
2016 Parallel GMRES solver for fast analysis of large linear dynamic systems on GPU platforms
Sheldon X.-D. Tan, Hengyang Zhao, Xuexin Liu, Hai Wang 0002, Guoyong Shi
Integr.5
2016 Hierarchical Dynamic Thermal Management Method for High-Performance Many-Core Microprocessors
abstract
It is challenging to manage the thermal behavior of many-core microprocessors while still keeping them running at high performance since the control complexity increases as the core number increases. In this article, a novel hierarchical dynamic thermal management method is proposed to overcome this challenge. The new method employs model predictive control (MPC) with task migration and a DVFS scheme to ensure smooth control behavior and negligible computing performance sacrifice. In order to be scalable to many-core systems, the hierarchical control scheme is designed with two levels. At the lower level, the cores are spatially clustered into blocks, and local task migration is used to match current power distribution with the optimal distribution calculated by MPC. At the upper level, global task migration is used with the unmatched powers from the lower level. A modified iterative minimum cut algorithm is used to assist the task migration decision making if the power number is large at the upper level. Finally, DVFS is applied to regulate the remaining unmatched powers. Experiments show that the new method outperforms existing methods and is very scalable to manage many-core microprocessors with small performance degradation.
Hai Wang 0002, Jian Ma 0002, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Keheng Huang, Zhenghong Zhang
ACM Trans. Design Autom. Electr. Syst.1
2016 Statistical Rare-Event Analysis and Parameter Guidance by Elite Learning Sample Selection
abstract
Accurately estimating the failure region of rare events for memory-cell and analog circuit blocks under process variations is a challenging task. In this article, we propose a new statistical method, called EliteScope , to estimate the circuit failure rates in rare-event regions and to provide conditions of parameters to achieve targeted performance. The new method is based on the iterative blockade framework to reduce the number of samples, but consists of two new techniques to improve existing methods. First, the new approach employs an elite-learning sample-selection scheme, which can consider the effectiveness of samples and well coverage for the parameter space. As a result, it can reduce additional simulation costs by pruning less effective samples while keeping the accuracy of failure estimation. Second, the EliteScope identifies the failure regions in terms of parameter spaces to provide a good design guidance to accomplish the performance target. It applies variance-based feature selection to find the dominant parameters and then determine the in-spec boundaries of those parameters. We demonstrate the advantage of our proposed method using several memory and analog circuits with different numbers of process parameters. Experiments on four circuit examples show that EliteScope achieves a significant improvement on failure-region estimation in terms of accuracy and simulation cost over traditional approaches. The 16b 6T-SRAM column example also demonstrates that the new method is scalable for handling large problems with large numbers of process variables.
Taeyoung Kim 0001, Hosoon Shin, Sheldon X.-D. Tan, Xin Li 0001, Haibao Chen, Hai Wang 0002
ACM Trans. Design Autom. Electr. Syst.7
2016 GPU-Accelerated Parallel Sparse LU Factorization Method for Fast Circuit Analysis
abstract
Lower upper (LU) factorization for sparse matrices is the most important computing step for circuit simulation problems. However, parallelizing LU factorization on the graphic processing units (GPUs) turns out to be a difficult problem due to intrinsic data dependence and irregular memory access, which diminish GPU computing power. In this paper, we propose a new sparse LU solver on GPUs for circuit simulation and more general scientific computing. The new method, which is called GPU accelerated LU factorization (GLU) solver (for GPU LU), is based on a hybrid right-looking LU factorization algorithm for sparse matrices. We show that more concurrency can be exploited in the right-looking method than the left-looking method, which is more popular for circuit analysis, on GPU platforms. At the same time, the GLU also preserves the benefit of column-based left-looking LU method, such as symbolic analysis and columnlevel concurrency. We show that the resulting new parallel GPU LU solver allows the parallelization of all three loops in the LU factorization on GPUs. While in contrast, the existing GPU-based left-looking LU factorization approach can only allow parallelization of two loops. Experimental results show that the proposed GLU solver can deliver 5.71χ and 1.46x speedup over the single-threaded and the 16-threaded PARDISO solvers, respectively, 19.56x speedup over the KLU solver, 47.13x over the UMFPACK solver, and 1.47x speedup over a recently proposed GPU-based left-looking LU solver on the set of typical circuit matrices from the University of Florida (UFL) sparse matrix collection. Furthermore, we also compare the proposed GLU solver on a set of general matrices from the UFL, GLU achieves 6.38x and 1.12x speedup over the singlethreaded and the 16-threaded PARDISO solvers, respectively, 39.39x speedup over the KLU solver, 24.04x over the UMFPACK solver, and 2.35x speedup over the same GPU-based left-looking LU solver. In addition, comparison on self-generated RLC mesh networks shows a similar trend, which further validates the advantage of the proposed method over the existing sparse LU solvers.
Sheldon X.-D. Tan, Hai Wang 0002, Guoyong Shi
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Learning Based Compact Thermal Modeling for Energy-Efficient Smart Building Management: (invited)
abstract
In this article, we propose a new behavioral thermal modeling method for fast building performance analysis, which is critical for energy-efficient smart building control and management. The new approach is based on two recurrent neutral network architecture to obtain the compact nonlinear thermal models for complicated building. We start with a more realistic building simulation program, EnergyPlus, from Department of Energy, to model some practical buildings such as office buildings and data centers. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). In this work, we apply two recurrent neural network (RNN) architectures to build the non-linear compact thermal model of the building: one is non-linear state-space RNN architecture (NLSS), which has global feedbacks, and the other one is Elman's RNN architecture (ELNN), which has local feedbacks in each layer. We give a simple formula to calculate the RNN layer number, layer size to configure RNN architecture to avoid overfitting and underfitting problems. A cross-validation based training technique is further applied to improve predictable accuracy of models. Experimental results from a case study of three buildings show that ELNN and NLSS can both build very accurate building thermal models for the 2-zone and 5-zone building cases: both of them have average errors from around 1% to 1.5% for the two buildings. For the more complex 6-zone building case, ELNN outperforms NLSS with maximum errors 16% against 23%. But both methods have 2.2% average errors.
Hengyang Zhao, Daniel Quach, Shujuan Wang, Hai Wang 0002, Haibao Chen, Xin Li 0001, Sheldon X.-D. Tan
ICCAD4
2015 H-Matrix-Based Finite-Element-Based Thermal Analysis for 3D ICs
abstract
In this article, we propose an efficient finite-element-based (FE-based) method for both steady and transient thermal analyses of high-performance integrated circuits based on the hierarchical matrix ( H -matrix) representation. H -matrix has been shown to provide a data-sparse way to approximate the matrices and their inverses with almost linear-space and time complexities. In this work, we apply the H -matrix concept for solving heating diffusion problems modeled by parabolic partial differential equations (PDEs) based on the finite element method. We show that the matrix from a FE-based steady and transient thermal analysis can be represented by H -matrix without any approximation, and its inverse and Cholesky factors can be evaluated by H -matrix with controlled accuracy. We then show and prove that the memory and time complexities of the solver are bounded by O ( k 1 N log N ) and O ( k 1 2 N log 2 N ), respectively, where k 1 is a small quantity determined by accuracy requirements and N is the number of unknowns in the system. The comparison with existing product-quality LU solvers, CSPARSE and UMFPACK, on a number of 3D IC thermal matrices, shows that the new method is much more memory efficient than these methods, which however prevents CPU time comparison with those methods on large examples. But the proposed method can solve all the given thermal circuits with decent scalabilities, which shows good agreement with the predicted theoretical results.
Haibao Chen, Ying-Chi Li, Sheldon X.-D. Tan, Xin Huang 0003, Hai Wang 0002, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.5
2015 Task Migrations for Distributed Thermal Management Considering Transient Effects
abstract
In this brief, a new distributed thermal management scheme using task migrations based on a new temperature metric called effective initial temperature is proposed to reduce the on-chip temperature variance and the occurrence of hot spots for many-core microprocessors. The new temperature metric derived from frequency domain moment matching technique incorporates both initial temperature and other transient effects to make optimized task migration decisions, which leads to more effective reduction of hot spots in the experiments on a 100-core microprocessor than the existing distributed thermal management methods.
Zao Liu, Sheldon X.-D. Tan, Xin Huang 0003, Hai Wang 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Compact thermal modeling for packaged microprocessor design with practical power maps
Zao Liu, Sheldon X.-D. Tan, Hai Wang 0002, Yingbo Hua, Ashish Gupta 0007
Integr.3
2014 Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs
abstract
Thermal issue is the leading design constraint for 3-D stacked integrated circuits (ICs) and through silicon vias (TSVs) are used to effectively reduce the temperature of 3-D ICs. Normally, TSV is considered as a good thermal conductor in its vertical direction, and its vertical thermal resistance has been well modeled. However, lateral heat transfer of TSVs, which is also important, was largely ignored in the past. In this paper, we propose an accurate physics-based model for lateral thermal resistance of TSVs in terms of physical and material parameters, and study the conditions for model accuracy. For TSV arrays or farm, we show that the space or pitch between TSVs has a significant impact on TSV thermal behavior and should be properly considered in the TSV models. The proposed lateral thermal resistance model is fully compatible with the existing modeling approaches, and thus we could build a more accurate complete TSV thermal model. The new TSV thermal model can be easily integrated into a finite difference (FD) based thermal analysis framework to improve analysis efficiency. The accuracy of the model is validated against a commercial finite element tool-COMSOL. Experimental results show that the improved TSV thermal model (with proposed lateral thermal model) could greatly improve the accuracy of FD method in thermal simulation comparing with the existing method.
Zao Liu, Sahana Swarup, Sheldon X.-D. Tan, Haibao Chen, Hai Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2013 Compact nonlinear thermal modeling of packaged integrated systems
abstract
This paper proposes a new thermal nonlinear modeling technique for packaged integrated systems. Thermal behavior of complicated systems like packaged electronic systems may exhibit nonlinear and temperature dependent properties. As a result, it is difficult to use a low order linear model to approximate the thermal behavior of the packaged integrated systems without accuracy loss. In this paper, we try to mitigate this problem by using piecewise linear (PWL) approach to characterizing the thermal behavior of those systems. The new method (called ThermSubPWL), which is the first proposed approach to nonlinear thermal modeling problem, identifies the linear local models for different temperature ranges using the subspace identification method. A linear transformation method is proposed to transform all the identified linear local models to the common state basis to build the continuous piecewise linear model. Experimental results validate the proposed method on a realistic packaged integrated system modeled via the multi-domain/physics commercial tool, COMSOL, under practical power signal inputs. The new piecewise models can lead to much smaller model order without accuracy loss, which translates to significant savings in both the simulation time and the time required to identify the reduced models compared to applying the high order models.
Zao Liu, Sheldon X.-D. Tan, Hai Wang 0002, Sahana Swarup, Ashish Gupta 0007
ASP-DAC3
2013 Dynamic thermal management for multi-core microprocessors considering transient thermal effects
abstract
Dynamic thermal management method is a viable way to effectively mitigate the thermal emergences. In this paper, a new thermal management scheme is proposed to reduce the on-chip temperature variance and the occurrence of hot spots by considering more transient thermal effects. The new method performs the task migrations to reduce the temperature variations across the chip. Instead of intuitively assigning the heavy tasks to the low temperature cores to balance the thermal profile based on steady state thermal analysis, the proposed method applies moment matching based transient thermal analysis techniques for fast thermal estimation and prediction to guide the migration process. We show that by considering the dominant temperature moment component, the resulting algorithm can lead to significant reduction of hot spots without full transient thermal simulation. Our experimental results on a 16 core microprocessor demonstrate that the proposed method can reduce the number of the hot spots by 50% compared to the simple lowest temperature based task scheduling method, leading to more uniform on-chip temperature distribution across the microprocessor cores.
Zao Liu, Tailong Xu, Sheldon X.-D. Tan, Hai Wang 0002
ASP-DAC4
2013 A power-driven thermal sensor placement algorithm for dynamic thermal management
abstract
On-chip physical thermal sensors play a vital role for accurately estimating the full-chip thermal profile. How to place physical sensors such that both the number of thermal sensors and the temperature estimation errors are minimized becomes important for on-chip dynamic thermal management of today's high-performance microprocessors. In this paper, we present a new systematic thermal sensor placement algorithm. Different from the traditional thermal sensor placement algorithms where only the temperature information is explored, the new placement method takes advantage of functional unit power information by exploiting the correlation of power estimation errors among functional blocks. The new power-driven placement algorithm applies the correlation clustering algorithm to determine both the locations of sensors and the number of sensors automatically such that the temperature estimation errors can be minimized. Experimental results on a dual-core architecture show that the new thermal sensor placements yield more accurate full-chip temperature estimation compared to the uniform and the k-means based placement approaches.
Hai Wang 0002, Sheldon X.-D. Tan, Sahana Swarup, Xuexin Liu
DATE1
2013 Parallel power grid analysis using preconditioned GMRES solver on CPU-GPU platforms
abstract
In this paper, we propose an efficient parallel dynamic linear solver, called GPU-GMRES, for transient analysis of large power grid networks. The new method is based on the preconditioned generalized minimum residual (GMRES) iterative method implemented on heterogeneous CPU-GPU platforms. The new solver is very robust and can be applied to power grids with different structures and other applications like thermal analysis. The proposed GPU-GMRES solver adopts the very general and robust incomplete LU (ILU) based preconditioner. We show that by properly selecting the right amount of fill-ins in the incomplete LU factors, a good trade-off between GPU efficiency and GMRES convergence rate can be achieved for the best overall performance. Such a tunable feature makes this algorithm very adaptive to different problems. Furthermore, we properly partition the major computing tasks in GMRES solver to minimize the data traffic between CPU and GPU, which further boosts performance of the proposed method. Experimental results on the set of published IBM benchmark circuits and mesh-structured power grid networks show that the GPU-GMRES solver can deliver order of magnitudes speedup over the direct LU solver UMFPACK. GPU-GMRES can also deliver 3-10× speedup over the CPU implementation of the same GMRES method on transient analysis.
Xuexin Liu, Hai Wang 0002, Sheldon X.-D. Tan
ICCAD2
2013 Composable thermal modeling and simulation for architecture-level thermal designs of multicore microprocessors
abstract
Efficient temperature estimation is vital for designing thermally efficient, lower power and robust integrated circuits in nanometer regime. Thermal simulation based on the detailed thermal structures no longer meets the demanding tasks for efficient design space exploration. The compact and composable model-based simulation provides a viable solution to this difficult problem. However, building such thermal models from detailed thermal structures was not well addressed in the past. In this article, we propose a new compact thermal modeling technique, called ThermComp , standing for thermal modeling with composable modules. ThermComp can be used for fast thermal design space exploration for multicore microprocessors. The new approach builds the composable model from detailed structures for each basic module using the finite difference method and reduces the model complexity by the sampling-based model order reduction technique. These composable models are then used to assemble different multicore architecture thermal models and realized into SPICE-like netlists. The resulting thermal models can be simulated by the general circuit simulator SPICE. ThermComp tries to preserve the accuracy of fine-grained models with the speed of coarse-grained models. Experimental results on a number of multicore microprocessor architectures show the new approach can easily build accurate thermal systems from compact composable models for fast architecture thermal analysis and optimization and is much faster than the existing HotSpot method with similar accuracy.
Hai Wang 0002, Sheldon X.-D. Tan, Ashish Gupta 0007, Yuan Yuan 0030
ACM Trans. Design Autom. Electr. Syst.1
2012 Parallel statistical analysis of analog circuits by GPU-accelerated graph-based approach
abstract
In this paper, we propose a new parallel statistical analysis method for large analog circuits using determinant decision diagram (DDD) based graph technique based on GPU platforms. DDD-based symbolic analysis technique enables exact symbolic analysis of vary large analog circuits. But we show that DDD-based graph analysis is very amenable for massively threaded based parallel computing based on GPU platforms. We design novel data structures to represent the DDD graphs in the GPUs to enable fast memory access of massive parallel threads for computing the numerical values of DDD graphs. The new method is inspired by inherent data parallelism and simple data independence in the DDD-based numerical evaluation process. Experimental results show that the new evaluation algorithm can achieve about one to two order of magnitudes speedup over the serial CPU based evaluations and 2-3 times speedup over numerical SPICE-based simulation method on some large analog circuits.
Xuexin Liu, Sheldon X.-D. Tan, Hai Wang 0002
DATE3
2012 A GPU-accelerated envelope-following method for switching power converter simulation
abstract
In this paper, we propose a new envelope-following parallel transient analysis method for the general switching power converters. The new method first exploits the parallelisim in the envelope-following method and parallelize the Newton update solving part, which is the most computational expensive, in GPU platforms to boost the simulation performance. To further speed up the iterative GMRES solving for Newton update equation in the envelope-following method, we apply the matrix-free Krylov basis generation technique, which was previously used for RF simulation. Last, the new method also applies more robust Gear-2 integration to compute the sensitivity matrix instead of traditional integration methods. Experimental results from several integrated on-chip power converters show that the proposed GPU envelope-following algorithm leads to about 10× speedup compared to its CPU counterpart, and 100× faster than the traditional envelop-following methods while still keeps the similar accuracy.
Xuexin Liu, Sheldon X.-D. Tan, Hai Wang 0002, Hao Yu 0001
DATE3
2012 Runtime power estimator calibration for high-performance microprocessors
abstract
Accurate runtime power estimation is important for on-line thermal/power regulation on today's high performance processors. In this paper, we introduce a power calibration approach with the assistance of on-chip physical thermal sensors. It is based on a new error compensation method which corrects the errors of power estimations using the feedback from physical thermal sensors. To deal with the problem of limited number of physical thermal sensors, we propose a statistical power correlation extraction method to estimate powers for places without thermal sensors. Experimental results on standard SPEC benchmarks show the new method successfully calibrates the power estimator with very low overhead introduced.
Hai Wang 0002, Sheldon X.-D. Tan, Xuexin Liu, Ashish Gupta 0007
DATE1
2012 Fast timing analysis of clock networks considering environmental uncertainty
Hai Wang 0002, Hao Yu 0001, Sheldon X.-D. Tan
Integr.1
2012 Fast Statistical Full-Chip Leakage Analysis for Nanometer VLSI Systems
abstract
In this article, we present a new full-chip statistical leakage estimation considering the spatial correlation condition (strong or weak). The new algorithm can deliver linear time, O ( N ), time complexity, where N is the number of grids on chip. The proposed algorithm adopts a set of uncorrelated virtual variables over grid cells to represent the original physical random variables and the cell size is determined by the spatial correlation length. In this way, each physical variable is always represented by virtual variables locally. We prove the number of neighbor cells for each grid cell is not related to the condition of spatial correlation (from no correlation to 100% correlated), which leads to linear time complexity in terms of number of gates. We compute the gate leakage by the orthogonal polynomials-based collocation method. The total leakage of a whole chip can be computed by simply summing up the coefficients of corresponding orthogonal polynomials in each grid cell. Furthermore, we develop a look-up table to cache statistical information for each type of gate instead of calculating leakage for every single instance of gate on a chip. As a result, a new statistical leakage characterization in Standard Cell Library (SCL) is put forward. Furthermore, an incremental analysis algorithm is proposed to update the chip-level statistical leakage information efficiently after a few changes are made. The proposed method has no restrictions on static leakage models, or types of leakage distributions. The large circuit examples in 45nm CMOS process demonstrate the proposed algorithm is 1000X faster than a recently proposed grid-based method with similar accuracy and many orders of magnitude times speedup over the Monte Carlo method. Experimental results also show the incremental analysis provides about 10X further speedup. We expect the incremental analysis could achieve more speedup over the full leakage analysis for larger problem sizes.
Ruijing Shen, Sheldon X.-D. Tan, Hai Wang 0002, Jinjun Xiong
ACM Trans. Design Autom. Electr. Syst.3
2012 Compact Modeling of Interconnect Circuits over Wide Frequency Band by Adaptive Complex-Valued Sampling Method
abstract
In this article, we propose a new model order-reduction method for compact modeling of interconnect circuits over wide frequency band using a novel complex-valued adaptive sampling and error estimation scheme. We address the outstanding error control problems in the existing sampling-based reduction framework over a frequency band. Our new method, WBMOR , explicitly and efficiently computes the exact residual errors to guide the sampling process. We show by sampling along the imaginary axis and performing a new complex-valued reduction that the reduced model will match exactly with the original model at the sample points. Additionally, we show in theory that the proposed method can achieve the error bound over a given frequency range. In practice, the new algorithm can help designers choose the best order of the reduced model for the given frequency range and error bound via the adaptive sampling scheme. In addition, WBMOR can perform wideband accurate reductions of interconnect circuits for analog and RF applications where model accuracy needs to be maintained over a wide frequency range. We compare several sampling schemes such as Monte Carlo, logarithmic, recently proposed resampling, and ARMS methods. Experimental results on a number of RLC circuits show that WBMOR is much more efficient than all the other sampling methods, including the recently proposed resampling and ARMS schemes with the same reduction orders. Compared with the traditional real-valued sampling methods, the complex-valued sampling method is more accurate for the same computational cost.
Hai Wang 0002, Sheldon X.-D. Tan, Ryan Rakib
ACM Trans. Design Autom. Electr. Syst.1
2011 Full-chip runtime error-tolerant thermal estimation and prediction for practical thermal management
abstract
Temperature estimation and prediction are critical for online regulation of temperature and hot spots on today's high performance processors. In this paper, we present a new method, called FRETEP, to accurately estimate and predict the full-chip temperature at runtime under more practical conditions where we have inaccurate thermal model, less accurate power estimations and limited number of on-chip physical thermal sensors. FRETEP employs a number of new techniques to address this problem. First, we propose a new thermal sensor based error compensation method to correct the errors due to the inaccuracies in thermal model and power estimations. Second, we raise a new correlation based method for error compensation estimation with limited number of thermal sensors. Third, we optimize the compact modeling technique and integrate it into the error compensation process in order to perform the thermal estimation with error compensation at runtime. Last but not least, to enable accurate temperature prediction for the emerging predictive thermal management, we design a full-chip thermal prediction framework employing time series prediction method. Experimental results show FRETEP accurately estimates and predicts the full-chip thermal behavior with very low overhead introduced and compares very favorably with the Kalman filter based approach on standard SPEC benchmarks.
Hai Wang 0002, Sheldon X.-D. Tan, Guangdeng Liao, Rafael Quintanilla, Ashish Gupta 0007
ICCAD1
2010 Wideband reduced modeling of interconnect circuits by adaptive complex-valued sampling method
abstract
In this paper, we propose a new wideband model order reduction method for interconnect circuits by using a novel adaptive sampling and error estimation scheme. We try to address the outstanding error control problems in the existing sampling-based reduction framework. In the new method, called WBMOR, we explicitly compute the exact residual errors to guide the sampling process. We show that by sampling along the imaginary axis and performing a new complex-valued reduction, the reduced model will match exactly with the original model at the sample points. We show theoretically that the proposed method can achieve the error bound over a given frequency range. Practically the new algorithm can help designers choose the best order of the reduced model for the given frequency range and error bound via adaptive sampling scheme. As a result, it can perform wideband accurate reductions of interconnect circuits for analog and RF applications. We compare several sampling schemes such as linear, logarithmic, and recently proposed re-sampling methods. Experimental results on a number of RLC circuits show that WBMOR is much more accurate than all the other simple sampling methods and the recently proposed re-sampling scheme with the same reduction orders. Compared with the real-valued sampling methods, the complex-valued sampling method is more accurate for the same computational costs.
Hai Wang 0002, Sheldon X.-D. Tan, Gengsheng Chen
ASP-DAC1
2010 A fast analog mismatch analysis by an incremental and stochastic trajectory piecewise linear macromodel
abstract
To cope with an increasing complexity when analyzing analog mismatch in sub-90nm designs, this paper presents a fast non-Monte-Carlo method to calculate mismatch in time domain. The local random mismatch is described by a noise source with an explicit dependence on geometric parameters, and is further expanded by stochastic orthogonal polynomials (SOPs). This forms a stochastic differential-algebra-equation (SDAE). To deal with large-scale problems, the SDAE is linearized at a number of snapshots along the nominal transient trajectory, and hence is naturally embedded into a trajectory-piecewise-linear (TPWL) macromodeling. The TPWL is improved with a novel incremental aggregation of subspaces identified at those snapshots. Experiments show that the proposed method, isTPWL, is hundreds of times faster than Monte-Carlo method with a similar accuracy. In addition, our macromodel further reduces runtime by up to 25X, and is faster to build and more accurate to simulate compared to existing approaches.
Hao Yu 0001, Xuexin Liu, Hai Wang 0002, Sheldon X.-D. Tan
ASP-DAC3
2009 Fast analysis of nontree-clock network considering environmental uncertainty by parameterized and incremental macromodeling
abstract
It is challenging to verify clock-skew for large-scale nontree clock network with environmental uncertainties such as supply voltage fluctuation and thermal temperature gradient. This paper presents a fast clock-skew analysis via parameterized incremental truncated-balanced-realization, called piTBR method. Environmental uncertainties are parametrically and structurally added into the state equation of clock network. A compact macromodel is obtained by the subspace projection constructed from the singular value decomposition (SVD) of circuit output waveforms. To reduce the computational cost, we propose an incremental SVD method that only needs to partially update the projection matrix by analyzing the perturbed output waveform owning to environmental uncertainties. Experiments on a number of clock networks show that compared with the macromodeling by the fast TBR method, our method reduces the computational cost in the order of 100× with a similar accuracy. In addition, compared with the macromodeling by the Krylov-subspace-based method, our method reduces the waveform error by 2× with a similar runtime.
Hai Wang 0002, Hao Yu 0001, Sheldon X.-D. Tan
ASP-DAC1