Norman Chang

dblp:38/3350 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2524-0935ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2025 Efficient ML-Based Transient Thermal Prediction for 3D-ICs
abstract
Thermal issues of 3D-ICs have become increasingly severe in recent years. Thus, thermal simulation is needed to ensure thermal safety during the design stage. However, performing thermal simulation iteratively requires a significant amount of time. As a result, a fast and accurate method for thermal prediction is a promising alternative to improve the turnaround time. In this paper, we propose a fast thermal prediction method using machine learning models. In the training phase, we employ two models: one for the initial three time steps and another for the subsequent time steps. To enhance prediction accuracy, we introduce two types of features: spaced-windowed features and time-decayed features. These features help us to capture spatial and temporal information effectively. In our experiment, the mean absolute error for the predicted temperature is 1.12°C, and the maximum error is 7.27 °C. In the prediction phase, we achieve a 116X speed-up compared to a commercial tool. With our proposed method, users can predict transient thermal profiles quickly and accurately to ensure thermal safety.
Yun-Feng Yang, Wei-Shen Wang, Yung-Jen Lee, Chien-Mo James Li, Norman Chang, Ying-Shiun Li, Jessica Yen, Lang Lin
ASP-DAC5
2025 Automatic IR-Informed Timing and Timing-Aware IR Optimization
abstract
This paper presents an integrated IR-Informed Timing and Timing-Aware IR Optimization flow with an IR-drop predictor. The proposed flow couples an IR-Informed Timing Optimizer with a Timing-Aware IR Optimizer to consider the mutual impact between IR-drop and timing during optimization. Then, we leverage a fast ML-based IR-drop predictor to quickly estimate the IR-drop after each iteration of optimization, which enables fast switching between the IR optimizer and timing optimizer. We further propose Feature Approximation to speed up the inference time of the IR-drop predictor. On two 7nm designs, the proposed flow closes timing and eliminates at least 90.6% of IR-drop violations. The Feature Approximation achieves 67% speed up in the runtime of the overall flow. Our optimization flow can be applied to a 945k-cell design with 7,578 IR-drop violations within 3 hours, demonstrating its practicality.
Po-Chieh Yen, Wei-Shen Wang, Shao-Yu Wu, Bing-Chen Li, Chien-Mo James Li, Norman Chang, Ying-Shiun Li, Lang Lin
ITC-Asia6
2024 Thermal-Aware Test Frequency Optimization
abstract
Thermal issues during testing of Very Large Scale Integration (VLSI) chips have become more severe as design complexity increases. Test frequency optimization is needed because high test frequencies can cause thermal damage to circuits under test (CUT), while low test frequencies can result in long test time. In this paper, we propose three techniques to minimize the test time of ATPG scan tests without peak temperature violation. First, we propose a single test frequency optimization using machine learning predicted power maps. Second, we partition a test schedule into subschedules and perform multiple test frequency optimization for each subschedule to further reduce test time. Third, we show that we can partition a test schedule by our proposed Power Gap to obtain an even shorter test time. Our experimental results show that the total test time at our optimized multiple test frequencies is 45.91% shorter than the total test time at the original single test frequency.
Wei-Shen Wang, Zhe-Jia Liang, Chien-Mo James Li, Norman Chang, Ying-Shiun Li
ITC-Asia4
2023 Invited Paper: Solving Fine-Grained Static 3DIC Thermal with ML Thermal Solver Enhanced with Decay Curve Characterization
abstract
Static chip thermal analysis provides detailed and accurate thermal profile on chip. The chip power map, commonly modeled as rectangular regions of distinct heat sources, significantly impacts the chip thermal profile. Since the heat sources result from numerous cells in functional blocks, the design space of chip power map is prohibitively enormous. Numerical simulations can be reliable for solving complex power maps; however, it could be very time-consuming when simulating a large SoC and/or 3DIC designs. Thus, there is an urgent need for speeding up the static chip thermal analysis to tackle various power maps. In this paper, we propose an approach of integrating our developed machine learning thermal solver [1] and decay curve characterization for solving static chip thermal with diverse power maps. The machine learning thermal solver would first solve the power maps on a coarse level (e.g., 200 um). The thermal results are further enhanced using the decay curve algorithm which would fine tune the solution locally provided by the machine learning thermal solver and calculate the local temperature variations at a finer level (e.g., 10 um). The deep learning models are trained on augmented artificial power maps and tested on realistic chip power maps. Experimental results validate the effectiveness of the proposed approach of offering fast and accurate chip thermal profile.
Norman Chang, Jie Yang 0023, Wenbo Xia, Lang Lin, Rishikesh Ranade
ICCAD2
2023 High-Speed, Low-Storage Power and Thermal Predictions for ATPG Test Patterns
abstract
High test power causes thermal damage to chips under test. We need power and thermal analyses to ensure thermal safety of ATPG patterns. This requires long runtime and large disk storage because there are many cycles in ATPG patterns. In this paper, we propose power and thermal predictions for test applications. To save runtime, we use multiple ML models and decay surface models for power and thermal predictions, respectively. To save storage, we build features from flip-flop values, so we don't need internal logic values from gate-level simulation. Our mean absolute percentage error (MAPE) for power prediction is less than 8%. Our mean absolute error (MAE) for thermal prediction is less than 1.2°C. We enable transient thermal analysis of long ATPG patterns, with 75X runtime speedup and 118X storage reduction. Our predictions are scalable with test speed, so they can be used to optimize test time while ensuring thermal safety.
Zhe-Jia Liang, Yu-Tsung Wu, Yun-Feng Yang, Chien-Mo James Li, Norman Chang, Ying-Shiun Li
ITC5
2023 Silicon-correlated Simulation Methodology of EM Side-channel Leakage Analysis
abstract
Cryptography hardware is vulnerable to side-channel (SC) attacks on power supply current flow and electromagnetic (EM) emission. This article proposes simulation-based power and EM side-channel leakage analysis (SCLA) techniques on a cryptographic integrated circuit (IC) chip in system level assembly. SCLA measures SC leakage metrics including T-score, SC leakage score, and the number of measurement traces to disclosure, leveraged by a secure system-on-chip design flow toward SC attack resiliency and SC leakage sign off. Power SCLA features the tracking of security sensitive registers within cryptographic logic paths and the automatic assignments of probe points on associated physical power nets. Power supply current traces are efficiently simulated for the large set of input payloads, with direct vector-based and vector-less random switching controls. EM SCLA evaluates magnetic fields created by every piece of metal wiring in metal stacks where power supply current of cryptographic processing flows. The EM emission and EM SCLA from the backside Si surface of an IC chip in flip-chip packaging are experimentally examined with a 0.13 μm test chip. The proposed simulation-based SCLA exhibits the SC leakage metrics of on-chip location and direction dependency as accurately as in the measurements.
Kazuki Monta, Lang Lin, Jimin Wen, Harsh Shrivastav, Calvin Chow, Joao Geada, Sreeja Chowdhury, Nitin Pundir, Norman Chang, Makoto Nagata
ACM J. Emerg. Technol. Comput. Syst.10
2022 Vector-based Dynamic IR-drop Prediction Using Machine Learning
abstract
Vector-based dynamic IR-drop analysis of the entire vector set is infeasible due to long runtime. In this paper, we use machine learning to perform vector-based IR drop prediction for all logic cells in the circuit. We extract important features, such as toggle counts and arrival time, directly from the logic simulation waveform so that we can perform vector-based IR-drop prediction quickly. We also propose a feature engineering method, density map, to increase correlation by 0.1. Our method is scalable because the feature dimension is fixed (72), independent of design size and cell library. Our experiments show that the mean absolute error of the predictor is less than 3% of the nominal supply voltage. We achieve more than 495 speedups compared to a popular commercial tool. Our machine learning prediction can be used to identify IR-drop risky vectors from the entire test vector set, which is infeasible using traditional IR-drop analysis.
Jia-Xian Chen, Shi-Tang Liu, Yu-Tsung Wu, Mu-Ting Wu, Chien-Mo James Li, Norman Chang, Ying-Shiun Li, Wentze Chuang
ASP-DAC6
2021 ML-augmented Methodology for Fast Thermal Side-channel Emission Analysis
abstract
Accurate side-channel attacks can non-invasively or semi-invasively extract secure information from hardware devices using "side- channel" measurements. The thermal profile of an IC is one class of side channel that can be used to exploit the security weaknesses in a design. Measurement of junction temperature from an on-chip thermal sensor or top metal layer temperature using an infrared thermal image of an IC with the package being removed can disclose secret keys of a cryptographic design through correlation power analysis. In order to identify the design vulnerabilities to thermal side channel attacks, design time simulation tools are highly important. However, simulation of thermal side-channel emission is highly complex and computationally intensive due to the scale of simulation vectors required and the multi-physics simulation models involved. Hence, in this paper, we have proposed a fast and comprehensive Machine Learning (ML) augmented thermal simulation methodology for thermal Side-Channel emission Analysis (SCeA). We have developed an innovative tile-based Delta-T Predictor using a data-driven DNN-based thermal solver. The developed tile based Delta-T Predictor temperature is used to perform the thermal side-channel analysis which models the scenario of thermal attacks with the measurement of junction temperature. This method can be 100-1000x faster depending on the size of the chip compared to traditional FEM-based thermal solvers with the same level of accuracy. Furthermore, this simulation allows for the determination of location- dependent wire temperature on the top metal layer to validate the scenario of thermal attack with top metal layer temperature. We have demonstrated the leakage of the encryption key in an 128-bit AES chip using both proposed tile-based temperature calculations and top metal wire temperature calculations, quantified by simulation MTD (Measurements-to-Disclosure).
Norman Chang, Deqi Zhu, Lang Lin, Dinesh Selvakumaran, Jimin Wen, Stephen H. Pan, Wenbo Xia, Calvin Chow, Gary Chen
ASP-DAC1
2018 Machine learning based generic violation waiver system with application on electromigration sign-off
abstract
Manually analyzing the results generated by EDA tools to waive or fix any violations is a tedious, error-prone and time-consuming process. By automating these time-consuming rigorous manual procedures by aggregating key insights across different designs using continuing and prior simulation data, a design team can speed up the tape-out process, optimize resources and significantly minimize the risk of overlooking must fix violations that are prone to cause field failures. In this paper, a machine learning based generic waiver system is proposed which continuously learns to improve with new design data using K-means clustering and nearest neighbor algorithms for risk scoring. The system has been used on new designs to demonstrate on-chip Electromigration (EM) waiver (EMWaiver) mechanism that yielded highly confident results.
Norman Chang, Ajay Baranwal, Ming-Chih Shih, Rahul Rajan, Yaowei Jia, Hui-Lun Liao, Ying-Shiun Li, Ting Ku, Rex Lin
ASP-DAC1
2014 Dynamic tail packing to optimize space utilization of file systems in embedded computing systems
abstract
Embedded computing systems usually have limited computing power, RAM space, and storage capacity due to the consideration of their cost, energy consumption, and physical size. Some of them such as sensor nodes and embedded consumer electronics only have a small-sized flash memory as their storage with a (simple) file system to manage their data, which are usually of small sizes. However, the existing file systems usually have low space utilization on managing small files and the tail data of large files. In this work, we propose a dynamic tail packing scheme to optimize the space utilization of file systems by dynamically aggregating/packing the tail data of (small) files together. The proposed scheme was implemented in the file system of Linux operating systems to evaluate its capability. The results demonstrate that the proposed scheme could significantly improve the space utilization of existing file systems.
Nien-I Hsu, Tseng-Yi Chen, Yuan-Hao Chang 0001, Hsin-Wen Wei, Wei-Kuan Shih, Norman Chang
RTCSA6
2014 Heterogeneous and Elastic Computation Framework for Mobile Cloud Computing
abstract
The number and variety of applications for mobile devices continue to grow. However, the resources on mobile devices including computation and storage do not keep pace with the growth. How to incorporate the computation capacity on cloud servers into mobile computing has been desired and challenge issues to resolve. In this work, we design an elastic computation framework to take advantage the heterogeneous computation capacity on cloud servers, which consist of CPUs and GPGPUs, to meet the computation demands of ever growing mobile applications. The computation framework extends OpenCL framework to link remote processors with local mobile applications. The framework is flexible in the sense that the computation can be stopped at any time and gains results, which is called imprecise computation in real-time computing literature. The framework has been evaluated against OpenCL benchmark and physical computation engine for gaming. The results show that the framework supports OpenCL benchmark, RODINIA, without modifying the codes with few exceptions. The elastic computation framework allows the cloud servers to support more mobile clients without sacrificing their QoS requirements. The experiment results also show that IO intensive applications do not perform well when the network capacity is insufficient or unreliable.
Chi-Sheng Shih 0001, Joen Chen, Yu-Hsin Wang, Norman Chang
Int. J. Softw. Eng. Knowl. Eng.4
2009 Early analysis for power distribution networks
abstract
An efficient and effective power distribution network is crucial to the function and performance of chip and package. As semiconductor processing technology advances to 90nm node and below, system-on-chip (SoC) designs face complex power supply challenges driven by changes such as higher placement and power density, smaller wire and via geometries, and lower supply voltages, in sophisticated, multi-layered packages and boards. The design and verification of power distribution networks is becoming more and more difficult, requiring that these issues be addressed at an early stage of design. In this talk, we present a methodology for verifying power distribution networks early in the design process, when complete placement and routing information is not yet available. The usage model is extremely flexible and allows designers to run the analysis with different types of design abstraction. We demonstrate that the proposed methodology can help designers ensure that the power distribution network meets the performance guideline all through the design cycle.
Aveek Sarkar, Norman Chang
ISPD3
2001 Challenges in Power-Ground Integrity
abstract
With the advance of semiconductor manufacturing, EDA, and VLSI design technologies, circuits with increasingly higher speed are being integrated at an increasingly higher density. This trend causes correspondingly larger voltage fluctuations in the on-chip power distribution network due to IR-drop, L di/dt noise, or LC resonance. Therefore, power-ground integrity becomes a serious challenge in designing future high-performance circuits. In this paper, we introduce power-ground integrity, addressing its importance, verification methodology, and problem solution.
Norman Chang
ICCAD2
2000 Clocktree RLC Extraction with Efficient Inductance Modeling
abstract
In this paper we present an efficient yet accurate inductance extraction methodology and also apply it to clocktree RLC extraction. We first show that without loss of accuracy, the inductance extraction problem of n traces with or without ground planes can be reduced to a number of one-trace and two-trace subproblems. We then solve one-trace and two-trace subproblems via a table-based approach. We finally validate the linear cascading assumption that enables us to apply our inductance extraction approach to clocktree RLC extraction and optimization.
Norman Chang, O. Sam Nakagawa, Weize Xie, Lei He 0001
DATE1
1997 Fast Generation of Statistically-based Worst-Case Modeling of On-Chip Interconnect
abstract
In this paper, we describe a novel methodology for obtaining statistically-based worst case (i.e. 3-/spl sigma/) R (resistance), C (capacitance), and delay given variations in interconnect-related process parameters. Our approach is based on a weighted root-sum square method to derive 3-/spl sigma/ C. A Monte Carlo-based method is used for the generation of 3-/spl sigma/ R as well as randomized distributed RC nets to obtain realistic 3-/spl sigma/ delays for long interconnect nets such as global critical paths. Using this methodology for a long critical net analysis on a 0.35 /spl mu/m process, a more than 70% improvement in 3-/spl sigma/ delay estimation compared with the traditional skew-corner worst case delay can be realized.
Norman Chang, Valery Kanevsky, O. Sam Nakagawa, Khalid Rahmat, Soo-Young Oh
ICCD1
1992 Interconnect Modeling and Design in High-Speed VLSI/ULSI Systems
abstract
A batch-oriented interconnect modeling system, IPDA (Interconnect Performance Design Assistant), developed for interconnect modeling and design in high-speed VLSI/ULSI systems, is presented. Intrachip communication is degraded by RC delay and crosstalk, and electromigration generates a serious reliability problem. Several approaches for alleviating these problems are quantitatively analyzed, and optimum approaches are recommended. In the interchip communication, reflection, crosstalk, and simultaneous switching noise, degrade the system performance. Approaches for solving these problems are introduced and analyzed for practical applications.>
Soo-Young Oh, Keh-Jeng Chang, Norman Chang, Ken Lee
ICCD3