EDBT 2026 Demo / reviewers in the wild / expert
Haibao Chen
dblp:68/10364 · also Hai-Bao Chen
· DBLP profile ↗
65ranked-venue papers
8as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 6 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Simulation of IC Packages with TEC Based on Adaptive Segmented Method and Spatially-Aware Thermal Neural Network
Shunxiang Lan, Haibao Chen |
ASP-DAC | 5 |
| 2026 | Energy-Efficient Short-Packet Covert Communications for Full-Duplex Wireless Systems With AoI Constraint
Yangfan Xu, Bin Yang 0010, Xiuwen Sun, Shikai Shen, Haibao Chen, Bao Gui, Tarik Taleb |
IEEE Internet Things J. | 6 |
| 2025 | Dual-branch cross-modal fusion with local-to-global learning for UAV object detectionabstractDue to the significant differences between unmanned aerial vehicle (UAV) images and natural scene images in terms of lighting, scale, and viewing angle, existing multispectral detection techniques often fail to fully utilize the remote dependencies between global and local information, resulting in poor performance in complex UAV scenarios. In this paper, we propose a novel two-branch cross-modal fusion network that integrates a dual cross attention transformer fusion block (CTF) for global feature dependency and an adaptive mask convolution fusion block (MCF) for underlying feature extraction. This achieves a unified representation with both global and local receptive fields. Our local-global training strategy utilizes a shallow global fusion network and a deep local fusion network, which operate on the entire image while also focusing on detailed local features. Additionally, we integrate an asymptotic feature pyramid network that employs adaptive spatial fusion to refine features, enhancing the accuracy of small object detection in UAV scenes. Evaluating our work with the DroneVehicle dataset for vehicle detection using infrared and visible light, our network outperformed existing methods, improving [email protected] by 7.01% compared to CAL-Net. APs for YOLOv8m-based single-modal infrared detection and visible light detection increased by 3.1% and 11.6%, respectively. Binyi Fang, Yixin Yang 0004, Jingjing Chang, Ziyang Gao, Haibao Chen |
ASP-DAC | 5 |
| 2025 | DuQTTA: Dual Quantized Tensor-Train Adaptation with Decoupling Magnitude-Direction for Efficient Fine-Tuning of LLMsabstractRecent parameter-efficient fine-tuning (PEFT) techniques have enabled large language models (LLMs) to be efficiently fine-tuned for specific tasks, while maintaining model performance with minimal additional trainable parameters. However, existing PEFT techniques continue to face challenges in balancing both accuracy and efficiency, especially when addressing scalability and the demands of lightweight deployment for LLMs. In this paper, we propose an efficient fine-tuning method of LLMs based on dual quantized Tensor-Train adaptation with decoupling magnitude-direction (DuQTTA). The proposed DuQTTA method employs Tensor-Train decomposition and dual-stage quantization to minimize model size and resource consumption. Additionally, it employs an adaptive optimization strategy and a decoupled update mechanism to improve model performance, thereby minimizing suboptimal outcomes and ensuring alignment with the full-parameter fine-tuning goals. Experimental results indicate that the proposed DuQTTA method outperforms existing PEFT methods, achieving up to a $65 \times$ compression rate compared to the LLaMA2-7B models, meanwhile delivering improvements of $4.44 \%, 3.14 \%$, and 0.97% over LoRA on LLaMA2-7B, LLaMA3-8B, and LLaMA2-13B, respectively. The proposed DuQTTA method is effective in compressing LLMs for deployment on resource-constrained edge devices. Haoyan Dong, Haibao Chen, Jingjing Chang, Yixin Yang 0004, Ziyang Gao, Zhigang Ji, Runsheng Wang, Ru Huang 0001 |
DAC | 2 |
| 2025 | Equivalent Lumped Element Model for Electromigration Considering Thermal EffectsabstractElectromigration (EM) remains a critical reliability concern in advanced integrated circuit design. Traditional physics-based approaches, which solve partial differential equations (PDEs), are computationally intensive, particularly in multi-physics scenarios. To address the issue, we propose a self-consistent lumped element modeling framework that leverages the equivalence between electrical behavior and stress evolution to forecast EM-induced stress under coupled electro-thermomechanical effects. Thermomechanical interactions driven by temperature gradients are explicitly modeled using embedded controlled sources. A threshold-activated switching mechanism is proposed to dynamically reconfigure circuit topology, enabling seamless simulation across both void nucleation and post-voiding phases. The proposed adaptive non-uniform spatial discretization framework can be used to enhance computational efficiency without sacrificing accuracy. Numerical results demonstrate >50× speedup against the finite element simulation for small interconnects with <1.5% error, and 3.11× acceleration over conventional equivalent circuits for large-scale structures while maintaining <0.5% error. Fully compatible with standard SPICE solver, the proposed approach exhibits strong potential for temperature-aware EM analysis and void prediction in full-chip VLSI applications. Hengyi Zhu, Tianshu Hou, Zhigang Ji, Runsheng Wang, Haibao Chen |
ICCAD | 7 |
| 2025 | Covert THz Communication against Randomly Distributed WardensabstractBy exploiting the high directivity of directional antennas and severe propagation loss, covert Terahertz (THz) communication can better suppress detection by wardens while achieving high-performance legitimate transmission. This emerging technology can achieve a higher level of covertness and can be applied in Internet of Things networks, such as smart senior care. In this work, we explore a one-hop system of covert THz communication, comprising a transmitter-receiver pair and multiple randomly distributed wardens. Utilizing the high directivity of antennas and the molecular absorption effect of THz signals, we can achieve transmissions with high covertness over THz bands. We divide the propagation area around Alice into several regions and propose a covert THz communication protocol in the considered scenario. Specifically, a circular area around Alice is divided into three regions based on a transmission model of directional antennas and the distance to Alice, and Alice decides to conduct transmissions if no warden exists in the insecure region (IR). Meanwhile, the theoretical expression of detection error probability (DEP) at a warden is derived and given. Numerical results demonstrate that our proposed protocol can increase the overall DEP of multiple wardens in both noncolluding and colluding modes compared to the case where no protocol is used. Xinzhe Pi, Bin Yang 0010, Lisheng Ma, Haibao Chen, Guozhu Zhao, Bao Gui |
ICPADS | 4 |
| 2025 | OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image GenerationabstractText-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations, especially for text rendering and style. Notably, recent state-of-the-art models, with their rich knowledge modeling capabilities, show potential in reasoning-driven image generation, yet existing evaluation systems have not adequately addressed this frontier. To systematically address these gaps, we introduce $\textbf{OneIG-Bench}$, a meticulously designed comprehensive benchmark framework for fine-grained evaluation of T2I models across multiple dimensions, including subject-element alignment, text rendering precision, reasoning-generated content, stylization, and diversity. By structuring the evaluation, this benchmark enables in-depth analysis of model performance, helping researchers and practitioners pinpoint strengths and bottlenecks in the full pipeline of image generation. Our codebase and dataset are now publicly available to facilitate reproducible evaluation studies and cross-model comparisons within the T2I research community. Jingjing Chang, Yixiao Fang, Peng Xing, Shuhan Wu, Rui Wang 0099, Xianfang Zeng, Gang Yu 0002, Haibao Chen |
NeurIPS | 9 |
| 2025 | Novel Partitioning-Based Approach for Electromigration Assessment With Neural NetworksabstractDue to continuing technology scaling, electromigration (EM) remains a prominent reliability concern in integrated circuit design. Traditional empirical methods often result in over-design in very large scale integration (VLSI) due to model inaccuracy. Recently, researchers have focused on analyzing EM susceptibility by tracking hydrostatic stress evolution in metal lines, governed by computationally expensive partial differential equations (PDEs). In this paper, we propose a partitioning-based approach using neural networks to efficiently forecast the stress evolution along interconnect trees during the void nucleation and growth phases. This approach begins by decomposing the interconnect tree into subcomponents, providing computationally efficient analytical solutions for predicting stress evolution within each subtree. Subsequently, we employ a lightweight neural network to reassemble these components with their corresponding solutions to the original structure, ensuring accurate stress prediction. This divide-and-conquer strategy can accommodate various tree structures, with offshoots at arbitrary junctions, and holds substantial promise for using NN-based methods to solve mesh-free stress evolution on much larger interconnect trees than previously possible, with reduced computational overhead and heightened accuracy. The proposed approach eliminates the need for time discretization and grid meshing typically required in numerical methods. Numerical results confirm its advantages in accuracy and computational efficiency. Tianshu Hou, Farid N. Najm, Ngai Wong 0001, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | MemATr: An Efficient and Lightweight Memory-augmented Transformer for Video Anomaly DetectionabstractAnomaly detection in videos is a long-standing and challenging problem. Previous methods often adopt deep and large neural networks to achieve the best detection accuracy; however, the high computational costs prevent them from being used in real-world applications with constrained computational resources. In this paper, we develop a mem ory- a ugmented tr ansformer named MemATr, which is capable of detecting video anomalies effectively. The proposed network is lightweight and can be easily deployed on mobile devices. Furthermore, we propose a memory transformer module to make predictions that are closer to normal inputs, thereby leading to a higher error for abnormal input patterns. Memory-attention is the main component of the proposed memory transformer, which can retrieve the features from learnable values rather than from the backbone like previous methods. Extensive experiments on the UCSD Ped2, CUHK Avenue, and ShanghaiTech benchmarks can demonstrate that our model has a significantly smaller model size while still achieving competitive detection accuracy. Our model has only 1/12 the number of parameters of the baseline model. Besides, our model achieves a 4.6% increase in accuracy on the ShanghaiTech dataset and has roughly the same accuracy compared with the baseline on the other two datasets. We validate the performance of the proposed model on the mobile device and the result shows it only has 49.8ms latency. The effectiveness of the proposed method on mobile devices is further supported by experimental results. A new quantitative parameter AMD (Applicability for Mobile Devices) is proposed to offer a novel approach to assist in making trade-offs for mobile devices. The proposed model obtains state-of-the-art results in terms of AMD. Jingjing Chang, Peining Zhen, Xiaotao Yan, Yixin Yang 0004, Ziyang Gao, Haibao Chen |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | CaS2M: A Calibrated Single-to-Multiple Framework for Real-World Partial Fingerprint RecognitionabstractWith the reducing size of fingerprint collection modules in mobile devices, partial fingerprints are increasingly characterized by smaller overlapping areas and higher self-similarity. Existing methods either aggregate similarity scores from individual Single-to-Single recognition or directly employ a Single-to-Multiple network to verify the match between the query and templates. However, these methods either lack sufficient interaction between templates, or fail to provide adequate supervision for the alignment process, a crucial step in fingerprint recognition, thereby limiting overall accuracy. In this paper, we propose a novel partial fingerprint recognition strategy termed Calibrated Single-to-Multiple (CaS2M), which first calibrates template fingerprints individually, then combines them with the query fingerprint in a matcher network for feature fusion. Building upon this strategy, we develop a dual-stage framework tailored to real-world applications. During enrollment, a lightweight patch-based feature indexing algorithm and a template selection strategy are employed accounting for limited hardware resources. For authentication, independent calibration is first applied, followed by an attention-based matcher network to verify identity consistency. Experimental results on multiple public datasets (NIST 302, NIST SD4, SpoofGAN, FVC2002 DB1A & DB3A) and a self-build dataset demonstrate that our framework achieves superior performance over state-of-the-art algorithms, providing new insights for multi-template partial fingerprint recognition. Ziyang Gao, Tianfan Peng, Jingjing Chang, Yixin Yang 0004, Haibao Chen |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Physics-Informed Learning Based Multiphysics Simulation for Fast Transient TSV Electromigration AnalysisabstractThrough Silicon Vias (TSVs) are vulnerable to electromigration (EM) degradation due to their high local current densities, thereby reducing the reliability of 3D ICs with stack dies and TSVs. Due to the broad application of 3D ICs, it is necessary to analyze the electromigration reliability of TSVs. To overcome the weakness of traditional method for EM modeling of TSVs, we propose a physics-informed learning approach for transient analysis of electromigration modeling in TSV by solving the conventional mass balance equation. The proposed method allows simultaneous consideration of atomic depletion and accumulation, effective resistance degradation, electric current evolution, and stress distribution. In particular, we propose a customized neural network to simulate the EM process in TSV without the need for fine grid meshing and temporal iteration in traditional methods. Considering that the loss function of the proposed model is a combination of different loss terms, we propose a modified self-adaptive loss balanced method to automatically adjust the weights of multiple loss terms to enhance network performance. Given the prediction uncertainty due to data randomness or model architecture constraints, Gaussian probabilistic model is constructed to define the self-adaptive weights and update the dynamic weights per epoch built on maximum likelihood estimation. Compared with the finite element method, the proposed physics informed neural network method can lead to a speedup with less than 0.1% mean square error. Experimental results also show that the proposed model achieves excellent performance over other competing methods and high robustness under values of initial weights, different numbers of hidden layers and neurons per layer. Xiaoman Yang, Haibao Chen, Yuhan Zhang 0005, Tianshu Hou, Pengpeng Ren, Runsheng Wang, Zhigang Ji, Ru Huang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | Physics-Informed Learning for EPG-Based TDDB AssessmentabstractTime-dependent dielectric breakdown (TDDB) is one of the important failure mechanisms for copper (Cu) interconnects. Many TDDB models have been proposed based on different physics kinetics in the past. Recently, a physics-based TDDB model, which is based on the breakdown concept of electric path generation (EPG), has been proposed and has shown advantage over widely accepted existing electrostatic field-based TDDB assessment. However, the determination of the time-to-failure from this EPG based TDDB model includes solving partial differential equation (PDE) with time-consuming finite-element method (FEM). In recent years, deep neural networks have been proposed to predict numerical solutions of PDEs. In this paper, we use physics-informed neural network to solve the diffusion equation of ions in an electric field extracted from EPG based TDDB model. The continuous definite condition and hard constrain optimization methods are used for improving the performance of PINN in terms of accuracy and speed. Compared with the FEM method, the proposed PINN method can lead to about 100 times speedup with less than 0.1% mean squared error. Dinghao Chen, Xiaoman Yang, Pengpeng Ren, Zhigang Ji, Haibao Chen |
ASPDAC | 6 |
| 2024 | Physics-Informed Learning for Versatile RRAM Reset and Retention SimulationabstractResistive random-access memory (RRAM) constitutes an emerging and promising platform for compute-inmemory (CIM) edge AI. However, the switching mechanism and controllability of RRAM are still under debate owing to the influence of multiphysics. Although physics-informed neural networks (PINNs) are successful in achieving mesh-free multiphysics solutions in many applications, the resultant accuracy is not satisfactory in RRAM analyses. This work investigates the characteristics of RRAM devices - retention and reset transition which are described in terms of the dissolution of a conductive filament (CF) in 3-D axis-symmetric geometry. Specifically, we provide a novel neural network characterization of ion migration, Joule heating, and carrier transport, governed by the solutions of partial differential equations (PDEs). Motivated by physics-informed learning, the separation of variables (SOV) method and the neural tangent kernel (NTK) theory, we propose a customized 3-channel fully-connected network and a modified random Fourier feature (mRFF) embedding strategy to capture multiscale properties and appropriate frequency features of the self-consistent multiphysics solutions. The proposed model eliminates the need for grid meshing and temporal iterations widely used in RRAM analysis. Experiments then confirm its superior accuracy over competing physics-informed methods. Tianshu Hou, Wenyong Zhou, Can Li 0024, Haibao Chen, Ngai Wong 0001 |
ASPDAC | 6 |
| 2024 | Enforcing hard constraints in physics-informed learning for transient TSV electromigration analysisabstractDue to the high local current densities, Through Silicon Vias (TSVs) are susceptible to electromigration (EM) degradation, which reduces the reliability of integrated circuits. Unlike traditional methods for TSV modeling and simulation, this paper introduces a unified hard constraint physics-informed learning neural network approach, called HCPINN, for the transient analysis of electromigration in TSVs by solving the conventional mass balance equation. The proposed method allows simultaneous consideration of atomic depletion and accumulation, effective resistance degradation, electric current evolution, and stress distribution. Specifically, we propose a hard constraint method for solving partial differential equations (PDEs) with general boundary conditions (BCs) for transient TSV electromigration analysis. By using the extra fields derived from the mixed finite element method, we reconstruct the corresponding PDEs by transforming general BCs into linear forms. Based on this derivation, we embed general BCs of mass balance equation into the proposed ansatz and employ sub-networks for the approximation on general BCs. The main neural network is responsible for training the internal part of the problem domain without adding loss terms with BCs, overcoming the convergence issue due to unbalanced gradients among different loss terms. Besides, we theoretically demonstrate that this reformulation of general BCs can stabilize the training process. Experimental results indicate that the proposed HCPINN exhibits superior performance and reduces boundary error in TSV electromigration analysis. Compared to the finite element method, the proposed network achieves approximately 100 times faster inference with a minimal mean squared error increase of less than 0.1%. Xiaoman Yang, Haibao Chen, Yuhan Zhang 0005, Yongkang Xue, Pengpeng Ren, Runsheng Wang, Zhigang Ji, Ru Huang 0001 |
ICCAD | 2 |
| 2024 | DRGA-Based Second-Order Block Arnoldi Method for Model Order Reduction of MIMO RCS CircuitsabstractWith the escalating demand for fast simulation of large-scale multi-input multi-output (MIMO) RCS circuits formulated as second-order differential systems, the need arises for more effective decentralized second-order model order reduction (MOR) methods, while providing a desired approximation of the original system. Dynamic relative gain array (DRGA) that takes into account both the steady-state and dynamic system information has shown promising efficacy in measuring the degree of each loop interaction, which is crucial for decoupling a MIMO system into several multi-input single-output (MISO) subsystems. Although several decentralized MOR methods have been introduced for dimension reduction to linear MIMO networks, hardly has any research explored second-order decentralized MOR methods with regard to MIMO RCS circuits. Besides, the existing DRGA method based on first-order state feedback predictive control greatly increases the computational complexity when directly applying to second-order RCS systems. Hence, we develop a second-order block Arnoldi method based on DRGA, termed DRGA-SOBAR, which enables the extension of the SOAR method and the second-order DRGA method to MIMO scenarios. Experimental results on RCS networks show that most input-output interactions are negligible in terms of the magnitude-wise insignificance, and our proposed DRGA-SOBAR based reduced systems perform with higher accuracy compared to the PRIMA and the generalized block SOAR (SOBAR) methods, and higher efficiency compared to the decentralized SOBAR algorithm based on RGA method as well. Haibao Chen, Jie Chen 0005, Pengpeng Ren, Zhigang Ji, Junhua Liu 0001, Runsheng Wang, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Analytical Post-Voiding Modeling and Efficient Characterization of EM Failure Effects Under Time-Dependent Current StressingabstractElectromigration (EM) has become the major concern for integrated circuits (ICs) in advanced technology nodes. Traditional empirical EM models, such as Black’s equation, show inaccurate estimation for the time-to-failure of ICs, thus resulting in unnecessary over-design. To address this drawback, we propose a few analytical solutions for calculating the transient stress evolution and void volume in straight multisegment interconnect trees during the post-voiding phase. By employing the Laplace transform, the proposed method aims at solving coupled partial differential equations (PDEs) governed by physics-based EM modeling. The analytical solutions can be tailored to expressions with required accuracy and computational savings, leading to a compact end-to-end system providing results of EM failure effects at arbitrary time instances and locations of interconnect trees with varying geometry under time-dependent current and temperature stressing. The EM lifetime such as the incubation time, related to the void volume evolution, at any desired precision, can be calculated by the analytical solutions. The proposed method shows its accuracy, scalability, and computational savings through results compared with the finite element method (FEM) tool COMSOL and the competing methods and can achieve up to$593\times $speedup with < 10% error in EM failure time estimation. Tianshu Hou, Ngai Wong 0001, Quan Chen 0007, Zhigang Ji, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Multilayer Perceptron-Based Stress Evolution Analysis Under DC Current Stressing for Multisegment WiresabstractElectromigration (EM) is one of the major concerns in the reliability analysis of very large-scale integration (VLSI) systems due to the continuous technology scaling. Accurately predicting the time-to-failure of integrated circuits (ICs) becomes increasingly important for modern IC design. However, traditional methods are often not sufficiently accurate, leading to undesirable over-design especially in advanced technology nodes. In this article, we propose an approach using multilayer perceptrons (MLPs) to compute stress evolution in the interconnect trees during the void nucleation phase. The availability of a customized trial function for neural network training holds the promise of finding dynamic mesh-free stress evolution on complex interconnect trees under time-varying temperatures. Specifically, we formulate a new objective function considering the EM-induced coupled partial differential equations (PDEs), boundary conditions (BCs), and initial conditions to enforce the physics-based constraints in the spatial–temporal domain. The proposed model avoids meshing and reduces temporal iterations compared with conventional numerical approaches like finite element method. Numerical results confirm its advantages on accuracy and computational performance. Tianshu Hou, Peining Zhen, Ngai Wong 0001, Quan Chen 0007, Guoyong Shi, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Equiprobability-Based Local Response Surface Method for High-Sigma Yield Estimation With Both High Accuracy and EfficiencyabstractWith the ever-increasing transistor density and memory capability in integrated circuits, the high-sigma yield estimation has become a growing concern. This work presents an equiprobability-based local response surface (ELRS) method that can perform a high-sigma yield estimation with both high accuracy and efficiency. Demonstrating with 6T-SRAM, the proposed method exhibits more than ten times improvement in accuracy when compared with the state-of-the-art while maintaining the efficiency to the best record in the literature. Pengpeng Ren, Haibao Chen, Zhigang Ji, Junhua Liu 0001, Runsheng Wang, Jianfu Zhang 0001, Ru Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | A Highly Compressed Accelerator With Temporal Optical Flow Feature Fusion and Tensorized LSTM for Video Action Recognition on Terminal DeviceabstractDeep learning-based action recognition has become ubiquitous in the video analysis area; however, large neural networks require enormous computations to achieve high performance, which hinder them from mobile applications that are tightly constrained by hardware resources. In this work, we introduce a highly compact and fast neural network-based action recognition accelerator named ARA on the terminal device. We build an LSTM-based spatio-temporal action recognition model with extracted time-series features from RGB frames and flow features from optical flow fields. Then the LSTM-based spatio-temporal model is deeply compressed with tensor decomposition to further reduce redundant parameters and lessen computation overhead. Based on the datasets UCF-11, UCF-101, and HMDB51, our proposed method achieves 95.87%, 94.08%, and 75.71% classification accuracy, being comparable with other state-of-the-art methods. In particular, our proposed method significantly compresses the parameter of the LSTM model$215\times $on the UCF-101 dataset. The proposed system can also achieve a fast running speed of 157.7 FPS on GPU. Furthermore, we validate the performance of the proposed system on an ARM-based terminal device; the results show it only has 0.017-s latency and 4.73-W power consumption. Peining Zhen, Xiaotao Yan, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Toward Compact Transformers for End-to-End Object Detection With Decomposed Chain Tensor StructureabstractDEtection TRansformer (DETR) is a recently proposed method that streamlines the detection pipeline and achieves competitive results against two-stage detectors such as Faster-RCNN. The DETR models get rid of complex anchor generation and post-processing procedures thereby making the detection pipeline more intuitive. However, the numerous redundant parameters in transformers make the computation and storage of the DETR models intensive, which seriously hinder them to be deployed on the resources-constrained devices. In this paper, to obtain a compact end-to-end detection framework, we propose to deeply compress the transformers with low-rank tensor decomposition. The basic idea of our tensor-based compression method is to represent the large-scale weight matrix in one network layer with a chain of low-order matrices. Furthermore, we show that redundant attention heads will hinder the performance of detection transformers. We thus propose a gated multi-head attention (GMHA) module to suppress the redundant attention information by normalizing the attention heads. In GMHA, each attention head has an independent gate to determine the passed attention value, thereby down-weighting the uninformative heads. The accuracy drop of the tensor-compressed DETR models can be mitigated by applying GMHA modules. Lastly, to obtain fully compressed DETR models, a low-bitwidth quantization technique is introduced for further reducing the model storage size. Based on the proposed methods, we can achieve significant parameter and model size reduction while maintaining high detection performance. We conduct extensive experiments on the COCO and PASCAL VOC datasets to validate the effectiveness of our tensor-compressed (tensorized) DETR models. The experimental results on the COCO benchmark show that we can attain$3.7\times $full model compression with$482\times $feed forward network (FFN) parameter reduction and only 0.6 points accuracy drop. Peining Zhen, Xiaotao Yan, Tianshu Hou, Haibao Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | A Deep Learning Framework for Solving Stress-based Partial Differential Equations in Electromigration AnalysisabstractThe electromigration-induced reliability issues (EM) in very large scale integration (VLSI) circuits have attracted continuous attention due to technology scaling. Traditional EM methods lead to inaccurate results incompatible with the advanced technology nodes. In this article, we propose a learning-based model by enforcing physical constraints of EM kinetics to solve the EM reliability problem. The method aims at solving stress-based partial differential equations (PDEs) to obtain the hydrostatic stress evolution on interconnect trees during the void nucleation phase, considering varying atom diffusivity on each segment, which is one of the EM random characteristics. The approach proposes a crafted neural network-based framework customized for the EM phenomenon and provides mesh-free solutions benefiting from the employment of automatic differentiation (AD). Experimental results obtained by the proposed model are compared with solutions obtained by competing methods, showing satisfactory accuracy and computational savings. Tianshu Hou, Peining Zhen, Zhigang Ji, Haibao Chen |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2023 | Towards Accurate Oriented Object Detection in Aerial Images with Adaptive Multi-level Feature FusionabstractDetecting objects in aerial images is a long-standing and challenging problem since the objects in aerial images vary dramatically in size and orientation. Most existing neural network based methods are not robust enough to provide accurate oriented object detection results in aerial images since they do not consider the correlations between different levels and scales of features. In this paper, we propose a novel two-stage network-based detector with a daptive f eature f usion towards highly accurate oriented object det ection in aerial images, named AFF-Det . First, a multi-scale feature fusion module (MSFF) is built on the top layer of the extracted feature pyramids to mitigate the semantic information loss in the small-scale features. We also propose a cascaded oriented bounding box regression method to transform the horizontal proposals into oriented ones. Then the transformed proposals are assigned to all feature pyramid network (FPN) levels and aggregated by the weighted RoI feature aggregation (WRFA) module. The above modules can adaptively enhance the feature representations in different stages of the network based on the attention mechanism. Finally, a rotated decoupled-RCNN head is introduced to obtain the classification and localization results. Extensive experiments are conducted on the DOTA and HRSC2016 datasets to demonstrate the advantages of our proposed AFF-Det. The best detection results can achieve 80.73% mAP and 90.48% mAP, respectively, on these two datasets, outperforming recent state-of-the-art methods. Peining Zhen, Suming Zhang, Xiaotao Yan, Zhigang Ji, Haibao Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | Deeply Tensor Compressed Transformers for End-to-End Object DetectionabstractDEtection TRansformer (DETR) is a recently proposed method that streamlines the detection pipeline and achieves competitive results against two-stage detectors such as Faster-RCNN. The DETR models get rid of complex anchor generation and post-processing procedures thereby making the detection pipeline more intuitive. However, the numerous redundant parameters in transformers make the DETR models computation and storage intensive, which seriously hinder them to be deployed on the resources-constrained devices. In this paper, to obtain a compact end-to-end detection framework, we propose to deeply compress the transformers with low-rank tensor decomposition. The basic idea of the tensor-based compression is to represent the large-scale weight matrix in one network layer with a chain of low-order matrices. Furthermore, we propose a gated multi-head attention (GMHA) module to mitigate the accuracy drop of the tensor-compressed DETR models. In GMHA, each attention head has an independent gate to determine the passed attention value. The redundant attention information can be suppressed by adopting the normalized gates. Lastly, to obtain fully compressed DETR models, a low-bitwidth quantization technique is introduced for further reducing the model storage size. Based on the proposed methods, we can achieve significant parameter and model size reduction while maintaining high detection performance. We conduct extensive experiments on the COCO dataset to validate the effectiveness of our tensor-compressed (tensorized) DETR models. The experimental results show that we can attain 3.7 times full model compression with 482 times feed forward network (FFN) parameter reduction and only 0.6 points accuracy drop. Peining Zhen, Ziyang Gao, Tianshu Hou, Haibao Chen |
AAAI | 5 |
| 2022 | Automatic Monitoring System for Engines Motion Attitude Based on Video Image DetectionabstractIn some special equipment, the engine assembles nozzles to drive the equipment and adjust movement posture. The nozzles action test is an essential step during the whole test session. The current engine action test relies on the judgment observed by testers. It is difficult to identify some actions when the nozzles swing too fast or slightly. This paper focuses on the research in automatic engine motion tests by using computer vision technology instead of manual observation. We are also dedicated to developing an integrated system used to solve the problems during the test, which can effectively solve the difficulties of inefficiency and inability to record. Algorithms based on the image detection techniques are designed to accomplish this work. At the same time, the algorithms integrate into the hardware and software platform. The proposed system is equipped with some hardware control and monitoring functions. The algorithm designed in this paper achieves 100% accuracy in nozzle target detection and motion detection in real scenarios. The automated monitoring platform supports more than eight cameras for motion monitoring, the algorithm computation speed is higher than 20FPS, and the delay time is not longer than 3s. Yiwei Fang, Ruilin Zeng, Haibao Chen |
COMPSAC | 5 |
| 2022 | FASSST: Fast Attention Based Single-Stage Segmentation Net for Real-Time Instance SegmentationabstractReal-time instance segmentation is crucial in various AI applications. This work designs a network named Fast Attention based Single-Stage Segmentation NeT (FASSST) that performs instance segmentation with video-grade speed. Using an instance attention module (IAM), FASSST quickly locates target instances and segments with region of interest (ROI) feature fusion (RFF) aggregating ROI features from pyramid mask layers. The module employs an efficient single-stage feature regression, straight from features to instance coordinates and class probabilities. Experiments on COCO and CityScapes datasets show that FASSST achieves state-of-the-art performance under competitive accuracy: real-time inference of 47.5FPS on a GTX1080Ti GPU and 5.3FPS on a Jetson Xavier NX board with only 71.6 GFLOPs. Peining Zhen, Tianshu Hou, Chiu Wa Ng, Haibao Chen, Hao Yu 0001, Ngai Wong 0001 |
WACV | 6 |
| 2022 | A Space-Time Neural Network for Analysis of Stress Evolution Under DC Current StressingabstractThe electromigration (EM)-induced reliability issues in very large-scale integration (VLSI) circuits have attracted increased attention due to the continuous technology scaling. Traditional EM models often lead to overly pessimistic predictions incompatible with the shrinking design margin in future technology nodes. Motivated by the latest success of neural networks in solving differential equations in physical problems, we propose a novel mesh-free model to compute EM-induced stress evolution in VLSI circuits. The model utilizes a specifically crafted space–time physics-informed neural network (STPINN) as the solver for EM analysis. By coupling the physics-based EM analysis with dynamic temperature incorporating Joule heating and via effect, we can observe stress evolution along multisegment interconnect trees under constant, time-dependent, and space–time-dependent temperature during the void nucleation phase. The proposed STPINN method obviates the time discretization and meshing required in conventional numerical stress evolution analysis and offers significant computational savings. Numerical comparison with competing schemes demonstrates a$2\times $–$52\times $speedup with a satisfactory accuracy. Tianshu Hou, Ngai Wong 0001, Quan Chen 0007, Zhigang Ji, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | S3-Net: A Fast and Lightweight Video Scene Understanding Network by Single-shot SegmentationabstractReal-time understanding in video is crucial in various AI applications such as autonomous driving. This work presents a fast single-shot segmentation strategy for video scene understanding. The proposed net, called S3-Net, quickly locates and segments target sub-scenes, meanwhile extracts structured time-series semantic features as inputs to an LSTM-based spatio-temporal model. Utilizing ten-sorization and quantization techniques, S3-Net is intended to be lightweight for edge computing. Experiments using CityScapes, UCF11, HMDB51 and MOMENTS datasets demonstrate that the proposed S3-Net achieves an accuracy improvement of 8.1% versus the 3D-CNN based approach on UCF11, a storage reduction of 6.9× and an inference speed of 22.8 FPS on CityScapes with a GTX1080Ti GPU. Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
WACV | 3 |
| 2021 | Unsupervised soft-label feature selection
Fei Wang 0055, Lei Zhu 0002, Jingjing Li 0001, Haibao Chen, Huaxiang Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2021 | TEANS: A Target Enhancement and Attenuated Nonmaximum Suppression Object Detector for Remote Sensing ImagesabstractIn this letter, we propose an effective approach to learn a convolutional neural network (CNN) model with target enhancement and attenuated nonmaximum suppression (NMS) technique (TEANS) for object detection in optical remote sensing images. TEANS mainly consists of two steps. First, the target enhancement architecture, including target upsampling and reconvolution, is designed into a given deep ResNet-101 model for accurate object detection, especially for small ones. Second, the attenuated NMS technique is used for overcoming wrong eliminations of serried object proposals. For verifying the effectiveness of the TEANS method, evaluations are implemented on a publicly available 15-class optical remote sensing object detection data set. Experimental results show that TEANS can achieve 5.55%, 18.77%, 26.81%, 55.07%, 28.48%, 6.01%, and 5.51% improvements in mean Average Precision (mAP), respectively, compared with standard Faster R-CNN, R-FCN, YOLOv2, SSD, USB-BBR, YOLOv3, and MS-VANs frameworks. Haibao Chen, Guanghui He 0002, Bingyi Zhang, Hao Yu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Fast Video Facial Expression Recognition by a Deeply Tensor-Compressed LSTM Neural Network for Mobile DevicesabstractMobile devices usually suffer from limited computation and storage resources, which seriously hinders them from deep neural network applications. In this article, we introduce a deeply tensor-compressed long short-term memory (LSTM) neural network for fast video-based facial expression recognition on mobile devices. First, a spatio-temporal facial expression recognition LSTM model is built by extracting time-series feature maps from facial clips. The LSTM-based spatio-temporal model is further deeply compressed by means of quantization and tensorization for mobile device implementation. Based on datasets of Extended Cohn-Kanade (CK+), MMI, and Acted Facial Expression in Wild 7.0, experimental results show that the proposed method achieves 97.96%, 97.33%, and 55.60% classification accuracy and significantly compresses the size of network model up to 221× with reduced training time per epoch by 60%. Our work is further implemented on the RK3399Pro mobile device with a Neural Process Engine. The latency of the feature extractor and LSTM predictor can be reduced 30.20× and 6.62× , respectively, on board with the leveraged compression methods. Furthermore, the spatio-temporal model costs only 57.19 MB of DRAM and 5.67W of power when running on the board. Peining Zhen, Haibao Chen, Zhigang Ji, Hao Yu 0001 |
ACM Trans. Internet Things | 2 |
| 2021 | S3-Net: A Fast Scene Understanding Network by Single-Shot Segmentation for Autonomous DrivingabstractReal-time segmentation and understanding of driving scenes are crucial in autonomous driving. Traditional pixel-wise approaches extract scene information by segmenting all pixels in a frame, and hence are inefficient and slow. Proposal-wise approaches only learn from the proposed object candidates, but still require multiple steps on the expensive proposal methods. Instead, this work presents a fast single-shot segmentation strategy for video scene understanding. The proposed net, called S3-Net, quickly locates and segments target sub-scenes , and meanwhile extracts attention-aware time-series sub-scene features ( ats-features ) as inputs to an attention-aware spatio-temporal model (ASM) . Utilizing tensorization and quantization techniques, S3-Net is intended to be lightweight for edge computing. Experiments results on CityScapes, UCF11, HMDB51, and MOMENTS datasets demonstrate that the proposed S3-Net achieves an accuracy improvement of 8.1% versus the 3D-CNN based approach on UCF11, a storage reduction of 6.9× and an inference speed of 22.8 FPS on CityScapes with a GTX1080Ti GPU. Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | Precise Power Capping for Latency-Sensitive Applications in DatacenterabstractPower capping is widely used in cloud datacenters to mitigate power over-provisioning problem, thus improve datacenter capacity and cut off their operation cost. However, inappropriate or aggressive power capping may lead to performance degradation of applications (especially latency-sensitive ones), and there are few effective methods that can accurately evaluate and control such negative impact caused by aggressive power capping. In this paper, we proposeFine-Grained Differential Method(FGD) to quantitatively analyze how inappropriate power capping degrades the performance of latency-sensitive applications. By using FGD, we can minimize the provisioned power for each server by setting a precise power budget according to application’sService Level Agreement(SLA). And we further proposePrecise Power Capping(PPCapping) which is designed to increase the datacenter capacity with a fixed power supply by means of FGD. Our research also provides an insight of precise tradeoff between applications’ SLAs and datacenter capacity. We verify FGD and PPCapping by using real world traces from Tencent’s datecenter with 25,328 servers. The experimental results show that FGD can accurately analyze the impact of power capping on the performance of latency-sensitive applications, and PPCapping can effectively increase datacenter capacity compared with the typical power provisioning strategy. Song Wu 0001, Xinhou Wang, Hai Jin 0001, Fangming Liu, Haibao Chen, Chuxiong Yan |
IEEE Trans. Sustain. Comput. | 6 |
| 2021 | A 3.85-Gb/s 8 × 8 Soft-Output MIMO Detector With Lattice-Reduction-Aided Channel PreprocessingabstractThis article presents an 8 × 8 lattice-reduction-aided (LRA) soft-output multiple-input multiple-output (MIMO) detector for Chinese enhanced ultrahigh throughput (EUHT) wireless local area network (LAN) standard. The preprocessing algorithm combining simplified-sorting Cholesky decomposition and low-complexity decoupled lattice reduction (LDLR) is proposed to reduce computational complexity and latency with parallelism improvement. In addition, K-best detection adopts a sorting-reduced strategy utilizing approximate ordered sequence. Compared with other published LRA K-best detection algorithms, simulation results show that our proposed algorithm has performance improvement. In addition, in order to save hardware resources, a folded K-best architecture and an optimized intermediate storage strategy are introduced. Furthermore, a fully pipelined VLSI architecture is designed in Semiconductor Manufacturing International Corporation (SMIC) 40-nm 1P9M technology to support the 8 × 8.64 -QAM MIMO-OFDM system. The detector can achieve 3.85-Gb/s data throughput at 641-MHz clock frequency with 0.71-μs latency. The proposed detector is competitive in terms of latency, throughput, and area efficiency to state-of-the-art works and can meet the data-rate requirement of the EUHT standard. Zhuojun Liang, Dongxu Lv, Chao Cui, Haibao Chen, Weifeng He, Weiguang Sheng, Naifeng Jing, Zhigang Mao, Guanghui He 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | An Anomaly Comprehension Neural Network for Surveillance Videos on Terminal DevicesabstractAnomaly comprehension in surveillance videos is more challenging than detection. This work introduces the design of a lightweight and fast anomaly comprehension neural network. For comprehension, a spatio-temporal LSTM model is developed based on the structured, tensorized time-series features extracted from surveillance videos. Deep compression of network size is achieved by tensorization and quantization for the implementation on terminal devices. Experiments on large-scale video anomaly dataset UCF-Crime demonstrate that the proposed network can achieve an impressive inference speed of 266 FPS on a GTX-1080Ti GPU, which is 4.29 faster than ConvLSTM-based method; a 3.34% AUC improvement with 5.55% accuracy niche versus the 3D-CNN based approach; and at least 15k× parameter reduction and 228× storage compression over the RNN-based approaches. Moreover, the proposed framework has been realized on an ARM-core based IOT board with only 2.4W power consumption. Guangtai Huang, Peining Zhen, Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
DATE | 5 |
| 2020 | DEEPEYE: A Deeply Tensor-Compressed Neural Network for Video Comprehension on Terminal DevicesabstractVideo object detection and action recognition typically require deep neural networks (DNNs) with huge number of parameters. It is thereby challenging to develop a DNN video comprehension unit in resource-constrained terminal devices. In this article, we introduce a deeply tensor-compressed video comprehension neural network, called DEEPEYE, for inference on terminal devices. Instead of building a Long Short-Term Memory (LSTM) network directly from high-dimensional raw video data input, we construct an LSTM-based spatio-temporal model from structured, tensorized time-series features for object detection and action recognition. A deep compression is achieved by tensor decomposition and trained quantization of the time-series feature-based LSTM network. We have implemented DEEPEYE on an ARM-core-based IOT board with 31 FPS consuming only 2.4W power. Using the video datasets MOMENTS, UCF11 and HMDB51 as benchmarks, DEEPEYE achieves a 228.1× model compression with only 0.47% mAP reduction; as well as 15 k × parameter reduction with up to 8.01% accuracy improvement over other competing approaches. Guangya Li, Ngai Wong 0001, Haibao Chen, Hao Yu 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2020 | BIA: Behavior Identification Algorithm Using Unsupervised Learning Based on Sensor Data for Home ElderlyabstractBehavior identification plays an important role in supporting homecare for the elderly living alone. In literature, plenty of algorithms have been designed to identify behaviors of the elderly by learning features or extracting patterns from sensor data. However, most of them adopted probabilistic models or supervised learning to identify behaviors based on labeled sensor data. This paper proposes a behavior identification algorithm (BIA) using unsupervised learning based on unlabeled sensor data for the elderly living alone in smart home. This paper presents the observation of elder behaviors with three features: Event Order, Time Length Similarity and Time Interval Similarity features. Based on these features of behavior observations, two properties of behaviors, including the Event Shift and Histogram Shape Similarity properties, are presented. According to these properties, the proposed BIA is developed. Finally, performance results show that the proposed BIA outperforms the existing unsupervised machine learning mechanisms in terms of the behavior identification precision and recall. Cuijuan Shang, Chih-Yung Chang, Guilin Chen, Haibao Chen |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | DEEPEYE: A Deeply Tensor-Compressed Neural Network Hardware Accelerator: Invited PaperabstractVideo detection and classification constantly involve high dimensional data that requires a deep neural network (DNN) with huge number of parameters. It is thereby quite challenging to develop a DNN video comprehension at terminal devices. In this paper, we introduce a deeply tensor compressed video comprehension neural network called DEEPEYE for inference at terminal devices. Instead of building a Long Short-Term Memory (LSTM) network directly from raw video data, we build a LSTM-based spatio-temporal model from tensorized time-series features for object detection and action recognition. Moreover, a deep compression is achieved by tensor decomposition and trained quantization of the time-series feature-based spatio-temporal model. We have implemented DEEPEYE on an ARM-core based IOT board with only 2.4W power consumption. Using the video datasets MOMENTS and UCF11 as benchmarks, DEEPEYE achieves a 228.1× model compression with only 0.47% mAP deduction; as well as 15k× parameter reduction yet 16.27% accuracy improvement. Guangya Li, Ngai Wong 0001, Haibao Chen, Hao Yu 0001 |
ICCAD | 4 |
| 2019 | Design for reliability with the advanced integrated circuit (IC) technology: challenges and opportunities
Zhigang Ji, Haibao Chen, Xiuyan Li |
Sci. China Inf. Sci. | 2 |
| 2019 | A large-scale in-memory computing for deep neural network with trained quantization
Haibao Chen, Hao Yu 0001 |
Integr. | 3 |
| 2019 | Scale Adaptive Proposal Network for Object Detection in Remote Sensing ImagesabstractObject detection in aerial images is widely applied in many applications. In recent years, faster region convolutional neural network shows a great improvement on object detecting in natural images. Considering the size and distribution characteristic of object in remote sensing images, the region proposal network (RPN) should be changed before being adopted. In this letter, a scale adaptive proposal network (SAPNet) is proposed to improve the accuracy of multiobject detection in remote sensing images. The SAPNet consists of multilayer RPNs which are designed to generate multiscale object proposals, and a final detection subnetwork in which fusion feature layer has been applied for better multiobject detection. Comparative experimental results show that the proposed SAPNet significantly improves the accuracy of multiobject detection. Guanghui He 0002, Haibao Chen, Naifeng Jing, Qin Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Thermal-Sensor-Based Occupancy Detection for Smart Buildings Using Machine-Learning MethodsabstractIn this article, we propose a novel approach to detect the occupancy behavior of a building through the temperature and/or possible heat source information. The new method can be used for energy reduction and security monitoring for emerging smart buildings. Our work is based on a building simulation program, EnergyPlus, from the Department of Energy. EnergyPlus can model various time-series inputs to a building such as ambient temperature; heating, ventilation, and air-conditioning (HVAC) inputs; power consumption of electronic equipment; lighting; and number of occupants in a room, sampled each hour, and produce resulting temperature traces of zones (rooms). Two machine-learning-based approaches for detecting human occupancy of a smart building are applied herein, namely support vector regression (SVR) and recurrent neural network (RNN). Experimental results with SVR show that the four-feature model provides accurate detection rates, giving a 0.638 average error and 5.32% error rate, and the five-feature model delivers a 0.317 average error and 2.64% error rate. This indicates that SVR is a viable option for occupancy detection. In the RNN method, Elman’s RNN can estimate occupancy information of each room of a building with high accuracy. It has local feedback in each layer and, for a five-zone building, it is very accurate for occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum, considering ambient, room temperatures, and HVAC powers as detectable information. Without knowing HVAC powers, the estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5. Our article further shows that both methods deliver similar accuracy in the occupancy detection. But the SVR model is more stable for adding or removing features of the system, while the RNN method can deliver more accuracy when the features used in the model do not change a lot. Hengyang Zhao, Qi Hua, Haibao Chen, Yaoyao Ye, Hai Wang 0002, Sheldon X.-D. Tan, Esteban Tlelo-Cuautle |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | Physics-Based Compact TDDB Models for Low-k BEOL Copper Interconnects With Time-Varying Voltage StressingabstractTime-dependent dielectric breakdown (TDDB) is one of the important failure mechanisms for copper (Cu) interconnects. This problem becomes more severe as the pitch between wires is shrinking and low-k dielectric materials (low electrical and mechanical strength) are used. Many TDDB models have been proposed based on different physics kinetics in the past. Recently, a physics-based TDDB model, which is based on the breakdown concept of electric path generation, has been proposed and has shown advantage over widely accepted existing electrostatic field-based TDDB assessment. However, determination of the time-to-failure from this model includes time-consuming finite-element method (FEM). In this paper, we try to mitigate this problem by developing fast time to failure evaluation method based on the closed form solution of the ion diffusion partial differential equations. We show that the location of the minimum concentration can be determined by the dominant terms with sufficient accuracy and the time to failure can also be computed with a few dominant terms. On top of this, we also consider the time-varying stressing voltages, which is commonly seen in practical VLSI chips. We propose to develop the equivalent dc stressing voltage, which is parameterized in terms of amplitude, duty cycle, and period for periodic stressing voltage waveforms using regression-based method. We further validate the proposed analytic TDDB concentration and time to failure formula, and the equivalent dc stressing voltage compact model against the results of an FEM analysis using COMSOL. Numerical results further show that the new compact TDDB model can lead to three orders of magnitude speedup with less than 1% error against the existing FEM results. Shaoyi Peng, Han Zhou 0002, Taeyoung Kim 0001, Haibao Chen, Sheldon X.-D. Tan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | ACStor: Optimizing Access Performance of Virtual Disk Images in CloudsabstractIn virtualized data centers, virtual disk images (VDIs) serve as the containers in virtual environment, so their access performance is critical for the overall system performance. Some distributed VDI chunk storage systems have been proposed in order to alleviate the I/O bottleneck for VM management. As the system scales up to a large number of running VMs, however, the overall network traffic would become unbalanced with hot spots on some VMs inevitably, leading to I/O performance degradation when accessing the VMs. In this paper, we propose an adaptive and collaborative VDI storage system (ACStor) to resolve the above performance issue. In comparison with the existing research, our solution is able to dynamically balance the traffic workloads in accessing VDI chunks, based on the run-time network state. Specifically, compute nodes with lightly loaded traffic will be adaptively assigned more chunk access requests from remote VMs and vice versa, which can effectively eliminate the above problem and thus improves the I/O performance of VMs. We implement a prototype based on our ACStor design, and evaluate it by various benchmarks on a real cluster with 32 nodes and a simulated platform with 256 nodes. Experiments show that under different network traffic patterns of data centers, our solution achieves up to 2-8× performance gain on VM booting time and VM's I/O throughput, in comparison with the other state-of-the-art approaches. Song Wu 0001, Sheng Di, Haibao Chen, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Energy and Lifetime Optimizations for Dark Silicon Manycore Microprocessor Considering Both Hard and Soft ErrorsabstractIn this paper, we propose a new energy and lifetime optimization techniques for emerging dark silicon manycore microprocessors considering both hard long-term reliability effects (hard errors) and transient soft errors, which have been studied less in the past. We consider a recently proposed physics-based electromigration (EM) reliability model to predict the EM-induced reliability. We employ both dynamic voltage and frequency scaling (DVFS) and dark silicon core state using ON/OFF switching action as the two control knobs. We show that on-chip power consumption has different (even contradicting) impacts on soft and hard reliability effects. This paper also shows that soft error should be mitigated by other techniques if aggressive low power and high long-term reliability are pursued. We focus on two optimization techniques for improving lifetime and reducing energy. To optimize EM-induced lifetime, we first apply the adaptive Q-learning-based method, which is suitable for dynamic runtime operation as it can provide cost-effective yet good solutions. The second lifetime optimization approach is the mixed-integer linear programming (MILP) method, which typically yields better solutions but at higher computational costs. To optimize the energy of a dark silicon chip subject to the both hard and soft reliability effects, power budgets, and performance limits, the Q-learning method has been applied as well. A large class of multithreaded applications is used as our benchmarks to validate and compare the proposed dynamic reliability management methods. Experimental results on a 64-core dark silicon chip show that the proposed DRM algorithm can effectively manage and optimize the lifetime of a dark silicon microprocessor under the given power budget and performance limit. Also, the proposed energy optimization can effectively manage and optimize energy consumption subject to both hard and soft-error rates, power budget, and performance limits as constraints. We also show that the under tightened power and performance constraints, we cannot satisfy both hard and soft errors at the same time as there is no simple tradeoff between performance/power and reliability in this case. Some other soft-error mitigation techniques are required in this case. Taeyoung Kim 0001, Zeyu Sun 0001, Haibao Chen, Hai Wang 0002, Sheldon X.-D. Tan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Electromigration recovery modeling and analysis under time-dependent current and temperature stressingabstractElectromigration (EM) has been considered to be the major reliability issue for current and future VLSI technologies. Current EM reliability analysis is overloaded by over-conservative and simplified EM models. Particularly the transient recovery effect in the EM-induced stress evolution kinetics has never been treated properly in all the existing analytical EM models. In this article, we propose a new physics-based dynamic compact EM model, which for the first time, can accurately predict the transient hydrostatic stress recovery effect in a confined metal wire. The new dynamic EM model is based on the direct analytical solution of one-dimensional Korhonen's equation with load driven by any unipolar or bipolar current waveforms under varying temperature. We show that the EM recovery effect can be quite significant even under unidirectional current loads. This healing process is sensitive to temperature, and higher temperatures lead to faster and more complete recovery. Such effect can be further exploited to significantly extend the lifetime of the interconnect wires if the chip current or power can be properly regulated and managed. As a result, the new dynamic EM model can be incorporated with existing dynamic thermal/power/reliability management and optimization approaches, devoted to reliability-aware optimization at multiple system levels (chip/server/rack/data centers). Presented results show that the proposed EM model agrees very well with the numerical analysis results under any time-varying current density and temperature profiles. Xin Huang 0003, Valeriy Sukharev, Taeyoung Kim 0001, Haibao Chen, Sheldon X.-D. Tan |
ASP-DAC | 4 |
| 2016 | Thermal modeling for energy-efficient smart building with advanced overfitting mitigation techniqueabstractBuilding energy accounts large amount of the total energy consumption, and smart building energy control leads to high energy efficiency and significant energy savings. A compact and accurate building thermal model is important for designing the efficient energy control system. In this paper, we propose an accurate thermal behavior modeling technique for general and complicated buildings. This new modeling technique builds compact thermal model by system identification using temperature and power data obtained from EnergyPlus software, which can provide realistic temperature, weather and power data for buildings. In order to make the best use of data from EnergyPlus and avoid the overfitting problem associated with the system identification method, a cross-validation technique is employed to generate multiple thermal models to find the optimal model order. The final model is then generated by performing a regular system identification using the previously selected order. Experimental results from a case study of a 5-zone building have shown that the proposed method is able to find the optimal model order, and the building models built by the proposed method can achieve 1-3% average errors and less than 10-18% maximum errors for the estimation of zone temperatures for about a one year period. Wandi Liu, Hai Wang 0002, Hengyang Zhao, Shujuan Wang, Haibao Chen, Yuzhuo Fu, Jian Ma 0002, Xin Li 0001, Sheldon X.-D. Tan |
ASP-DAC | 5 |
| 2016 | Learning-based dynamic reliability management for dark silicon processor considering EM effects
Taeyoung Kim 0001, Xin Huang 0003, Haibao Chen, Valeriy Sukharev, Sheldon X.-D. Tan |
DATE | 3 |
| 2016 | Dynamic reliability management for near-threshold dark silicon processorsabstractIn this article, we propose a new dynamic reliability management (DRM) techniques at the system level for emerging low power dark silicon manycore microprocessors operating in near-threshold region. We mainly consider the electromigration (EM) failures. To leverage the EM recovery effects, which was ignored in the past, at the system-level, we propose a new equivalent DC current model to consider recovery effects for general time-varying current waveforms so that existing compact EM model can be applied. The new equivalent DC current is calculated in two steps: firstly, the equivalent square waveform is calculated so that peak and terminal stresses are matched, secondly, the parameterized equivalent DC current is derived in terms of the parameters of the periodic fitted square waveforms from the first step. The new recovery EM model can allow EM-induced lifetime to be better managed at the system level. The system level energy optimization problem considering EM lifetime subject to power and performance constraints is framed by seeking the best dark silicon cores' voltage and on/off status. The resulting problem is solved by the State-Action-Reward-State-Action (SARSA) reinforcement learning algorithm. Experimental results on a 64-core near-threshold dark silicon processor show that the new equivalent EM DC currents can fully exhibit the recovery effects at the system-level so that trade-off between EM lifetime and energy/performance can be easily made. We further show that the proposed learning-based energy optimization can effectively manage and optimize energy subject to reliability, given power budget and performance limits. When the recovery effects are considered, the new optimization method can achieve 8.6× longer lifetime at the costs of 2.0× more energy and 3.3× more performance degradation. Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Jagadeesh Gaddipati, Hai Wang 0002, Haibao Chen, Sheldon X.-D. Tan |
ICCAD | 6 |
| 2016 | Dynamic Acceleration of Parallel Applications in Cloud Platforms by Adaptive Time-Slice ControlabstractTightly-coupled parallel applications in cloud systems may suffer from significant performance degradation because of the resource over-commitment issue. In this paper, we propose a dynamic approach based on the adaptive control over time-slice for virtual clusters, in order to mitigate the performance degradation for parallel applications in cloud and avoid the negative impact effectively on other non-parallel applications meanwhile. The key idea is to reduce the synchronization overhead inside and across virtual machines (VMs) in cloud systems, by dynamically adjusting the time-slices of VMs in terms of the spinlock latency at runtime. Such a design is motivated by our experimental finding that VM's time slice is a key factor determining the synchronization overhead as well as the parallel execution performance. We perform the evaluation on a real cluster environment deployed with XEN, using five well-known benchmarks with 10+ applications. Experiments show that our approach obtains 1.5-10× performance gain for running parallel applications, than other state-of-the-art solutions (including Credit Scheduling of Xen and the well-known methods like Co-Scheduling and Balance Scheduling), with nearly unaffected impact on the performance of non-parallel applications. Song Wu 0001, Zhenjiang Xie, Haibao Chen, Sheng Di, Hai Jin 0001 |
IPDPS | 3 |
| 2016 | Learning-based occupancy behavior detection for smart buildingsabstractIn this article, we propose a novel method to detect the occupancy behavior of a building through the temperature and/or possible heat source information, which can be used for energy reduction, security monitoring for emerging smart buildings. Our work is based on a realistic building simulation program, EnergyPlus, from Department of Energy. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). The new approach is based on a learning based approach in which a recurrent neutral network (RNN) is trained to detect the number of people in a room based on the room temperature and other information such as ambient temperature, and other related heat sources. We applied the Elman's recurrent neural network (ELNN), which has local feedbacks in each layer. We use an empirical formula to calculate the RNN layer number and layer size to configure RNN architecture to avoid overfitting and under-fitting problems. Experimental results from a case study of a 5-zone building show that ELNN can lead to very accurate occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum when we consider ambient, room temperatures and HVAC powers as detectable information. Without knowing HVAC powers, estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5. Hengyang Zhao, Zhongdong Qi, Shujuan Wang, Kambiz Vafai, Hai Wang 0002, Haibao Chen, Sheldon X.-D. Tan |
ISCAS | 6 |
| 2016 | A survey of cloud resource management for complex engineering applications
Haibao Chen, Song Wu 0001, Hai Jin 0001, Jidong Zhai, Yingwei Luo, Xiaolin Wang 0001 |
Frontiers Comput. Sci. | 1 |
| 2016 | Analytical Modeling and Characterization of Electromigration Effects for Multibranch Interconnect TreesabstractElectromigration (EM) in very large scale integration (VLSI) interconnects has become one of the major reliability issues for current and future VLSI technologies. However, existing EM modeling and analysis techniques are mainly developed for a single wire. For practical VLSI chips, the elemental EM reliability unit called interconnect tree is a multibranch interconnect segment consisting of a continuously connected, highly conductive metal (Cu) lines terminated by diffusion barriers and located within the single level of metallization. The EM effects in those branches are not independent and have to be considered simultaneously. In this paper, we demonstrate, for the first time, a first principle-based analytical solution of this problem. We have derived the analytical expressions describing the hydrostatic stress evolution in several typical interconnect trees: 1) the straight-line three-terminal wires; 2) the T-shaped four-terminal wires; and 3) the cross-shaped five-terminal wires. The new approach solves the stress evolution in a multibranch tree by de-coupling the individual segments through the proper boundary conditions (BCs) accounting the interactions between different branches. By using Laplace transformation technique, analytical solutions are obtained for each type of the interconnect trees. The analytical solutions in terms of a set of auxiliary basis functions using the complementary error function agree well with the numerical analysis results. Our analysis further demonstrates that using the first two dominant basis functions can lead to 0.5% error, which is sufficient for practical EM analysis. Haibao Chen, Sheldon X.-D. Tan, Xin Huang 0003, Taeyoung Kim 0001, Valeriy Sukharev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Statistical Rare-Event Analysis and Parameter Guidance by Elite Learning Sample SelectionabstractAccurately estimating the failure region of rare events for memory-cell and analog circuit blocks under process variations is a challenging task. In this article, we propose a new statistical method, called EliteScope , to estimate the circuit failure rates in rare-event regions and to provide conditions of parameters to achieve targeted performance. The new method is based on the iterative blockade framework to reduce the number of samples, but consists of two new techniques to improve existing methods. First, the new approach employs an elite-learning sample-selection scheme, which can consider the effectiveness of samples and well coverage for the parameter space. As a result, it can reduce additional simulation costs by pruning less effective samples while keeping the accuracy of failure estimation. Second, the EliteScope identifies the failure regions in terms of parameter spaces to provide a good design guidance to accomplish the performance target. It applies variance-based feature selection to find the dominant parameters and then determine the in-spec boundaries of those parameters. We demonstrate the advantage of our proposed method using several memory and analog circuits with different numbers of process parameters. Experiments on four circuit examples show that EliteScope achieves a significant improvement on failure-region estimation in terms of accuracy and simulation cost over traditional approaches. The 16b 6T-SRAM column example also demonstrates that the new method is scalable for handling large problems with large numbers of process variables. Taeyoung Kim 0001, Hosoon Shin, Sheldon X.-D. Tan, Xin Li 0001, Haibao Chen, Hai Wang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2015 | New electromigration modeling and analysis considering time-varying temperature and current densitiesabstractElectromigration (EM) is projected to be the major reliability issue for current and future VLSI technologies. However, existing EM models and assessment techniques are mainly based on the constant current density and temperature. Such models will not work well at the system level as the current density (power) and temperature are changing with time due to different tasks (their loans) applied at run time. Existing EM approaches using average current density or temperature, however, will lead to significant errors as shown in this work. In this paper, we propose a new physics-based EM model considering time-varying temperature and current density, which reflects a more practical chip working conditions especially for multi-core and emerging 3D ICs. We study the impacts of the time-varying current densities and temperature profiles on EM-induced lifetime of a wire for both nucleation phase and growth phase. We propose a fast stress calculation method for given time-varying temperature and current densities for the nucleation phase. We further develop new formulae to compute the resistance changes in growth phase due to changing temperature and current densities. Experimental results show that the proposed method shows an excellent agreement with the detailed numerical analysis but with much improved efficiency. Haibao Chen, Sheldon X.-D. Tan, Xin Huang 0003, Valeriy Sukharev |
ASP-DAC | 1 |
| 2015 | Interconnect reliability modeling and analysis for multi-branch interconnect treesabstractElectromigration (EM) in VLSI interconnects has become one of the major reliability issues for current and future VLSI technologies. However, existing EM modeling and analysis techniques are mainly developed for a single wire. For practical VLSI chips, the interconnects such as clock and power grid networks typically consist of multi-branch metal segments representing a continuously connected, highly conductive metal (Cu) lines within one layer of metallization, terminating at diffusion barriers. The EM effects in those branches are not independent and they have to be considered simultaneously. In this paper, we demonstrate, for the first time, a first principle based analytic solution of this problem. We investigate the analytic expressions describing the hydrostatic stress evolution in several typical interconnect trees: the straight-line 3-terminal wires, the T-shaped 4-terminal wires and the cross-shaped 5-terminal wires. The new approach solves the stress evolution in a multi-branch tree by de-coupling the individual segments through the proper boundary conditions accounting the interactions between different branches. By using Laplace transformation technique, analytical solutions are obtained for each type of the interconnect trees. The analytical solutions in terms of a set of auxiliary basis functions using the complementary error function agree well with the numerical analysis results. Our analysis further demonstrates that using the first two dominant basis functions can lead to 0.5% error, which is sufficient for practical EM analysis. Haibao Chen, Sheldon X.-D. Tan, Valeriy Sukharev, Xin Huang 0003, Taeyoung Kim 0001 |
DAC | 1 |
| 2015 | Learning Based Compact Thermal Modeling for Energy-Efficient Smart Building Management: (invited)abstractIn this article, we propose a new behavioral thermal modeling method for fast building performance analysis, which is critical for energy-efficient smart building control and management. The new approach is based on two recurrent neutral network architecture to obtain the compact nonlinear thermal models for complicated building. We start with a more realistic building simulation program, EnergyPlus, from Department of Energy, to model some practical buildings such as office buildings and data centers. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). In this work, we apply two recurrent neural network (RNN) architectures to build the non-linear compact thermal model of the building: one is non-linear state-space RNN architecture (NLSS), which has global feedbacks, and the other one is Elman's RNN architecture (ELNN), which has local feedbacks in each layer. We give a simple formula to calculate the RNN layer number, layer size to configure RNN architecture to avoid overfitting and underfitting problems. A cross-validation based training technique is further applied to improve predictable accuracy of models. Experimental results from a case study of three buildings show that ELNN and NLSS can both build very accurate building thermal models for the 2-zone and 5-zone building cases: both of them have average errors from around 1% to 1.5% for the two buildings. For the more complex 6-zone building case, ELNN outperforms NLSS with maximum errors 16% against 23%. But both methods have 2.2% average errors. Hengyang Zhao, Daniel Quach, Shujuan Wang, Hai Wang 0002, Haibao Chen, Xin Li 0001, Sheldon X.-D. Tan |
ICCAD | 5 |
| 2015 | Evaluating Latency-Sensitive Applications: Performance Degradation in Datacenters with Restricted Power BudgetabstractFor data centers with limited power supply, restricting the servers' power budget (i.e., The maximal power provided to servers) is an efficient approach to increase the server density (the server quantity per rack), which can effectively improve the cost-effectiveness of the data centers. However, this approach may also affect the performance of applications in servers. Hence, the prerequisite of adopting the approach in data centers is to precisely evaluate the application performance degradation caused by restricting the servers' power budget. Unfortunately, existing evaluation methods are inaccurate because they are either improper or coarse-grained, especially for the latency-sensitive applications widely deployed in data centers. In this paper, we analyze the reasons why state-of-the-art methods are not appropriate for evaluating the performance degradation of latency-sensitive applications in case of power restriction, and we propose a new evaluation method which can provide a fine-grained way to precisely describe and evaluate such degradation. We verify our proposed method by a real-world application and the traces from Ten cent's date enter with 25328 servers. The experimental results show that our method is much more accurate compared with the state of the art, and we can significantly increase datacenter efficiency by saving servers' power budget while maintaining the applications' performance degradation within controllable and acceptable range. Song Wu 0001, Chuxiong Yan, Haibao Chen, Hai Jin 0001, Deqing Zou |
ICPP | 3 |
| 2015 | Towards establishing a meaningful and practical dynamics results for the unified RNN model
Chen Qiao, Haibao Chen, Wenfeng Jing, Ke-Feng Sun |
Neurocomputing | 2 |
| 2015 | H-Matrix-Based Finite-Element-Based Thermal Analysis for 3D ICsabstractIn this article, we propose an efficient finite-element-based (FE-based) method for both steady and transient thermal analyses of high-performance integrated circuits based on the hierarchical matrix ( H -matrix) representation. H -matrix has been shown to provide a data-sparse way to approximate the matrices and their inverses with almost linear-space and time complexities. In this work, we apply the H -matrix concept for solving heating diffusion problems modeled by parabolic partial differential equations (PDEs) based on the finite element method. We show that the matrix from a FE-based steady and transient thermal analysis can be represented by H -matrix without any approximation, and its inverse and Cholesky factors can be evaluated by H -matrix with controlled accuracy. We then show and prove that the memory and time complexities of the solver are bounded by O ( k 1 N log N ) and O ( k 1 2 N log 2 N ), respectively, where k 1 is a small quantity determined by accuracy requirements and N is the number of unknowns in the system. The comparison with existing product-quality LU solvers, CSPARSE and UMFPACK, on a number of 3D IC thermal matrices, shows that the new method is much more memory efficient than these methods, which however prevents CPU time comparison with those methods on large examples. But the proposed method can solve all the given thermal circuits with decent scalabilities, which shows good agreement with the predicted theoretical results. Haibao Chen, Ying-Chi Li, Sheldon X.-D. Tan, Xin Huang 0003, Hai Wang 0002, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2015 | Synchronization-Aware Scheduling for Virtual Clusters in CloudabstractDue to high flexibility and cost-effectiveness, cloud computing is increasingly being explored as an alternative to local clusters by academic and commercial users. Recent research already confirmed the feasibility of running tightly-coupled parallel applications with virtual clusters. However, such types of applications suffer from significant performance degradation, especially as the over-commitment is common in cloud. That is, the number of executable Virtual CPUs (VCPUs) is often larger than that of available Physical CPUs (PCPUs) in the system. The performance degradation is mainly due to the fact that the current virtual machine monitors (VMMs) are unaware of the synchronization requirements of the VMs which are running parallel applications. In this paper, There are two key contributions. (1) We propose an autonomous synchronization-aware VM scheduling (SVS) algorithm, which can effectively mitigate the performance degradation of tightly-coupled parallel applications running atop them in over-committed situation. (2) We integrate the SVS algorithm into Xen VMM scheduler, and rigorously implement a prototype. We evaluate our design on a real cluster environment with NPB benchmark and real-world trace. Experiments show that our solution attains better performance for tightly-coupled parallel applications than the state-of-the-art approaches like Xen's Credit scheduler, balance scheduling, and hybrid scheduling. Song Wu 0001, Haibao Chen, Sheng Di, Bing Bing Zhou, Zhenjiang Xie, Hai Jin 0001, Xuanhua Shi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Communication-driven scheduling for virtual clusters in cloudabstractDue to high flexibility and cost-effectiveness, cloud computing is increasingly being explored as an alternative to local clusters by academic and commercial users. Recent research already confirmed the feasibility of running tightly-coupled parallel applications with virtual clusters. However, such types of applications suffer from significant performance degradation, especially as the overcommitment is common in cloud. That is, the number of executable Virtual CPUs (VCPUs) is often larger than that of available Physical CPUs (PCPUs) in the system. The performance degradation mainly results from that the current Virtual Machine Monitors (VMMs) cannot co-schedule (or coordinate at the same time) the VCPUs that host parallel application threads/processes with synchronization requirements. Haibao Chen, Song Wu 0001, Sheng Di, Bing Bing Zhou, Zhenjiang Xie, Hai Jin 0001, Xuanhua Shi |
HPDC | 1 |
| 2014 | Lifetime optimization for real-time embedded systems considering electromigration effectsabstractIn this article, we propose a new lifetime task optimization technique for real-time embedded processors considering the electromigration-induced reliability. The new approach is based on a recently proposed physics-based electromigration (EM) model for more accurate EM assessment of a power grid network at the chip level. We apply the dynamic voltage and frequency scaling (DVFS) (by selecting the performance states or p-states of the tasks to manage the power) and thus the lifetime of the processor running different tasks over their periods. We consider both single-rate and multi-rate embedded systems with preemption. To model the mean-time-to-failure (MTTF) of a task for a given p-state, response surface modeling is applied. We then frame the reliability optimization problem as the continuous constrained nonlinear optimization problem in which the system EM-induced reliability is maximized subject to the timing constraints, which is further solved by simulated annealing method. Experimental results show that for low utilization systems, significant reliability improvement can be achieved with even smaller power consumption than existing reliability-ignore scheduling method. The proposed method can lead to near Pareto's front trade-off between the power/energy and the lifetime compared to the existing task scheduling method. Taeyoung Kim 0001, Bowen Zheng 0001, Haibao Chen, Qi Zhu 0002, Valeriy Sukharev, Sheldon X.-D. Tan |
ICCAD | 3 |
| 2014 | IR-drop based electromigration assessment: parametric failure chip-scale analysisabstractThis paper presents a novel approach and techniques for electromigration (EM) assessment in power delivery networks. An increase in the voltage drop above the threshold level, caused by EM-induced increase in resistances of the individual interconnect segments, is considered as a failure criterion. This criterion replaces a currently employed conservative weakest segment criterion, which does not account an essential redundancy for current propagation existing in the power-ground (p/g) networks. EM-induced increase in the resistance of the individual grid segments is described in the approximation of the physics-based formalism for void nucleation and growth. A developed technique for calculating the hydrostatic stress distribution inside a multi branch interconnect tree allows to avoid over optimistic prediction of the time to failure made with the Blech-Black analysis of individual branches of interconnect segment. Experimental results obtained on the IBM benchmark circuit validate the proposed methods. Valeriy Sukharev, Xin Huang 0003, Haibao Chen, Sheldon X.-D. Tan |
ICCAD | 3 |
| 2014 | Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICsabstractThermal issue is the leading design constraint for 3-D stacked integrated circuits (ICs) and through silicon vias (TSVs) are used to effectively reduce the temperature of 3-D ICs. Normally, TSV is considered as a good thermal conductor in its vertical direction, and its vertical thermal resistance has been well modeled. However, lateral heat transfer of TSVs, which is also important, was largely ignored in the past. In this paper, we propose an accurate physics-based model for lateral thermal resistance of TSVs in terms of physical and material parameters, and study the conditions for model accuracy. For TSV arrays or farm, we show that the space or pitch between TSVs has a significant impact on TSV thermal behavior and should be properly considered in the TSV models. The proposed lateral thermal resistance model is fully compatible with the existing modeling approaches, and thus we could build a more accurate complete TSV thermal model. The new TSV thermal model can be easily integrated into a finite difference (FD) based thermal analysis framework to improve analysis efficiency. The accuracy of the model is validated against a commercial finite element tool-COMSOL. Experimental results show that the improved TSV thermal model (with proposed lateral thermal model) could greatly improve the accuracy of FD method in thermal simulation comparing with the existing method. Zao Liu, Sahana Swarup, Sheldon X.-D. Tan, Haibao Chen, Hai Wang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Application of General Orthogonal Polynomials to Fast Simulation of Nonlinear Descriptor Systems Through Piecewise-Linear ApproximationabstractIn this letter, we report an approach combining piecewise-linear (PWL) approximation with general orthogonal polynomials to efficiently simulate large nonlinear descriptor systems in time-domain. The main idea of this approach is first to approximate a nonlinear function by a piecewise-linear representation. Then, using the recursive formulae of general orthogonal polynomials, orthonormal bases can be produced for fast simulation of the PWL model. The effectiveness of our approach is demonstrated on two nonlinear circuit models. Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |