VLDB 2026 Research / reviewers in the wild / expert
Zhigang Ji
dblp:60/265
· DBLP profile ↗
30ranked-venue papers
1as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 17 since 2021Computer networks · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRACE: A Transferable Framework for Aging-aware Cell Delay EstimationabstractWith the continuous scaling of integrated circuits and the miniaturization of semiconductor devices, reliability issues have become increasingly critical. Aging delay prediction based on standard cells is essential for accurate circuit timing analysis. However, the growing diversity of process technology combinations poses significant challenges to the generalization capability of existing AI-based prediction methods. To address this, we propose a novel framework that first employs a graph neural network (GNN) to train a pre-trained model for delay prediction. Building upon this pre-trained model, we introduce a multi-task learning strategy combined with transfer learning to accelerate the training process and enhance adaptability across varying process conditions. This approach culminates in a unified model capable of accurate and efficient post-aging delay estimation. Experiments show that our method accelerates the simulation process by 17,025× compared to SPICE. At the same time, it achieves prediction accuracy comparable to the current state-of-the-art, while requiring 250× less data for training, substantially reducing computational resources. Muyan Jin, Yunlin Liu, Zejian Cai, Pengpeng Ren, Zhigang Ji |
DATE | 6 |
| 2026 | Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsabstractTensor parallelism (TP) in large-scale LLM inference and training introduces frequent collective operations that dominate inter-GPU communication. While in-switch computing, exemplified by NVLink SHARP (NVLS), accelerates collective operations by reducing redundant data transfer, its communication-centric design philosophy introduces the mismatch between its communication mode and the memory semantic requirement of LLM's computation kernel. Such a mismatch isolates the compute and communication phases, resulting in underutilized resources and limited overlap in multi-GPU systems. To address the limitation, we propose CAIS, the first ComputeAware In-Switch computing framework that aligns communication modes with computation's memory semantics requirement. CAIS consists of three integral techniques: (1) compute-aware ISA and microarchitecture extension to enable compute-aware in-switch computing. (2) merging-aware TB (Thread Block) coordination to improve the temporal alignment for efficient request merging. (3) graph-level dataflow optimizer to achieve a tight cross-kernel overlap. Evaluations on LLM workloads show that CAIS achieves$1.38 \times$average end-to-end training speedup over the SOTA NVLS-enabled solution, and$1.61 \times$over T3, the SOTA compute-communicate overlap solutions but do not leverage NVLS, demonstrating its effectiveness in accelerating TP on multi-GPU systems. Chen Zhang 0001, Qijun Zhang, Zhuoshan Zhou, Yijia Diao, Zhipeng Tu, Zhuoran Song, Zhigang Ji, Jingwen Leng, Minyi Guo |
HPCA | 11 |
| 2026 | Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUsabstractMixture-of-Experts (MoE) has been adopted by many leading large models to reduce computational requirements. However, frequent inter-GPU communication in MoE expert parallelism (EP) becomes a performance challenge. We observe substantial redundant inter-GPU data transfers in MoE that can be potentially addressed by in-switch computing. Unfortunately, the existing solution, NVLink SHARP (NVLS), can only support static collectives with regular patterns, incapable of dynamic communication with irregular patterns in MoE. To bridge the functionality gap, we propose DySHARP, an integral dynamic in-switch computing solution to accelerate MoE, encompassing both communication primitives and communication-aware scheduling: 1) Dynamic multimem addressing co-designs ISA, architecture, and runtime, as a dynamic extension to NVLS, reducing redundant traffic. However, the resulting traffic reduction is inherently asymmetric between two directions, preventing it from directly translating into speedup. 2) Token-centric kernel fusion deeply fuses the dispatch-computation-combine pipeline, resolving this asymmetry to translate traffic reduction into actual speedup. Compared with the state-of-the-art solution, DySHARP achieves up to 1.79× speedup. Qijun Zhang, Chen Zhang 0001, Zhuoshan Zhou, Zhipeng Tu, Guangyu Sun 0003, Zhiyao Xie, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He 0002, Minyi Guo |
ISCA | 10 |
| 2026 | MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
Zhuoshan Zhou, Chen Zhang 0001, Qijun Zhang, Zhe Zhou 0002, Zhipeng Tu, Guangyu Sun 0003, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He 0002, Minyi Guo |
ISCA | 10 |
| 2026 | Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center NetworksabstractHigh-performance artificial intelligence (AI) applications impose stringent reliability requirements on AI data center networks (DCNs), yet link failures are almost inevitable and can severely disrupt AI workloads such as large language model (LLM) training and inference. Existing deployed link failure detection and recovery mechanisms suffer from slow execution speed and limited failure coverage, failing to meet the demands of production AI DCNs. To address these issues, we propose Delphinus, an ultra-fast link failure detection and recovery solution built on the data-plane of programmable switches. It achieves ultra-fast failure detection via hardware-based port state monitoring, extends recoverable failure coverage through remote failure notification and relay, and enables fast recovery by path switchover. Delphinus can serve as a key generic function of switches, providing host-transparent link failure handling for Ethernet fabrics. We implement Delphinus on commercial hardware switches, and deploy it in large-scale production AI DCNs for over a year. Extensive evaluations demonstrate that Delphinus can complete link failure detection and recovery within sub-milliseconds, with negligible impact on application performance and imperceptible service interruption. Junye Zhang, Zhigang Ji, Kefei Liu, Rui Zhuang, Ruixue Wang, Weiqiang Cheng, Zixuan Guan |
SIGCOMM | 3 |
| 2025 | DuQTTA: Dual Quantized Tensor-Train Adaptation with Decoupling Magnitude-Direction for Efficient Fine-Tuning of LLMsabstractRecent parameter-efficient fine-tuning (PEFT) techniques have enabled large language models (LLMs) to be efficiently fine-tuned for specific tasks, while maintaining model performance with minimal additional trainable parameters. However, existing PEFT techniques continue to face challenges in balancing both accuracy and efficiency, especially when addressing scalability and the demands of lightweight deployment for LLMs. In this paper, we propose an efficient fine-tuning method of LLMs based on dual quantized Tensor-Train adaptation with decoupling magnitude-direction (DuQTTA). The proposed DuQTTA method employs Tensor-Train decomposition and dual-stage quantization to minimize model size and resource consumption. Additionally, it employs an adaptive optimization strategy and a decoupled update mechanism to improve model performance, thereby minimizing suboptimal outcomes and ensuring alignment with the full-parameter fine-tuning goals. Experimental results indicate that the proposed DuQTTA method outperforms existing PEFT methods, achieving up to a $65 \times$ compression rate compared to the LLaMA2-7B models, meanwhile delivering improvements of $4.44 \%, 3.14 \%$, and 0.97% over LoRA on LLaMA2-7B, LLaMA3-8B, and LLaMA2-13B, respectively. The proposed DuQTTA method is effective in compressing LLMs for deployment on resource-constrained edge devices. Haoyan Dong, Haibao Chen, Jingjing Chang, Yixin Yang 0004, Ziyang Gao, Zhigang Ji, Runsheng Wang, Ru Huang 0001 |
DAC | 6 |
| 2025 | Equivalent Lumped Element Model for Electromigration Considering Thermal EffectsabstractElectromigration (EM) remains a critical reliability concern in advanced integrated circuit design. Traditional physics-based approaches, which solve partial differential equations (PDEs), are computationally intensive, particularly in multi-physics scenarios. To address the issue, we propose a self-consistent lumped element modeling framework that leverages the equivalence between electrical behavior and stress evolution to forecast EM-induced stress under coupled electro-thermomechanical effects. Thermomechanical interactions driven by temperature gradients are explicitly modeled using embedded controlled sources. A threshold-activated switching mechanism is proposed to dynamically reconfigure circuit topology, enabling seamless simulation across both void nucleation and post-voiding phases. The proposed adaptive non-uniform spatial discretization framework can be used to enhance computational efficiency without sacrificing accuracy. Numerical results demonstrate >50× speedup against the finite element simulation for small interconnects with <1.5% error, and 3.11× acceleration over conventional equivalent circuits for large-scale structures while maintaining <0.5% error. Fully compatible with standard SPICE solver, the proposed approach exhibits strong potential for temperature-aware EM analysis and void prediction in full-chip VLSI applications. Hengyi Zhu, Tianshu Hou, Zhigang Ji, Runsheng Wang, Haibao Chen |
ICCAD | 4 |
| 2025 | Enhanced memory window and reliability of α-IGZO FeFET enabled by atomic-layer-deposited HfO2 interfacial layer
Yinchi Liu, Jining Yang, Yeye Guo, Dmitriy Anatolyevich Golosov, Chenjie Gu, Zhigang Ji, Shijin Ding |
Sci. China Inf. Sci. | 10 |
| 2025 | DSTC: Dual-Side Sparse Tensor Core for DNNs Acceleration on Modern GPU ArchitecturesabstractLeveraging sparsity in deep neural network (DNN) models holds significant promise for accelerating model inference. However, current GPUs can only harness sparsity in model weights, leaving activations unutilized due to their dynamic and unpredictable nature, which poses a considerable challenge for exploitation. In our research, we introduce a novel architectural approach aimed at effectively leveraging dual-side sparsity, encompassing both weight and activation sparsity. Our methodology involves a systematic examination of previous sparsity-related architectures, and culminating in the proposal of an uncharted paradigm that combines outer-product computation primitive and bitmap-based encoding format. Our approach showcases feasibility through minimal modifications to existing production-scale inner-product-based Tensor Cores. We introduce a set of innovative ISA extensions and carefully co-design matrix-matrix multiplication and convolution algorithms, the two predominant computation patterns in contemporary DNN models, to exploit our novel dual-side sparse Tensor Core. Our evaluation demonstrates the efficacy of our design, unlocking the full potential of dual-side DNN sparsity and delivering performance enhancements of up to an order of magnitude while incurring only modest hardware overhead. Chen Zhang 0001, Yang Wang 0053, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng, Zhigang Ji, Yuan Xie 0001, Ru Huang 0001 |
IEEE Trans. Computers | 7 |
| 2025 | Fine-Grained Structured Sparse Computing for FPGA-Based AI InferenceabstractWith the explosive growth in the number of parameters in deep neural networks (DNNs), sparsity-centric algorithm and hardware designs have become critical for low-latency AI serving systems. However, the inherent randomness in pruning methods often leads to fragmented data access and irregular computation patterns in sparse matrices, resulting in significantly reduced hardware efficiency. Addressing the balance between the ‘randomness’ required to maintain model accuracy and the ‘regularity’ needed for efficient hardware design is crucial for realizing effective sparse computing in AI. This article proposes a fine-grained structured sparsity (FSS) paradigm. The pruned sparse matrices in this paradigm exhibit characteristics of ‘local randomness’ and ‘global regularity’. This dual-feature design allows AI accelerator hardware based on the FSS paradigm to maintain both high model accuracy and efficient hardware design. We implemented this novel accelerator on the Xilinx Alveo U280 and validated our concept across three different AI models, including CNN, RNN, and LLM, demonstrating performance that significantly outperforms prior methods. Chen Zhang 0001, Shijie Cao, Guohao Dai 0001, Chenbo Geng, Zhuliang Yao, Wencong Xiao, Yunxin Liu 0001, Ming Wu 0007, Guangyu Sun 0003, Zhigang Ji, Runsheng Wang, Ru Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2025 | Physics-Informed Learning Based Multiphysics Simulation for Fast Transient TSV Electromigration AnalysisabstractThrough Silicon Vias (TSVs) are vulnerable to electromigration (EM) degradation due to their high local current densities, thereby reducing the reliability of 3D ICs with stack dies and TSVs. Due to the broad application of 3D ICs, it is necessary to analyze the electromigration reliability of TSVs. To overcome the weakness of traditional method for EM modeling of TSVs, we propose a physics-informed learning approach for transient analysis of electromigration modeling in TSV by solving the conventional mass balance equation. The proposed method allows simultaneous consideration of atomic depletion and accumulation, effective resistance degradation, electric current evolution, and stress distribution. In particular, we propose a customized neural network to simulate the EM process in TSV without the need for fine grid meshing and temporal iteration in traditional methods. Considering that the loss function of the proposed model is a combination of different loss terms, we propose a modified self-adaptive loss balanced method to automatically adjust the weights of multiple loss terms to enhance network performance. Given the prediction uncertainty due to data randomness or model architecture constraints, Gaussian probabilistic model is constructed to define the self-adaptive weights and update the dynamic weights per epoch built on maximum likelihood estimation. Compared with the finite element method, the proposed physics informed neural network method can lead to a speedup with less than 0.1% mean square error. Experimental results also show that the proposed model achieves excellent performance over other competing methods and high robustness under values of initial weights, different numbers of hidden layers and neurons per layer. Xiaoman Yang, Haibao Chen, Yuhan Zhang 0005, Tianshu Hou, Pengpeng Ren, Runsheng Wang, Zhigang Ji, Ru Huang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2024 | CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding ResiduesabstractAccurate identification of protein nucleic acid binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a single model that could ignore either the semantic context of the protein or the global 3D geometric information. Consequently, these approaches may result in incomplete or inaccurate protein analysis. To address the above issue, in this paper, we present CrossBind, a novel collaborative cross modal approach for identifying binding residues by exploiting both protein geometric structure and its sequence prior knowledge extracted from a large scale protein language model. Specifically, our multi modal approach leverages a contrastive learning technique and atom wise attention to capture the positional relationships between atoms and residues, thereby incorporating fine grained local geometric knowledge, for better binding residue prediction. Extensive experimental results demonstrate that our approach outperforms the next best state of the art methods, GraphSite and GraphBind, on DNA and RNA datasets by 10.8/17.3% in terms of the harmonic mean of precision and recall (F1 Score) and 11.9/24.8% in Matthews correlation coefficient (MCC), respectively. We release the code at https://github.com/BEAM-Labs/CrossBind. Linglin Jing, Yifan Wang 0008, Zhigang Ji, Hui Fang 0003, Zhen Li 0026 |
AAAI | 6 |
| 2024 | Physics-Informed Learning for EPG-Based TDDB AssessmentabstractTime-dependent dielectric breakdown (TDDB) is one of the important failure mechanisms for copper (Cu) interconnects. Many TDDB models have been proposed based on different physics kinetics in the past. Recently, a physics-based TDDB model, which is based on the breakdown concept of electric path generation (EPG), has been proposed and has shown advantage over widely accepted existing electrostatic field-based TDDB assessment. However, the determination of the time-to-failure from this EPG based TDDB model includes solving partial differential equation (PDE) with time-consuming finite-element method (FEM). In recent years, deep neural networks have been proposed to predict numerical solutions of PDEs. In this paper, we use physics-informed neural network to solve the diffusion equation of ions in an electric field extracted from EPG based TDDB model. The continuous definite condition and hard constrain optimization methods are used for improving the performance of PINN in terms of accuracy and speed. Compared with the FEM method, the proposed PINN method can lead to about 100 times speedup with less than 0.1% mean squared error. Dinghao Chen, Xiaoman Yang, Pengpeng Ren, Zhigang Ji, Haibao Chen |
ASPDAC | 5 |
| 2024 | Enforcing hard constraints in physics-informed learning for transient TSV electromigration analysisabstractDue to the high local current densities, Through Silicon Vias (TSVs) are susceptible to electromigration (EM) degradation, which reduces the reliability of integrated circuits. Unlike traditional methods for TSV modeling and simulation, this paper introduces a unified hard constraint physics-informed learning neural network approach, called HCPINN, for the transient analysis of electromigration in TSVs by solving the conventional mass balance equation. The proposed method allows simultaneous consideration of atomic depletion and accumulation, effective resistance degradation, electric current evolution, and stress distribution. Specifically, we propose a hard constraint method for solving partial differential equations (PDEs) with general boundary conditions (BCs) for transient TSV electromigration analysis. By using the extra fields derived from the mixed finite element method, we reconstruct the corresponding PDEs by transforming general BCs into linear forms. Based on this derivation, we embed general BCs of mass balance equation into the proposed ansatz and employ sub-networks for the approximation on general BCs. The main neural network is responsible for training the internal part of the problem domain without adding loss terms with BCs, overcoming the convergence issue due to unbalanced gradients among different loss terms. Besides, we theoretically demonstrate that this reformulation of general BCs can stabilize the training process. Experimental results indicate that the proposed HCPINN exhibits superior performance and reduces boundary error in TSV electromigration analysis. Compared to the finite element method, the proposed network achieves approximately 100 times faster inference with a minimal mean squared error increase of less than 0.1%. Xiaoman Yang, Haibao Chen, Yuhan Zhang 0005, Yongkang Xue, Pengpeng Ren, Runsheng Wang, Zhigang Ji, Ru Huang 0001 |
ICCAD | 8 |
| 2024 | Memory-facilitated Joint-space Shift Adaptation in Traffic ForecastingabstractTraffic forecasting, crucial for intelligent transport systems, faces significant challenges from distribution shifts due to the dynamic nature of traffic patterns. Although normalisation approaches have been proposed to address distribution shifts in other time-series forecasting tasks such as predicting electricity consumption load prediction or influenza-like illness patient number estimation, they fall short in handling the complex spatial and temporal shifts in traffic data. In this paper, we propose a novel memory-facilitated joint-space shift adaptation framework, ST-Align, to address this problem in traffic forecasting. ST-Align comprises two key components targeting the input and latent space, respectively: a memory-based data alignment module in the input space, and an end-to-end memory network structure dedicated to alignment within the latent space. This joint-space design enables our ST-Align framework to effectively capture and adapt to dynamic distribution shifts in both spatial and temporal dimensions, thus enhancing model performance. Extensive experiments on various real-world datasets and prediction backbones convincingly demonstrate the robustness and generalisability of our method. He Haitao, Gerald Schaefer, Zhigang Ji, Yifan Wang 0008, Hui Fang 0003 |
IJCNN | 4 |
| 2024 | A strong physical unclonable function with machine learning immunity for Internet of Things application
Pengpeng Ren, Yongkang Xue, Linglin Jing, Lining Zhang, Runsheng Wang, Zhigang Ji |
Sci. China Inf. Sci. | 6 |
| 2024 | A Memory-augmented Conditional Neural Process model for traffic predictionabstractThis paper presents the first neural process-based model for traffic prediction, the Memory-augmented Conditional Neural Process (MemCNP). Spatio-temporal traffic prediction involves predicting future traffic patterns based on historical traffic data and the road network structure. This problem remains a challenge due to the dynamic and heterogeneous nature of urban traffic. Existing models often struggle to capture these complexities, particularly in data-limited scenarios. To address these limitations, our model presents a novel framework for uncertainty estimation based on the conditional neural process, and further incorporates a memory network module designed to acquire a representative contextual reference, thereby improving model performance under complex data distributions. By integrating the conditional neural process and the memory network, MemCNP enables the learning of the most representative contexts through iterative updates, enhancing the model’s generalisability. This allows our model to be applicable beyond car traffic, effectively handling diverse real-world traffic scenarios, including urban non-motorised traffic such as cycling, which is essential for advancing more sustainable transportation systems. This is demonstrated by comprehensive experimental results on six benchmark datasets (PeMS04, PeMS07, PeMS08, NYCTaxi, CHIBike, and T-Drive) against existing state-of-the-art traffic prediction models, where MemCNP demonstrates superior performance. Additionally, through ablation and reliability studies, we provide a comprehensive analysis of the model’s effectiveness. • The first neural process-based model, MemCNP, for traffic prediction with limited data. • MemCNP introduces a novel framework for uncertainty estimation. • A novel memory network module acquires a representative contextual reference. • MemCNP is effective across diverse scenarios, including non-motorised traffic. He Haitao, Kunhao Yuan, Gerald Schaefer, Zhigang Ji, Hui Fang 0003 |
Knowl. Based Syst. | 5 |
| 2024 | Fast Aging-Aware Timing Analysis Framework With Temporal-Spatial Graph Neural NetworkabstractWith the downscaling of CMOS technology, device aging induced by hot carrier injection and bias temperature instability effects poses severe challenges to timing analysis of digital circuits. In this work, a fast aging-aware timing analysis framework based on temporal–spatial graph neural network (GNN) is proposed for the first time. The temporal–spatial GNN takes gated tanh unit (GTU) as the temporal network to extract devices’ degradation from dynamic biases, and takes inductive GraphSAGE as the spatial network to obtain whole graph information from circuit topology and output circuit aging delay. With comprehensive comparison among the network candidates, the combination of GTU and GraphSAGE presents the highest accuracy in predicting the standard cell aging delay. Owing to the superior features capture capability, this framework significantly improves the aging prediction efficiency under various operation conditions, especially facing the iterations of usage scenario, design version and process design kit. Compared with the conventional flow, the average acceleration ratio of our temporal–spatial network in predicting aging delay is more than 200 times. Furthermore, this framework is demonstrated with ADDER and FIFO circuits in timing analysis at the end of life. Thus, this work is helpful to the aging-aware circuit design in nano-scale technology. Jinfeng Ye, Pengpeng Ren, Yongkang Xue, Hui Fang 0003, Zhigang Ji |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | DRGA-Based Second-Order Block Arnoldi Method for Model Order Reduction of MIMO RCS CircuitsabstractWith the escalating demand for fast simulation of large-scale multi-input multi-output (MIMO) RCS circuits formulated as second-order differential systems, the need arises for more effective decentralized second-order model order reduction (MOR) methods, while providing a desired approximation of the original system. Dynamic relative gain array (DRGA) that takes into account both the steady-state and dynamic system information has shown promising efficacy in measuring the degree of each loop interaction, which is crucial for decoupling a MIMO system into several multi-input single-output (MISO) subsystems. Although several decentralized MOR methods have been introduced for dimension reduction to linear MIMO networks, hardly has any research explored second-order decentralized MOR methods with regard to MIMO RCS circuits. Besides, the existing DRGA method based on first-order state feedback predictive control greatly increases the computational complexity when directly applying to second-order RCS systems. Hence, we develop a second-order block Arnoldi method based on DRGA, termed DRGA-SOBAR, which enables the extension of the SOAR method and the second-order DRGA method to MIMO scenarios. Experimental results on RCS networks show that most input-output interactions are negligible in terms of the magnitude-wise insignificance, and our proposed DRGA-SOBAR based reduced systems perform with higher accuracy compared to the PRIMA and the generalized block SOAR (SOBAR) methods, and higher efficiency compared to the decentralized SOBAR algorithm based on RGA method as well. Haibao Chen, Jie Chen 0005, Pengpeng Ren, Zhigang Ji, Junhua Liu 0001, Runsheng Wang, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | ChameleMon: Shifting Measurement Attention as Network State ChangesabstractNetwork measurement is critical to many network applications. There are mainly two kinds of flow-level measurement tasks: 1) packet accumulation tasks and 2) packet loss tasks. In practice, the two kinds of tasks are often required at the same time, but existing works seldom handle both. In this paper, we design ChameleMon to support the two kinds of tasks simultaneously. The key design of ChameleMon is to shift measurement attention as network state changes, through two dimensions of dynamics: 1) dynamically allocating memory between the two kinds of tasks; 2) dynamically monitoring the flows of importance. To realize the key design, we propose a key technique, leveraging Fermat's little theorem to devise a flexible data structure, namely FermatSketch. FermatSketch is dividable, additive, and subtractive, supporting the two kinds of tasks. We have implemented a ChameleMon prototype on a testbed with a Fat-tree topology. We conduct extensive experiments and the results show ChameleMon supports the two kinds of tasks with low memory/bandwidth overhead, and more importantly, it can automatically shift measurement attention as network state changes. Kaicheng Yang 0001, Yuhan Wu 0001, Ruijie Miao, Tong Yang 0003, Zirui Liu 0002, Zicang Xu, Yikai Zhao 0001, Hanglong Lv, Zhigang Ji, Gaogang Xie |
SIGCOMM | 10 |
| 2023 | Analytical Post-Voiding Modeling and Efficient Characterization of EM Failure Effects Under Time-Dependent Current StressingabstractElectromigration (EM) has become the major concern for integrated circuits (ICs) in advanced technology nodes. Traditional empirical EM models, such as Black’s equation, show inaccurate estimation for the time-to-failure of ICs, thus resulting in unnecessary over-design. To address this drawback, we propose a few analytical solutions for calculating the transient stress evolution and void volume in straight multisegment interconnect trees during the post-voiding phase. By employing the Laplace transform, the proposed method aims at solving coupled partial differential equations (PDEs) governed by physics-based EM modeling. The analytical solutions can be tailored to expressions with required accuracy and computational savings, leading to a compact end-to-end system providing results of EM failure effects at arbitrary time instances and locations of interconnect trees with varying geometry under time-dependent current and temperature stressing. The EM lifetime such as the incubation time, related to the void volume evolution, at any desired precision, can be calculated by the analytical solutions. The proposed method shows its accuracy, scalability, and computational savings through results compared with the finite element method (FEM) tool COMSOL and the competing methods and can achieve up to$593\times $speedup with < 10% error in EM failure time estimation. Tianshu Hou, Ngai Wong 0001, Quan Chen 0007, Zhigang Ji, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Equiprobability-Based Local Response Surface Method for High-Sigma Yield Estimation With Both High Accuracy and EfficiencyabstractWith the ever-increasing transistor density and memory capability in integrated circuits, the high-sigma yield estimation has become a growing concern. This work presents an equiprobability-based local response surface (ELRS) method that can perform a high-sigma yield estimation with both high accuracy and efficiency. Demonstrating with 6T-SRAM, the proposed method exhibits more than ten times improvement in accuracy when compared with the state-of-the-art while maintaining the efficiency to the best record in the literature. Pengpeng Ren, Haibao Chen, Zhigang Ji, Junhua Liu 0001, Runsheng Wang, Jianfu Zhang 0001, Ru Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | A Deep Learning Framework for Solving Stress-based Partial Differential Equations in Electromigration AnalysisabstractThe electromigration-induced reliability issues (EM) in very large scale integration (VLSI) circuits have attracted continuous attention due to technology scaling. Traditional EM methods lead to inaccurate results incompatible with the advanced technology nodes. In this article, we propose a learning-based model by enforcing physical constraints of EM kinetics to solve the EM reliability problem. The method aims at solving stress-based partial differential equations (PDEs) to obtain the hydrostatic stress evolution on interconnect trees during the void nucleation phase, considering varying atom diffusivity on each segment, which is one of the EM random characteristics. The approach proposes a crafted neural network-based framework customized for the EM phenomenon and provides mesh-free solutions benefiting from the employment of automatic differentiation (AD). Experimental results obtained by the proposed model are compared with solutions obtained by competing methods, showing satisfactory accuracy and computational savings. Tianshu Hou, Peining Zhen, Zhigang Ji, Haibao Chen |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Towards Accurate Oriented Object Detection in Aerial Images with Adaptive Multi-level Feature FusionabstractDetecting objects in aerial images is a long-standing and challenging problem since the objects in aerial images vary dramatically in size and orientation. Most existing neural network based methods are not robust enough to provide accurate oriented object detection results in aerial images since they do not consider the correlations between different levels and scales of features. In this paper, we propose a novel two-stage network-based detector with a daptive f eature f usion towards highly accurate oriented object det ection in aerial images, named AFF-Det . First, a multi-scale feature fusion module (MSFF) is built on the top layer of the extracted feature pyramids to mitigate the semantic information loss in the small-scale features. We also propose a cascaded oriented bounding box regression method to transform the horizontal proposals into oriented ones. Then the transformed proposals are assigned to all feature pyramid network (FPN) levels and aggregated by the weighted RoI feature aggregation (WRFA) module. The above modules can adaptively enhance the feature representations in different stages of the network based on the attention mechanism. Finally, a rotated decoupled-RCNN head is introduced to obtain the classification and localization results. Extensive experiments are conducted on the DOTA and HRSC2016 datasets to demonstrate the advantages of our proposed AFF-Det. The best detection results can achieve 80.73% mAP and 90.48% mAP, respectively, on these two datasets, outperforming recent state-of-the-art methods. Peining Zhen, Suming Zhang, Xiaotao Yan, Zhigang Ji, Haibao Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Towards a more efficient few-shot learning-based human gesture recognition via dynamic vision sensors
Linglin Jing, Yifan Wang 0008, Tailin Chen, Shirin Dora, Zhigang Ji, Hui Fang 0003 |
BMVC | 5 |
| 2022 | Insights of VG-dependent threshold voltage fluctuations from dual-point random telegraph noise characterization in nanoscale transistors
Xuepeng Zhan, Jiezhi Chen, Zhigang Ji |
Sci. China Inf. Sci. | 3 |
| 2022 | A Space-Time Neural Network for Analysis of Stress Evolution Under DC Current StressingabstractThe electromigration (EM)-induced reliability issues in very large-scale integration (VLSI) circuits have attracted increased attention due to the continuous technology scaling. Traditional EM models often lead to overly pessimistic predictions incompatible with the shrinking design margin in future technology nodes. Motivated by the latest success of neural networks in solving differential equations in physical problems, we propose a novel mesh-free model to compute EM-induced stress evolution in VLSI circuits. The model utilizes a specifically crafted space–time physics-informed neural network (STPINN) as the solver for EM analysis. By coupling the physics-based EM analysis with dynamic temperature incorporating Joule heating and via effect, we can observe stress evolution along multisegment interconnect trees under constant, time-dependent, and space–time-dependent temperature during the void nucleation phase. The proposed STPINN method obviates the time discretization and meshing required in conventional numerical stress evolution analysis and offers significant computational savings. Numerical comparison with competing schemes demonstrates a$2\times $–$52\times $speedup with a satisfactory accuracy. Tianshu Hou, Ngai Wong 0001, Quan Chen 0007, Zhigang Ji, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Fast Video Facial Expression Recognition by a Deeply Tensor-Compressed LSTM Neural Network for Mobile DevicesabstractMobile devices usually suffer from limited computation and storage resources, which seriously hinders them from deep neural network applications. In this article, we introduce a deeply tensor-compressed long short-term memory (LSTM) neural network for fast video-based facial expression recognition on mobile devices. First, a spatio-temporal facial expression recognition LSTM model is built by extracting time-series feature maps from facial clips. The LSTM-based spatio-temporal model is further deeply compressed by means of quantization and tensorization for mobile device implementation. Based on datasets of Extended Cohn-Kanade (CK+), MMI, and Acted Facial Expression in Wild 7.0, experimental results show that the proposed method achieves 97.96%, 97.33%, and 55.60% classification accuracy and significantly compresses the size of network model up to 221× with reduced training time per epoch by 60%. Our work is further implemented on the RK3399Pro mobile device with a Neural Process Engine. The latency of the feature extractor and LSTM predictor can be reduced 30.20× and 6.62× , respectively, on board with the leveraged compression methods. Furthermore, the spatio-temporal model costs only 57.19 MB of DRAM and 5.67W of power when running on the board. Peining Zhen, Haibao Chen, Zhigang Ji, Hao Yu 0001 |
ACM Trans. Internet Things | 4 |
| 2019 | Design for reliability with the advanced integrated circuit (IC) technology: challenges and opportunities
Zhigang Ji, Haibao Chen, Xiuyan Li |
Sci. China Inf. Sci. | 1 |
| 2014 | EMD-Based Multi-Model Prediction for Network Traffic in Software-Defined NetworksabstractAccurately predicting for network traffic is significant for network operation and maintenance in software-defined networks (SDN). In this paper, Multi-frequency characteristic of complex network traffic is considered, and a new algorithm named EMD-based multi-model Prediction (EMD-MMP) for network prediction is proposed. The main idea in this algorithm is to decompose the network traffic series into different modes with different frequency by Empirical Mode Decomposition (EMD). According to the characteristics and the cross correlation coefficient of the modes, we reconstruct new components for de-noising by summing up parts of the high frequency modes. Then the new components and the remaining old modes are predicted by ARMA and SVR methods. Finally, the historical traffic data of Internet2 is employed for our experiments to demonstrate the precision of our new prediction algorithm compared with the Auto-Regressive and Moving Average (ARMA) and Support Vector Regression (SVR) models. On average, the EMD-MMP method improves ARMA and SVR by 0.62% and 10.6% at the Mean Absolute Percentage Error (MAPE) statistic indicator, and the Mean Square Error (MSE) of EMD-MMP is 12060.92 while the ARMA and SVR are 13968.8 and 47588.3. Besides, the EMD-MMP algorithm gives a better understanding of the nature of the network traffic. Longfei Dai, Wenguo Yang, Suixiang Gao, Yinben Xia, Mingming Zhu, Zhigang Ji |
MASS | 6 |