EDBT 2026 Demo / reviewers in the wild / expert
Hao Yan 0002
dblp:77/6310-2
· DBLP profile ↗
29ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0002-5312-4483ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 2 first-author · 24 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GoG-Predict: IR-aware Path Waveform Prediction with Structural Entity Interaction LearningabstractIn advanced technology nodes, Current Source Models (CSMs) are widely adopted due to their high accuracy. However, IR-drop–induced supply voltage fluctuations increase the computational burden of CSM-based timing analysis. Existing machine-learning (ML) approaches accelerate CSM evaluation by assuming fixed receiver capacitance along the path, which contradicts the variable-load dependency inherent to CSM formulation. In this work, we propose GoG-Predict, a fast and accurate IR-aware path waveform prediction framework. GoG-Predict employs a Graph-of-Graphs (GoG) Neural Network to model the structured interactions among driver cells, load nets, and receiver cells. In addition, a Feature-wise Linear Modulation (FiLM) mechanism is incorporated to account for the impact of IR drop on output waveform. We evaluate GoG-Predict on open-source designs using a commercial 12nm technology. Experimental results show an average waveform error of 2.03% and a 1426 × runtime speedup over HSPICE, while achieving 1.63 × higher accuracy compared with state-of-the-art ML-based prediction methods. Yanglong Mao, Ziyue Han, Yunfan Zuo, Chenpu Shi, Yaning Jia, Hao Yan 0002, Longxing Shi |
ACM Great Lakes Symposium on VLSI | 7 |
| 2026 | DCTDSE: A Bimodal Design Space Exploration Flow via Discrete-Continuous TransformationabstractThe conservation core accelerator presents a promising avenue for improving the computational efficiency of specific applications. However, its design space is ultra-high-dimensional, significantly increasing the exploratory effort required to identify the optimal design across performance, power, and area metrics. Furthermore, the discrete nature of the microarchitecture space renders conventional search methods ineffective. To tackle these challenges, we propose the Discrete Continuous Transformation to speedup Design Space Exploration, namely DCTDSE. It can operate in either offline or online mode. In offline mode, it transforms the original discrete design space into a continuous space, builds predictive models, performs parallel gradient-based optimization, and maps the results back to the discrete domain. In the online mode, DCTDSE refines the models by iteratively resampling previously found solutions, thereby enhancing exploration quality while maintaining moderate runtime overhead. Experimental results indicate that DCTDSE achieves a 3.9× to 40× speedup over benchmark methods in offline mode. In online mode, it provides a 2.5× speedup, with a 21% reduction in exploration quality relative to the most accurate comparison method. Shuaibo Huang, Liangji Wu, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Learning-Driven Hierarchical Particle Swarm Optimization Framework for Power Delivery Network SynthesisabstractAs process nodes shrink, intensified IR drop and electromigration (EM) issues, driven by the power delivery network (PDN), affect reliability and further compress the routability of the entire design. Since physical design is time-consuming, effectively predicting and optimizing PDN during the power planning stage is crucial. This effort is complicated by complex interactions between multiple parameters and the challenges associated with nonconvex optimization. To overcome these limitations, we propose a learning-driven hierarchical particle swarm optimization framework for PDN, integrated with a residual TransUNet-based prediction model for post-placement congestion, EM, and IR drop prediction. In addition, we introduce the gate network for loss fusion, which excels in balancing output predictions and enhancing the accuracy of critical tasks. Analogous to hyperparameter optimization in deep learning, we can utilize GPU parallelization to search for the global optimum of the PDN structure efficiently. In circuits at the 22nmand 16nmtechnology nodes, with cell counts ranging from 560 to nearly 200, 000, our method achieves reductions in mean and maximum congestion by 15% and 49%, respectively, compared to state-of-the-art methods. In addition, it achieves an overall optimization time of around 22 seconds. Yunfan Zuo, Yuwei Sun, Pinquan Li, Yan Li 0056, Shuo Cui, Hao Yan 0002, Longxing Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | FACT: Fast and Accurate Multi-Corner Predictor for Timing Closure in Commercial EDA FlowsabstractWith technology scaling progressing well into deep nanometer region, the number of technology corners surges from dozens to hundreds. This dramatic increase in corners significantly complicates timing closure during Engineering Change Orders (ECO) stages, as performing full-corner static timing analysis (STA) becomes increasingly time-consuming and challenging. Existing methodologies often leverage machine learning (ML) techniques to predict unknown corners based on a subset of known corners. However, as the total number of corners expands, these methods not only require a larger set of known corners but also become highly sensitive to known corner selection. Additionally, as designers push designs into near-threshold voltage regions to enhance energy efficiency, the number of corners increases and the nonlinearity between corners becomes more pronounced. This intensifies the difficulty for current ML-based methods to accurately predict full-corner timing metrics. In this work, we propose FACT, a fast and accurate multi-corner predictor for timing closure optimization. Our approach simplifies the process by necessitating timing analysis under only one known corner. By effectively capturing correlations across diverse technology files, FACT robustly infers full-corner timing metrics, even under challenging near-threshold conditions. Moreover, our framework seamlessly integrates with commercial EDA design flows, making it practical in industrial environments. Experimental results on open-source designs indicate the superior stability of our method, coupled with a significant runtime speed-up over both traditional and prior ML-based timing ECO flows. Ziyue Han, Shiyang Wu, Hao Yan 0002, Longxing Shi |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2026 | Enhanced TransUNet Framework for Predicting Static IR Drop and Chip RoutabilityabstractAs semiconductor processes advance, the power delivery network (PDN) increasingly affects the power supply from the pads to the cells. Significant IR drops in standard cells can lead to timing violations, while suboptimal PDN topologies can lead to increased congestion. Together, these factors degrade overall chip performance and reliability. To accelerate design iteration, accurately and efficiently predicting unevenly distributed IR drop and congestion, especially in hotspot areas, has become a critical challenge. This article introduces an enhanced TransUNet-based framework for distribution-aware static IR drop and congestion prediction, treating both problems as separate but related image prediction tasks. Such an abstraction preserves the IR drop and congestion distribution patterns on the original real physical layout and retains local hot spots. Our proposed framework leverages image classification techniques to model IR drop prediction as a spatial pattern recognition task, effectively addressing the long-tail distribution in different regions. To enhance hotspot prediction, we incorporate wavelet transform and transformer-based analysis to enable multiscale feature fusion. In the open-source CircuitNet dataset, our method predicts static IR drop with a mean absolute error( MAE ) of 0.374 mV and a maximum error rate( Err m ) of 18.7%, reducing MAE and Err m by 77.4% and 79.2%, respectively, compared to the state-of-the-art method, all within 100 ms. The congestion prediction evaluations show 65.3% lower NRMSE scores and 18.4% higher SSIM scores relative to the existing SOTA approach. Our approach accurately and reliably predicts long-tail distributions and localized hotspots in both IR drop and congestion tasks. Yunfan Zuo, Pinquan Li, Yuwei Sun, Hao Yan 0002, Longxing Shi |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | Rank-based Multi-objective Approximate Logic Synthesis via Monte Carlo Tree SearchabstractApproximate Logic Synthesis (ALS) is an automated technique designed for error-tolerant applications, optimizing delay, area, and power under specified error constraints. However, existing methods typically focus on either delay reduction or area minimization, often leading to local optima in multi-objective optimization. This paper proposes a rankbased multi-objective ALS framework using Monte Carlo Tree Search (MCTS). It develops non-dominated circuit ranking, to guide MCTS in exploring local approximate changes (LACs) across the entire circuit and generate approximate circuit sets with great optimization potential. Additionally, a Rank-Transformer model is introduced to predict pathdomain ranks, enhancing the application of high-quality LACs within circuit paths. Experimental results show that our framework achieves faster and more efficient optimization in delay and area simultaneously compared to state-of-the-art methods. Yuyang Ye 0001, Xiangfei Hu, Peng Xu 0052, Yu Gong 0002, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
DAC | 7 |
| 2025 | Timing-Driven Approximate Logic Synthesis Based on Double-Chase Grey Wolf OptimizerabstractWith the shrinking technology nodes, timing optimization becomes increasingly challenging. Approximate logic synthesis (ALS) can perform local approximate changes (LACs) on circuits to optimize timing with the cost of slight inaccuracy. However, existing ALS methods that focus solely on critical path depth reduction or area minimization are not optimal in timing optimization. This paper proposes an effective timing-driven ALS framework, where we employ a double-chase grey wolf optimizer to explore and apply LACs, simultaneously bringing excellent critical path shortening and area reduction under error constraints. Subsequently, it utilizes post-optimization under area constraints to convert area reduction into further timing improvement, thus achieving maximum critical path delay reduction. According to experiments on open-source circuits with 28nm technology, compared to the SOTA method, our framework can generate approximate circuits with greater critical path delay reduction under different error and area constraints. Xiangfei Hu, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001 |
DATE | 4 |
| 2025 | An Imitation Augmented Reinforcement Learning Framework for CGRA Design Space ExplorationabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a promising architecture that warrants thorough design space exploration (DSE). However, traditional DSE methods for CGRAs often get trapped in local optima due to singularities, i.e., invalid design points caused by CGRA mapping failures. In this paper, we propose a singularity-aware framework based on the integration of reinforcement learning (RL) and imitation learning (IL) for DSE of CGRAs. Our approach learns from both valid and invalid points, substantially reducing the probability of sampling singularities and accelerating the escape from inefficient regions, ultimately achieving high-quality Pareto points. Experimental results demonstrate that our framework improves the hypervolume (HV) of the Pareto front by 23.56% compared to state-of-the-art methods, with a comparable time overhead. Liangji Wu, Shuaibo Huang, Shiyang Wu, Hao Yan 0002, Longxing Shi |
DATE | 6 |
| 2025 | ARS-Flow 2.0: An enhanced design space exploration flow for accelerator-rich system based on active learning
Shuaibo Huang, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi |
Integr. | 3 |
| 2025 | Robust optimization algorithm of RF MEMS switches considering uncertainties
Hao Yan 0002, Yaning Jia, Chuangyuan Zeng, Xiaoping Liao, Xiao Shi 0001 |
Integr. | 1 |
| 2025 | Learning-Driven Physically Aware Large-Scale Circuit Gate SizingabstractGate sizing plays an important role in timing optimization after physical design. Existing machine learning-based gate sizing works cannot optimize timing on multiple timing paths simultaneously and neglect the physical constraint on layouts. They cause suboptimal sizing solutions and low-efficiency issues when compared with commercial gate sizing tools. In this work, we propose a learning-driven physically aware gate sizing framework to optimize timing performance on large-scale circuits efficiently. In our gradient descent optimization-based work, for obtaining accurate gradients, a multimodal gate sizing-aware timing model is achieved via learning timing information on multiple timing paths and physical information on multiple-scaled layouts jointly. Then, gradient generation based on the sizing-oriented estimator and adaptive back-propagation are developed to update gate sizes. Our results demonstrate that our work achieves higher-timing performance improvements in a faster way compared with the commercial gate sizing tool. Yuyang Ye 0001, Peng Xu 0052, Lizheng Ren, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | ARS-Flow: A Design Space Exploration Flow for Accelerator-rich System based on Active LearningabstractSurrogate model-based design space exploration (DSE) is the mainstream method to search for optimal microarchitecture designs. However, it is hard to build accurate models for accelerator-rich systems within limited samples due to its high dimensional characteristic. Moreover, it is easy to fall into local optimal or difficult to converge. To solve these two problems, we propose a DSE flow based on active learning, namely ARS-Flow. It is featured with Pareto-region-oriented stochastic resampling method (PRSRS) and multiobjective genetic algorithm with self-adaptive hyperparameter control (SAMOGA). Taking the gem5-SALAM system for illustration, the proposed method can build more accurate models and find better microarchitecture designs with acceptable runtime costs. Shuaibo Huang, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi |
ASPDAC | 3 |
| 2024 | A Graph-Learning-Driven Prediction Method for Combined Electromigration and Thermomigration Stress on Multi-Segment InterconnectsabstractAs technology advances, the temperature gradient in the interconnects becomes more significant, which causes serious thermomigration. Simulating the coupling effects of thermomigration (TM) and electromigration (EM) on large-scale circuits is very time-consuming caused by a substantial increase in computational complexity. Recently, some researchers utilized graph learning-based methods to predict EM stress in medium-scale cases. Unfortunately, these works overlooked the effects of TM. To predict the EM - TM stress of large-scale interconnects accurately and efficiently, we propose a framework based on Graph Attention Networks (GATs) with a customized alternating aggregation method for collecting information in junctions and branches of interconnects jointly. The experimental results show that our work achieves an average relative error of less than 1 % compared to the commercial software COMSOL for inter-connects consisting of fewer than 200 segments. Furthermore, our method also achieves 9037 x speedup in predicting the OpenROAD test circuit with a maximum segment number reaching 10807. Yunfan Zuo, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Longxing Shi |
DATE | 5 |
| 2024 | Timing-Driven Technology Mapping Approximation Based on Reinforcement LearningabstractAs the transistor technology nodes shrink into the nano-scales, timing guardbands caused by aging effects and process variations continue to increase. Approximate computing can eliminate aging-and-variation-induced timing guardbands without sacrificing the design performance. It can apply local approximate changes (LACs) automatically in circuits to reduce critical path delay. However, efficiently achieving timing optimization under error distance constraints is still tricky. This work proposes an automated timing-driven technology mapping approximation framework based on reinforcement learning (RL). The framework uses path-weighted graph neural networks (PGNNs) to embed RL states and timing path-aware LAC candidates to construct RL action spaces. It can efficiently eliminate timing guardbands induced by aging and variation. Our proposed circuit-agnostic framework operates on the gate-level netlists. According to experiments on the open-source circuits using TSMC 28nm and 16nm technology under aging and variation conditions, our framework can achieve an average 24.78% critical path delay reduction under 5 different error distance constraints and 4.83× speedup, compared with a state-of-the-art method. Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Graph-Learning-Driven Path-Based Timing Analysis Results Predictor from Graph-Based Timing AnalysisabstractWith diminishing margins in advanced technology nodes, the performance of static timing analysis (STA) is a serious concern, including accuracy and runtime. The STA can generally be divided into graph-based analysis (GBA) and path-based analysis (PBA). For GBA, the timing results are always pessimistic, leading to overdesign during design optimization. For PBA, the timing pessimism is reduced via propagating real path-specific slews with the cost of severe runtime overheads relative to GBA. In this work, we present a fast and accurate predictor of post-layout PBA timing results from inexpensive GBA based on deep edge-featured graph attention network, namely deep EdgeGAT. Compared with the conventional machine and graph learning methods, deep EdgeGAT can learn global timing path information. Experimental results demonstrate that our predictor has the potential to substantially predict PBA timing results accurately and reduce timing pessimism of GBA with maximum error reaching 6.81 ps, and our work achieves an average 24.80× speedup faster than PBA using the commercial STA tool. Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
ASP-DAC | 4 |
| 2023 | A Novel Delay Calibration Method Considering Interaction between Cells and WiresabstractIn the advanced technology, the accuracy of cell and wire delay modeling are the key metrics for timing analysis. However, when the supply voltage decreases to the near-threshold regime, the complicated process variation effect causes the cell delay and the wire delay hard to model. Most researchers study cell or wire delay separately, ignoring the coefficients between them. In this paper, we propose an N-sigma delay model by characterizing different sigma levels$\mathbf{(-3\sigma {to}+3\sigma)}$of the cell and wire delay distribution. The N-sigma cell delay model is represented by the first four moments and calibrated by the operating conditions (input slew, output load). Meanwhile, based on the Elmore model, the wire delay variability is calculated by considering the effect of drive and load cells. The delay models are verified through the ISCAS85 benchmarks and the functional units of PULPino processor with TSMC 28 nm technology. Compared to the SPICE results, the average errors for estimating the$+/-\mathbf{3\sigma}$cell delay are 2.1 % and 2.7% and those of the wire delay are 2.4% and 1.6%, respectively. The errors of path delay analysis keep below 6.6% and the speed is 103X over SPICE MC simulations. Leilei Jin, Wenjie Fu 0003, Hao Yan 0002, Xiao Shi 0001, Longxing Shi |
DATE | 4 |
| 2023 | Fast and Accurate Wire Timing Estimation Based on Graph LearningabstractAccurate wire timing estimation has become a bottleneck in timing optimization since it needs a long turn-around time using a sign-off timer. The gate timing can be calculated accurately using lookup tables in cell libraries. In comparison, the accuracy and efficiency of wire timing calculation for complex RC nets are extremely hard to trade-off. The limited number of wire paths opens a door for the graph learning method in wire timing estimation. In this work, we present a fast and accurate wire timing estimator based on a novel graph learning architecture, namely GNNTrans. It can generate wire path representations by aggregating local structure information and global relationships of whole RC nets, which cannot be collected with traditional graph learning work efficiently. Experimental results on both tree-like and non-tree nets demonstrate improved accuracy, with the max error of wire delay being lower than 5 ps. In addition, our estimator can predict the timing of over 200K nets in less than 100 secs. The fast and accurate work can be integrated into incremental timing optimization for routed designs. Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
DATE | 4 |
| 2023 | FPGNN-ATPG: An Efficient Fault Parallel Automatic Test Pattern GeneratorabstractThe advanced multi-core technology enables parallel computing to speed up Automatic Test Pattern Generation (ATPG). The main challenge is to solve an increasing number of hard-to-solve faults effectively. In this paper, we develop an efficient parallel computing system for the ATPG program, i.e., FPGNN-ATPG, which is consisted of two parts: graph-neural-networks-based (GNN-based) fault classification and fault-driven deterministic test pattern generator (DTPG). The end-to-end GNN-based classifier can predict fault types with superior accuracy compared with classical machine learning methods. And the fault-driven DTPG can solve different types of faults in parallel without runtime overhead. According to the experimental results on an 8-core machine, our FPGNN-ATPG framework obtains an average of 7.56X speedup while reducing 14.13% pattern count ratio with full 100% fault coverage for 10 industrial instances. Yuyang Ye 0001, Zonghui Wang, Zun Xue, Hao Yan 0002 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2023 | Optimized matrix ordering of sparse linear solver using a few-shot model for circuit simulation
Yuyang Ye 0001, Hao Yan 0002, Longxing Shi |
Integr. | 4 |
| 2023 | An efficient SRAM yield analysis method based on scaled-sigma adaptive importance sampling with meta-model accelerated
Liang Pang 0002, Mengyun Yao, Xiao Shi 0001, Hao Yan 0002, Longxing Shi |
Integr. | 6 |
| 2023 | Aging-Aware Critical Path Selection via Graph Attention NetworksabstractIn advanced technology nodes, aging effects like negative and positive bias temperature instability (NBTI and PBTI) become increasingly significant, making timing closure and optimization more challenging. Unfortunately, conventional critical path (CP) selection tools used in reliability-aware design flow cannot accurately identify CPs under different aging conditions. To address this issue, we propose an aging-aware CP selection flow comprising two parts: 1) critical cell detection and 2) path criticality (PC) computation. We employ graph-attention (GAT) networks to predict the critical cells in the aged circuits, and a PC computation algorithm that takes into account circuit-level and transistor-level parameters to generate PC rank lists. Our experimental results demonstrate that our GAT model outperforms classical machine learning models in detecting critical cells. Additionally, compared with the commercial tool, our aging-aware flow achieves an average accuracy of 99.52%, 98.69%, and 97.20% for top-10%, top-5%, and top-1% path sets, respectively, in five industrial designs subjected to different aging conditions and workloads. Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Quality Driven Systematic Approximation for Binary-Weight Neural Network DeploymentabstractNeural networks (NNs) with large scales of artificial neurons are increasingly used in recognition and classification tasks. In power-constrained scenarios, the tradeoff between performance and hardware consumptions must be carefully evaluated before silicon tape-out. In this paper, we proposed a systematic approach to design ultra-low power NN system. This work is motivated by the facts that NNs are resilient to approximation in many of the computations and NNs are outputting statistical tensors which are acceptable to less-than-perfect results. We resort to the front-back end approach with a twofold aim: (1) a fast and accurate design approach is proposed by estimating the computing quality of low-power approximate adder arrays, and it is adopted to evaluate the neural network system; (2) a quality configurable engine with different approximation degrees while processing NNs is implemented. The proposed work is demonstrated with a comprehensive keyword spotting (KWS) system as an ultra-low power NN engine. The experimental environment is setup with ten keywords from the google speech command dataset (GSCD) using an industrial 22-nm ultra-low-leakage (ULL) process. Comparing to the state-of-the-art KWS processors, the proposed approximate NN engine can demonstrate over 60% improvement in power efficiency and$1.1\times $area efficiency while achieving similar recognition accuracy. Yu Gong 0002, Hao Cai 0001, Haige Wu, Hao Yan 0002, Zhen Wang 0019, Longxing Shi, Bo Liu 0019 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | A Compact High-Dimensional Yield Analysis Method using Low-Rank Tensor Approximationabstract“Curse of dimensionality” has become the major challenge for existing high-sigma yield analysis methods. In this article, we develop a meta-model using Low-Rank Tensor Approximation (LRTA) to substitute expensive SPICE simulation. The polynomial degree of our LRTA model grows linearly with the circuit dimension. This makes it especially promising for high-dimensional circuit problems. Our LRTA meta-model is solved efficiently with a robust greedy algorithm and calibrated iteratively with a bootstrap-assisted adaptive sampling method. We also develop a novel global sensitivity analysis approach to generate a reduced LRTA meta-model which is more compact. It further accelerates the procedure of model calibration and yield estimation. Experiments on memory and analog circuits validate that the proposed LRTA method outperforms other state-of-the-art approaches in terms of accuracy and efficiency. Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Chengzhen Xuan, Lei He 0001, Longxing Shi |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | An Adaptive Delay Model for Timing Yield Estimation under Wide-Voltage RangeabstractYield analysis for wide-voltage circuit design is a strong nonlinear integration problem. The most challenging task is how to accurately estimate the yield of long-tail distribution. This paper proposes an adaptive delay model to substitute expensive transistor-level simulation for timing yield estimation. We use the Low-Rank Tensor Approximation (LRTA) to model the delay variation from a large number of process parameters. Moreover, an adaptive nonlinear sampling algorithm is adopted to calibrate the model iteratively, which can capture the larger variability of delay distribution for different voltage regions. The proposed method is validated on benchmark circuits of TAU15 in 45nm free PDK. The experiment results show that our method achieves 20-100X speedup compared to Monte Carlo simulation at the same accuracy level. Hao Yan 0002, Xiao Shi 0001, Chengzhen Xuan, Peng Cao 0002, Longxing Shi |
ASP-DAC | 1 |
| 2020 | A Non-Gaussian Adaptive Importance Sampling Method for High-Dimensional and Multi-Failure-Region Yield AnalysisabstractRare-event yield analysis is challenging for high-dimensional circuit cases. In this paper, we propose a non-Gaussian adaptive importance sampling (NGAIS) method. In order to approximate the failure region in high-dimensional space, we model it as a mixture of von Mises-Fisher distributions. We formulate the parameter estimation problem as a maximum likelihood estimation problem, and then solve with expectation-maximization algorithm. Experiments on bit cell, amplifier and SRAM column circuit validate that the proposed NGAIS method outperforms other state-of-the-art approaches in terms of accuracy and efficiency. Xiao Shi 0001, Hao Yan 0002, Chuwen Li, Jianli Chen, Longxing Shi, Lei He 0001 |
ICCAD | 2 |
| 2020 | An Efficient Adaptive Importance Sampling Method for SRAM and Analog Yield AnalysisabstractPerformance failure has become a major threat for various memory and analog circuits. It is challenging to estimate the extremely small failure probability when failed samples are distributed in multiple disjoint regions. In this article, we propose an adaptive importance sampling (AIS) algorithm. AIS has several iterations of sampling region adjustments, while existing methods predecide a static sampling distribution. We design two adaptive frameworks based on resampling and population Metropolis-Hastings (MH) to iteratively search for failure regions. The experimental results of the AIS method exhibit better efficiency and higher accuracy. For SRAM bit cell with single failure region, the AIS method uses 2-$27{\times }$ fewer samples and reaches better accuracy when compared to several recent methods. For a two-stage amplifier circuit with multiple failure regions, the AIS method is $90{\times }$ faster than Monte Carlo and 7-23 ${\times }$ over other methods. For charge pump circuit and $C^{2}MOS$ master-slave latch circuit, the AIS method can reach 6-$18{\times }$ and 4-$6{\times }$ speedup over other methods, respectively. Xiao Shi 0001, Hao Yan 0002, Longxing Shi, Lei He 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Meta-Model based High-Dimensional Yield Analysis using Low-Rank Tensor Approximationabstract"Curse of dimensionality" has become the major challenge for existing high-sigma yield analysis methods. In this paper, we develop a meta-model using Low-Rank Tensor Approximation (LRTA) to substitute expensive SPICE simulation. The polynomial degree of our LRTA model grows linearly with circuit dimension. This makes it especially promising for high-dimensional circuit problems. Our LRTA meta-model is solved efficiently with a robust greedy algorithm, and calibrated iteratively with an adaptive sampling method. Experiments on bit cell and SRAM column validate that proposed LRTA method outperforms other state-of-the-art approaches in terms of accuracy and efficiency. Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Longxing Shi, Lei He 0001 |
DAC | 2 |
| 2019 | Efficient Yield Analysis for SRAM and Analog Circuits using Meta-Model based Importance Sampling MethodabstractPerformance failure has become the major threat to the robustness and reliability of various memory and analog circuits. It is challenging to accurately estimate the extremely small failure probability when failed samples are distributed in multiple disjoint failure regions. In this paper, we develop a novel meta-model based importance sampling (MIS) method. MIS utilizes Gaussian Process meta-model to construct quasi-optimal importance sampling distribution, and performs Markov Chain Monte Carlo (MCMC) simulation to generate new samples from the proposed distribution. By updating our global Importance Sampling estimator in an iterated framework, MIS leads to better efficiency and higher accuracy. For SRAM bit cell with single failure region, MIS uses 4-6X fewer samples and reaches better accuracy when compared to several recent methods. For a two-stage amplifier circuit with multiple failure schemes, MIS is 213X faster than MC without compromising accuracy, while other methods fail to cover all failure regions in our experiment. Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Longxing Shi, Lei He 0001 |
ICCAD | 2 |
| 2019 | Adaptive Clustering and Sampling for High-Dimensional and Multi-Failure-Region SRAM Yield AnalysisabstractStatistical circuit simulation is exhibiting increasing importance for memory circuits under process variation. It is challenging to accurately estimate the extremely low failure probability as it becomes a high-dimensional and multi-failure-region problem. In this paper, we develop an Adaptive Clustering and Sampling (ACS) method. ACS proceeds iteratively to cluster samples and adjust sampling distribution, while most existing approaches pre-decide a static sampling distribution. By adaptively searching in multiple cone-shaped subspaces, ACS obtains better accuracy and efficiency. This result is validated by our experiments. For SRAM bit cell with single failure region, ACS requires 3-5X fewer samples and achieves better accuracy compared with existing approaches. For 576-dimensional SRAM column circuit with multiple failure regions, ACS is 2050X faster than MC without compromising accuracy, while other methods fail to converge to correct failure probability in our experiment. Xiao Shi 0001, Hao Yan 0002, Xiaofen Xu, Longxing Shi, Lei He 0001 |
ISPD | 2 |