Peng Cao 0002

dblp:06/5143-2 · DBLP profile ↗
← Back
33ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0003-2039-9031ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Physical-Aware Multicorner Timing Prediction for Pre-Placement Optimization Using BiLSTM and GNN Networks
abstract
The timing inpredictability at early design stage poses significant challenge to design convergence of integrated circuits, especially under multiple PVT corners. Prior works have concentrated exclusively on research of timing correlation and the related impact to circuit design across multiple design stages or corners, without addressing the influence of both on the final prediction accurately and efficiently. In this work, a physical aware timing prediction framework for multiple corners is proposed for pre-placement prediction and optimization by employing the bidirectional LSTM network and MLP network to learn the correlation from the sequence features and global features at logic synthesis stage, in which a GNN-based model is utilized to estimate the HPWL as physical feature, addressing the timing prediction challenge suffering from physical information absence. Besides, a corner selection mechanism is introduced to reduce the post-synthesis feature extraction overhead from multiple corners. The proposed pre-placement timing prediction and optimization framework was validated under TSMC 22nm process with the ISCAS’89, OpenCores, and Gaisler benchmark circuits. Significant prediction accuracy improvement is achieved to estimate the post-placement path delay for 12 PVT corners by post-synthesis timing analysis results under 5 corners with an average rRMSE of 8.75% for unseen designs, demonstrating a 19×∼53× runtime speedup. Significant placement quality improvement is achieved in terms of 25.96% ADP reduction and 20.52% PDP reduction in average compared to traditional design flow.
Peng Cao 0002, Pengcheng Fan, Jiaqi Lyu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 StatCHAR: Statistical Timing Characterization Framework via Heterogeneous Graph Attention Network and Active Learning With Parasitic RC Reduction
abstract
Statistical timing characterization for standard cell library poses significant challenges to accuracy and runtime cost. Prior analytical and learning-based methods neglect the profound influence induced by the layout-dependent parasitic resistor and capacitor (RC) network in cell netlist as well as the timing correlation between the topological structures of cells and process, voltage, and temperature (PVT) corners for model training, resulting in tremendous simulation effort and poor accuracy. In this work, a Statistical timing Characterization framework via Heterogeneous graph attention network and Active learning with parasitic RC Reduction (StatCHAR) is proposed, where the transistors and parasitic RC in cell are represented as heterogeneous nodes for graph learning and redundant RC nodes are removed to alleviate node imbalance issue and improve accuracy. The significant training data are selected from the full characterization set with active learning strategy to achieve the optimal balance between simulation overhead for the training set and prediction precision for the remaining test set. The proposed framework was validated with typical standard cells under multiple PVT corners with TSMC 22nm process, which achieves an excellent prediction with a relative Root Mean Square Error (rRMSE) of only 2.43% with only 11.9% of total characterization data for training, demonstrating an accuracy improvement of 2.7$\times \sim 12.1\times $for statical timing analysis on benchmark circuits compared to competitive learning based methods and a characterization runtime reduction by 7.5$\times $.
Peng Cao 0002, Zeyuan Deng, Yuhan Dong, Yuyang Ye 0001, Jun Yang 0006
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 Late Breaking Results: BLAST: Bisection-Free Learning Approach for Statistical Timing Characterization
abstract
Statistical timing characterization for standard cells faces significant computational challenges due to the laborious bisection analysis for setup/hold constraint of sequential cells. To address this issue, we propose a Bisection-Free Learning Approach for Statistical Timing Characterization (BLAST) by extracting inherent delay for data path and clock path in sequential cells as specific features. Multi-task learning is implemented with a multi-gate mixture-of-experts (MMoE) model to exploit the profound interdependency between setup and hold constraint for different timing arcs, where the active learning strategy is incorporated to improve learning efficiency. Experimental results under 135 PVT corners with TSMC 12nm process demonstrate that the proposed BLAST achieves considerable acceleration by avoiding the iterative bisection search for statistical constraint prediction with 76.9% runtime reduction compared to the commercial tool. Excellent prediction accuracy is achieved for various flip-flops by BLAST with the relative root mean square error (rRMSE) of $\mathbf{2. 2 1 \%}$ and worst-case absolute error (WCAE) of 0.82 ps.
Kai Jing, Zeyuan Deng, Junming Jiao, Peng Cao 0002
DAC5
2025 An Optimization-Aware Prerouting Timing Prediction Framework Based on Multimodal Learning
abstract
Accurate and efficient prerouting timing estimation is particularly crucial during placement to alleviate time-consuming design iterations. Machine-learning (ML)-based methods have been introduced recently to predict the post-routing timing results at placement stage, but most of them neglect the impact of timing optimization during physical design, suffering from accuracy loss due to inconsistent circuit netlist. In this work, an optimization-aware prerouting timing prediction framework based on multimodal learning is proposed to calibrate the timing changes between placement and routing stages, where the local netlist and layout information are extracted by graph neural network (GNN) and convolutional neural network (CNN), respectively, while the global information along the path is further extracted by Transformer network. Based on the predicted post-routing timing results by the proposed framework, timing optimization guidance is generated to enhance traditional design flow with better physical implementation quality. Experimental results demonstrate that for the OpenCores benchmark circuits under TSMC 22nm process, the proposed framework achieves significant correlation and accuracy improvement with an average of 0.9219 in terms of R2 score and 2.12% of mean absolute percentage error (MAPE) as well as an average runtime acceleration of$645\times $compared with traditional design flow on testing designs. With the timing optimization guidance, significant worst negative slack (WNS) and total negative slack (TNS) improvement are achieved compared with traditional flow after placement and routing, respectively, without noticeable area, power, wire length, and the number of design rule check (DRC) violations increase.
Peng Cao 0002, Yusen Qin, Guoqing He, Zhanhua Zhang, Yuyang Ye 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 An On-Chip-Training Keyword-Spotting Chip Using Interleaved Pipeline and Computation-in-Memory Cluster in 28-nm CMOS
abstract
To improve the precision of keyword spotting (KWS) for individual users on edge devices, we propose an on-chip-training KWS (OCT-KWS) chip for private data protection while also achieving ultralow -power inference. Our main contributions are: 1) identity interchange and interleaved pipeline methods during backpropagation (BP), enabling the pipelined execution of operations that traditionally had to be performed sequentially, reducing cache requirements for loss values by 95.8%; 2) all-digital isolated-bitline (BL)-based computation-in-memory (CIM) macro, eliminating ineffective computations caused by glitches, achieving 2.03$\times$higher energy efficiency; and 3) multisize CIM cluster-based BP data flow, designing each CIM macro collaboratively to achieve all-time full utilization, reducing 47.2% of output feature map (Ofmap) access. Fabricated in 28-nm CMOS and enhanced with a refined library characterization methodology, this chip achieves both the highest training energy efficiency of 101.5 TOPS/W and the lowest inference energy of 9.9nJ/decision among current KWS chips. By retraining a three-class depthwise-separable convolutional neural network (DSCNN), detection accuracy on the private dataset increases from 80.8% to 98.9%.
Junyi Qian, Peng Cao 0002, Xin Si, Weiwei Shan
IEEE Trans. Very Large Scale Integr. Syst.6
2024 Heterogeneous Graph Attention Network Based Statistical Timing Library Characterization with Parasitic RC Reduction
abstract
Statistical timing characterization for standard cell library poses significant challenge to accuracy and runtime cost. Prior analytical and machine learning-based methods neglect the profound influence induced by layout-dependent parasitic resistor and capacitor (RC) network in cell netlist as well as the timing correlation between topological structures of cells and process, voltage, and temperature (PVT) corners, resulting in tremendous simulation effort and/or poor accuracy. In this work, an accurate and efficient statistical cell timing library characterization framework is proposed based on heterogeneous graph attention network (HGAT) assisted with parasitic RC reduction approach, where the transistors and parasitic RC in cell are represented as heterogeneous nodes for graph learning and redundant RC nodes are removed to alleviate node imbalance issue and improve prediction accuracy. The proposed framework was validated with TSMC 22nm standard cells under multiple PVT corners to predict the standard deviation of cell delay with the error of 2.67% on average for all validated cells in terms of relative Root Mean Squared Error (rRMSE) with $3 \times $ characterization runtime speedup, achieving $2.7 \sim 6.9 \times $ accuracy improvement compared with prior works. The predicted statistical timing libraries were further validated with ISCAS’89 benchmark circuits for statistical static timing analysis (SSTA), where the critical path delay at $3 \sigma$ percentile point is reported with the average mismatch of $1.34 ps$ compared with foundry-provided library, showing $10.7 \sim 14.5 \times $ better accuracy than the competitive approaches.
Yuyang Ye 0001, Guoqing He, Peng Cao 0002
ASPDAC5
2024 An Optimization-aware Pre-Routing Timing Prediction Framework Based on Heterogeneous Graph Learning
abstract
Accurate and efficient pre-routing timing estimation is particularly crucial in timing-driven placement, as design iterations caused by timing divergence are time-consuming. However, existing machine learning prediction models overlook the impact of timing optimization techniques during routing stage, such as adjusting gate sizes or swapping threshold voltage types to fix routing-induced timing violations. In this work, an optimization-aware pre-routing timing prediction framework based on heterogeneous graph learning is proposed to calibrate the timing changes introduced by wire parasitic and optimization techniques. The path embedding generated by the proposed framework fuses learned local information from graph neural network and global information from transformer network to perform accurate endpoint arrival time prediction. Experimental results demonstrate that the proposed framework achieves an average accuracy improvement of 0.10 in terms of R2score on testing designs and brings average runtime acceleration of three orders of magnitude compared with the design flow.
Guoqing He, Yuyang Ye 0001, Peng Cao 0002
ASPDAC6
2024 A Physical and Timing Aware Placement Optimization Framework Based on Graph Neural Network
abstract
Timing-driven placement is crucial in physical design flow with significant impact on later routability and ultimate manufacturability, which may deviate from finding the optimal solution and/or lead to unnecessary iterations, suffering from interleaved optimization steps and the corresponding inaccurate timing estimation. To solve this issue, we propose a Physical and Timing Aware framework with Graph Neural Network, PTA-GNN, which provides the candidate gate sizing and buffer insertion solutions as well as the timing constraint for potential violated paths as guidance to improve placement quality significantly. Experimental results on the OpenCores benchmarks with 22nm technology demonstrate that the proposed placement optimization framework achieves up to 89.09% worst negative slack (WNS), 55.47% total negative slack (TNS) improvement and 25.36% reduction on the number of violating paths (#VP). Our framework benefits the later routing stage with 2.19% wire-length decrease and 22% runtime reduction compared to standard physical design flow.
Zhanhua Zhang, Guoqing He, Peng Cao 0002
ICCAD4
2024 LAG-Sizer: A Novel Gate Sizer Based on Leak Generative Adversarial Network with Feature Fusion
abstract
Gate sizing is an NP-hard problem to achieve Performance, Power and Area (PPA) optimization. Recently proposed learning-based approaches struggle to overcome the runtime issue of traditional heuristics, but lack the consideration of the intrinsic features for candidate gates in library and could not address the inequality issue of candidate sizes for different gates properly, suffering from insufficient design space exploration and inaccurate sizing assignment. In this work, based on a variant of generative adversarial network, Leak Adversarial Generation (LAG), a novel LAG-Sizer is proposed to model gate sizing as sequence generation problem, which breaks the traditional adversarial network by leaking the discriminator feature information into the generator to guide sizing generation. Feature fusion technique is introduced to comprehensively consider circuit feature and cell library feature while a unified classification is proposed to perfectly solve the inequality issue for sizing. The proposed sizer was validated with IWLS2005 and Opencores benchmark circuits under 22nm process. Experimental results demonstrate that an average of 4.6% Total Negative Slack (TNS) improvement and 15.6% number of violating endpoints (NVE) reduction are achieved by this work with similar area and power consumption compared to commercial tools as well as significant runtime speedup of 47.8×.
Zhanhua Zhang, Guoqing He, Peng Cao 0002
ICCAD4
2024 Ultra-low-power one-hot transmission-gate multiplexer (OTG-MUX) scalable into large fan-in circuits in 28 nm CMOS
Yuqiang Cui, Weiwei Shan, Peng Cao 0002
Integr.3
2023 An efficient path delay variability model for wide-voltage-range digital circuits
Weiwei Shan, Yuqiang Cui, Wentao Dai, Xinning Liu, Peng Cao 0002, Jun Yang 0006
Sci. China Inf. Sci.6
2023 TF-Predictor: Transformer-Based Prerouting Path Delay Prediction Framework
abstract
Timing mismatch between different stages of physical design poses great challenges for circuit optimization to achieve the desired performance, power, and area (PPA) tradeoff. The inaccurate timing estimation prior to routing may lead to over-design with unwanted power and area consumption or iterating back to cell placement at the cost of design turn-around time. Existing learning models could not predict post-routing circuit timing with satisfying accuracy and efficiency due to the limitations of the ignorance of delay correlation along the timing path and the empirical feature selection solutions. In this work, an accurate and efficient prerouting path delay prediction framework is proposed by utilizing a transformer network and residual model with an ensemble feature selection mechanism. Owing to the combined filter and wrapper methods, an ensemble feature selection mechanism is implemented to determine the optimal feature subset based on the timing and physical information at the placement stage for path delay prediction, which is extracted as feature sequences for each cell along the timing path to be trained by transformer network. With the residual model, the predicted timing mismatch between the placement and routing stages by the transformer network is further calibrated to estimate the post-routing path delay. The proposed framework has been validated with ISCAS’85 and OpenCores benchmark circuits for the prediction of post-routing path delay, where the perdition error in terms of relative root mean squared error is limited within 1.3% and 3.0% and the correlation coefficient$R$is higher than 0.999 and 0.995 for seen and unseen circuits, respectively, indicating an error reduction by 2.3–10.6 times compared by prior learning-based models. In addition, the framework achieves average three orders of magnitude speedup compared with the commercial tools and is accelerated by a factor of 14–128 as against the competitive learning models, which is promising to be applied to guide design optimization prior to time-consuming routing stage.
Peng Cao 0002, Guoqing He, Tai Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A Timing Yield Model for SRAM Cells at Sub/Near-Threshold Voltages Based on a Compact Drain Current Model
abstract
Sub/near-threshold static random-access memory (SRAM) design is crucial for addressing the memory bottleneck in power-constrained applications. However, the high integration density and reliability under process variations demand an accurate estimation of extremely small failure probabilities. To capture such a “rare event” in memory circuits, the time and storage overhead of conventional simulations based on the Monte Carlo (MC) analysis cannot be tolerated. On the other hand, classic analytical methods predicting failure probabilities from a physical expression become inaccurate in the sub/near-threshold voltage domain due to the hypothetical distribution or the oversimplified drain current ($I_{ds}$) model for nanoscale devices. This work first proposes a simple but efficient empirical$I_{ds}$model to describe the drain-induced barrier lowering (DIBL) effect. Based on that, the probability density functions of the interest metrics in SRAM are derived. Two analytical models are then put forward to evaluate SRAM dynamic stabilities, including the access time failure and the write failure. The proposed models can be extended easily to different types of SRAM with different read/write-assist circuits. The models are validated against MC simulations across different operating voltages and temperatures. The average relative errors at 0.5-V$V_{\mathrm{ DD}}$are only 8.8% for the access-time failure model and 10.4% for the write failure model. The size of the required sample data set is$43.6\times $smaller than that of the state-of-the-art method.
Shan Shen, Peng Cao 0002, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Efficient and Accurate ECO Leakage Optimization Framework With GNN and Bidirectional LSTM
abstract
Engineering change order (ECO) plays an important role in design flow to perform leakage optimization with gate-sizing and$V_{\mathrm{ th}}$assignment approaches. Unfortunately, it is extremely time consuming due to the iterative nature of cell swap and timing check. Many learning-based methods, especially, graph neural networks (GNNs), have been utilized in leakage optimization to predict$V_{\mathrm{ th}}$assignment, but most of them treat the cells and their neighborhood cells uniformly when aggregating cell-level topology information to gather design-level information and discard the path-level information, suffering from accuracy loss, which could be exploited by bidirectional long short-term memory (BiLSTM) network. In this work, a GNN-BiLSTM-based framework is proposed to perform commercial-quality$V_{\mathrm{ th}}$assignment for leakage optimization by learning design-level and path-level information and is validated with the benchmarks from Opencores and IWLS 2005 under TSMC 28 nm technology. The experimental results demonstrate that the proposed framework achieves the most accurate$V_{\mathrm{ th}}$assignment prediction compared with the competitive models with F1-score ranging from 0.954 to 0.975 for seen designs and from 0.945 to 0.965 for unseen designs, respectively. The divergence between the leakage optimization results of this work and the commercial tool is limited to be between 8.5% and 26.1%, which is reduced by at least$2.2\times $compared with prior works. Owing to efficient training convergence and inference speed, our approach achieves significant runtime improvement by up to$10\times $over commercial tool with similar leakage optimization results.
Peng Cao 0002, Guoqing He, Zhanhua Zhang, Jun Yang 0006
IEEE Trans. Very Large Scale Integr. Syst.1
2022 A Graph Neural Network Method for Fast ECO Leakage Power Optimization
abstract
In modern design, engineering change order (ECO) is often utilized to perform power optimization including gate-sizing and Vth-assignments, which is efficient but highly timing consuming. Many graph neural network (GNN) based methods are recently proposed for fast and accurate ECO power optimization by considering neighbors' information. Nonetheless, these works fail to learn high-quality node representations on directed graph since they treat all neighbors uniformly when gathering their information and lack local topology information from neighbors one or two-hop away. In this paper, we introduce a directed GNN based method which learns information from different neighbors respectively and contains rich local topology information, which was validated by the Opencores and IWLS 2005 benchmarks with TSMC 28nm technology. Experimental results show that our approach outperforms prior GNN based methods with at least 7.8% and 7.6% prediction accuracy improvement for seen and unseen designs respectively as well as 8.3% to 29.0% leakage optimization improvement. Compared with commercial EDA tool PrimeTime, the proposed framework achieves similar power optimization results with up to 12X runtime improvement.
Peng Cao 0002
ASP-DAC2
2022 Pre-Routing Path Delay Estimation Based on Transformer and Residual Framework
abstract
Timing estimation prior to routing is of vital importance for optimization at placement stage and timing closure. Existing wire- or net-oriented learning-based methods limits the accuracy and efficiency of prediction due to the neglect of the delay correlation along path and computational complexity for delay accumulation. In this paper, an efficient and accurate pre-routing path delay prediction framework is proposed by employing transformer network and residual model, where the timing and physical information at placement stage is extracted as sequence features while the residual of path delay is modeled to calibrate the mismatch between the pre- and post-routing path delay. Experimental results demonstrate that with the proposed framework, the prediction error of post-routing path delay is less than 1.68% and 3.12% for seen and unseen circuits in terms of rRMSE, which is reduced by 2.3~5.0 times compared with exiting learning-based method for pre-routing prediction. Moreover, this framework produces at least three orders of magnitude speedup compared with the traditional design flow, which is promising to guide circuit optimization with satisfying prediction accuracy prior to time-consuming routing and timing analysis.
Tai Yang, Guoqing He, Peng Cao 0002
ASP-DAC3
2021 A Timing Prediction Framework for Wide Voltage Design with Data Augmentation Strategy
abstract
Wide voltage design has been widely used to achieve power reduction and energy efficiency improvement. The consequent increasing number of PVT corners poses severe challenges to timing analysis in terms of accuracy and efficiency. The data insufficiency issue during path delay acquisition raises the difficulty for the training of machine learning models, especially at low voltage corners due to tremendous library characterization effort and/or simulation cost. In this paper, a learning-based timing prediction framework is proposed to predict path delays across wide voltage region by LightGBM (Light Gradient Boosting Machine) with data augmentation strategies including CTGAN (Conditional Generative Adversarial Networks) and SMOTER (Synthetic Minority Oversampling Technique for Regression), which generate realistic synthetic data of circuit delays to improve prediction precision and reduce data sampling effort. Experimental results demonstrate that with the proposed framework, the path delays at low voltage could be predicted by their delays at high voltage corners with rRMSE of less than 5%, owing to the data augmentation strategies which achieve significant prediction error reduction by up to 12x.
Peng Cao 0002, Tai Yang
ASP-DAC1
2021 An Adaptive Delay Model for Timing Yield Estimation under Wide-Voltage Range
abstract
Yield analysis for wide-voltage circuit design is a strong nonlinear integration problem. The most challenging task is how to accurately estimate the yield of long-tail distribution. This paper proposes an adaptive delay model to substitute expensive transistor-level simulation for timing yield estimation. We use the Low-Rank Tensor Approximation (LRTA) to model the delay variation from a large number of process parameters. Moreover, an adaptive nonlinear sampling algorithm is adopted to calibrate the model iteratively, which can capture the larger variability of delay distribution for different voltage regions. The proposed method is validated on benchmark circuits of TAU15 in 45nm free PDK. The experiment results show that our method achieves 20-100X speedup compared to Monte Carlo simulation at the same accuracy level.
Hao Yan 0002, Xiao Shi 0001, Chengzhen Xuan, Peng Cao 0002, Longxing Shi
ASP-DAC4
2021 Semi-Analytical Path Delay Variation Model With Adjacent Gates Decorrelation for Subthreshold Circuits
abstract
The subthreshold circuit is a practical design style for the ultralow-power applications, but its timing estimation is a challenge due to the increasing local variation effects. The delay variation of adjacent gates is not independent because of input slew variation caused by the precedent gate, so their correlation effects are difficult to model and estimate. This article proposes a semi-analytical statistical delay model considering local variation for the subthreshold region, that is, the combination of analytical and simulation-based method. First, it decorrelates the slew influence between adjacent stages by dividing delay and output slew model into fast/slow input cases and dividing delay variation model into process variation and input slew variation. Then, it can be applied into multi-PVT conditions with a one-time SPICE nominal simulation by analyzing the independence of variability and relative variability of step input gate delay variance with output load capacitance and process, voltage, and temperature (PVT). Finally, experiments are carried out for different benchmarks, processes, voltages, and temperatures (BPVTs). The average errors of variance on different BPVTs are 4.8%, 3.1%, 4.0%, and 4.7%. Compared with other analytical works, the accuracies' improvements of three metrics (variance, variability, and max delay) are 8.3×, 9.6×, and 2.7× by the mean error at all test benchmarks. Compared with industrial method LVF, it has a comparable error in max path delay and runtime, and less three orders of magnitudes than LVF in characterization time and stored data (from TB to GB) at all test benchmarks.
Peng Cao 0002, Mengxiao Li, Yu Gong 0002, Zhiyuan Liu 0011, Geng Bai, Jun Yang 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 A Learning-Based Timing Prediction Framework for Wide Supply Voltage Design
abstract
Wide voltage design provides the tremendous benefits for state-of-the-art circuit design in terms of power consumption reduction and energy efficiency enhancement. The traditional design and verification flow depends on the standard cell libraries, which are only available from foundries for limited PVT (Process-Voltage-Temperature) corners near the nominal voltages, leading to remarkable characterization effort and storage overhead. In this paper, a learning-based framework is proposed to predict circuit path delays across multiple voltages and process corners without the requirement of cell library for each PVT corner, which consists of dilated-CNN (Conventional Neural Network) based feature engineering and ensemble model. The proposed method was verified with the supply voltages ranging from 0.5V to 0.9V under FF, SS and TT corners. Experimental results demonstrate that the prediction error is limited by 4.9% and 7.9% respectively within and across process corners for various working temperatures, which achieves significant precision enhancement compared with related learning-based methods.
Peng Cao 0002, Hao Cai 0001, Aiguo Bu
ACM Great Lakes Symposium on VLSI2
2020 Statistical Timing Model for Subthreshold Circuit with Correlated Variation Consideration
abstract
Subthreshold circuit has the significant advantage in low-power applications but suffers from serious variation increase. Traditional EDA tools could not achieve balance between accuracy and simulation effort for statistical timing analysis in subthreshold region. Many researches have been devoted to statistical timing model to reveal the relation between delay variation and process variation with physical insight. However, the statistical correlation of gate delay in circuit path is hard to capture and not considered appropriately in most prior works. In this paper, a statistical timing model for subthreshold circuit is proposed with the consideration of local process variation and correlated variation from adjacent gates, which is established by deriving variance models for gate delay and output waveform analytically for fast and slow input waveform separately, so that the path delay variation can be translated from the accumulation of the correlated gate delays with input slew to a linear combination of independent step input delays. The proposed model was verified under the process of TSMC28nm technology at the subthreshold supply voltage with the estimation error of less than 6% for circuit path delay variation in benchmark ISCAS99 compared with Monto Carlo simulation results, which outperforms prior works with 5~10X accuracy increase and acceptable simulation cost.
Peng Cao 0002, Mengxiao Li, Zhiyuan Liu 0011, Jun Yang 0006
ISCAS2
2019 A Statistical Current and Delay Model Based on Log-Skew-Normal Distribution for Low Voltage Region
abstract
The increasing performance variation and non-Gaussian distribution pose remarkable challenges to timing analysis for circuits operating in low voltage region. Accurate modeling of the statistical characteristics is urgently required with process variation consideration. In this paper, the statistical models for drain current and gate delay in low voltage region are established in analytical form based on the log-skew-normal (LSN) distribution via moment matching technique. Experimental results show that the probability distribution function (PDF) curves obtained from the proposed models for drain current and gate delay are highly fitted with Monte Carlo (MC) simulation results in sub/near-threshold regions. Moreover, owing to the proposed LSN-based statistical model, less than 8% error is introduced in the predicted sensitivity of gate delay and the maximum/minimum delay indicated by ±3σ percentile points can be calculated more precisely than the LN-based method with up to 3× accuracy improvement for low supply voltage.
Peng Cao 0002, Jiangping Wu, Zhiyuan Liu 0011, Jun Yang 0006, Longxing Shi
ACM Great Lakes Symposium on VLSI1
2019 A Statistical Timing Model for Low Voltage Design Considering Process Variation
abstract
Near-threshold voltage (NTV) design suffers severe challenge due to the dramatic increase in performance uncertainty introduced by process variation. This paper proposes an analytical approach based on Log-Normal (LN) distribution to characterize the statistical delay for NTV design considering the dominant threshold voltage variation from gate-level to circuit-level. At gate-level, the multivariable threshold voltage variation issue is solved by the equivalent threshold voltage method and equivalent drain current method for generic gates with stack topology and parallel topology, respectively. At circuit-level, a statistical timing model is proposed as the linear combination of the independent statistical delays of all gates in the path with step input by considering the varied correlation between adjacent gates. To the best of our knowledge, we firstly propose a statistical timing model analytically for practical circuit path with physical insights of supply voltage, transistor size, and load capacitance. The characterization effort for each path is only one-time SPICE simulation, which is negligible compared with Monte Carlo (MC) simulation in statistical static timing analysis (SSTA) methods. Experimental results under a commercial 28-nm CMOS process show the proposed models have high accuracy at low supply voltage compared with MC simulations, where the modeling errors for the mean and variance of gate delay can be limited within 1.54% and 11.2%, respectively. Moreover, as for the practical paths in ITC'99 benchmark, the maximum modeling errors of mean, variance, minimum delay, and maximum delay is less than 3.04%, 11.40%, 4.50%, and 2.87%, respectively.
Peng Cao 0002, Zhiyuan Liu 0011, Jiangping Wu, Jun Yang 0006, Longxing Shi
ICCAD1
2017 Context Management Scheme Optimization of Coarse-Grained Reconfigurable Architecture for Multimedia Applications
abstract
Due to the combination of flexibility and efficiency, coarse-grained reconfigurable architectures (CGRAs) are suitable for the implementation of computing-intensive applications. However, with the growing performance requirements, the scale of CGRA increases exponentially, which leads to configuration performance degradation and configuration power rise. Based on the analysis of configuration context features, we optimize the context management scheme of CGRA from the aspects of context cache structure and replacement strategy. The context cache is structured hierarchically to reduce the memory overhead without configuration performance degradation and a hybrid context replacement algorithm is proposed to further increase the configuration efficiency with a novel context frequency weight factor. Experimental results show that the proposed context management scheme improves the configuration performance of the base CGRA significantly by 13.6%-20.5% for H.264 decoding and 13.6%-20.5% for MPEG2 decoding with only 43% context cache cost. Compared with other works, the proposed context management scheme shows the advantages of 2.3-6× less normalized context cache size and 2.3-2.7× cache efficiency.
Peng Cao 0002, Bo Liu 0019, Jinjiang Yang, Jun Yang 0006, Meng Zhang 0010, Longxing Shi
IEEE Trans. Very Large Scale Integr. Syst.1
2015 An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding
abstract
A coarse-grained reconfigurable processing unit (RPU) consisting of 16 ×16 multi-functional processing elements (PEs) interconnected by an area-efficient line-switched mesh connect (LSMC) routing is implemented on a 5.4 mm ×3.1 mm die in TSMC 65 nm LP1P8M CMOS technology. A hierarchical configuration context (HCC) organization scheme is proposed to reduce the implementation overhead and the energy dissipation spent on fast reconfiguration. The proposed RPU is integrated into two system-on-a-chips (SoCs), targeting multiple-standard video decoding. The high-performance chip, comprising two RPU processors (named REMUS_HPP), can decode 1920 ×1080 H.264 video streams at 30 frames per second (fps) under 200 MHz. REMUS_HPP achieves a 25% performance gain over the XPP-III reconfigurable processor with only 280 mW power consumption, resulting in a 14.3 × improvement on energy efficiency. The other chip (named REMUS_LPP), targeting low power applications, integrates only one RPU processor. REMUS_LPP can decode 720 ×480 H.264 video streams at 35fps with 24.5 mW under 75 MHz, achieving a 76% reduction in power dissipation and a 3.96 × improvement on energy efficiency compared with the ADRES reconfigurable processor.
Leibo Liu, Dong Wang 0040, Min Zhu 0001, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Jun Yang 0006, Shaojun Wei
IEEE Trans. Multim.6
2015 Correction to "An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding"
Leibo Liu, Dong Wang 0040, Min Zhu 0001, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Jun Yang 0006, Shaojun Wei
IEEE Trans. Multim.6
2014 Configuration approaches to improve computing efficiency of coarse-grained reconfigurable multimedia processor
abstract
This paper proposes three configuration approaches to improve computing efficiency of a coarse-grained reconfigurable array, including input data relocation, line-based context switching, and loop interval minimization. These proposed approaches fully exploit the parallelism and pipelining of the reconfigurable array, which reduce interval latency when switching the configuration contexts, and therefore greatly enhance computing efficiency. These proposed techniques are used in a coarse-grained reconfigurable multimedia system (REMUS). Measured results show that, owing to the proposed approaches, REMUS can achieve 1080p@30fps performance for H.264 high profile video decoding under 200MHz working frequency. When normalized to the same technology, REMUS outperforms XPP-III 6.98x in energy efficiency.
Chen Yang 0005, Leibo Liu, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Shaojun Wei
FPL5
2014 Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor
Leibo Liu, Victor Y. Chen, Dong Wang 0040, Shouyi Yin, Peng Cao 0002, Shaojun Wei
Sci. China Inf. Sci.8
2014 On-Chip Memory Hierarchy in One Coarse-Grained Reconfigurable Architecture to Compress Memory Space and to Reduce Reconfiguration Time and Data-Reference Time
abstract
The coarse-grained reconfigurable architecture (CGRA) is proven to be energy efficient in several specific domains. In CGRAs, the on-chip memory hierarchy, which contains the context memory and the data memory organizations, should be well considered to achieve appropriate tradeoffs among three aspects: 1) performance; 2) area; and 3) power. In this paper, two techniques called the hierarchical configuration context (HCC) and the lifetime-based data-memory organization (LDO) focusing on the context memory and the data memory organizations are proposed to compress the on-chip memory space and to reduce the reconfiguration time and the data-reference time. In the HCC, the contexts are constructed in a hierarchical fashion to completely eliminate the repetitive portions of the contexts, not only reducing the overall context storage, but also alleviating the context transportation overhead. A fast context-indexing mechanism in the HCC is proposed to achieve fast reconfiguration, as the hierarchically organized contexts can be located and accessed conveniently. In the LDO, the on-chip data are classified into two types, based on the lifetime of data. The short-lifetime data are stored in the first in first out to increase the reuse ratio of memory space automatically, whereas the long-lifetime data are stored in the radom access memory for several time references. The HCC and the LDO are used in a CGRA core called as reconfigurable processing unit (RPU). Two RPUs are integrated in a reconfigurable computing processor (RCP) called as REconfigurable MUlti-media System, High-Performance Processor (REMUS_HPP). Because of the HCC, compared with a traditional nonhierarchical system, the total context storage required in H.264 decoding is reduced by 77%. Because of the LDO, the normalized on-chip data memory size at same performance level in the REMUS_HPP is only 23.8% and 14.8% of those in XPP-III (a high-performance RCP) and ADRES (a low-power RCP). REMUS_HPP is implemented on a 48.9-mm2silicon with TSMC 65-nm technology, using a 200-MHz working frequency to achieve 1920 × 1088 at 30 fps H.264 high-profile decoding. Compared with XPP-III, the performance of the REMUS_HPP is 1.81× boosted, whereas the energy efficiency is 4.75× higher.
Yansheng Wang, Leibo Liu, Shouyi Yin, Min Zhu 0001, Peng Cao 0002, Jun Yang 0006, Shaojun Wei
IEEE Trans. Very Large Scale Integr. Syst.5
2013 Implementation of multi-standard video decoding algorithms on a coarse-grained reconfigurable multimedia processor
abstract
This paper proposed a THPHP (Task-based Hybrid Parallels and Hybrid Pipelines) scheme to implement multistandard video decoding algorithms, i.e. MPEG-2, H.264 and AVS (Audio Video coding Standard), on a heterogeneous coarsegrained reconfigurable multimedia processor called REMUS (REconfigurable MUltimedia System). Multiple level parallelism and multiple level pipeline techniques are proposed in this scheme. Simulation results show that the video decoder can support H.264 HP (High Profile) 1920×1080@30fps (frame per second) streams, AVS JP (Jizhun Profile) 1920×1080@39fps streams, and MPEG-2 MP (Main Profile) 1920×1080@41fps streams when exploiting a 200MHz working frequency.
Leibo Liu, Victor Y. Chen, Shouyi Yin, Dong Wang 0040, Shaojun Wei, Li Zhou 0015, Peng Cao 0002
ISCAS9
2013 Hierarchical representation of on-chip context to reduce reconfiguration time and implementation area for coarse-grained reconfigurable architecture
Yansheng Wang, Leibo Liu, Shouyi Yin, Min Zhu 0001, Peng Cao 0002, Jun Yang 0006, Shaojun Wei
Sci. China Inf. Sci.5
2009 Area-efficient line-based two-dimensional discrete wavelet transform architecture without data buffer
abstract
An area-efficient architecture for 2D DWT is proposed in this paper based on novel decomposed lifting scheme, where no data buffer is required to preserve and reorder the intermediate data between the row and column processor. Compared with the reported research, the proposed design could benefit from the reduction of internal memory size and the number of multipliers, adders and registers. The design was implemented for 2D 9/7 and 5/3 DWT in SMIC 0.18 mum CMOS logic fabrication with 15 K equivalent 2-input NAND gates under 150 MHz, which can accommodate up to 512times512 image size with 4 K bytes on-chip dual-port RAM.
Peng Cao 0002, Chao Wang 0068, Jun Yang 0006, Longxing Shi
ICME1
2007 An Efficient VLSI Architecture for Lifting-Based Discrete Wavelet Transform
abstract
In this paper, we propose an efficient VLSI architecture which performs the two-dimensional (2-D) discrete wavelet transform (DWT) of 9/7 filter for JPEG2000. Based on the modified lifting-based DWT algorithm, an efficient VLSI architecture for one-dimensional (1-D) DWT is derived to reduce the hardware cost and shorten the critical path. The proposed 2-D DWT architecture is composed of two 1-D processors (row and column processors). Based on the line-based architecture, the column processor can start column-wise transform while only two rows have been processed For an M×N image, only 5.5N internal memory is required for the 9/7 filter to perform the 2-D DWT with the critical path of one multiplier. Finally, Verilog simulation results are presented to show that the proposed architecture in comparison with other existing architectures is fast and efficient for the 2-D DWT computation.
Chao Wang 0068, Wu Zhilin, Peng Cao 0002, Li Jie
ICME3