Longxing Shi

dblp:89/1700 · also Longxin Shi · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
28since 2021 · last 2026
0000-0002-0629-7154ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 28 since 2021Software engineering, systems software and programming languages · 7 · 5 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 GoG-Predict: IR-aware Path Waveform Prediction with Structural Entity Interaction Learning
abstract
In advanced technology nodes, Current Source Models (CSMs) are widely adopted due to their high accuracy. However, IR-drop–induced supply voltage fluctuations increase the computational burden of CSM-based timing analysis. Existing machine-learning (ML) approaches accelerate CSM evaluation by assuming fixed receiver capacitance along the path, which contradicts the variable-load dependency inherent to CSM formulation. In this work, we propose GoG-Predict, a fast and accurate IR-aware path waveform prediction framework. GoG-Predict employs a Graph-of-Graphs (GoG) Neural Network to model the structured interactions among driver cells, load nets, and receiver cells. In addition, a Feature-wise Linear Modulation (FiLM) mechanism is incorporated to account for the impact of IR drop on output waveform. We evaluate GoG-Predict on open-source designs using a commercial 12nm technology. Experimental results show an average waveform error of 2.03% and a 1426 × runtime speedup over HSPICE, while achieving 1.63 × higher accuracy compared with state-of-the-art ML-based prediction methods.
Yanglong Mao, Ziyue Han, Yunfan Zuo, Chenpu Shi, Yaning Jia, Hao Yan 0002, Longxing Shi
ACM Great Lakes Symposium on VLSI8
2026 Harnessing Spatiotemporal Redundancy for Fast Diffusion Models on FPGA
Dongge Qin, Junhong Qian, Shaoqiang Lu, Yangbo Wei, Ruizhe Deng, Xiao Shi 0001, Longxing Shi, Lei He 0001
ISCAS8
2026 DCTDSE: A Bimodal Design Space Exploration Flow via Discrete-Continuous Transformation
abstract
The conservation core accelerator presents a promising avenue for improving the computational efficiency of specific applications. However, its design space is ultra-high-dimensional, significantly increasing the exploratory effort required to identify the optimal design across performance, power, and area metrics. Furthermore, the discrete nature of the microarchitecture space renders conventional search methods ineffective. To tackle these challenges, we propose the Discrete Continuous Transformation to speedup Design Space Exploration, namely DCTDSE. It can operate in either offline or online mode. In offline mode, it transforms the original discrete design space into a continuous space, builds predictive models, performs parallel gradient-based optimization, and maps the results back to the discrete domain. In the online mode, DCTDSE refines the models by iteratively resampling previously found solutions, thereby enhancing exploration quality while maintaining moderate runtime overhead. Experimental results indicate that DCTDSE achieves a 3.9× to 40× speedup over benchmark methods in offline mode. In online mode, it provides a 2.5× speedup, with a 21% reduction in exploration quality relative to the most accurate comparison method.
Shuaibo Huang, Liangji Wu, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 Learning-Driven Hierarchical Particle Swarm Optimization Framework for Power Delivery Network Synthesis
abstract
As process nodes shrink, intensified IR drop and electromigration (EM) issues, driven by the power delivery network (PDN), affect reliability and further compress the routability of the entire design. Since physical design is time-consuming, effectively predicting and optimizing PDN during the power planning stage is crucial. This effort is complicated by complex interactions between multiple parameters and the challenges associated with nonconvex optimization. To overcome these limitations, we propose a learning-driven hierarchical particle swarm optimization framework for PDN, integrated with a residual TransUNet-based prediction model for post-placement congestion, EM, and IR drop prediction. In addition, we introduce the gate network for loss fusion, which excels in balancing output predictions and enhancing the accuracy of critical tasks. Analogous to hyperparameter optimization in deep learning, we can utilize GPU parallelization to search for the global optimum of the PDN structure efficiently. In circuits at the 22nmand 16nmtechnology nodes, with cell counts ranging from 560 to nearly 200, 000, our method achieves reductions in mean and maximum congestion by 15% and 49%, respectively, compared to state-of-the-art methods. In addition, it achieves an overall optimization time of around 22 seconds.
Yunfan Zuo, Yuwei Sun, Pinquan Li, Yan Li 0056, Shuo Cui, Hao Yan 0002, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2026 Scalable Yield Analysis of SRAM and Analog Circuits Using Multi-Kernel Sparse Representation
abstract
With the advancement of technology nodes and the increasingly stringent requirements for stability, general yield analysis of customized circuits in the early stages of design has become a key bottleneck in manufacturing. In this article, we propose a multi-kernel sparse representation-based classification (MKSRC) method to enhance the efficiency and scalability of failure probability estimation by classifying tail samples. It employs class-balanced sampling to address data imbalance issues and utilizes multi-kernel features with adaptive kernel weights to enhance the accuracy and robustness of the classifier. Experimental results on 32-bit SRAM columns and analog circuits demonstrate that the proposed MKSRC method achieves higher classification accuracy and efficiency compared to other state-of-the-art methods, particularly in scenarios with limited training data. Compared to SOTA yield estimation methods, the MKSRC method achieves an average 2.57–3.29× improvement in both accuracy and efficiency, highlighting its ability to provide efficient and scalable yield analysis solutions for both SRAM and analog circuits.
Liangji Wu, Zhongxi Guo, Xiao Shi 0001, Longxing Shi
ACM Trans. Design Autom. Electr. Syst.6
2026 FACT: Fast and Accurate Multi-Corner Predictor for Timing Closure in Commercial EDA Flows
abstract
With technology scaling progressing well into deep nanometer region, the number of technology corners surges from dozens to hundreds. This dramatic increase in corners significantly complicates timing closure during Engineering Change Orders (ECO) stages, as performing full-corner static timing analysis (STA) becomes increasingly time-consuming and challenging. Existing methodologies often leverage machine learning (ML) techniques to predict unknown corners based on a subset of known corners. However, as the total number of corners expands, these methods not only require a larger set of known corners but also become highly sensitive to known corner selection. Additionally, as designers push designs into near-threshold voltage regions to enhance energy efficiency, the number of corners increases and the nonlinearity between corners becomes more pronounced. This intensifies the difficulty for current ML-based methods to accurately predict full-corner timing metrics. In this work, we propose FACT, a fast and accurate multi-corner predictor for timing closure optimization. Our approach simplifies the process by necessitating timing analysis under only one known corner. By effectively capturing correlations across diverse technology files, FACT robustly infers full-corner timing metrics, even under challenging near-threshold conditions. Moreover, our framework seamlessly integrates with commercial EDA design flows, making it practical in industrial environments. Experimental results on open-source designs indicate the superior stability of our method, coupled with a significant runtime speed-up over both traditional and prior ML-based timing ECO flows.
Ziyue Han, Shiyang Wu, Hao Yan 0002, Longxing Shi
ACM Trans. Design Autom. Electr. Syst.6
2026 Enhanced TransUNet Framework for Predicting Static IR Drop and Chip Routability
abstract
As semiconductor processes advance, the power delivery network (PDN) increasingly affects the power supply from the pads to the cells. Significant IR drops in standard cells can lead to timing violations, while suboptimal PDN topologies can lead to increased congestion. Together, these factors degrade overall chip performance and reliability. To accelerate design iteration, accurately and efficiently predicting unevenly distributed IR drop and congestion, especially in hotspot areas, has become a critical challenge. This article introduces an enhanced TransUNet-based framework for distribution-aware static IR drop and congestion prediction, treating both problems as separate but related image prediction tasks. Such an abstraction preserves the IR drop and congestion distribution patterns on the original real physical layout and retains local hot spots. Our proposed framework leverages image classification techniques to model IR drop prediction as a spatial pattern recognition task, effectively addressing the long-tail distribution in different regions. To enhance hotspot prediction, we incorporate wavelet transform and transformer-based analysis to enable multiscale feature fusion. In the open-source CircuitNet dataset, our method predicts static IR drop with a mean absolute error( MAE ) of 0.374 mV and a maximum error rate( Err m ) of 18.7%, reducing MAE and Err m by 77.4% and 79.2%, respectively, compared to the state-of-the-art method, all within 100 ms. The congestion prediction evaluations show 65.3% lower NRMSE scores and 18.4% higher SSIM scores relative to the existing SOTA approach. Our approach accurately and reliably predicts long-tail distributions and localized hotspots in both IR drop and congestion tasks.
Yunfan Zuo, Pinquan Li, Yuwei Sun, Hao Yan 0002, Longxing Shi
ACM Trans. Design Autom. Electr. Syst.5
2025 Rank-based Multi-objective Approximate Logic Synthesis via Monte Carlo Tree Search
abstract
Approximate Logic Synthesis (ALS) is an automated technique designed for error-tolerant applications, optimizing delay, area, and power under specified error constraints. However, existing methods typically focus on either delay reduction or area minimization, often leading to local optima in multi-objective optimization. This paper proposes a rankbased multi-objective ALS framework using Monte Carlo Tree Search (MCTS). It develops non-dominated circuit ranking, to guide MCTS in exploring local approximate changes (LACs) across the entire circuit and generate approximate circuit sets with great optimization potential. Additionally, a Rank-Transformer model is introduced to predict pathdomain ranks, enhancing the application of high-quality LACs within circuit paths. Experimental results show that our framework achieves faster and more efficient optimization in delay and area simultaneously compared to state-of-the-art methods.
Yuyang Ye 0001, Xiangfei Hu, Peng Xu 0052, Yu Gong 0002, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
DAC9
2025 Truly Pre-Routing Timing Prediction via Considering Power Delivery Network
abstract
Fast and accurate pre-routing timing prediction is essential in the chip design flow. However, existing machine learning (ML)assisted pre-routing timing methods often overlook the impact of power delivery networks (PDNs), which contribute to IR drop and routing congestion. This limitation can make these methods less practical for realworld circuit design flows. To address this, we propose two specialized encoders-an IR drop-aware encoder and a routing congestion-aware encoder-that effectively capture PDN effects through multimodal fusion of netlist, layout, and PDN data. To mitigate the challenges of imbalanced multimodal fusion, we further develop a Pareto optimization approach to ensure balanced utilization of all modalities, enhancing timing prediction accuracy. Comprehensive experiments on large-scale open-source designs using TSMC’s 16 nm technology node validate the superiority of our model over state-of-the-art pre-routing timing prediction methods.
Yuyang Ye 0001, Mingwei He, Lizheng Ren, Jianwang Zhai, Tinghuan Chen, Jun Yang 0006, Longxing Shi
DAC7
2025 An Imitation Augmented Reinforcement Learning Framework for CGRA Design Space Exploration
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) are a promising architecture that warrants thorough design space exploration (DSE). However, traditional DSE methods for CGRAs often get trapped in local optima due to singularities, i.e., invalid design points caused by CGRA mapping failures. In this paper, we propose a singularity-aware framework based on the integration of reinforcement learning (RL) and imitation learning (IL) for DSE of CGRAs. Our approach learns from both valid and invalid points, substantially reducing the probability of sampling singularities and accelerating the escape from inefficient regions, ultimately achieving high-quality Pareto points. Experimental results demonstrate that our framework improves the hypervolume (HV) of the Pareto front by 23.56% compared to state-of-the-art methods, with a comparable time overhead.
Liangji Wu, Shuaibo Huang, Shiyang Wu, Hao Yan 0002, Longxing Shi
DATE7
2025 ARS-Flow 2.0: An enhanced design space exploration flow for accelerator-rich system based on active learning
Shuaibo Huang, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
Integr.4
2025 Learning-Driven Physically Aware Large-Scale Circuit Gate Sizing
abstract
Gate sizing plays an important role in timing optimization after physical design. Existing machine learning-based gate sizing works cannot optimize timing on multiple timing paths simultaneously and neglect the physical constraint on layouts. They cause suboptimal sizing solutions and low-efficiency issues when compared with commercial gate sizing tools. In this work, we propose a learning-driven physically aware gate sizing framework to optimize timing performance on large-scale circuits efficiently. In our gradient descent optimization-based work, for obtaining accurate gradients, a multimodal gate sizing-aware timing model is achieved via learning timing information on multiple timing paths and physical information on multiple-scaled layouts jointly. Then, gradient generation based on the sizing-oriented estimator and adaptive back-propagation are developed to update gate sizes. Our results demonstrate that our work achieves higher-timing performance improvements in a faster way compared with the commercial gate sizing tool.
Yuyang Ye 0001, Peng Xu 0052, Lizheng Ren, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 ARS-Flow: A Design Space Exploration Flow for Accelerator-rich System based on Active Learning
abstract
Surrogate model-based design space exploration (DSE) is the mainstream method to search for optimal microarchitecture designs. However, it is hard to build accurate models for accelerator-rich systems within limited samples due to its high dimensional characteristic. Moreover, it is easy to fall into local optimal or difficult to converge. To solve these two problems, we propose a DSE flow based on active learning, namely ARS-Flow. It is featured with Pareto-region-oriented stochastic resampling method (PRSRS) and multiobjective genetic algorithm with self-adaptive hyperparameter control (SAMOGA). Taking the gem5-SALAM system for illustration, the proposed method can build more accurate models and find better microarchitecture designs with acceptable runtime costs.
Shuaibo Huang, Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
ASPDAC4
2024 A Deep-Learning-Based Statistical Timing Prediction Method for Sub-16nm Technologies
abstract
Pre-routing timing estimation is vital but challenging since accurate net information is available only after routing and parasitic extraction. Existing methodologies predict the timing metrics with the help of the placement information of standard cells. However, neglecting the analysis of process variation effects hinders the precision of those methodologies, especially in sub-16nm technologies as delay distributions become asymmetric. Therefore, a deep-learning-based statistical timing prediction method is proposed to model process variation effects in the pre-routing stage. Congestion features and pin-to-pin features are fed into graph neural networks for post-routing interconnect parasitic and arc delay prediction. Moreover, a calibration method is proposed to compensate for the precision loss of the delay propagation. We evaluate our methods using open-source designs and EDA tools, which demonstrate improved accuracy in pre-routing timing prediction methods and a remarkable speed-up compared to traditional routing and timing analysis process.
Leilei Jin, Wenjie Fu 0003, Longxing Shi
DATE4
2024 A Graph-Learning-Driven Prediction Method for Combined Electromigration and Thermomigration Stress on Multi-Segment Interconnects
abstract
As technology advances, the temperature gradient in the interconnects becomes more significant, which causes serious thermomigration. Simulating the coupling effects of thermomigration (TM) and electromigration (EM) on large-scale circuits is very time-consuming caused by a substantial increase in computational complexity. Recently, some researchers utilized graph learning-based methods to predict EM stress in medium-scale cases. Unfortunately, these works overlooked the effects of TM. To predict the EM - TM stress of large-scale interconnects accurately and efficiently, we propose a framework based on Graph Attention Networks (GATs) with a customized alternating aggregation method for collecting information in junctions and branches of interconnects jointly. The experimental results show that our work achieves an average relative error of less than 1 % compared to the commercial software COMSOL for inter-connects consisting of fewer than 200 segments. Furthermore, our method also achieves 9037 x speedup in predicting the OpenROAD test circuit with a maximum segment number reaching 10807.
Yunfan Zuo, Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Longxing Shi
DATE6
2024 Timing-Driven Technology Mapping Approximation Based on Reinforcement Learning
abstract
As the transistor technology nodes shrink into the nano-scales, timing guardbands caused by aging effects and process variations continue to increase. Approximate computing can eliminate aging-and-variation-induced timing guardbands without sacrificing the design performance. It can apply local approximate changes (LACs) automatically in circuits to reduce critical path delay. However, efficiently achieving timing optimization under error distance constraints is still tricky. This work proposes an automated timing-driven technology mapping approximation framework based on reinforcement learning (RL). The framework uses path-weighted graph neural networks (PGNNs) to embed RL states and timing path-aware LAC candidates to construct RL action spaces. It can efficiently eliminate timing guardbands induced by aging and variation. Our proposed circuit-agnostic framework operates on the gate-level netlists. According to experiments on the open-source circuits using TSMC 28nm and 16nm technology under aging and variation conditions, our framework can achieve an average 24.78% critical path delay reduction under 5 different error distance constraints and 4.83× speedup, compared with a state-of-the-art method.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 Graph-Learning-Driven Path-Based Timing Analysis Results Predictor from Graph-Based Timing Analysis
abstract
With diminishing margins in advanced technology nodes, the performance of static timing analysis (STA) is a serious concern, including accuracy and runtime. The STA can generally be divided into graph-based analysis (GBA) and path-based analysis (PBA). For GBA, the timing results are always pessimistic, leading to overdesign during design optimization. For PBA, the timing pessimism is reduced via propagating real path-specific slews with the cost of severe runtime overheads relative to GBA. In this work, we present a fast and accurate predictor of post-layout PBA timing results from inexpensive GBA based on deep edge-featured graph attention network, namely deep EdgeGAT. Compared with the conventional machine and graph learning methods, deep EdgeGAT can learn global timing path information. Experimental results demonstrate that our predictor has the potential to substantially predict PBA timing results accurately and reduce timing pessimism of GBA with maximum error reaching 6.81 ps, and our work achieves an average 24.80× speedup faster than PBA using the commercial STA tool.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
ASP-DAC6
2023 A Novel Delay Calibration Method Considering Interaction between Cells and Wires
abstract
In the advanced technology, the accuracy of cell and wire delay modeling are the key metrics for timing analysis. However, when the supply voltage decreases to the near-threshold regime, the complicated process variation effect causes the cell delay and the wire delay hard to model. Most researchers study cell or wire delay separately, ignoring the coefficients between them. In this paper, we propose an N-sigma delay model by characterizing different sigma levels$\mathbf{(-3\sigma {to}+3\sigma)}$of the cell and wire delay distribution. The N-sigma cell delay model is represented by the first four moments and calibrated by the operating conditions (input slew, output load). Meanwhile, based on the Elmore model, the wire delay variability is calculated by considering the effect of drive and load cells. The delay models are verified through the ISCAS85 benchmarks and the functional units of PULPino processor with TSMC 28 nm technology. Compared to the SPICE results, the average errors for estimating the$+/-\mathbf{3\sigma}$cell delay are 2.1 % and 2.7% and those of the wire delay are 2.4% and 1.6%, respectively. The errors of path delay analysis keep below 6.6% and the speed is 103X over SPICE MC simulations.
Leilei Jin, Wenjie Fu 0003, Hao Yan 0002, Xiao Shi 0001, Longxing Shi
DATE7
2023 Fast and Accurate Wire Timing Estimation Based on Graph Learning
abstract
Accurate wire timing estimation has become a bottleneck in timing optimization since it needs a long turn-around time using a sign-off timer. The gate timing can be calculated accurately using lookup tables in cell libraries. In comparison, the accuracy and efficiency of wire timing calculation for complex RC nets are extremely hard to trade-off. The limited number of wire paths opens a door for the graph learning method in wire timing estimation. In this work, we present a fast and accurate wire timing estimator based on a novel graph learning architecture, namely GNNTrans. It can generate wire path representations by aggregating local structure information and global relationships of whole RC nets, which cannot be collected with traditional graph learning work efficiently. Experimental results on both tree-like and non-tree nets demonstrate improved accuracy, with the max error of wire delay being lower than 5 ps. In addition, our estimator can predict the timing of over 200K nets in less than 100 secs. The fast and accurate work can be integrated into incremental timing optimization for routed designs.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
DATE6
2023 Optimized matrix ordering of sparse linear solver using a few-shot model for circuit simulation
Yuyang Ye 0001, Hao Yan 0002, Longxing Shi
Integr.5
2023 An efficient SRAM yield analysis method based on scaled-sigma adaptive importance sampling with meta-model accelerated
Liang Pang 0002, Mengyun Yao, Xiao Shi 0001, Hao Yan 0002, Longxing Shi
Integr.7
2023 A Timing Yield Model for SRAM Cells at Sub/Near-Threshold Voltages Based on a Compact Drain Current Model
abstract
Sub/near-threshold static random-access memory (SRAM) design is crucial for addressing the memory bottleneck in power-constrained applications. However, the high integration density and reliability under process variations demand an accurate estimation of extremely small failure probabilities. To capture such a “rare event” in memory circuits, the time and storage overhead of conventional simulations based on the Monte Carlo (MC) analysis cannot be tolerated. On the other hand, classic analytical methods predicting failure probabilities from a physical expression become inaccurate in the sub/near-threshold voltage domain due to the hypothetical distribution or the oversimplified drain current ($I_{ds}$) model for nanoscale devices. This work first proposes a simple but efficient empirical$I_{ds}$model to describe the drain-induced barrier lowering (DIBL) effect. Based on that, the probability density functions of the interest metrics in SRAM are derived. Two analytical models are then put forward to evaluate SRAM dynamic stabilities, including the access time failure and the write failure. The proposed models can be extended easily to different types of SRAM with different read/write-assist circuits. The models are validated against MC simulations across different operating voltages and temperatures. The average relative errors at 0.5-V$V_{\mathrm{ DD}}$are only 8.8% for the access-time failure model and 10.4% for the write failure model. The size of the required sample data set is$43.6\times $smaller than that of the state-of-the-art method.
Shan Shen, Peng Cao 0002, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Aging-Aware Critical Path Selection via Graph Attention Networks
abstract
In advanced technology nodes, aging effects like negative and positive bias temperature instability (NBTI and PBTI) become increasingly significant, making timing closure and optimization more challenging. Unfortunately, conventional critical path (CP) selection tools used in reliability-aware design flow cannot accurately identify CPs under different aging conditions. To address this issue, we propose an aging-aware CP selection flow comprising two parts: 1) critical cell detection and 2) path criticality (PC) computation. We employ graph-attention (GAT) networks to predict the critical cells in the aged circuits, and a PC computation algorithm that takes into account circuit-level and transistor-level parameters to generate PC rank lists. Our experimental results demonstrate that our GAT model outperforms classical machine learning models in detecting critical cells. Additionally, compared with the commercial tool, our aging-aware flow achieves an average accuracy of 99.52%, 98.69%, and 97.20% for top-10%, top-5%, and top-1% path sets, respectively, in five industrial designs subjected to different aging conditions and workloads.
Yuyang Ye 0001, Tinghuan Chen, Hao Yan 0002, Bei Yu 0001, Longxing Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 Quality Driven Systematic Approximation for Binary-Weight Neural Network Deployment
abstract
Neural networks (NNs) with large scales of artificial neurons are increasingly used in recognition and classification tasks. In power-constrained scenarios, the tradeoff between performance and hardware consumptions must be carefully evaluated before silicon tape-out. In this paper, we proposed a systematic approach to design ultra-low power NN system. This work is motivated by the facts that NNs are resilient to approximation in many of the computations and NNs are outputting statistical tensors which are acceptable to less-than-perfect results. We resort to the front-back end approach with a twofold aim: (1) a fast and accurate design approach is proposed by estimating the computing quality of low-power approximate adder arrays, and it is adopted to evaluate the neural network system; (2) a quality configurable engine with different approximation degrees while processing NNs is implemented. The proposed work is demonstrated with a comprehensive keyword spotting (KWS) system as an ultra-low power NN engine. The experimental environment is setup with ten keywords from the google speech command dataset (GSCD) using an industrial 22-nm ultra-low-leakage (ULL) process. Comparing to the state-of-the-art KWS processors, the proposed approximate NN engine can demonstrate over 60% improvement in power efficiency and$1.1\times $area efficiency while achieving similar recognition accuracy.
Yu Gong 0002, Hao Cai 0001, Haige Wu, Hao Yan 0002, Zhen Wang 0019, Longxing Shi, Bo Liu 0019
IEEE Trans. Circuits Syst. I Regul. Pap.7
2022 A Compact High-Dimensional Yield Analysis Method using Low-Rank Tensor Approximation
abstract
“Curse of dimensionality” has become the major challenge for existing high-sigma yield analysis methods. In this article, we develop a meta-model using Low-Rank Tensor Approximation (LRTA) to substitute expensive SPICE simulation. The polynomial degree of our LRTA model grows linearly with the circuit dimension. This makes it especially promising for high-dimensional circuit problems. Our LRTA meta-model is solved efficiently with a robust greedy algorithm and calibrated iteratively with a bootstrap-assisted adaptive sampling method. We also develop a novel global sensitivity analysis approach to generate a reduced LRTA meta-model which is more compact. It further accelerates the procedure of model calibration and yield estimation. Experiments on memory and analog circuits validate that the proposed LRTA method outperforms other state-of-the-art approaches in terms of accuracy and efficiency.
Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Chengzhen Xuan, Lei He 0001, Longxing Shi
ACM Trans. Design Autom. Electr. Syst.6
2022 A Fast Cross-Layer Dynamic Power Estimation Method by Tracking Cycle-Accurate Activity Factors With Spark Streaming
abstract
The advent of autonomous power-limited systems poses a new challenge for early design space exploration. The existing architecture-level power evaluation tools lose accuracy due to ignoring features of circuit-level behaviors and influences of process, voltage, and temperature variations. Although power estimations based on SPICE or PrimeTime PX (PTPX) are accurate enough, they come at the cost of long simulation time and are available only in very late phases of design flow. In this article, a fast and accurate dynamic power evaluation method is proposed, which estimates activity factors at the circuit level. The impact of process variation at the gate level is considered through the proposed effective capacitance model. Activity factors are then estimated by the model and input vectors of the circuit. Input vectors are generated by architecture-level simulations in the form of streaming. For real-time and high-speed power evaluation, a data streaming framework is proposed for massive parallelism. The cross-layer estimation is verified based on the functional units of PULPino processor running SPEC CPU2006 benchmarks. Compared with the SPICE results using SMIC 28-nm PDK, our cycle-by-cycle dynamic power analysis shows an average error of 5.4%. Meanwhile, our approach realizes 65.2% faster than the traditional PTPX simulation and 48.8% faster compared with the state-of-art cross-level evaluation method.
Leilei Jin, Wenjie Fu 0003, Longxing Shi
IEEE Trans. Very Large Scale Integr. Syst.4
2021 An Adaptive Delay Model for Timing Yield Estimation under Wide-Voltage Range
abstract
Yield analysis for wide-voltage circuit design is a strong nonlinear integration problem. The most challenging task is how to accurately estimate the yield of long-tail distribution. This paper proposes an adaptive delay model to substitute expensive transistor-level simulation for timing yield estimation. We use the Low-Rank Tensor Approximation (LRTA) to model the delay variation from a large number of process parameters. Moreover, an adaptive nonlinear sampling algorithm is adopted to calibrate the model iteratively, which can capture the larger variability of delay distribution for different voltage regions. The proposed method is validated on benchmark circuits of TAU15 in 45nm free PDK. The experiment results show that our method achieves 20-100X speedup compared to Monte Carlo simulation at the same accuracy level.
Hao Yan 0002, Xiao Shi 0001, Chengzhen Xuan, Peng Cao 0002, Longxing Shi
ASP-DAC5
2021 A Novel Digital Control Method of Primary-Side Regulated Flyback With Active Clamping Technique
abstract
Because of the different working principles between conventional flyback and active-clamped flyback (ACF), the existing primary-side regulation (PSR) technology cannot be applied to ACF. And in order to achieve PSR of ACF, a low-cost digital sampling method based on the auxiliary winding of the transformer is proposed. In the conventional peak current mode (PCM) controlled ACF, the sampling resistor of primary current may introduce additional power loss, and high-frequency oscillations during sampling may reduce system's stability. Therefore, a digital feedback-feedforward control (FFC) method is proposed to eliminate the sampling resistor. The stability of the proposed control method and its influence on the dynamic response are studied. In addition, since GaN device has higher power loss during reverse conduction than its Si counterpart, design considerations of reverse conduction time of power switches are also discussed in this paper. All control methods and design considerations are verified on a GaN-based 12V-3A ACF, and the proposed control strategy is implemented by FPGA.
Minggang Chen, Linlin Huang 0002, Weifeng Sun 0001, Longxing Shi
IEEE Trans. Circuits Syst. I Regul. Pap.5
2020 A Cross-Layer Power and Timing Evaluation Method for Wide Voltage Scaling
abstract
Wide supply voltage scaling is critical to enable worthwhile dynamic adjustment of the processor efficiency against varying workloads. In this paper, a cross-layer power and timing evaluation method is proposed to estimate the processor energy efficiency using both circuit and architectural information in a wide voltage range. The process variations are considered through statistical static timing analysis while the voltage effect is modeled through secondary iterated fittings. The error for estimating processor energy efficiency decreases to 8.29% when the supply voltage is scaled from 1.1V to 0.6V, while traditional architectural evaluations behave more than 40% errors.
Wenjie Fu 0003, Leilei Jin, Longxing Shi
DAC5
2020 TYMER: A Yield-based Performance Model for Timing-speculation SRAM
abstract
In low power designs, timing-speculative techniques are proposed to boost the SRAM frequency and throughput. This paper proposes TYMER, a unified yield-based performance model for timing-speculative SRAM. In TYMER, the first sub-model evaluates access-time yield at different worldline enable time for a general 6T SRAM under low supply voltages, while the second one uses the yield results to estimate optimal sensing time and the overall read latency for speculative SRAM. TYMER is not only compared with simulation results but also the measurements from 28nm fabricated speculative SRAM chips. Both cases show precise evaluation results under different operating conditions.
Shan Shen, Liang Pang 0002, Tianxiang Shao, Xiao Shi 0001, Longxing Shi
DAC6
2020 Modeling and Designing of a PVT Auto-tracking Timing-speculative SRAM
abstract
In the low supply voltage region, the performance of 6T cell SRAM degrades seriously, which takes more time to achieve the sufficient voltage difference on bitlines. Timing- speculative techniques are proposed to boost the SRAM frequency and the throughput with speculatively reading data in an aggressive timing and correcting timing failures in one or more extended cycles. However, the throughput gains of timing- speculative SRAM are affected by the process, voltage and temperature (PVT) variations, which causes the timing design of speculative SRAM to be either too aggressive or too conservative. This paper first proposes a statistical model to abstract the characteristics of speculative SRAM and shows the presence of an optimal sensing time that maximizes the overall throughput. Then, with the guidance of the performance model, a PVT auto-tracking speculative SRAM is designed and fabricated, which can dynamically self-tune the bitline sensing to the optimal time as the working condition changes. According to the measurement results, the maximum throughput gain of the proposed 28nm SRAM is 1.62X compared to the baseline at 0.6V VDD.
Shan Shen, Tianxiang Shao, Jun Yang 0006, Longxing Shi
DATE5
2020 A Non-Gaussian Adaptive Importance Sampling Method for High-Dimensional and Multi-Failure-Region Yield Analysis
abstract
Rare-event yield analysis is challenging for high-dimensional circuit cases. In this paper, we propose a non-Gaussian adaptive importance sampling (NGAIS) method. In order to approximate the failure region in high-dimensional space, we model it as a mixture of von Mises-Fisher distributions. We formulate the parameter estimation problem as a maximum likelihood estimation problem, and then solve with expectation-maximization algorithm. Experiments on bit cell, amplifier and SRAM column circuit validate that the proposed NGAIS method outperforms other state-of-the-art approaches in terms of accuracy and efficiency.
Xiao Shi 0001, Hao Yan 0002, Chuwen Li, Jianli Chen, Longxing Shi, Lei He 0001
ICCAD5
2020 An Efficient Adaptive Importance Sampling Method for SRAM and Analog Yield Analysis
abstract
Performance failure has become a major threat for various memory and analog circuits. It is challenging to estimate the extremely small failure probability when failed samples are distributed in multiple disjoint regions. In this article, we propose an adaptive importance sampling (AIS) algorithm. AIS has several iterations of sampling region adjustments, while existing methods predecide a static sampling distribution. We design two adaptive frameworks based on resampling and population Metropolis-Hastings (MH) to iteratively search for failure regions. The experimental results of the AIS method exhibit better efficiency and higher accuracy. For SRAM bit cell with single failure region, the AIS method uses 2-$27{\times }$ fewer samples and reaches better accuracy when compared to several recent methods. For a two-stage amplifier circuit with multiple failure regions, the AIS method is $90{\times }$ faster than Monte Carlo and 7-23 ${\times }$ over other methods. For charge pump circuit and $C^{2}MOS$ master-slave latch circuit, the AIS method can reach 6-$18{\times }$ and 4-$6{\times }$ speedup over other methods, respectively.
Xiao Shi 0001, Hao Yan 0002, Longxing Shi, Lei He 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 TS Cache: A Fast Cache With Timing-Speculation Mechanism Under Low Supply Voltages
abstract
To mitigate the ever-worsening “power wall” problem, more and more applications need to expand their working voltage to the wide-voltage range including the nearthreshold region. However, the read delay distribution of the static random access memory (SRAM) cells under the nearthreshold voltage shows a more serious long-tail characteristic than that under the nominal voltage due to the process fluctuation. Such degradation of SRAM delay makes the SRAM-based cache a performance bottleneck of systems as well. To avoid unreliable data reading, circuit-level studies use larger/more transistors in a bitcell by sacrificing chip area and the static power of cache arrays. Architectural studies propose the auxiliary error correction or block disabling/remapping methods in fault-tolerant caches, which worsen both the hit latency and energy efficiency due to the complex accessing logic. This article proposes a timing-speculation (TS) cache to boost the cache frequency and improve energy efficiency under low supply voltages. In the TS cache, the voltage differences of bitlines (BLs) are continuously evaluated twice by a sense amplifier (SA), and the access timing error can be detected much earlier than that in prior methods. According to the measurement results from the fabricated chips, the TS L1 cache aggressively increases its frequency to 1.62× and 1.92× compared with the conventional scheme at 0.5- and 0.6-V supply voltages, respectively.
Shan Shen, Tianxiang Shao, Xiaojing Shang, Jun Yang 0006, Longxing Shi
IEEE Trans. Very Large Scale Integr. Syst.7
2019 Meta-Model based High-Dimensional Yield Analysis using Low-Rank Tensor Approximation
abstract
"Curse of dimensionality" has become the major challenge for existing high-sigma yield analysis methods. In this paper, we develop a meta-model using Low-Rank Tensor Approximation (LRTA) to substitute expensive SPICE simulation. The polynomial degree of our LRTA model grows linearly with circuit dimension. This makes it especially promising for high-dimensional circuit problems. Our LRTA meta-model is solved efficiently with a robust greedy algorithm, and calibrated iteratively with an adaptive sampling method. Experiments on bit cell and SRAM column validate that proposed LRTA method outperforms other state-of-the-art approaches in terms of accuracy and efficiency.
Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Longxing Shi, Lei He 0001
DAC5
2019 A Statistical Current and Delay Model Based on Log-Skew-Normal Distribution for Low Voltage Region
abstract
The increasing performance variation and non-Gaussian distribution pose remarkable challenges to timing analysis for circuits operating in low voltage region. Accurate modeling of the statistical characteristics is urgently required with process variation consideration. In this paper, the statistical models for drain current and gate delay in low voltage region are established in analytical form based on the log-skew-normal (LSN) distribution via moment matching technique. Experimental results show that the probability distribution function (PDF) curves obtained from the proposed models for drain current and gate delay are highly fitted with Monte Carlo (MC) simulation results in sub/near-threshold regions. Moreover, owing to the proposed LSN-based statistical model, less than 8% error is introduced in the predicted sensitivity of gate delay and the maximum/minimum delay indicated by ±3σ percentile points can be calculated more precisely than the LN-based method with up to 3× accuracy improvement for low supply voltage.
Peng Cao 0002, Jiangping Wu, Zhiyuan Liu 0011, Jun Yang 0006, Longxing Shi
ACM Great Lakes Symposium on VLSI6
2019 A Statistical Timing Model for Low Voltage Design Considering Process Variation
abstract
Near-threshold voltage (NTV) design suffers severe challenge due to the dramatic increase in performance uncertainty introduced by process variation. This paper proposes an analytical approach based on Log-Normal (LN) distribution to characterize the statistical delay for NTV design considering the dominant threshold voltage variation from gate-level to circuit-level. At gate-level, the multivariable threshold voltage variation issue is solved by the equivalent threshold voltage method and equivalent drain current method for generic gates with stack topology and parallel topology, respectively. At circuit-level, a statistical timing model is proposed as the linear combination of the independent statistical delays of all gates in the path with step input by considering the varied correlation between adjacent gates. To the best of our knowledge, we firstly propose a statistical timing model analytically for practical circuit path with physical insights of supply voltage, transistor size, and load capacitance. The characterization effort for each path is only one-time SPICE simulation, which is negligible compared with Monte Carlo (MC) simulation in statistical static timing analysis (SSTA) methods. Experimental results under a commercial 28-nm CMOS process show the proposed models have high accuracy at low supply voltage compared with MC simulations, where the modeling errors for the mean and variance of gate delay can be limited within 1.54% and 11.2%, respectively. Moreover, as for the practical paths in ITC'99 benchmark, the maximum modeling errors of mean, variance, minimum delay, and maximum delay is less than 3.04%, 11.40%, 4.50%, and 2.87%, respectively.
Peng Cao 0002, Zhiyuan Liu 0011, Jiangping Wu, Jun Yang 0006, Longxing Shi
ICCAD6
2019 Efficient Yield Analysis for SRAM and Analog Circuits using Meta-Model based Importance Sampling Method
abstract
Performance failure has become the major threat to the robustness and reliability of various memory and analog circuits. It is challenging to accurately estimate the extremely small failure probability when failed samples are distributed in multiple disjoint failure regions. In this paper, we develop a novel meta-model based importance sampling (MIS) method. MIS utilizes Gaussian Process meta-model to construct quasi-optimal importance sampling distribution, and performs Markov Chain Monte Carlo (MCMC) simulation to generate new samples from the proposed distribution. By updating our global Importance Sampling estimator in an iterated framework, MIS leads to better efficiency and higher accuracy. For SRAM bit cell with single failure region, MIS uses 4-6X fewer samples and reaches better accuracy when compared to several recent methods. For a two-stage amplifier circuit with multiple failure schemes, MIS is 213X faster than MC without compromising accuracy, while other methods fail to cover all failure regions in our experiment.
Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Longxing Shi, Lei He 0001
ICCAD5
2019 Adaptive Clustering and Sampling for High-Dimensional and Multi-Failure-Region SRAM Yield Analysis
abstract
Statistical circuit simulation is exhibiting increasing importance for memory circuits under process variation. It is challenging to accurately estimate the extremely low failure probability as it becomes a high-dimensional and multi-failure-region problem. In this paper, we develop an Adaptive Clustering and Sampling (ACS) method. ACS proceeds iteratively to cluster samples and adjust sampling distribution, while most existing approaches pre-decide a static sampling distribution. By adaptively searching in multiple cone-shaped subspaces, ACS obtains better accuracy and efficiency. This result is validated by our experiments. For SRAM bit cell with single failure region, ACS requires 3-5X fewer samples and achieves better accuracy compared with existing approaches. For 576-dimensional SRAM column circuit with multiple failure regions, ACS is 2050X faster than MC without compromising accuracy, while other methods fail to converge to correct failure probability in our experiment.
Xiao Shi 0001, Hao Yan 0002, Xiaofen Xu, Longxing Shi, Lei He 0001
ISPD6
2019 An embedded implementation of CNN-based hand detection and orientation estimation algorithm
Zeheng Liu, Hao Liu 0013, Longxing Shi, Xinning Liu
Mach. Vis. Appl.6
2018 A Light CNN based Method for Hand Detection and Orientation Estimation
abstract
Hand detection is an essential step to support many tasks including HCI applications. However, detecting various hands robustly under conditions of cluttered backgrounds, motion blur or changing light is still a challenging problem. Recently, object detection methods using CNN models have significantly improved the accuracy of hand detection yet at a high computational expense. In this paper, we propose a light CNN network, which uses a modified MobileNet as the feature extractor in company with the SSD framework to achieve a robust and fast detection of hand location and orientation. The network generates a set of feature maps of various resolutions to detect hands of different sizes. In order to improve the robustness, we also employ a top-down feature fusion architecture that integrates context information across levels of features. For an accurate estimation of hand orientation by CNN, we manage to estimate two orthogonal vectors' projections along the horizontal and vertical axes then recover the size and orientation of a bounding box exactly enclosing the hand. Evaluated on the challenging Oxford hand dataset, our method reaches 83.2% average precision (AP) at 139 FPS on a Nvidia Titan X, outperforming the previous methods both in accuracy and efficiency.
Zeheng Liu, Yang Zhang 0132, Hao Liu 0013, Jianhui Wu 0001, Longxing Shi
ICPR8
2018 Detecting the phase behavior on cache performance using the reuse distance vectors
Shan Shen, Longxing Shi
J. Syst. Archit.4
2018 An Analytical Cache Performance Evaluation Framework for Embedded Out-of-Order Processors Using Software Characteristics
abstract
Utilizing analytical models to evaluate proposals or provide guidance in high-level architecture decisions is been becoming more and more attractive. A certain number of methods have emerged regarding cache behaviors and quantified insights in the last decade, such as the stack distance theory and the memory level parallelism (MLP) estimations. However, prior research normally oversimplified the factors that need to be considered in out-of-order processors, such as the effects triggered by reordered memory instructions, and multiple dependences among memory instructions, along with the merged accesses in the same MSHR entry. These ignored influences actually result in low and unstable precisions of recent analytical models. By quantifying the aforementioned effects, this article proposes a cache performance evaluation framework equipped with three analytical models, which can more accurately predict cache misses, MLPs, and the average cache miss service time, respectively. Similar to prior studies, these analytical models are all fed with profiled software characteristics in which case the architecture evaluation process can be accelerated significantly when compared with cycle-accurate simulations. We evaluate the accuracy of proposed models compared with gem5 cycle-accurate simulations with 16 benchmarks chosen from Mobybench Suite 2.0, Mibench 1.0, and Mediabench II. The average root mean square errors for predicting cache misses, MLPs, and the average cache miss service time are around 4%, 5%, and 8%, respectively. Meanwhile, the average error of predicting the stall time due to cache misses by our framework is as low as 8%. The whole cache performance estimation can be sped by about 15 times versus gem5 cycle-accurate simulations and 4 times when compared with recent studies. Furthermore, we have shown and studied the insights between different performance metrics and the reorder buffer sizes by using our models. As an application case of the framework, we also demonstrate how to use our framework combined with McPAT to find out Pareto optimal configurations for cache design space explorations.
Kecheng Ji, Longxing Shi, Jianping Pan 0001
ACM Trans. Embed. Comput. Syst.3
2017 AFEC: An analytical framework for evaluating cache performance in out-of-order processors
abstract
Evaluating cache performance is becoming critically important to predict the overall performance of out-of-order processors. Non-blocking caches, which are very common in out-of-order CPUs, can reduce the average cache miss penalty by overlapping multiple outstanding memory requests and merging different cache misses with the same cacheline address into one memory request. Normally, memory-level-parallelism (MLP) has been used as a metric to describe the concurrency of memory access. Unfortunately, due to the extremely dynamic dependences among the program memory references, it is very difficult to quantify MLP without time-consuming simulations. Moreover, the merging of multiple cache misses, which makes the average cache miss service time less than the physical DDR access latency, is seldom considered in the existing researches. In this paper, we propose a cache performance evaluation framework based on program trace analysis and analytical models to fast estimate MLP and the effective cache miss service time without simulations. Comparing with the results by Gem5 simulations of MobyBench 2.0, Mibench 1.0 and Mediabench II, the average accuracy of the modeled MLP and the average cache miss service time is higher than 91% and 92%, respectively. Combined with cache misses calculated by the stack distance theory, the average absolute error of CPU stall time (due to cache misses) is lower than 10%, while the evaluation time can be sped up by 35 times relative to the Gem5 full simulations.
Kecheng Ji, Longxing Shi, Jianping Pan 0001
DATE4
2017 Context Management Scheme Optimization of Coarse-Grained Reconfigurable Architecture for Multimedia Applications
abstract
Due to the combination of flexibility and efficiency, coarse-grained reconfigurable architectures (CGRAs) are suitable for the implementation of computing-intensive applications. However, with the growing performance requirements, the scale of CGRA increases exponentially, which leads to configuration performance degradation and configuration power rise. Based on the analysis of configuration context features, we optimize the context management scheme of CGRA from the aspects of context cache structure and replacement strategy. The context cache is structured hierarchically to reduce the memory overhead without configuration performance degradation and a hybrid context replacement algorithm is proposed to further increase the configuration efficiency with a novel context frequency weight factor. Experimental results show that the proposed context management scheme improves the configuration performance of the base CGRA significantly by 13.6%-20.5% for H.264 decoding and 13.6%-20.5% for MPEG2 decoding with only 43% context cache cost. Compared with other works, the proposed context management scheme shows the advantages of 2.3-6× less normalized context cache size and 2.3-2.7× cache efficiency.
Peng Cao 0002, Bo Liu 0019, Jinjiang Yang, Jun Yang 0006, Meng Zhang 0010, Longxing Shi
IEEE Trans. Very Large Scale Integr. Syst.6
2014 A Side-channel Analysis Resistant Reconfigurable Cryptographic Coprocessor Supporting Multiple Block Cipher Algorithms
abstract
A side-channel analysis resistant reconfigurable cryptographic coprocessor is designed and fabricated in 0.18μm CMOS with 1.8V supply and 100MHz frequency, supporting multiple block cipher algorithms of AES, DES, RC6 and IDEA. Our countermeasure utilizes idle processing elements existed in reconfigurable array to do dummy operations to hide leakage information. This method has little impact on area and frequency, and it is flexible after silicon. It resists SPA and DPA without distinguishing the encryption region. And by correlation-based electromagnetic analysis, measurement to disclosure of DES enhances 36 times with partial countermeasures and AES discloses no subkey after more than one million electromagnetic traces with full countermeasures.
Weiwei Shan, Longxing Shi, Xingyuan Fu, Chaoxuan Tian, Jun Yang 0006, Jie Li 0057
DAC2
2014 An energy efficient OpenCL implementation of a fingerprint verification system on heterogeneous mobile device
abstract
With the increasing concerns over the personal privacy of mobile devices, biometrics algorithms plays an important role to enhance the security. As one of the most popular approaches, fingerprint verification as a personal identification interface is widely recognized and adopted by many commercial devices. However, its inherent computational complexity make the algorithm of fingerprint verification difficult to achieve high performance on mobile platforms, such a battery powered, size limited, and producing cost controlled device. In addition to the performance, energy efficiency is also of significant consideration of such a fingerprint verification system. In this paper, we present an energy efficient OpenCL based heterogeneous implementation of the fingerprint verification system on a commercial mobile platform, taking advantage of mobile CPUs and GPUs. We carefully analyze the workloads through system profiling to identify the parallelism then to partition the algorithm between the CPU and GPU . The experimental results show that our GPU implementation of DFT analysis achieves a 1.4X speedup and 36.87% energy reduction compared to the CPU only implementation in the mobile platform. This heterogenous implementation of the entire fingerprint verification system accomplishes 1.32X speedup and 16.70% energy superiority above the CPU only solution. To the best of the authors' knowledge, this work is the first published implementation of OpenCL based fingerprint verification system accelerated by mobile GPUs on a heterogeneous mobile device. We believe our mapping methodology of this fingerprint verification system can be generalized to map more similar applications onto heterogeneous mobile devices.
Wen Wen 0003, Longxing Shi
RTCSA5
2012 All-Digital Wide Range Precharge Logic 50% Duty Cycle Corrector
abstract
A novel all-digital 50% duty cycle corrector (DCC) is pro- posed in this paper. The DCC features include a delay unit based on precharge logic gates with low delay time and a robust SR latch under process voltage and temperature variations for final edge combination over wide frequency and duty-cycle ranges. The rising edge of the output clock has a constant delay when comparing to the input clock, which makes it easy to cooperate with a delay locked loop. The circuit is fabricated in Chartered 0.18-μm CMOS process. The acceptable input clock frequency ranges from 400 MHz to 2 GHz. The correcting error is ±3.5% at 1 GHz or ±1% at 400 MHz.
Junhui Gu, Jianhui Wu 0001, Danhong Gu, Meng Zhang 0010, Longxing Shi
IEEE Trans. Very Large Scale Integr. Syst.5
2011 Traffic-aware game based channel assignment in wireless sensor networks
abstract
Game theory has been shown to be a promising way to solve the multi-channel assignment problem in wireless sensor networks. However, without considering different traffic carried by nodes, the state-of-the-art game based channel assignment algorithm cannot reduce interference effectively. In this paper, we analyze interference from the network's global perspective and propose a new interference metric considering different traffic carried by nodes. Then we propose a distributed channel assignment algorithm that jointly exploits topology, routing and traffic information to minimize the total interference of the network. Simulation results show that the proposed algorithm achieves significant performance improvement in throughput, delivery ratio and energy efficiency.
Fulong Jiang, Hao Liu 0013, Longxing Shi
IWCMC3
2011 A Fast Locking All-Digital Phase-Locked Loop via Feed-Forward Compensation Technique
abstract
A fast locking all-digital phase-locked loop (ADPLL) via feed-forward compensation technique is proposed in this paper. The implemented ADPLL has two operation modes which are frequency acquisition mode and phase acquisition mode. In frequency acquisition mode, the ADPLL achieves a fast frequency locking via the proposed feed-forward compensation algorithm. In phase acquisition mode, the ADPLL achieves a finer phase locking. To verify the proposed algorithm and architecture, the ADPLL design is implemented by SMIC 0.18-μm 1P6M CMOS technology. The core size of the ADPLL is 582.2 μm * 343 μm. The frequency range of the ADPLL is from 4 to 416 MHz. The measurement results show that the ADPLL can achieve a frequency locking in two reference cycles when locking to 376 MHz. The corresponding power consumption is 11.394 mW.
Xin Chen 0039, Jun Yang 0006, Longxing Shi
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Area-efficient line-based two-dimensional discrete wavelet transform architecture without data buffer
abstract
An area-efficient architecture for 2D DWT is proposed in this paper based on novel decomposed lifting scheme, where no data buffer is required to preserve and reorder the intermediate data between the row and column processor. Compared with the reported research, the proposed design could benefit from the reduction of internal memory size and the number of multipliers, adders and registers. The design was implemented for 2D 9/7 and 5/3 DWT in SMIC 0.18 mum CMOS logic fabrication with 15 K equivalent 2-input NAND gates under 150 MHz, which can accommodate up to 512times512 image size with 4 K bytes on-chip dual-port RAM.
Peng Cao 0002, Chao Wang 0068, Jun Yang 0006, Longxing Shi
ICME4
2006 Energy-optimal dynamic voltage scaling for sporadic tasks
abstract
Reducing power consumption is a challenge to system designers. Dynamic power management (DPM) and dynamic voltage scaling (DVS) have emerged as efficient solutions to reduce the power consumption of embedded and portable systems. This paper proposes an online dynamic voltage scaling policy expressed as energy-optimal sporadic task scheduling (EOSTS) for sporadic tasks. EOSTS differs from other existing DVS policies in the fact that the former essentially consists of an energy-optimal model, based on strict theorems proving. Four theorems are proposed in this paper altogether, which is helpful to power management policy optimization research. Experimental results demonstrate that EOSTS policy is superior in power saving to general DPM policies
Bu Aiguo, Longxing Shi, Li Jie, Chao Wang 0068
ISCAS2