EDBT 2026 Demo / reviewers in the wild / expert
Xiaolang Yan
dblp:40/4081
· DBLP profile ↗
57ranked-venue papers
2as first author
4since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5Software engineering, systems software and programming languages · 3Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Structured Term Pruning for Computational Efficient Neural Networks InferenceabstractThe state-of-the-art convolutional neural network accelerators are showing a growing interest in exploiting the bit-level sparsity and eliminating the ineffectual computations of zero bits. However, the excessive redundancy and the irregular distribution of nonzero bits limit the real speedup in the accelerators. To address this, we propose an algorithm-architecture codesign, named structured term pruning (STP), to boost the computation efficiency of neural networks inference. Specifically, we enhance the bit sparsity by guiding the weights toward the value with fewer power-of-two terms. Then, we structure the terms with layer-wise group budgets. Retraining is adopted to recover the accuracy drop. We also design the hardware of the group processing element and the fast signed-digital encoder for efficient implementation of STP networks. The system design of STP is realized with some easy alterations on an input stationary systolic array design. Extensive evaluation results demonstrate that STP can reduce significant inference computation costs, and achieve$2.35\times $computational energy saving for the ResNet18 network on the ImageNet dataset. Kai Huang 0002, Bowen Li 0017, Siang Chen, Luc Claesen, Wei Xi 0001, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong, Xiaolang Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2023 | Structured Dynamic Precision for Deep Neural Networks QuantizationabstractDeep Neural Networks (DNNs) have achieved remarkable success in various Artificial Intelligence applications. Quantization is a critical step in DNNs compression and acceleration for deployment. To further boost DNN execution efficiency, many works explore to leverage the input-dependent redundancy with dynamic quantization for different regions. However, the sensitive regions in the feature map are irregularly distributed, which restricts the real speed up for existing accelerators. To this end, we propose an algorithm-architecture co-design, named Structured Dynamic Precision (SDP). Specifically, we propose a quantization scheme in which the high-order bit part and the low-order bit part of data can be masked independently. And a fixed number of term parts are dynamically selected for computation based on the importance of each term in the group. We also present a hardware design to enable the algorithm efficiently with small overheads, whose inference time mainly scales with the precision proportionally. Evaluation experiments on extensive networks demonstrate that compared to the state-of-the-art dynamic quantization accelerator DRQ, our SDP can achieve 29% performance gain and 51% energy reduction for the same level of model accuracy. Kai Huang 0002, Bowen Li 0017, Dongliang Xiong, Haitian Jiang, Xiaowen Jiang 0001, Xiaolang Yan, Luc Claesen, Dehong Liu, Junjian Chen, Zhili Liu |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2022 | Predicting the Output Structure of Sparse Matrix Multiplication with Sampled Compression RatioabstractSparse general matrix multiplication (SpGEMM) is a fundamental building block in numerous scientific applications. One critical task of SpGEMM is to compute or predict the structure of the output matrix (i.e., the number of nonzero elements per output row) for efficient memory allocation and load balance, which impact the overall performance of SpGEMM. Existing work either precisely calculates the output structure or adopts upper-bound or sampling-based methods to predict the output structure. However, these methods either take much execution time or are not accurate enough. In this paper, we propose a novel sampling-based method with better accuracy and low costs compared to the existing sampling-based method. The proposed method first predicts the compression ratio of SpGEMM by leveraging the number of intermediate products (denoted as FLOP) and the number of nonzero elements (denoted as NNZ) of the same sampled result matrix. And then, the predicted output structure is obtained by dividing the FLOP per output row by the predicted compression ratio. We also propose a reference design of the existing sampling-based method with optimized computing overheads to demonstrate the better accuracy of the proposed method. We construct 623 test cases with various matrix dimensions and sparse structures to evaluate the prediction accuracy. Experimental results show that the absolute relative errors of the proposed method and the reference design are 1.30% and 7.93%, respectively, on average, and 25% and 158%, respectively, in the worst case. Zhaoyang Du, Yijin Guan, Tianchan Guan, Dimin Niu, Nianxiong Tan, Xiaopeng Yu 0002, Hongzhong Zheng, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001 |
ICPADS | 9 |
| 2021 | Expected Energy Optimization for Real-Time Multiprocessor SoCs Running Periodic Tasks with Uncertain Execution TimeabstractEnergy optimization plays an increasingly critical role in designing an embedded real-time multiprocessor System on Chip (MPSoC). Dynamic Voltage Frequency Scaling (DVFS) and Dynamic Power Management (DPM) are preferable techniques to optimize energy consumption. However, previous DVFS and DPM algorithms were mostly designed for inter-task scheduling, without sufficient exploration on intra-task scheduling for further energy reduction. This paper presents a new intra-task scheduling approach considering the probabilistic distribution of task execution time, and it optimizes the mathematical expectation of power consumption (expected power consumption) for periodic dependent tasks with uncertain execution time running on MPSoCs using DVFS and DPM. The energy-efficient scheduling problem can be formulated by means of mixed integer linear programming (MILP) with the proposed technique. Moreover, we also propose a technique to compress the exploration space by reorganizing the probabilistic profiling information of all tasks. Our experimental results on synthetic and realistic benchmarks show that the proposed approach achieves up to 30 percent energy savings compared with other existing methods. Kai Huang 0002, Ke Wang 0034, Dandan Zheng 0001, Xiaowen Jiang 0001, Rongjie Yan, Xiaolang Yan |
IEEE Trans. Sustain. Comput. | 7 |
| 2020 | Xuantie-910: Innovating Cloud and Edge Computing by RISC-VabstractThis article consists only of a collection of slides from the author's conference presentation. Chen Chen 0058, Xiaoyan Xiang, Chang Liu 0021, Yunhai Shang, Ren Guo, Dongqi Liu 0003, Ziyi Hao, Chunqiang Li, Yu Pu, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001, Xiaoning Qi |
Hot Chips Symposium | 14 |
| 2020 | Xuantie-910: A Commercial Multi-Core 12-Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension : Industrial ProductabstractThe open source RISC-V ISA has been quickly gaining momentum. This paper presents Xuantie-910, an industry leading 64-bit high performance embedded RISC-V processor from Alibaba T-Head division. It is fully based on the RV64GCV instruction set and it features custom extensions to arithmetic operation, bit manipulation, load and store, TLB and cache operations. It also implements the 0.7.1 stable release of RISCV vector extension specification for high efficiency vector processing. Xuantie-910 supports multi-core multi-cluster SMP with cache coherence. Each cluster contains 1 to 4 core(s) capable of booting the Linux operating system. Each single core utilizes the state-of-the-art 12-stage deep pipeline, out-of-order, multi-issue superscalar architecture, achieving a maximum clock frequency of 2.5 GHz in the typical process, voltage and temperature condition in a TSMC 12nm FinFET process technology. Each single core with the vector execution unit costs an area of 0.8 mm2, (excluding the L2 cache). The toolchain is enhanced significantly to support the vector extension and custom extensions. Through hardware and toolchain co-optimization, to date Xuantie-910 delivers the highest performance (in terms of IPC, speed, and power efficiency) for a number of industrial control flow and data computing benchmarks, when compared with its predecessors in the RISC-V family. Xuantie-910 FPGA implementation has been deployed in the data centers of Alibaba Cloud, for applicationspecific acceleration (e.g., blockchain transaction). The ASIC deployment at low-cost SoC applications, such as IoT endpoints and edge computing, is planned to facilitate Alibaba’s end-to-end and cloud-to-edge computing infrastructure. Chen Chen 0058, Xiaoyan Xiang, Chang Liu 0021, Yunhai Shang, Ren Guo, Dongqi Liu 0003, Ziyi Hao, Chunqiang Li, Yu Pu, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001, Xiaoning Qi |
ISCA | 14 |
| 2019 | A Scalable and Adaptable ILP-Based Approach for Task Mapping on MPSoC Considering Load Balance and Communication OptimizationabstractTask mapping has been a hot topic in multiprocessor system-on-chip software design for decades. During the mapping process, load balance (LB) and communication optimization have been two important performance optimization factors. This paper studies the relations between LB, interprocessor communications, and communication pipeline technique during the mapping process, and proposes an integer linear programming (ILP)-based static task mapping approach, which considers both LB and communication optimization. The approach consists of an optimized ILP model for task mapping with fewer variables compared to previous ILP mapping works. Moreover, to enhance the scalability of the ILP task mapping, the task-processor-cluster algorithm is proposed to reduce the scale of the task graph and the number of processors and then solve the coarse-grained input by the ILP mapping. To increase the adaptability of the ILP task mapping, the improved augmented E-constraint method is further integrated with the ILP formulations to select the best mapping for different applications. Experimental results on a 2/4/8/16/24-CPU platform of both synthetic and real-life benchmarks demonstrate the efficiency of the proposed approach. Kai Huang 0002, Dandan Zheng 0001, Min Yu 0006, Xiaowen Jiang 0001, Xiaolang Yan, Lisane B. de Brisolara, Ahmed Amine Jerraya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Providing Predictable Performance via a Slowdown Estimation ModelabstractInterapplication interference at shared main memory slows down different applications differently. A few slowdown estimation models have been proposed to provide predictable performance by quantifying memory interference, but they have relatively low accuracy. Thus, we propose a more accurate slowdown estimation model called SEM at main memory. First, SEM unifies the slowdown estimation model by measuring IPC directly. Second, SEM uses the per-bank structure to monitor memory interference and improves estimation accuracy by considering write interference, row-buffer interference, and data bus interference. The evaluation results show that SEM has significantly lower slowdown estimation error (4.06%) compared to STFM (30.15%) and MISE (10.1%). Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan |
ACM Trans. Archit. Code Optim. | 4 |
| 2017 | Eliminating Timing Errors Through Collaborative Design to Maximize the ThroughputabstractIn advanced technology nodes, large timing margins must be added to allow for worse process, voltage, temperature, and aging variations. The error detection and correction (EDAC) technique effectively eliminates these margins by timing speculation, but the high design complexity and large hardware cost make many existing EDAC systems unsuitable for commercial processors. Based on the instruction-level locality of timing errors, a collaborative EDAC approach is proposed to address this issue. The hardware layer adopts simple and low cost EDAC circuits to ensure correct operation when timing error occurs, while a runtime software layer prevents recurring errors of the same instruction by sending timing error alarms to the hardware layer. Cooperation of both layers, accompanied with the proposed profile-guided timing error avoidance algorithm, eliminates more than 95% of errors with small runtime overhead. This significantly improves overall performance and alleviates pressure on the EDAC circuits. Experimental results based on the three-stage commercial CK802 processor in SMIC 40LL process present that the approach has improved the peak performance of the baseline EDAC system (Razor-Lite + half-frequency replay) by 8% and reduced the energy consumption by 25%, with less than 1.4% area overhead. Zhan-Hui Li, Tao-Tao Zhu, Zhi-Jian Chen, Jian-Yi Meng, Xiaoyan Xiang, Xiaolang Yan |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2017 | Error-Resilient Integrated Clock Gate for Clock-Tree Power Optimization on a Wide Voltage IOT ProcessorabstractEnergy-efficiency optimization occupies an important position in the Internet of Things application. The error-resilience technique has begun to emerge and brought the performance and energy benefits as a new vision for alternative computing, because it eliminates the overconstrained margin in current processor design flow and protects the system from process, supply voltage, temperature, and aging variations through an error-resilient mechanism rather than expensive guardbands. However, as a traditional clock-tree power optimization technique, the clock gating mechanism cannot work in such a system when it faces the timing violation problem. In this paper, we propose an error-resilient integrated clock gate (ERICG) and its automatic integration methodology in error detection and correction (EDAC) system design flow. ERICG can provide the ability of in situ timing EDAC with only four additional transistors compared with a conventional integrated clock gate. The SPICE simulation shows that it is a metastable-hardened cell and can work well in the wide voltage operation (0.5~ 1.1 V) including the near-threshold region. We implement it in a commercial C-SKY CK802 processor based on an SMIC 40-nm technology. The result shows that it improves the energy efficiency by 68% compared with the non-EDAC design and lowers the total power by 28.72% over the conventional EDAC design at 0.6 V. Tao-Tao Zhu, Jian-Yi Meng, Xiaoyan Xiang, Xiaolang Yan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Memory Access Scheduling Based on Dynamic Multilevel Priority in Shared DRAM SystemsabstractInterapplication interference at shared main memory severely degrades performance and increasing DRAM frequency calls for simple memory schedulers. Previous memory schedulers employ a per-application ranking scheme for high system performance or a per-group ranking scheme for low hardware cost, but few provide a balance. We propose DMPS, a memory scheduler based on dynamic multilevel priority. First, DMPS uses “memory occupancy” to measure interference quantitatively. Second, DMPS groups applications, favors latency-sensitive groups, and dynamically prioritizes applications by employing a per-level ranking scheme. The simulation results show that DMPS has 7.2% better system performance and 22% better fairness over FRFCFS at low hardware complexity and cost. Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan |
ACM Trans. Archit. Code Optim. | 4 |
| 2015 | Fast Level-Set-Based Inverse Lithography Algorithm for Process Robustness Improvement and Its Application
Zhen Geng, Zheng Shi 0002, Xiaolang Yan, Kai-sheng Luo |
J. Comput. Sci. Technol. | 3 |
| 2015 | Profiling and annotation combined method for multimedia application specific MPSoC performance estimationabstractAccurate and fast performance estimation is necessary to drive design space exploration and thus support important design decisions. Current techniques are either time consuming or not accurate enough. In this paper, we solve these problems by presenting a hybrid method for multimedia multiprocessor system-on-chip (MPSoC) performance estimation. A general coverage analysis tool GNU gcov is employed to profile the execution statistics during the native simulation. To tackle the complexity and keep the analysis and simulation manageable, the orthogonalization of communication and computation parts is adopted. The estimation result of the computation part is annotated to a transaction accurate model for further analysis, by which a gradual refinement of MPSoC performance estimation is supported. The implementation and its experimental results prove the feasibility and efficiency of the proposed method. Kai Huang 0002, Siwen Xiu, Dandan Zheng 0001, Min Yu 0006, De Ma, Kai Huang 0001, Gang Chen 0023, Xiaolang Yan |
Frontiers Inf. Technol. Electron. Eng. | 9 |
| 2015 | Communication Optimizations for Multithreaded Code Generation from Simulink ModelsabstractCommunication frequency is increasing with the growing complexity of emerging embedded applications and the number of processors in the implemented multiprocessor SoC architectures. In this article, we consider the issue of communication cost reduction during multithreaded code generation from partitioned Simulink models to help designers in code optimization to improve system performance. We first propose a technique combining message aggregation and communication pipeline methods, which groups communications with the same destinations and sources and parallelizes communication and computation tasks. We also present a method to apply static analysis and dynamic emulation for efficient communication buffer allocation to further reduce synchronization cost and increase processor utilization. The existing cyclic dependency in the mapped model may hinder the effectiveness of the two techniques. We further propose a set of optimizations involving repartition with strongly connected threads to maximize the degree of communication reduction and preprocessing strategies with available delays in the model to reduce the number of communication channels that cannot be optimized. Experimental results demonstrate the advantages of the proposed optimizations with 11--143% throughput improvement. Kai Huang 0002, Min Yu 0006, Rongjie Yan, Xiaolang Yan, Lisane B. de Brisolara, Ahmed Amine Jerraya, Jiong Feng |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2014 | Analysis and evaluation of per-flow delay bound for multiplexing modelsabstractMultiplexing models are common in resource sharing communication media such as buses, crossbars and networks. While sending packets over a multiplexing node, the packet delay bound can be computed using network calculus models. The tightness of such delay bound remains an open problem. This paper studies the multiplexing models for weighted round robin scheduling with different traffic arrival curves, and analyzes per-flow packet delay bounds with different service properties. We empirically evaluate the tightness of the delay bounds. Our results show the quality of different analysis models, and how influential each parameter is to tightness. Yanchen Long, Zhonghai Lu, Xiaolang Yan |
DATE | 3 |
| 2014 | SoC processor for real-time object labeling in life camera streams with low line level latencyabstractImage recognition systems implement a number of processing stages: preprocessing, segmentation and classification. In camera based video processing chains, usually several frame delays are incurred between the moment of capture and the actual availability of the classification results. Hardware architectures for stream based video processing have already been widely employed. In this paper, a new hardware architecture for accelerating the generic task of connected component analysis and object labeling in the segmentation step is presented. The architecture is specifically optimized for very low latency between image component capture by a camera and the detection in hardware. This latency constitutes only a few delay lines, thereby shortening the response time by a few orders of magnitude in comparison to traditional frame-buffer based methods. Zhengqiang Yu, Luc Claesen, Andy Motten, Yimu Wang, Xiaolang Yan |
ISCAS | 6 |
| 2014 | Perceptual image quality assessment metric using mutual information of Gabor features
Yong Ding 0003, Xiaolang Yan, Andrey S. Krylov |
Sci. China Inf. Sci. | 4 |
| 2014 | SVM based layout retargeting for fast and regularized inverse lithographyabstractInverse lithography technology (ILT), also known as pixel-based optical proximity correction (PB-OPC), has shown promising capability in pushing the current 193 nm lithography to its limit. By treating the mask optimization process as an inverse problem in lithography, ILT provides a more complete exploration of the solution space and better pattern fidelity than the traditional edge-based OPC. However, the existing methods of ILT are extremely time-consuming due to the slow convergence of the optimization process. To address this issue, in this paper we propose a support vector machine (SVM) based layout retargeting method for ILT, which is designed to generate a good initial input mask for the optimization process and promote the convergence speed. Supervised by optimized masks of training layouts generated by conventional ILT, SVM models are learned and used to predict the initial pixel values in the ‘undefined areas’ of the new layout. By this process, an initial input mask close to the final optimized mask of the new layout is generated, which reduces iterations needed in the following optimization process. Manufacturability is another critical issue in ILT; however, the mask generated by our layout retargeting method is quite irregular due to the prediction inaccuracy of the SVM models. To compensate for this drawback, a spatial filter is employed to regularize the retargeted mask for complexity reduction. We implemented our layout retargeting method with a regularized level-set based ILT (LSB-ILT) algorithm under partially coherent illumination conditions. Experimental results show that with an initial input mask generated by our layout retargeting method, the number of iterations needed in the optimization process and runtime of the whole process in ILT are reduced by 70.8% and 69.0%, respectively. Kai-sheng Luo, Zheng Shi 0002, Xiaolang Yan, Zhen Geng |
J. Zhejiang Univ. Sci. C | 3 |
| 2013 | A New Level-Set-Based Inverse Lithography Algorithm for Process Robustness Improvement with Attenuated Phase Shift MaskabstractInverse lithography technology (ILT) is one of the promising resolution enhancement techniques (RET), as the advanced integrated circuits (IC) technology nodes still use the 193nm light source. Among all the algorithms for ILT, the level-set-based ILT (LSB-ILT) is a feasible choice with good production result in practice. However, existing ILT algorithms optimize mask at nominal process condition without giving sufficient attention to the process variations, and thus the optimized masks show poor performance with focus and dose variations. In this paper, we put forward a new LSB-ILT algorithm for process robustness improvement with attenuated Phase Shift Mask (att-PSM) which is extensively used in the semiconductor foundries. In order to account for the process variations in the optimization, we adopt a new form of the cost function by adding the objective function of process variation band (PV band) to the nominal cost. The test patterns are from the M1 layer of a 28nm layout. Experimental results show that our new algorithm has a larger process window (PW) and reduces the process manufacturability index (PMI) by 41.37% compared with the LSB-ILT algorithm without PV band consideration. Zhen Geng, Zheng Shi 0002, Xiaolang Yan, Kai-sheng Luo |
CAD/Graphics | 3 |
| 2013 | Particle state compression scheme for centralized memory-efficient particle filtersabstractIn this paper, particle state compression scheme is proposed together with its architecture for centralized implementation of particle filters. In the scheme, state values are processed in original bit-width, stored in a compressed way and recovered before sampling of next iteration. The advantage of the scheme is that particle states memory requirement can be greatly reduced while the trade-off is the deviations between original and recovered states introduced by the process. A case study in Nearly Constant Turn (NCT) scenario shows that while achieving the same level of filtering accuracy, proposed scheme can save up to 49.69% memory overhead for storing particle state values compared to traditional realizations. Qinglin Tian, Xiaolang Yan, Ruohong Huan |
ICASSP | 3 |
| 2013 | Novel serpentine structure design method considering confidence level and estimation precisionabstractDue to the importance of metal layers in the product yield, serpentine test structures are usually fabricated on test chips to extract parameters for yield prediction. In this paper, the confidence level and estimation precision of the average defect density on metal layers are investigated to minimize the randomness of experimental results and make the measured parameters more convincing. On the basis of the Poisson yield model, the method to determine the total area of all serpentine test structures is obtained using the law of large numbers and the Lindeberg-Levy theorem. Furthermore, the method to determine an adequate area of each serpentine test structure is proposed under a specific requirement of confidence level and estimation precision. The results of Monte Carlo simulation show that the proposed method is consistent with theoretical analyses. It is also revealed by wafer experimental results that the method of designing serpentine test structure proposed in this paper has better performance. Li-sheng Chen, Xiao-hua Luo, Jiao-jiao Zhu, Fan-chao Jie, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 5 |
| 2013 | Regularized level-set-based inverse lithography algorithm for IC mask synthesisabstractInverse lithography technology (ILT) is one of the promising resolution enhancement techniques, as the advanced IC technology nodes still use the 193 nm light source. In ILT, optical proximity correction (OPC) is treated as an inverse imaging problem to find the optimal solution using a set of mathematical approaches. Among all the algorithms for ILT, the level-set-based ILT (LSB-ILT) is a feasible choice with good production in practice. However, the manufacturability of the optimized mask is one of the critical issues in ILT; that is, the topology of its result is usually too complicated to manufacture. We put forward a new algorithm with high pattern fidelity called regularized LSB-ILT implemented in partially coherent illumination (PCI), which has the advantage of reducing mask complexity by suppressing the isolated irregular holes and protrusions in the edges generated in the optimization process. A new regularization term named the Laplacian term is also proposed in the regularized LSB-ILT optimization process to further reduce mask complexity in contrast with the total variation (TV) term. Experimental results show that the new algorithm with the Laplacian term can reduce the complexity of mask by over 40% compared with the ordinary LSB-ILT. Zhen Geng, Zheng Shi 0002, Xiaolang Yan, Kai-sheng Luo |
J. Zhejiang Univ. Sci. C | 3 |
| 2013 | High throughput VLSI architecture for H.264/AVC context-based adaptive binary arithmetic coding (CABAC) decodingabstractContext-based adaptive binary arithmetic coding (CABAC) is the major entropy-coding algorithm employed in H.264/AVC. In this paper, we present a new VLSI architecture design for an H.264/AVC CABAC decoder, which optimizes both decode decision and decode bypass engines for high throughput, and improves context model allocation for efficient external memory access. Based on the fact that the most possible symbol (MPS) branch is much simpler than the least possible symbol (LPS) branch, a newly organized decode decision engine consisting of two serially concatenated MPS branches and one LPS branch is proposed to achieve better parallelism at lower timing path cost. A look-ahead context index (ctxIdx) calculation mechanism is designed to provide the context model for the second MPS branch. A head-zero detector is proposed to improve the performance of the decode bypass engine according to UEG k encoding features. In addition, to lower the frequency of memory access, we reorganize the context models in external memory and use three circular buffers to cache the context models, neighboring information, and bit stream, respectively. A pre-fetching mechanism with a prediction scheme is adopted to load the corresponding content to a circular buffer to hide external memory latency. Experimental results show that our design can operate at 250 MHz with a 20.71k gate count in SMIC18 silicon technology, and that it achieves an average data decoding rate of 1.5 bins/cycle. Kai Huang 0002, De Ma, Rongjie Yan, Haitong Ge, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 5 |
| 2013 | Performance Estimation Techniques With MPSoC Transaction-Accurate ModelsabstractEfficient design of multiprocessor system-on-chip (MPSoC) requires early, fast, and accurate performance estimation techniques. In this paper, we present new techniques based on fine-grained code analysis to estimate accurate performance during simulation of MPSoC transaction accurate models. First, a GCC profiling tool is applied in the native simulation process. Based on the profiling result, an instruction analyzer of the target CPU architecture is proposed to analyze the cycle cost of C code under estimation. In addition, a memory analyzer is used to further estimate memory access latency including both instruction/data cache time cost and global memory access cycles. Both data and instruction cache models are proposed to estimate cache miss penalty, and a segment-based strategy is adopted to update the cache models more efficiently. Furthermore, an equalized access model is presented to imitate the memory access behavior of processors for estimating global memory access latency caused by bus contention and memory bandwidth. We have applied these techniques on an H.264 decoder application with different hardware architectures. The experimental results show that applying these techniques can obviously improve estimation accuracy of transaction accurate models close to that of the virtual prototype models, with a tolerable overhead on simulation speed. De Ma, Rongjie Yan, Kai Huang 0002, Min Yu 0006, Siwen Xiu, Haitong Ge, Xiaolang Yan, Ahmed Amine Jerraya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2012 | Hierarchical resampling architecture for distributed particle filtersabstractIn this paper, a hierarchical resampling (HR) architecture has been presented for distributed particle filters (PFs). The proposed architectures decomposes the resampling step into two hierarchies, of which the first one, called intermediate resampling, is conducted consecutively among processing elements (PEs) the moment new particles and their weights are generated by each PE, and the second one, named unitary resampling, is performed sequentially after the whole intermediate resampling procedure and shared by all PEs. Compared with traditional distributed architectures, the HR architecture eliminates the particle redistribution step, and has such advantages as short execution time, high memory efficiency and well scalability. Xiaolang Yan, Ruohong Huan |
ICASSP | 3 |
| 2012 | Weight sorting based scheme and architecture for distributed particle filtersabstractThis paper presents an efficient weight sorting based scheme and architecture for distributed particle filters (PFs). Instead of redistributing particles among processing elements (PEs) after the resampling step, the proposed scheme sorts the newly generated weights, along with the corresponding propagated particles, in reverse order of the current partial weight sums (PWSs) of PEs so that the final local weight sums in each PE are as close to each other as possible. Resampling is then performed locally in each PE without the redistribution procedure before the next iteration. The corresponding architecture uses two levels of multiplexers to realize the weight/state sorting operation, and features regular structure, low execution time, high memory efficiency and well scalability. Xiaolang Yan, Ruohong Huan |
ISCAS | 3 |
| 2012 | A new via chain design method considering confidence level and estimation precisionabstractFor accurate prediction of via yield, via chains are usually fabricated on test chips to investigate issues about vias. To minimize the randomness of experiments and make the testing results more convincing, the confidence level and estimation precision of the via failure rate are investigated in this paper. Based on the Poisson yield model, the method of determining an adequate number of total vias is obtained using the law of large numbers and the de Moivre-Laplace theorem. Moreover, for a specific confidence level and estimation precision, the method of determining a suitable via chain length is proposed. For area minimization, an optimal combination of total vias and via chain length is further determined. Monte Carlo simulation results show that the method is in good accordance with theoretical analyses. Results of via failure rates measured on test chips also reveal that via chains designed using the proposed method has a better performance. In addition, the proposed methodology can be extended to investigate statistical significance for other failure modes. Xiao-hua Luo, Li-sheng Chen, Jiao-jiao Zhu, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 4 |
| 2012 | Array based HV/VH tree: an effective data structure for layout representationabstractWe present a new data structure for the representation of an integrated circuit layout. It is a modified HV/VH tree using arrays as the primary container in bisector lists and leaf nodes. By grouping and sorting objects within these arrays together with a customized binary search algorithm, our new data structure provides excellent performance in both memory usage and region query speed. Experimental results show that in comparison with the original HV/VH tree, which has been regarded as the best layout data structure to date, the new data structure uses much less memory and can become 30% faster on region query. Yongjun Zheng, Zheng Shi 0002, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 5 |
| 2012 | Scratch-concerned yield modeling for IC manufacturing involved with a chemical mechanical polishing processabstractIn existing integrated circuit (IC) fabrication methods, the yield is typically limited by defects generated in the manufacturing process. In fact, the yield often shows a good correlation with the type and density of the defect. As a result, an accurate defect limited yield model is essential for accurate correlation analysis and yield prediction. Since real defects exhibit a great variety of shapes, to ensure the accuracy of yield prediction, it is necessary to select the most appropriate defect model and to extract the critical area based on the defect model. Considering the realistic outline of scratches introduced by the chemical mechanical polishing (CMP) process, we propose a novel scratch-concerned yield model. A linear model is introduced to model scratches. Based on the linear model, the related critical area extraction algorithm and defect density distribution are discussed. Owing to higher correspondence with the realistic outline of scratches, the linear defect model enables a more accurate yield prediction caused by scratches and results in a more accurate total product yield prediction as compared to the traditional circular model. Jiao-jiao Zhu, Xiao-hua Luo, Li-sheng Chen, Yi Ye, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 5 |
| 2011 | Behavioral modeling of direct sampling mixerabstractThis paper presents a behavioral model of the direct sampling mixer (DSM) using MATLAB SIMULINK environment. The proposed model is integrated into a SIMULINK block with tunable parameters, which allows fast and convenient time domain simulations. Nonidealities, including transconductance nonlinearity, jitter effects and on-resistance thermal noise were addressed. Among these nonidealities, jitter problem is most difficult problem. We focused our attention on jitter effects in the SC filters which, to the best of our knowledge, is the first attempt to address this issue. In handling this problem, we adopted a quantified probability distribution method. Xiaolang Yan, Ruohong Huan |
ISCAS | 3 |
| 2011 | An Efficient Compatibility-Based Test Data Compression and Its Decoder Architecture
Min-yong Wan, Yong Ding 0003, Xiaolang Yan |
J. Electron. Test. | 4 |
| 2011 | A general communication performance evaluation model based on routing path decompositionabstractThe network-on-chip (NoC) architecture is a main factor affecting the system performance of complicated multi-processor systems-on-chips (MPSoCs). To evaluate the effects of the NoC architectures on communication efficiency, several kinds of techniques have been developed, including various simulators and analytical models. The simulators are accurate but time consuming, especially in large space explorations of diverse network configurations; in contrast, the analytical models are fast and flexible, providing alternative methods for performance evaluation. In this paper, we propose a general analytical model to estimate the communication performance for arbitrary NoCs with wormhole routing and virtual channel flow control. To resolve the inherent dependency of successive links occupied by one packet in wormhole routing, we propose the routing path decomposition approach to generating a series of ordered link categories. Then we use the traditional queuing system to derive the fine-grained transmission latency for each network component. According to our experiments, the proposed analytical model provides a good approximation of the average packet latency to the simulation results, and estimates the network throughput precisely under various NoC configurations and workloads. Also, the analytical model runs about 10 5 times faster than the cycle-accurate NoC simulator. Practical applications of the model including bottleneck detection and virtual channel allocation are also presented. Ai-lian Cheng, Xiaolang Yan, Ruohong Huan |
J. Zhejiang Univ. Sci. C | 3 |
| 2011 | Current oscillations and low-frequency noises in GaAs MESFET channels with sidegating biasabstractLow-frequency noises (LFN) and noise-like oscillations (NLO) in GaAs metal semiconductor field effect transistor (MESFET) channel current were investigated under sidegating bias conditions. It was found that the fluctuations of the channel current were directly dependent upon the sidegating bias. As the sidegating bias decreased, the amplitudes of the oscillations would increase correspondingly. Furthermore, the LFN and NLO would attenuate sharply when the sidegating bias increased to more than a certain voltage. Two mechanisms are presented to demonstrate that the effective substrate resistivity or the channel-substrate junction modulated by sidegating bias and deep level traps would take responsibilities for the LFN and NLO. Yong Ding 0003, Xiao-hua Luo, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 3 |
| 2011 | A sparse matrix model-based optical proximity correction algorithm with model-based mapping between segments and control sitesabstractOptical proximity correction (OPC) is a key step in modern integrated circuit (IC) manufacturing. The quality of model-based OPC (MB-OPC) is directly determined by segment offsets after OPC processing. However, in conventional MB-OPC, the intensity of a control site is adjusted only by the movement of its corresponding segment; this scheme is no longer accurate enough as the lithography process advances. On the other hand, matrix MB-OPC is too time-consuming to become practical. In this paper, we propose a new sparse matrix MB-OPC algorithm with model-based mapping between segments and control sites. We put forward the concept of ‘sensitive area’. When the Jacobian matrix used in the matrix MB-OPC is evaluated, only the elements that correspond to the segments in the sensitive area of every control site need to be calculated, while the others can be set to 0. The new algorithm can effectively improve the sparsity of the Jacobian matrix, and hence reduce the computations. Both theoretical analysis and experiments show that the sparse matrix MB-OPC with model-based mapping is more accurate than conventional MB-OPC, and much faster than matrix MB-OPC while maintaining high accuracy. Bin Lin 0003, Xiaolang Yan, Zheng Shi 0002, Yiwei Yang 0003 |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | Erratum to: A sparse matrix model-based optical proximity correction algorithm with model-based mapping between segments and control sitesabstractOptical proximity correction (OPC) is a key step in modern integrated circuit (IC) manufacturing. The quality of model-based OPC (MB-OPC) is directly determined by segment offsets after OPC processing. However, in conventional MB-OPC, the intensity of a control site is adjusted only by the movement of its corresponding segment; this scheme is no longer accurate enough as the lithography process advances. On the other hand, matrix MB-OPC is too time-consuming to become practical. In this paper, we propose a new sparse matrix MB-OPC algorithm with model-based mapping between segments and control sites. We put forward the concept of ‘sensitive area’. When the Jacobian matrix used in the matrix MB-OPC is evaluated, only the elements that correspond to the segments in the sensitive area of every control site need to be calculated, while the others can be set to 0. The new algorithm can effectively improve the sparsity of the Jacobian matrix, and hence reduce the computations. Both theoretical analysis and experiments show that the sparse matrix MB-OPC with model-based mapping is more accurate than conventional MB-OPC, and much faster than matrix MB-OPC while maintaining high accuracy. Bin Lin 0003, Xiaolang Yan, Zheng Shi 0002, Yiwei Yang 0003 |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | Efficient implementation of a cubic-convolution based image scaling engineabstractIn video applications, real-time image scaling techniques are often required. In this paper, an efficient implementation of a scaling engine based on 4×4 cubic convolution is proposed. The cubic convolution has a better performance than other traditional interpolation kernels and can also be realized on hardware. The engine is designed to perform arbitrary scaling ratios with an image resolution smaller than 2560×1920 pixels and can scale up or down, in horizontal or vertical direction. It is composed of four functional units and five line buffers, which makes it more competitive than conventional architectures. A strict fixed-point strategy is applied to minimize the quantization errors of hardware realization. Experimental results show that the engine provides a better image quality and a comparatively lower hardware cost than reference implementations. Yong Ding 0003, Ming-Yu Liu 0002, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 4 |
| 2011 | A robust motion estimation with center-biased diamond search and its parallel architecture for motion-compensated de-interlace
Yong Ding 0003, Xiaolang Yan |
J. Supercomput. | 2 |
| 2010 | A high efficient memory architecture for H.264/AVC motion compensationabstractIn H.264/AVC decoding system, motion compensation operation occupies about 80% of the total memory access and becomes the system bottleneck. In this paper, a high efficient memory architecture for H.264/AVC motion compensation is proposed to extremely reduce external memory access bandwidth. A four-level hierarchical memory organization scheme is utilized to explore the reusability of neighboring blocks at an acceptable area cost. To improve the system processing throughput, five optimization techniques are adopted in motion compensation operation, which enable video decoder to achieve real-time decoding of HD 1080p video stream when operating at 110 MHz. Compared with the existing works, the proposed architecture is able to reduce the memory bandwidth requirement in motion compensation progress by 83.7% and performs better in the real-time application. Chunshu Li, Kai Huang 0002, Xiaolang Yan, Jiong Feng, De Ma, Haitong Ge |
ASAP | 3 |
| 2010 | COSMO: CO-Simulation with MATLAB and OMNeT++ for Indoor Wireless NetworksabstractSimulations are widely used to design and evaluate new protocols and applications of indoor wireless networks. However, the available network simulation tools face the challenges of providing accurate indoor channel models, three-dimensional (3-D) models, model portability, and effective validation. In order to overcome these challenges, this paper presents a new CO-Simulation framework based on MATLAB and OMNeT++ (COSMO) to rapidly build credible simulations for indoor wireless networks. A hierarchical ad hoc passive RFID network for indoor tag locating is described as a case study, demonstrating the significance and efficiency of COSMO compared with other network simulators. COSMO surpasses other network simulators in terms of workload and validity. Zhonghai Lu, Qiang Chen 0014, Xiaolang Yan, Lirong Zheng 0001 |
GLOBECOM | 4 |
| 2010 | A Low Delay Multiple Reader Passive RFID System Using Orthogonal TH-PPM IR-UWBabstractNA Zhonghai Lu, Zhibo Pang, Xiaolang Yan, Qiang Chen 0014, Lirong Zheng 0001 |
ICCCN | 4 |
| 2010 | Discrete-time charge analysis for a digital RF charge sampling mixerabstractThis paper presents an approach for analyzing the key parts of a general digital radio frequency (RF) charge sampling mixer based on discrete-time charge values. The cascade sampling and filtering stages are analyzed and expressed in theoretical formulae. The effects of a pseudo-differential structure and CMOS switch-on resistances on the transfer function are addressed in detail. The DC-gain is restrained by using the pseudo-differential structure. The transfer gain is reduced because of the charge-sharing time constant when taking CMOS switch-on resistances into account. The unfolded transfer gains of a typical digital RF charge sampling mixer are analyzed in different cases using this approach. A circuit-level model of the typical mixer is then constructed and simulated in Cadence SpectreRF to verify the results. This work informs the design of charge-sampling, infinite impulse response (IIR) filtering, and finite impulse response (FIR) filtering circuits. The discrete-time approach can also be applied to other multi-rate receiver systems based on charge sampling techniques. Ning Ge 0001, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 3 |
| 2009 | Simulink®-based heterogeneous multiprocessor SoC design flow for mixed hardware/software refinement and simulation
Sangil Han, Soo-Ik Chae, Lisane B. de Brisolara, Luigi Carro, Katalin Popovici, Xavier Guerin, Ahmed Amine Jerraya, Kai Huang 0002, Xiaolang Yan |
Integr. | 10 |
| 2008 | A quasi fixed frequency constant on time controlled boost converterabstractContinuous conduction mode (CCM) boost converter is difficult to compensate due to the right half plane (RHP) zero. This paper presents a quasi fixed frequency constant on time (COT) controlled boost converter, it has no loop compensation requirement. With the proposed COT control, pulse frequency modulation (PFM) comes free at light load, which improves the light load efficiency compared with traditional PWM control. A compact on-time generator that controls the operation frequency is proposed. And a current sink based feedback circuit is used to improve the quality of feedback signal. The COT controller is designed with 1.5 μm BCD process. Simulation results demonstrate that the COT boost converter has fast load transient response and quasi fixed frequency operation. Xiaoru Xu, Xiaolang Yan |
ISCAS | 3 |
| 2007 | An Efficient SIMD Architecture with Parallel Memory for 2D Cosine Transforms of Video CodingabstractThis paper proposes an efficient SIMD architecture with parallel memory for 2D cosine transforms of multiple video standards. A novel parallel memory scheme is employed to provide conflict-free parallel access in both horizontal and vertical directions with the successive or even/odd mode, as well as to eliminate data permutation and matrix transposition. Furthermore, application specific instructions are presented to accelerate the transform kernels, such as butterfly and rotate operations with scaling, rounding and clipping. The simulation results show that proposed architecture achieves significant performance improvement with low hardware cost of 3.2K equivalent gate count for parallel memory subsystem (not including SRAMs) and 19.8K for arithmetic units@250MHz in 0.18 μm process. Jianying Peng, Xing Qin, Dexian Li, Xiaolang Yan, Xiexiong Chen |
ASAP | 4 |
| 2007 | An Efficient Diagnostic Test Pattern Generation Framework Using Boolean SatisfiabilityabstractThis paper presents a diagnostic test pattern generation (DTPG) framework based upon a Boolean Satisfiability engine. We first propose an enhanced miter-based model for distinguishing fault candidates that can achieve greater efficiency as well as can prove a group of undifferentiable faults. The model can also be used to generate diagnostic tests for distinguishing faults of different fault types. Based on this model, we propose a diagnostic pattern compaction strategy. By exploring "don't cares " at the primary inputs, the number of required diagnostic patterns can be reduced. Experimental results show that the proposed method achieves a greater diagnosis resolution when combined with existing approaches. Also, fewer diagnostic test patterns are needed. Feijun Zheng, Kwang-Ting Cheng, Xiaolang Yan, John Moondanos, Ziyad Hanna |
ATS | 3 |
| 2007 | Simulink-Based MPSoC Design Flow: Case Study of Motion-JPEG and H.264abstractSystem-level design methodologies have been introduced as a solution to handle the design complexity of embedded multiprocessor SoC (MPSoC) systems. In this paper we describe a system-level design flow starting from Simulink specification, focusing on concurrent hardware and software design and verification at four different abstraction levels: Simulink Combined Algorithm and Architecture Model (CAAM), Virtual Architecture, Transaction-accurate Model and Virtual Prototype. We used two multimedia applications, Motion-JPEG and H.264, to evaluate this design flow. Experimental results show that our design flow can generate various MPSoC architectures from Simulink CAAM correctly and efficiently, allowing processor and task design space exploration at different abstraction levels. Kai Huang 0002, Sangil Han, Katalin Popovici, Lisane B. de Brisolara, Xavier Guerin, Xiaolang Yan, Soo-Ik Chae, Luigi Carro, Ahmed Amine Jerraya |
DAC | 7 |
| 2007 | An optimized linear skewing interleave scheme for on-chip multi-access memory systemsabstractAn optimized linear skewing interleave scheme for on-chip multi-access memory systems is proposed in this paper. The proposed scheme can support simultaneous access of multiple subarray types of data elements in a 2-D data space with modulo addressing. 2pq (pq is the number of data elements in a subarray) memory modules are used without redundancy to save the on-chip memory. It uses linear skewing in the horizontal direction and uses nonlinear skewing in the vertical direction. Fast implementation method for the proposed scheme is also described. Results show that compared to previous linear skewing schemes, the proposed scheme can reduce 13.6%, on average, of the on-chip memory for cases of pq = 4 or 8 and reduce 35.5%, on average, of the external memory bandwidth for benchmark of motion estimation due to modulo addressing. Chunyue Liu, Xiaolang Yan, Xing Qin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Phase noise analysis of oscillators with Sylvester representation for periodic time-varying modulus matrix by regular perturbations
JianXing Fan, Huazhong Yang, Hui Wang 0004, Xiaolang Yan, Chaohuan Hou |
Sci. China Ser. F Inf. Sci. | 4 |
| 2005 | A novel data processing circuit in high-speed serial communicationabstractA novel data processing circuit in high-speed serial communication has been demonstrated in this work. The circuit, including a serializer and a frequency divider, was developed to convert the transmission signals into the desired format. The chip design is based on TSMC 0.25μm mixed signal model, and semi-custom design methodology is used. Pre- and post-layout simulation results indicated that the speed of circuit has reached 480MHz. Moreover, the data are processed properly in agreement with USB2.0 specification. Yongjian Tang, Lenian He, Xiaolang Yan |
ASP-DAC | 3 |
| 2005 | A new method for model based frugal OPCabstractImprovements on Resolution Enhancement Technologies (RETs) enable minimum feature size of IC to shrink consistently with Moore's Law. However growing mask data volume also tremendously increases manufacture cost. The cost increase is partially due to the complicated optical proximity corrections applied on mask design. Frugal OPC methods have been introduced to reduce the complexity. In this paper, a new method for frugal OPC is presented. Based on recognition of critical spots under yield related constraints, the new correction flow keeps fidelity on critical sites while still retaining the frugality of modified designs. Xiaolang Yan, Zheng Shi 0002 |
ASP-DAC | 1 |
| 2005 | Q-DPM: An Efficient Model-Free Dynamic Power Management TechniqueabstractWhen applying dynamic power management (DPM) techniques to pervasively deployed embedded systems, the technique needs to be very efficient so that it is feasible to implement the technique on low end processors and tight-budget memory. Furthermore, it should have the capability to track time varying behavior rapidly, because the time variance is an inherent characteristic of real world systems. Existing methods, which are usually model-based, may not satisfy the aforementioned requirements. In this paper, we propose a model-free DPM technique based on Q-learning. Q-DPM is much more efficient because it removes the overhead of parameter estimator and mode-switch controller. Furthermore, its policy optimization is performed via consecutive online trialing, which also leads to very rapid response to time varying behavior. Richard Yao, Xiaolang Yan |
DATE | 4 |
| 2005 | Processor Load Analysis for Mobile Multimedia Streaming: The Implication of Power ReductionabstractThe software codec on mobile device introduces significant power consumption because the energy efficiency of general processor based system is much lower than that of the dedicated hardware such as ASIC based accelerator. Dynamical Voltage Scaling (DVS) is one of the most efficient techniques to promote the energy efficiency. Most existing papers on this topic use simple heuristics to predict processor load, and poor prediction accuracy is observed in experiments. We advocate intensive analysis on processor load before designing DVS framework and algorithm. Hence, we conduct load analysis on more than 600 processor load trace files for 57 test sequences and 98 representative clips from Internet. Basic statistical analysis and time series analysis are applied intensively to identify major characteristics of the processor load. The analysis shows that it is feasible to predict processor load using low order linear time series model if the load is sampled using feature period. Moreover, there is indeed significant potential to reduce the energy consumption. Based on the analysis results, we develop a fully adaptive DVS technique to adjust supply voltage online with controllable penalty. Zihua Guo, Richard Yao, Xiaolang Yan |
ICME | 5 |
| 2005 | Full-IC manufacturability check based on dense silicon imaging
Xiaolang Yan, Zheng Shi 0002, Gensheng Gao |
Sci. China Ser. F Inf. Sci. | 1 |
| 2004 | Power Consumption of Wireless NIC and Its Impact on Joint Routing and Power Control in Ad Hoc Network
Menglian Zhao, Xiaolang Yan |
EUC | 5 |
| 2004 | Heterogeneous Grid Computing for Energy Constrained Mobile Device
Menglian Zhao, Xiaolang Yan |
EUC | 5 |
| 2004 | Tiling artifact reduction for JPEG2000 image at low bit-rateabstractJPEG2000 is a very promising still image standard because of its excellent performance. However, it also causes a higher complexity to implement. In practice, in JPEG2000 coding systems, an image is segmented into serial tiles and each tile is compressed or transformed independently, thus creating tiling artifacts in the tile boundary. At low bit rates, tiling artifacts in JPEG2000 images are annoying. This paper introduced a new post-processing method to reduce this artifact, where max-lift wavelet subband decomposition was used and then subband coefficients were adaptively filtered and soft-thresholded. The post-processing is done out of JPEG2000 image coding systems, so it has no compatible problem with the JPEG2000 standard. Experiments showed that the post-processing was effective to obviously enhance the visual quality of the JPEG2000 image at low bit rates. Xing Qin, Xiaolang Yan, Chong-Peng Yang |
ICME | 2 |
| 2003 | Equivalence Checking Using Independent CutsabstractWith the increase in the complexity of present day systems, proving the correctness of a design has become a major concern. This paper describes a novel implementation of a BDD-based combinational equivalence checking (CEC) tool, which is distinguished from others by one heuristic. It is proposed to select an effective cut, with no dependence remaining. In addition, successfully verification of all the ISCAS'85 benchmark circuits demonstrates the efficiency of our approach. Xiaolang Yan, Yongjiang Lu, Haitong Ge |
Asian Test Symposium | 2 |