Xiaojian Yang

dblp:99/5425 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 12 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 A Diagonal Block Memory-Aware Polynomial Preconditioner for Linear and Eigenvalue Solvers
abstract
Krylov subspace methods are widely used in scientific computing to solve large sparse linear systems and eigenvalue problems. Their performance bottleneck is often dominated by high-order matrix-power kernels (MPK), especially in polynomial preconditioners that must scale to millions or billions of variables. We present Diagonal Block MPK (DBMPK), a lightweight and parallel-friendly optimization that partitions the input matrix into diagonal blocks and off-diagonal regions. This design enables efficient intra-block data reuse and eliminates inter-block dependencies. It improves cache locality, parallelism, and reduces preprocessing overheads, compared to existing techniques. Our evaluation on x86 and Arm HPC platforms shows that DBMPK improves MPK performance by 26.6%-38.4%. When applied to polynomial preconditioners for linear systems and eigenvalue problems, it achieves consistent end-to-end speedups of 18.6%-34.0%, including in weak scaling tests on 128 nodes, demonstrating strong scalability and practical impact.
Xiaojian Yang, Yuhui Ni, Shengguo Li, Dezun Dong, Chuanfu Xu, Haipeng Jia, Jie Liu 0002
PPoPP1
2026 A Memory-Aware Sparse Matrix-Matrix Multiplication on Multicore Architectures
abstract
Sparse matrix–matrix multiplication (SpMM) is a fundamental operation in scientific computing with broad applications across numerous domains. Tiling is a key optimization technique for improving data locality and is widely adopted in high-performance computing. However, the irregular data access patterns inherent to SpMM make it challenging to exploit tiling effectively for data reuse. In this article, we propose MaSpMM , a memory-aware SpMM framework that integrates cache-aware tiling with a segment-oriented data layout. MaSpMM stores matrices as continuous segments to enhance data locality within each tile. Moreover, since many sparse matrices in real-world applications exhibit symmetry, we further develop MaSpMM-Sym, an extension that recursively partitions symmetric matrices to eliminate write conflicts and further improve locality. To adapt to diverse scenarios, we finally introduce MaSpMM-Adap, which adaptively selects the most suitable approach for each input matrix. Comprehensive evaluations on both x86 and ARM CPUs demonstrate that MaSpMM-Adap achieves average speedups of up to 1.86× over Intel oneMKL, 1.84× over ASpT, and 1.75× over J-Stream.
Deshun Bi, Shengguo Li, Haozhong Qiu, Chuanfu Xu, Xiaojian Yang, Dezun Dong, Tiaojie Xiao, Jie Liu 0002
ACM Trans. Archit. Code Optim.5
2025 CRAMG: A Communication-Reduced Algebraic Multigrid Method
abstract
Algebraic multigrid (AMG) is widely used to accelerate largescale sparse linear solvers.In distributed environments, neighboring communication overhead in AMG significantly impacts overall solution time.We propose Communication-Reduced Algebraic Multigrid (CRAMG) methods to minimize inter-process data exchange and message count by fusing interpolation/restriction operators with residual computations.This reduces communication frequency from four per level to as few as two.Experiments show up to 45% reduction in data exchange and 35% fewer messages.Performance evaluations on an Intel platform demonstrate significant improvements
Xiaojian Yang, Yunqing Huang, Dezun Dong, Chuanfu Xu, Jie Liu 0002, Xiaoqiang Yue, Shengguo Li
ICS2
2025 DAS-ILU: A Distributed Asynchronous Parallel ILU Factorization Based on Domain Decomposition
abstract
This paper presents DAS-ILU, a Distributed Asynchronous parallel Incomplete LU factorization method based on domain decomposition. DAS-ILU partitions the computational domain into independently processed interior nodes and asynchronously updated separator nodes, thereby reducing cross-processor dependencies and halving the separator size compared to conventional methods. To further improve performance, it employs optimized data exchange patterns to minimize communication overhead and extends support to block-structured sparse matrices via exact block inversions. Comprehensive evaluations on a range of problem types—including structural mechanics, computational fluid dynamics, and reservoir simulation demonstrate the superior performance of DAS-ILU. Compared to state-of-the-art ILU implementations, DAS-ILU achieves solve time speedups of up to 2.07 × over Chow-Patel’s fine-grained parallel ILU and up to 4.11 × over HYPRE’s ILU. Moreover, DAS-ILU exhibits strong robustness when applied to challenging nonsymmetric and indefinite systems.
Shengguo Li, Xiaojian Yang, Yunqing Huang, Chuanfu Xu, Dezun Dong, Jianchun Wang, Jie Liu 0002
SC3
2024 DBSR: An Efficient Storage Format for Vectorizing Sparse Triangular Solvers on Structured Grids
abstract
The Sparse Triangular Solver (SPTRSV) plays a critical role in solving structured grid problems. Yet, the commonly used sparse matrix storage formats for structured grid methods do not efficiently support SPTRSV in utilizing the instruction parallelism offered by modern multi-core CPUs. We introduce DBSR, a new sparse storage format to enable SPTRSV to take advantage of the SIMD instructions. DBSR promotes contiguous memory access and vectorized computation, while also optimizing memory usage. We evaluate DBSR by applying it within multigrid algorithms and the zero fill-in incomplete $\mathbf{L U}$ preconditioner. Our evaluation, conducted on four architectures - three ARMv8 systems and one x86 system - demonstrates that DBSR consistently outperforms mainstreamed storage formats across evaluation workloads and platforms.
Xiaojian Yang, Shengguo Li, Dezun Dong
SC1
2024 Optimizing Multi-Grid Preconditioned Conjugate Gradient Method on Multi-Cores
abstract
Multigrid preconditioned conjugate gradient (MGPCG) is commonly used in high-performance computing (HPC) workloads. However, MGPCG is notoriously challenging to optimize since most of its computation kernels are memory-bounded with low arithmetic intensity and non-trivial communication patterns among parallel processes. This article presents new techniques to improve the data locality and reduce the communication overhead of MGPCG by first merging the kernels of multigrid (MG). We then develop an asynchronous neighboring communication algorithm to reduce the data communications across parallel processes. We demonstrated the benefits of our approach by applying it to the high-performance conjugate gradient (HPCG) benchmark and integrating it with a real-life algebraic multigrid package. We test the resulting software implementations on three ARMv8 and one Intel Xeon system. Experimental results show that our approach leads to a 1.62x-2.54x speedup over the engineer- and vendor-tuned HPCG implementations across various workloads and platforms.
Xiaojian Yang, Shengguo Li, Dezun Dong, Chun Huang 0006, Zheng Wang 0079
IEEE Trans. Parallel Distributed Syst.2
2023 Efficiently Running SpMV on Multi-core DSPs for Banded Matrix
Deshun Bi, Shengguo Li, Xiaojian Yang, Dezun Dong
ICA3PP (5)4
2023 Optimizing Multi-grid Computation and Parallelization on Multi-cores
abstract
Multigrid algorithms are widely used to solve large-scale sparse linear systems, which is essential for many high-performance workloads. The symmetric Gauss-Seidel (SYMGS) method is often responsible for the performance bottleneck of MG. This paper presents new methods to parallelize and enhance the computation and parallelization efficiency of the SYMGS and MG algorithms on multi-core CPUs. Our solution employs a matrix splitting strategy and a revised computation formula to decrease the computation operations and memory accesses in SYMGS. With this new SYMGS strategy, we can then merge the two most time-consuming components of MG. On top of these, we propose a new asynchronous parallelization scheme to reduce the synchronization overhead when parallelizing SYMGS. We demonstrate the benefit of our techniques by integrating them with the HPCG benchmark and two real-life applications. Evaluation conducted on four architectures, including three ARMv8 and one x86, shows that our techniques greatly surpass the performance of engineer- and vendor-tuned implementations across various workloads and platforms.
Xiaojian Yang, Shengguo Li, Dezun Dong, Chun Huang 0006, Zheng Wang 0001
ICS1
2023 Memory-aware Optimization for Sequences of Sparse Matrix-Vector Multiplications
abstract
This paper presents a novel approach to optimize multiple invocations of a sparse matrix-vector multiplication (SpMV) kernel performed on the same sparse matrix A and dense vector x, like Ax, A2x, ⋯, Akx, and their linear combinations such as Ax + A2x. Such computations are frequently used in scientific applications for solving linear equations and in multi-grid methods. Existing SpMV optimization techniques typically focus on a single SpMV invocation and do not consider opportunities for optimization across a sequence of SpMV operations (SSpMV), leaving much room for performance improvement. Our work aims to bridge this performance gap. It achieve this by partitioning the sparse matrix into submatrices and devising a new computation pipeline that reduces memory access to the sparse matrix and exploits the data locality of the dense vector of SpMV. Additionally, we demonstrate how our approach can be integrated with parallelization schemes to further improve performance. We evaluate our approach on four distinct multi-core systems, including three ARM and one Intel platform. Experimental results show that our techniques improve the standard implementation and the highly-optimized Intel math kernel library (MKL) by a large margin.
Shengguo Li, Dezun Dong, Xiaojian Yang, Zheng Wang 0001
IPDPS5
2012 Multiple-platform data integration method with application to combined analysis of microarray and proteomic data
abstract
BACKGROUND: It is desirable in genomic studies to select biomarkers that differentiate between normal and diseased populations based on related data sets from different platforms, including microarray expression and proteomic data. Most recently developed integration methods focus on correlation analyses between gene and protein expression profiles. The correlation methods select biomarkers with concordant behavior across two platforms but do not directly select differentially expressed biomarkers. Other integration methods have been proposed to combine statistical evidence in terms of ranks and p-values, but they do not account for the dependency relationships among the data across platforms. RESULTS: In this paper, we propose an integration method to perform hypothesis testing and biomarkers selection based on multi-platform data sets observed from normal and diseased populations. The types of test statistics can vary across the platforms and their marginal distributions can be different. The observed test statistics are aggregated across different data platforms in a weighted scheme, where the weights take into account different variabilities possessed by test statistics. The overall decision is based on the empirical distribution of the aggregated statistic obtained through random permutations. CONCLUSION: In both simulation studies and real biological data analyses, our proposed method of multi-platform integration has better control over false discovery rates and higher positive selection rates than the uncombined method. The proposed method is also shown to be more powerful than rank aggregation method.
Shicheng Wu, Yawen Xu, Zeny Z. Feng, Xiaojian Yang, Xiaogang Wang 0007, Xin Gao 0033
BMC Bioinform.4
2012 The LASSO and Sparse Least Squares Regression Methods for SNP Selection in Predicting Quantitative Traits
abstract
Recent work concerning quantitative traits of interest has focused on selecting a small subset of single nucleotide polymorphisms (SNPs) from amongst the SNPs responsible for the phenotypic variation of the trait. When considered as covariates, the large number of variables (SNPs) and their association with those in close proximity pose challenges for variable selection. The features of sparsity and shrinkage of regression coefficients of the least absolute shrinkage and selection operator (LASSO) method appear attractive for SNP selection. Sparse partial least squares (SPLS) is also appealing as it combines the features of sparsity in subset selection and dimension reduction to handle correlations amongst SNPs. In this paper we investigate application of the LASSO and SPLS methods for selecting SNPs that predict quantitative traits. We evaluate the performance of both methods with different criteria and under different scenarios using simulation studies. Results indicate that these methods can be effective in selecting SNPs that predict quantitative traits but are limited by some conditions. Both methods perform similarly overall but each exhibit advantages over the other in given situations. Both methods are applied to Canadian Holstein cattle data to compare their performance.
Zeny Z. Feng, Xiaojian Yang, Sanjeena Subedi, Paul D. McNicholas
IEEE ACM Trans. Comput. Biol. Bioinform.2
2007 An Instruction Folding Solution to a Java Processor
Yiyu Tan, Anthony Shi-Sheung Fong, Xiaojian Yang
NPC3
2006 Dragon2006: blockage-aware congestion-controlling mixed-size placer
abstract
In this paper, we develop a mixed-size placement tool, Dragon2006, to solve large scale placement problems effectively. A top-down hierarchical approach based on min-cut partitioning and simulated annealing is used to place very large SoC-style designs containing fixed blockage, movable macro blocks of various sizes and standard cells. Moreover, we have applied several techniques for wirelength optimization, congestion estimation in the presence of blockage and white space allocation for congestion removal.
Taraneh Taghavi, Xiaojian Yang, Bo-Kyung Choi, Maogang Wang, Majid Sarrafzadeh
ISPD2
2005 Dragon2005: large-scale mixed-size placement tool
abstract
In this paper, we develop a mixed-size placement tool, Dragon2005, to solve large scale placement problems effectively. A top-down hierarchical approach based on min-cut partitioning and simulated annealing is used to place very large SoC-style designs containing thousands of macro blocks of various sizes and millions of standard cells. Macro aware partitioning and techniques to properly handle different bin sizes are required, because of the existence of large macro blocks. Our tool is also a congestion and timing aware placement tool.
Taraneh Taghavi, Xiaojian Yang, Bo-Kyung Choi
ISPD2
2003 Routability-driven white space allocation for fixed-die standard-cell placement
abstract
The use of white space in fixed-die standard-cell placement is an effective way to improve routability. In this paper, we present a white space allocation approach that dynamically assigns white space according to the congestion distribution of the placement. In the top-down placement flow, white space is assigned to congested regions using smooth allocating functions. A post-allocation optimization step is taken to further improve placement quality. Experimental results show that the proposed allocation approach, combined with a multilevel placement flow, significantly improves placement routability and layout quality. A set of approaches for white space allocation has been presented and compared in this paper. All of them are based on routability-driven methods. However, these approaches vary in the allocation function and allocation aggressiveness. All the placement results are investigated by feeding them into a widely used industrial router (Warp Route of Cadence). Comparisons have been made between: 1) placement with or without white space allocation; 2) different white space allocation approaches; and 3) our placement flow, industrial placement tool, and the other state-of-the-art academic placement tool.
Xiaojian Yang, Bo-Kyung Choi, Majid Sarrafzadeh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Congestion reduction during placement with provably good approximation bound
abstract
This paper presents a novel method to reduce routing congestion during placement stage. The proposed approach is used as a post-processing step in placement. Congestion reduction is based on local improvement on the existing layout. However, the approach has a global view of the congestion over the entire design. It uses integer linear programming (ILP) to formulate the problem of conflicts between multiple congested regions, and performs local improvement according to the solution of the ILP problem. The approximation algorithm of the formulated ILP problem is studied and good approximation bounds are given and proved. Experiments show that the proposed approach can effectively alleviate the congestion of global routing results. The low computational complexity of the proposed approach indicates its scalability on large designs.
Xiaojian Yang, Maogang Wang, Ryan Kastner, Soheil Ghiasi, Majid Sarrafzadeh
ACM Trans. Design Autom. Electr. Syst.1
2002 Timing-driven placement using design hierarchy guided constraint generation
abstract
Design hierarchy plays an important role in timing-driven placement for large circuits. In this paper, we present a new methodology for delay budgeting based timing-driven placement. A novel slack assignment approach is described as well as its application on delay budgeting with design hierarchy information. The proposed timing-driven placement flow is implemented into a placement tool named Dragon (timing-driven mode), and evaluated using an industrial place and route flow. Compared to Cadence QPlace, timing-driven Dragon generates placement results with shorter clock cycle and better routability.
Xiaojian Yang, Bo-Kyung Choi, Majid Sarrafzadeh
ICCAD1
2002 A Standard-Cell Placement Tool for Designs with High Row Utilization
abstract
In this paper we study the correlation between wirelength and routability for standard-cell placement problem, under the modern place-and-route environment. We present a placement tool named Dragon (version 2.1), and show its ability to produce good quality placement for designs with high row utilization. Compared to an industrial placer and an academic state-of-the-art placer, Dragon can produce placement with better routability and shorter total wirelength. We describe many novel algorithmic details and implementation details of this placement tool. Experimental results show that minimizing wirelength improves routability and layout quality.
Xiaojian Yang, Bo-Kyung Choi, Majid Sarrafzadeh
ICCD1
2002 Routability driven white space allocation for fixed-die standard-cell placement
abstract
The use of white space in fixed-die standard-cell placement is an effective way to improve routability. In this paper, we present a white space allocation approach that dynamically assigns white space according to the congestion distribution of the placement. In the top-down placement flow, white space is assigned to congested regions using a smooth allocating function. A post allocation optimization step is taken to further improve placement quality. Experimental results show that the proposed allocation approach, combined with a multilevel placement flow, significantly improves placement routability and layout quality.In our experiments, we compared our placement tool with two other fixed-die placers using an industrial place and route flow. Placements created by all three tools have been routed with an industrial router (Warp Route of Cadence). Compared with a leading-edge industrial tool, our placer produces placements with similar or better routability and on average 8.8% shorter routed wirelength. Furthermore, our tool produces placement that runs faster through the Warp Route compared with the industrial tool. Compared with a state-of-the-art academic placement tool (Capo/MetaPlacer), our placer shows ability to produce more routable placements: for 15 out of all 16 benchmarks our placer's outputs are routable while Capo/MetaPlacer only creates 4 routable placements.
Xiaojian Yang, Bo-Kyung Choi, Majid Sarrafzadeh
ISPD1
2002 Predicting potential performance for digital circuits
abstract
Presents a new concept of potential slack to measure so-called potential performance of digital circuits. Potential means how much improvement could be made in the future in terms of timing, area, and power dissipation. Predicting potential performance helps the circuit designers make good design decisions at a specific level of abstraction. The authors describe two algorithms for potential slack: an optimal algorithm and a fast greedy algorithm. The former is based on a maximal-independent set on transitive graphs, while the latter focuses on potential slack estimation in a greedy manner. Applications to gate- and physical-level design problems are provided to show the effectiveness of potential slack in predicting the potential performance of digital circuits.
Chunhong Chen, Xiaojian Yang, Majid Sarrafzadeh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Congestion estimation during top-down placement
abstract
Congestion is one of the fundamental issues in very large scale integration physical design. In this paper, we propose two congestion-estimation approaches for early placement stages. First, we theoretically analyze the peak-congestion value of the design and experimentally validate the estimation approach. Second, we estimate regional congestion at the early stages of top-down placement. This is done by combining the wire-length distribution model and interregion wire estimation. Both approaches are based on the well-known Rent's rule, which is previously used for wirelength estimation. This is the first attempt to predict congestion using Rent's rule. The estimation results are compared with the layout after placement and global routing. Experiments on large industry circuits show that the early congestion estimation based on Rent's rule is a promising approach.
Xiaojian Yang, Ryan Kastner, Majid Sarrafzadeh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2001 Congestion Reduction During Placement Based on Integer Programming
abstract
This paper presents a novel method to reduce routing congestion during placement stage. The proposed approach is used as a post-processing step in placement. Congestion reduction is based on local improvement on the existing layout. However, the approach has a global view of the congestion over the entire design. It uses integer linear programming (ILP) to formulate the conflicts between multiple congested regions, and performs local improvement according to the solution of ILP. Experiments show that the proposed approach can effectively reduce the total overflow of global routing result. The short running time of the algorithm indicates good scalability on large designs.
Xiaojian Yang, Ryan Kastner, Majid Sarrafzadeh
ICCAD1
2001 Congestion estimation during top-down placement
abstract
Congestion is one of the fundamental issues in VLSI physical design. In this paper, we propose two congestion estimation approaches for early placement stages. First, we theoretically analyze the peak congestion value of the design and experimentally validate the estimation approach. Second, we estimate regional congestion in the early top-down placement. This is done by combining the wirelength distribution model and inter-region wire estimation. Both approaches are based on the well known Rent's rule, which is previously used for wirelength estimation. This is the first attempt to predict congestion using Rent's rule. The estimation results are compared with the layout after placement and global routing. Experiments on large industry circuits show that the early congestion estimation based on Rent's rule is a promising approach.
Xiaojian Yang, Ryan Kastner, Majid Sarrafzadeh
ISPD1
2000 Potential Slack: An Effective Metric of Combinational Circuit Performance
abstract
This paper proposes the concept of potential slack and shows that it is an effective metric of combinational circuit performance. We provide several methods for estimating potential slack and prove one (a maximal-independent-set based algorithm) in particular which works best. Experiments in gate sizing show that potential slack provides 100% correct prediction for circuit area optimization. We also explore the role of potential slack in timing-driven placement.
Chunhong Chen, Xiaojian Yang, Majid Sarrafzadeh
ICCAD2
2000 DRAGON2000: Standard-Cell Placement Tool for Large Industry Circuits
abstract
In this paper, we develop a new standard cell placement tool, Dragon2000, to solve large scale placement problem effectively. A top-down hierarchical approach is used in Dragon2000. State-of-the-art partitioning tools are tightly integrated with wirelength minimization techniques to achieve superior performance. We argue that net-cut minimization is a good and important shortcut to solve the large scale placement problem. Experimental results show that minimizing net-cut is more important than greedily obtain a wirelength optimal placement at intermediate hierarchical levels. We run Dragon2000 on recently released large benchmark suite ISPD98 as well as MCNC circuits. For circuits which have more than 100 k cells, comparing to iToolsl.4.0, Dragon2000 can produce slightly better placement results (1.4%) while spending much less amount of time (2/spl times/ speedup). This is also the first published placement result on the publicly available large industrial circuits.
Maogang Wang, Xiaojian Yang, Majid Sarrafzadeh
ICCAD2
2000 Multi-center congestion estimation and minimization during placement
abstract
As technology advances, more and more issues need to be considered in the placement stage, e.g., wirelength, congestion, timing, coupling. It is very hard to consider all of them together at the same time. Thus it is good if we can optimize one cost function without a ecting others. In this paper, we will study methods to optimize congestion in placement without in icting degradations/violations in other objectives or constraint. We give a mathematical equation to predict the over ow within a region using a normal distribution approximation. According to experiments, this equation does give a good estimation of over ow. We used this equation to nd the smallest regions whichhave enough routing resource to alleviate the congestion and propose the exible expansion scheme in our multi-center congestion reduction (MC 2 R) algorithm. Experimental results show that generally there is a correlation between the amount of reduction in congestion and the amount ofchange made to the placement: the more we change the placement, the more reduction in congestion we will get. However, the exible expansion scheme is very e ective in helping us reduce congestion while make only little change to the placement. Comparing to the full expansion scheme (49 % congestion reduction and 6:5 % change in placement), the exible expansion scheme together with MC 2 R algorithm can reduce congestion by almost the same amount (42%) with much less change made to the placement (1:8%). 1.
Maogang Wang, Xiaojian Yang, Kenneth Eguro, Majid Sarrafzadeh
ISPD2
2000 A snap-on placement tool
abstract
The standard cell placement problem has been extensively studied in the past twenty years. Many approaches were proposed and proven effiective in practice. However, successful placement tools need enormous time in the course of development. In this paper we propose a new snap-on placement tool, which is based on multilevel hierarchical placement method. It has great flexibility to combine existing packages and techniques in its top-down framework. In addition, it can be used to build a good placement tool in a short amount of time. Some important issues in multilevel hierarchical placement are discussed here. We investigate the behavior of net-cut and wirelength objectives in global placement problem, propose a + level clustering technique and design a new topdown placement method based on partitioning, annealing and + level technique. We alsowork on the trade-off between solution quality and running time during the hierarchical placement. Experimental results show the strength of proposed placement tool, it produces very good results on all benchmarks and the best known result on the largest MCNC benchmark (avql).
Xiaojian Yang, Maogang Wang, Kenneth Eguro, Majid Sarrafzadeh
ISPD1
2000 Congestion minimization during placement
abstract
Typical placement objectives involve reducing net-cut cost or minimizing wirelength. Congestion minimization is the least understood, however, it models routability most accurately. In this paper, we study the congestion minimization problem during placement. First, we show that a global placement with minimum wirelength has minimum total congestion. We show that minimizing wirelength may (and in general, will) create locally congested regions. We test seven different congestion minimization objectives. We also propose a post processing stage to minimize congestion. Our main contribution and results can be summarized as follows. (1) Among a variety of cost functions and methods for congestion minimization (including several currently used in industry), wirelength alone followed by a post processing congestion minimization works the best and is one of the fastest. (2) Cost functions such as a hybrid length plus congestion (commonly believed to be very effective) do not always work very well. (3) Net-centric post-processing techniques are among the best congestion alleviation approaches. (4) Congestion at the global placement level, correlates well with congestion of detailed placement.
Maogang Wang, Xiaojian Yang, Majid Sarrafzadeh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2