Jun Tao 0001

dblp:35/5170-1 · DBLP profile ↗
← Back
50ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-8742-687XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 49 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 BitPair: An Efficient 2-Bit Serial Precision-Scalable Accelerator for GEMM in Deep Neural Networks
abstract
As deep neural networks grow in size, their computational and energy demands pose significant challenges for efficient hardware acceleration. Quantization reduces storage, data movement, and computation costs by lowering numerical precision. While 8-bit integer (INT8) is widely adopted, lower precisions such as INT6 and INT4 can also maintain acceptable accuracy with specialized quantization techniques, motivating accelerators that support multiple precisions. This paper presents BitPair, a precision-scalable accelerator that employs a 2-bit serial dataflow to perform general matrix multiplications (GEMM). BitPair supports signed and unsigned integer operations at 2/4/6/8-bit precision for both input matrices. Implemented in Verilog and synthesized in TSMC 12nm technology, BitPair is compared with two prior precision-scalable architectures, Loom and BitShare. Evaluation shows that BitPair achieves up to 1.63 × lower bandwidth, 1.19 × smaller area, and 1.14 × lower power compared with Loom and BitShare.
Jun Tao 0001, Jun Han 0003
ACM Great Lakes Symposium on VLSI3
2026 Sensitivity Analysis of Process Variables for Large-Scale Circuits based on Graph Attention Networks
Zhe Zhu, Jun Tao 0001
ISCAS4
2025 MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware System
abstract
The rapid advancement of visual autonomous systems, especially in autonomous driving, underscores the critical role of Image Signal Processors (ISPs) as they convert RAW sensor data into RGB images suited for visual interpretation. Traditional ISPs rely on tuning hyperparameters to adapt to varying imaging conditions; however, the vast parameter space and intricate tuning process pose significant challenges for realtime autonomous applications. Existing autonomous ISP hyperparameter optimization methods rely largely on offline or proxybased online tuning, limiting their accuracy and responsiveness to real-time environmental changes. In response, we propose an online ISP hyperparameter optimization framework based on Deep Reinforcement Learning (DRL), marking the first proxyfree, real-time optimization approach. Our design exhibits a master-slave Multi-Agent System (MAS), enabling rapid and cooperative parameter optimization with improved inter-frame consistency. Furthermore, we design the MAS-ISP automated visual system, incorporating innovative hardware designs such as Strip Convolution Kernel and Stride-Aware Dual-Buffer Memory, which drastically reduce resource consumption in CNN hardware. MAS-ISP achieves 1080P@75FPS/240FPS on FPGA/ASIC platforms, supporting real-time and reliable visual systems.
Zhijian Hao, Ruoxi Zhu, Qi Zheng 0004, Shuocheng Wang, Shushi Chen, Leilei Huang, Jun Tao 0001, Yibo Fan
DAC10
2025 Robust analog/RF circuit design via Cycle-Consistent Generative Adversarial Networks
Nanlin Guo, Jun Tao 0001, Xuan Zeng 0001, Xin Li 0001
Integr.2
2025 Efficient Design Optimization for Diffractive Deep Neural Networks
abstract
Since diffractive deep neural network (D2NN) provides a full optical solution to implement deep neural networks (DNNs), it offers ultrafast operation speed and virtually unlimited bandwidth, yielding an alternative-yet-competitive approach for computer-based neural networks. A D2NN is composed of several 3D-printed phase masks as hidden layers and a number of optical detectors at the output. To enable automatic and efficient design of D2NNs, we propose an iterative optimization method to determine the optimal design parameters of D2NNs. During each iteration step, we first optimize the physical parameters for masks (e.g., thicknesses) while fixing the detector parameters (e.g., locations). Next, we exhaustively search the detector parameters with fixed masks. These two steps are repeated until convergence is reached. Our numerical experiments demonstrate that the proposed optimization algorithm can produce a high-performance D2NN achieving 97% accuracy for recognizing handwritten digits.
Yuncheng Liu, Jun Tao 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Coordinating Binary Trait: Accurate and Lightweight Runtime On-Chip Power Meter Design
abstract
As heterogeneous platforms scale, power management issues grow increasingly critical. On-chip power awareness is fundamental to effective power management on such platforms. However, the integration of diverse components leads to higher power consumption and more complex power patterns, complicating the perception of power distribution. To address this challenge, this article presents COordinating BInary Trait (COBIT), which focuses on the design and optimization of on-chip power meters (OPMs) for heterogeneous design, enabling power management in large-scale systems. COBIT extracts essential bitwise wires as features, learns a boosting model based on binary trees, and implements the OPM for runtime power prediction. Additionally, multiobjective optimization algorithms are employed in the design space exploration of the power meter to ensure an accurate and low-overhead integration. Experimental results demonstrate the capabilities of COBIT on heterogeneous hardware. Using only 0.028% bitwise power proxies of all RTL wires in the NVIDIA deep learning accelerator (NVDLA) and 0.036% in the Berkeley out-of-order machine (BOOM) processor core, the per-cycle models achieve a mean absolute percentage error (MAPE) of 1.69% and 2.49%, respectively. Regarding the success rate of predictions on power peaks in NVDLA, COBIT achieves a maximum performance improvement of 12.93% and an average improvement of 8.11% over existing methods. Moreover, the gate-area/power overhead of our OPM on BOOM and NVDLA is 1.25%/0.217% and 0.49%/0.097%, respectively, while performing per-cycle power prediction in just three cycles. Unlike previous approaches, which struggle with balancing accuracy and efficiency, COBIT effectively addresses both challenges, delivering unprecedented benefits for large-scale systems.
Weixi Lyu, Yifan Liu 0017, Jun Tao 0001, Jun Han 0003
IEEE Trans. Very Large Scale Integr. Syst.4
2024 Auto-ISP: An Efficient Real-Time Automatic Hyperparameter Optimization Framework for ISP Hardware System
abstract
Image Signal Processor (ISP) is widely used in intelligent edge devices across various scenarios. The intricate and time-consuming tuning process demands substantial expertise. Current AI-based auto-tuning operates discretely offline, relying on predefined scenes with human intervention, leading to inconvenient manipulation, with potentially fatal impacts on downstream tasks in unforeseen scenes. We propose a real-time automatic hyperparameter optimization ISP hardware system to address real-world scenarios. Our design features a tri-step framework and a hardware accelerator, demonstrating superior performance in human and computer vision tasks, even in real-time unforeseen scenes. Experiments showcase its practicality, achieving 1080P@75FPS/240FPS in FPGA/ASIC, respectively.
Zihao Liu 0015, Ruoxi Zhu, Qi Zheng 0004, Zhijian Hao, Tao Liu 0023, Jun Tao 0001, Yibo Fan
DAC8
2024 FPIA: Communication-Aware Multi-Chiplet Integration With Field-Programmable Interconnect Fabric on Reusable Silicon Interposer
abstract
Silicon interposer re-usage is drawing attention for cost-effective multi-chiplet integrated systems. To address the communication awareness of inter/off-chiplet interconnect, the paper proposes a field-programmable interconnect fabric and develops its corresponding automatic physical integration tool. The tile-based fabric consists of turnout, cross-over boxes and parallel tracks. It features micro-bump-wise connecting flexibility and hardware efficiency. The automation flow performs chiplet location optimization and efficient bump-to-bump routing, supporting multi-lane bus interconnect and miscellaneous external ports. The methodology is validated by 9 different integration scenarios, where the routability is guaranteed when the local resource utilization ratio approaches 94.5%. The data’s maximum interconnect latency is 2.2 ns and the energy consumption is 1.18 pJ/bit at a bitrate of 1 Gbps. The latency consumes$16.5\times \sim ~53.4\times $fewer clock cycles than the state-of-the-art network-on-package-based reusable interposer architectures.
Bo Jiao 0003, Haozhe Zhu, Jundong Zhu, Dexin Wen, Lingli Wang, Jun Tao 0001, Chixiao Chen, Yinhe Han 0001, Qi Liu 0010, Ninghui Sun, Ming Liu 0022
IEEE Trans. Circuits Syst. I Regul. Pap.10
2024 Yield Optimization for Analog Circuits over Multiple Corners via Bayesian Neural Networks: Enhancing Circuit Reliability under Environmental Variation
abstract
The reliability of circuits is significantly affected by process variations in manufacturing and environmental variation during operation. Current yield optimization algorithms take process variations into consideration to improve circuit reliability. However, the influence of environmental variations (e.g., voltage and temperature variations) is often ignored in current methods because of the high computational cost. In this article, a novel and efficient approach named BNN-BYO is proposed to optimize the yield of analog circuits in multiple environmental corners. First, we use a Bayesian Neural Network (BNN) to simultaneously model the yields and performances of interest in multiple corners efficiently. Next, the multi-corner yield optimization can be performed by embedding BNN into a Bayesian optimization framework. Since the correlation among yields and performances of interest in different corners is implicitly encoded in the BNN model, it provides great modeling capabilities for yields and their uncertainties to improve the efficiency of yield optimization. Our experimental results demonstrate that the proposed method can save up to 45.3% of simulation cost compared to other baseline methods to achieve the same target yield. In addition, for the same simulation cost, our proposed method can find better design points with 3.2% yield improvement.
Nanlin Guo, Fulin Peng, Jiahe Shi, Fan Yang 0001, Jun Tao 0001, Xuan Zeng 0001
ACM Trans. Design Autom. Electr. Syst.5
2023 TPNoC: An Efficient Topology Reconfigurable NoC Generator
abstract
With the core count increasing in Chip to support various data-intensive workloads, Network-on-chip (NoC) has become the better solution for addressing on-chip interconnection. Various data-intensive workloads have different traffic patterns that require NoC with different topologies and microarchitectures. On the one hand, topology type selection has a great influence on the final performance, area, and energy. However, it is difficult to change the topology type in the traditional NoC RTL design process once it is determined. On the other hand, NoC platforms have many tunable micro-architecture design parameters, which require careful design space exploration to trade off performance advantages and overhead. Designing and validating each microarchitecture of NoCs to account for various trade-offs will greatly exacerbate the design cost issue.
Jiangnan Yu, Fan Yang 0001, Xiaoling Yi, Chixiao Chen, Jun Tao 0001, Dong Xu 0015, Xiankui Xiong
ACM Great Lakes Symposium on VLSI5
2023 A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D Integration
abstract
Chiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost.
Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen
ISCAS6
2023 Correlated Bayesian Model Fusion: Efficient High-Dimensional Performance Modeling of Analog/RF Integrated Circuits Over Multiple Corners
abstract
Efficient high-dimensional performance modeling of analog/RF circuits over multiple corners is an important-yet-challenging task. In this article, we propose a novel performance modeling approach for analog/RF circuits, referred to as correlated Bayesian model fusion (C-BMF). The key idea is to encode the correlation information for both model template and coefficient magnitude among different corners by using a unified prior distribution. Next, the prior distribution is combined with a few simulation samples via Bayesian inference to efficiently determine the unknown model coefficients. Two circuit examples designed in a commercial 40-nm CMOS process demonstrate that C-BMF achieves about$2\times $cost reduction over the traditional state-of-the-art modeling technique without surrendering any accuracy.
Zhengqi Gao, Fa Wang, Jun Tao 0001, Yangfeng Su, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Self-Supervised On-Device Federated Learning From Unlabeled Streams
abstract
The ubiquity of edge devices has led to a growing amount of unlabeled data produced at the edge. Deep learning models deployed on edge devices are required to learn from these unlabeled data to continuously improve accuracy. Self-supervised representation learning has achieved promising performances using centralized unlabeled data. However, the increasing awareness of privacy protection limits centralizing the distributed unlabeled image data on edge devices. While federated learning has been widely adopted to enable distributed machine learning with privacy preservation, without a data selection method to efficiently select streaming data, the traditional federated learning framework fails to handle these huge amounts of decentralized unlabeled data with limited storage resources on edge. To address these challenges, we propose a self-supervised on-device federated learning framework with coreset selection, which we call SOFed, to automatically select a coreset that consists of the most representative samples into the replay buffer on each device. It preserves data privacy as each client does not share raw data while learning good visual representations. Experiments demonstrate the effectiveness and significance of the proposed method in visual representation learning.
Jiahe Shi, Yawen Wu, Dewen Zeng, Jun Tao 0001, Jingtong Hu, Yiyu Shi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Efficient Statistical Parameter Extraction for Modeling MOSFET Mismatch
abstract
In this article, we propose an efficient statistical parameter extraction method to accurately model the random device mismatch of MOSFETs. The key idea is to approximate the performance variations as mathematical functions of device mismatch. Based on these approximated functions and the electrical test data, we solve the unknown statistical parameters by nonlinear optimization. Our numerical experiments demonstrate that the proposed method can remarkably improve the modeling accuracy with affordable computational cost, compared against the state-of-the-art techniques.
Nanlin Guo, Nengyong Zhu, Jun Tao 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 An Automated Compiler for RISC-V Based DNN Accelerator
abstract
Multifarious hardware accelerators are developed for the widely used Deep neural networks (DNN). Nowadays the SoCs composed of a general processor and a coupled accelerator are becoming prevalent. Compared to the specialized DNN accelerator for one specific DNN, this kind of coupled architecture is programmable and supports diverse DNNs. However, for the low-level programming interface of the co-processor-like accelerator and the multi-hierarchy memory structure, programming for the DNN accelerator is not easy work. Meanwhile, there are a couple of tensor compilers that deploy the DNN on various hardware. In this work, we combine the flexibility of the tensor compiler and the high efficiency of the hardware accelerator by proposing an automated compiler that can compile tensor programs and generate high-performance programs for programmable DNN accelerators. Our compiler is based on TVM [1] and target at Rocket Chip Coprocessor (RoCC) [2]. The compiler is flexible and supports many kinds of RISC-V instructions. The programmer can define the hardware constraints in the proposed compiler which makes the generated code more efficient. Our compiler can lower the program with the ping-pong strategy and the generated code can achieve 26% speed up compared to the baseline.
Wuzhen Xie, Xiaoling Yi, Ruiyao Pu, Xiankui Xiong, Haidong Yao, Chixiao Chen, Jun Tao 0001, Fan Yang 0001
ISCAS9
2022 NNASIM: An Efficient Event-Driven Simulator for DNN Accelerators with Accurate Timing and Area Models
abstract
In this paper, we propose NNASIM, an efficient timing and area accurate event-driven simulator for custom DNN accelerators. NNASIM is a highly-modular and highly parameterized modeling framework. We build accurate timing and area models for common accelerator modules like GEMM, ALU array, and crossbar using ASIC synthesis flows. These models are fed into the event-driven simulator for fast simulation. NNASIM is integrated with a RISC-V simulator. This approach guarantees the functional correctness of the accelerator simulation at the instruction level. The experimental results show that our model evaluates the performance and area of DNN accelerators with less than 0.76% and 2.83% error, respectively, compared to RTL implementations. NNASIM allows designers to model the performance and area of the accelerator at a high level, and thus enables the systematic microarchitecture design space exploration of the custom accelerators. Index Terms accelerators.
Xiaoling Yi, Jiangnan Yu, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Jun Tao 0001, Fan Yang 0001
ISCAS7
2022 Unsupervised deep domain adaptation framework in remote acoustic time parametric imaging
abstract
Abstract Intelligent waveform analysis surveying geological structures is a challenging task in remote acoustic measurement for formation detection, which has two problems: (1) time parametric imaging is disturbed by noisy environments and (2) manually annotated data for machine learning are unattainable. These restrict the deployment of advanced parameter extraction methods in imaging instruments. As a potential theory, domain adaptation makes the intelligent prediction implementable in the above situations. Hence, to counter high precision travel time extraction for acoustic imaging, a deep adaptation‐picking network (DAPN) framework is proposed, which consists of three modules: signal‐to‐signal generator, domain adaptation encoder, and feature recognition decoder. The translation module preprocesses the different datasets to improve the training accuracy and confuses the distribution between source and target domains based on the adversarial network. The core convolution neural network achieves travel time measurement, where the backbone extracts content features with modified maximum mean discrepancy embedding. Meanwhile, the decoding structure preserves the signal features in the training process. The effectiveness and feasibility of DAPN are verified by experimental results, suggesting that DAPN accurately extracts time parameters. Compared with traditional parametric imaging schemes, the proposed method has the advantages of high accuracy and anti‐noise ability.
Yuezhou Wu, Jun Tao 0001
IET Signal Process.2
2022 Fast Statistical Analysis of Rare Failure Events With Truncated Normal Distribution in High-Dimensional Variation Space
abstract
In this article, to accurately estimate the rare failure rates for large-scale circuits (e.g., SRAM) where process variations are modeled as truncated normal distributions in high-dimensional space, we propose a novel truncated scaled-sigma sampling (T-SSS) method. Similar to scaled-sigma sampling (SSS), T-SSS distorts the truncated normal distributions by a scaling factor, resulting in an analytical model for failure rate estimation. By drawing random samples from the distorted distribution and estimating a sequence of scaled failure rates, we can solve all unknown model coefficients and predict the original failure rate by extrapolation. The accuracy of T-SSS is further assessed by estimating its confidence interval (CI) based on resampling. Our numerical results demonstrate that the proposed T-SSS method can achieve superior accuracy over the state-of-the-art method without increasing the computational cost.
Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Correlated Rare Failure Analysis via Asymptotic Probability Evaluation
abstract
In this article, a novel asymptotic probability evaluation (APE) method is proposed to estimate the probability of correlated rare failure events for complex integrated systems containing a large number of replicated cells. The key idea is to approximate the failure rate of the entire system by solving a set of nonlinear equations derived from a general analytical model. An error refinement method based on look-up table is further developed to improve numerical stability and, hence, reduce estimation error. Furthermore, a statistical algorithm based on resampling is developed to accurately estimate the confidence interval of APE. Our numerical experiments demonstrate that compared to the state-of-the-art method, APE can reduce the estimation error by up to$30\times $without increasing the computational cost.
Jun Tao 0001, Handi Yu, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Adaptable Approximate Multiplier Design Based on Input Distribution and Polarity
abstract
Approximate computing is an efficient approach to reduce the design complexity for error-resilient applications. Multipliers are key arithmetic units in many applications, such as deep neural networks (DNNs) and digital signal processing (DSP) systems. In this article, an open-source adaptable approximate multiplier design driven by input distribution and polarity is proposed to generate optimized approximate multipliers to trade off between the application-level performance and the hardware cost. The proposed method minimizes the average square of the absolute error of an approximate multiplier according to the probability distributions of operands extracted from the target application with consideration of input polarity, achieving low hardware cost and negligible application-level performance loss. The proposed method can generate unsigned multipliers (or signed multipliers) based on the Braun multiplier (or Baugh–Wooley multiplier). To demonstrate the effectiveness of the method, three different-scale quantized DNNs, including LeNet, AlexNet, and VGG16 with 8$\times $8 unsigned multiplication and an adaptive least mean square (LMS)-based finite impulse response (FIR) filter with 16$\times $16 fixed-point signed multiplication, are evaluated. In the DNN training process, a noise training technique is adopted to reduce the accuracy loss due to the approximation. When compared to the state-of-the-art approximate multipliers, the generated multipliers can achieve up to 26.4% and 27.1% product of power, delay, and area gains with negligible application-level performance loss in VGG16 and FIR applications, respectively.
Zhen Li 0059, Su Zheng, Jide Zhang, Jingbo Gao, Jun Tao 0001, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.6
2021 Bayesian Inference on Introduced General Region: An Efficient Parametric Yield Estimation Method for Integrated Circuits
abstract
In this paper, we propose an efficient parametric yield estimation method based on Bayesian Inference. By observing that nowadays analog and mixed-signal circuit is designed via a multi-stage flow, and that the circuit performance correlation of early stage and late stage is naturally symmetrical, we introduce a general region to capture the common features of the early and late stage. Meanwhile, two private regions are also incorporated to represent the unique features of these two stages respectively. Afterwards, we introduce classifiers one for each region to explicitly encode the correlation information. Next, we set up a graphical model, and consequently adopt Bayesian Inference to calculate the model parameters. Finally, based on the obtained optimal model parameters, we can accurately and efficiently estimate the parametric yield with a simple sampling method. Our numerical experiments demonstrate that compared to the state-of-the-art algorithms, our proposed method can better estimate the yield while significantly reducing the number of circuit simulations.
Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001
ASP-DAC3
2020 Exploring a Bayesian Optimization Framework Compatible with Digital Standard Flow for Soft-Error-Tolerant Circuit
abstract
Soft error is a major reliability concern in advanced technology nodes. Although mitigating Soft Error Rate (SER) will inevitably sacrifice area and power, few studies paid attention to optimization methods to explore trade-offs between area, power and SER. This paper proposes an optimization framework based on Bayesian approach for soft-error-tolerant circuit design. It comprises two steps:1) data preprocessing and 2) Bayesian optimization. In the preprocessing step, a strategy incorporating k-means algorithm and a novel sequencing algorithm is used to cluster Flip-Flops (FFs) with similar SER in order to reduce the dimensionality for the subsequent step. Bayesian Neural Network (BNN) is the applied surrogate model for acquiring the posterior distribution of three design metrics, while the Lower confidence bound (LCB) functions are employed as acquisition functions to select the next point based on BNN when optimizing. Finally, the non-dominated sorting genetic algorithm (NSGA-II) is used to search the Pareto Optimal Front (POF) solutions of three LCB functions. Experimental results demonstrate the proposed framework has a 1.4x improvement in accuracy and a 70% reduction in SER with acceptable increases in power and area.
Yan Li 0084, Xiaoyoung Zeng, Zhengqi Gao, Liyu Lin, Jun Tao 0001, Jun Han 0003, Xu Cheng 0002, Mehdi Baradaran Tahoori, Xiaoyang Zeng
DAC5
2020 Multi-Corner Parametric Yield Estimation via Bayesian Inference on Bernoulli Distribution with Conjugate Prior
abstract
To efficiently estimate parametric yields over multiple process, voltage, temperature corners for binary output circuits, we propose a novel Bayesian Inference method based on Bernoulli distribution with conjugate prior in this paper. The key idea is to adopt a product of Beta distributions as the conjugate prior for the yields and encode circuit performance correlations among different corners into this prior. Next, the hyper-parameters are optimized by using multi-start Quasi-Newton method, and the yields over different corners are estimated via maximum-a-posteriori. Two circuit examples demonstrate that the proposed method achieves up to 3.0× cost reduction over the state-of-the-art methods without surrendering any accuracy.
Jiahe Shi, Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001
ISCAS3
2020 Efficient Rare Failure Analysis Over Multiple Corners via Correlated Bayesian Inference
abstract
In this article, we propose an efficient correlated Bayesian inference (CBI) method to estimate the system-level failure rates for large-scale circuit systems over multiple process corners. The key idea is to encode the correlations of circuit performances among the different corners into the prior distributions of several carefully defined failure events. The hyper-parameters of these distributions can be learned from a few simulation samples via Bayesian inference and, next, the system-level failure rates over different corners can be simultaneously estimated by taking into account these prior distributions. An iteratively constrained inference method is further developed to guarantee the numerical stability of the proposed method and legalize all estimated failure rates. The numerical experiments demonstrate that compared to the state-of-the-art algorithm, the proposed method can achieve around 10× runtime reduction without surrendering any accuracy.
Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Efficient Parametric Yield Estimation Over Multiple Process Corners via Bayesian Inference Based on Bernoulli Distribution
abstract
Parametric yield estimation over multiple process corners plays an important role in robust circuit design. In this article, we propose a novel Bayesian inference method based on Bernoulli distribution (BI-BD) to efficiently estimate the multicorner yields for binary output circuit. The key idea is to encode the circuit performance correlation among different corners as our prior knowledge. Consequently, after combining a few simulation samples, the yield estimation over all corners can be calibrated via Bayesian inference based on iterative reweighted least squares (IRLS) and expectation maximization (EM). A circuit example demonstrates that the proposed BI-BD method can achieve up to 2.0 × cost reduction over the conventional Monte Carlo method without surrendering any accuracy.
Zhengqi Gao, Jun Tao 0001, Dian Zhou, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Efficient Statistical Analysis for Correlated Rare Failure Events via Asymptotic Probability Approximation
abstract
In this article, a novel asymptotic probability approximation (APA) method is proposed to estimate the overall rare probability of correlated failure events for complex circuits containing a large number of replicated cells (e.g., SRAM bit-cells). The key idea of APA is to approximate the overall circuit failure rate based on a set of carefully defined failure events. An efficient hierarchical subset simulation (H-SUS) method is developed to calculate the aforementioned failure rate and a statistical methodology is further proposed to estimate the confidence interval of APA. Our numerical experiments demonstrate that APA can accurately and reliably estimate the overall failure rate of correlated rare failure events involving more than 20 000 independent random variables.
Fulin Peng, Handi Yu, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 A Synthesizable Constant Tuning Gain Technique for Wideband LC-VCO Design
abstract
In this paper, an iterative method is proposed for the optimization of wideband voltage-controlled oscillators (VCOs). The proposed method attempts to improve two significant performance metrics of wideband VCOs, i.e., the nonuniform band distribution and the nonconstant tuning gain KVCO. To address these nonideality issues, we introduce the deviation as a prior knowledge for calibration in each iteration step. The optimal capacitor array weighting plan can be asymptotically approached through multiple iterations. To demonstrate the efficacy of the proposed method, we use it to implement a 1-to-2.4-GHz LC-VCO and achieve a frequency tuning range of 82.4%. The optimized LC-tank involves 64 mixed cells, and each cell contains 2-bit binary-weighted switching capacitors, leading to 256 tuning bands. The LC-VCO with the synthesized tank achieves uniform frequency steps around 5 MHz with the relative error within ±4.4%. Meanwhile, the tuning gain KVCO is around 10 MHz/V with the relative error within ±3.2%. For this LC-VCO design, each iteration step costs about 55.9 s. After seven iterations, we approach the convergent solution where both the relative errors of the frequency steps and the tuning gains satisfy the given specifications.
Xiaoxue Sun, Chenyang Kong, Jun Tao 0001, Zhangwen Tang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Analog/RF Post-silicon Tuning via Bayesian Optimization
abstract
Tunable analog/RF circuit has emerged as a promising technique to address the significant performance uncertainties caused by process variations. To optimize these tunable circuits after fabrication, most existing post-silicon programming methods are developed by using real-valued performance metrics. However, when measuring a performance of interest on silicon, it is often substantially more expensive to obtain a real-valued measurement than a binary testing outcome (i.e., pass or fail). In this article, we propose a Gaussian Process Classification model to capture the binary performance metrics of tunable analog/RF circuits. Based on these models, post-silicon programming is cast into an optimization problem that can be solved by a novel Bayesian optimization algorithm. Moreover, measurement noises are further incorporated into our proposed post-silicon programming to produce a robust circuit. Two circuit examples demonstrate that the proposed approach can efficiently program tunable circuits with binary performance metrics while other conventional methods are not applicable.
Renjian Pan, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
ACM Trans. Design Autom. Electr. Syst.2
2019 Efficient Performance Trade-off Modeling for Analog Circuit based on Bayesian Neural Network
abstract
In this paper, we propose an efficient performance trade-off modeling method for analog circuit based on Bayesian Neural Network (BNN). First, we use a single BNN to simultaneously model multiple performances of interest (PoIs) of an analog circuit. This BNN model can be trained by using a novel automatic differential variational inference (ADVI) method with affordable computational cost. Next, the performance trade-off model can be extracted by embedding BNN into Bayesian optimization framework combined with a modified multi-objective evolutionary method. Since the correlations among different PoIs are implicitly encoded in the BNN model, the proposed method can capture the performance trade-off model efficiently and accurately. The numerical experiments demonstrate that compared to the state-of-the-art algorithms, the proposed method can achieve up to 2× runtime reduction without surrendering any accuracy.
Zhengqi Gao, Jun Tao 0001, Fan Yang 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001
ICCAD2
2019 Graph-Constrained Sparse Performance Modeling for Analog Circuit Optimization via SDP Relaxation
abstract
In this paper, a graph-constrained sparse performance modeling method is proposed for analog circuit optimization. It builds sparse polynomial models constrained by an acyclic graph. These models can be used to solve analog optimization problems within local design spaces by using convex semidefinite programming relaxation both efficiently and robustly. Our numerical examples demonstrate that the proposed modeling and optimization method can quickly and accurately converge to a superior solution for analog circuits while the conventional method fails to work.
Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Correlated Rare Failure Analysis via Asymptotic Probability Evaluation
abstract
In this paper, a novel Asymptotic Probability Estimation (APE) method is proposed to estimate the probability of correlated rare failure events for complex integrated systems containing a large number of replicated cells. The key idea is to approximate the failure rate of the entire system by solving a set of nonlinear equations derived from a general analytical model. An error refinement method based on Look-up Table (LUT) is further developed to improve numerical stability and, hence, reduce estimation error. Our numerical experiments demonstrate that compared to the state-of-the-art method, APE can reduce the estimation error by up to 45x without increasing the computational cost.
Jun Tao 0001, Handi Yu, Dian Zhou, Yangfeng Su, Xuan Zeng 0001, Xin Li 0001
DAC1
2017 Efficient programming of reconfigurable radio frequency (RF) systems
abstract
Reconfigurable radio frequency (RF) system has recently emerged as a promising solution to cope with multiple communication standards and high spectrum density. In this paper, we propose a novel optimization framework to efficiently program a reconfigurable RF system. In particular, two novel techniques, including (i) search space reduction by adaptive resolution and (ii) global polynomial optimization based on branch and bound, are developed. When combined with a relaxation iteration scheme, our proposed method offers superior performance when programming a large-scale reconfigurable RF system designed for the WLAN 802.11g standard.
Mohamed Baker Alawieh, Fa Wang, Jun Tao 0001, Shihui Yin, Minhee Jun, Xin Li 0001, Tamal Mukherjee, Rohit Negi
ICCAD3
2016 Efficient spatial variation modeling via robust dictionary learning
Changhai Liao, Jun Tao 0001, Xuan Zeng 0001, Yangfeng Su, Dian Zhou, Xin Li 0001
DATE2
2016 Efficient statistical analysis for correlated rare failure events via asymptotic probability approximation
abstract
In this paper, a novel Asymptotic Probability Approximation (APA) method is proposed to estimate the overall rare probability of correlated failure events for complex circuits containing a large number of replicated cells (e.g., SRAM bit-cells). The key idea of APA is to approximate the overall circuit failure rate based on a set of carefully defined failure events. An efficient Hierarchal Subset Simulation (H-SUS) method is developed to calculate the aforementioned failure rate and a statistical methodology is further proposed to estimate the confidence interval of APA. Our numerical experiments demonstrate that APA can accurately and reliably estimates the overall failure rate of correlated rare failure events involving more than 20,000 independent random variables.
Handi Yu, Jun Tao 0001, Changhai Liao, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
ICCAD2
2016 Efficient Hybrid Performance Modeling for Analog Circuits Using Hierarchical Shrinkage Priors
abstract
Efficient performance modeling is an extremely important task for yield analysis and design optimization of analog circuits. In this paper, a novel regression modeling method based on hierarchical shrinkage priors is proposed to construct hybrid performance models with both high accuracy and low computational cost. In particular, the user-defined model templates derived from design equations and the general-purpose orthogonal polynomials are combined together to set up a hybrid dictionary. Next, in order to avoid over-shrinking large model coefficients, a novel regression method based on hierarchical shrinkage priors and variational Bayesian inference is adopted for model fitting. A rail-to-rail operational amplifier example demonstrates that the proposed method achieves up to 40% error reduction over other state-of-the-art approaches without increasing the modeling cost.
Changhai Liao, Jun Tao 0001, Handi Yu, Zhangwen Tang, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Efficient Spatial Variation Modeling of Nanoscale Integrated Circuits Via Hidden Markov Tree
abstract
In this paper, we propose a novel spatial variation modeling method based on hidden Markov tree (HMT) for nanoscale integrated circuits, which could efficiently improve the accuracy of full-wafer/chip spatial variations recovery at extremely low measurement cost. Applying this method, HMT is introduced to set up a statistical model for coefficients after exploring the underlying correlated representation of the spatial variation in the frequency domain. Accordingly, two key inherent properties of the modeling coefficients, i.e., correlations and sparse presentations in the frequency domain, can be captured exactly and the modeling accuracy can be improved evidently. Then, maximum-a-posteriori estimation is applied to formulate the original problem as a convex optimization that could be solved efficiently and robustly. Numerical results based on industrial data demonstrate that the proposed method can achieve superior accuracy over other existing approaches including orthogonal matching pursuit, l1-norm regularization, and reweighted l1-norm regularization.
Changhai Liao, Jun Tao 0001, Xuan Zeng 0001, Yangfeng Su, Dian Zhou, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Harvesting Design Knowledge From the Internet: High-Dimensional Performance Tradeoff Modeling for Large-Scale Analog Circuits
abstract
Efficiently optimizing large-scale, complex analog systems requires to know the performance tradeoffs for various analog circuit blocks. In this paper, we propose a radically new approach for analog performance tradeoff modeling. Our key idea is to broadly search the rich design knowledge from the Internet, and then mathematically encode the knowledge as high-dimensional performance tradeoff curves that are referred to as Pareto fronts in the literature. Toward this goal, several novel numerical algorithms, such as sparse regression and semi-infinite programming, are developed in order to construct the high-dimensional Pareto front model while guaranteeing its monotonicity. Our numerical examples demonstrate that the proposed modeling technique can accurately capture the high-dimensional Pareto fronts for large-scale analog systems (e.g., analog-to-digital converter) while most traditional methods are limited to low-dimensional Pareto front modeling of small circuit blocks without considering layout parasitics and manufacturing nonidealities.
Jun Tao 0001, Changhai Liao, Xuan Zeng 0001, Xin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 Toward efficient programming of reconfigurable radio frequency (RF) receivers
abstract
Reconfigurable radio frequency (RF) system is an emerging component to mitigate the growing engineering cost for wireless chip design. In this paper, we propose a new methodology for efficient programming of reconfigurable RF receiver. The proposed method is facilitated by two novel techniques: two-phase relaxation search and Pareto-based search space reduction. Our numerical experiments demonstrate that the proposed methodology is more robust (i.e., close to global optimum) and/or efficient (i.e., with low computational cost) than other traditional algorithms based on either local relaxation or simulated annealing.
Jun Tao 0001, Ying-Chih Wang, Minhee Jun, Xin Li 0001, Rohit Negi, Tamal Mukherjee, Lawrence T. Pileggi
ASP-DAC1
2014 Integrated Algorithm for 3-D IC Through-Silicon Via Assignment
abstract
Through-silicon via (TSV) with flip-chip packaging is a technology that enables vertical integration of silicon dies, forming a single 3-D IC stack. A practical model for preplaced TSV assignment of 3-D nets is proposed for this technology. We prove that the general preplaced 3-D IC TSV assignment problem with more than two dies is NP-complete. An integrated algorithm that combines shortest path search, bipartite matching, min-cost max-flow calculation, and postprocessing is developed. Experimental results using actual testing silicon data demonstrate that our flow achieves good results with reasonable runtime when compared to other existing works.
Xiaodong Liu 0018, Gary K. Yeap, Jun Tao 0001, Xuan Zeng 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2013 An efficient method for gradient-aware dummy fill synthesis
Hai Zhou 0001, Changhao Yan, Jun Tao 0001, Xuan Zeng 0001
Integr.4
2011 Efficient Approximation Algorithms for Chemical Mechanical Polishing Dummy Fill
abstract
To reduce chip-scale topography variation in chemical mechanical polishing process, dummy fill is widely used to improve the layout density uniformity. Previous researches formulated the density-driven dummy fill problem as a standard linear program (LP). However, solving the huge linear program formed by real-life designs is very expensive and has become the hurdle in deploying the technology. Even though there exist efficient heuristics, their performance cannot be guaranteed. Furthermore, dummy fill can also change the interconnect coupling capacitance which might lead to a significant influence on circuit delay, crosstalk, and power consumption. In this paper, we develop a dummy fill algorithm that can be applied to solve both the traditional density-driven problem and the problem considering fill-induced coupling capacitance impact. The proposed algorithm is both efficient and with provably good performance, which is based on a fully polynomial time approximation scheme by Fleischer for covering LP problems. Moreover, based on the approximation algorithm, we also propose a new greedy iterative algorithm to achieve high quality solutions more efficiently than previous Monte Carlo based heuristic methods. Final experimental results demonstrate the effectiveness and efficiency of our algorithms.
Chunyang Feng, Hai Zhou 0001, Changhao Yan, Jun Tao 0001, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2011 Binning Optimization for Transparently-Latched Circuits
abstract
With increasing process variation, binning has become an important technique to improve the values of fabricated chips, especially in high performance microprocessors where transparent latches are widely used. In this paper, we formulate and solve the binning optimization problem that decides the bin boundaries and their testing order to maximize the profit (considering the test cost) for a transparently-latched circuit. The problem is decomposed into four sub-problems. First, to compute the clock period distribution of the transparently-latched circuit, a sample-based statistical static timing analysis (SSTA) approach is developed which is based on the generalized stochastic collocation method with the sparse grid technique. The minimal clock period on each sample point is found by solving a minimal cycle ratio problem in the constraint graph. Second, a greedy method is proposed to maximize profit considering both the sales revenue and the test cost by iteratively assigning each boundary to its optimal position. Third, an optimal algorithm of O(n log n) runtime is used to generate the optimal testing order to minimize the test cost, based on alphabetic tree. Last, a simple approach is presented to decide the optimal number of bins, which helps to complete the whole binning scheme with maximal profit. Experiments on all the ISCAS'89 sequential benchmarks with 65 nm technology show 10.68% profit improvement in average. Some comparisons with other methods suggest the advantage of our method. The results also demonstrate that the proposed SSTA method achieves an error of 0.70% and speedup of 110X in average compared with the Monte Carlo simulation.
Hai Zhou 0001, Li Li 0021, Jun Tao 0001, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 Provably good and practically efficient algorithms for CMP dummy fill
abstract
Abstract—To reduce chip-scale topography variation in Chemical Me-chanical Polishing (CMP) process, dummy fill is widely used to improve the layout density uniformity. Previous researches formulated the dummy fill problem as a standard Linear Program (LP). However, solving the huge linear program formed by real-life designs is very expensive and has become the hurdle in deploying the technology. Even though there exist efficient heuristics, their performance cannot be guaranteed. In this paper, we develop a dummy fill algorithm that is both efficient and with provably good performance. It is based on a fully polynomial time approximation scheme by Fleischer [4] for covering LP problems. Furthermore, based on the approximation algorithm, we also propose a new greedy iterative algorithm to achieve high quality solutions more efficiently than previous Monte-Carlo based heuristic methods. Experimental results demonstrate the effectiveness and efficiency of our algorithms.
Chunyang Feng, Hai Zhou 0001, Changhao Yan, Jun Tao 0001, Xuan Zeng 0001
DAC4
2009 Binning optimization based on SSTA for transparently-latched circuits
abstract
With increasing process variation, binning has become an important technique to improve the values of fabricated chips, especially in high performance microprocessors where transparent latches are widely used. In this paper, we formulate and solve the binning optimization problem that decides the bin boundaries and their testing order to maximize the benefit (considering the test cost) for a transparently-latched circuit. The problem is decomposed into three sub-problems which are solved sequentially. First, to compute the clock period distribution of the transparently-latched circuit, a sample-based SSTA approach is developed which is based on the generalized stochastic collocation method (gSCM) with Sparse Grid technique. The minimal clock period on each sample point is found by solving a minimal cycle ratio problem in the constraint graph. Second, a greedy algorithm is proposed to maximize the sales profit by iteratively assigning each boundary to its optimal position. Then, an optimal algorithm of O(n log n) runtime is used to generate the optimal testing order of bin boundaries to minimize the test cost, based on alphabetic tree. Experiments on all the ISCAS'89 sequential benchmarks with 65-nm technology show 6.69% profit improvement and 14.00% cost reduction in average. The results also demonstrate that the proposed SSTA method achieves an error of 0.70% and speedup of 110X in average compared with the Monte Carlo simulation.
Hai Zhou 0001, Jun Tao 0001, Xuan Zeng 0001
ICCAD3
2008 Timing yield driven clock skew scheduling considering non-Gaussian distributions of critical path delays
abstract
In nanometer technologies, process variations possess growing nonlinear impacts on circuit performance, which causes critical path delays of combinatorial circuits variate randomly with non-Gaussian distribution. In this paper, we propose a novel clock skew scheduling methodology that optimizes timing yield by handling non-Gaussian distributions of critical path delays. Firstly a general formulation of the optimization problem is proposed, which covers most of the previous formulations and indicates their limitations with statistical interpretations. Then a generalized minimum balancing algorithm is proposed for effectively solving the skew scheduling problem. Experimental results show that the proposed method significantly outperforms some representative methods previously proposed for yield optimization, and could obtain timing yield improvements up to 33.6% and averagely 17.7%.
Wai-Shing Luk, Xuan Zeng 0001, Jun Tao 0001, Changhao Yan, Jiarong Tong, Wei Cai 0003, Jia Ni
DAC4
2007 Stochastic Sparse-grid Collocation Algorithm (SSCA) for Periodic Steady-State Analysis of Nonlinear System with Process Variations
abstract
In this paper, stochastic collocation algorithm combined with sparse grid technique (SSCA) is proposed to deal with the periodic steady-state analysis for nonlinear systems with process variations. Compared to the existing approaches, SSCA has several considerable merits. Firstly, compared with the moment-matching parameterized model order reduction (PMOR), which equally treats the circuit response on process variables and frequency parameter by Taylor approximation, SSCA employs homogeneous chaos to capture the impact of process variations with exponential convergence rate and adopts Fourier series or wavelet bases to model the steady-state behavior in time domain. Secondly, contrary to stochastic Galerkin algorithm (SGA), which is efficient for stochastic linear system analysis, the complexity of SSCA is much smaller than that of SGA for nonlinear case. Thirdly, different from efficient collocation method, the heuristic approach which may results in "rank deficient problem" and "Runge phenomenon", sparse grid technique is developed to select the collocation points in SSCA in order to reduce the complexity while guaranteing the approximation accuracy. Furthermore, though SSCA is proposed for the stochastic nonlinear steady-state analysis, it can be applied for any other kinds of nonlinear system simulation with process variations, such as transient analysis, etc.
Jun Tao 0001, Xuan Zeng 0001, Wei Cai 0003, Yangfeng Su, Dian Zhou, Charles C. Chiang
ASP-DAC1
2006 A one-shot projection method for interconnects with process variations
abstract
With the development of IC technology, it becomes urgent to investigate model reduction method for interconnects with process variations. In this paper, a one-shot projection algorithm (OPM) is proposed to generate a projection matrix that is independent of statistically varying parameters. As a result, construction of the reduced system can be decoupled with the Monte Carlo analysis in either frequency domain or time domain. Therefore, without loss of accuracy, OPM can obtain a reduced system in much less CPU time compared with the previous perturbation scheme. Numerical results have demonstrated the advantages of the proposed OPM
Jun Tao 0001, Xuan Zeng 0001, Fan Yang 0001, Yangfeng Su, Lihong Feng, Wei Cai 0003, Dian Zhou, Charles C. Chiang
ISCAS1
2005 Block SAPOR: block Second-order Arnoldi method for Passive Order Reduction of multi-input multi-output RCS interconnect circuits
abstract
Recently model order reduction techniques for second-order systems have obtained many research interests for the simulation of RCS interconnect circuits employing susceptance elements. In this paper, we propose a Block SAPOR (Block Second-order Arnoldi method for Passive Order Reduction) for Multi-Input Multi-Output RCS Circuits. The proposed Block SAPOR algorithm can simultaneously guarantee passivity and achieve higher accuracy than the first order reduction technique PRIMA. Most importantly, the reduced system matrices obtained by the proposed method can preserve the structure of the original system matrices. Such a nice property makes it possible to construct an equivalent RCS circuit for the reduced system.
Xuan Zeng 0001, Yangfeng Su, Jun Tao 0001, Zhaojun Bai, Charles C. Chiang, Dian Zhou
ASP-DAC4
2005 A novel wavelet method for noise analysis of nonlinear circuits
abstract
In this paper, a novel wavelet method is proposed for noise analysis of nonlinear circuits. Compared with the existing algorithms capable of accessing circuit performance in the present of noise, the proposed method presents several merits. First, it fully accounts for nonlinearities. Second, it can handle signals with continuous frequency spectra. Third, by taking advantage of the properties of the wavelet bases, such as local compactness and multi-resolution, it holds high simulation speed and high accuracy. Furthermore, an adaptive scheme exists to automatically select the wavelet basis functions for a desired accuracy. All these merits make the novel wavelet method outperforms its previous techniques.
Xuan Zeng 0001, Jun Tao 0001, Charles C. Chiang, Dian Zhou
ASP-DAC3
2004 Analog circuit behavioral modeling via wavelet collocation method with auto-companding
Jun Tao 0001, Xuan Zeng 0001, Charles C. Chiang, Dian Zhou
ASP-DAC2