EDBT 2026 Demo / reviewers in the wild / expert
Xin Li 0001
dblp:09/1365-1
· DBLP profile ↗
190ranked-venue papers
30as first author
28since 2021 · last 2026
0000-0002-4510-2436ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 159 · 27 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Software engineering, systems software and programming languages · 5Computer networks · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image EditingabstractText-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification. Kaixiang Yang 0004, Boyang Shen, Xin Li 0001, Yuchen Dai, Yueran Ma, Qiang Li 0018, Zhiwei Wang 0002 |
AAAI | 3 |
| 2025 | Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional DiffusionabstractDual-view mammography, including craniocaudal (CC) and mediolateral oblique (MLO) projections, offers complementary anatomical views crucial for breast cancer diagnosis. However, in real-world clinical workflows, one view may be missing, corrupted, or degraded due to acquisition errors or compression artifacts, limiting the effectiveness of downstream analysis. View-to-view translation can help recover missing views and improve lesion alignment. Unlike natural images, this task in mammography is highly challenging due to large non-rigid deformations and severe tissue overlap in X-ray projections, which obscure pixel-level correspondences. In this paper, we propose Column-Aware and Implicit 3D Diffusion (CA3D-Diff), a novel bidirectional mammogram view translation framework based on conditional diffusion model. To address cross-view structural misalignment, we first design a column-aware Cross-Attention mechanism that leverages the geometric property that anatomically corresponding regions tend to lie in similar column positions across views. Furthermore, we introduce an implicit 3D structure reconstruction module that back-projects noisy 2D latents into a coarse 3D feature volume based on breast-view projection geometry. Extensive experiments demonstrate that CA3D-Diff achieves superior performance in bidirectional tasks, outperforming state-of-the-art methods in visual fidelity and structural consistency. Furthermore, the synthesized views effectively improve single-view malignancy classification in screening settings, demonstrating the practical value of our method in realworld diagnostics. Our code is available at https://github.com/lixinHUST/CA3D-Diff. Xin Li 0001, Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 1 |
| 2025 | Reviving Mural Art through Generative AI: A Comparative Study of AI-Generated and Hand-Crafted Recreations
Shuo Zhao 0017, Xiaoyang He, Xin Tong 0004, Xin Li 0001, Dan Wu 0003 |
CHI | 5 |
| 2025 | CoStoDet-DDPM: Collaborative Training of Stochastic and Deterministic Models Improves Surgical Workflow Anticipation and RecognitionabstractAnticipating and recognizing surgical workflows are critical for intelligent surgical assistance systems. However, existing methods rely on deterministic decision-making, struggling to generalize across the large anatomical and procedural variations inherent in real-world surgeries.In this paper, we introduce an innovative framework that incorporates stochastic modeling through a denoising diffusion probabilistic model (DDPM) into conventional deterministic learning for surgical workflow analysis. At the heart of our approach is a collaborative co-training paradigm: the DDPM branch captures procedural uncertainties to enrich feature representations, while the task branch focuses on predicting surgical phases and instrument usage.Theoretically, we demonstrate that this mutual refinement mechanism benefits both branches: the DDPM reduces prediction errors in uncertain scenarios, and the task branch directs the DDPM toward clinically meaningful representations. Notably, the DDPM branch is discarded during inference, enabling real-time predictions without sacrificing accuracy.Experiments on the Cholec80 dataset show that for the anticipation task, our method achieves a 16% reduction in eMAE compared to state-of-the-art approaches, and for phase recognition, it improves the Jaccard score by 1.0%. Additionally, on the AutoLaparo dataset, our method achieves a 1.5% improvement in the Jaccard score for phase recognition, while also exhibiting robust generalization to patient-specific variations. Our code and weight are available at https://github.com/kk42yy/CoStoDet-DDPM. Kaixiang Yang 0004, Xin Li 0001, Qiang Li 0018, Zhiwei Wang 0002 |
ICCV | 2 |
| 2025 | Joint Holistic and Lesion Controllable Mammogram Synthesis via Gated Conditional Diffusion ModelabstractMammography is the most commonly used imaging modality for breast cancer screening, driving an increasing demand for deep-learning techniques to support large-scale analysis. However, the development of accurate and robust methods is often limited by insufficient data availability and a lack of diversity in lesion characteristics. While generative models offer a promising solution for data synthesis, current approaches often fail to adequately emphasize lesion-specific features and their relationships with surrounding tissues. In this paper, we propose Gated Conditional Diffusion Model (GCDM), a novel framework designed to jointly synthesize holistic mammogram images and localized lesions. GCDM is built upon a latent denoising diffusion framework, where the noised latent image is concatenated with a soft mask embedding that represents breast, lesion, and their transitional regions, ensuring anatomical coherence between them during the denoising process. To further emphasize lesion-specific features, GCDM incorporates a gated conditioning branch that guides the denoising process by dynamically selecting and fusing the most relevant radiomic and geometric properties of lesions, effectively capturing their interplay. Experimental results demonstrate that GCDM achieves precise control over small lesion areas while enhancing the realism and diversity of synthesized mammograms. These advancements position GCDM as a promising tool for clinical applications in mammogram synthesis. Our code is available at https://github.com/lixinHUST/Gated-Conditional-Diffusion-Model/ Xin Li 0001, Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002 |
ACM Multimedia | 1 |
| 2025 | Robust analog/RF circuit design via Cycle-Consistent Generative Adversarial Networks
Nanlin Guo, Jun Tao 0001, Xuan Zeng 0001, Xin Li 0001 |
Integr. | 4 |
| 2025 | Efficient Design Optimization for Diffractive Deep Neural NetworksabstractSince diffractive deep neural network (D2NN) provides a full optical solution to implement deep neural networks (DNNs), it offers ultrafast operation speed and virtually unlimited bandwidth, yielding an alternative-yet-competitive approach for computer-based neural networks. A D2NN is composed of several 3D-printed phase masks as hidden layers and a number of optical detectors at the output. To enable automatic and efficient design of D2NNs, we propose an iterative optimization method to determine the optimal design parameters of D2NNs. During each iteration step, we first optimize the physical parameters for masks (e.g., thicknesses) while fixing the detector parameters (e.g., locations). Next, we exhaustively search the detector parameters with fixed masks. These two steps are repeated until convergence is reached. Our numerical experiments demonstrate that the proposed optimization algorithm can produce a high-performance D2NN achieving 97% accuracy for recognizing handwritten digits. Yuncheng Liu, Jun Tao 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Root-Cause Analysis with Semi-Supervised Co-Training for Integrated SystemsabstractRoot-cause analysis for integrated systems has become increasingly challenging due to their growing complexity. To tackle these challenges, machine learning (ML) has been applied to enhance root-cause analysis. Nonetheless, ML-based root-cause analysis usually requires abundant training data with root causes labeled by human experts, which are difficult or even impossible to obtain. To overcome this drawback, a semi-supervised co-training method is proposed for root-cause analysis in this article, which only requires a small portion of labeled data. First, a random forest is trained with labeled data. Next, we propose a co-training technique to learn from unlabeled data with semi-supervised learning, which pre-labels a subset of these data automatically and then retrains each decision tree in the random forest. In addition, a robust framework is proposed to avoid over-fitting. We further apply initialization by clustering and feature selection to improve the diagnostic performance. With two case studies from industry, the proposed approach shows superior performance against other state-of-the-art methods by saving up to 67% of labeling efforts. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | Dual-view Correlation Hybrid Attention Network for Robust Holistic Mammogram ClassificationabstractMammogram image is important for breast cancer screening, and typically obtained in a dual-view form, i.e., cranio-caudal (CC) and mediolateral oblique (MLO), to provide complementary information for clinical decisions. However, previous methods mostly learn features from the two views independently, which violates the clinical knowledge and ignores the importance of dual-view correlation in the feature learning. In this paper, we propose a dual-view correlation hybrid attention network (DCHA-Net) for robust holistic mammogram classification. Specifically, DCHA-Net is carefully designed to extract and reinvent deep feature maps for the two views, and meanwhile to maximize the underlying correlations between them. A hybrid attention module, consisting of local relation and non-local attention blocks, is proposed to alleviate the spatial misalignment of the paired views in the correlation maximization. A dual-view correlation loss is introduced to maximize the feature similarity between corresponding strip-like regions with equal distance to the chest wall, motivated by the fact that their features represent the same breast tissues, and thus should be highly-correlated with each other. Experimental results on the two public datasets, i.e., INbreast and CBIS-DDSM, demonstrate that the DCHA-Net can well preserve and maximize feature correlations across views, and thus outperforms previous state-of-the-art methods for classifying a whole mammogram as malignant or not. Zhiwei Wang 0002, Junlin Xian, Kangyi Liu, Xin Li 0001, Qiang Li 0018, Xin Yang 0008 |
IJCAI | 4 |
| 2023 | Correlated Bayesian Model Fusion: Efficient High-Dimensional Performance Modeling of Analog/RF Integrated Circuits Over Multiple CornersabstractEfficient high-dimensional performance modeling of analog/RF circuits over multiple corners is an important-yet-challenging task. In this article, we propose a novel performance modeling approach for analog/RF circuits, referred to as correlated Bayesian model fusion (C-BMF). The key idea is to encode the correlation information for both model template and coefficient magnitude among different corners by using a unified prior distribution. Next, the prior distribution is combined with a few simulation samples via Bayesian inference to efficiently determine the unknown model coefficients. Two circuit examples designed in a commercial 40-nm CMOS process demonstrate that C-BMF achieves about$2\times $cost reduction over the traditional state-of-the-art modeling technique without surrendering any accuracy. Zhengqi Gao, Fa Wang, Jun Tao 0001, Yangfeng Su, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Unsupervised Two-Stage Root-Cause Analysis With Transfer Learning for Integrated SystemsabstractThe growing complexity of integrated systems makes root-cause analysis increasingly difficult. To address this challenge, advances in machine learning (ML) have been leveraged in recent years to design ML-based techniques for root-cause analysis. However, most of these methods require root-cause labels for defective samples obtained based on the analysis by human experts. In this article, we propose a multialgorithm two-stage clustering method with transfer learning for unsupervised root-cause analysis. First, a two-stage clustering method is proposed by applying multiple clustering methods to accommodate both numerical and categorical data and leveraging Silhouette score for model selection. Next, a double-bootstrapping method is proposed for data selection, transferring valuable information from a source product to a target product with insufficient data. In the first bootstrapping step, a random forest model is built to select effective source data. In the second bootstrapping step, clustering ensemble is applied to two-stage clustering to further improve the accuracy for root-cause analysis. Two case studies based on network products demonstrate the superior performance of the proposed approach compared to other state-of-the-art methods. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Efficient Statistical Parameter Extraction for Modeling MOSFET MismatchabstractIn this article, we propose an efficient statistical parameter extraction method to accurately model the random device mismatch of MOSFETs. The key idea is to approximate the performance variations as mathematical functions of device mismatch. Based on these approximated functions and the electrical test data, we solve the unknown statistical parameters by nonlinear optimization. Our numerical experiments demonstrate that the proposed method can remarkably improve the modeling accuracy with affordable computational cost, compared against the state-of-the-art techniques. Nanlin Guo, Nengyong Zhu, Jun Tao 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Robust Wafer Classification With Imperfectly Labeled Data Based on Self-Boosting Co-TeachingabstractWafer classification is a critical task for semiconductor manufacturing. Most conventional algorithms require a large-scale perfectly labeled dataset to train accurate classifiers. In practice, it is usually difficult or even impossible to collect perfect labels without errors, and the classification accuracy in the presence of imperfectly labeled data would degrade. To facilitate robust wafer classification with noisy labels, we propose a novel self-boosting co-teaching (SB-CT) approach. Specifically, we iteratively correct the wrong labels by using the predictions of two classifiers that are jointly trained with noisily labeled data. To make the proposed method of practical utility, we develop a novel method to accurately estimate the noise rate, and adopt a probability scaling technique to further improve the classification accuracy. As demonstrated by the experimental results based on two industrial datasets, the proposed SB-CT approach achieves superior accuracy over other conventional methods. Shuo Zhao 0017, Zikun Zhu, Xin Li 0001, Ying-Chi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Data-Driven Parameterized Corner Synthesis for Efficient Validation of Perception Systems for Autonomous DrivingabstractToday's automotive cyber-physical systems for autonomous driving aim to enhance driving safety by replacing the uncertainties posed by human drivers with standard procedures of automated systems. However, the accuracy of in-vehicle perception systems may significantly vary under different operational conditions (e.g., fog density, light condition, etc.) and consequently degrade the reliability of autonomous driving. A perception system for autonomous driving must be carefully validated with an extremely large dataset collected under all possible operational conditions in order to ensure its robustness. The aforementioned dataset required for validation, however, is expensive or even impossible to acquire in practice, since most operational corners rarely occur in a real-world environment. In this paper, we propose to generate synthetic datasets at a variety of operational corners by using a parameterized cycle-consistent generative adversarial network (PCGAN) . The proposed PCGAN is able to learn from an image dataset recorded at real-world operational conditions with only a few samples at corners and synthesize a large dataset at a given operational corner. By taking STOP sign detection as an example, our numerical experiments demonstrate that the proposed approach is able to generate high-quality synthetic datasets to facilitate accurate validation. Handi Yu, Xin Li 0001 |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2023 | Unsupervised Cross-Modality Adaptation via Dual Structural-Oriented Guidance for 3D Medical Image SegmentationabstractDeep convolutional neural networks (CNNs) have achieved impressive performance in medical image segmentation; however, their performance could degrade significantly when being deployed to unseen data with heterogeneous characteristics. Unsupervised domain adaptation (UDA) is a promising solution to tackle this problem. In this work, we present a novel UDA method, named dual adaptation-guiding network (DAG-Net), which incorporates two highly effective and complementary structural-oriented guidance in training to collaboratively adapt a segmentation model from a labelled source domain to an unlabeled target domain. Specifically, our DAG-Net consists of two core modules: 1) Fourier-based contrastive style augmentation (FCSA) which implicitly guides the segmentation network to focus on learning modality-insensitive and structural-relevant features, and 2) residual space alignment (RSA) which provides explicit guidance to enhance the geometric continuity of the prediction in the target modality based on a 3D prior of inter-slice correlation. We have extensively evaluated our method with cardiac substructure and abdominal multi-organ segmentation for bidirectional cross-modality adaptation between MRI and CT images. Experimental results on two different tasks demonstrate that our DAG-Net greatly outperforms the state-of-the-art UDA approaches for 3D medical image segmentation on unlabeled target images. Junlin Xian, Dandan Tu, Senhua Zhu, Changzheng Zhang, Xiaowu Liu, Xin Li 0001, Xin Yang 0008 |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Semi-Supervised Root-Cause Analysis with Co-Training for Integrated SystemsabstractThe increasing complexity of integrated systems has exacerbated the challenges associated with system diagnosis. To tackle these challenges, intelligent root-cause-analysis facilitated by machine learning has been proposed in recent years. However, most of these methods rely on a large amount of data with root-cause labels, which are often either not available or difficult to obtain. In this paper, we propose a semi-supervised root-cause-analysis method with co-training, where only a small set of labeled data is required. Using random forest as the learning kernel, a co-training technique is proposed to leverage the unlabeled data by automatically pre-labeling a subset of them and retraining each decision tree. In addition, several novel techniques are proposed to avoid over-fitting and determine hyper-parameters. Two case studies based on industrial designs demonstrate that the proposed approach significantly outperforms state-of-the-art methods by saving up to 43% of labeling efforts by human experts. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
VTS | 2 |
| 2022 | Convolutional-capsule network for gastrointestinal endoscopy image classificationabstractAutomated diagnosis of digestive tract diseases from gastrointestinal endoscopy images is of high importance for improving the diagnosis accuracy and efficiency. The current mainstream methods for image classification of digestive tract endoscopy images are based on Convolutional Neural Networks (CNNs). However, due to their inherent defects, CNNs are not strong enough in learning deformation-invariant global features which is essential in gastrointestinal endoscopic image classification. To solve this problem, in this paper we present a two-stage endoscopic image classification method which can effectively combine complementary advantages of midlevel CNN features and a capsule network. Specifically, the core of our method is a lesion-aware CNN feature extraction module which can encode sufficiently detailed information of lesions in midlevel CNN features and in turn enable the subsequent capsule classification network to effectively learn deformation-invariant relationships between image entities. Extensive experiments demonstrate the superiority of our method to the state-of-the-art methods with the classification accuracy of 94.83% on the Kvasir v2 data set and the classification accuracy of 85.99% on the HyperKvasir data set. Wei Wang 0355, Xin Yang 0008, Xin Li 0001, Jinhui Tang 0001 |
Int. J. Intell. Syst. | 3 |
| 2022 | Fast Statistical Analysis of Rare Failure Events With Truncated Normal Distribution in High-Dimensional Variation SpaceabstractIn this article, to accurately estimate the rare failure rates for large-scale circuits (e.g., SRAM) where process variations are modeled as truncated normal distributions in high-dimensional space, we propose a novel truncated scaled-sigma sampling (T-SSS) method. Similar to scaled-sigma sampling (SSS), T-SSS distorts the truncated normal distributions by a scaling factor, resulting in an analytical model for failure rate estimation. By drawing random samples from the distorted distribution and estimating a sequence of scaled failure rates, we can solve all unknown model coefficients and predict the original failure rate by extrapolation. The accuracy of T-SSS is further assessed by estimating its confidence interval (CI) based on resampling. Our numerical results demonstrate that the proposed T-SSS method can achieve superior accuracy over the state-of-the-art method without increasing the computational cost. Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Knowledge Transfer in Board-Level Functional Fault Diagnosis Enabled by Domain AdaptationabstractHigh integration densities and design complexity make board-level functional fault diagnosis extremely difficult. Machine-learning techniques can identify functional faults with high accuracy, but they require a large volume of data to achieve high-prediction accuracy. This drawback limits the effectiveness of traditional machine-learning algorithms for training a model in the early stage of manufacturing, when only a limited amount of fail data and repair records are available. We propose a board-level diagnosis workflow that utilizes domain adaptation (DA) to transfer the knowledge learned from mature boards to a new board in the ramp-up phase. First, based on the requirement of fault diagnosis, we select an appropriate domain-adaptation method to reduce differences between mature boards and the new board. Second, these DA methods utilize information from both the mature and the new boards with carefully designed domain-alignment rules and train a functional fault diagnosis classifier. Experimental results using three complex boards in volume production and one new board in the ramp-up phase show that, with the help of DA and the proposed workflow, the diagnosis accuracy is improved. Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Unsupervised Two-Stage Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems have placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligence and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. We propose a two-stage unsupervised root-cause-analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to cluster the data in a coarse-grained manner. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. The proposed method can accommodate both numerical and categorical test items. A combination of the L-method, cross validation, and Silhouette score enables us to automatically determine all hyperparameters. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause-analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Correlated Rare Failure Analysis via Asymptotic Probability EvaluationabstractIn this article, a novel asymptotic probability evaluation (APE) method is proposed to estimate the probability of correlated rare failure events for complex integrated systems containing a large number of replicated cells. The key idea is to approximate the failure rate of the entire system by solving a set of nonlinear equations derived from a general analytical model. An error refinement method based on look-up table is further developed to improve numerical stability and, hence, reduce estimation error. Furthermore, a statistical algorithm based on resampling is developed to accurately estimate the confidence interval of APE. Our numerical experiments demonstrate that compared to the state-of-the-art method, APE can reduce the estimation error by up to$30\times $without increasing the computational cost. Jun Tao 0001, Handi Yu, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Robust Classification with Noisy Labels for Manufacturing Applications: A Hybrid Approach Based on Active Learning and Data CleaningabstractClassification is an important machine learning technique that attracts growing interests in various manufacturing applications. Learning an accurate classifier generally requires a large-scale perfectly-labeled training dataset. However, such "golden" labels are not only expensive but also difficult to collect in practice. To facilitate accurate classification in the presence of noisy labels, we propose a novel hybrid method based on active learning and data cleaning. Specifically, we first train an initial classifier with noisily- labeled data. Based on its prediction outcomes, a set of most informative samples is queried for manual annotation. To effectively correct other incorrect labels, we further self-label the unqueried samples based on the true labels provided by human experts and the estimated labels predicted by the initial classifier. As demonstrated by the experimental results based on two industrial datasets, the proposed approach achieves superior accuracy over other conventional methods. Shuo Zhao 0017, Xin Li 0001, Ying-Chi Chen |
IECON | 2 |
| 2021 | Unsupervised Root-Cause Analysis with Transfer Learning for Integrated SystemsabstractThe increasing complexity of integrated systems has exacerbated the problems associated with root-cause analysis. Leveraging advances artificial intelligence, a large amount of intelligent root-cause-analysis methods have been proposed in recent years. However, most of these methods rely on root-cause labels from repair history for defective samples, which are often expensive to obtain. In this paper, we propose an unsupervised root-cause-analysis method that utilizes transfer learning. A two-stage clustering method is first developed by exploiting model selection based on the concept of Silhouette score. Next, a data-selection method based on ensemble learning is proposed to transfer valuable information from a source product to improve the root-cause-analysis accuracy on the target product with insufficient data. Two case studies based on industry designs demonstrate that the proposed approach significantly outperforms other state-of-the-art unsupervised root-cause-analysis methods. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
VTS | 2 |
| 2021 | Facial Expression Recognition with Identity and Emotion Joint LearningabstractDifferent subjects may express a specific expression in different ways due to inter-subject variabilities. In this work, besides training deep-learned facial expression feature (emotional feature), we also consider the influence of latent face identity feature such as the shape or appearance of face. We propose an identity and emotion joint learning approach with deep convolutional neural networks (CNNs) to enhance the performance of facial expression recognition (FER) tasks. First, we learn the emotion and identity features separately using two different CNNs with their corresponding training data. Second, we concatenate these two features together as a deep-learned Tandem Facial Expression (TFE) Feature and feed it to the subsequent fully connected layers to form a new model. Finally, we perform joint learning on the newly merged network using only the facial expression training data. Experimental results show that our proposed approach achieves 99.31 and 84.29 percent accuracy on the CK+ and the FER+ database, respectively, which outperforms the residual network baseline as well as many other state-of-the-art methods. Ming Li 0026, Xingchang Huang, Zhanmei Song, Xin Li 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2021 | Board-Level Functional Fault Identification Using Streaming DataabstractHigh integration densities and design complexity of printed-circuit boards make board-level functional fault identification extremely difficult. Machine learning provides an opportunity to identify functional faults with high accuracy and thereby reduce repair cost. However, the large volume of manufacturing data comes in a streaming format and exhibits time-dependent concept drift in a production environment. These drawbacks limit the effectiveness of traditional machine-learning algorithms. We propose a diagnosis workflow that utilizes online learning to train classifiers incrementally with a small chunk of data at each step. These online-learning algorithms adapt to concept drift quickly with carefully designed update rules. A hybrid algorithm is also proposed to handle the scenario that data for varying numbers of boards are collected at different times. This hybrid algorithm concurrently implements two basic models. For each data chunk, this algorithm chooses the better model with high probability. The experimental results using two boards in high-volume production show that, with the help of online learning and the proposed hybrid algorithm, the F1-score for diagnosis based on binary classifiers can be improved from 57.3% to 81.0%. The top-3 accuracy for diagnosis based on multiclass classifiers can be improved from 78.3% to 91.4%. Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Black-Box Test-Cost Reduction Based on Bayesian Network ModelsabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. In this article, we propose a novel black-box test selection method based on Bayesian networks (BNs), which extract the strong relationship among tests. First, the problem of reducing the black-box test cost is formulated as a constrained optimization problem. Next, multiple structure learning and transfer learning algorithms are implemented to construct BN models. Based on these BN models, we propose an iterative test selection method with a new metric, Bayesian index, for test-cost reduction. In addition, averaging strategies are applied to enhance the reduction performance. Finally, a robust model selection framework is proposed to select the optimal BN model for test-cost reduction. Two case studies with production test data demonstrate that when no prior information is provided, our proposed approach effectively reduces the test cost by up to 14.7%, compared to the state-of-the-art greedy algorithm. Moreover, our proposed approach further reduces the test cost by up to 7.1% when prior information is provided from similar products. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | A Survey on Edge and Edge-Cloud Computing Assisted Cyber-Physical SystemsabstractIn recent years, the investigations on cyber-physical systems (CPS) have become increasingly popular in both academia and industry. A primary obstruction against the booming deployment of CPS applications lies in how to process and manage large amounts of generated data for decision making. To tackle this predicament, researchers advocate the idea of coupling edge computing, or edge-cloud computing into the design of CPS. However, this coupling process raises a diversity of challenges to the quality-of-services (QoS) of CPS applications. In this article, we present a survey on edge computing or edge-cloud computing assisted CPS designs from the QoS optimization perspective. We first discuss critical challenges in service latency, energy consumption, security, privacy, and reliability during the integration of CPS with edge computing or edge-cloud computing. Afterwards, we give an overview on the state-of-the-art works tackling different challenges for QoS optimization, and present a systematic classification during outlining literature for highlighting their similarities and differences. We finally summarize the experiences learned from surveyed works and envision future research directions on edge computing or edge-cloud computing assisted CPS optimization. Kun Cao 0001, Shiyan Hu 0001, Yang Shi 0001, Armando W. Colombo, Stamatis Karnouskos, Xin Li 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | Guest Editorial: Cloud-Edge Computing for Cyber-Physical Systems and Internet of ThingsabstractThis Special Section on ‘`Cloud-Edge Computing for Cyber-Physical Systems and Internet-of-Things’' is oriented to the dissemination of a few of those latest research and innovation results, covering many aspects of design, optimization, implementation, and evaluation of emerging cloud-edge solutions for CPS and IoT applications. The selected high-quality contributions cover a broad range of novel technologies and application scenarios in CPS and IoT. We hope that these accepted papers will produce long-lasting impacts, as well as stimulating and encouraging the international community to work on this exciting and impactful topic. Shiyan Hu 0001, Yang Shi 0001, Armando W. Colombo, Stamatis Karnouskos, Xin Li 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Efficient Classification via Partial Co-Training for Virtual MetrologyabstractDeveloping accurate and cost-effective classification techniques to facilitate virtual metrology is a critical task for modern manufacturing. In this paper, we consider the scenario in which labeling data is expensive, causing a shortage of labeled data. As a consequence, conventional classification methods suffer from a high risk of overfitting. To address this issue, we develop a novel semi-supervised classification method, namely Partial Cotraining with Logistic Regression (PCT-LR). PCT-LR finds a subset of the original features to generate a partial view, and uses this partial view to provide side information to support the complete view that includes all features. Both views are cooptimized in a Bayesian inference with a Gaussian process prior and a logistic regression classifier. The proposed method is validated with two industrial examples. Experiment results suggest that the amount of required labeled data can be reduced by up to 18% without loss in accuracy. Xin Li 0001, R. D. (Shawn) Blanton, Xiang Li 0001 |
ETFA | 2 |
| 2020 | A Classification Framework Using Imperfectly Labeled Data for Manufacturing ApplicationsabstractIn recent years, classification techniques have been broadly adopted for a variety of smart manufacturing applications, including system health maintenance, defect detection and diagnosis, etc. However, most classification methods require a large set of training data that are accurately labeled by human experts in a specific domain. Collecting these training data is time-consuming and prohibitively expensive in practical applications. To overcome this challenge, we develop a classification framework using imperfectly labeled data. First, a statistical model is proposed to derive a set of probabilistic labels with consideration of labeling errors. Next, an accurate classifier is trained from these inaccurate labels. As demonstrated by the experimental results of two industrial examples, the proposed framework achieves superior classification accuracy over other conventional approaches. Shuo Zhao 0017, Xin Li 0001, Ying-Chi Chen |
ETFA | 2 |
| 2020 | Exploring Inter-Sensor Correlation for Missing Data EstimationabstractData mining techniques have been widely applied to various fields including industrial, business, and governmental applications. Missing data is a common occurrence in a number of real-world databases, which may substantially affect the accuracy of data processing. In this paper, we propose a novel approach for missing data estimation by efficiently exploring inter-sensor correlation. Namely, given multiple sensors for data collection, we attempt to recover the missing data of a few sensors by using the measurement data from other sensors. Towards this goal, we develop an iterative solver for missing data estimation. Our numerical experiments on two industrial datasets demonstrate that the proposed method can reduce the imputation error by up to 7.25× compared to a conventional method in the literature. Liying Li 0002, Yang Liu 0064, Tongquan Wei, Xin Li 0001 |
IECON | 4 |
| 2020 | Unsupervised Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems has placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligent and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. In this paper, we propose a two-stage unsupervised root-cause analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to roughly cluster the data. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. In additional, L-method and cross validation are applied to automatically determine the hyper-parameters of our algorithm. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 3 |
| 2020 | Multi-phase and Multi-level Selective Feature Fusion for Automated Pancreas Segmentation from CT Images
Xixi Jiang, Qingqing Luo, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (4) | 6 |
| 2020 | Big Data for Cyber-Physical SystemsabstractCyber-physical systems (CPS) are characterized by deep and complex intertwining among cyber components and physical components. Due to the fast increase in system complexities, the operations of CPS involve sensing, processing and storage of massive amount of data. This nature of “big data” imposes fundamental challenges on the design and management of CPS in multiple aspects such as performance, energy efficiency, security, privacy, reliability, sustainability, fault tolerance, scalability and flexibility. Tackling these challenges necessitates innovative big data techniques for handling massive data in CPS. The articles in this special section include a few selected state-of-the-art research results on the topic of big data sensing, processing and storage for CPS, and stimulates a broad range of researchers to participate in the interdisciplinary CPS research in the future. This special issue has received a significant number of submissions while only a small portion of them are selected for publications. The selected papers showcase how interesting data analytics techniques can be leveraged to optimize different metrics in CPS, such as timing, efficiency, schedulability, power, reliability, and security, etc. Shiyan Hu 0001, Xin Li 0001, Haibo He, Shuguang Cui, Manish Parashar |
IEEE Trans. Big Data | 2 |
| 2020 | Efficient Rare Failure Analysis Over Multiple Corners via Correlated Bayesian InferenceabstractIn this article, we propose an efficient correlated Bayesian inference (CBI) method to estimate the system-level failure rates for large-scale circuit systems over multiple process corners. The key idea is to encode the correlations of circuit performances among the different corners into the prior distributions of several carefully defined failure events. The hyper-parameters of these distributions can be learned from a few simulation samples via Bayesian inference and, next, the system-level failure rates over different corners can be simultaneously estimated by taking into account these prior distributions. An iteratively constrained inference method is further developed to guarantee the numerical stability of the proposed method and legalize all estimated failure rates. The numerical experiments demonstrate that compared to the state-of-the-art algorithm, the proposed method can achieve around 10× runtime reduction without surrendering any accuracy. Zhengqi Gao, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Efficient Statistical Analysis for Correlated Rare Failure Events via Asymptotic Probability ApproximationabstractIn this article, a novel asymptotic probability approximation (APA) method is proposed to estimate the overall rare probability of correlated failure events for complex circuits containing a large number of replicated cells (e.g., SRAM bit-cells). The key idea of APA is to approximate the overall circuit failure rate based on a set of carefully defined failure events. An efficient hierarchical subset simulation (H-SUS) method is developed to calculate the aforementioned failure rate and a statistical methodology is further proposed to estimate the confidence interval of APA. Our numerical experiments demonstrate that APA can accurately and reliably estimate the overall failure rate of correlated rare failure events involving more than 20 000 independent random variables. Fulin Peng, Handi Yu, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2020 | Partial Bayesian Co-training for Virtual MetrologyabstractBuilding accurate regression models using limited data is a challenging problem in manufacturing data analysis. In this paper, we study a particular semisupervised learning problem where labeled data are limited, while unlabeled data are plentiful. In these conditions, conventional single-view learning methods are prone to overfitting. To tackle this problem, we develop a novel co-training technique, namely partial Bayesian co-training (PBCT). PBCT scales down the original set of features to create a partial view, and then exploit side information from the partial view to enhance the complete model. The PBCT model also allows integrating domain knowledge to enhance model accuracy. The proposed method is validated with experiments on industrial manufacturing data. The experimental results show that under a reduction of labeled data by up to 50%, a robust estimation is still attainable. This suggests that the PBCT model is a promising solution to a broad spectrum of applications. Cuong Manh Nguyen, Xin Li 0001, R. D. (Shawn) Blanton, Xiang Li 0040 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Bi-Modality Medical Image Synthesis Using Semi-Supervised Sequential Generative Adversarial NetworksabstractIn this paper, we propose a bi-modality medical image synthesis approach based on sequential generative adversarial network (GAN) and semi-supervised learning. Our approach consists of two generative modules that synthesize images of the two modalities in a sequential order. A method for measuring the synthesis complexity is proposed to automatically determine the synthesis order in our sequential GAN. Images of the modality with a lower complexity are synthesized first, and the counterparts with a higher complexity are generated later. Our sequential GAN is trained end-to-end in a semi-supervised manner. In supervised training, the joint distribution of bi-modality images are learned from real paired images of the two modalities by explicitly minimizing the reconstruction losses between the real and synthetic images. To avoid overfitting limited training images, in unsupervised training, the marginal distribution of each modality is learned based on unpaired images by minimizing the Wasserstein distance between the distributions of real and fake images. We comprehensively evaluate the proposed model using two synthesis tasks based on three types of evaluate metrics and user studies. Visual and quantitative results demonstrate the superiority of our method to the state-of-the-art methods, and reasonable visual quality and clinical significance. Code is made publicly available at https://github.com/hust- linyi/Multimodal-Medical-Image-Synthesis. Xin Yang 0008, Yi Lin 0009, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Fine-grained Adaptive Testing Based on Quality PredictionabstractThe ever-increasing complexity of integrated circuits inevitably leads to high test cost. Adaptive testing provides an effective solution for test-cost reduction; this testing framework selects the important test items for each set of chips. However, adaptive testing methods designed for digital circuits are coarse-grained, and they are targeted only at systematic defects. To incorporate fabrication variations and random defects in the testing framework, we propose a fine-grained adaptive testing method based on machine learning. We use the parametric test results from the previous stages of test to train a quality-prediction model for use in subsequent test stages. Next, we partition a given lot of chips into two groups based on their predicted quality. A test-selection method based on statistical learning is applied to the chips with high predicted quality. An ad hoc test-selection method is proposed and applied to the chips with low predicted quality. Experimental results using a large number of fabricated chips and the associated test data show that to achieve the same defect level as in prior work on adaptive testing, the fine-grained adaptive testing method reduces test cost by 90% for low-quality chips and up to 7% for all the chips in a lot. Renjian Pan, Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2020 | Analog/RF Post-silicon Tuning via Bayesian OptimizationabstractTunable analog/RF circuit has emerged as a promising technique to address the significant performance uncertainties caused by process variations. To optimize these tunable circuits after fabrication, most existing post-silicon programming methods are developed by using real-valued performance metrics. However, when measuring a performance of interest on silicon, it is often substantially more expensive to obtain a real-valued measurement than a binary testing outcome (i.e., pass or fail). In this article, we propose a Gaussian Process Classification model to capture the binary performance metrics of tunable analog/RF circuits. Based on these models, post-silicon programming is cast into an optimization problem that can be solved by a novel Bayesian optimization algorithm. Moreover, measurement noises are further incorporated into our proposed post-silicon programming to produce a robust circuit. Two circuit examples demonstrate that the proposed approach can efficiently program tunable circuits with binary performance metrics while other conventional methods are not applicable. Renjian Pan, Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2019 | Knowledge Transfer in Board-Level Functional Fault Identification using Domain AdaptationabstractHigh integration densities and design complexity make board-level functional fault identification extremely difficult. Machine-learning techniques can identify functional faults with high accuracy, but they require a large volume of data to achieve high prediction accuracy. This drawback limits the effectiveness of traditional machine-learning algorithms for training a model in the early stage of manufacturing, when only a limited amount of fail data and repair records are available. We propose a board-level diagnosis workflow that utilizes domain adaptation to transfer the knowledge learned from a mature board to a new board in the ramp-up phase. First, a metric is designed to evaluate the similarity between products, and based on the calculated value of the similarity, either a homogeneous or a heterogeneous domain adaptation algorithm is selected. Second, these domain adaptation algorithms utilize information from both the mature and the new boards with carefully designed domain-alignment rules and train a functional fault identification classifier. Three complex boards in volume production and one new board in the ramp-up phase are used to validate the proposed domain-adaptation approach in terms of the diagnosis accuracy. Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2019 | Board-Level Functional Fault Identification using Streaming DataabstractHigh integration densities and design complexity of printed-circuit boards make board-level functional fault identification extremely difficult. Machine learning provides an opportunity to identify functional faults with high accuracy and thereby reduce repair cost. However, the large volume of manufacturing data comes in a streaming format and exhibits time-dependent concept drift in a production environment. These drawbacks limit the effectiveness of traditional machine-learning algorithms. We propose a diagnosis workflow that utilizes online learning to train classifiers incrementally with a small chunk of data at each step. These online learning algorithms adapt to concept drift quickly with carefully designed update rules. A hybrid algorithm is also proposed to handle the scenario that data for varying numbers of boards are collected at different times. Experimental results using two boards in high-volume production show that, with the help of online learning and the proposed hybrid algorithm, the F1-score for diagnosis can be improved from 57.3% to 78.9%. Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 3 |
| 2019 | Black-Box Test-Coverage Analysis and Test-Cost Reduction Based on a Bayesian Network ModelabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. The conventional greedy algorithm, which selects the most important tests by considering both strong and weak relationships among tests, suffers from overfitting. In order to overcome overfitting, we propose a novel black-box test selection method based on a Bayesian network model. First, the problem of reducing black-box test cost is formulated as a constrained optimization problem. Next, a score-based algorithm is implemented to construct the Bayesian network for black-box tests. Finally, we propose a Bayesian index with the property of Markov blankets, and then an iterative test selection method is developed based on our proposed Bayesian index. The proposed approach ensures that only the strong relationships among black-box tests are used for test selection so that this approach is more robust to overfitting. Two case studies with production test data demonstrate that the proposed approach effectively reduces test cost by up to 14.7%, compared to a conventional greedy algorithm. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 3 |
| 2019 | Computer vision algorithms and hardware implementations: A surveyabstractThe field of computer vision is experiencing a great-leap-forward development today. This paper aims at providing a comprehensive survey of the recent progress on computer vision algorithms and their corresponding hardware implementations. In particular, the prominent achievements in computer vision tasks such as image classification, object detection and image segmentation brought by deep learning techniques are highlighted. On the other hand, review of techniques for implementing and optimizing deep-learning-based computer vision algorithms on GPU, FPGA and other new generations of hardware accelerators are presented to facilitate real-time and/or energy-efficient operations. Finally, several promising directions for future research are presented to motivate further development in the field. Youni Jiang, Xuejiao Yang, Xin Li 0001 |
Integr. | 5 |
| 2019 | Guest Editors' Introduction to the Special Section on Hardware and Algorithms for Energy-Constrained On-chip Machine LearningabstractNo abstract available. Jae-sun Seo, Yu Cao 0001, Xin Li 0001, Paul N. Whatmough |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Guest Editors' Introduction: Hardware and Algorithms for Energy-Constrained On-Chip Machine Learning (Part 2)abstractNo abstract available. Jae-sun Seo, Yu Cao 0001, Xin Li 0001, Paul N. Whatmough |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Graph-Constrained Sparse Performance Modeling for Analog Circuit Optimization via SDP RelaxationabstractIn this paper, a graph-constrained sparse performance modeling method is proposed for analog circuit optimization. It builds sparse polynomial models constrained by an acyclic graph. These models can be used to solve analog optimization problems within local design spaces by using convex semidefinite programming relaxation both efficiently and robustly. Our numerical examples demonstrate that the proposed modeling and optimization method can quickly and accurately converge to a superior solution for analog circuits while the conventional method fails to work. Jun Tao 0001, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Dependable Visual Light-Based Indoor Localization with Automatic Anomaly Detection for Location-Based Service of Mobile Cyber-Physical SystemsabstractIndoor localization has become popular in recent years due to the increasing need of location-based services in mobile cyber-physical systems (CPS). The massive deployment of light emitting diodes (LEDs) further promotes the indoor localization using visual light. As a key enabling technique for mobile CPS, accurate indoor localization based on visual light communication remains nontrivial due to various non-idealities such as attenuation induced by unexpected obstacles. The anomalies of localization can potentially reduce the dependability of location-based services. In this article, we develop a novel indoor localization framework based on relative received signal strength. Most importantly, an efficient method is derived from the triangle inequality to automatically detect the abnormal LED lamps that are blocked by obstacles. These LED lamps are then ignored by our localization algorithm so that they do not bias the localization results, which improves the dependability of our localization framework. As demonstrated by the simulation results, the proposed techniques can achieve superior accuracy over the conventional approaches, especially when there exist abnormal LED lamps. Yang Liu 0064, Xiaoming Chen 0003, Dileep Kadambi, Ajinkya Bari, Xin Li 0001, Shiyan Hu 0001, Pingqiang Zhou |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2019 | Cross-Scale Predictive DictionariesabstractSparse representations using data dictionaries provide an efficient model particularly for signals that do not enjoy alternate analytic sparsifying transformations. However, solving inverse problems with sparsifying dictionaries can be computationally expensive, especially when the dictionary under consideration has a large number of atoms. In this paper, we incorporate additional structure on to dictionary-based sparse representations for visual signals to enable speedups when solving sparse approximation problems. The specific structure that we endow onto sparse models is that of a multi-scale modeling where the sparse representation at each scale is constrained by the sparse representation at coarser scales. We show that this cross-scale predictive model delivers significant speedups, often in the range of , with little loss in accuracy for linear inverse problems associated with images, videos, and light fields. Vishwanath Saragadam, Xin Li 0001, Aswin C. Sankaranarayanan |
IEEE Trans. Image Process. | 2 |
| 2018 | Intelligent corner synthesis via cycle-consistent generative adversarial networks for efficient validation of autonomous driving systemsabstractToday's automotive vehicles are often equipped with powerful data processing systems for driver assistance and/or autonomous driving. To meet the rigorous safety standard, one critical task is to ensure extremely small failure rate over all possible operation conditions. Such a validation task requires a large amount of on-road testing data to cover all possible corners. In this paper, we describe a novel general-purpose methodology to synthetically and efficiently generate a broad spectrum of corner cases for validation purpose. Our proposed method is based upon cycle-consistent generative adversarial networks (CycleGANs) trained by a small set of image samples to mathematically map a nominal case to other corner cases. By taking STOP sign detection as an example, our numerical experiments demonstrate that the proposed approach is able to reduce the validation error by up to 100× given a limited data set for corner cases. Handi Yu, Xin Li 0001 |
ASP-DAC | 2 |
| 2018 | Predictive Modeling for Advanced Virtual Metrology: A Tree-Based ApproachabstractThe rapid development of industry 4.0 has promoted the extensive adoption of big data analytics for manufacturing industry. In this domain, virtual metrology is a critical technique that is able to reduce manufacturing cost over a large amount of practical applications. In this paper, we propose a novel tree-based approach for simultaneous feature selection and predictive modeling to facilitate efficient virtual metrology. The proposed method accurately identifies multiple feature sets and then chooses the best candidate to minimize modeling error. As demonstrated by the experimental results based on two industrial examples, the proposed method can achieve higher modeling accuracy and find a more complete feature set than the conventional approach implemented with orthogonal matching pursuit (OMP). Yang Liu 0064, Xin Li 0001 |
ETFA | 2 |
| 2018 | Model-based and data-driven approaches for building automation and controlabstractSmart buildings in the future are complex cyber-physical-human systems that involve close interactions among embedded platform (for sensing, computation, communication and control), mechanical components, physical environment, building architecture, and occupant activities. The design and operation of such buildings require a new set of methodologies and tools that can address these heterogeneous domains in a holistic, quantitative and automated fashion. In this paper, we will present our design automation methods for improving building energy efficiency and offering comfortable services to occupants at low cost. In particular, we will highlight our work in developing both model-based and data-driven approaches for building automation and control, including methods for co-scheduling heterogeneous energy demands and supplies, for integrating intelligent building energy management with grid optimization through a proactive demand response framework, for optimizing HVAC control with deep reinforcement learning, and for accurately measuring in-building temperature by combining prior modeling information with few sensor measurements based upon Bayesian inference. Tianshu Wei, Xiaoming Chen 0003, Xin Li 0001, Qi Zhu 0002 |
ICCAD | 3 |
| 2018 | Fine-Grained Adaptive Testing Based on Quality PredictionabstractThe ever-increasing complexity of integrated circuits inevitably leads to high test cost. Adaptive testing provides an effective solution for test-cost reduction; this testing framework selects the important test items for each set of chips. However, adaptive testing methods designed for digital circuits are coarse-grained, and they are targeted only at systematic defects. In order to incorporate fabrication variations and random defects in the testing framework, we propose a fine-grained adaptive testing method based on machine learning. We use the parametric test results from the previous stages of test to train a quality-prediction model for use in subsequent test stages. Next, we partition a given lot of chips into two groups based on their predicted quality. A test-selection method based on statistical learning is applied to the chips with high predicted quality. An ad hoc test-selection method is proposed and applied to the chips with low predicted quality. Experimental results using a large number of fabricated chips and the associated test data show that to achieve the same defect level as in prior work on adaptive testing, the fine-grained adaptive testing method reduces test cost by 90% for low-quality chips, and up to 7% for all the chips in a lot. Renjian Pan, Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2018 | Guest Editors' Introduction: Frontiers of Hardware and Algorithms for On-chip LearningabstractNo abstract available. Yu Cao 0001, Xin Li 0001, Jae-sun Seo, Ganesh Dasika |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | Design Automation for Cyber-Physical Systems [Scanning the Issue]abstractCyber-physical systems (CPSs) are characterized by the seamless integration and close interaction of cyber components (e.g., sensors, computation nodes, communication networks) and physical processes (e.g., mechanical devices, physical environment, humans). The cyber components monitor, analyze, and control the physical processes, and react to their changes through feedback loops. A classic example of CPSs is autonomous vehicles. These vehicles collect information of the surrounding physical environment via heterogeneous sensors such as cameras, radar, and LIDAR; process and analyze the multi-modal information at real time with advanced computing devices such as GPUs, application-specific SoCs and multicore CPUs; automatically make planning and control decisions; and continuously actuate the corresponding mechanical components. The cyber components of autonomous vehicles are much more intelligent and complex than those of traditional vehicles, and interact more directly and closely with the physical environment. Qi Zhu 0002, Alberto L. Sangiovanni-Vincentelli, Shiyan Hu 0001, Xin Li 0001 |
Proc. IEEE | 4 |
| 2018 | Identifying Wafer-Level Systematic Failure Patterns via Unsupervised LearningabstractIn this paper, we propose a novel methodology for detecting systematic failure patterns at the wafer level for yield learning. Our proposed methodology takes the binary testing results (i.e., pass or fail) of all dies over multiple wafers, cluster these wafers according to their spatial signatures of failures, and eventually identify the underlying systematic failure patterns. Several data processing techniques, including singular value decomposition, hierarchical clustering, etc., are adopted to make the proposed methodology robust to random failures. In addition, a Pseudo-Boolean satisfiability solver is used to extract a minimal set of systematic failure patterns that explain all wafer-level spatial signatures. These patterns help process engineers identify the root causes of failures and accelerate yield learning. The efficacy of our proposed approach is demonstrated by one synthetic data set and one industrial data set. Mohamed Baker Alawieh, Fa Wang, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Efficient Hierarchical Performance Modeling for Analog and Mixed-Signal Circuits via Bayesian Co-LearningabstractWith the continuous drive toward integrated circuits scaling, efficient performance modeling is becoming more crucial yet more challenging. In this paper, we propose a novel method of hierarchical performance modeling based on Bayesian co-learning. We exploit the hierarchical structure of a circuit to establish a Bayesian framework where unlabeled data samples are generated to improve modeling accuracy without running additional simulation. Consequently, our proposed method only requires a small number of labeled samples, along with a large number of unlabeled samples obtained at almost no-cost, to accurately learn a performance model. Our numerical experiments demonstrate that the proposed approach achieves up to 3.6× runtime speed-up over the state-of-the-art modeling technique without surrendering any accuracy. Mohamed Baker Alawieh, Fa Wang, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Guest Editorial Circuit and System Design Automation for Internet of ThingsabstractInternet-of-Things (IoT) is the technical backbone of smart cities which are envisioned to cope up with rapid urbanization of human population with limited resources. IoT provides three key features of smart cities such as intelligence, interconnection, and instrumentation. IoT is essentially a system-of-systems which can be considered as a configurable dynamic global network of networks. The main components of IoT include the following: 1) The Things; 2) Internet; 3) LAN; and 4) The Cloud. IoT is built by various diverse components including electronics, sensors, actuators, controllers, networks, firmware, and software. However, the existing electronics, controllers, and processors do not meet IoT requirements, such as multiple sensors, communication protocols, and security requirements. The existing computer-aided design (CAD) or electronic design automation tools are not enough to meet diverse challenges such as time-to-market, complexity, and cost of IoT. The required electronic circuits and systems need to be developed by handling and solving specific requirements. Real-time and ultralow power plays a major role since mobile devices in the IoT have to provide a long availability with a relative small energy budget. At the same time, reliability, availability, real-time constraints, and performance requirements pose significant challenges, and therefore, lead to a high interest in research. In this special issue, different approaches to design novel devices, circuits, and systems for solving the challenges with IoT are targeted. Various novel design automation components including modeling, design flows, simulation methods, and optimizations for designing of modern IoT are targeted, from system level down to device level. The current special issue was envisioned with the above technical considerations. After a rigorous review process, a set of articles were selected for this special issue. These papers are briefly discussed in the rest of the editorial. Saraju P. Mohanty, Michael Hübner 0001, Chun Jason Xue, Xin Li 0001, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Improving Diagnostic Resolution of Failing ICs Through LearningabstractDiagnosis is the first analysis step for uncovering the root cause of failure for a defective integrated logic circuit. The conventional objective of identifying failure locations has been augmented with various physically-aware diagnosis techniques that are intended to improve both resolution and accuracy. Despite these advances, it is often the case, however, that resolution, i.e., the number of locations or candidates reported by diagnosis, exceeds the number of actual failing locations. To address this major challenge, a novel, machine-learning-based resolution improvement methodology named physically-aware diagnostic resolution enhancement (PADRE) is described. PADRE uses easily-available tester and simulation data to extract features that uniquely characterize each candidate. PADRE applies machine learning to the features to identify candidates that correspond to the actual failure locations. Through various experiments, PADRE is shown to significantly improve resolution with virtually no negative impact on accuracy. Additional experiments demonstrate that PADRE is robust against data set variation and feature-data availability. Xin Li 0001, R. D. (Shawn) Blanton |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Efficient Hierarchical Performance Modeling for Integrated Circuits via Bayesian Co-LearningabstractWith the continuous drive towards integrated circuits scaling, efficient performance modeling is becoming more crucial yet, more challenging. In this paper, we propose a novel method of hierarchical performance modeling based on Bayesian co-learning. We exploit the hierarchical structure of a circuit to establish a Bayesian framework where unlabeled data samples are generated to improve modeling accuracy without running additional simulation. Consequently, our proposed method only requires a small number of labeled samples, along with a large number of unlabeled samples obtained at almost no-cost, to accurately learn a performance model. Our numerical experiments demonstrate that the proposed approach achieves up to 3.66x runtime speed-up over the state-of-the-art modeling technique without surrendering any accuracy. Mohamed Baker Alawieh, Fa Wang, Xin Li 0001 |
DAC | 3 |
| 2017 | Correlated Rare Failure Analysis via Asymptotic Probability EvaluationabstractIn this paper, a novel Asymptotic Probability Estimation (APE) method is proposed to estimate the probability of correlated rare failure events for complex integrated systems containing a large number of replicated cells. The key idea is to approximate the failure rate of the entire system by solving a set of nonlinear equations derived from a general analytical model. An error refinement method based on Look-up Table (LUT) is further developed to improve numerical stability and, hence, reduce estimation error. Our numerical experiments demonstrate that compared to the state-of-the-art method, APE can reduce the estimation error by up to 45x without increasing the computational cost. Jun Tao 0001, Handi Yu, Dian Zhou, Yangfeng Su, Xuan Zeng 0001, Xin Li 0001 |
DAC | 6 |
| 2017 | Partial co-training for virtual metrologyabstractVirtual metrology is an important tool for industrial automation. To accurately build regression models for virtual metrology, we consider semi-supervised learning where labeled data are expensive to collect, but unlabeled data are abundant. In such a scenario, due to the scarcity of labeled data, traditional single-view learning methods face the risk of overfitting. To address the overfitting issue, we develop a Partial Co-training framework, which is an extension of the original co-training approach by means of an undirected probabilistic graphical model. Unlike other co-training techniques, this model creates a partial view by shrinking the original feature space, and makes use of this partial-view to provide guidance information for improving the complete-view model. Our approach is validated with data from two manufacturing applications. The results indicate that a consistent and robust estimation is achievable with very limited labeled data. Xin Li 0001, R. D. (Shawn) Blanton, Xiang Li 0040 |
ETFA | 2 |
| 2017 | ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA
Song Han 0003, Junlong Kang, Huizi Mao, Yiming Hu, Xin Li 0001, Dongliang Xie, Yu Wang 0002, Huazhong Yang, William J. Dally |
FPGA | 5 |
| 2017 | Efficient programming of reconfigurable radio frequency (RF) systemsabstractReconfigurable radio frequency (RF) system has recently emerged as a promising solution to cope with multiple communication standards and high spectrum density. In this paper, we propose a novel optimization framework to efficiently program a reconfigurable RF system. In particular, two novel techniques, including (i) search space reduction by adaptive resolution and (ii) global polynomial optimization based on branch and bound, are developed. When combined with a relaxation iteration scheme, our proposed method offers superior performance when programming a large-scale reconfigurable RF system designed for the WLAN 802.11g standard. Mohamed Baker Alawieh, Fa Wang, Jun Tao 0001, Shihui Yin, Minhee Jun, Xin Li 0001, Tamal Mukherjee, Rohit Negi |
ICCAD | 6 |
| 2017 | Impact of circuit-level non-idealities on vision-based autonomous driving systemsabstractWe describe a novel methodology to validate vision-based autonomous driving systems over different circuit corners with consideration of temperature variation and circuit aging. The proposed work is motivated by the fact that low-level circuit implementation may have a significant impact on system performance, even though such effects have not been appropriately taken into account today. Our approach seamlessly integrates the image data recorded under nominal conditions with comprehensive statistical circuit models to synthetically generate the critical corner cases for which an autonomous driving system is likely to fail. As such, a given automotive system can be robustly validated for these worst-case scenarios that cannot be easily captured by physical experiments. Handi Yu, Changhao Yan, Xuan Zeng 0001, Xin Li 0001 |
ICCAD | 4 |
| 2017 | Compressive spectral anomaly detectionabstractWe propose a novel compressive imager for detecting anomalous spectral profiles in a scene. We model the background spectrum as a low-dimensional subspace while assuming the anomalies to form a spatially-sparse set of spectral profiles different from the background. Our core contributions are in the form of a two-stage sensing mechanism. In the first stage, we estimate the subspace for the background spectrum by acquiring spectral measurements at a few randomly-selected pixels. In the second stage, we acquire spatially-multiplexed spectral measurements of the scene. We remove the contributions of the background spectrum from the spatially-multiplexed measurements by projecting onto the complementary subspace of the background spectrum; the resulting measurements are of a sparse matrix that encodes the presence and spectra of anomalies, which can be recovered using a Multiple Measurement Vector formulation. Theoretical analysis and simulations show significant speed up in acquisition time over other anomaly detection techniques. A lab prototype based on a DMD and a visible spectrometer validates our proposed imager. Vishwanath Saragadam, Jian Wang 0100, Xin Li 0001, Aswin C. Sankaranarayanan |
ICCP | 3 |
| 2017 | An Audio Based Piano Performance Evaluation Method Using Deep Neural Network Based Acoustic Modeling
Ming Li 0026, Zhanmei Song, Xin Li 0001, Hua Yi, Manman Zhu |
INTERSPEECH | 4 |
| 2017 | Energy-aware morphable cache management for self-powered non-volatile processorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In order to survive the power failures, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup, however, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Evaluation shows that the proposed scheme can achieve 16.7% energy reduction with comparative performance with the single level cell (SLC) based hybrid cache. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Xin Li 0001, Zhiping Jia |
RTCSA | 5 |
| 2017 | Algorithm and hardware implementation for visual perception system in autonomous vehicle: A survey
Weijing Shi, Mohamed Baker Alawieh, Xin Li 0001, Huafeng Yu |
Integr. | 3 |
| 2017 | Guest Editors' Introduction: Hardware and Algorithms for On-Chip Learningabstracteditorial Free Access Share on Guest Editors’ Introduction: Hardware and Algorithms for On-Chip Learning Authors: Yu Cao Arizona State University, Tempe, Arizona Arizona State University, Tempe, ArizonaView Profile , Xin Li Carnegie Mellon University, Pittsburgh, PA Carnegie Mellon University, Pittsburgh, PAView Profile , Taemin Kim Intel, Hillsboro, OR Intel, Hillsboro, ORView Profile , Suyog Gupta Google, Mountain View, CA Google, Mountain View, CAView Profile Authors Info & Claims ACM Journal on Emerging Technologies in Computing SystemsVolume 13Issue 3July 2017 Article No.: 30pp 1–3https://doi.org/10.1145/3022193Published:09 February 2017Publication History 0citation335DownloadsMetricsTotal Citations0Total Downloads335Last 12 Months17Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Yu Cao 0001, Xin Li 0001, Suyog Gupta |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2017 | DFM Evaluation Using IC Diagnosis DataabstractDesign for manufacturability rule evaluation using manufactured silicon (DREAMS) is a comprehensive methodology for evaluating the yield-preserving capabilities of a set of design for manufacturability (DFM) rules using the results of logic diagnosis performed on failed ICs. DREAMS is an improvement over prior art in that the distribution of rule violations over the diagnosis candidates and the entire design are taken into account along with the nature of the failure (e.g., bridge versus open) to appropriately weight the rules. Silicon and simulation results demonstrate the efficacy of the DREAMS methodology. Specifically, virtual data is used to demonstrate that the DFM rule most responsible for failure can be reliably identified even in light of the ambiguity inherent to a nonideal diagnostic resolution, and a corresponding rule-violation distribution that is counter-intuitive. We also show that the combination of physically aware diagnosis and the nature of the violated DFM rule can be used together to improve rule evaluation even further. Application of DREAMS to the diagnostic results from an in-production chip provides valuable insight in how specific DFM rules improve yield (or not) for a given design manufactured in particular facility. Finally, we also demonstrate that a significant artifact of DREAMS is a dramatic improvement in diagnostic resolution. This means that in addition to identifying the most ineffective DFM rule(s), validation of that outcome via physical failure analysis of failed chips can be eased due to the corresponding improvement in diagnostic resolution. R. D. (Shawn) Blanton, Fa Wang, Pranab K. Nag, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Machine Learning for Noise Sensor Placement and Full-Chip Voltage Emergency DetectionabstractPower supply fluctuation can be potential threat to the correct operations of processors, in the form of voltage emergency that happens when supply voltage drops below a certain threshold. Noise sensors (with either analog or digital outputs) can be placed in the nonfunction area of processors to detect voltage emergencies by monitoring the runtime voltage fluctuations. Our work addresses two important problems related to building a sensor-based voltage emergency detection system: 1) offline sensor placement, i.e., where to place the noise sensors so that the number and locations of sensors are optimized in order to strike a balance between design cost and chip reliability and 2) online voltage emergency detection, i.e., how to use these placed sensors to detect voltage emergencies in the hotspot locations. In this paper, we propose integrated solutions to these two problems, respectively, for analog and digital (more specifically, binary) sensor outputs, by exploiting the voltage correlation between the sensor candidate locations and the hotspot locations. For the analog case, we use the Group Lasso and an ordinary least squares approach; for the binary case, we integrate the Group Lasso and the SVM approach. Experimental results show that, our approach can achieve 2.3X-2.7X better voltage emergency detection results on average for analog outputs when compared to the state-of-the-art work; and for the binary case, on average our methodology can achieve up to 21% improvement in prediction accuracy compared to an approach called max-probability-no-prediction. Shupeng Sun, Xin Li 0001, Haifeng Qian, Pingqiang Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | C-YES: An Efficient Parametric Yield Estimation Approach for Analog and Mixed-Signal Circuits Based on Multicorner-Multiperformance CorrelationsabstractParametric yield estimation is a critical task for design and validation of analog and mixed-signal (AMS) circuits. However, the computational cost for yield estimation based on Monte Carlo (MC) analysis is often prohibitively high, especially when multiple circuit performances and/or environmental corners (e.g., voltage and temperature corners) are considered. In this paper, a novel statistical method named correlation-aided yield estimation (C-YES) is proposed to reduce the computational cost for parametric yield estimation. Our proposed approach exploits the fact that multiple circuit performances over different environmental corners are often correlated. Hence, we can accurately predict the performance value at one corner from the simulation results for other performances and/or corners. Based upon this observation, instead of running a large number of MC simulations to cover all performances and corners, an efficient algorithm is developed to select a small set of the most “informative” simulations that should be performed for yield estimation. Our numerical experiments show that for parametric yield estimation with multiple circuit performances and environmental corners, C-YES achieves 6.5-9.3× runtime speedups over other conventional methods. Hengliang Zhu, Xuan Zeng 0001, Dian Zhou, Ruey-Wen Liu, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Guest Editorial: Special Issue on Smart Homes, Buildings and InfrastructuresabstractNo abstract available. Xin Li 0001, Shiyan Hu 0001, Qi Zhu 0002 |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2017 | Training Fixed-Point Classifiers for On-Chip Low-Power ImplementationabstractIn this article, we develop several novel algorithms to train classifiers that can be implemented on chip with low-power fixed-point arithmetic with extremely small word length. These algorithms are based on Linear Discriminant Analysis (LDA), Support Vector Machine (SVM), and Logistic Regression (LR), and are referred to as LDA-FP, SVM-FP, and LR-FP, respectively. They incorporate the nonidealities (i.e., rounding and overflow) associated with fixed-point arithmetic into the offline training process so that the resulting classifiers are robust to these nonidealities. Mathematically, LDA-FP, SVM-FP, and LR-FP are formulated as mixed integer programming problems that can be robustly solved by the branch-and-bound methods described in this article. Our numerical experiments demonstrate that LDA-FP, SVM-FP, and LR-FP substantially outperform the conventional approaches for the emerging biomedical applications of brain decoding. Hassan Albalawi, Yuanning Li, Xin Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2017 | An FPGA-Based Hardware Accelerator for Traffic Sign DetectionabstractTraffic sign detection plays an important role in a number of practical applications, such as intelligent driver assistance and roadway inventory management. In order to process the large amount of data from either real-time videos or large off-line databases, a high-throughput traffic sign detection system is required. In this paper, we propose an FPGA-based hardware accelerator for traffic sign detection based on cascade classifiers. To maximize the throughput and power efficiency, we propose several novel ideas, including: 1) rearranged numerical operations; 2) shared image storage; 3) adaptive workload distribution; and 4) fast image block integration. The proposed design is evaluated on a Xilinx ZC706 board. When processing high-definition (1080p) video, it achieves the throughput of 126 frames/s and the energy efficiency of 0.041 J/frame. Weijing Shi, Xin Li 0001, Zhiyi Yu, Gary Overett |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | High-Dimensional and Multiple-Failure-Region Importance Sampling for SRAM Yield AnalysisabstractThe failure rate of static RAM (SRAM) cells is restricted to be extremely low to ensure sufficient high yield for the entire chip. In addition, multiple performances of interest and influences from peripherals make SRAM failure rate estimation a high-dimensional multiple-failure-region problem. This paper proposes a new method featuring a multistart-point sequential quadratic programming (SQP) framework to extend minimized norm importance sampling (IS) to address this problem. Failure regions in the variation space are first found by the low-discrepancy sampling sequence. Afterward, start points are generated in all identified failure regions and local optimizations based on SQP are invoked from these start points searching for the optimal shift vectors (OSVs). Based on the OSVs, a Gaussian mixture distorted distribution is constructed for IS. To further reduce the computational cost of IS while fully considering the influence of increasing dimensionality, an adaptive model training framework is proposed to keep high efficiency for both low- and high-dimensional problems. The experimental results show that the proposed method can not only approximate failure rate with high accuracy and efficiency in low-dimensional cases but also keep these features in high-dimensional ones. Mengshuo Wang, Changhao Yan, Xin Li 0001, Dian Zhou, Xuan Zeng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Thermal modeling for energy-efficient smart building with advanced overfitting mitigation techniqueabstractBuilding energy accounts large amount of the total energy consumption, and smart building energy control leads to high energy efficiency and significant energy savings. A compact and accurate building thermal model is important for designing the efficient energy control system. In this paper, we propose an accurate thermal behavior modeling technique for general and complicated buildings. This new modeling technique builds compact thermal model by system identification using temperature and power data obtained from EnergyPlus software, which can provide realistic temperature, weather and power data for buildings. In order to make the best use of data from EnergyPlus and avoid the overfitting problem associated with the system identification method, a cross-validation technique is employed to generate multiple thermal models to find the optimal model order. The final model is then generated by performing a regular system identification using the previously selected order. Experimental results from a case study of a 5-zone building have shown that the proposed method is able to find the optimal model order, and the building models built by the proposed method can achieve 1-3% average errors and less than 10-18% maximum errors for the estimation of zone temperatures for about a one year period. Wandi Liu, Hai Wang 0002, Hengyang Zhao, Shujuan Wang, Haibao Chen, Yuzhuo Fu, Jian Ma 0002, Xin Li 0001, Sheldon X.-D. Tan |
ASP-DAC | 8 |
| 2016 | Re-thinking polynomial optimization: Efficient programming of reconfigurable radio frequency (RF) systems by convexificationabstractReconfigurable radio frequency (RF) system has emerged as a promising avenue to achieve high communication performance while adapting to versatile commercial wireless environment. In this paper, we propose a novel technique to optimally program a reconfigurable RF system in order to achieve maximum performance and/or minimum power. Our key idea is to adopt an equation-based optimization method that relies on general-purpose, non-convex polynomial performance models to determine the optimal configurations of all tunable circuit blocks. Most importantly, our proposed approach guarantees to find the globally optimal solution of the non-convex polynomial programming problem by solving a sequence of convex semi-definite programming (SDP) problems based on convexification. A reconfigurable RF front-end example designed for WLAN 802.11g demonstrates that the proposed method successfully finds the globally optimal configuration, while other traditional techniques often converge to local optima. Fa Wang, Shihui Yin, Minhee Jun, Xin Li 0001, Tamal Mukherjee, Rohit Negi, Lawrence T. Pileggi |
ASP-DAC | 4 |
| 2016 | Efficient performance modeling of analog integrated circuits via kernel density based sparse regressionabstractWith the aggressive scaling of integrated circuit technology, analog performance modeling is facing enormous challenges due to high-dimensional variation space and expensive transistor-level simulation. In this paper, we propose a kernel density based sparse regression algorithm (KDSR) to accurately fit analog performance models where the modeling error is not simply Gaussian due to strong nonlinearity. The key idea of KDSR is to approximate the non-Gaussian likelihood function by using non-parametric kernel density estimation. Furthermore, we adopt Laplace distribution as our prior knowledge to enforce a sparse pattern for model coefficients. The unknown model coefficients are finally determined by using an EM type algorithm for maximum-a-posteriori (MAP) estimation. Our proposed method can be viewed as an iterative and weighted sparse regression algorithm that aims to reduce the estimation bias for model coefficients due to outliers. Our experimental results demonstrate that our proposed KDSR method can achieve superior accuracy over the conventional sparse regression method. Chenlei Fang, Qicheng Huang, Fan Yang 0001, Xuan Zeng 0001, Dian Zhou, Xin Li 0001 |
DAC | 6 |
| 2016 | Efficient performance modeling via Dual-Prior Bayesian Model Fusion for analog and mixed-signal circuitsabstractIn this paper, we propose a novel Dual-Prior Bayesian Model Fusion (DP-BMF) algorithm for performance modeling. Different from the previous BMF methods which use only one source of prior knowledge, DP-BMF takes advantage of multiple sources of prior knowledge to fully exploit the available information and, hence, further reduce the modeling cost. Based on a graphical model, an efficient Bayesian inference is developed to fuse two different prior models and combine the prior information with a small number of training samples to achieve high modeling accuracy. Several circuit examples demonstrate that the proposed method can achieve up to 1.83× cost reduction over the traditional one-prior BMF method without surrendering any accuracy. Qicheng Huang, Chenlei Fang, Fan Yang 0001, Xuan Zeng 0001, Dian Zhou, Xin Li 0001 |
DAC | 6 |
| 2016 | Correlated Bayesian Model Fusion: efficient performance modeling of large-scale tunable analog/RF integrated circuitsabstractTunable circuit has emerged as a promising methodology to address the grand challenge posed by process variations. Efficient high-dimensional performance modeling of tunable analog/RF circuits is an important yet challenging task. In this paper, we propose a novel performance modeling approach for tunable circuits, referred to as Correlated Bayesian Model Fusion (C-BMF). The key idea is to encode the correlation information for both model template and coefficient magnitude among different knob configurations by using a unified prior distribution. The prior distribution is then combined with a few simulation samples via Bayesian inference to efficiently determine the unknown model coefficients. Two circuit examples designed in a commercial 32nm SOI CMOS process demonstrate that C-BMF achieves more than 2× cost reduction over the traditional state-of-the-art modeling technique without surrendering any accuracy. Fa Wang, Xin Li 0001 |
DAC | 2 |
| 2016 | Efficient spatial variation modeling via robust dictionary learning
Changhai Liao, Jun Tao 0001, Xuan Zeng 0001, Yangfeng Su, Dian Zhou, Xin Li 0001 |
DATE | 6 |
| 2016 | Efficient statistical validation of machine learning systems for autonomous drivingabstractToday's automotive industry is making a bold move to equip vehicles with intelligent driver assistance features. A modern automobile is now equipped with a powerful computing platform to run multiple machine learning algorithms for environment perception (e.g., pedestrian detection) and motion control (e.g., vehicle stabilization). These machine learning systems must be highly robust with extremely small failure rate in order to ensure safe and reliable driving. In this paper, we propose a novel Subset Sampling (SUS) algorithm to efficiently validate a machine learning system. In particular, a Markov Chain Monte Carlo algorithm based on graph mapping is developed to accurately estimate the rare failure rate with a minimal amount of test data, thereby minimizing the validation cost. Our numerical experiments show that SUS achieves 15.2× runtime speed-up over the conventional brute-force Monte Carlo method. Weijing Shi, Mohamed Baker Alawieh, Xin Li 0001, Huafeng Yu, Nikos Aréchiga, Nobuyuki Tomatsu |
ICCAD | 3 |
| 2016 | Efficient statistical analysis for correlated rare failure events via asymptotic probability approximationabstractIn this paper, a novel Asymptotic Probability Approximation (APA) method is proposed to estimate the overall rare probability of correlated failure events for complex circuits containing a large number of replicated cells (e.g., SRAM bit-cells). The key idea of APA is to approximate the overall circuit failure rate based on a set of carefully defined failure events. An efficient Hierarchal Subset Simulation (H-SUS) method is developed to calculate the aforementioned failure rate and a statistical methodology is further proposed to estimate the confidence interval of APA. Our numerical experiments demonstrate that APA can accurately and reliably estimates the overall failure rate of correlated rare failure events involving more than 20,000 independent random variables. Handi Yu, Jun Tao 0001, Changhai Liao, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
ICCAD | 7 |
| 2016 | Cross-scale predictive dictionaries for image and video restorationabstractWe propose a novel signal model, based on sparse representations, that captures cross-scale features for visual signals. We show that cross-scale predictive model enables faster solutions to sparse approximation problems. This is achieved by first solving the sparse approximation problem for the downsampled signal and using the support of the solution to constrain the support at the original resolution. The speedups obtained are especially compelling for high-dimensional signals that require large dictionaries to provide precise sparse approximations. We demonstrate speedups in the order of 10-20x for denoising and up to 9x speed-ups for compressive sensing of images and videos. Vishwanath Saragadam, Aswin C. Sankaranarayanan, Xin Li 0001 |
ICIP | 3 |
| 2016 | Identifying systematic spatial failure patterns through wafer clusteringabstractIn this paper, we propose a novel methodology for detecting systematic spatial failure patterns at wafer level for yield learning. Our proposed methodology takes the testing results (i.e., pass or fail) of a number of dies over different wafers, cluster all these wafers according to their failures, and eventually identify the underlying spatial failure patterns. Several novel machine learning algorithms, including singular value decomposition, hierarchical clustering, dictionary learning, etc., are developed in order to make the proposed methodology robust to random failures. The efficacy of our proposed approach is demonstrated by an industrial data set. Mohamed Baker Alawieh, Fa Wang, Xin Li 0001 |
ISCAS | 3 |
| 2016 | Virtual temperature measurement for smart buildings via Bayesian model fusionabstractOne important goal of creating smart buildings is to offer highly comfortable services to the occupants at low cost. Real-time temperature measurement and monitoring is a critical task to facilitate high-quality service with low energy consumption and, hence, cost. In this paper, we propose a novel framework to accurately measure in-building temperature by using a small number of sensors. The key idea is to combine the prior knowledge on temperature statistics with a few sensor measurements and then predict the spatial temperature distribution by maximum-a-posteriori estimation. Our experimental results demonstrate that the average estimation error is less than 0.3 degree with very few sensors. Xiaoming Chen 0003, Xin Li 0001 |
ISCAS | 2 |
| 2016 | Diagnostic resolution improvement through learning-guided physical failure analysisabstractAn accurate and high-resolution diagnosis enables physical failure analysis (PFA) to identify and understand the root-cause of integrated-circuit failure. Despite many existing techniques for improving diagnosis, resolution is still far from ideal, which hinders PFA and other analyses. To address this challenge, we extend the capability of PADRE (physically-aware diagnostic resolution enhancement), a powerful machine learning based diagnosis resolution improvement technique, with a novel, active learning (AL) based PFA selection approach. An active-learning based PADRE (AL PADRE) selects the most useful defects for PFA in order to improve diagnostic resolution. AL PADRE provides an alternative to the normal PFA selection procedure, it improves the the accuracy of PADRE, and thus enables a more accurately improved resolution. AL PADRE is validated by both simulation-based experiment and silicon experiment. Simulation-based experiments show that by using AL PADRE, the number of PFAs required for increasing the accuracy to a stable level of 90% is reduced by more than 60% on average compared to baseline approach, and AL PADRE consistently outperforms the baseline approach for accuracy improvement in various scenarios. In the silicon experiment, by using AL PADRE, the number of chips needed to undergo PFA was reduced by more than 6x in order to increase diagnosis accuracy by more than 20%. Carlston Lim, Xin Li 0001, R. D. (Shawn) Blanton, M. Enamul Amyeen |
ITC | 3 |
| 2016 | Preface
Hongbo Fu 0001, Xin Li 0001, Lizhuang Ma, Jun-Hai Yong |
Comput. Graph. | 2 |
| 2016 | Editorial: Special Issue on The 14th International Conference on Computer-Aided Design and Computer Graphics (CAD/Graphics 2015)
Xin Li 0001, Sheldon X.-D. Tan, Yu Wang 0002 |
Integr. | 1 |
| 2016 | Energy-Constrained Distributed Learning and Classification by Exploiting Relative Relevance of Sensors' DataabstractWe consider the problem of communicating data from energy-constrained distributed sensors. To reduce energy requirements, we go beyond the source reconstruction problem classically addressed, and focus on the problem where the recipient wants to perform supervised learning and classification on the data received from the sensors. Restricting our attention to a noiseless communication setting under simplistic Gaussian source assumptions, we study supervised learning and classification under total energy limitations. The energy constraints are modeled in two ways: 1) a linear scaling and 2) an exponential scaling of energy with number of bits used for compression at sensors. We first assume that the underlying parameters for Gaussian distributions have already been learned, and obtain (with linear scaling, reverse-waterfilling-type) strategies for allocating energy, and thus, bits, across different sensors under these two models. Intuitively, these strategies allocate larger rates and energies to sensors that are more “relevant” for the classification goal. These strategies are used to obtain an achievable bound on the tradeoff between energy and error-probability (classification risk). We then provide an algorithm for learning the distribution-parameters of the sensor-data under energy constraints to arrive at high-reliability energy-allocation strategies, while enabling the energy-allocation algorithm to backtrack when the underlying distributions change, or when there is noise in sensed data that can push the algorithm toward a local minimum. Finally, we provide numerical results on energy-savings for classification of simulated data as well as neural data acquired from electrocorticography (ECoG) experiments. Majid Mahzoon, Christy Li, Xin Li 0001, Pulkit Grover |
IEEE J. Sel. Areas Commun. | 3 |
| 2016 | Modeling Random Telegraph Noise as a Randomness Source and its Application in True Random Number GenerationabstractThe random telegraph noise (RTN) is becoming more serious in advanced technologies. Due to the unpredictability of the physical phenomenon, RTN is a good randomness source for true random number generators (TRNG). In this paper, we build fundamental randomness models for TRNGs based on single trap- and multiple traps-induced RTN. We theoretically derive the autocorrelation coefficient, bias, and bit rate for RTN-based TRNGs. Two representative RTN-based TRNG schemes are simulated to verify the proposed randomness models. An oscillator-based TRNG is also studied based on the theoretical randomness model of multiple traps-induced RTN. We also provide basic guidelines for designing RTN-based TRNGs. Xiaoming Chen 0003, Boxun Li, Yu Wang 0002, Xin Li 0001, Yongpan Liu, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Efficient Hybrid Performance Modeling for Analog Circuits Using Hierarchical Shrinkage PriorsabstractEfficient performance modeling is an extremely important task for yield analysis and design optimization of analog circuits. In this paper, a novel regression modeling method based on hierarchical shrinkage priors is proposed to construct hybrid performance models with both high accuracy and low computational cost. In particular, the user-defined model templates derived from design equations and the general-purpose orthogonal polynomials are combined together to set up a hybrid dictionary. Next, in order to avoid over-shrinking large model coefficients, a novel regression method based on hierarchical shrinkage priors and variational Bayesian inference is adopted for model fitting. A rail-to-rail operational amplifier example demonstrates that the proposed method achieves up to 40% error reduction over other state-of-the-art approaches without increasing the modeling cost. Changhai Liao, Jun Tao 0001, Handi Yu, Zhangwen Tang, Yangfeng Su, Dian Zhou, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2016 | Efficient Spatial Variation Modeling of Nanoscale Integrated Circuits Via Hidden Markov TreeabstractIn this paper, we propose a novel spatial variation modeling method based on hidden Markov tree (HMT) for nanoscale integrated circuits, which could efficiently improve the accuracy of full-wafer/chip spatial variations recovery at extremely low measurement cost. Applying this method, HMT is introduced to set up a statistical model for coefficients after exploring the underlying correlated representation of the spatial variation in the frequency domain. Accordingly, two key inherent properties of the modeling coefficients, i.e., correlations and sparse presentations in the frequency domain, can be captured exactly and the modeling accuracy can be improved evidently. Then, maximum-a-posteriori estimation is applied to formulate the original problem as a convex optimization that could be solved efficiently and robustly. Numerical results based on industrial data demonstrate that the proposed method can achieve superior accuracy over other existing approaches including orthogonal matching pursuit, l1-norm regularization, and reweighted l1-norm regularization. Changhai Liao, Jun Tao 0001, Xuan Zeng 0001, Yangfeng Su, Dian Zhou, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | Harvesting Design Knowledge From the Internet: High-Dimensional Performance Tradeoff Modeling for Large-Scale Analog CircuitsabstractEfficiently optimizing large-scale, complex analog systems requires to know the performance tradeoffs for various analog circuit blocks. In this paper, we propose a radically new approach for analog performance tradeoff modeling. Our key idea is to broadly search the rich design knowledge from the Internet, and then mathematically encode the knowledge as high-dimensional performance tradeoff curves that are referred to as Pareto fronts in the literature. Toward this goal, several novel numerical algorithms, such as sparse regression and semi-infinite programming, are developed in order to construct the high-dimensional Pareto front model while guaranteeing its monotonicity. Our numerical examples demonstrate that the proposed modeling technique can accurately capture the high-dimensional Pareto fronts for large-scale analog systems (e.g., analog-to-digital converter) while most traditional methods are limited to low-dimensional Pareto front modeling of small circuit blocks without considering layout parasitics and manufacturing nonidealities. Jun Tao 0001, Changhai Liao, Xuan Zeng 0001, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Bayesian Model Fusion: Large-Scale Performance Modeling of Analog and Mixed-Signal Circuits by Reusing Early-Stage DataabstractEfficient performance modeling of today's analog and mixed-signal circuits is an important yet challenging task, due to the high-dimensional variation space and expensive circuit simulation. In this paper, we propose a novel performance modeling algorithm that is referred to as Bayesian model fusion (BMF) to address this challenge. The key idea of BMF is to borrow the information collected from an early stage (e.g., schematic level) to facilitate efficient performance modeling at a late stage (e.g., post layout). Such a goal is achieved by statistically modeling the performance correlation between early and late stages through Bayesian inference. Furthermore, to make the proposed BMF method of practical utility, four implementation issues, including: 1) prior mapping; 2) missing prior knowledge; 3) fast solver; and 4) prior and hyper-parameter selection, are carefully considered in this paper. Two circuit examples designed in a commercial 32 nm CMOS silicon on insulator process demonstrate that the proposed BMF method achieves up to 9× runtime speed-up over the traditional modeling technique without surrendering any accuracy. Fa Wang, Paolo Cachecho, Wangyang Zhang, Shupeng Sun, Xin Li 0001, Rouwaida Kanj, Chenjie Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Parasitic-Aware Common-Centroid FinFET Placement and Routing for Current-Ratio MatchingabstractThe FinFET technology is regarded as a better alternative for modern high-performance and low-power integrated-circuit design due to more effective channel control and lower power consumption. However, the gate-misalignment problem resulting from process variation and the parasitic resistance resulting from interconnecting wires based on the FinFET technology becomes even more severe compared with the conventional planar CMOS technology. Such gate misalignment and unwanted parasitic resistance may increase the threshold voltage and decrease the drain current of transistors. When applying the FinFET technology to analog circuit design, the variation of drain currents can destroy current-ratio matching among transistors and degrade circuit performance. In this article, we present the first FinFET placement and routing algorithms for layout generation of a common-centroid FinFET array to precisely match the current ratios among transistors. Experimental results show that the proposed matching-driven FinFET placement and routing algorithms can obtain the best current-ratio matching compared with the state-of-the-art common-centroid placer. Po-Hsun Wu, Mark Po-Hung Lin, Xin Li 0001, Tsung-Yi Ho |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Statistical Rare-Event Analysis and Parameter Guidance by Elite Learning Sample SelectionabstractAccurately estimating the failure region of rare events for memory-cell and analog circuit blocks under process variations is a challenging task. In this article, we propose a new statistical method, called EliteScope , to estimate the circuit failure rates in rare-event regions and to provide conditions of parameters to achieve targeted performance. The new method is based on the iterative blockade framework to reduce the number of samples, but consists of two new techniques to improve existing methods. First, the new approach employs an elite-learning sample-selection scheme, which can consider the effectiveness of samples and well coverage for the parameter space. As a result, it can reduce additional simulation costs by pruning less effective samples while keeping the accuracy of failure estimation. Second, the EliteScope identifies the failure regions in terms of parameter spaces to provide a good design guidance to accomplish the performance target. It applies variance-based feature selection to find the dominant parameters and then determine the in-spec boundaries of those parameters. We demonstrate the advantage of our proposed method using several memory and analog circuits with different numbers of process parameters. Experiments on four circuit examples show that EliteScope achieves a significant improvement on failure-region estimation in terms of accuracy and simulation cost over traditional approaches. The 16b 6T-SRAM column example also demonstrates that the new method is scalable for handling large problems with large numbers of process variables. Taeyoung Kim 0001, Hosoon Shin, Sheldon X.-D. Tan, Xin Li 0001, Haibao Chen, Hai Wang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2015 | SIPredict: Efficient post-layout waveform prediction via System IdentificationabstractIn this paper, we propose a post-layout waveform prediction method by System Identification (SI) based on the fact that the waveforms of pre-layout and post-layout are always correlated. Mathematical models are built to describe the relationships between the pre-layout and post-layout simulation results via SI techniques. The model parameters are calibrated by using the simulation results of the first few data points of pre-layout and post-layout stages. By taking the corresponding pre-layout simulation results as inputs of the calibrated models, the rest post-layout waveforms can thus be predicted as the output of the models. Several examples demonstrate the efficiency of the prediction, which helps the designers have a quick view of the post-layout waveforms in the design process. Qicheng Huang, Xiao Li 0002, Fan Yang 0001, Xuan Zeng 0001, Xin Li 0001 |
ASP-DAC | 5 |
| 2015 | Fast statistical analysis of rare failure events for memory circuits in high-dimensional variation spaceabstractAccurately estimating the rare failure rates for nanoscale memory circuits is a challenging task, especially when the variation space is high-dimensional. In this paper, we summarize two novel techniques to address this technical challenge. First, we describe a subset simulation (SUS) technique to estimate the rare failure rates for continuous performance metrics. The key idea of SUS is to express the rare failure probability of a given circuit as the product of several large conditional probabilities by introducing a number of intermediate failure events. These conditional probabilities can be efficiently estimated with a set of Markov chain Monte Carlo samples generated by a modified Metropolis algorithm. Second, to efficiently estimate the rare failure rates for discrete performance metrics, scaled-sigma sampling (SSS) can be used. SSS aims to generate random samples from a distorted probability distribution for which the standard deviation (i.e., sigma) is scaled up. Next, the failure rate is accurately estimated from these scaled random samples by using an analytical model derived from the theorem of “soft maximum”. Our experimental results of several nanoscale circuit examples demonstrate that SUS and SSS achieve significantly improved accuracy over other traditional techniques when the dimensionality of the variation space is more than a few hundred. Shupeng Sun, Xin Li 0001 |
ASP-DAC | 2 |
| 2015 | Accurate passivity-enforced macromodeling for RF circuits via iterative zero/pole update based on measurement dataabstractPassive macromodeling for RF circuit blocks is a critical task to facilitate efficient system-level simulation for large-scale RF systems (e.g., wireless transceivers). In this paper we propose a novel algorithm to find the optimal macromodel that minimizes the modeling error based on measurement data, while simultaneously guaranteeing passivity. The key idea is to attack the passive macromodeling problem by solving a sequence of convex semi-definite programming (SDP) problems. As such, the proposed method can iteratively find the optimal poles and zeros for macromodeling. Our experimental results with several commercial RF circuit examples demonstrate that the proposed macromodeling method reduces the modeling error by 1.31-2.74× over other conventional approaches. Ying-Chih Wang, Shihui Yin, Minhee Jun, Xin Li 0001, Lawrence T. Pileggi, Tamal Mukherjee, Rohit Negi |
ASP-DAC | 4 |
| 2015 | Efficient multivariate moment estimation via Bayesian model fusion for analog and mixed-signal circuitsabstractA critical-yet-challenging problem of analog/mixed-signal circuit validation in either pre-silicon or post-silicon stage is to estimate the parametric yield of the performances. In this paper, we propose a novel Bayesian model fusion method for efficient multivariate moment estimation of multiple correlated performance metrics by borrowing the prior knowledge from the early stage. The key idea is to model the multiple performance metrics as a jointly Gaussian distribution and encode the prior knowledge as a normal-Wishart distribution according to the theory of conjugate prior. The late-stage multivariate moments can be accurately estimated by Bayesian inference with very few late-stage samples. Several circuit examples demonstrate that the proposed method can achieve up to 16× cost reduction over the traditional method without surrendering any accuracy. Qicheng Huang, Chenlei Fang, Fan Yang 0001, Xuan Zeng 0001, Xin Li 0001 |
DAC | 5 |
| 2015 | Vortex: variation-aware training for memristor X-barabstractRecent advances in development of memristor devices and crossbar integration allow us to implement a low-power on-chip neuromorphic computing system (NCS) with small footprint. Training methods have been proposed to program the memristors in a crossbar by following existing training algorithms in neural network models. However, the robustness of these training methods has not been well investigated by taking into account the limits imposed by realistic hardware implementations. In this work, we present a quantitative analysis on the impact of device imperfections and circuit design constraints on the robustness of two popular training methods -- "close-loop on-device" (CLD) and "open-loop off-device" (OLD). A novel variation-aware training scheme, namely, Vortex, is then invented to enhance the training robustness of memristor crossbar-based NCS by actively compensating the impact of device variations and optimizing the mapping scheme from computations to crossbars. On average, Vortex can significantly improve the test rate by 29.6% and 26.4%, compared to the traditional OLD and CLD, respectively. Beiye Liu, Hai Li 0001, Yiran Chen 0001, Xin Li 0001, Qing Wu 0002, Tingwen Huang |
DAC | 4 |
| 2015 | A statistical methodology for noise sensor placement and full-chip voltage map generationabstractNoise margin violation, also known as voltage emergency induced by continuously reducing noise margin and increasing magnitude of current swings, is becoming a severe threat to the correct execution of applications in processors. Noise sensors can be placed in the non-function area of processors to detect such emergencies by monitoring runtime voltage fluctuations. In this work, we aim to accurately predict the voltage droops using a small set of sensors. We achieve our goal in two steps: We first propose a methodology via group lasso approach to select the optimal set of noise sensors, then build a practical model via ordinary least-squares fitting approach to predict the voltage in the function area of the chip, using the selected sensors in non-function area. Experiment results show that when compared to the full-chip voltage transient simulation, the prediction error of our model is much less than 0.01, and compared to prior work, our approach can achieve better error rates of voltage emergency detection (less than half). Shupeng Sun, Pingqiang Zhou, Xin Li 0001, Haifeng Qian |
DAC | 4 |
| 2015 | An EDA framework for large scale hybrid neuromorphic computing systemsabstractIn implementations of neuromorphic computing systems (NCS), memristor and its crossbar topology have been widely used to realize fully connected neural networks. However, many neural networks utilized in real applications often have a sparse connectivity, which is hard to be efficiently mapped to a crossbar structure. Moreover, the scale of the neural networks is normally much larger than that can be offered by the latest integration technology of memristor crossbars. In this work, we propose AutoNCS -- an EDA framework that can automate the NCS designs that combine memristor crossbars and discrete synapse modules. The connections of the neural networks are clustered to improve the utilization of the memristor elements in crossbar structures by taking into account the physical design cost of the NCS. Our results show that AutoNCS can substantially enhance the utilization efficiency of memristor crossbars while reducing the wirelength, area and delay of the physical designs of the NCS. Wei Wen 0003, Chi-Ruo Wu, Beiye Liu, Tsung-Yi Ho, Xin Li 0001, Yiran Chen 0001 |
DAC | 6 |
| 2015 | mTunes: efficient post-silicon tuning of mixed-signal/RF integrated circuits based on Markov decision processabstractUncertainty prevails in IC manufacturing and circuit operation. In particular, process variability has a huge impact on circuit performance, especially for mixed-signal/RF circuits, leading to unacceptable yields. Additionally, environmental uncertainties, such as temperature fluctuation and channel variation, further deteriorate performances in field. To combat variability, circuits are often made reconfigurable by adding tunable knobs to recover circuit performance in the post-manufacturing stage. However, as the number of knobs increases, knob tuning becomes challenging due to the huge search space. In fact, knob-tuning policies can have an observable impact on final performance and power consumption. In this paper, we propose mTunes, a method based on the Markov decision process for dynamically choosing the "right" knob tuning sub-routine from a pre-defined set achieving a balance between performance and power constraints. The proposed method has been applied to a reconfigurable RF front-end design, showing 60% improvement in yield compared to static tuning policies. Manzil Zaheer, Fa Wang, Chenjie Gu, Xin Li 0001 |
DAC | 4 |
| 2015 | Efficient bit error rate estimation for high-speed link by Bayesian model fusion
Chenlei Fang, Qicheng Huang, Fan Yang 0001, Xuan Zeng 0001, Xin Li 0001, Chenjie Gu |
DATE | 5 |
| 2015 | A fast spatial variation modeling algorithm for efficient test cost reduction of analog/RF circuits
Hugo R. Gonçalves, Xin Li 0001, Miguel Correia 0002, Vítor Grade Tavares, John M. Carulli Jr., Kenneth M. Butler |
DATE | 2 |
| 2015 | Phase Noise Impairment and Environment-Adaptable Fast (EAF) Optimization for Programming of Reconfigurable Radio Frequency (RF) ReceiversabstractIn order to support a multi-standard platform, a reconfigurable RF front-end needs an optimal configuration that adapts to a dynamic communication condition. To find an optimal configuration efficiently, we previously proposed the Environment-Adaptable Fast (EAF) optimization in terms of the RF impairments of gain, nonlinearity and noise figure. However, this preliminary study did not include the important impairment of phase noise. In this paper, we extend the EAF optimization algorithm to phase noise impairment in a reconfigurable RF front-end. In this study, we will propose a novel statistical estimation tool for obtaining phase noise spectrum information with the Interpolated FIR (IFIR) model and the least mean squares (LMS) adaptive algorithm. We formulate the calculation of the Signal-to-Interference-and- Noise Ratio (SINR) which hastens the optimization process. Phase noise is included in the SINR calculation. We demonstrate the efficient performance of the EAF optimization method even with phase noise impairment. This study shows that while finding an optimal configuration, the EAF optimization significantly reduces simulation time compared to the other four conventional optimization methods. Minhee Jun, Rohit Negi, Shihui Yin, Fa Wang, Megha Sunny, Tamal Mukherjee, Xin Li 0001 |
GLOBECOM | 7 |
| 2015 | EDA Challenges for Memristor-Crossbar based Neuromorphic ComputingabstractThe increasing gap between the high data processing capability of modern computing systems and the limited memory bandwidth motivated the recent significant research on neuromorphic computing systems (NCS), which are inspired from the working mechanism of human brains. Discovery of memristor further accelerates engineering realization of NCS by leveraging the similarity between synaptic connections in neural networks and programming weight of the memristor. However, to achieve a stable large-scale NCS for practical applications, many essential EDA design challenges still need to be overcome especially the state-of-the-art memristor crossbar structure is adopted. In this paper, we summarize some of our recent published works about enhancing the design robustness and efficiency of memristor crossbar based NCS. The experiments show that the impacts of noises generated by process variations and the IR-drop over the crossbar can be effectively suppressed by our noise-eliminating training method and IR-drop compensation technique. Moreover, our network clustering techniques can alleviate the challenges of limited crossbar scale and routing congestion in NCS implementations. Beiye Liu, Wei Wen 0003, Yiran Chen 0001, Xin Li 0001, Chi-Ruo Wu, Tsung-Yi Ho |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | Statistical Learning in Chip (SLIC)abstractDespite best efforts, integrated systems are “born” (manufactured) with a unique `personality' that stems from our inability to precisely fabricate their underlying circuits, and create software a priori for controlling the resulting uncertainty. It is possible to use sophisticated test methods to identify the best-performing systems but this would result in unacceptable yields and correspondingly high costs. The system personality is further shaped by its environment (e.g., temperature, noise and supply voltage) and usage (i.e., the frequency and type of applications executed), and since both can fluctuate over time, so can the system's personality. Systems also “grow old” and degrade due to various wear-out mechanisms (e.g., negative-bias temperature instability), and unexpectedly due to various early-life failure sources. These “nature and nurture” influences make it extremely difficult to design a system that will operate optimally for all possible personalities. To address this challenge, we propose to develop statistical learning in-chip (SLIC). SLIC is a holistic approach to integrated system design based on continuously learning key personality traits on-line, for self-evolving a system to a state that optimizes performance hierarchically across the circuit, platform, and application levels. SLIC will not only optimize integrated-system performance but also reduce costs through yield enhancement since systems that would have before been deemed to have weak personalities (unreliable, faulty, etc.) can now be recovered through the use of SLIC. R. D. (Shawn) Blanton, Xin Li 0001, Ken Mai, Diana Marculescu, Radu Marculescu, Jeyanandh Paramesh, Jeff G. Schneider, Donald E. Thomas |
ICCAD | 2 |
| 2015 | From Robust Chip to Smart Building: CAD Algorithms and Methodologies for Uncertainty Analysis of Building PerformanceabstractBuildings consume about 40% of the total energy use in the U.S. and, hence, accurately modeling, analyzing and optimizing building energy is considered as an extremely important task today. Towards this goal, uncertainty/sensitivity analysis has been proposed to identify the critical physical and environmental parameters contributing to building energy consumption. In this paper, we propose to apply sparse regression techniques to uncertainty/sensitivity analysis of smart buildings. We consider the orthogonal matching pursuit (OMP) algorithm as a case study to demonstrate its superior efficacy over other conventional approaches. Experimental results reveal that OMP achieves up to 18.6× runtime speedups over the conventional least-squares fitting method without surrendering any accuracy. Xiaoming Chen 0003, Xin Li 0001, Sheldon X.-D. Tan |
ICCAD | 2 |
| 2015 | Co-Learning Bayesian Model Fusion: Efficient Performance Modeling of Analog and Mixed-Signal Circuits Using Side InformationabstractEfficient performance modeling of today's analog and mixed-signal (AMS) circuits is an important yet challenging task. In this paper, we propose a novel performance modeling algorithm that is referred to as Co-Learning Bayesian Model Fusion (CL-BMF). The key idea of CL-BMF is to take advantage of the additional information collected from simulation and/or measurement to reduce the performance modeling cost. Different from the traditional performance modeling approaches which focus on the prior information of model coefficients (i.e. the coefficient side information) only, CL-BMF takes advantage of another new form of prior knowledge: the performance side information. In particular, CL-BMF combines the coefficient side information, the performance side information and a small number of training samples through Bayesian inference based on a graphical model. Two circuit examples designed in a commercial 32nm SOI CMOS process demonstrate that CL-BMF achieves up to 5× speed-up over other state-of-the-art performance modeling techniques without surrendering any accuracy. Fa Wang, Manzil Zaheer, Xin Li 0001, Jean-Olivier Plouchart, Alberto Valdes-Garcia |
ICCAD | 3 |
| 2015 | Learning Based Compact Thermal Modeling for Energy-Efficient Smart Building Management: (invited)abstractIn this article, we propose a new behavioral thermal modeling method for fast building performance analysis, which is critical for energy-efficient smart building control and management. The new approach is based on two recurrent neutral network architecture to obtain the compact nonlinear thermal models for complicated building. We start with a more realistic building simulation program, EnergyPlus, from Department of Energy, to model some practical buildings such as office buildings and data centers. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). In this work, we apply two recurrent neural network (RNN) architectures to build the non-linear compact thermal model of the building: one is non-linear state-space RNN architecture (NLSS), which has global feedbacks, and the other one is Elman's RNN architecture (ELNN), which has local feedbacks in each layer. We give a simple formula to calculate the RNN layer number, layer size to configure RNN architecture to avoid overfitting and underfitting problems. A cross-validation based training technique is further applied to improve predictable accuracy of models. Experimental results from a case study of three buildings show that ELNN and NLSS can both build very accurate building thermal models for the 2-zone and 5-zone building cases: both of them have average errors from around 1% to 1.5% for the two buildings. For the more complex 6-zone building case, ELNN outperforms NLSS with maximum errors 16% against 23%. But both methods have 2.2% average errors. Hengyang Zhao, Daniel Quach, Shujuan Wang, Hai Wang 0002, Haibao Chen, Xin Li 0001, Sheldon X.-D. Tan |
ICCAD | 6 |
| 2015 | Optimizing Boolean embedding matrix for compressive sensing in RRAM crossbarabstractThe emerging resistive random-access-memory (RRAM) crossbar provides an intrinsic fabric for matrix-vector multiplication, which can be leveraged as power efficient linear embedding hardware for data analytics such as compressive sensing. As the matrix elements are represented by resistance of RRAM cells, it imposes constraints for the embedding matrix due to limited RRAM programming resolution. A random Boolean embedding can be efficiently mapped to the RRAM crossbar but suffers from poor performance. Learning-based embedding matrices can deliver optimized performance but are continuous-valued which prevents it from being mapped to RRAM crossbar structure directly. In this paper, we have proposed one algorithm that can find an optimal Boolean embedding matrix for a given learned real-valued embedding matrix, so that it can be effectively mapped to the RRAM crossbar structure while high performance is preserved. The numerical experiments demonstrate that the proposed optimized Boolean embedding can reduce the embedding distortion by 2.7x, and image recovery error by 2.5x compared to the random Boolean embedding, both mapped on RRAM crossbar. In addition, optimized Boolean embedding on RRAM crossbar exhibits 10x faster speed, 17x better energy efficiency, and three orders of magnitude smaller area with slight accuracy penalty, when compared to the optimized real-valued embedding on CMOS ASIC platform. Yuhao Wang 0002, Xin Li 0001, Hao Yu 0001, Leibin Ni, Chuliang Weng, Junfeng Zhao 0003 |
ISLPED | 2 |
| 2015 | Common-Centroid FinFET Placement Considering the Impact of Gate MisalignmentabstractThe FinFET technology has been regarded as a better alternative among different device technologies at 22nm node and beyond due to more effective channel control and lower power consumption. However, the gate misalignment problem resulting from process variation based on the FinFET technology becomes even severer compared with the conventional planar CMOS technology. Such misalignment may increase the threshold voltage and decrease the drain current of a single transistor. When applying the FinFET technology to analog circuit design, the variation of drain currents will destroy the current matching among transistors and degrade the circuit performance. In this paper, we present the first FinFET placement technique for analog circuits considering the impact of gate misalignment together with systematic and random mismatch. Experimental results show that the proposed algorithms can obtain an optimized common-centroid FinFET placement with much better current matching. Po-Hsun Wu, Mark Po-Hung Lin, Xin Li 0001, Tsung-Yi Ho |
ISPD | 3 |
| 2015 | Fast Statistical Analysis of Rare Circuit Failure Events via Scaled-Sigma Sampling for High-Dimensional Variation SpaceabstractAccurately estimating the rare failure rates for nanoscale circuit blocks (e.g., static random-access memory, D flip-flop, etc.) is a challenging task, especially when the variation space is high-dimensional. In this paper, we propose a novel scaled-sigma sampling (SSS) method to address this technical challenge. The key idea of SSS is to generate random samples from a distorted distribution for which the standard deviation (i.e., sigma) is scaled up. Next, the failure rate is accurately estimated from these scaled random samples by using an analytical model derived from the theorem of “soft maximum.” Our proposed SSS method can simultaneously estimate the rare failure rates for multiple performances and/or specifications with only a single set of transistor-level simulations. To quantitatively assess the accuracy of SSS, we estimate the confidence interval of SSS based on bootstrap. Several circuit examples designed in nanoscale technologies demonstrate that the proposed SSS method achieves significantly better accuracy than the traditional importance sampling technique when the dimensionality of the variation space is more than a few hundred. Shupeng Sun, Xin Li 0001, Hongzhou Liu, Kangsheng Luo, Ben Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | A Novel Analog Physical Synthesis Methodology Integrating Existent Design ExpertiseabstractAnalog layout design has been a manual, time-consuming, and error-prone task for decades. To speed up layout design time for a new design, analog layout designers prefer referring to legacy designs and layouts rather than starting from scratch, or thoroughly applying placement and routing tools because legacy layouts contain pretty much design expertise. Motivated by such layout design process, this paper presents the first knowledge-based physical synthesis methodology to generate new layouts by integrating existent design expertise. The proposed approach can automatically analyze legacy design data including circuits, layouts, and constraints, extract matched sub-circuits between new and legacy designs, and generate multiple layouts for the new design by utilizing the quality-approved legacy layouts as much as possible. Experimental results show that the proposed methodology can achieve high layout reusage rate, and hence the designers' layout preference can be successfully reserved. Po-Hsun Wu, Mark Po-Hung Lin, Tung-Chieh Chen, Ching-Feng Yeh, Xin Li 0001, Tsung-Yi Ho |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2015 | Efficient Transient Analysis of Power Delivery Network With Clock/Power Gating by Sparse ApproximationabstractTransient analysis of large-scale power delivery network (PDN) is a critical task to ensure the functional correctness and desired performance of today's integrated circuits (ICs), especially if significant transient noises are induced by clock and/or power gating due to the utilization of extensive power management. In this paper, we propose an efficient algorithm for PDN transient analysis based on sparse approximation. The key idea is to exploit the fact that the transient response caused by clock/power gating is often localized and the voltages at many other “inactive” nodes are almost unchanged, thereby rendering a unique sparse structure. By taking advantage of the underlying sparsity of the solution structure, a modified conjugate gradient algorithm is developed and tuned to efficiently solve the PDN analysis problem with low computational cost. Our numerical experiments based on standard benchmarks demonstrate that the proposed transient analysis with sparse approximation offers up to 2.2× runtime speedup over other traditional methods, while simultaneously achieving similar accuracy. Hengliang Zhu, Yuanzhe Wang, Frank Liu 0001, Xin Li 0001, Xuan Zeng 0001, Peter Feldmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Toward efficient programming of reconfigurable radio frequency (RF) receiversabstractReconfigurable radio frequency (RF) system is an emerging component to mitigate the growing engineering cost for wireless chip design. In this paper, we propose a new methodology for efficient programming of reconfigurable RF receiver. The proposed method is facilitated by two novel techniques: two-phase relaxation search and Pareto-based search space reduction. Our numerical experiments demonstrate that the proposed methodology is more robust (i.e., close to global optimum) and/or efficient (i.e., with low computational cost) than other traditional algorithms based on either local relaxation or simulated annealing. Jun Tao 0001, Ying-Chih Wang, Minhee Jun, Xin Li 0001, Rohit Negi, Tamal Mukherjee, Lawrence T. Pileggi |
ASP-DAC | 4 |
| 2014 | Computer-Aided Design of Machine Learning Algorithm: Training Fixed-Point Classifier for On-Chip Low-Power ImplementationabstractIn this paper, we propose a novel linear discriminant analysis algorithm, referred to as LDA-FP, to train on-chip classifiers that can be implemented with low-power fixed-point arithmetic with extremely small word length. LDA-FP incorporates the non-idealities (i.e., rounding and overflow) associated with fixed-point arithmetic into the training process so that the resulting classifiers are robust to these non-idealities. Mathematically, LDA-FP is formulated as a mixed integer programming problem that can be efficiently solved by a novel branch-and-bound method proposed in this paper. Our numerical experiments demonstrate that LDA-FP substantially outperforms the conventional approach for the emerging biomedical application of brain computer interface. Hassan Albalawi, Yuanning Li, Xin Li 0001 |
DAC | 3 |
| 2014 | BMF-BD: Bayesian Model Fusion on Bernoulli Distribution for Efficient Yield Estimation of Integrated CircuitsabstractAccurate yield estimation is one of the important yet challenging tasks for both pre-silicon verification and post-silicon validation. In this paper, we propose a novel method of Bayesian model fusion on Bernoulli distribution (BMF-BD) for efficient yield estimation at the late stage by borrowing the prior knowledge from an early stage. BMF-BD is particularly developed to handle the cases where the pre-silicon simulation and/or post-silicon measurement results are binary: either "pass" or "fail". The key idea is to model the binary simulation/measurement outcome as a Bernoulli distribution and then encode the prior knowledge as a Beta distribution based on the theory of conjugate prior. As such, the late-stage yield can be accurately estimated through Bayesian inference with very few late-stage samples. Several circuit examples demonstrate that BMF-BD achieves up to 10× cost reduction over the conventional estimator without surrendering any accuracy. Chenlei Fang, Fan Yang 0001, Xuan Zeng 0001, Xin Li 0001 |
DAC | 4 |
| 2014 | Verification based ECG biometrics with cardiac irregular conditions using heartbeat level and segment level information fusionabstractWe propose an ECG based robust human verification system for both healthy and cardiac irregular conditions using the heartbeat level and segment level information fusion. At the heartbeat level, we first propose a novel beat normalization and outlier removal algorithm after peak detection to extract normalized representative beats. Then after principal component analysis (PCA), we apply linear discriminant analysis (LDA) and within-class covariance normalization (WCCN) for beat variability compensation followed by cosine similarity and Snorm as scoring. At the segment level, we adopt the hierarchical Dirichlet process auto-regressive hidden Markov model (HDP-AR-HMM) in the Bayesian non-parametric framework for unsupervised joint segmentation and clustering without any peak detection. It automatically decodes each raw signal into a string vector. We then apply n-gram language model and hypothesis testing for scoring. Combining the aforementioned two subsystems together further improved the performance and outperformed the PCA baseline by 25% relatively on the PTB database. Ming Li 0026, Xin Li 0001 |
ICASSP | 2 |
| 2014 | Reduction and IR-drop compensations techniques for reliable neuromorphic computing systemsabstractNeuromorphic computing system (NCS) is a promising architecture to combat the well-known memory bottleneck in Von Neumann architecture. The recent breakthrough on memristor devices made an important step toward realizing a low-power, small-footprint NCS on-a-chip. However, the currently low manufacturing reliability of nano-devices and the voltage IR-drop along metal wires and memristors arrays severely limits the scale of memristor crossbar based NCS and hinders the design scalability. In this work, we propose a novel system reduction scheme that significantly lowers the required dimension of the memristor crossbars in NCS while maintaining high computing accuracy. An IR-drop compensation technique is also proposed to overcome the adverse impacts of the wire resistance and the sneak-path problem in large memristor crossbar designs. Our simulation results show that the proposed techniques can improve computing accuracy by 27.0% and 38.7% less circuit area compared to the original NCS design. Beiye Liu, Hai Li 0001, Yiran Chen 0001, Xin Li 0001, Tingwen Huang, Qing Wu 0002, Mark Barnell |
ICCAD | 4 |
| 2014 | Fast statistical analysis of rare circuit failure events via subset simulation in high-dimensional variation spaceabstractIn this paper, we propose a novel subset simulation (SUS) technique to efficiently estimate the rare failure rate for nanoscale circuit blocks (e.g., SRAM, DFF, etc.) in high-dimensional variation space. The key idea of SUS is to express the rare failure probability of a given circuit as the product of several large conditional probabilities by introducing a number of intermediate failure events. These conditional probabilities can be efficiently estimated with a set of Markov chain Monte Carlo samples generated by a modified Metropolis algorithm, and then used to calculate the rare failure rate of the circuit. To quantitatively assess the accuracy of SUS, a statistical methodology is further proposed to accurately estimate the confidence interval of SUS based on the theory of Markov chain Monte Carlo simulation. Our experimental results of two nanoscale circuit examples demonstrate that SUS achieves significantly enhanced accuracy over other traditional techniques when the dimensionality of the variation space is more than a few hundred. Shupeng Sun, Xin Li 0001 |
ICCAD | 2 |
| 2014 | MPME-DP: multi-population moment estimation via dirichlet process for efficient validation of analog/mixed-signal circuitsabstractMoment estimation is one of the most important tasks to appropriately characterize the performance variability of today's nanoscale integrated circuits. In this paper, we propose an efficient algorithm of multi-population moment estimation via Dirichlet Process (MPME-DP) for validation of analog and mixed-signal circuits with extremely small sample size. The key idea is to partition all populations (e.g., different environmental conditions, setup configurations, etc.) into groups. The populations within the same group are similar and their common knowledge can be extracted to improve the accuracy of moment estimation. As will be demonstrated by the silicon measurement data of a high-speed I/O link, MPME-DP reduces the moment estimation error by up to 65% compared to other conventional estimators. Manzil Zaheer, Xin Li 0001, Chenjie Gu |
ICCAD | 2 |
| 2014 | DALM-SVD: Accelerated sparse coding through singular value decomposition of the dictionaryabstractSparse coding techniques have seen an increasing range of applications in recent years, especially in the area of image processing. In particular, sparse coding using ℓ1-regularization has been efficiently solved with the Augmented Lagrangian (AL) applied to its dual formulation (DALM). This paper proposes the decomposition of the dictionary matrix in its Singular Value/Vector form in order to simplify and speed-up the implementation of the DALM algorithm. Furthermore, we propose an update rule for the penalty parameter used in AL methods that improves the convergence rate. The SVD of the dictionary matrix is done as a pre-processing step prior to the sparse coding, and thus the method is better suited for applications where the same dictionary is reused for several sparse recovery steps, such as block image processing. Hugo R. Gonçalves, Miguel Correia 0002, Xin Li 0001, Aswin C. Sankaranarayanan, Vítor Grade Tavares |
ICIP | 3 |
| 2014 | Bayesian model fusion: Enabling test cost reduction of analog/RF circuits via wafer-level spatial variation modelingabstractIn this paper, a novel Bayesian model fusion (BMF) method is proposed for test cost reduction based on wafer-level spatial variation modeling. BMF relies on the assumption that a large number of wafers of the same circuit design (e.g., all wafers from the same lot) share a similar spatial pattern. Hence, the measurement data from one wafer can be borrowed to model the spatial variation of other wafers via Bayesian inference. By applying the Sherman-Morrison-Woodbury formula, a fast numerical algorithm is derived to reduce the computational cost of BMF for practical test applications. Furthermore, a new test methodology is developed based on BMF and it closely monitors the escape rate and yield loss. As is demonstrated by the wafer probe measurement data of an industrial RF transceiver, BMF achieves 1.125× reduction in test cost and 2.6× reduction in yield loss, compared to the conventional approach based on virtual probe (VP). Shanghang Zhang, Xin Li 0001, R. D. (Shawn) Blanton, José Machado da Silva, John M. Carulli Jr., Kenneth M. Butler |
ITC | 2 |
| 2014 | Multiple-Population Moment Estimation: Exploiting Interpopulation Correlation for Efficient Moment Estimation in Analog/Mixed-Signal ValidationabstractMoment estimation is an important problem during circuit validation, in both presilicon and postsilicon stages. From the estimated moments, the probability of failure and parametric yield can be estimated at each circuit configuration and corner, and these metrics are used for design optimization and making product qualification decisions. The problem is especially difficult if only a very small sample size is allowed for measurement or simulation, as is the case for complex analog/mixed-signal circuits. In this paper, we propose an efficient moment estimation method, called multiple-population moment estimation (MPME), that significantly improves estimation accuracy under small sample size. The key idea is to leverage the data collected under different corners/configurations to improve the accuracy of moment estimation at each individual corner/configuration. Mathematically, we employ the hierarchical Bayesian framework to exploit the underlying correlation in the data. We apply the proposed method to several datasets including postsilicon measurements of a commercial high-speed I/O link, and demonstrate an average error reduction of up to 2×, which can be equivalently translated to significant reduction of validation time and cost. Chenjie Gu, Manzil Zaheer, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | Efficient moment estimation with extremely small sample size via bayesian inference for analog/mixed-signal validationabstractA critical problem in pre-Silicon and post-Silicon validation of analog/mixed-signal circuits is to estimate the distribution of circuit performances, from which the probability of failure and parametric yield can be estimated at all circuit configurations and corners. With extremely small sample size, traditional estimators are only capable of achieving a very low confidence level, leading to either over-validation or under-validation. In this paper, we propose a multi-population moment estimation method that significantly improves estimation accuracy under small sample size. In fact, the proposed estimator is theoretically guaranteed to outperform usual moment estimators. The key idea is to exploit the fact that simulation and measurement data collected under different circuit configurations and corners can be correlated, and are conditionally independent. We exploit such correlation among different populations by employing a Bayesian framework, i.e., by learning a prior distribution and applying maximum a posteriori estimation using the prior. We apply the proposed method to several datasets including post-silicon measurements of a commercial high-speed I/O link, and demonstrate an average error reduction of up to 2x, which can be equivalently translated to significant reduction of validation time and cost. Chenjie Gu, Eli Chiprout, Xin Li 0001 |
DAC | 3 |
| 2013 | Bayesian model fusion: large-scale performance modeling of analog and mixed-signal circuits by reusing early-stage dataabstractEfficient high-dimensional performance modeling of today's complex analog and mixed-signal (AMS) circuits with large-scale process variations is an important yet challenging task. In this paper, we propose a novel performance modeling algorithm that is referred to as Bayesian Model Fusion (BMF). Our key idea is to borrow the simulation data generated from an early stage (e.g., schematic level) to facilitate efficient high-dimensional performance modeling at a late stage (e.g., post layout) with low computational cost. Such a goal is achieved by statistically modeling the performance correlation between early and late stages through Bayesian inference. Several circuit examples designed in a commercial 32nm CMOS process demonstrate that BMF achieves up to 9x runtime speedup over the traditional modeling technique without surrendering any accuracy. Fa Wang, Wangyang Zhang, Shupeng Sun, Xin Li 0001, Chenjie Gu |
DAC | 4 |
| 2013 | Automatic clustering of wafer spatial signaturesabstractIn this paper, we propose a methodology based on unsupervised learning for automatic clustering of wafer spatial signatures to aid yield improvement. Our proposed methodology is based on three steps. First, we apply sparse regression to automatically capture wafer spatial signatures by a small number of features. Next, we apply an unsupervised hierarchical clustering algorithm to divide wafers into a few clusters where all wafers within the same cluster are similar. Finally, we develop a modified L-method to determine the appropriate number of clusters from the hierarchical clustering result. The accuracy of the proposed methodology is demonstrated by several industrial data sets of silicon measurements. Wangyang Zhang, Xin Li 0001, Sharad Saxena, Andrzej J. Strojwas, Rob A. Rutenbar |
DAC | 2 |
| 2013 | DREAMS: DFM rule EvAluation using manufactured siliconabstractDREAMS (DFM Rule EvAluation using Manufactured Silicon) is a comprehensive methodology for evaluating the yield-preserving capabilities of a set of DFM (design for manufacturability) rules using the results of logic diagnosis performed on failed ICs. DREAMS is an improvement over prior art in that the distribution of rule violations over the diagnosis candidates and the entire design are taken into account along with the nature of the failure (e.g., bridge versus open) to appropriately weight the rules. Silicon and simulation results demonstrate the efficacy of the DREAMS methodology. Specifically, virtual data is used to demonstrate that the DFM rule most responsible for failure can be reliably identified even in light of the ambiguity inherent to a nonideal diagnostic resolution, and a corresponding rule-violation distribution that is counter-intuitive. We also show that the combination of physically-aware diagnosis and the nature of the violated DFM rule can be used together to improve rule evaluation even further. Application of DREAMS to the diagnostic results from an in-production chip provides valuable insight in how specific DFM rules improve yield (or not) for a given design manufactured in particular facility. Finally, we also demonstrate that a significant artifact of DREAMS is a dramatic improvement in diagnostic resolution. This means that in addition to identifying the most ineffective DFM rule(s), validation of that outcome via physical failure analysis of failed chips can be eased due to the corresponding improvement in diagnostic resolution. R. D. (Shawn) Blanton, Fa Wang, Pranab K. Nag, Xin Li 0001 |
ICCAD | 6 |
| 2013 | Bayesian model fusion: a statistical framework for efficient pre-silicon validation and post-silicon tuning of complex analog and mixed-signal circuitsabstractIn this paper, we describe a novel statistical framework, referred to as Bayesian Model Fusion (BMF), that allows us to minimize the simulation and/or measurement cost for both pre-silicon validation and post-silicon tuning of analog and mixed-signal (AMS) circuits with consideration of large-scale process variations. The BMF technique is motivated by the fact that today's AMS design cycle typically spans multiple stages (e.g., schematic design, layout design, first tape-out, second tape-out, etc.). Hence, we can reuse the simulation and/or measurement data collected at an early stage to facilitate efficient validation and tuning of AMS circuits with a minimal amount of data at the late stage. The efficacy of BMF is demonstrated by using several industrial circuit examples. Xin Li 0001, Fa Wang, Shupeng Sun, Chenjie Gu |
ICCAD | 1 |
| 2013 | Fast statistical analysis of rare circuit failure events via scaled-sigma sampling for high-dimensional variation spaceabstractAccurately estimating the rare failure rates for nanoscale circuit blocks (e.g., SRAM, DFF, etc.) is a challenging task, especially when the variation space is high-dimensional. In this paper, we propose a novel scaled-sigma sampling (SSS) method to address this technical challenge. The key idea of SSS is to generate random samples from a distorted distribution for which the standard deviation (i.e., sigma) is scaled up. Next, the failure rate is accurately estimated from these scaled random samples by using an analytical model derived from the theorem of “soft maximum”. Several circuit examples designed in nanoscale technologies demonstrate that the proposed SSS method achieves superior accuracy over the traditional importance sampling technique when the dimensionality of the variation space is more than a few hundred. Shupeng Sun, Xin Li 0001, Hongzhou Liu, Kangsheng Luo, Ben Gu |
ICCAD | 2 |
| 2013 | Test data analytics - Exploring spatial and test-item correlations in production test dataabstractThe discovery of patterns and correlations hidden in the test data could help reduce test time and cost. In this paper, we propose a methodology and supporting statistical regression tools that can exploit and utilize both spatial and inter-test-item correlations in the test data for test time and cost reduction. We first describe a statistical regression method, called group lasso, which can identify inter-test-item correlations from test data. After learning such correlations, some test items can be identified for removal from the test program without compromising test quality. An extended version of this method, weighted group lasso, allows taking into account the distinct test time/cost of each individual test item in the formulation as a weighted optimization problem. As a result, its solution would favor more costly test items for removal from the test program. We further integrate weighted group lasso with another statistical regression technique, virtual probe, which can learn spatial correlations of test data across a wafer. The integrated method could then utilize both spatial and inter-test-item correlations to maximize the number of test items whose values can be predicted without measurement. Experimental results of a high-volume industrial device show that utilizing both spatial and inter-test-item correlations can help reduce test time by up to 55%. Chun-Kai Hsu, Fan Lin, Kwang-Ting Cheng, Wangyang Zhang, Xin Li 0001, John M. Carulli Jr., Kenneth M. Butler |
ITC | 5 |
| 2013 | PADRE: Physically-Aware Diagnostic Resolution EnhancementabstractDiagnosis is the first step of IC failure analysis. The conventional objective of identifying the failure locations has been augmented with various physically-aware techniques that are intended to improve both diagnostic resolution and accuracy. Despite these advances, it is often the case however that resolution, i.e., the number of locations or candidates reported by diagnosis, exceeds the number of actual failing locations. Imperfect resolution greatly hinders any follow-on, information-extraction analyses (e.g., physical failure analysis, volume diagnosis, etc.) due to the resulting ambiguity. To address this major challenge, a novel, unsupervised learning methodology that uses ordinarily-available tester and simulation data is described that significantly improves resolution with virtually no negative impact on accuracy. Simulation experiments using a variety of fault types (SSL, MSL, bridges, opens and cell-level input-pattern faults) reveal that the number of failed ICs that have perfect resolution can be more than doubled, and overall resolution is improved by 22%. Application to silicon data also demonstrates significant improvement in resolution (38% overall and the number of chips with ideal resolution is nearly tripled) and verification using PFA demonstrates that accuracy is maintained. Osei Poku, Xin Li 0001, R. D. (Shawn) Blanton |
ITC | 3 |
| 2013 | Efficient Spatial Pattern Analysis for Variation Decomposition Via Robust Sparse RegressionabstractIn this paper, we propose a new technique to achieve accurate decomposition of process variation by efficiently performing spatial pattern analysis. We demonstrate that the spatially correlated systematic variation can be accurately represented by the linear combination of a small number of templates. Based on this observation, an efficient sparse regression algorithm is developed to accurately extract the most adequate templates to represent spatially correlated variation. In addition, a robust sparse regression algorithm is proposed to automatically remove measurement outliers. We further develop a fast numerical algorithm that may reduce the computational time by several orders of magnitude over the traditional direct implementation. Our experimental results based on both synthetic and silicon data demonstrate that the proposed sparse regression technique can capture spatially correlated variation patterns with high accuracy and efficiency. Wangyang Zhang, Karthik Balakrishnan, Xin Li 0001, Duane S. Boning, Sharad Saxena, Andrzej J. Strojwas, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | A learning-based autoregressive model for fast transient thermal analysis of chip-multiprocessorsabstractThermal issues have become critical roadblocks for the development of advanced chip-multiprocessors (CMPs). In this paper, we introduce a new angle to view transient thermal analysis - based on predicting thermal profile, instead of calculating it. We develop a systematic framework that can learn different thermal profiles of a CMP by using an autoregressive (AR) model. The proposed AR model can serve as a fast alternative for predicting the transient temperature of a CMP with reasonably good accuracy. Experimental results show that the proposed AR model can achieve approximately 113X speed-up over existing thermal profile estimation methods, while introducing an error of only 0.8°C on average. Da-Cheng Juan, Huapeng Zhou, Diana Marculescu, Xin Li 0001 |
ASP-DAC | 4 |
| 2012 | Statistical design and optimization for adaptive post-silicon tuning of MEMS filtersabstractLarge-scale process variations can significantly limit the practical utility of microelectro-mechanical systems (MEMS) for RF (radio frequency) applications. In this paper we describe a novel technique of adaptive post-silicon tuning to reliably design MEMS filters that are robust to process variations. Our key idea is to implement a number of redundant MEMS resonators to form an array and then optimally select a subset of these resonators to achieve the desired frequency response. Several new CAD algorithms and methodologies are proposed to optimize and configure the design variables of the proposed MEMS resonator array. A MEMS design example demonstrates that the proposed post-silicon tuning is able to reduce the ripple of the channel filter gain by 7x over other traditional approaches. Fa Wang, Gökçe Keskin, Andrew Phelps 0001, Jonathan Rotner, Xin Li 0001, Gary K. Fedder, Tamal Mukherjee, Lawrence T. Pileggi |
DAC | 5 |
| 2012 | An information-theoretic framework for optimal temperature sensor allocation and full-chip thermal monitoringabstractFull-chip thermal monitoring is an important and challenging issue in today's microprocessor design. In this paper, we propose a new information-theoretic framework to quantitatively model the uncertainty of on-chip temperature variation by differential entropy. Based on this framework, an efficient optimization scheme is developed to find the optimal spatial locations for temperature sensors such that the full-chip thermal map can be accurately captured with a minimum number of on-chip sensors. In addition, several efficient numerical algorithms are proposed to minimize the computational cost of the proposed entropy calculation and optimization. As will be demonstrated by our experimental examples, the proposed entropy-based method achieves superior accuracy (1.4x error reduction) for full-chip thermal monitoring over prior art. Huapeng Zhou, Xin Li 0001, Chen-Yong Cher, Eren Kursun, Haifeng Qian, Shi-Chune Yao |
DAC | 2 |
| 2012 | Post-silicon performance modeling and tuning of analog/mixed-signal circuits via Bayesian Model FusionabstractPost-silicon tuning has recently emerged as an important technique to combat large-scale uncertainties (e.g., process variation, device modeling errors, etc) for today's nanoscale circuits. This talk presents a novel Bayesian Model Fusion (BMF) technique for efficient post-silicon performance modeling and tuning of analog and mixed-signal (AMS) circuits. The key idea is to borrow the simulation or measurement data from an early stage (e.g., pre-silicon) to accurately build AMS performance models at a late stage (e.g., post-silicon). The post-silicon models are then used to facilitate efficient tuning of AMS circuits. A circuit example designed in a commercial 32 nm CMOS process is used to demonstrate the efficacy of the proposed post-silicon performance modeling and tuning methodology based on BMF. Xin Li 0001 |
ICCAD | 1 |
| 2012 | Efficient parametric yield estimation of analog/mixed-signal circuits via Bayesian model fusionabstractParametric yield estimation is one of the most critical-yet-challenging tasks for designing and verifying nanoscale analog and mixed-signal circuits. In this paper, we propose a novel Bayesian model fusion (BMF) technique for efficient parametric yield estimation. Our key idea is to borrow the simulation data from an early stage (e.g., schematic-level simulation) to efficiently estimate the performance distributions at a late stage (e.g., post-layout simulation). BMF statistically models the correlation between early-stage and late-stage performance distributions by Bayesian inference. In addition, a convex optimization is formulated to solve the unknown late-stage performance distributions both accurately and robustly. Several circuit examples designed in a commercial 32 nm CMOS process demonstrate that the proposed BMF technique achieves up to 3.75X runtime speedup over the traditional kernel estimation method. Xin Li 0001, Wangyang Zhang, Fa Wang, Shupeng Sun, Chenjie Gu |
ICCAD | 1 |
| 2012 | Efficient SRAM Failure Rate Prediction via Gibbs SamplingabstractStatistical analysis of SRAM has emerged as a challenging issue because the failure rate of SRAM cells is extremely small. In this paper, we develop an efficient importance sampling algorithm to capture the rare failure event of SRAM cells. In particular, we adapt the Gibbs sampling technique from the statistics community to find the optimal probability distribution for importance sampling with a low computational cost (i.e., a small number of transistor-level simulations). The proposed Gibbs sampling method applies an integrated optimization engine to adaptively explore the failure region in a Cartesian or spherical coordinate system by sampling a sequence of 1-D probability distributions. Several implementation issues such as 1-D random sampling and starting point selection are carefully studied to make the Gibbs sampling method efficient and accurate for SRAM failure rate prediction. Our experimental results of a 90 nm SRAM cell demonstrate that the proposed Gibbs sampling method achieves 1.4-4.9× runtime speedup over other state-of-the-art techniques when a high prediction accuracy is required (e.g., the relative error defined by the 99% confidence interval reaches 5%). In addition, we further demonstrate an important example for which the proposed Gibbs sampling algorithm accurately estimates the correct failure probability, while the traditional techniques fail to work. Shupeng Sun, Yamei Feng, Changdao Dong, Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | Efficient SRAM failure rate prediction via Gibbs samplingabstractStatistical analysis of SRAM has emerged as a challenging issue because the failure rate of SRAM cells is extremely small. In this paper, we develop an efficient importance sampling algorithm to capture the rare failure event of SRAM cells. In particular, we adapt the Gibbs sampling technique from the statistics community to find the optimal probability distribution for importance sampling with minimum computational cost (i.e., a small number of transistor-level simulations). The proposed Gibbs sampling method applies an integrated optimization engine to adaptively explore the failure region by sampling a sequence of one-dimensional probability distributions. Several implementation issues such as one-dimensional random sampling and starting point selection are carefully studied to make the Gibbs sampling method efficient and accurate for SRAM failure rate prediction. Our experimental results of a commercial 65nm SRAM cell demonstrate that the proposed Gibbs sampling method achieves 3~10x runtime speed-up over other state-of-the-art techniques without surrendering any accuracy. Changdao Dong, Xin Li 0001 |
DAC | 2 |
| 2011 | Rethinking memory redundancy: optimal bit cell repair for maximum-information storageabstractSRAM design has been a major challenge for nanoscale manufacturing technology. We propose a new bit cell repair scheme for designing maximum-information memory system (MIMS). Unlike the traditional memory repair that attempts to replace all failed bit cells by redundant columns and/or rows, we propose to repair the important bits (e.g., the most significant bit) only so that the information density (i.e., the number of information bits per unit area) is maximized. Towards this goal, an efficient statistical algorithm is derived to efficiently estimate the information density and then optimize the memory system for maximum-information storage. Our experimental results demonstrate that with a traditional 6-T SRAM cell designed in a commercial 45nm CMOS process, the proposed MIMS design can successfully operate at an extremely low power supply voltage (i.e., 0.6 V) and improve the signal-to-noise ratio (SNR) by more than 20 dB compared to the traditional SRAM design. Xin Li 0001 |
DAC | 1 |
| 2011 | Efficient incremental analysis of on-chip power grid via sparse approximationabstractIn this paper, a new sparse approximation technique is proposed for incremental power grid analysis. Our proposed method is motivated by the observation that when a power grid network is locally updated during circuit design, its response changes locally and, hence, the incremental "change" of the power grid voltage is almost zero at many internal nodes, resulting in a unique sparse pattern. An efficient Orthogonal Matching Pursuit (OMP) algorithm is adopted to solve the proposed sparse approximation problem. In addition, several numerical techniques are proposed to improve the numerical stability of the proposed solver, while simultaneously maintaining its high efficiency. Several industrial circuit examples demonstrate that when applied to incremental power grid analysis, our proposed approach achieves up to 130× runtime speed-up over the traditional Algebraic Multi-Grid (AMG) method, without surrendering any accuracy. Xin Li 0001, Ming Yuan Ting |
DAC | 2 |
| 2011 | Formal verification of phase-locked loops using reachability analysis and continuizationabstractWe present an approach for verifying locking of charge-pump phase-locked loops by performing reachability analysis on a behavioral model of the circuit. Bounded uncertain parameters in the behavioral model make it possible to represent all possible behaviors of more detailed models. The dynamics of the behavioral model is hybrid (i.e., discrete and continuous) due to the switching of charge pumps that drive the analog control circuits. A unique feature of phase-locked loops compared to most other hybrid systems is that they require thousands of switchings in the continuous dynamics to converge sufficiently close to a limit cycle. This makes reachability analysis a challenging task since switches in the dynamics are expensive to compute and result in conservative overapproximations. We solve this problem by overapproximating the effects of the switching conditions with uncertain parameters in linear continuous models, a method we call continuization. Using efficient reachability algorithms for discrete-time linear systems, locking is verified over the complete range of possible initial states of a charge-pump PLL designed in 32nm CMOS SOI technology in comparable time required for Monte Carlo simulations of the same behavioral model. Matthias Althoff, Soner Yaldiz, Akshay Rajhans, Xin Li 0001, Bruce H. Krogh, Lawrence T. Pileggi |
ICCAD | 4 |
| 2011 | Toward efficient spatial variation decomposition via sparse regressionabstractIn this paper, we propose a new technique to accurately decompose process variation into two different components: (1) spatially correlated variation, and (2) uncorrelated random variation. Such variation decomposition is important to identify systematic variation patterns at wafer and/or chip level for process modeling, control and diagnosis. We demonstrate that spatially correlated variation carries a unique sparse signature in frequency domain. Based upon this observation, an efficient sparse regression algorithm is applied to accurately separate spatially correlated variation from uncorrelated random variation. An important contribution of this paper is to develop a fast numerical algorithm that reduces the computational time of sparse regression by several orders of magnitude over the traditional implementation. Our experimental results based on silicon measurement data demonstrate that the proposed sparse regression technique can capture spatially correlated variation patterns with high accuracy. The estimation error is reduced by more than 3.5× compared to other traditional methods. Wangyang Zhang, Karthik Balakrishnan, Xin Li 0001, Duane S. Boning, Rob A. Rutenbar |
ICCAD | 3 |
| 2011 | Test cost reduction through performance prediction using virtual probeabstractThe virtual probe (VP) technique, based on recent breakthroughs in compressed sensing, has demonstrated its ability for accurate prediction of spatial variations from a small set of measurement data. In this paper, we explore its application to cost reduction of production testing. For a number of test items, the measurement data from a small subset of chips can be used to accurately predict the performance of other chips on the same wafer without explicit measurement. Depending on their statistical characteristics, test items can be classified into three categories: highly predictable, predictable, and un-predictable. A case study of an industrial RF radio transceiver with more than 50 production test items shows that a good fraction of these test items (39 out of 51 items) are predictable or highly predictable. In this example, the 3σ error of VP prediction is less than 12% for predictable or highly predictable test items. Applying the VP technique can on average replace 59% of test measurement by prediction and, consequently, reduce the overall test time by 57.6%. Hsiu-Ming Chang 0001, Kwang-Ting Cheng, Wangyang Zhang, Xin Li 0001, Kenneth M. Butler |
ITC | 4 |
| 2011 | Virtual Probe: A Statistical Framework for Low-Cost Silicon Characterization of Nanoscale Integrated CircuitsabstractIn this paper, we propose a new technique, referred to as virtual probe (VP), to efficiently measure, characterize, and monitor spatially-correlated inter-die and/or intra-die variations in nanoscale manufacturing process. VP exploits recent breakthroughs in compressed sensing to accurately predict spatial variations from an exceptionally small set of measurement data, thereby reducing the cost of silicon characterization. By exploring the underlying sparse pattern in spatial frequency domain, VP achieves substantially lower sampling frequency than the well-known Nyquist rate. In addition, VP is formulated as a linear programming problem and, therefore, can be solved both robustly and efficiently. Our industrial measurement data demonstrate the superior accuracy of VP over several traditional methods, including 2-D interpolation, Kriging prediction, and k-LSE estimation. Wangyang Zhang, Xin Li 0001, Frank Liu 0001, Emrah Acar, Rob A. Rutenbar, R. D. (Shawn) Blanton |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Toward efficient large-scale performance modeling of integrated circuits via multi-mode/multi-corner sparse regressionabstractIn this paper, we propose a novel multi-mode/multi-corner sparse regression (MSR) algorithm to build large-scale performance models of integrated circuits at multiple working modes and environmental corners. Our goal is to efficiently extract multiple performance models to cover different modes/corners with a small number of simulation samples. To this end, an efficient Bayesian inference with shared prior distribution (i.e., model template) is developed to explore the strong performance correlation among different modes/corners in order to achieve high modeling accuracy with low computational cost. Several industrial circuit examples demonstrate that the proposed MSR achieves up to 185 × speedup over least-squares regression [14] and 6.7 × speedup over least-angle regression [7] without surrendering any accuracy. Wangyang Zhang, Tsung-Hao Chen, Ming Yuan Ting, Xin Li 0001 |
DAC | 4 |
| 2010 | Bayesian virtual probe: minimizing variation characterization cost for nanoscale IC technologies via Bayesian inferenceabstractThe expensive cost of testing and characterizing parametric variations is one of the most critical issues for today's nanoscale manufacturing process. In this paper, we propose a new technique, referred to as Bayesian Virtual Probe (BVP), to efficiently measure, characterize and monitor spatial variations posed by manufacturing uncertainties. In particular, the proposed BVP method borrows the idea of Bayesian inference and information theory from statistics to determine an optimal set of sampling locations where test structures should be deployed and measured to monitor spatial variations with maximum accuracy. Our industrial examples with silicon measurement data demonstrate that the proposed BVP method offers superior accuracy (1.5x error reduction) over the VP approach that was recently developed in [12]. Wangyang Zhang, Xin Li 0001, Rob A. Rutenbar |
DAC | 2 |
| 2010 | Maximum-information storage system: Concept, implementation and applicationabstractThe aggressive technology scaling has made it increasingly difficult to design high-performance, high-density SRAM circuits. In this paper, we propose a new SRAM design methodology that is referred to as maximum-information storage system (MISS). Unlike most traditional SRAM circuits that are designed for maximum cell density, MISS aims to maximize the information density (i.e., the number of information bits per unit area). Towards this goal, an information model is derived to quantitatively measure the information bits stored in a given SRAM system. In addition, a convex optimization framework is developed to optimize SRAM cells to achieve maximum information storage. Our design example in a commercial 65nm CMOS process demonstrates that MISS achieves more than 3.5× area reduction over the traditional SRAM design, while storing the same amount of information. Furthermore, two real-life signal processing examples show that given the same area constraint, MISS can increase signal-to-noise ratio by more than 30 dB compared to the traditional SRAM system. Xin Li 0001 |
ICCAD | 1 |
| 2010 | Multi-Wafer Virtual Probe: Minimum-cost variation characterization by exploring wafer-to-wafer correlationabstractIn this paper, we propose a new technique, referred to as Multi-Wafer Virtual Probe (MVP) to efficiently model wafer-level spatial variations for nanoscale integrated circuits. Towards this goal, a novel Bayesian inference is derived to extract a shared model template to explore the wafer-to-wafer correlation information within the same lot. In addition, a robust regression algorithm is proposed to automatically detect and remove outliers (i.e., abnormal measurement data with large error) so that they do not bias the modeling results. The proposed MVP method is extensively tested for silicon measurement data collected from 200 wafers at an advanced technology node. Our experimental results demonstrate that MVP offers superior accuracy over other traditional approaches such as VP and EM, if a limited number of measurement data are available. Wangyang Zhang, Xin Li 0001, Emrah Acar, Frank Liu 0001, Rob A. Rutenbar |
ICCAD | 2 |
| 2010 | Finding Deterministic Solution From Underdetermined Equation: Large-Scale Performance Variability Modeling of Analog/RF CircuitsabstractThe aggressive scaling of integrated circuit technology results in high-dimensional, strongly-nonlinear performance variability that cannot be efficiently captured by traditional modeling techniques. In this paper, we adapt a novel L0-norm regularization method to address this modeling challenge. Our goal is to solve a large number of (e.g., 104-106) model coefficients from a small set of (e.g., 102-103) sampling points without over-fitting. This is facilitated by exploiting the underlying sparsity of model coefficients. Namely, although numerous basis functions are needed to span the high-dimensional, strongly-nonlinear variation space, only a few of them play an important role for a given performance of interest. An efficient orthogonal matching pursuit (OMP) algorithm is applied to automatically select these important basis functions based on a limited number of simulation samples. Several circuit examples designed in a commercial 65 nm process demonstrate that OMP achieves up to 25× speedup compared to the traditional least-squares fitting method. Xin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Finding deterministic solution from underdetermined equation: large-scale performance modeling by least angle regressionabstractThe aggressive scaling of IC technology results in high-dimensional, strongly-nonlinear performance variability that cannot be efficiently captured by traditional modeling techniques. In this paper, we adapt a novel L1-norm regularization method to address this modeling challenge. Our goal is to solve a large number of (e.g., 104~106) model coefficients from a small set of (e.g., 102~103) sampling points without over-fitting. This is facilitated by exploiting the underlying sparsity of model coefficients. Namely, although numerous basis functions are needed to span the high-dimensional, strongly-nonlinear variation space, only a few of them play an important role for a given performance of interest. An efficient algorithm of least angle regression (LAR) is applied to automatically select these important basis functions based on a limited number of simulation samples. Several circuit examples designed in a commercial 65nm process demonstrate that LAR achieves up to 25x speedup compared with the traditional least-squares fitting. Xin Li 0001 |
DAC | 1 |
| 2009 | SRAM parametric failure analysisabstractWith aggressive technology scaling, SRAM design has been seriously challenged by the difficulties in analyzing rare failure events. In this paper we propose to create statistical performance models with accuracy sufficient to facilitate probability extraction for SRAM parametric failures. A piecewise modeling technique is first proposed to capture the performance metrics over the large variation space. A controlled sampling scheme and a nested Monte Carlo analysis method are then applied for the failure probability extraction at cell-level and array-level respectively. Our 65nm SRAM example demonstrates that by combining the piecewise model and the fast probability extraction methods, we have significantly accelerated the SRAM failure analysis. Jian Wang 0100, Soner Yaldiz, Xin Li 0001, Lawrence T. Pileggi |
DAC | 3 |
| 2009 | Efficient design-specific worst-case corner extraction for integrated circuitsabstractWhile statistical analysis has been considered as an important tool for nanoscale integrated circuit design, many IC designers would like to know the design-specific worst-case corners for circuit debugging and failure diagnosis. In this paper, we propose a novel algorithm to efficiently extract the worst-case corners for nanoscale ICs. Our proposed approach mathematically formulates a quadratically constrained quadratic programming (QCQP) problem for corner extraction. Next, it applies the Lagrange duality theory to convert the non-convex QCQP problem to a convex semi-definite programming (SDP) problem that is easier to solve. Our circuit example designed in a commercial CMOS process demonstrates that the proposed SDP formulation can find the worst-case corners both efficiently and robustly, while the traditional QCQP fails to achieve global convergence. Tsung-Hao Chen, Ming Yuan Ting, Xin Li 0001 |
DAC | 4 |
| 2009 | Virtual probe: A statistically optimal framework for minimum-cost silicon characterization of nanoscale integrated circuitsabstractIn this paper, we propose a new technique, referred to as virtual probe (VP), to efficiently measure, characterize and monitor both inter-die and spatially-correlated intra-die variations in nanoscale manufacturing process. VP exploits recent breakthroughs in compressed sensing [15]-[17] to accurately predict spatial variations from an exceptionally small set of measurement data, thereby reducing the cost of silicon characterization. By exploring the underlying sparse structure in (spatial) frequency domain, VP achieves substantially lower sampling frequency than the well-known (spatial) Nyquist rate. In addition, VP is formulated as a linear programming problem and, therefore, can be solved both robustly and efficiently. Our industrial measurement data demonstrate that by testing the delay of just 50 chips on a wafer, VP accurately predicts the delay of the other 219 chips on the same wafer. In this example, VP reduces the estimation error by up to 10× compared to other traditional methods. Categories and Subject Descriptors B.7.2 [Integrated Circuits]: Design Aids — Verification General Terms Algorithms Xin Li 0001, Rob A. Rutenbar, R. D. (Shawn) Blanton |
ICCAD | 1 |
| 2009 | Regular Analog/RF Integrated Circuits Design Using Optimization With Recourse Including Ellipsoidal UncertaintyabstractLong design cycles due to the inability to predict silicon realities are a well-known problem that plagues analog/RF integrated circuit product development. As this problem worsens for nanoscale IC technologies, the high cost of design and multiple manufacturing spins causes fewer products to have the volume required to support full-custom implementation. Design reuse and analog synthesis make analog/RF design more affordable; however, the increasing process variability and lack of modeling accuracy remain extremely challenging for nanoscale analog/RF design. We propose a regular analog/RF IC using metal-mask configurability design methodology Optimization with Recourse of Analog Circuits including Layout Extraction (ORACLE), which is a combination of reuse and shared-use by formulating the synthesis problem as an optimization with recourse problem. Using a two-stage geometric programming with recourse approach, ORACLE solves for both the globally optimal shared and application-specific variables. Furthermore, robust optimization is proposed to treat the design with variability problem, further enhancing the ORACLE methodology by providing yield bound for each configuration of regular designs. The statistical variations of the process parameters are captured by a confidence ellipsoid. We demonstrate ORACLE for regular Low Noise Amplifier designs using metal-mask configurability, where a range of applications share common underlying structure and application-specific customization is performed using the metal-mask layers. Two RF oscillator design examples are shown to achieve robust designs with guaranteed yield bound. Yang Xu 0017, Kan-Lin Hsiung, Xin Li 0001, Lawrence T. Pileggi, Stephen P. Boyd |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Statistical regression for efficient high-dimensional modeling of analog and mixed-signal performance variationsabstractThe continuous technology scaling brings about high-dimensional performance variations that cannot be easily captured by the traditional response surface modeling. In this paper we propose a new statistical regression (STAR) technique that applies a novel strategy to address this high dimensionality issue. Unlike most traditional response surface modeling techniques that solve model coefficients from over-determined linear equations, STAR determines all unknown coefficients by moment matching. As such, a large number of (e.g., 103~105) model coefficients can be extracted from a small number of (e.g., 102~103) sampling points without over-fitting. In addition, a novel recursive estimator is proposed to accurately and efficiently predict the moment values. The proposed recursive estimator is facilitated by exploiting the interaction between different moment estimators and formulating the moment estimation problem into a special form that can be iteratively solved. Several circuit examples designed in commercial CMOS processes demonstrate that STAR achieves more than 20x runtime speedup compared with the traditional response surface modeling. Xin Li 0001, Hongzhou Liu |
DAC | 1 |
| 2008 | Digital Circuit Design Challenges and Opportunities in the Era of Nanoscale CMOSabstractWell-designed circuits are one key ldquoinsulatingrdquo layer between the increasingly unruly behavior of scaled complementary metal-oxide-semiconductor devices and the systems we seek to construct from them. As we move forward into the nanoscale regime, circuit design is burdened to ldquohiderdquo more of the problems intrinsic to deeply scaled devices. How this is being accomplished is the subject of this paper. We discuss new techniques for logic circuits and interconnect, for memory, and for clock and power distribution. We survey work to build accurate simulation models for nanoscale devices. We discuss the unique problems posed by nanoscale lithography and the role of geometrically regular circuits as one promising solution. Finally, we look at recent computer-aided design efforts in modeling, analysis, and optimization for nanoscale designs with ever increasing amounts of statistical variation. Benton H. Calhoun, Yu Cao 0001, Xin Li 0001, Ken Mai, Lawrence T. Pileggi, Rob A. Rutenbar, Kenneth L. Shepard |
Proc. IEEE | 3 |
| 2008 | Defining Statistical Timing Sensitivity for Logic Circuits With Large-Scale Process and Environmental VariationsabstractThe large-scale process and environmental variations for today's nanoscale ICs require statistical approaches for timing analysis and optimization. In this paper, we demonstrate why the traditional concept of slack and critical path becomes ineffective under large-scale variations and propose a novel sensitivity framework to assess the ldquocriticalityrdquo of every path, arc, and node in a statistical timing graph. We theoretically prove that the path sensitivity is exactly equal to the probability that a path is critical and that the arc (or node) sensitivity is exactly equal to the probability that an arc (or a node) sits on the critical path. An efficient algorithm with incremental analysis capability is developed for fast sensitivity computation that has linear runtime complexity in circuit size. The efficacy of the proposed sensitivity analysis is demonstrated on both standard benchmark circuits and large industrial examples. Xin Li 0001, Jiayong Le, Mustafa Celik, Lawrence T. Pileggi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Quadratic Statistical MAX Approximation for Parametric Yield Estimation of Analog/RF Integrated CircuitsabstractIn this paper, we propose an efficient numerical algorithm for estimating the parametric yield of analog/RF circuits, considering large-scale process variations. Unlike many traditional approaches that assume normal performance distributions, the proposed approach is particularly developed to handle multiple correlated nonnormal performance distributions, thereby providing better accuracy than the traditional techniques. Starting from a set of quadratic performance models, the proposed parametric yield estimation conceptually maps multiple correlated performance constraints to a single auxiliary constraint by using a MAX operator. As such, the parametric yield is uniquely determined by the probability distribution of the auxiliary constraint and, therefore, can easily be computed. In addition, two novel numerical algorithms are derived from moment matching and statistical Taylor expansion, respectively, to facilitate efficient quadratic statistical MAX approximation. We prove that these two algorithms are mathematically equivalent if the performance distributions are normal. Our numerical examples demonstrate that the proposed algorithm provides an error reduction of 6.5 times compared to a normal-distribution-based method while achieving a runtime speedup of 10-20 times over the Monte Carlo analysis with 103samples. Xin Li 0001, Yaping Zhan, Lawrence T. Pileggi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Efficient Parametric Yield Extraction for Multiple Correlated Non-Normal Performance Distributions of Analog/RF CircuitsabstractIn this paper we propose an efficient numerical algorithm to estimate the parametric yield of analog/RF circuits with consideration of large-scale process variations. Unlike many traditional approaches that assume Normal performance distributions, the proposed approach is especially developed to handle multiple correlated non-Normal performance distributions, thereby providing better accuracy than other traditional techniques. Starting from a set of quadratic performance models, the proposed parametric yield extraction conceptually maps multiple correlated performance constraints to a single auxiliary constraint using a MAX(·) operator. As such, the parametric yield is uniquely determined by the probability distribution of the auxiliary constraint and, therefore, can be easily computed. In addition, a novel second-order statistical Taylor expansion is proposed for an analytical MAX(·) approximation, facilitating fast yield estimation. Our numerical examples in a commercial BiCMOS process demonstrate that the proposed algorithm provides 2--3x error reduction compared with a Normal-distribution-based method, while achieving orders of magnitude more efficiency than the Monte Carlo analysis with 104 samples. Xin Li 0001, Lawrence T. Pileggi |
DAC | 1 |
| 2007 | Parameterized Macromodeling for Analog System-Level Design ExplorationabstractIn this paper we propose a novel parameterized macromodeling technique for analog circuits. Unlike traditional macromodels that are only extracted for a small variation space, our proposed approach captures a significantly larger analog design space to facilitate system-level design exploration. Combining a novel piece-wise approximation algorithm and a new multi-point model-order-reduction approach, the proposed method generates compact macromodels covering the entire feasible design space. Our experiments demonstrate that using such models can achieve more than 60 x speed-up while incurring less than 4% overall error when varying design parameters by an order of magnitude. Jian Wang 0100, Xin Li 0001, Lawrence T. Pileggi |
DAC | 2 |
| 2007 | Adaptive post-silicon tuning for analog circuits: concept, analysis and optimizationabstractThe well-known Pelgrom model [14] has demonstrated that the variation between two devices on the same die due to random mismatch is inversely proportional to the square root of the device area: σ ∼ 1/sqrt(Area). Based on the Pelgrom model, analog devices are sized to be large enough to average out random variations. Importantly, with CMOS scaling, variations due to random doping fluctuations are making it exceedingly difficult to control device mismatches by sizing alone; namely, the devices have to be made so large that the benefits of CMOS scaling are not realized for analog and RF circuits. In this paper we propose a novel post-silicon tuning methodology to reduce random mismatches for analog circuits in sub-90nm CMOS. A novel dynamic programming algorithm is incorporated into a fast Monte Carlo simulation flow for statistical analysis and optimization of the proposed tunable analog circuits. We apply the proposed postsilicon tuning methodology to several commonly-used analog circuit blocks. We demonstrate that with the post-silicon tuning, device mismatch exponentially decreases as area increases: σ ∼ exp(-α·Area). Xin Li 0001, YuTsun Chien, Lawrence T. Pileggi |
ICCAD | 1 |
| 2007 | Robust Analog/RF Circuit Design With Projection-Based Performance ModelingabstractIn this paper, a robust analog design (ROAD) tool for post-tuning (i.e., locally optimizing) analog/RF circuits is proposed. Starting from an initial design derived from hand analysis or analog circuit optimization based on simplified models, ROAD extracts accurate performance models via transistor-level simulation and iteratively improves the circuit performance by a sequence of geometric programming steps. Importantly, ROAD sets up all design constraints to include large-scale process and environmental variations, thereby facilitating the tradeoff between yield and performance. A crucial component of ROAD is a novel projection-based scheme for quadratic (both polynomial and posynomial) performance modeling, which allows our approach to scale well to large problem sizes. A key feature of this projection-based scheme is a new implicit power iteration algorithm to find the optimal projection space and extract the unknown model coefficients with robust convergence. The efficacy of ROAD is demonstrated on several circuit examples Xin Li 0001, Padmini Gopalakrishnan, Yang Xu 0017, Lawrence T. Pileggi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Asymptotic Probability Extraction for Nonnormal Performance DistributionsabstractWhile process variations are becoming more significant with each new IC technology generation, they are often modeled via linear regression models so that the resulting performance variations can be captured via normal distributions. Nonlinear response surface models (e.g., quadratic polynomials) can be utilized to capture larger scale process variations; however, such models result in nonnormal distributions for circuit performance. These performance distributions are difficult to capture efficiently since the distribution model is unknown. In this paper, an asymptotic-probability-extraction (APEX) method for estimating the unknown random distribution when using a nonlinear response surface modeling is proposed. The APEX begins by efficiently computing the high-order moments of the unknown distribution and then applies moment matching to approximate the characteristic function of the random distribution by an efficient rational function. It is proven that such a moment-matching approach is asymptotically convergent when applied to quadratic response surface models. In addition, a number of novel algorithms and methods, including binomial moment evaluation, PDF/CDF shifting, nonlinear companding and reverse evaluation, are proposed to improve the computation efficiency and/or approximation accuracy. Several circuit examples from both digital and analog applications demonstrate that APEX can provide better accuracy than a Monte Carlo simulation with 104samples and achieve up to 10times more efficiency. The error, incurred by the popular normal modeling assumption for several circuit examples designed in standard IC technologies, is also shown Xin Li 0001, Jiayong Le, Padmini Gopalakrishnan, Lawrence T. Pileggi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Architecture-aware FPGA placement using metric embeddingabstractSince performance on FPGAs is dominated by the routing architecture rather than wirelength, we propose a new ar-chitecture-aware approach to initial FPGA placement that models the relationship between performance and the routing grid, using concepts from graph embedding and metric geometry. Our approach, CAPRI, can be viewed as an embedding of a graph representing the netlist into a metric space that is representative of the FPGA. First, we develop an analytic metric of distance that models delays along the FPGA routing grid. We then embed a netlist into the defined metric space using matrix projections and online bipartite matching. Experimental comparisons with the popular FPGA tool, VPR, show that with CAPRI's initial solution, the resulting placements show median improvements of 10% in critical path delays for the larger MCNC benchmarks. Total placement runtime is also improved by 2x on average. Padmini Gopalakrishnan, Xin Li 0001, Lawrence T. Pileggi |
DAC | 2 |
| 2006 | Projection-based statistical analysis of full-chip leakage power with non-log-normal distributionsabstractIn this paper we propose a novel projection-based algorithm to estimate the full-chip leakage power with consideration of both inter-die and intra-die process variations. Unlike many traditional approaches that rely on log-Normal approximations, the proposed algorithm applies a novel projection method to extract a low-rank quadratic model of the logarithm of the full-chip leakage current and, therefore, is not limited to log-Normal distributions. By exploring the underlying sparse structure of the problem, an efficient algorithm is developed to extract the non-log-Normal leakage distribution with linear computational complexity in circuit size. In addition, an incremental analysis algorithm is proposed to quickly update the leakage distribution after changes to a circuit are made. Our numerical examples in a commercial 90nm CMOS process demonstrate that the proposed algorithm provides 4x error reduction compared with the previously proposed log-Normal approximations, while achieving orders of magnitude more efficiency than a Monte Carlo analysis with 10 4 samples. Xin Li 0001, Jiayong Le, Lawrence T. Pileggi |
DAC | 1 |
| 2005 | OPERA: optimization with ellipsoidal uncertainty for robust analog IC designabstractAs the design-manufacturing interface becomes increasingly complicated with IC technology scaling, the corresponding process variability poses great challenges for nanoscale analog/RF design. Design optimization based on the enumeration of process corners has been widely used , but can suffer from inefficiency and overdesign. In this paper we propose to formulate the analog and RF design with variability problem as a special type of robust optimization problem, namely robust geometric programming. The statistical variations in both the process parameters and design variables are captured by a pre-specified confidence ellipsoid. Using such optimization with ellipsoidal uncertainy approach, robust design can be obtained with guaranteed yield bound and lower design cost, and most importantly, the problem size grows linearly with number of uncertain parameters. Numerical examples demonstrate the efficiency and reveal the trade-off between the design cost versus the yield requirement. We will also demonstrate significant improvement in the design cost using this approach compared with corner-enumeration optimization. Yang Xu 0017, Kan-Lin Hsiung, Xin Li 0001, Ivan Nausieda, Stephen P. Boyd, Lawrence T. Pileggi |
DAC | 3 |
| 2005 | Correlation-aware statistical timing analysis with non-gaussian delay distributionsabstractProcess variations have a growing impact on circuit performance for today's integrated circuit (IC) technologies. The Non-Gaussian delay distributions as well as the correlations among delays make statistical timing analysis more challenging than ever. In this paper, we present an efficient block-based statistical timing analysis approach with linear complexity with respect to the circuit size, which can accurately predict Non-Gaussian delay distributions from realistic nonlinear gate and interconnect delay models. This approach accounts for all correlations, from manufacturing process dependence, to re-convergent circuit paths to produce more accurate statistical timing predictions. With this approach, circuit designers can have increased confidence in the variation estimates, at a low additional computation cost. Yaping Zhan, Andrzej J. Strojwas, Xin Li 0001, Lawrence T. Pileggi, David Newmark, Mahesh Sharma |
DAC | 3 |
| 2005 | Modeling Interconnect Variability Using Efficient Parametric Model Order ReductionabstractAssessing IC manufacturing process fluctuations and their impacts on IC interconnect performance has become unavoidable for modern DSM designs. However, the construction of parametric interconnect models is often hampered by the rapid increase in computational cost and model complexity. In this paper we present an efficient yet accurate parametric model order reduction algorithm for addressing the variability of IC interconnect performance. The efficiency of the approach lies in a novel combination of low-rank matrix approximation and multi-parameter moment matching. The complexity of the proposed parametric model order reduction is as low as that of a standard Krylov subspace method when applied to a nominal system. Under the projection-based framework, our algorithm also preserves the passivity of the resulting parametric models. Peng Li 0001, Frank Liu 0001, Xin Li 0001, Lawrence T. Pileggi, Sani R. Nassif |
DATE | 3 |
| 2005 | Defining statistical sensitivity for timing optimization of logic circuits with large-scale process and environmental variationsabstractThe large-scale process and environmental variations for today's nanoscale ICs are requiring statistical approaches for timing analysis and optimization. Significant research has been recently focused on developing new statistical timing analysis algorithms, but often without consideration for how one should interpret the statistical timing results for optimization. In this paper (Li et al., 2005) we demonstrate why the traditional concepts of slack and critical path become ineffective under large-scale variations, and we propose a novel sensitivity-based metric to assess the "criticality" of each path and/or arc in the statistical timing graph. We define the statistical sensitivities for both paths and arcs, and theoretically prove that our path sensitivity is equivalent to the probability that a path is critical, and our arc sensitivity is equivalent to the probability that an arc sits on the critical path. An efficient algorithm with incremental analysis capability is described for fast sensitivity computation that has a linear runtime complexity in circuit size. The efficacy of the proposed sensitivity analysis is demonstrated on both standard benchmark circuits and large industry examples. Xin Li 0001, Jiayong Le, Mustafa Celik, Lawrence T. Pileggi |
ICCAD | 1 |
| 2005 | Parameterized interconnect order reduction with explicit-and-implicit multi-parameter moment matching for inter/intra-die variationsabstractIn this paper we propose a novel parameterized interconnect order reduction algorithm, CORE, to efficiently capture both inter-die and intra-die variations. CORE applies a two-step explicit-and-implicit scheme for multiparameter moment matching. As such, CORE can match significantly more moments than other traditional techniques using the same model size. In addition, a recursive Arnoldi algorithm is proposed to quickly construct the Krylov subspace that is required for parameterized order reduction. Applying the recursive Arnoldi algorithm significantly reduces the computation cost for model generation. Several RC and RLC interconnect examples demonstrate that CORE can provide up to 10/spl times/ better modeling accuracy than other traditional techniques, while achieving smaller model complexity (i.e. size). It follows that these interconnect models generated by CORE can provide more accurate simulation result with cheaper simulation cost, when they are utilized for gate-interconnect co-simulation. Xin Li 0001, Peng Li 0001, Lawrence T. Pileggi |
ICCAD | 1 |
| 2005 | Projection-based performance modeling for inter/intra-die variationsabstractLarge-scale process fluctuations in nano-scale IC technologies suggest applying high-order (e.g., quadratic) response surface models to capture the circuit performance variations. Fitting such models requires significantly more simulation samples and solving much larger linear equations. In this paper, we propose a novel projection-based extraction approach, PROBE, to efficiently create quadratic response surface models and capture both inter-die and intra-die variations with affordable computation cost. PROBE applies a novel projection scheme to reduce the response surface modeling cost (i.e., both the required number of samples and the linear equation size) and make the modeling problem tractable even for large problem sizes. In addition, a new implicit power iteration algorithm is developed to find the optimal projection space and solve for the unknown model coefficients. Several circuit examples from both digital and analog circuit modeling applications demonstrate that PROBE can generate accurate response surface models while achieving up to 12/spl times/ speedup compared with the traditional methods. Xin Li 0001, Jiayong Le, Lawrence T. Pileggi, Andrzej J. Strojwas |
ICCAD | 1 |
| 2005 | Performance-centering optimization for system-level analog design explorationabstractIn this paper we propose a novel analog design optimization methodology to address two key aspects of top-down system-level design: (1) how to optimally compare and select analog system architectures in the early phases of design; and (2) how to hierarchically propagate performance specifications from system level to circuit level to enable independent circuit block design. Importantly, due to the inaccuracy of early-stage system-level models, and the increasing magnitude of process and environmental variations, the system-level exploration must leave sufficient design margin to ensure a successful late-stage implementation. Therefore, instead of minimizing a design objective function, and thereby converging on a constraint boundary, we apply a novel performance centering optimization. Our proposed methodology centers the analog design in the performance space, and maximizes the distance to all constraint boundaries. We demonstrate that this early-stage design margin, which is measured by the volume of the inscribed ellipsoid lying inside the performance constraints, provides an excellent quality measure for comparing different system architectures. The efficacy of our performance centering approach is shown for analog design examples, including a complete clock data recovery system design and implementation. Xin Li 0001, Jian Wang 0100, Lawrence T. Pileggi, Tun-Shih Chen, Wanju Chiang |
ICCAD | 1 |
| 2004 | STAC: statistical timing analysis with correlationabstractCurrent technology trends have led to the growing impact of both inter-die and intra-die process variations on circuit performance. While it is imperative to model parameter variations for sub-100nm technologies to produce an upper bound prediction on timing, it is equally important to consider the correlation of these variations for the bound to be useful. In this paper we present an efficient block-based statistical static timing analysis algorithm that can account for correlations from process parameters and re-converging paths. The algorithm can also accommodate dominant interconnect coupling effects to provide an accurate compilation of statistical timing information. The generality and efficiency for the proposed algorithm is obtained from a novel simplification technique that is derived from the statistical independence theories and principal component analysis (PCA) methods. The technique significantly reduces the cost for mean, variance and covariance computation of a set of correlated random variables. Jiayong Le, Xin Li 0001, Lawrence T. Pileggi |
DAC | 2 |
| 2004 | A frequency relaxation approach for analog/RF system-level simulationabstractThe increasing complexity of today's mixed-signal integrated circuits necessitates both top-down and bottom-up system-level verification. Time-domain state-space modeling and simulation approaches have been successfully applied for such purposes (e.g. Simulink); however, analog circuits are often best analyzed in the frequency domain. Circuit-level analyses, such as harmonic balance, have been successfully extended to the frequency domain [2], but these algorithms are impractical for simulating large systems with wide-band input and noise signals. In this paper we proposed a frequency-domain approach for analog/RF system-level simulation that is capable of capturing various second order effects (e.g. nonlinearity, noise, etc.) for both time-invariant and time-varying systems with wide-band inputs. The simulator directly evaluates the frequency domain response at each node via a relaxation scheme that is proven to be convergent under typical circuit conditions. Our experimental results demonstrate the accuracy and efficiency of the proposed simulator under various wide-band input and noise excitations. Xin Li 0001, Yang Xu 0017, Peng Li 0001, Padmini Gopalakrishnan, Lawrence T. Pileggi |
DAC | 1 |
| 2004 | Robust analog/RF circuit design with projection-based posynomial modelingabstractWe propose a robust analog design tool (ROAD) for post-tuning analog/RF circuits. Starting from an initial design derived from hand analysis or analog circuit synthesis based on simplified models, ROAD extracts accurate posynomial performance models via transistor-level simulation and optimizes the circuit by geometric programming. Importantly, ROAD sets up all design constraints to include large-scale process variations to facilitate the tradeoff between yield and performance. A novel convex formulation of the robust design problem is utilized to improve the optimization efficiency and to produce a solution that is superior to other local tuning methods. In addition, a novel projection-based approach for posynomial fitting is used to facilitate scaling to large problem sizes. A new implicit power iteration algorithm is proposed to find the optimal projection space and extract the posynomial coefficients with robust convergence. The efficacy of ROAD is demonstrated on several circuit examples. Xin Li 0001, Padmini Gopalakrishnan, Yang Xu 0017, Lawrence T. Pileggi |
ICCAD | 1 |
| 2004 | Asymptotic probability extraction for non-normal distributions of circuit performanceabstractWhile process variations are becoming more significant with each new IC technology generation, they are often modeled via linear regression models so that the resulting performance variations can be captured via normal distributions. Nonlinear (e.g. quadratic) response surface models can be utilized to capture larger scale process variations; however, such models result in non-normal distributions for circuit performance which are difficult to capture since the distribution model is unknown. In this paper we propose an asymptotic probability extraction method, APEX, for estimating the unknown random distribution when using nonlinear response surface modeling. APEX first uses a binomial moment evaluation to efficiently compute the high order moments of the unknown distribution, and then applies moment matching to approximate the characteristic function of the random circuit performance by an efficient rational function. A simple statistical timing example and an analog circuit example demonstrate that APEX can provide better accuracy than Monte Carlo simulation with 10 samples and achieve orders of magnitude more efficiency. We also show the error incurred by the popular normal modeling assumption using standard IC technologies. Xin Li 0001, Jiayong Le, Padmini Gopalakrishnan, Lawrence T. Pileggi |
ICCAD | 1 |
| 2003 | A frequency separation macromodel for system-level simulation of RF circuitsabstractIn this paper we propose a frequency-separation methodology to generate system-level macromodels for analog and RF circuits. The proposed macromodels are similar in form to those based on Volterra kernel calculations, but are much simpler in terms of characterization and overall model complexity, and can be derived from existing device models. This simplicity is realized by applying some basic assumptions on the form of the input excitations, and via separation of the nonlinearities from the dynamic behavior. In addition, by further separating the ideal model functionality, this macromodel is applicable to strongly nonlinear components such as mixers. While time-varying Volterra series models have been proposed for mixers with a fixed local oscillation (LO) signal, the proposed frequency separation model is completely general and can capture the variations of the LO input during a system-level simulation. The proposed macromodels are demonstrated in a system-level simulation tool based on Simulink for efficient evaluation of the entire RF system and associated components. A GSM receiver system in 0.25μm CMOS process is used to demonstrate the efficacy of these macromodels in our system-level simulation environment. Xin Li 0001, Peng Li 0001, Yang Xu 0017, Robert Dimaggio, Lawrence T. Pileggi |
ASP-DAC | 1 |
| 2003 | Analog and RF circuit macromodels for system-level analysisabstractDesign and validation of mixed-signal integrated systems require system-level model abstractions. Generalized Volterra series based models have been successfully applied for analog and RF component macromodels, but their complexity can sometimes limit their utility for time-varying systems and large circuits with complex device models or numerous parasitics. In this paper we propose simple and efficient analog and RF circuit macromodels that provide accurate model abstractions for large, complex time-varying circuits over frequency bands of interest. By starting with the system-level block diagram model structures and focusing on the narrow RF bands, the proposed macromodels can efficiently capture the nonlinear behavior as well as the impact of RLC coupling parasitics via compact reduced-order model forms. While the macromodel can trade accuracy for simplicity in terms of the number of frequency expansion points, we find that expansion about one frequency point provides the accuracy required for system-level analysis of most RF and narrow-band analog components. The macromodel form corresponds to block diagram structures that are easily incorporated into our system-level simulation tool based on Simulink. Xin Li 0001, Peng Li 0001, Yang Xu 0017, Lawrence T. Pileggi |
DAC | 1 |
| 2003 | Noise Macromodel for Radio Frequency Integrated CircuitsabstractNoise performance is a critical analog and RF circuit design constraint, and can impact the selection of the IC system-level architecture. It is therefore imperative that some model of the noise is represented at the highest levels of abstraction during the design process. In this paper we propose a noise macromodel for analog circuits and demonstrate it by way of implementation in a system level simulator based on MATLAB. We also explain our process of macromodel extraction via reformulation of frequency-domain noise analysis results, and the corresponding steps of model order reduction. The results demonstrate the efficacy of this macromodel for frequency domain system level simulation. Yang Xu 0017, Xin Li 0001, Peng Li 0001, Lawrence T. Pileggi |
DATE | 2 |
| 2003 | A Hybrid Approach to Nonlinear Macromodel Generation for Time-Varying Analog Circuits
Peng Li 0001, Xin Li 0001, Yang Xu 0017, Lawrence T. Pileggi |
ICCAD | 2 |
| 2001 | The autocorrelation matching method for distributed MIMO communications over unknown FIR channelsabstractThe autocorrelation matching method is a blind signal separation and channel equalization technique for distributed MIMO communication systems over unknown FIR channels using only second order statistics. This method is based on a theoretical discovery, ie, under the condition that the autocorrelation functions of the (multiple) inputs are linearly shift-independent, an input is recovered, up to a unitary factor and a delay, by an output of an MIMO-FIR equalizer if and only if the autocorrelation function of the output matches that of the input. An optimal zero-forcing equalizer is developed to maximize the SNR for the outputs, ie, the recovered inputs. Some preliminary simulation results show that the BER in the recovered inputs is about 3/spl times/10/sup -5/ at the SNR=15 dB. This method has the potential to be applied to cellular wireless communications for the purpose of boosting spectrum efficiency or suppressing co-channel interference. Ruey-Wen Liu, Xieting Ling, Xin Li 0001 |
ICASSP | 4 |
| 2001 | Behavioral Modeling of Analog Circuits by Wavelet Collocation MethodabstractIn this paper, we develop a wavelet collocation method with nonlinear companding for behavioral modeling of analog circuits. To construct the behavioral models, the circuit is first partitioned into building blocks and the input-output function of each block is then approximated by wavelets. As the blocks are mathematically represented by sets of simple wavelet basis functions, the computation cost for the behavioral simulation is significantly reduced. The proposed method presents several merits compared with those conventional techniques. First, the algorithm for expanding input-output functions by wavelets is a general-purpose approach, which can be applied in automatically modeling of different analog circuit blocks with different structures. Second, both the small signal effect and the large signal effect are modeled in a unified formulation, which eases the process of modeling and simulation. Third, a nonlinear companding method is developed to control the modeling error distribution, To demonstrate the promising features of the proposed method, a 4th order switched-current filter is employed to build the behavioral model. Xin Li 0001, Xuan Zeng 0001, Dian Zhou, Xieting Ling |
ICCAD | 1 |