EDBT 2026 Demo / reviewers in the wild / expert
Xinli Gu
dblp:55/816
· DBLP profile ↗
66ranked-venue papers
12as first author
4since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 12 first-author · 4 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Knowledge Transfer in Board-Level Functional Fault Diagnosis Enabled by Domain AdaptationabstractHigh integration densities and design complexity make board-level functional fault diagnosis extremely difficult. Machine-learning techniques can identify functional faults with high accuracy, but they require a large volume of data to achieve high-prediction accuracy. This drawback limits the effectiveness of traditional machine-learning algorithms for training a model in the early stage of manufacturing, when only a limited amount of fail data and repair records are available. We propose a board-level diagnosis workflow that utilizes domain adaptation (DA) to transfer the knowledge learned from mature boards to a new board in the ramp-up phase. First, based on the requirement of fault diagnosis, we select an appropriate domain-adaptation method to reduce differences between mature boards and the new board. Second, these DA methods utilize information from both the mature and the new boards with carefully designed domain-alignment rules and train a functional fault diagnosis classifier. Experimental results using three complex boards in volume production and one new board in the ramp-up phase show that, with the help of DA and the proposed workflow, the diagnosis accuracy is improved. Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Unsupervised Two-Stage Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems have placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligence and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. We propose a two-stage unsupervised root-cause-analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to cluster the data in a coarse-grained manner. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. The proposed method can accommodate both numerical and categorical test items. A combination of the L-method, cross validation, and Silhouette score enables us to automatically determine all hyperparameters. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause-analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Board-Level Functional Fault Identification Using Streaming DataabstractHigh integration densities and design complexity of printed-circuit boards make board-level functional fault identification extremely difficult. Machine learning provides an opportunity to identify functional faults with high accuracy and thereby reduce repair cost. However, the large volume of manufacturing data comes in a streaming format and exhibits time-dependent concept drift in a production environment. These drawbacks limit the effectiveness of traditional machine-learning algorithms. We propose a diagnosis workflow that utilizes online learning to train classifiers incrementally with a small chunk of data at each step. These online-learning algorithms adapt to concept drift quickly with carefully designed update rules. A hybrid algorithm is also proposed to handle the scenario that data for varying numbers of boards are collected at different times. This hybrid algorithm concurrently implements two basic models. For each data chunk, this algorithm chooses the better model with high probability. The experimental results using two boards in high-volume production show that, with the help of online learning and the proposed hybrid algorithm, the F1-score for diagnosis based on binary classifiers can be improved from 57.3% to 81.0%. The top-3 accuracy for diagnosis based on multiclass classifiers can be improved from 78.3% to 91.4%. Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Black-Box Test-Cost Reduction Based on Bayesian Network ModelsabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. In this article, we propose a novel black-box test selection method based on Bayesian networks (BNs), which extract the strong relationship among tests. First, the problem of reducing the black-box test cost is formulated as a constrained optimization problem. Next, multiple structure learning and transfer learning algorithms are implemented to construct BN models. Based on these BN models, we propose an iterative test selection method with a new metric, Bayesian index, for test-cost reduction. In addition, averaging strategies are applied to enhance the reduction performance. Finally, a robust model selection framework is proposed to select the optimal BN model for test-cost reduction. Two case studies with production test data demonstrate that when no prior information is provided, our proposed approach effectively reduces the test cost by up to 14.7%, compared to the state-of-the-art greedy algorithm. Moreover, our proposed approach further reduces the test cost by up to 7.1% when prior information is provided from similar products. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Unsupervised Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems has placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligent and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. In this paper, we propose a two-stage unsupervised root-cause analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to roughly cluster the data. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. In additional, L-method and cross validation are applied to automatically determine the hyper-parameters of our algorithm. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 5 |
| 2020 | Hierarchical Symbol-Based Health-Status Analysis Using Time-Series Data in a Core Router SystemabstractTo ensure high reliability and rapid error recovery in commercial core router systems, a health-status analyzer is essential to monitor the different features of core routers. However, traditional health analyzers need to store a large amount of historical data in order to identify health status. The storage requirement becomes prohibitively high when we attempt to carry out long-term health-status analysis for a large number of core routers. We describe the design of a symbol-based health status analyzer that first encodes, as a symbol sequence, the long-term complex time series collected from a number of core routers, and then utilizes the symbol sequence to do health analysis. The symbolic aggregation approximation (SAX), 1d-SAX, moving-average-based trend approximation, and nonparametric symbolic approximation representation methods are implemented to encode complex time series in a hierarchical way. Hierarchical agglomerative clustering and sequitur rule discovery are implemented to learn important global and local patterns. Three classification methods including a vector-space-model-based approach are then utilized to identify the health status of core routers. Data collected from a set of commercial core router systems are used to validate the proposed health-status analyzer. The experimental results show that our symbol-based health status analyzer requires much lower storage than traditional methods, but can still maintain comparable diagnosis accuracy. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Self-Learning and Efficient Health-Status Analysis for a Core Router SystemabstractThe health status of core router systems needs to be analyzed efficiently in order to ensure high reliability and timely error recovery. Although a large amount operational data is collected from core routers, due to high computational complexity and expensive labor cost, only a small part of this data is labeled by experts. The lack of labels is an impediment toward the adoption of supervised learning. We present an iterative self-learning procedure for assessing the health status of a core router. This procedure first computes a representative feature matrix to capture different characteristics of time-series data. Not only statistical-modeling-based features are computed from three general categories but also a recurrent neural network-based autoencoder is utilized to capture a wider range of hidden patterns. Moreover, both minimum-redundancy-maximum-relevance (mRMR) method and fully connected feedforward autoencoder are applied to further reduce dimensionality of extracted feature matrix. Hierarchical clustering is then utilized to infer labels for the unlabeled dataset. Finally, a classifier is built and iteratively updated using both labeled and unlabeled dataset. Field data collected from a set of commercial core routers are used to experimentally validate the proposed health-status analyzer. The experimental results show that the proposed feature-based self-learning health analyzer achieves higher precision and recall than the traditional supervised health analyzer as well the currently deployed rule-based health analyzer. Moreover, it achieves better performance than the three anomaly detection baseline methods under the transformed binary classification scenario. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Fine-grained Adaptive Testing Based on Quality PredictionabstractThe ever-increasing complexity of integrated circuits inevitably leads to high test cost. Adaptive testing provides an effective solution for test-cost reduction; this testing framework selects the important test items for each set of chips. However, adaptive testing methods designed for digital circuits are coarse-grained, and they are targeted only at systematic defects. To incorporate fabrication variations and random defects in the testing framework, we propose a fine-grained adaptive testing method based on machine learning. We use the parametric test results from the previous stages of test to train a quality-prediction model for use in subsequent test stages. Next, we partition a given lot of chips into two groups based on their predicted quality. A test-selection method based on statistical learning is applied to the chips with high predicted quality. An ad hoc test-selection method is proposed and applied to the chips with low predicted quality. Experimental results using a large number of fabricated chips and the associated test data show that to achieve the same defect level as in prior work on adaptive testing, the fine-grained adaptive testing method reduces test cost by 90% for low-quality chips and up to 7% for all the chips in a lot. Renjian Pan, Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2019 | Knowledge Transfer in Board-Level Functional Fault Identification using Domain AdaptationabstractHigh integration densities and design complexity make board-level functional fault identification extremely difficult. Machine-learning techniques can identify functional faults with high accuracy, but they require a large volume of data to achieve high prediction accuracy. This drawback limits the effectiveness of traditional machine-learning algorithms for training a model in the early stage of manufacturing, when only a limited amount of fail data and repair records are available. We propose a board-level diagnosis workflow that utilizes domain adaptation to transfer the knowledge learned from a mature board to a new board in the ramp-up phase. First, a metric is designed to evaluate the similarity between products, and based on the calculated value of the similarity, either a homogeneous or a heterogeneous domain adaptation algorithm is selected. Second, these domain adaptation algorithms utilize information from both the mature and the new boards with carefully designed domain-alignment rules and train a functional fault identification classifier. Three complex boards in volume production and one new board in the ramp-up phase are used to validate the proposed domain-adaptation approach in terms of the diagnosis accuracy. Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2019 | Board-Level Functional Fault Identification using Streaming DataabstractHigh integration densities and design complexity of printed-circuit boards make board-level functional fault identification extremely difficult. Machine learning provides an opportunity to identify functional faults with high accuracy and thereby reduce repair cost. However, the large volume of manufacturing data comes in a streaming format and exhibits time-dependent concept drift in a production environment. These drawbacks limit the effectiveness of traditional machine-learning algorithms. We propose a diagnosis workflow that utilizes online learning to train classifiers incrementally with a small chunk of data at each step. These online learning algorithms adapt to concept drift quickly with carefully designed update rules. A hybrid algorithm is also proposed to handle the scenario that data for varying numbers of boards are collected at different times. Experimental results using two boards in high-volume production show that, with the help of online learning and the proposed hybrid algorithm, the F1-score for diagnosis can be improved from 57.3% to 78.9%. Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 5 |
| 2019 | Black-Box Test-Coverage Analysis and Test-Cost Reduction Based on a Bayesian Network ModelabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. The conventional greedy algorithm, which selects the most important tests by considering both strong and weak relationships among tests, suffers from overfitting. In order to overcome overfitting, we propose a novel black-box test selection method based on a Bayesian network model. First, the problem of reducing black-box test cost is formulated as a constrained optimization problem. Next, a score-based algorithm is implemented to construct the Bayesian network for black-box tests. Finally, we propose a Bayesian index with the property of Markov blankets, and then an iterative test selection method is developed based on our proposed Bayesian index. The proposed approach ensures that only the strong relationships among black-box tests are used for test selection so that this approach is more robust to overfitting. Two case studies with production test data demonstrate that the proposed approach effectively reduces test cost by up to 14.7%, compared to a conventional greedy algorithm. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 5 |
| 2019 | IP Session on Machine Learning Applications in IC Test-Related TasksabstractOver the last decade there has been a surge of activity in employing advanced statistical analysis and machine learning methods to various test-related tasks. The topic is no longer simply a matter of academic curiosity but, rather, a pressing need of the industry as it seeks to address various challenges. In this session, three industry experts have been invited to give their perspective, describe machine learning use cases, and discuss challenges and future work ideas. The three talks will cover the use of deep learning for hotspot detection, the challenge of rendering machine learning-based decisions in the semiconductor industry trustable and explainable, and data analytics across the the complete product cycle towards improved product reliability. Ghada Sokar, Yassin Zakaria, Asmaa Rabie, Kareem Madkour, Ira Leventhal, Jochen Rivoir, Xinli Gu, Haralampos-G. D. Stratigopoulos |
VTS | 7 |
| 2019 | Changepoint-Based Anomaly Detection for Prognostic Diagnosis in a Core Router SystemabstractPrognostic diagnosis is desirable for commercial core router systems to ensure early failure prediction and fast error recovery. The effectiveness of prognostic diagnosis depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the statistical properties of the monitored data change significantly as time proceeds. We describe the design of a changepoint (CP)-based anomaly detector that first detects CPs from collected time-series data, and then utilizes these CPs to detect anomalies. Different CP detection approaches are implemented to detect various types of CPs. A clustering method is then developed to identify normal/abnormal patterns from CP windows. Data collected from a set of commercial core router systems are used to validate the proposed anomaly detector. Experimental results show that our CP-based anomaly detector achieves better performance than traditional methods in terms of two metrics, namely success ratio and nonfalse-alarm ratio. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Failure prediction based on anomaly detection for complex core routersabstractData-driven prognostic health management is essential to ensure high reliability and rapid error recovery in commercial core router systems. The effectiveness of prognostic health management depends on whether failures can be accurately predicted with sufficient lead time. This paper describes how time-series analysis and machine-learning techniques can be used to detect anomalies and predict failures in complex core router systems. First both a feature-categorization-based hybrid method and a changepoint-based method have been developed to detect anomalies in time-varying features with different statistical characteristics. Next, a SVM-based failure predictor is developed to predict both categories and lead time of system failures from collected anomalies. A comprehensive set of experimental results is presented for data collected during 30 days of field operation from over 20 core routers deployed by customers of a major telecom company. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ICCAD | 4 |
| 2018 | Self-Learning Health-Status Analysis for a Core Router SystemabstractThe health status of core router systems needs to be analyzed efficiently in order to ensure high reliability and timely error recovery. Although a large amount operational data is collected from core routers, only a small part of this data is labeled by experts. The lack of labels is an impediment towards the adoption of supervised learning. We present an iterative self-learning procedure for assessing the health status of a core router. This procedure first computes a representative feature matrix to capture different characteristics of time-series data. Hierarchical clustering is then utilized to infer labels for the unlabeled dataset. Finally, a classifier is built and iteratively updated using both labeled and unlabeled dataset. Field data collected from a set of commercial core routers are used to experimentally validate the proposed health-status analyzer. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2018 | Fine-Grained Adaptive Testing Based on Quality PredictionabstractThe ever-increasing complexity of integrated circuits inevitably leads to high test cost. Adaptive testing provides an effective solution for test-cost reduction; this testing framework selects the important test items for each set of chips. However, adaptive testing methods designed for digital circuits are coarse-grained, and they are targeted only at systematic defects. In order to incorporate fabrication variations and random defects in the testing framework, we propose a fine-grained adaptive testing method based on machine learning. We use the parametric test results from the previous stages of test to train a quality-prediction model for use in subsequent test stages. Next, we partition a given lot of chips into two groups based on their predicted quality. A test-selection method based on statistical learning is applied to the chips with high predicted quality. An ad hoc test-selection method is proposed and applied to the chips with low predicted quality. Experimental results using a large number of fabricated chips and the associated test data show that to achieve the same defect level as in prior work on adaptive testing, the fine-grained adaptive testing method reduces test cost by 90% for low-quality chips, and up to 7% for all the chips in a lot. Renjian Pan, Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 6 |
| 2018 | Toward Predictive Fault Tolerance in a Core-Router System: Anomaly Detection Using Correlation-Based Time-Series AnalysisabstractFault tolerance is used in communication systems to ensure high reliability and rapid error recovery. The effectiveness of most proactive fault-tolerant mechanism depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the monitored data involves temporal measurements and exhibits significantly different statistical characteristics for its constituent features. We describe the design of an anomaly detector that monitors the time-series data of a complex core router system. Anomaly detection techniques are compared in terms of their effectiveness for detecting different types of anomalies. A feature-categorizing-based hybrid method is proposed to overcome the difficulty of detecting anomalies in features with different statistical characteristics. Furthermore, a correlation analyzer is implemented to remove irrelevant and redundant features. Three types of synthetic anomalies, generated using a small amount of real data for a commercial telecom system, are used to validate the proposed anomaly detector. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | RetroDMR: Troubleshooting non-deterministic faults with retrospective DMRabstractThe most notorious faults for diagnosis in post-silicon validation are those that manifest themselves in a non-deterministic manner with system-level functional tests, where errors randomly appear from time to time even when applying the same workloads. In this work, we propose a novel diagnostic framework that resorts to dual-modular redundancy (DMR) for troubleshooting non-deterministic faults, namely RetroDMR. To be specific, we log the essential events (e.g., the sequence of thread migration) in the faulty run to record the mapping relationship between threads and their corresponding execution units. Then in the following diagnosis runs, we apply redundant multithreading (RMT) technique to reduce error detection latency, while at the same time we try to follow the thread migration sequence of the original run whenever possible. By doing so, RetroDMR significantly improves the reproduction rate and diagnosis resolution for non-deterministic faults, as demonstrated in our experimental results. Ting Wang 0008, Yannan Liu, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
DATE | 6 |
| 2017 | Data-driven fault diagnosis with missing syndromes imputation for functional test through conditional specificationabstractIn the electronic system manufacturing process, the board-level functional test is recognized as the most significant step to prevent defective products from entering the market. In recent years, machine learning and data mining have proven to be efficient techniques in determining root cause from the problematic functional test result, especially when the integrated circuits (IC) are becoming increasingly highly-integrated. However, the test results are sometimes unavailable due to either abnormal ending of the test sequence or occasional system failures, which results in a decreased performance of data-driven diagnosis systems. In this paper, we propose a data imputation algorithm to predict the missing entries in the functional test result, by considering the correlation between test items with conditional specification. We evaluate our data imputation algorithm over the test results collected from three different stages of functional test on a line card used in the telecommunication system. The result shows that our proposed data imputation algorithm consistently outperforms other imputation techniques with various data-driven approaches in terms of diagnosing the root cause, increasing the diagnosis accuracy by an average of 28.13% compared to none data imputation, and 9.74% compared to the naive pass imputation. Tong Guan, Zhaobo Zhang, Wen Dong 0001, Chunming Qiao, Xinli Gu |
ETS | 5 |
| 2017 | Changepoint-based anomaly detection in a core router systemabstractPrognostic diagnosis is desirable for commercial core router systems to ensure early failure prediction and fast error recovery. The effectiveness of prognostic diagnosis depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the statistical properties of the monitored data change significantly as time proceeds. We describe the design of a changepoint-based anomaly detector that first detects changepoints from collected time-series data, and then utilizes these changepoints to detect anomalies. Two approaches based on maximum-likelihood estimation are implemented to detect different types of changepoints. A clustering method is then developed to identify a wide range of normal/abnormal patterns from changepoint windows. Data collected from a set of commercial core router systems are used to validate the proposed anomaly detector. Experimental results show that our changepoint-based anomaly detector achieves better performance than traditional methods. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2017 | Symbol-based health-status analysis in a core router systemabstractTo ensure high reliability and rapid error recovery in commercial core router systems, a health-status analyzer is essential to monitor the different features of core routers. However, traditional health analyzers need to store a large amount of historical data in order to identify health status. The storage requirement becomes prohibitively high when we attempt to carry out long-term health-status analysis for a large number of core routers. We describe the design of a symbol-based health status analyzer that first encodes, as a symbol sequence, the long-term complex time series collected from a number of core routers, and then utilizes the symbol sequence to do health analysis. The symbolic aggregation approximation (SAX) and moving-average-based trend approximation methods are implemented to encode complex time series in a hierarchical way. Hierarchical agglomerative clustering and sequitur rule discovery are implemented to learn important global and local patterns. Two classification methods are then utilized to identify the health status of core routers. Data collected from a set of commercial core router systems are used to validate the proposed health-status analyzer. The experimental results show that our symbol-based health status analyzer requires much lower storage than traditional methods, but can still maintain comparable diagnosis accuracy. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2016 | Accurate anomaly detection using correlation-based time-series analysis in a core router systemabstractFault tolerance is used in communication systems to ensure high reliability and rapid error recovery. The effectiveness of most proactive fault-tolerant mechanism depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the monitored data involves temporal measurements and exhibits significantly different statistical characteristics for its constituent features. We describe the design of an anomaly detector that monitors the time-series data of a complex core router system. Anomaly detection techniques are compared in terms of their effectiveness for detecting different types of anomalies. A feature-categorizing-based hybrid method is proposed to overcome the difficulty of detecting anomalies in features with different statistical characteristics. Furthermore, a correlation analyzer is implemented to remove irrelevant and redundant features. Three types of synthetic anomalies, generated using a small amount of real data for a commercial telecom system, are used to validate the proposed anomaly detector. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2016 | Efficient Board-Level Functional Fault Diagnosis With Missing SyndromesabstractFunctional fault diagnosis is widely used in board manufacturing to ensure product quality and improve product yield. Advanced machine-learning techniques have recently been advocated for reasoning-based diagnosis; these techniques are based on the historical record of successfully repaired boards. However, traditional diagnosis systems fail to provide appropriate repair suggestions when the diagnostic logs are fragmented and some error outcomes, or syndromes, are not available during diagnosis. We describe the design of a diagnosis system that can handle missing syndromes and can be applied to four widely used machine-learning techniques. Several imputation methods are discussed and compared in terms of their effectiveness for addressing missing syndromes. Moreover, a syndrome-selection technique based on the minimum-redundancy-maximum-relevance criteria is also incorporated to further improve the efficiency of the proposed methods. Two large-scale synthetic data sets generated from the log information of complex industrial boards in volume production are used to validate the proposed diagnosis system in terms of diagnosis accuracy and training time. Shi Jin 0001, Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Adaptive Board-Level Functional Fault Diagnosis Using Incremental Decision TreesabstractBoard-level functional fault diagnosis is needed for high-volume production to improve product yield. However, to ensure diagnosis accuracy and effective board repair, a large number of syndromes must be used. Therefore, the diagnosis cost can be prohibitively high due to the increase in diagnosis time and the complexity of test execution and analysis. We propose an adaptive diagnosis method based on incremental decision trees (DTs). Faulty components are classified according to the discriminative ability of the syndromes in DT training. The diagnosis procedure is constructed as a binary tree, with the most discriminative syndrome as the root and final repair suggestions are available as the leaf nodes of the tree. The syndrome to be used in the next step is determined based on the observation of syndromes thus far in the diagnosis procedure. The number of syndromes required for diagnosis can be significantly reduced compared to the total number of syndromes used for system training. Moreover, online learning is facilitated in the proposed diagnosis system using an incremental version of DTs, so as to bridge the knowledge obtained at test-design stage with the knowledge gained during volume production. The diagnosis system can thus adapt to occurrences of new error scenarios on-the-fly. Diagnosis results for three complex boards from industry, currently in volume production, highlight the effectiveness of the proposed approach. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | On test syndrome merging for reasoning-based board-level functional fault diagnosisabstractMachine learning algorithms are advocated for automated diagnosis of board-level functional failures due to the extreme complexity of the problem. Such reasoning-based solutions, however, remain ineffective at the early stage of the product cycle, simply because there are insufficient historical data for training the diagnostic system that has a large number of test syndromes. In this paper, we present a novel test syndrome merging methodology to tackle this problem. That is, by leveraging the domain knowledge of the diagnostic tests and the board structural information, we adaptively reduce the feature size of the diagnostic system by selectively merging test syndromes such that it can effectively utilize the available training cases. Experimental results demonstrate the effectiveness of the proposed solution. Zelong Sun, Li Jiang 0002, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
ASP-DAC | 6 |
| 2015 | Self-learning and adaptive board-level functional fault diagnosisabstractFunctional fault diagnosis is necessary for board-level product qualification. However, ambiguous diagnosis results can lead to long debug times and wrong repair actions, which significantly increase repair cost and adversely impact yield. A state-of-the-art functional fault diagnosis system involves several key components: (1) design of functional test programs, (2) collection of functional-failure syndromes, (3) building of the diagnosis engine, (4) isolation of root causes, and (5) evaluation of the diagnosis engine. Advances in each of these components can pave the way for a more effective diagnosis system, thus improving diagnosis accuracy and reducing diagnosis time. Machine-learning and data analysis techniques offer an unprecedented opportunity to develop an automated and adaptive diagnosis system to increase diagnosis accuracy and reduce diagnosis time. This paper describes how all the above components of an advanced diagnosis system can benefit from machine learning and information theory. Topics discussed include incremental learning, decision trees, root-cause analysis and evaluation metrics, data acquisition, and knowledge transfer. Fangming Ye, Krishnendu Chakrabarty, Zhaobo Zhang, Xinli Gu |
ASP-DAC | 4 |
| 2015 | Information-Theoretic Syndrome Evaluation, Statistical Root-Cause Analysis, and Correlation-Based Feature Selection for Guiding Board-Level Fault DiagnosisabstractReasoning-based functional-fault diagnosis has recently been advocated to achieve high diagnosis accuracy, low defect escapes, and reducing manufacturing cost. However, such diagnosis method requires a rich set of test items (syndromes) and a sizable database of faulty boards to learn from. An insufficient number of failed boards, ambiguous root-cause identification, and redundant or irrelevant syndromes can render reasoning-based diagnosis ineffective. Periodic evaluation and analysis can help locate weaknesses in a diagnosis system and thereby provide guidelines for redesigning the tests, which facilitates better diagnosis. We propose an information-theoretic framework for evaluating the effectiveness of and providing guidance to a reasoning-based functional-fault diagnosis system. Syndrome analysis based on feature selection methods provides a representative set of syndromes and suggests irrelevant syndromes in diagnosis. Root-cause analysis measures the discriminative ability of differentiating a given root cause from others. Results are presented for four types of diagnosis systems for three complex boards that are in volume production. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Knowledge discovery and knowledge transfer in board-level functional fault diagnosisabstractDiagnosis of functional failures at the board level is critical for improving product yield and reducing manufacturing cost. Reasoning techniques increase the accuracy of functional-fault diagnosis based on the history of successfully repaired boards. However, depending on the complexity of the product, it usually takes several months to accumulate an adequate database for training a reasoning-based diagnosis system. During the initial product ramp-up phase, reasoning-based diagnosis is not feasible for yield learning, since the required database is not available due to lack of volume. We propose a knowledge-discovery method and a knowledge-transfer method for facilitating board-level functional fault diagnosis. First, an analysis technique based on machine learning is used to discover knowledge from syndromes, which can be used for training a diagnosis engine. Second, knowledge from diagnosis engines used for earlier-generation products can be automatically transferred through root-cause mapping and syndrome mapping based on keywords and board-structure similarities. Two complex boards in volume production and with a mature diagnosis system, and three new boards in the ramp-up phase, are used to validate the proposed knowledge-discovery and knowledge-transfer approach in terms of the diagnosis accuracy obtained using the new diagnosis systems. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2014 | Special session 8B - Panel: In-field testing of SoC devices: Which solutions by which players?abstractIn-field testing of SoC devices is increasingly important to face the dependability requirements of several application domains. Different solutions can be devised and adopted. We summarize the main solutions currently adopted by industry, identify the most critical open issues, and discuss important future trends. Jacob A. Abraham, Xinli Gu, Teresa MacLaurin, Janusz Rajski, Paul G. Ryan, Dimitris Gizopoulos, Matteo Sonza Reorda |
VTS | 2 |
| 2014 | Board-Level Functional Fault Diagnosis Using Multikernel Support Vector Machines and Incremental LearningabstractAdvanced machine learning techniques offer an unprecedented opportunity to increase the accuracy of board-level functional fault diagnosis and reduce product cost through successful repair. Ambiguous or incorrect diagnosis results lead to long debug times and even wrong repair actions, which significantly increase repair cost. We propose a smart diagnosis method based on multikernel support vector machines (MK-SVMs) and incremental learning. The MK-SVM method leverages a linear combination of single kernels to achieve accurate faulty-component classification based on the errors observed. The MK-SVMs thus generated can also be updated based on incremental learning, which allows the diagnosis system to quickly adapt to new error observations and provide even more accurate fault diagnosis. Two complex boards from industry, currently in volume production, are used to validate the proposed diagnosis approach in terms of diagnosis accuracy (success rate) and quantifiable improvements over previously proposed machine-learning methods based on several single-kernel SVMs and artificial neural networks. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2013 | Handling Missing Syndromes in Board-Level Functional-Fault DiagnosisabstractFunctional fault diagnosis is widely used in board manufacturing to ensure product quality and improve product yield. Advanced machine-learning techniques have recently been advocated for reasoning-based diagnosis, these technologies are based on historical data of successfully repaired boards. However, traditional diagnosis systems fail to provide appropriate repair suggestions when the diagnostic logs are fragmented and some error outcomes, or syndromes, are not available during diagnosis. We describe the design of a diagnosis system, based on support vector machines, that can handle missing syndromes by using the method of imputation. Several imputation methods are discussed and compared in terms of their efficiency in handling missing syndromes. Two large-scale synthetic data sets generated from the log information of complex industrial boards in volume production are used to validate the proposed diagnosis system in terms of diagnosis accuracy and training time. Fangming Ye, Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
Asian Test Symposium | 5 |
| 2013 | Panel session what is the electronics industry doing to win the battle against the expected scary failure rates in future technology nodes?abstractSummary form only given. The major bottleneck for technology scaling is the growing rate of hardware failures. Process variations are becoming extreme and sensitivity to radiation is becoming severe. In addition, intrinsic failures such as device parameter degradation are accelerating the wear-out. All of these are leading to higher random in-filed failures and shorter device lifetime. The 2011 ITRS (International Technology Roadmap for Semiconductors) projects very high bit failure rates of the order of 10-2for SRAM and of 10-3for latches for 16nm high performance technology. Hence, solving reliability challenges for future technologies requires new efficient and cost effective approaches not only to detect and recover from in-filed failures, but also to extend the device lifetime for targeted applications. Said Hamdioui, Davide Appello, Arnaud Grasset, Xinli Gu, Bram Kruseman, Riccardo Mariani, Hermann Obermeir, Srikanth Venkataraman |
ETS | 4 |
| 2013 | Information-theoretic syndrome and root-cause analysis for guiding board-level fault diagnosisabstractHigh-volume manufacturing of complex electronic products involves functional test at board level to ensure low defect escapes. Machine-learning techniques have recently been proposed for reasoning-based functional-fault diagnosis system to achieve high diagnosis accuracy. However, machine learning requires a rich set of test items (syndromes) and a sizable database of faulty boards. An insufficient number of failed boards, ambiguous root-cause identification, and redundant or irrelevant syndromes can render machine learning ineffective. We propose an evaluation and enhancement framework based on information theory for guiding diagnosis systems using syndrome and root-cause analysis. Syndrome analysis based on subset selection provides a representative set of syndromes with minimum redundancy and maximum relevance. Root-cause analysis measures the discriminative ability of differentiating a given root cause from others. The metrics obtained from the proposed framework can also provide guidelines for test redesign to enhance diagnosis. A real board from industry, currently in volume production, and an additional synthetic board, based on data extrapolated from another real board, are used to demonstrate the effectiveness of the proposed framework. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ETS | 4 |
| 2013 | AgentDiag: An agent-assisted diagnostic framework for board-level functional failuresabstractDiagnosing functional failures in complicated electronic boards is a challenging task, wherein debug technicians try to identify defective components by analyzing some syndromes obtained from the application of diagnostic tests. The diagnosis effectiveness and efficiency rely heavily on the quality of the in-house developed diagnostic tests and the debug technicians' knowledge and experience, which, however, have no guarantees nowadays. To tackle this problem, we propose a novel agent-assisted diagnostic framework for board-level functional failures, namely AgentDiag, which facilitates to evaluate the quality of the diagnostic tests and bridge the knowledge gap between the diagnostic programmers who write diagnostic tests and the debug technicians who conduct in-field diagnosis with a lightweight model of the boards and tests. Experimental results on a real industrial board and an OpenRISC design demonstrate the effectiveness of the proposed solution. Zelong Sun, Li Jiang 0002, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
ITC | 6 |
| 2013 | Board-Level Functional Fault Diagnosis Using Artificial Neural Networks, Support-Vector Machines, and Weighted-Majority VotingabstractIncreasing integration densities and high operating speeds lead to subtle manifestation of defects at the board level. Functional fault diagnosis is, therefore, necessary for board-level product qualification. However, ambiguous diagnosis results lead to long debug times and even wrong repair actions, which significantly increase repair cost and adversely impact yield. Advanced machine-learning (ML) techniques offer an unprecedented opportunity to increase the accuracy of board-level functional diagnosis and reduce high-volume manufacturing cost through successful repair. We propose a smart diagnosis method based on two ML classification models, namely, artificial neural networks (ANNs) and support-vector machines (SVMs) that can learn from repair history and accurately localize the root cause of a failure. Fine-grained fault syndromes extracted from failure logs and corresponding repair actions are used to train the classification models. We also propose a decision machine based on weighted-majority voting, which combines the benefits of ANNs and SVMs. Three complex boards from the industry, currently in volume production, and additional synthetic data, are used to validate the proposed methods in terms of diagnostic accuracy, resolution, and quantifiable improvement over current diagnostic software. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Session Summary II: Dependable VLSI for Product ReliabilityabstractSummary form only given, as follows. This session presents the importance of product quality and reliability from both a customer perspective and VLSI component perspective. DFT technologies ranging from VLSI component design, manufacturing, test to diagnosis and even to product customer site (in-field) test are discussed to help product reliability. Four presentations in this session will cover quality memory test, flash storage in-field test, special ASIC DFT feature design and board/system usage of component level DFT features. Xinli Gu |
Asian Test Symposium | 1 |
| 2012 | Session Summary V: Is Component Interconnection Test Enough for Board or System TestabstractSummary form only given, as follows. This panel discusses board/system level test requirements, current test technologies and challenges. We will also discuss what future new standards, technologies and tools are necessary to improve both the test quality and test efficiency for production boards. This panel will cover from an end-to-end standpoint to improve product quality and reliability, including DFT technologies in ASIC components, boards/systems, and intelligent diagnosis. Xinli Gu |
Asian Test Symposium | 1 |
| 2012 | In-Field Testing of NAND Flash Storage: Why and How?abstractNAND Flash memories have rapidly emerged as a storage class memory such as SSD (Solid State Disk), CF (Compact Flash) Card, SD (Secure Digital Memory) Card. Due to its distinct operation mechanisms, NAND Flash memory suffers from erase/program endurance, data retention and program/read disturbance problems. Specifically, erase and program operation keeps in developing bad blocks during the lifetime of memory chips. Bad blocks are blocks that contain faulty bits but the ECC (Error Correction Code) algorithm cannot correct them. Although wear leveling tries to balance the erase/program operations on different blocks so that all blocks can wear out at a similar pace, new bad blocks still inevitably occur.We propose an in-field testing technique which takes some pages in a block as predictors. Due to wear out faster than the other pages, the predictors will become bad before the other pages in the block become bad. The further questions are (1) how to detect those wearing fast pages so as to use them as predictors, (2) how many predictors are needed to achieve a satisfactory prediction accuracy, (3) misprediction will result in what negative impact on performance and endurance. Yu Hu 0001, Xinli Gu, Xiaowei Li 0001 |
Asian Test Symposium | 2 |
| 2012 | Adaptive Board-Level Functional Fault Diagnosis Using Decision TreesabstractFunctional fault diagnosis at board-level is desirable for high-volume production since it improves product yield. However, to ensure diagnosis accuracy and effective board repair, a large number of syndromes must be used. Therefore, the diagnosis cost can be prohibitively high due to the increase in diagnosis time and the complexity of syndrome collection/analysis. We propose an adaptive diagnosis method based on decision trees (DTs). Faulty components are classified according to the discriminative ability of the syndromes in DT training. The diagnosis procedure is constructed as a binary tree, with the most discriminative syndrome as the root and final repair suggestions are available as the leaf nodes of the tree. The syndrome to be collected in the next step is determined based on the observations of syndromes collected thus far in the diagnosis procedure. The number of syndromes required for diagnosis can also be significantly reduced compared to the number of syndromes used for system training. Diagnosis results for two complex boards from industry, currently in volume production, and additional synthetic data highlight the effectiveness of the proposed approach. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
Asian Test Symposium | 4 |
| 2012 | Board-Level Functional Fault Diagnosis Using Learning Based on Incremental Support-Vector MachinesabstractAdvanced machine learning techniques offer an unprecedented opportunity to increase the accuracy of board-level functional fault diagnosis based on the historical data of successfully repaired boards. However, the training complexity increases significantly in diagnosis systems due to the increasing amount of the historical data. We propose a smart learning method in the diagnosis system using incremental support-vector machines (SVMs). The SVMs updated using incremental learning allow the diagnosis system to quickly adapt to new error observations and provide more accurate fault diagnosis. Two sets of large-scale synthetic data generated from the log information of two complex industrial boards, in volume production, are used to validate the proposed diagnosis approach in terms of training time and diagnosis accuracy over a previously proposed diagnosis system based on simple support-vector machines. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
Asian Test Symposium | 4 |
| 2012 | Re-using chip level DFT at board levelabstractAs chips are getting increasingly complex, there is no surprise to find more and more built-in DFX. This built-in DFT is obviously beneficial for chip/silicon DFX engineers; however, board/system level DFX engineers often have limited access to the build in DFX features. There is currently an increasing demand from board/system level DFX engineers to reuse chip/silicon DFX at board/system level. This special session will discuss: What chip access is needed for board-level for test and diagnosis? How to accomplish the access? Will IEEE P1687 and IEEE 1149.1 solve these problems? Xinli Gu, Jeff Rearick, Bill Eklow, Martin Keim, Artur Jutman, Krishnendu Chakrabarty, Erik Larsson |
ETS | 1 |
| 2012 | Diagnostic system based on support-vector machines for board-level functional diagnosisabstractFault diagnosis is critical for improving product yield and reducing manufacturing cost. However, it is very challenging to identify the root cause of failures on a complex circuit board. Ambiguous diagnosis results lead to long debug times and even wrong repair actions, which significantly increases the repair cost. We propose an automatic diagnostic system using support vector machines (SVMs). The proposed system acquires debug knowledge from empirical data; this strategy avoids the difficulties involved in knowledge acquisition in traditional fault diagnosis methods. SVMs provide an optimal separating hyperplane in classification. The optimal solution and generalization ability of SVMs lead to higher diagnostic accuracy, compared to the classical learning approaches such as artificial neural networks (ANNs). An industrial board is used to validate the effectiveness of the proposed system. Extensive simulation results demonstrate that the SVMs-based diagnostic system provides quantifiable improvement over current diagnostic software and an ANN-based diagnostic system. Zhaobo Zhang, Xinli Gu, Yaohui Xie, Zhanglei Wang, Krishnendu Chakrabarty |
ETS | 2 |
| 2012 | Are industrial test problems real problems? I thought research has resolved them all!abstractIs there any real industrial test problem that requires research collaboration, or is there no such thing? We will hear from both industrial and research folks about their experiences, both successful and otherwise. How do we get past statements like these and move on to genuine and effective collaboration? █ Industry: You have to make it truly work for us … █ Research: You have to fund my students first … and give us your designs █ Industry: bye !! … █ Research: So we still have no industrial test bench, … and funding : Following this panel session, there will be an invited poster session from both industry and research demonstrating their problems, technologies and their willingness to look for cooperation. Visit these posters and tell them either you solved their problem 10 years ago, or ask for a check to solve them.… Xinli Gu |
ITC | 1 |
| 2012 | Reproduction and Detection of Board-Level Functional FailureabstractNo trouble found (NTF) due to functional failures is a common scenario today in board-level testing at system companies. A component on a board fails during the board-level functional test, but it passes the automatic test equipment (ATE) test when it is returned to the supplier for warranty replacement or service repair. To find the root cause of NTF, we propose an innovative functional test approach and discrete Fourier transform (DFT) methods for the detection of board-level functional failures. These DFT and test methods allow us to reproduce and detect functional failures in a controlled deterministic environment, and provide ATE tests to the supplier for early screening of defective parts. Experiments on an industry design show that the proposed functional scan test with appropriate functional constraints can adequately mimic the functional state space, as measured by appropriate coverage metrics. Experiments also show that most functional failures due to dominant bridging, crosstalk, and delay faults due to power supply noise can be reproduced and detected by functional scan test. We also describe two approaches to enhance the mimicking of the functional state space. The first approach allows us to select a given number of initial states in linear time and functional scan tests resulting from these selected states are used to mimic the functional state space. The second approach is based on controlled state injection. Hongxia Fang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Diagnosis of Board-Level Functional Failures Under Uncertainty Using Dempster-Shafer TheoryabstractDespite recent advances in structural test methods, the diagnosis of the root cause of board-level failures for functional tests remains a major challenge. A promising approach to address this problem is to carry out fault diagnosis in two phases-suspect faulty components on the board or modules within components (together referred to as blocks in this paper) are first identified and ranked, and then fine-grained diagnosis is used to target the suspect blocks in a ranked order. We propose a new method based on dataflow analysis and Dempster-Shafer (DS) theory for ranking faulty blocks in the first phase of diagnosis. The proposed approach transforms the information derived from one functional test failure into multiple-stage failures by partitioning the given functional test into multiple stages. A measure of “belief” is then assigned to each block based on the knowledge of each failing stage, and the DS theory is subsequently used to aggregate the beliefs from multiple failing stages. Blocks with higher beliefs are ranked on the top of the candidate list. Simulations on an industry design for a network interface application as well as on an open source system-on-a-chip show that the proposed method can provide accurate ranking for most board-level functional failures. Hongxia Fang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Physical-Defect Modeling and Optimization for Fault-Insertion TestabstractHardware fault insertion is a promising method for system reliability assessment and fault isolation. It provides feedback on the fault tolerance of a large system, creates artificial faulty scenarios that can be used as reference points for fault diagnosis, and leads to a quality diagnostic program. Optimization of fault insertion location is critical for accelerating the assessment of system reliability and constructing a complete knowledge base for fault diagnosis. In this work, we construct a pin-level fault model that is able to effectively mimic the errors (effects) caused by physical defects within the component. A simulation framework and optimization techniques are proposed to select a minimum subset of output pins that can represent as many physical defects as possible. The optimization results provide guidelines on the fault insertion locations and the appropriate fault types for insertion. In addition, three intrinsic characteristics of output pins, including testability number, fan-in size, and transition counts, are analyzed. The effectiveness of the proposed model is evaluated in terms of impact on system response and error-detection latency. Experimental results are presented for OpenCore benchmarks. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Deterministic test for the reproduction and detection of board-level functional failuresabstractA common scenario in industry today is “No Trouble Found” (NTF) due to functional failures. A component on a board fails during board-level functional test, but it passes the Automatic Test Equipment (ATE) test when it is returned to the supplier for warranty replacement or service repair. To find the root cause of NTF, we propose an innovative functional test approach and DFT methods for the detection of boardlevel functional failures. These DFT and test methods allow us to reproduce and detect functional failures in a controlled deterministic environment, which can provide ATE tests to the supplier for early screening of defective parts. Experiments on an industry design show that functional scan test with appropriate functional constraints can adequately mimic the functional state space well (measured by appropriate coverage metrics). Experiments also show that most functional failures due to stuck-at, dominant bridging, and crosstalk faults can be reproduced and detected by functional scan test. Hongxia Fang, Xinli Gu, Krishnendu Chakrabarty |
ASP-DAC | 3 |
| 2011 | Ranking of Suspect Faulty Blocks Using Dataflow Analysis and Dempster-Shafer Theory for the Diagnosis of Board-Level Functional FailuresabstractDespite recent advances in structural test methods, the diagnosis of the root cause of board-level failures for functional tests remains a major challenge. A promising approach to address this problem is to carry out fault diagnosis in two phases -- suspect faulty components on the board or modules within components (together referred to as blocks in this paper) are first identified and ranked, and then fine-grained diagnosis is used to target the suspect blocks in ranked order. We propose a new method based on dataflow analysis and Dempster-Shafer theory for ranking faulty blocks in the first phase of diagnosis. The proposed approach transforms the information derived from one functional test failure into multiple-stage failures by partitioning the given functional test into multiple stages. A measure of "belief" is then assigned to each block based on the knowledge of each failing stage, and Dempster-Shafer theory is subsequently used to aggregate the beliefs from multiple failing stages. Blocks with higher beliefs are ranked at the top of the candidate list. Simulations on an industry design for a network interface application show that the proposed method can provide accurate ranking for most board-level functional failures. Hongxia Fang, Xinli Gu, Krishnendu Chakrabarty |
ETS | 3 |
| 2011 | The gap: Test challenges in Asia manufacturing fieldabstractToday, more and more electronic manufacturing is being done in Asia. Test, as one of the important functions of manufacturing, is critical to guarantee the product quality and as a monitor of the manufacturing process. Through presenting the gap between the test challenges the Asia companies are facing and the tools they have today, we hope give the ITC community an opportunity to better understand the needs of innovation for test technologies and tools. This panel invites speakers from companies in Asia to present their test challenges and current solutions. It covers both design for test challenges and the complexities of the production test challenges. The focus is on the gap between what they need with today's challenges and the current capabilities. Xinli Gu |
ITC | 1 |
| 2011 | Smart diagnosis: Efficient board-level diagnosis and repair using artificial neural networksabstractDiagnosis of functional failures at the board level is critical for improving product yield and reducing manufacturing cost. State-of-the-art board-level diagnostic software is unable to cope with high complexity and ever-increasing clock frequencies, and the identification of the root cause of failure on a board is a major problem today. Ambiguous or incorrect repair suggestions lead to long debug times and even wrong repair actions, which significantly increases the repair cost and adversely impacts yield. We propose a smart diagnosis method based on artificial neural networks that can learn from repair history and accurately localize the root cause of a failure. Fine-grained fault syndromes extracted from failure logs and the corresponding repair actions are used to train the neural network. The proposed network structure is simple, it can be rapidly trained, and it is scalable to large datasets. Moreover, the relationship between typical syndromes and the most appropriate repair actions can be easily inferred from the network structure. An industrial board, which is currently in production, is used to validate the diagnosis approach in terms of diagnostic accuracy, resolution, and quantifiable improvement over current diagnostic software. Zhaobo Zhang, Krishnendu Chakrabarty, Zhanglei Wang, Xinli Gu |
ITC | 5 |
| 2010 | Mimicking of Functional State Space with Structural Tests for the Diagnosis of Board-Level Functional FailuresabstractA common scenario in industry today is “No Trouble Found” (NTF) due to functional failures. A component on a board fails during board-level functional test, but it passes the Automatic Test Equipment (ATE) test when it is returned to the supplier for warranty replacement or service repair. To find the root cause of NTF, we define an innovative deterministic test, namely functional scan test. We also propose two approaches for using functional scan test to adequately mimic functional state space. The first approach allows us to select a given number of initial states in linear time and functional scan tests resulting from these selected states are used to mimic the functional state space effectively. The second approach adopts a state-injection technique. Experiments on an industry design show that by using either multiple initial states or state injection, functional scan test with appropriate functional constraints can mimic the functional state space well, measured by appropriate coverage metrics. Therefore, it is feasible to use functional scan test to detect board-level functional failures in a controlled deterministic environment and diagnose the root cause to faulty wires/gates inside a component. It is also shown that the proposed method outperforms a random method in selecting the given number of effective initial states. Hongxia Fang, Xinli Gu, Krishnendu Chakrabarty |
Asian Test Symposium | 3 |
| 2010 | Optimization and Selection of Diagnosis-Oriented Fault-Insertion Points for System TestabstractHardware fault-insertion test is a promising method to diagnose functional failures and target ''no trouble found (NTF)" problems in electronic systems. However, it is costly and impractical to equip all the potential fault sites with fault-insertion hardware. We present an optimization method to select the most effective outputs of a module where fault insertion logic must be placed to facilitate diagnosis. Faults inserted at the selected outputs are able to generate fault syndromes that are most similar to the errors produced by defects inside the module. This approach also ensures that the ambiguous fault candidates from other modules are maximally removed from the set of suspects. A fault syndrome is defined by the order of error occurrence at the observation points, and it is referred as an error flow. The similarity between two error flows is measured by the metric of edit distance. An integer linear programming model is used to maximize diagnostic effectiveness with a small number of fault-insertion points. Results on diagnostic accuracy for an open-source RISC highlight the effectiveness of the proposed method compared to a baseline random fault-insertion scheme. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
Asian Test Symposium | 3 |
| 2010 | Board-level fault diagnosis using an error-flow dictionaryabstractDiagnosis of functional failures is critical for locating manufacturing defects, increasing yield, and reducing field returns. It is important to narrow down the defective module in a failed component during board-level diagnosis. In this paper, a generic fault-diagnosis method based on an error-flow dictionary is presented to identify the root cause of functional failures on a chip or board. Error propagation mimics actual dataflow in a circuit, thus it reflects the native (functional) mode of circuit operation. In contrast to conventional fault syndromes, error flow includes the failure information in terms of circuit functionality, which significantly facilitates the diagnosis of functional failures. In the proposed diagnosis procedure, error flow is first learned from a good circuit by intentionally inserting faults, and then the root cause of a failing circuit is determined by comparing the similarity between the pre-learned error flow and the error flow observed from the failing circuit. The similarity of two error flows is evaluated based on the length of the longest common subsequence in string matching. Results for an open-source RISC SoC and an industrial communication circuit highlight the effectiveness of the proposed method. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
ITC | 3 |
| 2010 | Board-level fault diagnosis using Bayesian inferenceabstractIncreasing integration densities and high operating speeds are leading to subtle manifestations of defects at the board level. Board-level functional test is therefore necessary for product qualification. The diagnosis of functional failures is especially challenging, and the cost associated with board-level diagnosis is escalating rapidly. An effective and cost-efficient board-level diagnosis strategy is needed to reduce manufacturing cost and time-to-market, as well as to improve product quality. In this paper, we use Bayesian inference to develop a new board-level diagnosis framework that allows us to identify faulty devices or faulty modules within a device on a failing board with high confidence. Bayesian inference offers a powerful probabilistic method for pattern analysis, classification, and decision making under uncertainty. We apply this inference technique by first generating a database of fault syndromes obtained using fault-insertion test at the module pin level on a fault-free board, and then use this database along with the observed erroneous behavior of a failing board to infer the most likely faulty device. Results on a case study using an open-source RISC system-on-chip highlight the effectiveness of the proposed framework in terms of fault-localization accuracy and correctness of diagnosis. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
VTS | 3 |
| 2009 | Physical defect modeling for fault insertion in system reliability testabstractHardware fault-insertion test (FIT) is a promising method for system reliability test and diagnosis coverage measurement. It improves the speed of releasing a quality diagnostic program before manufacturing and provides feedbacks of fault tolerance of a very complicated large system. Certain level insufficient fault tolerance can be fixed in the current system but others may require ASIC or overall system architectural modifications. The FIT is achieved by introducing an artificial fault (defect modeling) at the pin level of a module to mimic any physical defect behavior within the module, such as SEU (single event upset) or escaped delay defect. We present a hardware architectural solution for pin fault insertion. We also present a simulation framework and optimization techniques for a subset of module pin selection for FIT, such that desired coverage are obtained under the constraints of limited FIT pins due to the costs of the associated implementation. Experimental results are presented for selected ISCAS and OpenCore benchmarks, as well as for an industrial circuit. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
ITC | 3 |
| 2006 | Design for Board and System Level Structural Test and DiagnosisabstractThe success of system test is measured by test quality and cost. System test quality and cost rely on several factors, such as component and board test quality, system test completeness, the support of system diagnostics, and a process that controls overall quality, resource and cost balances. Traditional structural test techniques used at the component level can achieve both high test quality and low test costs. This paper describes an approach to extend the functionalities of structural test techniques to the board and system level to improve the test accessibility, test time, and diagnostic capability. This approach has become practice in a large telecommunication company and the benefits received from this practice are tremendous. Examples will be given at the end of the paper Toai Vo, Ted Eaton, Pradipta Ghosh, Huai Li, Hong Shin Jun, Rong Fang, Dan Singletary, Xinli Gu |
ITC | 11 |
| 2005 | A practical perspective on reducing ASIC NTFsabstractAs chip, board and system technologies scale towards higher speeds and greater logic density, the effect of defects becomes more subtle, but more pervasive. As hardware designers push technologies to the limit, system DPM rates continue to increase, yet it becomes increasingly difficult to determine the nature of the failures. More and more often, components/ASICs which fail at board and system test are sent to suppliers, only to have them returned "NTF" (no trouble found). This paper presents some of the issues that Cisco Systems has experienced with respect to NTFs, and how some of those issues were resolved. These issues span from chip to system and from process to test to debug. The paper discusses the importance of a process to deal with NTFs and the importance of accurate data to determine and fix unwanted trends. Ultimately, most problems were resolved once the trend data and the offending logic were completely understood. Not all NTFs resulted from test escapes. It was clear, however, that some sort of "correlation" between the ASIC test and the system test needed to be in place to resolve/prevent NTF issues. In its conclusion, this paper advocates for much better correlation between the ASIC test on the component tester and the functional test in the system chassis. Zoe Conroy, Geoff Richmond, Xinli Gu, Bill Eklow |
ITC | 3 |
| 2004 | Realizing High Test Quality Goals with Smart Test Resource UsageabstractGrowing ASIC design sizes and advanced deep sub-micron technologies require new fault models and more test vectors to meet high test quality goals. To realize these goals within given test resources and cost constraints, new DFT techniques must be used. This paper reports test quality metrics and the test cost of industrial designs for different fault models using three DFT techniques: ATPG for deterministic patterns, Logic BIST for pseudo-random patterns, and EDT for compressed deterministic patterns. It is shown how these techniques can be used to achieve the high quality goals within the test resources currently available for stuck-at tests. Xinli Gu, Cyndee Wang, Abby Lee, Bill Eklow, Kun-Han Tsai, Jan Arild Tofte, Mark Kassab, Janusz Rajski |
ITC | 1 |
| 2004 | At-Speed Interconnect Test and Diagnosis of External Memories on a SystemabstractThis work presents a built-in self test (BIST) implementation for external memories like DDR (double data rate), double DDR, QDR (quad data rate) SRAM, DDR FCRAM (fast cycle RAM), and RLDRAM (reduced latency DRAM). We utilize the memory controller in the functional block to design the BIST so that the BIST design can be simplified and executed at the functional speed. However, there are many different types of the memory controllers depending on the types of external memories, functional interface protocols, and implementation methodologies. In order to support the various memory controllers, we defined the latency of the memory controllers and classified them into three different categories: fixed latency, handshake, and both fixed latency, and handshake memory controllers. With these three models, we developed a general BIST architecture to support different types of memory controllers. During the boundary-scan driven BIST operation in the board and the system-level test and diagnosis, system clock, system hard reset, soft reset, and other programmable features were considered carefully to make the BIST operate properly. This work also presents a unique way of utilizing special BIST functions during the board and system level test, and also during the system mission operation. Heon C. Kim, Hong Shin Jun, Xinli Gu, Sung Soo Chung |
ITC | 3 |
| 2002 | Re-Using DFT Logic for Functional and Silicon Debugging TestabstractThis paper presents a technique of re-using DFT logic for system functional and silicon debugging. By re-configuring the existing DFT logic implemented on an ASIC, we are able to 1) test each part of an ASIC in a system environment separately and thus locate manufacturing defects, 2) control and observe any state elements of an ASIC to facilitate system function and silicon debugging, and 3) use structural tests to cover device and their interconnect tests on a board. Therefore, we can achieve debugging and test at both device level and system board level. Xinli Gu, Heon C. Kim, Sung Soo Chung |
ITC | 1 |
| 2001 | An effort-minimized logic BIST implementation methodabstractThis paper presents LBIST (Logic Built-In Self Test) design practice at Cisco Systems. It focuses on the LBIST design tasks that could affect design schedules and efforts. These are design timing closure and signature mismatch debugging. Our timing closure technique guarantees timing closure for LBIST insertion without any iteration between synthesis and LBIST insertion. In addition, it guarantees that only one iteration between static timing analysis and LBIST insertion is required to close all timing violations. The signature mismatch debugging technique effectively identifies the causes by indicating the pattern, the scan flip-flop and its operation mode, where the mismatch happens. These techniques save design efforts and the product-to-market time. We have integrated this method into an ASIC design flow. The results of using this flow in a large telecommunication design are described. Xinli Gu, Sung Soo Chung, Frank Tsang, Jan Arild Tofte, Hamid Rahmanian |
ITC | 1 |
| 1998 | A new approach to scan chain reordering using physical design informationabstractScan chain reordering based on physical design information helps in reducing routing bottleneck and in minimizing design constraint violations. This paper proposes integrating this capability into synthesis-based design reoptimization. It describes the benefits of such an approach, the design synthesis context, presents new ordering concepts and concludes with results on real designs. Mokhtar Hirech, James Beausang, Xinli Gu |
ITC | 3 |
| 1995 | An Efficient and Economic Partitioning Approach for TestabilityabstractThis paper presents an RT level partitioning approach for sequential circuits described as data path and control part. The data path of a circuit is partitioned at some hard-to-test points detected by an RT level testability analysis algorithm. These points are then made directly accessible by DFT techniques. The control part is also modified to control the circuit in normal mode and test mode. In the normal mode, the circuit is controlled to perform its function, while in the test mode, all partitions are controlled independently. As a result, test quality is improved by independent test generation and test application for every partition. The partitioning complexity is reduced by the use of testability analysis results and the area overhead is lower than that of full scan designs for most benchmarks we used. Experiments show results of the approach as compared with no scan, partial scan and full scan schemes. Xinli Gu, Krzysztof Kuchcinski, Zebo Peng |
ITC | 1 |
| 1995 | RT level testability-driven partitioningabstractThis paper presents a method of partitioning RT level designs based on testability analysis results. The partitioning is carried out in two steps: (1) the data path of a design is partitioned at some hard-to-test points detected by the testability analysis algorithm. These points are made directly accessible by some DFT techniques; and (2) the control part of a design is modified to operate in two modes. In normal mode, the design is controlled to fulfill the design function. In test mode, each partition is controlled independently. As a result, ATPG and test application for each partition can be done independently. In this approach, each partition is guaranteed to be acyclic, have good testability measurements and suitable size and depth for the ATPG tool to be used. When BIST technique is used, it also guarantees that these partitions are not random pattern resistant. Experiment with four benchmarks has shown the improvements on fault coverage, ATPG time and test application time after partitioning. Xinli Gu |
VTS | 1 |
| 1992 | An approach to testability analysis and improvement for VLSI systems
Xinli Gu, Krzysztof Kuchcinski, Zebo Peng |
Microprocess. Microprogramming | 1 |
| 1991 | Testability measure with reconvergent fanout analysis and its applications
Xinli Gu, Krzysztof Kuchcinski, Zebo Peng |
Microprocessing and Microprogramming | 1 |