EDBT 2026 Demo / reviewers in the wild / expert
Zhaobo Zhang
dblp:32/9056
· DBLP profile ↗
51ranked-venue papers
14as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M4EI: A Hierarchical Multimodal Memory Framework for Embodied Intelligence with Causal-Driven Retrieval
Zhaobo Zhang, Donggang Cao, Hangqi Ren |
KSEM (1) | 1 |
| 2025 | Experimental and Analytical Analysis of Turn-on Parasitic Oscillation and Dynamic Current Sharing of the Si/SiC Hybrid SwitchabstractThe Si/SiC hybrid switch (HyS), consisting of a low current rated SiC MOSFET paralleled with a high current rated IGBT, can achieve high output current at low cost and reduce switching losses caused by the tail current during the IGBT turn-off transition. When the device turns on, high-frequency ringing at the gate can induce voltage overshoots and current spikes, thereby increasing the device’s current stress. The high-frequency oscillation occurs between the transistor’s capacitance and the board’s parasitic inductance. Such high-frequency oscillations also increase electromagnetic interference (EMI), which can disrupt the performance of the parallel hybrid switch. When the load current is concentrated in the MOSFET, the generated heat accelerates its degradation. Reducing the transition time needed to achieve current balancing can enhance system reliability. In this paper, a detailed turn-on analytical model has been proposed to analyze the effect of parasitic inductances and capacitances, and the current dynamic sharing transition speed to offer guidelines for Si/SiC hybrid power module design. The symmetry of the power loop inductance is also analyzed in the model, providing a more systematic and comprehensive analysis of the turn-on characteristics of the hybrid switch compared to existing models. The validity of the model is verified through simulation and experiments. Mohan Zhang, Ian Laird, Saeed Jahdi, Wenzhi Zhou, Zhaobo Zhang |
IECON | 5 |
| 2025 | SymmPi: Exploiting Symmetry Removal for Fast Subgraph MatchingabstractAbstract Symmetry, a phenomenon of self-similarity, is common in many networks, which often incurs a lot of redundant accesses and computations, even duplicate results when executing graph matching tasks. Many approaches (e.g. symmetry-breaking methods) try to disrupt symmetry by translating symmetry into restrictions and then imposing restrictions on the exploration order. However, the restrictions are finer-grained. If the pattern graph is complex, more restrictions are generated from symmetry breaking methods, thus complicating the exploration process and degrading the performance. Here, we present novel SymmPi, which exploits symmetry removal for fast graph matching. SymmPi first identifies the coarse-grained axisymmetric subgraphs of the given pattern graphs instead of finer relationships. If a pattern graph is not axisymmetric, SymmPi will remove some of its edges until axisymmetric subgraphs are found. Thus, the original pattern graph is transformed to a set of axisymmetric subgraphs plus some edges. Then, SymmPi finds the matches of the axisymmetric subgraph and extends these matches to the original pattern graphs by permuting the matches with additional checks. Our experiments on both directed and undirected graphs, demonstrate that SymmPi achieves a significant performance improvement over the state-of-the-art undirected and directed graph matching methods and systems. Yujiang Wang 0007, Zhaobo Zhang, Pingpeng Yuan, Hai Jin 0001 |
Data Sci. Eng. | 3 |
| 2024 | Correcting Pronoun Homophones with Subtle Semantics in Chinese Speech RecognitionabstractSpeech recognition is becoming prevalent in daily life. However, due to the similar semantic context of the entities and the overlap of Chinese pronunciation, the pronoun homophone, especially “他/她/它 (he/she/it)”, (their pronunciation is “Tā”) is usually recognized incorrectly. It poses a challenge to automatically correct them during the post-processing of Chinese speech recognition. In this paper, we propose three models to address the common confusion issues in this domain, tailored to various application scenarios. We implement the language model, the LSTM model with semantic features, and the rule-based assisted Ngram model, enabling our models to adapt to a wide range of requirements, from high-precision to low-resource offline devices. The extensive experiments show that our models achieve the highest recognition rate for “Tā” correction with improvements from 70% in the popular voice input methods up to 90%. Further ablation analysis underscores the effectiveness of our models in enhancing recognition accuracy. Therefore, our models improve the overall experience of Chinese speech recognition of “Tā” and reduce the burden of manual transcription corrections. Zhaobo Zhang, Rui Gan, Pingpeng Yuan, Hai Jin 0001 |
LREC/COLING | 1 |
| 2024 | Evaluation of Machine Neutral Point Overvoltage Mitigation in SiC Motor Drives using Three Passive Filter DesignsabstractIn SiC inverter-fed motor drives, fast-switched power devices escalate the reflected wave phenomenon, causing overvoltage oscillations at the motor terminals with shorter cable lengths between the inverter and motor. Likewise, overvoltage oscillations manifest at motor neutral point due to the reflection of inverter common-mode voltage through the machine windings. Despite extensive research on motor terminal overvoltage mitigation, the motor neutral point overvoltage remains insufficiently addressed in the literature, although it can have a more detrimental impact on the machine winding insulation especially if the high switching frequency of SiC inverters aligns with the machine antiresonance frequency. This paper evaluates the effectiveness of three passive filter designs (R, RC, and C filter) connected between the motor neutral point and grounded frame for neutral point overvoltage mitigation. The paper provides mathematical design guidelines for the parameter selection of the three filters while assesses their performance in mitigating the motor neutral point overvoltage. An experimental evaluation using a 2.2kW SiC-based motor drive setup is presented, demonstrating the effectiveness of the three filter designs to eliminate the motor neutral point overvoltage while highlighting the pros and cons of each filter design. Mustafa Memon, Mohamed S. Diab, Xibo Yuan, Zhaobo Zhang |
IECON | 4 |
| 2023 | Demystifying deep learning in predictive monitoring for cloud-native SLOsabstractThe complexity inherent in managing cloud computing systems calls for novel solutions that can effectively enforce high-level Service Level Objectives (SLOs) promptly. Unfortunately, most of the current SLO management solutions rely on reactive approaches, i.e., correcting SLO violations only after they have occurred. Further, the few methods that explore predictive techniques to prevent SLO violations focus solely on forecasting low-level system metrics, such as CPU and Memory utilization. Although valid in some cases, these metrics do not necessarily provide clear and actionable insights into application behavior. This paper presents a novel approach that directly predicts high-level SLOs using low-level system metrics. We target this goal by training and optimizing two state-of-the-art neural network models, a Short-Term Long Memory - LSTM, and a Transformer-based model. Our models provide actionable insights into application behavior by establishing proper connections between the evolution of low-level workload-related metrics and the high-level SLOs. We demonstrate our approach to selecting and preparing the data. We show in practice how to optimize LSTM and Transformer by targeting efficiency as a high-level SLO metric and performing a comparative analysis. We show how these models behave when the input workloads come from different distributions. Consequently, we demonstrate their ability to generalize in heterogeneous systems. Finally, we operationalize our two models by integrating them into the Polaris framework we have been developing to enable a performance-driven SLO-native approach to Cloud computing. Andrea Morichetta 0002, Thomas W. Pusztai, Deepak Vij, Víctor Casamayor-Pujol, Philipp Raith, Stefan Nastic, Schahram Dustdar, Zhaobo Zhang |
CLOUD | 9 |
| 2023 | Exploring Word-Sememe Graph-Centric Chinese Antonym Detection
Zhaobo Zhang, Pingpeng Yuan, Hai Jin 0001 |
ECML/PKDD (3) | 1 |
| 2023 | Improving Entity Linking in Chinese Domain by Sense Embedding Based on Graph Clustering
Zhaobo Zhang, Zhi-Man Zhong, Pingpeng Yuan, Hai Jin 0001 |
J. Comput. Sci. Technol. | 1 |
| 2022 | Learning Chinese Word Embeddings By Discovering Inherent Semantic Relevance in Sub-charactersabstractLearning Chinese word embeddings is important in many tasks of Chinese language information processing, such as entity linking, entity extraction, and knowledge graph. A Chinese word consists of Chinese characters, which can be decomposed into sub-characters (radical, component, stroke, etc). Similar to roots in English words, sub-characters also indicate the origins and basic semantics of Chinese characters. So, many researches follow the approaches designed for learning embeddings of English words to improve Chinese word embeddings. However, some Chinese characters sharing the same sub-characters have different meanings. Furthermore, with more cultural interaction and the popularization of the Internet and web, many neologisms, such as transliterated loanwords and network terms, are emerging, which are only close to the pronunciation of their characters, but far from their semantics. Here, a tripartite weighted graph is proposed to model the semantic relationship among words, characters, and sub-characters, in which the semantic relationship is evaluated according to the Chinese linguistic information. So, the semantic relevance hidden in lower components (sub-characters, characters) can be used to further distinguish the semantics of corresponding higher components (characters, words). Then, the tripartite weighted graph is fed into our Chinese word embedding modelinsideCC to reveal the semantic relationship among different language components, and learn the embeddings of words. Extensive experimental results on multiple corpora and datasets verify that our proposed methods outperform the state-of-the-art counterparts by a significant margin. Zhaobo Zhang, Pingpeng Yuan, Hai Jin 0001, Qiang-Sheng Hua |
CIKM | 2 |
| 2022 | Improving Chinese Word Representation Using Four Corners FeaturesabstractIntuitively, word representation for logographic languages like Chinese can be enhanced by its internal characteristics. Several research endeavors tried to learn Chinese word embeddings with characters, radicals, or subcharacters containing rich semantic information. In this paper, motivated by Four-Corner Method for Character Indexation, we extract features from four corners of characters with important morphological charactertics. Based on the features from four corners, we propose a model to utilize characters and four corner features of words to capture both semantic and morphological information. Moreover, we apply an attention scheme to integrate internal information dynamically, which includes two strategies to assign different weights for elements according to the word frequency. Experimental results on social news corpus and Chinese Wikipedia Dump show exploiting the four corner morphological features is crucial for capturing the meanings of Chinese words. Meanwhile, the results on word analogy, word similarity, and text classification tasks demonstrate that our approach obtains better results than state-of-the-art approaches. Hai Jin 0001, Zhaobo Zhang, Pingpeng Yuan |
IEEE Trans. Big Data | 2 |
| 2022 | Unsupervised Two-Stage Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems have placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligence and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. We propose a two-stage unsupervised root-cause-analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to cluster the data in a coarse-grained manner. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. The proposed method can accommodate both numerical and categorical test items. A combination of the L-method, cross validation, and Silhouette score enables us to automatically determine all hyperparameters. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause-analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Black-Box Test-Cost Reduction Based on Bayesian Network ModelsabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. In this article, we propose a novel black-box test selection method based on Bayesian networks (BNs), which extract the strong relationship among tests. First, the problem of reducing the black-box test cost is formulated as a constrained optimization problem. Next, multiple structure learning and transfer learning algorithms are implemented to construct BN models. Based on these BN models, we propose an iterative test selection method with a new metric, Bayesian index, for test-cost reduction. In addition, averaging strategies are applied to enhance the reduction performance. Finally, a robust model selection framework is proposed to select the optimal BN model for test-cost reduction. Two case studies with production test data demonstrate that when no prior information is provided, our proposed approach effectively reduces the test cost by up to 14.7%, compared to the state-of-the-art greedy algorithm. Moreover, our proposed approach further reduces the test cost by up to 7.1% when prior information is provided from similar products. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Adaptive Addresses for Next Generation IP Protocol in Hierarchical NetworksabstractWe propose the adaptive addresses under a hierarchical network structure, which can be realized in a newer generation of IP protocol (i.e., IPvn). It minimizes the communication overhead, enables arbitrary address space extension, simplifies both the network data-plane and control-plane, and supports better network security. More importantly, it supports incremental deployment from the network edge and gradual growth towards the core. A clear boundary between IPvn domain and the existing IPv4/IPv6 networks enables transparent cross-domain communication. The clear evolution path makes pre-standard deployment possible. We design both control plane and data plane, prototype the routers within and on the edge of an IPvn domain, and evaluate the performance. We open source the project to encourage further investigation and development. Haoyu Song 0001, Zhaobo Zhang, Yingzhen Qu, James N. Guichard |
ICNP | 2 |
| 2020 | Unsupervised Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems has placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligent and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. In this paper, we propose a two-stage unsupervised root-cause analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to roughly cluster the data. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. In additional, L-method and cross validation are applied to automatically determine the hyper-parameters of our algorithm. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2020 | Hierarchical Symbol-Based Health-Status Analysis Using Time-Series Data in a Core Router SystemabstractTo ensure high reliability and rapid error recovery in commercial core router systems, a health-status analyzer is essential to monitor the different features of core routers. However, traditional health analyzers need to store a large amount of historical data in order to identify health status. The storage requirement becomes prohibitively high when we attempt to carry out long-term health-status analysis for a large number of core routers. We describe the design of a symbol-based health status analyzer that first encodes, as a symbol sequence, the long-term complex time series collected from a number of core routers, and then utilizes the symbol sequence to do health analysis. The symbolic aggregation approximation (SAX), 1d-SAX, moving-average-based trend approximation, and nonparametric symbolic approximation representation methods are implemented to encode complex time series in a hierarchical way. Hierarchical agglomerative clustering and sequitur rule discovery are implemented to learn important global and local patterns. Three classification methods including a vector-space-model-based approach are then utilized to identify the health status of core routers. Data collected from a set of commercial core router systems are used to validate the proposed health-status analyzer. The experimental results show that our symbol-based health status analyzer requires much lower storage than traditional methods, but can still maintain comparable diagnosis accuracy. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Self-Learning and Efficient Health-Status Analysis for a Core Router SystemabstractThe health status of core router systems needs to be analyzed efficiently in order to ensure high reliability and timely error recovery. Although a large amount operational data is collected from core routers, due to high computational complexity and expensive labor cost, only a small part of this data is labeled by experts. The lack of labels is an impediment toward the adoption of supervised learning. We present an iterative self-learning procedure for assessing the health status of a core router. This procedure first computes a representative feature matrix to capture different characteristics of time-series data. Not only statistical-modeling-based features are computed from three general categories but also a recurrent neural network-based autoencoder is utilized to capture a wider range of hidden patterns. Moreover, both minimum-redundancy-maximum-relevance (mRMR) method and fully connected feedforward autoencoder are applied to further reduce dimensionality of extracted feature matrix. Hierarchical clustering is then utilized to infer labels for the unlabeled dataset. Finally, a classifier is built and iteratively updated using both labeled and unlabeled dataset. Field data collected from a set of commercial core routers are used to experimentally validate the proposed health-status analyzer. The experimental results show that the proposed feature-based self-learning health analyzer achieves higher precision and recall than the traditional supervised health analyzer as well the currently deployed rule-based health analyzer. Moreover, it achieves better performance than the three anomaly detection baseline methods under the transformed binary classification scenario. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Black-Box Test-Coverage Analysis and Test-Cost Reduction Based on a Bayesian Network ModelabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. The conventional greedy algorithm, which selects the most important tests by considering both strong and weak relationships among tests, suffers from overfitting. In order to overcome overfitting, we propose a novel black-box test selection method based on a Bayesian network model. First, the problem of reducing black-box test cost is formulated as a constrained optimization problem. Next, a score-based algorithm is implemented to construct the Bayesian network for black-box tests. Finally, we propose a Bayesian index with the property of Markov blankets, and then an iterative test selection method is developed based on our proposed Bayesian index. The proposed approach ensures that only the strong relationships among black-box tests are used for test selection so that this approach is more robust to overfitting. Two case studies with production test data demonstrate that the proposed approach effectively reduces test cost by up to 14.7%, compared to a conventional greedy algorithm. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 2 |
| 2019 | Changepoint-Based Anomaly Detection for Prognostic Diagnosis in a Core Router SystemabstractPrognostic diagnosis is desirable for commercial core router systems to ensure early failure prediction and fast error recovery. The effectiveness of prognostic diagnosis depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the statistical properties of the monitored data change significantly as time proceeds. We describe the design of a changepoint (CP)-based anomaly detector that first detects CPs from collected time-series data, and then utilizes these CPs to detect anomalies. Different CP detection approaches are implemented to detect various types of CPs. A clustering method is then developed to identify normal/abnormal patterns from CP windows. Data collected from a set of commercial core router systems are used to validate the proposed anomaly detector. Experimental results show that our CP-based anomaly detector achieves better performance than traditional methods in terms of two metrics, namely success ratio and nonfalse-alarm ratio. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Failure prediction based on anomaly detection for complex core routersabstractData-driven prognostic health management is essential to ensure high reliability and rapid error recovery in commercial core router systems. The effectiveness of prognostic health management depends on whether failures can be accurately predicted with sufficient lead time. This paper describes how time-series analysis and machine-learning techniques can be used to detect anomalies and predict failures in complex core router systems. First both a feature-categorization-based hybrid method and a changepoint-based method have been developed to detect anomalies in time-varying features with different statistical characteristics. Next, a SVM-based failure predictor is developed to predict both categories and lead time of system failures from collected anomalies. A comprehensive set of experimental results is presented for data collected during 30 days of field operation from over 20 core routers deployed by customers of a major telecom company. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ICCAD | 2 |
| 2018 | Self-Learning Health-Status Analysis for a Core Router SystemabstractThe health status of core router systems needs to be analyzed efficiently in order to ensure high reliability and timely error recovery. Although a large amount operational data is collected from core routers, only a small part of this data is labeled by experts. The lack of labels is an impediment towards the adoption of supervised learning. We present an iterative self-learning procedure for assessing the health status of a core router. This procedure first computes a representative feature matrix to capture different characteristics of time-series data. Hierarchical clustering is then utilized to infer labels for the unlabeled dataset. Finally, a classifier is built and iteratively updated using both labeled and unlabeled dataset. Field data collected from a set of commercial core routers are used to experimentally validate the proposed health-status analyzer. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2018 | Innovative practices on machine learning for emerging applicationsabstractThe IP session focuses on using Machine Learning (ML) techniques on several emerging applications. The first contribution discusses hotspot detection by using ML. The second presentation then talks a data-driven health monitoring solution. The last contribution discusses using ML to emulate hardware Trojans. Kareem Madkour, Zhaobo Zhang, Alfred L. Crouch, Peter L. Levin, Eve Hunter, Yu Huang 0005 |
VTS | 2 |
| 2018 | Toward Predictive Fault Tolerance in a Core-Router System: Anomaly Detection Using Correlation-Based Time-Series AnalysisabstractFault tolerance is used in communication systems to ensure high reliability and rapid error recovery. The effectiveness of most proactive fault-tolerant mechanism depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the monitored data involves temporal measurements and exhibits significantly different statistical characteristics for its constituent features. We describe the design of an anomaly detector that monitors the time-series data of a complex core router system. Anomaly detection techniques are compared in terms of their effectiveness for detecting different types of anomalies. A feature-categorizing-based hybrid method is proposed to overcome the difficulty of detecting anomalies in features with different statistical characteristics. Furthermore, a correlation analyzer is implemented to remove irrelevant and redundant features. Three types of synthetic anomalies, generated using a small amount of real data for a commercial telecom system, are used to validate the proposed anomaly detector. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | RetroDMR: Troubleshooting non-deterministic faults with retrospective DMRabstractThe most notorious faults for diagnosis in post-silicon validation are those that manifest themselves in a non-deterministic manner with system-level functional tests, where errors randomly appear from time to time even when applying the same workloads. In this work, we propose a novel diagnostic framework that resorts to dual-modular redundancy (DMR) for troubleshooting non-deterministic faults, namely RetroDMR. To be specific, we log the essential events (e.g., the sequence of thread migration) in the faulty run to record the mapping relationship between threads and their corresponding execution units. Then in the following diagnosis runs, we apply redundant multithreading (RMT) technique to reduce error detection latency, while at the same time we try to follow the thread migration sequence of the original run whenever possible. By doing so, RetroDMR significantly improves the reproduction rate and diagnosis resolution for non-deterministic faults, as demonstrated in our experimental results. Ting Wang 0008, Yannan Liu, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
DATE | 4 |
| 2017 | Data-driven fault diagnosis with missing syndromes imputation for functional test through conditional specificationabstractIn the electronic system manufacturing process, the board-level functional test is recognized as the most significant step to prevent defective products from entering the market. In recent years, machine learning and data mining have proven to be efficient techniques in determining root cause from the problematic functional test result, especially when the integrated circuits (IC) are becoming increasingly highly-integrated. However, the test results are sometimes unavailable due to either abnormal ending of the test sequence or occasional system failures, which results in a decreased performance of data-driven diagnosis systems. In this paper, we propose a data imputation algorithm to predict the missing entries in the functional test result, by considering the correlation between test items with conditional specification. We evaluate our data imputation algorithm over the test results collected from three different stages of functional test on a line card used in the telecommunication system. The result shows that our proposed data imputation algorithm consistently outperforms other imputation techniques with various data-driven approaches in terms of diagnosing the root cause, increasing the diagnosis accuracy by an average of 28.13% compared to none data imputation, and 9.74% compared to the naive pass imputation. Tong Guan, Zhaobo Zhang, Wen Dong 0001, Chunming Qiao, Xinli Gu |
ETS | 2 |
| 2017 | Changepoint-based anomaly detection in a core router systemabstractPrognostic diagnosis is desirable for commercial core router systems to ensure early failure prediction and fast error recovery. The effectiveness of prognostic diagnosis depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the statistical properties of the monitored data change significantly as time proceeds. We describe the design of a changepoint-based anomaly detector that first detects changepoints from collected time-series data, and then utilizes these changepoints to detect anomalies. Two approaches based on maximum-likelihood estimation are implemented to detect different types of changepoints. A clustering method is then developed to identify a wide range of normal/abnormal patterns from changepoint windows. Data collected from a set of commercial core router systems are used to validate the proposed anomaly detector. Experimental results show that our changepoint-based anomaly detector achieves better performance than traditional methods. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2017 | Symbol-based health-status analysis in a core router systemabstractTo ensure high reliability and rapid error recovery in commercial core router systems, a health-status analyzer is essential to monitor the different features of core routers. However, traditional health analyzers need to store a large amount of historical data in order to identify health status. The storage requirement becomes prohibitively high when we attempt to carry out long-term health-status analysis for a large number of core routers. We describe the design of a symbol-based health status analyzer that first encodes, as a symbol sequence, the long-term complex time series collected from a number of core routers, and then utilizes the symbol sequence to do health analysis. The symbolic aggregation approximation (SAX) and moving-average-based trend approximation methods are implemented to encode complex time series in a hierarchical way. Hierarchical agglomerative clustering and sequitur rule discovery are implemented to learn important global and local patterns. Two classification methods are then utilized to identify the health status of core routers. Data collected from a set of commercial core router systems are used to validate the proposed health-status analyzer. The experimental results show that our symbol-based health status analyzer requires much lower storage than traditional methods, but can still maintain comparable diagnosis accuracy. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2016 | Accurate anomaly detection using correlation-based time-series analysis in a core router systemabstractFault tolerance is used in communication systems to ensure high reliability and rapid error recovery. The effectiveness of most proactive fault-tolerant mechanism depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the monitored data involves temporal measurements and exhibits significantly different statistical characteristics for its constituent features. We describe the design of an anomaly detector that monitors the time-series data of a complex core router system. Anomaly detection techniques are compared in terms of their effectiveness for detecting different types of anomalies. A feature-categorizing-based hybrid method is proposed to overcome the difficulty of detecting anomalies in features with different statistical characteristics. Furthermore, a correlation analyzer is implemented to remove irrelevant and redundant features. Three types of synthetic anomalies, generated using a small amount of real data for a commercial telecom system, are used to validate the proposed anomaly detector. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2016 | Efficient Board-Level Functional Fault Diagnosis With Missing SyndromesabstractFunctional fault diagnosis is widely used in board manufacturing to ensure product quality and improve product yield. Advanced machine-learning techniques have recently been advocated for reasoning-based diagnosis; these techniques are based on the historical record of successfully repaired boards. However, traditional diagnosis systems fail to provide appropriate repair suggestions when the diagnostic logs are fragmented and some error outcomes, or syndromes, are not available during diagnosis. We describe the design of a diagnosis system that can handle missing syndromes and can be applied to four widely used machine-learning techniques. Several imputation methods are discussed and compared in terms of their effectiveness for addressing missing syndromes. Moreover, a syndrome-selection technique based on the minimum-redundancy-maximum-relevance criteria is also incorporated to further improve the efficiency of the proposed methods. Two large-scale synthetic data sets generated from the log information of complex industrial boards in volume production are used to validate the proposed diagnosis system in terms of diagnosis accuracy and training time. Shi Jin 0001, Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Adaptive Board-Level Functional Fault Diagnosis Using Incremental Decision TreesabstractBoard-level functional fault diagnosis is needed for high-volume production to improve product yield. However, to ensure diagnosis accuracy and effective board repair, a large number of syndromes must be used. Therefore, the diagnosis cost can be prohibitively high due to the increase in diagnosis time and the complexity of test execution and analysis. We propose an adaptive diagnosis method based on incremental decision trees (DTs). Faulty components are classified according to the discriminative ability of the syndromes in DT training. The diagnosis procedure is constructed as a binary tree, with the most discriminative syndrome as the root and final repair suggestions are available as the leaf nodes of the tree. The syndrome to be used in the next step is determined based on the observation of syndromes thus far in the diagnosis procedure. The number of syndromes required for diagnosis can be significantly reduced compared to the total number of syndromes used for system training. Moreover, online learning is facilitated in the proposed diagnosis system using an incremental version of DTs, so as to bridge the knowledge obtained at test-design stage with the knowledge gained during volume production. The diagnosis system can thus adapt to occurrences of new error scenarios on-the-fly. Diagnosis results for three complex boards from industry, currently in volume production, highlight the effectiveness of the proposed approach. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | On test syndrome merging for reasoning-based board-level functional fault diagnosisabstractMachine learning algorithms are advocated for automated diagnosis of board-level functional failures due to the extreme complexity of the problem. Such reasoning-based solutions, however, remain ineffective at the early stage of the product cycle, simply because there are insufficient historical data for training the diagnostic system that has a large number of test syndromes. In this paper, we present a novel test syndrome merging methodology to tackle this problem. That is, by leveraging the domain knowledge of the diagnostic tests and the board structural information, we adaptively reduce the feature size of the diagnostic system by selectively merging test syndromes such that it can effectively utilize the available training cases. Experimental results demonstrate the effectiveness of the proposed solution. Zelong Sun, Li Jiang 0002, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
ASP-DAC | 4 |
| 2015 | Self-learning and adaptive board-level functional fault diagnosisabstractFunctional fault diagnosis is necessary for board-level product qualification. However, ambiguous diagnosis results can lead to long debug times and wrong repair actions, which significantly increase repair cost and adversely impact yield. A state-of-the-art functional fault diagnosis system involves several key components: (1) design of functional test programs, (2) collection of functional-failure syndromes, (3) building of the diagnosis engine, (4) isolation of root causes, and (5) evaluation of the diagnosis engine. Advances in each of these components can pave the way for a more effective diagnosis system, thus improving diagnosis accuracy and reducing diagnosis time. Machine-learning and data analysis techniques offer an unprecedented opportunity to develop an automated and adaptive diagnosis system to increase diagnosis accuracy and reduce diagnosis time. This paper describes how all the above components of an advanced diagnosis system can benefit from machine learning and information theory. Topics discussed include incremental learning, decision trees, root-cause analysis and evaluation metrics, data acquisition, and knowledge transfer. Fangming Ye, Krishnendu Chakrabarty, Zhaobo Zhang, Xinli Gu |
ASP-DAC | 3 |
| 2015 | Information-Theoretic Syndrome Evaluation, Statistical Root-Cause Analysis, and Correlation-Based Feature Selection for Guiding Board-Level Fault DiagnosisabstractReasoning-based functional-fault diagnosis has recently been advocated to achieve high diagnosis accuracy, low defect escapes, and reducing manufacturing cost. However, such diagnosis method requires a rich set of test items (syndromes) and a sizable database of faulty boards to learn from. An insufficient number of failed boards, ambiguous root-cause identification, and redundant or irrelevant syndromes can render reasoning-based diagnosis ineffective. Periodic evaluation and analysis can help locate weaknesses in a diagnosis system and thereby provide guidelines for redesigning the tests, which facilitates better diagnosis. We propose an information-theoretic framework for evaluating the effectiveness of and providing guidance to a reasoning-based functional-fault diagnosis system. Syndrome analysis based on feature selection methods provides a representative set of syndromes and suggests irrelevant syndromes in diagnosis. Root-cause analysis measures the discriminative ability of differentiating a given root cause from others. Results are presented for four types of diagnosis systems for three complex boards that are in volume production. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Knowledge discovery and knowledge transfer in board-level functional fault diagnosisabstractDiagnosis of functional failures at the board level is critical for improving product yield and reducing manufacturing cost. Reasoning techniques increase the accuracy of functional-fault diagnosis based on the history of successfully repaired boards. However, depending on the complexity of the product, it usually takes several months to accumulate an adequate database for training a reasoning-based diagnosis system. During the initial product ramp-up phase, reasoning-based diagnosis is not feasible for yield learning, since the required database is not available due to lack of volume. We propose a knowledge-discovery method and a knowledge-transfer method for facilitating board-level functional fault diagnosis. First, an analysis technique based on machine learning is used to discover knowledge from syndromes, which can be used for training a diagnosis engine. Second, knowledge from diagnosis engines used for earlier-generation products can be automatically transferred through root-cause mapping and syndrome mapping based on keywords and board-structure similarities. Two complex boards in volume production and with a mature diagnosis system, and three new boards in the ramp-up phase, are used to validate the proposed knowledge-discovery and knowledge-transfer approach in terms of the diagnosis accuracy obtained using the new diagnosis systems. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 2 |
| 2014 | Built-In Self-Test, Diagnosis, and Repair of MultiMode Power SwitchesabstractRecently proposed power-gating structures for intermediate power-off modes offer significant power saving benefits as they reduce the leakage power during short periods of inactivity. Even though they are very effective for reducing static power consumption, their reliable operation can be compromised by process variations and manufacturing defects. In this paper, we propose a signature analysis technique to efficiently test power-gating structures that provide intermediate power-off modes. Based on this technique, a methodology to repair catastrophic and parametric faults, and to tolerate process variations is presented. For testing and repairing multimode power switches, we propose a robust built-in self-test and built-in self-repair scheme that reduces test cost and obviates additional manufacturing steps for post-silicon repair. Simulation results highlight the low-cost and effectiveness of the proposed method for detecting, diagnosing, and repairing defects. Ran Wang 0002, Zhaobo Zhang, Xrysovalantis Kavousianos, Yiorgos Tsiatouhas, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Board-Level Functional Fault Diagnosis Using Multikernel Support Vector Machines and Incremental LearningabstractAdvanced machine learning techniques offer an unprecedented opportunity to increase the accuracy of board-level functional fault diagnosis and reduce product cost through successful repair. Ambiguous or incorrect diagnosis results lead to long debug times and even wrong repair actions, which significantly increase repair cost. We propose a smart diagnosis method based on multikernel support vector machines (MK-SVMs) and incremental learning. The MK-SVM method leverages a linear combination of single kernels to achieve accurate faulty-component classification based on the errors observed. The MK-SVMs thus generated can also be updated based on incremental learning, which allows the diagnosis system to quickly adapt to new error observations and provide even more accurate fault diagnosis. Two complex boards from industry, currently in volume production, are used to validate the proposed diagnosis approach in terms of diagnosis accuracy (success rate) and quantifiable improvements over previously proposed machine-learning methods based on several single-kernel SVMs and artificial neural networks. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Static Power Reduction Using Variation-Tolerant and Reconfigurable Multi-Mode Power SwitchesabstractMultithreshold CMOS is very effective for reducing standby leakage power during long periods of inactivity. Recently, a power-gating scheme was presented to support multiple power-off modes and reduce the leakage power during short periods of inactivity. However, this scheme can suffer from high sensitivity to process variations, which impedes manufacturability. We propose a new power-gating technique that is tolerant to process variations and scalable to more than two intermediate power-off modes. The proposed design requires less design effort and offers greater power reduction and smaller area cost than the previous method. In addition, it can be combined with existing techniques to offer further static power reduction benefits. Analysis and extensive simulation results demonstrate the effectiveness of the proposed design. Zhaobo Zhang, Xrysovalantis Kavousianos, Krishnendu Chakrabarty, Yiorgos Tsiatouhas |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Handling Missing Syndromes in Board-Level Functional-Fault DiagnosisabstractFunctional fault diagnosis is widely used in board manufacturing to ensure product quality and improve product yield. Advanced machine-learning techniques have recently been advocated for reasoning-based diagnosis, these technologies are based on historical data of successfully repaired boards. However, traditional diagnosis systems fail to provide appropriate repair suggestions when the diagnostic logs are fragmented and some error outcomes, or syndromes, are not available during diagnosis. We describe the design of a diagnosis system, based on support vector machines, that can handle missing syndromes by using the method of imputation. Several imputation methods are discussed and compared in terms of their efficiency in handling missing syndromes. Two large-scale synthetic data sets generated from the log information of complex industrial boards in volume production are used to validate the proposed diagnosis system in terms of diagnosis accuracy and training time. Fangming Ye, Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
Asian Test Symposium | 3 |
| 2013 | Information-theoretic syndrome and root-cause analysis for guiding board-level fault diagnosisabstractHigh-volume manufacturing of complex electronic products involves functional test at board level to ensure low defect escapes. Machine-learning techniques have recently been proposed for reasoning-based functional-fault diagnosis system to achieve high diagnosis accuracy. However, machine learning requires a rich set of test items (syndromes) and a sizable database of faulty boards. An insufficient number of failed boards, ambiguous root-cause identification, and redundant or irrelevant syndromes can render machine learning ineffective. We propose an evaluation and enhancement framework based on information theory for guiding diagnosis systems using syndrome and root-cause analysis. Syndrome analysis based on subset selection provides a representative set of syndromes with minimum redundancy and maximum relevance. Root-cause analysis measures the discriminative ability of differentiating a given root cause from others. The metrics obtained from the proposed framework can also provide guidelines for test redesign to enhance diagnosis. A real board from industry, currently in volume production, and an additional synthetic board, based on data extrapolated from another real board, are used to demonstrate the effectiveness of the proposed framework. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ETS | 2 |
| 2013 | AgentDiag: An agent-assisted diagnostic framework for board-level functional failuresabstractDiagnosing functional failures in complicated electronic boards is a challenging task, wherein debug technicians try to identify defective components by analyzing some syndromes obtained from the application of diagnostic tests. The diagnosis effectiveness and efficiency rely heavily on the quality of the in-house developed diagnostic tests and the debug technicians' knowledge and experience, which, however, have no guarantees nowadays. To tackle this problem, we propose a novel agent-assisted diagnostic framework for board-level functional failures, namely AgentDiag, which facilitates to evaluate the quality of the diagnostic tests and bridge the knowledge gap between the diagnostic programmers who write diagnostic tests and the debug technicians who conduct in-field diagnosis with a lightweight model of the boards and tests. Experimental results on a real industrial board and an OpenRISC design demonstrate the effectiveness of the proposed solution. Zelong Sun, Li Jiang 0002, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
ITC | 4 |
| 2013 | Board-Level Functional Fault Diagnosis Using Artificial Neural Networks, Support-Vector Machines, and Weighted-Majority VotingabstractIncreasing integration densities and high operating speeds lead to subtle manifestation of defects at the board level. Functional fault diagnosis is, therefore, necessary for board-level product qualification. However, ambiguous diagnosis results lead to long debug times and even wrong repair actions, which significantly increase repair cost and adversely impact yield. Advanced machine-learning (ML) techniques offer an unprecedented opportunity to increase the accuracy of board-level functional diagnosis and reduce high-volume manufacturing cost through successful repair. We propose a smart diagnosis method based on two ML classification models, namely, artificial neural networks (ANNs) and support-vector machines (SVMs) that can learn from repair history and accurately localize the root cause of a failure. Fine-grained fault syndromes extracted from failure logs and corresponding repair actions are used to train the classification models. We also propose a decision machine based on weighted-majority voting, which combines the benefits of ANNs and SVMs. Three complex boards from the industry, currently in volume production, and additional synthetic data, are used to validate the proposed methods in terms of diagnostic accuracy, resolution, and quantifiable improvement over current diagnostic software. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Adaptive Board-Level Functional Fault Diagnosis Using Decision TreesabstractFunctional fault diagnosis at board-level is desirable for high-volume production since it improves product yield. However, to ensure diagnosis accuracy and effective board repair, a large number of syndromes must be used. Therefore, the diagnosis cost can be prohibitively high due to the increase in diagnosis time and the complexity of syndrome collection/analysis. We propose an adaptive diagnosis method based on decision trees (DTs). Faulty components are classified according to the discriminative ability of the syndromes in DT training. The diagnosis procedure is constructed as a binary tree, with the most discriminative syndrome as the root and final repair suggestions are available as the leaf nodes of the tree. The syndrome to be collected in the next step is determined based on the observations of syndromes collected thus far in the diagnosis procedure. The number of syndromes required for diagnosis can also be significantly reduced compared to the number of syndromes used for system training. Diagnosis results for two complex boards from industry, currently in volume production, and additional synthetic data highlight the effectiveness of the proposed approach. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
Asian Test Symposium | 2 |
| 2012 | Board-Level Functional Fault Diagnosis Using Learning Based on Incremental Support-Vector MachinesabstractAdvanced machine learning techniques offer an unprecedented opportunity to increase the accuracy of board-level functional fault diagnosis based on the historical data of successfully repaired boards. However, the training complexity increases significantly in diagnosis systems due to the increasing amount of the historical data. We propose a smart learning method in the diagnosis system using incremental support-vector machines (SVMs). The SVMs updated using incremental learning allow the diagnosis system to quickly adapt to new error observations and provide more accurate fault diagnosis. Two sets of large-scale synthetic data generated from the log information of two complex industrial boards, in volume production, are used to validate the proposed diagnosis approach in terms of training time and diagnosis accuracy over a previously proposed diagnosis system based on simple support-vector machines. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
Asian Test Symposium | 2 |
| 2012 | Diagnostic system based on support-vector machines for board-level functional diagnosisabstractFault diagnosis is critical for improving product yield and reducing manufacturing cost. However, it is very challenging to identify the root cause of failures on a complex circuit board. Ambiguous diagnosis results lead to long debug times and even wrong repair actions, which significantly increases the repair cost. We propose an automatic diagnostic system using support vector machines (SVMs). The proposed system acquires debug knowledge from empirical data; this strategy avoids the difficulties involved in knowledge acquisition in traditional fault diagnosis methods. SVMs provide an optimal separating hyperplane in classification. The optimal solution and generalization ability of SVMs lead to higher diagnostic accuracy, compared to the classical learning approaches such as artificial neural networks (ANNs). An industrial board is used to validate the effectiveness of the proposed system. Extensive simulation results demonstrate that the SVMs-based diagnostic system provides quantifiable improvement over current diagnostic software and an ANN-based diagnostic system. Zhaobo Zhang, Xinli Gu, Yaohui Xie, Zhanglei Wang, Krishnendu Chakrabarty |
ETS | 1 |
| 2012 | Physical-Defect Modeling and Optimization for Fault-Insertion TestabstractHardware fault insertion is a promising method for system reliability assessment and fault isolation. It provides feedback on the fault tolerance of a large system, creates artificial faulty scenarios that can be used as reference points for fault diagnosis, and leads to a quality diagnostic program. Optimization of fault insertion location is critical for accelerating the assessment of system reliability and constructing a complete knowledge base for fault diagnosis. In this work, we construct a pin-level fault model that is able to effectively mimic the errors (effects) caused by physical defects within the component. A simulation framework and optimization techniques are proposed to select a minimum subset of output pins that can represent as many physical defects as possible. The optimization results provide guidelines on the fault insertion locations and the appropriate fault types for insertion. In addition, three intrinsic characteristics of output pins, including testability number, fan-in size, and transition counts, are analyzed. The effectiveness of the proposed model is evaluated in terms of impact on system response and error-detection latency. Experimental results are presented for OpenCore benchmarks. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Signature Analysis for Testing, Diagnosis, and Repair of Multi-mode Power SwitchesabstractPower-gating structures for intermediate power-off modes offer significant power saving benefits as they reduce the leakage power during short periods of inactivity. However, reliable operation of such devices must be ensured by using adequate test methods. We propose a signature analysis technique to efficiently test power-gating structures that provide intermediate power-off modes. In particular, the proposed technique can be used to test and diagnose an efficient multi-mode power-gating architecture that was proposed recently. In addition, we propose a methodology to repair catastrophic and parametric faults, and to tolerate process variations. Analysis and extensive simulations demonstrate the effectiveness of the proposed method. Zhaobo Zhang, Xrysovalantis Kavousianos, Yiorgos Tsiatouhas, Krishnendu Chakrabarty |
ETS | 1 |
| 2011 | A BIST scheme for testing and repair of multi-mode power switchesabstractIt was shown recently that signature analysis can be used for the test, diagnosis and repair of a robust multi-mode power-gating architecture. A drawback of this approach is that it requires a tester in a production-test environment, and potentially expensive manufacturing steps are necessary to repair defective power switches. We propose a built-in self-test (BIST) and built-in-self-repair (BISR) scheme for test and repair of multi-mode power switches. The proposed method reduces test cost and obviates additional manufacturing steps for post-silicon repair. In addition to eliminating the need for an external tester, it offers protection against latent defects that are manifested as errors in the field. In this way, the robust BIST/BISR solution for power switches enhances the reliability of multi-core chips that employ aggressive power management techniques. Simulation results highlight the low hardware overhead and effectiveness of the proposed method for detecting, diagnosing and repairing defects. Zhaobo Zhang, Xrysovalantis Kavousianos, Yiorgos Tsiatouhas, Krishnendu Chakrabarty |
IOLTS | 1 |
| 2011 | Smart diagnosis: Efficient board-level diagnosis and repair using artificial neural networksabstractDiagnosis of functional failures at the board level is critical for improving product yield and reducing manufacturing cost. State-of-the-art board-level diagnostic software is unable to cope with high complexity and ever-increasing clock frequencies, and the identification of the root cause of failure on a board is a major problem today. Ambiguous or incorrect repair suggestions lead to long debug times and even wrong repair actions, which significantly increases the repair cost and adversely impacts yield. We propose a smart diagnosis method based on artificial neural networks that can learn from repair history and accurately localize the root cause of a failure. Fine-grained fault syndromes extracted from failure logs and the corresponding repair actions are used to train the neural network. The proposed network structure is simple, it can be rapidly trained, and it is scalable to large datasets. Moreover, the relationship between typical syndromes and the most appropriate repair actions can be easily inferred from the network structure. An industrial board, which is currently in production, is used to validate the diagnosis approach in terms of diagnostic accuracy, resolution, and quantifiable improvement over current diagnostic software. Zhaobo Zhang, Krishnendu Chakrabarty, Zhanglei Wang, Xinli Gu |
ITC | 1 |
| 2010 | Optimization and Selection of Diagnosis-Oriented Fault-Insertion Points for System TestabstractHardware fault-insertion test is a promising method to diagnose functional failures and target ''no trouble found (NTF)" problems in electronic systems. However, it is costly and impractical to equip all the potential fault sites with fault-insertion hardware. We present an optimization method to select the most effective outputs of a module where fault insertion logic must be placed to facilitate diagnosis. Faults inserted at the selected outputs are able to generate fault syndromes that are most similar to the errors produced by defects inside the module. This approach also ensures that the ambiguous fault candidates from other modules are maximally removed from the set of suspects. A fault syndrome is defined by the order of error occurrence at the observation points, and it is referred as an error flow. The similarity between two error flows is measured by the metric of edit distance. An integer linear programming model is used to maximize diagnostic effectiveness with a small number of fault-insertion points. Results on diagnostic accuracy for an open-source RISC highlight the effectiveness of the proposed method compared to a baseline random fault-insertion scheme. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
Asian Test Symposium | 1 |
| 2010 | Board-level fault diagnosis using an error-flow dictionaryabstractDiagnosis of functional failures is critical for locating manufacturing defects, increasing yield, and reducing field returns. It is important to narrow down the defective module in a failed component during board-level diagnosis. In this paper, a generic fault-diagnosis method based on an error-flow dictionary is presented to identify the root cause of functional failures on a chip or board. Error propagation mimics actual dataflow in a circuit, thus it reflects the native (functional) mode of circuit operation. In contrast to conventional fault syndromes, error flow includes the failure information in terms of circuit functionality, which significantly facilitates the diagnosis of functional failures. In the proposed diagnosis procedure, error flow is first learned from a good circuit by intentionally inserting faults, and then the root cause of a failing circuit is determined by comparing the similarity between the pre-learned error flow and the error flow observed from the failing circuit. The similarity of two error flows is evaluated based on the length of the longest common subsequence in string matching. Results for an open-source RISC SoC and an industrial communication circuit highlight the effectiveness of the proposed method. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
ITC | 1 |
| 2010 | Board-level fault diagnosis using Bayesian inferenceabstractIncreasing integration densities and high operating speeds are leading to subtle manifestations of defects at the board level. Board-level functional test is therefore necessary for product qualification. The diagnosis of functional failures is especially challenging, and the cost associated with board-level diagnosis is escalating rapidly. An effective and cost-efficient board-level diagnosis strategy is needed to reduce manufacturing cost and time-to-market, as well as to improve product quality. In this paper, we use Bayesian inference to develop a new board-level diagnosis framework that allows us to identify faulty devices or faulty modules within a device on a failing board with high confidence. Bayesian inference offers a powerful probabilistic method for pattern analysis, classification, and decision making under uncertainty. We apply this inference technique by first generating a database of fault syndromes obtained using fault-insertion test at the module pin level on a fault-free board, and then use this database along with the observed erroneous behavior of a failing board to infer the most likely faulty device. Results on a case study using an open-source RISC system-on-chip highlight the effectiveness of the proposed framework in terms of fault-localization accuracy and correctness of diagnosis. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
VTS | 1 |
| 2009 | Physical defect modeling for fault insertion in system reliability testabstractHardware fault-insertion test (FIT) is a promising method for system reliability test and diagnosis coverage measurement. It improves the speed of releasing a quality diagnostic program before manufacturing and provides feedbacks of fault tolerance of a very complicated large system. Certain level insufficient fault tolerance can be fixed in the current system but others may require ASIC or overall system architectural modifications. The FIT is achieved by introducing an artificial fault (defect modeling) at the pin level of a module to mimic any physical defect behavior within the module, such as SEU (single event upset) or escaped delay defect. We present a hardware architectural solution for pin fault insertion. We also present a simulation framework and optimization techniques for a subset of module pin selection for FIT, such that desired coverage are obtained under the constraints of limited FIT pins due to the costs of the associated implementation. Experimental results are presented for selected ISCAS and OpenCore benchmarks, as well as for an industrial circuit. Zhaobo Zhang, Zhanglei Wang, Xinli Gu, Krishnendu Chakrabarty |
ITC | 1 |