Kaibo Liu

dblp:131/9740 · DBLP profile ↗
← Back
46ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0003-2863-5748ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DrWASI: LLM-assisted Differential Testing for WebAssembly System Interface Implementations
abstract
WebAssembly (Wasm) is an emerging binary format that serves as a compilation target for over 40 programming languages. Wasm runtimes provide execution environments that enhance portability by abstracting away operating systems and hardware details. A key component in these runtimes is the WebAssembly System Interface (WASI), which manages interactions with operating systems, like file operations. Considering the critical role of Wasm runtimes, the community has aimed to detect their implementation bugs. However, no work has focused on WASI-specific bugs that can affect the original functionalities of running Wasm binaries and cause unexpected results. To fill the void, we present DrWASI , the first general-purpose differential testing framework for WASI implementations. Our approach uses a large language model to generate seeds and applies variant and environment mutation strategies to expand and enrich the test case corpus. We then perform differential testing across major Wasm runtimes. By leveraging dynamic and static information collected during and after the execution, DrWASI can identify bugs. Our evaluation shows that DrWASI uncovered 33 unique bugs, with all confirmed and 7 fixed by developers. This research represents a pioneering step in exploring a promising yet under-explored area of the Wasm ecosystem, providing valuable insights for stakeholders.
Ningyu He, Jianting Gao, Shangtong Cao, Kaibo Liu, Haoyu Wang 0001, Yun Ma 0002, Gang Huang 0001, Xuanzhe Liu
ACM Trans. Softw. Eng. Methodol.5
2025 LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
abstract
Kaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang, Mark Harman, Yudong Han, Yun Ma, Yihong Dong, Ge Li, Gang Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaibo Liu, Zhenpeng Chen 0001, Jie Zhang 0050, Mark Harman, Yudong Han 0001, Yun Ma 0002, Yihong Dong, Ge Li 0001, Gang Huang 0001
ACL (1)1
2025 Speedy Error Reconciliation
Kaibo Liu, Xiaozhuo Gu, Peixin Ren, Xuwen Nie, Yunlv Lv
Inscrypt (1)1
2025 Degradation Modeling and Prognostic Analysis Under Unknown Failure Modes
abstract
Operating units often experience various failure modes in complex systems, leading to distinct degradation paths. Relying on a prognostic model trained on a single failure mode may result in poor generalization performance across multiple failure modes. Therefore, accurately identifying the failure mode is of critical importance. Current prognostic approaches either ignore failure modes during degradation or assume known failure mode labels, which can be challenging to acquire in practice. Moreover, the high dimensionality and complex relations of sensor signals make it challenging to identify the failure modes accurately. To address these issues, we propose a novel failure mode diagnosis method that leverages a dimension reduction technique called UMAP (Uniform Manifold Approximation and Projection) to project and visualize each unit’s degradation trajectory into a lower dimension. Then, using these degradation trajectories, we develop a time series-based clustering method to identify the training units’ failure modes. Finally, we introduce a monotonically constrained prognostic model to predict the failure mode labels and Remaining Useful Life (RUL) of the test units simultaneously using the obtained failure modes of the training units. The proposed prognostic model provides failure mode-specific RUL predictions while preserving the monotonic property of the RUL predictions across consecutive time steps. We evaluate the proposed model using a case study with the aircraft gas turbine engine dataset.Note to Practitioners—The paper aims to develop an unsupervised method for identifying potential failure modes based on multiple sensor signals during the degradation process. After obtaining the failure mode labels, we introduce a monotonically constrained prognostic model for jointly predicting the failure modes and the RUL for the test units. Implementing this method involves four steps: First, collecting multiple sensor signals, and the failure times of historical units. Second, applying the UMAP dimension reduction technique to transform the high-dimensional multisensor data into low-dimensional representations and then visualizing the degradation trajectories for each unit. Third, utilizing a time series-based clustering method to get the failure mode labels of each training unit. Fourth, constructing a joint prognostic model to predict the failure modes and provide monotonically constrained RUL predictions for the test units. The proposed data-driven method can capture complex data relationships and has a flexible model structure, allowing it to adapt to different types of sensor data and operating conditions. In addition, the proposed model provides RUL predictions that are consistent with prior domain knowledge. As a result, the proposed model can be applied in various practical scenarios that involve degradation systems characterized by complex structures and unknown failure modes, providing a valuable tool for practitioners in real-time decision-making and maintenance planning.
Ye Kwon Huh, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2025 A Bayesian Spike-and-Slab Sensor Selection Approach for High-Dimensional Prognostics
abstract
With recent advances in sensor technology, more and more sensors are being used to simultaneously monitor the degradation of a system. As the number of sensors increases, it becomes increasingly difficult to distinguish informative sensors from uninformative sensors when performing prognostics, especially under the presence of different sensor correlations, signal-to-noise ratios, measurement units, and data characteristics. Existing methods for sensor selection typically rely on penalized-likelihood methods, which are known to provide biased estimates and poor sensor selection results in such high-dimensional settings. To overcome this challenge, we propose a novel data-fusion method that simultaneously selects informative sensors using Bayesian spike-and-slab priors and fuses the informative sensors into a 1-D health index (HI) to better characterize the degradation process for further prognostic analysis. Compared to the existing literature, the proposed Bayesian spike-and-slab sensor selection approach provides several unique advantages: 1) superior sensor selection performance in high-dimensional scenarios; 2) consistent sensor selection results with correlated sensors; 3) guaranteeing weak and strong selection consistency under mild assumptions; and 4) higher RUL prediction accuracy in a wide range of simulation and case studies. Note to Practitioners—This paper is motivated by the practical challenge of selecting informative sensors from high-dimensional multisensor systems (i.e., when there are many sensors relative to the number of training units) when conducting prognostics. Informative sensors not only provide valuable insights on the system’s degradation process, but also can be used to construct degradation indicators for practitioners to monitor and interpret the system status. In order to select informative sensors, this paper proposes a novel approach involving Bayesian spike-and-slab priors. After we select the informative sensors, they are then fused into a 1-D HI to better characterize the degradation process. This approach is particularly useful for selecting informative sensors under high-dimensional scenarios with possibly correlated sensors, when all units degrade under a single failure mode and operating condition.
Ye Kwon Huh, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2025 An Integrated Uncertainty Quantification Model for Longitudinal and Time-to-Event Data
abstract
We present a novel joint prognostic framework for the integrated analysis and uncertainty quantification of longitudinal (i.e., multi-sensor degradation signals) data and time-to-event data. Specifically, the proposed method models longitudinal data using a functional principal component analysis (FPCA), while the time-to-event data is characterized by a Bayesian neural network-based Cox (BNN-Cox) model. The proposed method delivers several unique advantages: 1) Providing accurate remaining useful life (RUL) predictions while seamlessly integrating the uncertainties of both the longitudinal and time-to-event sub-models; 2) Demonstrating great flexibility in modeling both data types; 3) Allowing online, real-time updates of the RUL distribution as new measurements are collected; and 4) Making reliable predictions under limited data availability. Compared to existing methods that provide limited uncertainty information restricted to a single sub-model, the proposed approach offers more accurate and comprehensive uncertainty information via uncertainty propagation. The numerical evaluations on simulated and real-world data suggest that the proposed method achieves outstanding performance compared to existing benchmarks. Note to Practitioners—This paper is motivated by the practical issue of extracting prognostic insights from longitudinal and time-to-event data. There are two fundamental research questions involved: 1) How to accurately model both types of data without resorting to restrictive parametric assumptions; and 2) how to seamlessly integrate the uncertainties from both sub-models into the final RUL predictions. The proposed method is particularly useful in cases when there are modeling uncertainties in both longitudinal and time-to-event data, such as complex manufacturing or energy systems with multiple sensors, such as aircraft engines. There are four main steps involved when implementing the proposed method: 1) fit the historical longitudinal data using an FPCA-based degradation sub-model; 2) construct a BNN-Cox sub-model using the fitted longitudinal data and time-to-event data; 3) predict the degradation status and remaining useful life of the in-service units based on their longitudinal data and time-to-event data; and 4) provide uncertainty quantifications of the RUL estimates by integrating the uncertainties across the two sub-models. A key advantage of the proposed method is that practitioners can assess the reliability of RUL predictions by leveraging the well-quantified uncertainty estimates, allowing them to make well-informed maintenance decisions and avoid unnecessary operational expenses.
Ye Kwon Huh, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2025 Online Monitoring of Heterogeneous Partially Observable Data Streams Based on Q-Learning
abstract
With the rapid advances in Internet of Things (IoT) technology and computational infrastructure, heterogeneous data streams are becoming common in various manufacturing applications. Meanwhile, the resource constraints often restrict the full observability of data streams due to limited budget for deploying or turning on every sensor at a time, as well as limited transmission and processing time for collecting high frequency information from all data streams, which poses significant challenges for multivariate statistical process control (SPC) and quality improvement. In this article, diverging from conventional heuristic approaches, we propose a new algorithm based on Q-learning to online monitor and quickly detect mean shifts occurring to heterogeneous data streams in the context of limited resources, where only a subset of observations is available at each acquisition time. In particular, we integrate Q-learning with a nonparametric cumulative sum (CUSUM) procedure to effectively detect a wide range of possible mean shifts when data streams follow arbitrary distributions. Both simulations and a case study are thoroughly conducted to evaluate the performance and demonstrate the superiority of the proposed method.Note to Practitioners—This paper is motivated by the practical issue of online process monitoring and anomaly detection with resource limitations. In particular, we can only select a subset of data streams to monitor at each time epoch, and the challenges are to dynamically choose which ones to observe and when to raise an alarm. Unlike the existing methodologies which are heuristic and only consider short-term rewards from the dynamic sampling, this paper proposes a novel monitoring and sampling strategy that allows the practitioners to cost-effectively monitor heterogeneous data streams by considering the long-term rewards. Three main steps are involved in the proposed method: (i) construct a local nonparametric monitoring statistic for each data stream; (ii) train the proposed reinforcement learning framework; and (iii) at each time epoch, determine the most informative data stream to observe according to the Q-table and decide whether to raise the alarm. Experimental results through simulations and a case study have shown that the proposed method has better performance than the existing methods in reducing detection delay.
Haoqian Li, Honghan Ye, Jing-Ru C. Cheng, Kaibo Liu
IEEE Trans Autom. Sci. Eng.4
2025 Online Monitoring of High-Dimensional Data Streams With Deep Q-Network
abstract
With the fast advancements in Internet of Things (IoT) technology and sensing infrastructure, a wide range of systems continues to generate a massive amount of data. Meanwhile, practical resource constraints such as limited bandwidth or processing capability restrict the full observability of data streams in real time. As a result, the practitioners often need to dynamically decide which data streams to observe given the resource constraints in order to quickly detect any system anomaly as soon as possible. In this article, we propose a reinforcement learning framework based on deep Q-learning and combine it with statistical process control (SPC) techniques to effectively monitor high-dimensional data streams when only partial observations are available at each acquisition time due to resource constraints. To the best of our knowledge, this is the first work that integrates deep reinforcement learning, which considers long-term rewards associated with the dynamic data sampling strategy, into SPC charts; this integration effectively addresses the challenge of resource constraints in monitoring high-dimensional data streams. Specifically, we construct a nonparametric monitoring statistic for each data stream and develop a reinforcement learning framework to automatically identify the most informative data streams for observation at each time epoch. The state space, action space, and rewards in the reinforcement learning framework are carefully designed and a Double Dueling Q-network is trained accordingly. Unlike existing methods, which rely on heuristic approaches to determine the sampling strategy, the proposed framework maximizes the long-term reward, thus leading to superior performance compared to existing benchmarks. Numerical simulations and a case study are thoroughly conducted, showing that the proposed method outperforms the state-of-the-art algorithms by significantly reducing detection delay. Note to Practitioners—This paper is motivated by the practical issue of online process monitoring and anomaly detection with resource constraints. In particular, due to resource constraints, practitioners can only select a subset of data streams to monitor at each time epoch. Thus, the central challenges are to dynamically choose which data streams to observe and to decide when to raise an alarm. Unlike the existing methodologies which are heuristic and only consider short-term rewards from the dynamic sampling, this paper proposes a novel reinforcement learning framework that allows practitioners to monitor high-dimensional heterogeneous data streams more efficiently by considering the long-term rewards associated with the dynamic data sampling strategy. Four main steps are involved in the proposed method: (i) construct a nonparametric local monitoring statistic for each data stream; (ii) train the Double Dueling Q-network offline according to the proposed deep reinforcement learning framework; (iii) at each time epoch of the online monitoring, determine the most informative data streams to observe based on the deep Q-network; and (iv) determine whether to raise an alarm in the system.
Haoqian Li, Ziqian Zheng, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2025 Prediction of Condition Monitoring Signals Using Scalable Pairwise Gaussian Processes and Bayesian Model Averaging
abstract
Predicting condition monitoring signals has become a critical task for health status assessment and monitoring of industrial systems. It is crucial to incorporate correlated historical data when making predictions for the target signal. As a flexible nonparametric approach, the multi-output Gaussian process (MOGP) model can be employed for this nonlinear regression problem. One effective way to construct MOGP is through convolving a latent function drawn from a Gaussian process. While leveraging the convolution process makes MOGP expressive, there are several challenges that remain to be addressed. First, the scalability of MOGP is always an important issue since the computational demands would increase drastically as the dimension of output variables grows. Besides, the negative transfer should be mitigated when the target variable and the source variable share little commonality. In this study, a pairwise structure is adopted by decomposing the full multivariate model into a group of bi-output models. Furthermore, a Bayesian model averaging approach is utilized to combine the prediction results of the bi-output models. A model selection scheme based on Bayes factor is employed to alleviate negative transfer and facilitate model scalability further by discarding the weakly correlated outputs. The key advantage of the proposed model lies in the improved prediction and uncertainty quantification performance compared with traditional MOGP models. The superiority of the proposed method is validated by numerical studies and a case study.Note to Practitioners—This study addresses the challenge of predicting condition monitoring signals in a high-dimensional setting. Existing nonparametric approaches suffer from high computation complexity or ineffective information integration from different signals. We propose a novel approach using a scalable pairwise MOGP model based on Bayesian model averaging. Our method decomposes the full model into bi-output submodels and averages them in a Bayesian way. We also employ a model selection scheme based on the Bayes factor to alleviate negative transfer by discarding weakly correlated outputs. Numerical experiments suggest that this approach can improve prediction accuracy and uncertainty quantification performance for the prediction of condition monitoring signals. Our approach offers a promising solution for condition monitoring signal prediction in automatic and industrial systems where condition monitoring data are readily available.
Jinwen Sun, Dharmaraj Veeramani, Kaibo Liu
IEEE Trans Autom. Sci. Eng.4
2025 Self-Starting Monitoring and Dynamic Sampling of High-Dimensional Data Streams
abstract
In today’s manufacturing industries, the development of sensor technology and Internet of Things has made real-time process monitoring of high-dimensional data increasingly vital. However, resource constraints, such as limited power, budget, and transmission capacity, often prevent access to full data streams in real time. This means that practitioners need to effectively monitor the process based on only partially observed data by dynamically deciding the sampling layout in real time. Another common challenge of process monitoring in practice is the lack of historical reference data, which can occur due to process/system upgrades or equipment replacements. To address these critical challenges, this paper proposes MASS (Monitoring with Adaptive Sampling under Self-starting scheme), a novel self-starting monitoring approach tailored to monitor high-dimensional data streams when only limited resources and historical reference data are available. Our monitoring framework is based on a quantile-based nonparametric CUSUM procedure with likelihood ratio-based statistics, and then the Thompson Sampling (TS) algorithm is adopted to handle partially observed data in the self-starting scenario. A key feature of our proposed method is its adaptive estimation of the out-of-control distribution and data quantiles, which ensures robust detection for various shifts in data streams with arbitrary and heterogeneous distributions, even in cases with limited reference data. The outperformance of the proposed method is demonstrated through simulation experiments and a real-world case study. Note to Practitioners—This paper is motivated by the critical challenges of online process monitoring when only limited resources (e.g., limited power availability, limited number of sensors, and limited transmission capacity) and limited historical in-control reference data are available. For example, consider a scenario where a newly established production system requires online monitoring across multiple data streams. In such cases, there is often a deficiency of reference data crucial for constructing a reliable control chart. Additionally, due to resource constraints, it is frequently infeasible to gather information from all streams associated with the production line at each epoch in real time. Unlike previous methods which require either a sufficient amount of reference data, or fully observable data streams, this paper proposes a novel monitoring and dynamic sampling scheme to effectively monitor partially observable data streams with only a small amount of reference data. To implement the methodology, it requires: (i) to initiate the process by estimating quantiles with a small amount of reference data, (ii) to determine which data streams to observe at each time epoch, (iii) to adaptively update the estimation of process parameters during online monitoring, and (iv) to construct a set of local and global statistics that can be used to quickly detect the system anomaly in real time. Numerical experiments and a real-world case study suggest that our proposed method efficiently leverages available data to reduce detection delays and enhance effectiveness against various shifts, in comparison to the benchmark methods.
Ziqian Zheng, Jun Li 0023, Kaibo Liu
IEEE Trans Autom. Sci. Eng.4
2024 HITS: High-coverage LLM-based Unit Test Generation via Method Slicing
abstract
Large language models (LLMs) have behaved well in generating unit tests for Java projects. However, the performance for covering the complex focal methods within the projects is poor. Complex methods comprise many conditions and loops, requiring the test cases to be various enough to cover all lines and branches. However, existing test generation methods with LLMs provide the whole method-to-test to the LLM without assistance on input analysis. The LLM has difficulty inferring the test inputs to cover all conditions, resulting in missing lines and branches. To tackle the problem, we propose decomposing the focal methods into slices and asking the LLM to generate test cases slice by slice. Our method simplifies the analysis scope, making it easier for the LLM to cover more lines and branches in each slice. We build a dataset comprising complex focal methods collected from the projects used by existing state-of-the-art approaches. Our experiment results show that our method significantly outperforms current test case generation methods with LLMs and the typical SBST method Evosuite regarding both line and branch coverage scores.
Kaibo Liu, Ge Li 0001, Zhi Jin 0001
ASE2
2024 TrickyBugs: A Dataset of Corner-case Bugs in Plausible Programs
abstract
We call a program that passes existing tests but still contains bugs as a buggy plausible program. Bugs in such a program can bypass the testing environment and enter the production environment, causing unpredictable consequences. Therefore, discovering and fixing such bugs is a fundamental and critical problem. However, no existing bug dataset is purposed to collect this kind of bug, posing significant obstacles to relevant research. To address this gap, we introduce TrickyBugs, a bug dataset with 3,043 buggy plausible programs sourced from human-written submissions of 324 real-world competition coding tasks. We identified the buggy plausible programs from approximately 400,000 submissions, and all the bugs in TrickyBugs were not previously detected. We hope that TrickyBugs can effectively facilitate research in the fields of automated program repair, fault localization, test generation, and test adequacy.
Kaibo Liu, Yudong Han 0001, Jie Zhang 0050, Zhenpeng Chen 0001, Federica Sarro, Gang Huang 0001, Yun Ma 0002
MSR1
2024 An Integrated Deep Learning-Based Data Fusion and Degradation Modeling Method for Improving Prognostics
abstract
Accurate prognostics are crucially important to prevent unexpected failures in industrial and service systems. This process aims to monitor the degradation status of units and predict their remaining useful lifetime (RUL) by analyzing the data collected from multiple sensors. Existing studies for prognostics either focus on health index (HI)-based statistical fusion methods that are limited by restrictive assumptions or machine learning methods that model the HI and degradation status in two separate steps. However, the restrictive assumptions are often invalid in practice, and the intrinsic connection between the HI and degradation status is missing if the two parts are modeled separately, leading to poor prognostic results. This paper proposes an integrated deep learning-based data fusion and degradation modeling method by integrating a deep neural network (DNN) and a long short-term memory (LSTM) to characterize the nonlinear relationship between the HI and multiple sensor signals and to describe the underlying degradation status of units. In particular, our innovative idea is to develop an integrated backpropagation parameter estimation algorithm to solve the fusion procedure and the degradation modeling in an integrated manner by considering the properties of HI construction in the loss functions. Thus, the constructed HI is expected to better characterize the underlying degradation process and lead to a superior prognostic result. In the case study on the degradation of aircraft gas turbine engines, the proposed method achieves promising performance compared with the existing benchmarks of statistical models and other deep learning models. Note to Practitioners—This paper develops an integrated deep learning-based data fusion and degradation modeling method for improving prognostics when multiple sensors are available to monitor the degradation status of a unit. There are four steps for implementing this method in practice: 1) collecting multiple sensor signals of historical units; 2) constructing the HI and modeling the degradation status of units by combining a DNN model and an LSTM model; 3) solving the fusion procedure and the degradation modeling in an integrated manner by developing an integrated backpropagation parameter estimation algorithm; and 4) making prognostics for in-service units. The novelty of the proposed method is that it conducts prognostics by combining a DNN fusion model and an LSTM degradation model, and seamlessly integrates the fusion procedure with the degradation modeling to construct the HI for better characterizing the status of a unit. As a result, the proposed method has two main advantages: (i) capable of characterizing various degradation processes of different engineering systems; and (ii) superior prognostic results by constructing more suitable HI for the degradation process.
Di Wang 0019, Kaibo Liu
IEEE Trans Autom. Sci. Eng.2
2024 Transfer Learning-Based Independent Component Analysis
abstract
Understanding the underlying component structure is crucial for multivariate signal analysis. Among all the techniques that try to learn the latent structure, independent component analysis (ICA) is one of the most important and popular methods, which aims to extract independent components from multivariate signals and enables further analysis. For example, in electroencephalogram (EEG) analysis, artifacts filtering and disease detection are conducted based on the independent components of the signals. One critical challenge in existing ICA approaches is that the component extraction accuracy may degrade when the available data of a unit are limited. To address this issue, this paper proposes a transfer learning-based ICA method by innovatively transferring component distribution from a source domain, so that accurate component extraction results can be achieved even when only limited data are available in the target domain. To the best of our knowledge, this is the first work that leverages transfer learning to improve ICA accuracy with limited available data. In particular, we first extract all the independent components from the source domain by maximizing the log-likelihood function with a Newton-like method on a smooth manifold. Then for the target domain, the component with the largest negentropy is extracted in each round. To effectively leverage the knowledge from the source domain and to prevent the negative transfer, we try to find a component in the source domain that matches the component we are extracting. The probability density function of the matched component will then be used to improve the component extraction accuracy if such matched component can be found; otherwise, no knowledge will be transferred. Numerical simulations and a case study with electrocardiogram (ECG) data are conducted, showing the effectiveness of the proposed method in transferring knowledge and reducing negative transfer. Note to Practitioners—This paper is motivated by the practical issue of conducting independent component analysis when data are limited to derive a reliable result. To address this challenge, we propose to transfer knowledge from a source domain to the target domain. Specifically, there are two fundamental questions involved: 1) what knowledge can be transferred from the source domain; and 2) how to minimize the negative transfer when no useful knowledge is available. Our novel idea is to transfer the component distributions and the negative transfer is largely reduced through a component matching step as a result. There are four main steps involved when implementing the proposed method: 1) solve all the independent components in the source domain; 2) extract the independent component in the target domain with the largest negentropy; 3) decide whether a matched component can be found from the source domain; and 4) re- estimate the current independent component in the target domain by leveraging the distribution information of the matched component. Step 2 to step 4 are repeated several times until all the independent components in the target domain are extracted or some termination condition is reached.
Ziqian Zheng, Brock Hable, Yutao Gong, Robert W. Shannon, Kaibo Liu
IEEE Trans Autom. Sci. Eng.7
2024 Instance Selection via Voronoi Neighbors for Binary Classification Tasks
abstract
Large datasets available in many applications have enabled the training of binary classifiers to match or even outperform humans. However, the large volume of data introduces computational burden during the training and calibration of model parameters. Since the optimal decision surface (ODS) of a classification task is often determined by a few nearby instances, a novel PDOC-V method is proposed to identify them. A Bayesian probability model is adopted to describe the ODS. An instance is close to the ODS if its probability of belonging to the positive and negative classes is similar. The probabilities of an instance are estimated by partitioning the input space into cells containing a single instance via the Voronoi diagram and inspecting its Voronoi neighbors. A randomized ray shooting algorithm is adopted to accelerate our algorithm. In many natural datasets, the spatial distribution of instances is often uneven. For such datasets, our method is more robust than existing distance-based instance selection methods. Comprehensive experiments suggest that common classifiers trained on instances selected by PDOC-V can accurately recover the ODS. Moreover, for many natural datasets, common classifiers trained on 10% - 20% of instances can achieve more than 98% of the full set performance.
Kaibo Liu
IEEE Trans. Knowl. Data Eng.2
2024 Real-time Cyber-Physical Security Solution Leveraging an Integrated Learning-Based Approach
abstract
Cyber-Physical Systems (CPS) has emerged as a paradigm that connects cyber and physical worlds, which provides unprecedented opportunities to realize intelligent applications such as smart home, smart cities, and smart manufacturing. However, CPS faces a great number of information security challenges (e.g., attacks) due to the integration of CPS as well as the human behaviors and interactions. Therefore, accurate and real-time attack detection and identification are essential to ensure information security and reliability of CPS. In this paper, we propose a novel integrated learning method that accurately detects an attack of a CPS system and then identifies the attack type in real time. Specifically, we consider a One-Class Support Vector Machine (OCSVM) model that only relies on the data from the normal state for training to achieve a real-time and effective detection of a CPS system state (i.e., normal or under-attack). If the system is detected to be under-attack, we then develop a Pairwise Self-supervised Long Short-Term Memory (PSLSTM) approach to identify the attack type, which aims to accurately distinguish the known attack types and discover unknown new attacks. Lastly, experimental results show the proposed method achieves promising performances compared with conventional and state-of-the-art learning-based benchmarks.
Di Wang 0019, Fangyu Li 0002, Kaibo Liu, Xi Zhang 0006
ACM Trans. Sens. Networks3
2023 Who Judges the Judge: An Empirical Study on Online Judge Tests
abstract
Online Judge platforms play a pivotal role in education, competitive programming, recruitment, career training, and large language model training. They rely on predefined test suites to judge the correctness of submitted solutions. It is therefore important that the solution judgement is reliable and free from potentially misleading false positives (i.e., incorrect solutions that are judged as correct). In this paper, we conduct an empirical study of 939 coding problems with 541,552 solutions, all of which are judged to be correct according to the test suites used by the platform, finding that 43.4% of the problems include false positive solutions (3,440 bugs are revealed in total). We also find that test suites are, nevertheless, of high quality according to widely-studied test effectiveness measurements: 88.2% of false positives have perfect (100%) line coverage, 78.9% have perfect branch coverage, and 32.5% have a perfect mutation score. Our findings indicate that more work is required to weed out false positive solutions and to further improve test suite effectiveness. We have released the detected false positive solutions and the generated test inputs to facilitate future research.
Kaibo Liu, Yudong Han 0001, Jie Zhang 0050, Zhenpeng Chen 0001, Federica Sarro, Mark Harman, Gang Huang 0001, Yun Ma 0002
ISSTA1
2023 GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning
abstract
Training data pre-processing pipelines are essential to deep learning (DL). As the performance of model training keeps increasing with both hardware advancements (e.g., faster GPUs) and various software optimizations, the data pre-processing on CPUs is becoming more resource-intensive and a severe bottleneck of the pipeline. This problem is even worse in the cloud, where training jobs exhibit diverse CPU-GPU demands that usually result in mismatches with fixed hardware configurations and resource fragmentation, degrading both training performance and cluster utilization. We introduce GoldMiner, an input data processing service for stateless operations used in pre-processing data for DL model training. GoldMiner decouples data pre-processing from model training into a new role called the data worker. Data workers facilitate scaling of data pre-processing to anywhere in a cluster, effectively pooling the resources across the cluster to satisfy the diverse requirements of training jobs. GoldMiner achieves this decoupling in a fully automatic and elastic manner. The key insight is that data pre-processing is inherently stateless, thus can be executed independently and elastically. This insight guides GoldMiner to automatically extract stateless computation out of a monolithic training program, efficiently disaggregate it across data workers, and elastically scale data workers to tune the resource allocations across jobs to optimize cluster efficiency. We have applied GoldMiner to industrial workloads, and our evaluation shows that GoldMiner can transform unmodified training programs to use data workers, accelerating individual training jobs by up to 12.1x. GoldMiner also improves average job completion time and aggregate GPU utilization by up to 2.5x and 2.1x in a 64-GPU cluster, respectively, by scheduling data workers with elasticity.
Zhi Yang 0001, Yu Cheng 0030, Chao Tian 0001, Shiru Ren, Wencong Xiao, Man Yuan, Langshi Chen, Kaibo Liu, Yang Zhang 0102, Yong Li 0045, Wei Lin 0016
Proc. ACM Manag. Data9
2022 Special Issue on Automation Analytics Beyond Industry 4.0: From Hybrid Strategy to Zero-Defect Manufacturing
abstract
Most traditional industries or emerging countries may not be capable of directly transiting to Industry 4.0. To fill the gap between as-is Industry 3.0 and to-be Industry 4.0, some disruptive innovations from automation and industrial engineering identify best practice with adopting cost-effective semi-automated systems to manage the potential socio-economic impacts of infrastructure disruptions, while considering total resource management for sustainability. This is the so-called “hybrid strategy (HS),” or “Industry 3.5.” On the other hand, the current Industry 4.0-related technologies should also emphasize quality enhancement to achieve “zero-defect manufacturing (ZDM),” also referred to as “Industry 4.1.” ZDM is a systematic strategy to realize the goal of Zero Defects, which includes two phases. Phase I: accomplish Zero Defects of all thedeliverablesby applying efficient and economical total-quality-inspection techniques; and Phase II: further ensure Zero Defects of all theproductsgradually by improving the yield with big data analytics and continuous improvement. Both the challenges and opportunities from HS to ZDM have significantly expanded the scope of traditional automation science and engineering.
Fan-Tien Cheng, Chia-Yen Lee, Min-Hsiung Hung, Lars Mönch, James R. Morrison, Kaibo Liu
IEEE Trans Autom. Sci. Eng.6
2022 Individualized Degradation Modeling and Prognostics in a Heterogeneous Group via Incorporating Intrinsic Covariate Information
abstract
This article focuses on individualized degradation modeling and prognostics for a heterogeneous group, where each individual unit shows a distinct degradation process. Existing degradation models usually treat each unit separately and do not fully utilize the distinct characteristics of each individual. In this study, we propose a generic framework to handle the heterogeneity across units by effectively leveraging the intrinsic covariate information, which is closely related to the unit’s degradation process. Specifically, we employ a multivariate Gaussian process (MGP) to nonparametrically establish the relation between the covariate information and degradation process. Through modeling the unit similarities based on the covariates, efficient information transfer among units is enabled for better degradation modeling and prognostics, as the collected degradation signals from one unit can be shared with the entire heterogeneous group. A theoretical justification for the proposed model is also investigated. Simulation studies are presented to evaluate the parameter estimation accuracy and the sensitivity of the proposed method. A case study on the Alzheimer’s disease (AD) neuroimaging initiative data set is further conducted, which demonstrates the advantage of the proposed method over existing benchmark approaches.Note to Practitioners—This article is motivated by the practical issue of degradation modeling and prognostics for a heterogeneous group, where all units in a group share some similarities and each unit has its own distinct individual-level characteristics (covariates). The covariates in this study refer to static intrinsic characteristics instead of dynamic external environmental conditions. Several practical examples are explained inSection Iwith more details. There are two fundamental questions involved: 1) how to quantify the distinct individual characteristics of each unit while representing the group-level commonalities among all units and 2) how to perform degradation modeling and prognostics of a newly launched unit with few degradation signals available. The novelty of this article lies in encoding the available knowledge about individual covariates and group-level commonalities into the degradation modeling and prognostics. There are three main steps involved when implementing the proposed method: 1) collecting degradation signals, failure time, and covariate information of heterogeneous units; 2) constructing an MGP-based degradation model; and 3) predicting the degradation status and remaining useful life of the in-service units based on their covariates and signals. The proposed method is particularly useful when we collect various intrinsic covariates which can effectively represent the individual-level characteristics and when only sparse or no data are available for the units of interest.
Changyue Song, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2022 Building Local Models for Flexible Degradation Modeling and Prognostics
abstract
To avoid unexpected failures of engineering systems, sensors have been widely used to monitor the degradation process of the systems. A number of studies have been conducted to analyze the collected sensor signals and predict the failure time. However, the existing studies are usually restricted and cannot be adapted to different practical situations. In this paper, we propose a systematic method for degradation modeling and prognosis that can be widely applied in different scenarios. In particular, the proposed method is capable to handle one or multiple sensors, powerful to capture the nonlinear relations between sensor signals and the degradation process with few assumptions, generic to consider multiple failure modes, flexible to deal with unequally spaced sensor measurements or asynchronous signals, and easily understandable with little preprocessing required. The main idea is to predict the failure time of an in-service unit based on a subset of the nearest historical units, where features are extracted from each sensor to describe the progression of sensor signals and local linear regression models are constructed to establish the relation between failure time and the extracted features. The prediction variance is then used as the goodness-of-fit measure, based on which decision-level fusion and feature-level fusion are proposed to combine multiple sensors. A case study with two datasets on the degradation modeling of aircraft engines is conducted which shows satisfactory performance of the proposed method. Note to Practitioners—This paper aims at modeling the collected sensor signals to understand the degradation process of the monitored engineering systems and predict the failure time. The main idea is to measure the similarity of units and predict the failure time of an in-service unit based on a subset of the nearest historical units. The developed method is widely applicable in different practical situations such as multiple sensors, multiple failure modes, asynchronous signals, and missing data. Furthermore, the method requires little preprocessing. There are several steps involved for implementing the proposed method: 1) collecting the sensor signals for historical units and the in-service unit; 2) extracting features from each sensor signal; 3) constructing a local linear model to predict the failure time based on the extracted features, and obtaining the prediction variance on the in-service unit; and 4) combining the information of different sensors using the decision-level fusion or feature-level fusion, if each unit is monitored by multiple sensors.
Changyue Song, Ziqian Zheng, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2022 A Generic Indirect Deep Learning Approach for Multisensor Degradation Modeling
abstract
To monitor the degradation status of units and prevent unexpected failures in engineering systems, health index (HI)-based data fusion technologies have been rapidly developed by combining multiple sensor signals, which are helpful to understand the degradation processes of units and predict their remaining useful lifetime (RUL). Although promising, existing HI-based data fusion models for degradation modeling are still limited due to the restrictive assumptions made during the fusion or the degradation modeling processes, e.g., assuming the fusion model as a linear or kernel-based function from multiple sensor signals, or modeling the degradation process by a preselected basis function. Such assumptions are often invalid in industrial practice and may fail to accurately characterize the complicated relationships between multiple sensor signals and the underlying degradation process. To address the issue, this article proposes a generic indirect deep learning method that constructs an HI by combining multiple sensor signals to better characterize the degradation process. In particular, our innovative idea is to seamlessly integrate a deep neural network (DNN) and a long short term memory (LSTM) model to construct the HI by fusing multiple sensor signals and characterize the degradation process, which can be applied to the degradation modeling of various engineering systems. Domain knowledge including the concept of failure threshold and monotonicity of the degradation process is also considered to enhance the interpretability of the proposed method. For parameter estimation, we develop an indirect gradient descent (IGD) algorithm to train the proposed method. Simulation studies and a case study on the degradation of aircraft gas turbine engines are presented to validate the performance of the proposed method.Note to Practitioners—The article aims to develop a generic health index (HI)-based data fusion method for degradation modeling when multiple sensors are available to monitor the degradation status of a unit. Specifically, the developed method addresses two challenging questions in practice: 1) how to effectively combine multiple sensor signals to construct an HI that accurately characterizes the underlying degradation status and 2) how to flexibly model the degradation evolution based on the constructed HI. To implement this method in practice, four steps are included as follows:First, collecting multiple sensor signals and failure time of historical units.Second, constructing the HI and describing the underlying degradation process by training a deep neural network (DNN) model and a long short term memory (LSTM) model, respectively.Third, estimating model parameters using the proposed IGD algorithm.Fourth, constructing the HI of in-service units and predicting their remaining useful lifetime (RUL) using the constructed HI. The proposed method is expected to be able to characterize various degradation processes and be applied to the degradation modeling of different engineering systems.
Di Wang 0019, Kaibo Liu, Xi Zhang 0006
IEEE Trans Autom. Sci. Eng.2
2022 A Generic Online Nonparametric Monitoring and Sampling Strategy for High-Dimensional Heterogeneous Processes
abstract
With the rapid advancement of in-process measurements and sensor technology driven by zero-defect manufacturing applications, high-dimensional heterogeneous processes that continuously collect distinct physical characteristics frequently appear in modern industries. Such large-volume high-dimensional data place a heavy demand on data collection, transmission, and analysis in practice. Thus, practitioners often need to decide which informative data streams to observe given the resource constraints at each data acquisition time, which poses significant challenges for multivariate statistical process control and quality improvement. In this article, we propose a generic online nonparametric monitoring and sampling scheme to quickly detect mean shifts occurring in heterogeneous processes when only partial observations are available at each acquisition time. Our innovative idea is to seamlessly integrate the Thompson sampling (TS) algorithm with a quantile-based nonparametric cumulative sum (CUSUM) procedure to construct local statistics of all data streams based on the partially observed data. Furthermore, we develop a global monitoring scheme of using the sum of top-${r}$local statistics, which can quickly detect a wide range of possible mean shifts. Tailored to monitoring the heterogeneous data streams, the proposed method balances between exploration that searches unobserved data streams for possible mean shifts and exploitation that focuses on highly suspicious data streams for quick shift detection. Both simulations and a case study are comprehensively conducted to evaluate the performance and demonstrate the superiority of the proposed method.Note to Practitioners—This paper is motivated by the critical challenge of online process monitoring by considering the cost-effectiveness and resource constraints in practice (e.g., limited number of sensors, limited transmission bandwidth or energy constraint, and limited processing time). Unlike the existing methodologies which rely on the restrictive assumptions (e.g., normally distributed, exchangeable data streams) or require historical full observations of all data streams to be available offline for training, this paper proposes a novel monitoring and sampling strategy that allows the practitioners to cost-effectively monitor high-dimensional heterogeneous data streams that contain distinct physical characteristics and follow different distributions. To implement the proposed methodology, it is necessary: (i) to identify sample quantiles for each data stream based on historical in-control data offline; (ii) to determine which data streams to observe at each acquisition time based on the resource constraints; and (iii) to automatically screen out the suspicious data streams to form the global monitoring statistic. Experimental results through simulations and a case study have shown that the proposed method has much better performance than the existing methods in reducing detection delay and effectively dealing with heterogeneous data streams.
Honghan Ye, Kaibo Liu
IEEE Trans Autom. Sci. Eng.2
2021 Collusion Detection and Ground Truth Inference in Crowdsourcing for Labeling Tasks
abstract
Crowdsourcing has been a prompt and cost-effective way of obtaining labels in many machine learning applications. In the literature, a number of algorithms have been developed to infer the ground truth based on the collected labels. However, most existing studies assume workers to be independent and are vulnerable to worker collusion. This paper aims at detecting the collusive behaviors of workers in labeling tasks. Specifically, we consider collusion in a pairwise manner and propose a penalized pairwise profile likelihood method based on the adaptive LASSO penalty for collusion detection. Many models that describe the behavior of independent workers can be incorporated into our proposed framework as the baseline model. We further investigate the theoretical properties of the proposed method that guarantee the asymptotic performance. An algorithm based on expectation-maximization algorithm and coordinate descent is proposed to numerically maximize the penalized pairwise profile likelihood function for parameter estimation. To the best of our knowledge, this is the first statistical model that simultaneously detects collusion, learns workers’ capabilities, and infers the ground true labels. Numerical studies using synthetic and real data sets are also conducted to verify the performance of the method.
Changyue Song, Kaibo Liu, Xi Zhang 0006
J. Mach. Learn. Res.2
2021 Adaptive Preventive Maintenance for Flow Shop Scheduling With Resumable Processing
abstract
In this article, we focus on a joint scheduling problem that considers the corrective maintenance (CM) due to unexpected breakdowns and the scheduled preventive maintenance (PM) in a generic M-machine flow shop. The objective is to find the optimal job sequence and PM schedule such that the total of the tardiness cost, PM cost, and CM cost is minimized. Currently, most existing studies on the PM schedules are based on a fixed PM interval, which is rigid and may lead to poor performance, as the fixed strategy fails to effectively balance the trade-offs between the production scheduling and maintenance. To address this critical research issue, our novel idea is to dynamically update the PM interval based on the real-time machine age, such that the maintenance activity coordinates with the job scheduling to the maximum extent, which results in an overall cost saving. Specifically, a correction factor is introduced to dynamically update the PM interval and to help evaluate whether it is worthwhile to process the job first at the risk of the CM before performing the PM action. To demonstrate the effectiveness of the adaptive strategy, simulations and a case study on mining operations are conducted to show that the adaptive strategy outperforms the existing methods with a less total cost.
Honghan Ye, Xi Wang 0020, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2020 Simultaneous Translation Policies: From Fixed to Adaptive
abstract
Adaptive policies are better than fixed policies for simultaneous translation, since they can flexibly balance the tradeoff between translation quality and latency based on the current context information.But previous methods on obtaining adaptive policies either rely on complicated training process, or underperform simple fixed policies.We design an algorithm to achieve adaptive policies via a simple heuristic composition of a set of fixed policies.Experiments on Chinese→English and German→English show that our adaptive policies can outperform fixed ones by up to 4 BLEU points for the same latency, and more surprisingly, it even surpasses the BLEU score of full-sentence translation in the greedy mode (and very close to beam mode), but with much lower latency.
Baigong Zheng, Kaibo Liu, Renjie Zheng, Mingbo Ma, Hairong Liu, Liang Huang 0001
ACL2
2020 Opportunistic Decoding with Timely Correction for Simultaneous Translation
abstract
Simultaneous translation has many important application scenarios and attracts much attention from both academia and industry recently.Most existing frameworks, however, have difficulties in balancing between the translation quality and latency, i.e., the decoding policy is usually either too aggressive or too conservative.We propose an opportunistic decoding technique with timely correction ability, which always (over-)generates a certain mount of extra words at each step to keep the audience on track with the latest information.At the same time, it also corrects, in a timely fashion, the mistakes in the former overgenerated words when observing more source context to ensure high translation quality.Experiments show our technique achieves substantial reduction in latency and up to +3.1 increase in BLEU, with revision rate under 8% in Chinese-to-English and English-to-Chinese translation.
Renjie Zheng, Mingbo Ma, Baigong Zheng, Kaibo Liu, Liang Huang 0001
ACL4
2020 Spatiotemporal Thermal Field Modeling Using Partial Differential Equations With Time-Varying Parameters
abstract
Accurate modeling of a thermal field is one of the fundamental requirements in engineering thermal management in numerous industries. Existing studies have shown that using differential equations to model a thermal field delivers good performance when the parameters are predetermined through physical or experimental analysis. However, due to variations of the inner medium affected by certain latent factors, the parameters in differential equation models may not be treated as constants while the thermal field is estimated, and this fact poses a new challenge to field estimation by directly solving the differential equation models. In this study, a novel approach to thermal field modeling is developed by considering the parameters as functional variables that vary temporally in partial differential equations (PDEs). This approach provides a new perspective to model the dynamic thermal field by fully using the collected sensor data from the thermal system. Specifically, time-varying parameters can be constructed through a combination of basis functions whose coefficients can be efficiently estimated through the sensor data. A two-level iterative parameter estimation algorithm is also tailored to obtain the parameters in the PDE model. Both simulation and real case studies show that our proposed approach provides satisfactory estimation performance compared with the benchmark method that uses the constant parameter estimation. Note to Practitioners-The proposed method aims to model a thermal field using PDEs with time-varying parameters. To better implement this method in practice, three things are noteworthy: first, the proposed method models a thermal field by fully considering physics-specific engineering knowledge using PDEs and the collected sensor data from thermal systems. Second, because time-varying parameters in PDEs cannot be estimated directly, the proposed model represents the time-varying parameters by a combination of B-spline basis functions in terms of time. Estimating time-varying parameters is converted into estimating the constant coefficients of the basis functions. Because the derivatives of a thermal field might not have an analytical expression, the proposed model represents the thermal field by a combination of B-spline basis functions. Taking the derivatives of the thermal field is converted into taking the derivatives of the corresponding basis functions. Third, the proposed method can not only model a thermal field but can also be applied in other physics-specific engineering cases.
Di Wang 0019, Kaibo Liu, Xi Zhang 0006
IEEE Trans Autom. Sci. Eng.2
2020 Spatiotemporal Multitask Learning for 3-D Dynamic Field Modeling
abstract
3-D dynamic field modeling using data acquired from sensor networks is typically complex due to the data sparsity and missing problem. In this article, we consider the ubiquitous missing data problem in current sensor networks and aim to take complete advantage of the existing sensor data for thermal field modeling. In the common scenario, data from the target network are not always obtainable, but data from other neighboring networks with homogeneous fields are accessible. Thus, a novel method that captures the information acquired from these neighboring networks is proposed. To achieve accurate thermal field estimation using limited sensor observations, we develop a mixed-effect model framework in which the dynamic field is decomposed into a mean profile and local variability. In particular, we establish a spatiotemporal field multitask learning (FML) approach to identify the spatiotemporal correlation by integrating a multitask Gaussian process (MGP) framework into an autoregressive (AR) model using neighboring data sources from homogeneous fields. Our proposed method is verified through a real case study of thermal field estimation during grain storage. Note to Practitioners-The proposed method aims to obtain an accurate estimation of a thermal field when certain sensor data are inaccessible. To better implement this method in practice, three things are noteworthy: First, the mean profile of the thermal field should be extracted using the thermodynamic model, so that the remaining data are able to follow a Gaussian process. Second, the FML approach considers neighboring data sources from homogeneous thermal fields to achieve an accurate estimation of the target thermal field. Thus, the target thermal field and other thermal fields should be under similar external conditions, e.g., environmental surroundings, geographical location, and field size. Third, the proposed method can not only process the data from grid-based sensor networks, but also can be extended to other topological structures of sensor networks for field estimation.
Di Wang 0019, Kaibo Liu, Xi Zhang 0006, Hui Wang 0035
IEEE Trans Autom. Sci. Eng.2
2019 STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix-to-Prefix Framework
abstract
Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, Haifeng Wang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Mingbo Ma, Liang Huang 0001, Hao Xiong 0005, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Hua Wu 0003, Haifeng Wang 0001
ACL (1)5
2019 LinearFold: linear-time approximate RNA folding by 5'-to-3' dynamic programming and beam search
abstract
MOTIVATION: Predicting the secondary structure of an ribonucleic acid (RNA) sequence is useful in many applications. Existing algorithms [based on dynamic programming] suffer from a major limitation: their runtimes scale cubically with the RNA length, and this slowness limits their use in genome-wide applications. RESULTS: We present a novel alternative O(n3)-time dynamic programming algorithm for RNA folding that is amenable to heuristics that make it run in O(n) time and O(n) space, while producing a high-quality approximation to the optimal solution. Inspired by incremental parsing for context-free grammars in computational linguistics, our alternative dynamic programming algorithm scans the sequence in a left-to-right (5'-to-3') direction rather than in a bottom-up fashion, which allows us to employ the effective beam pruning heuristic. Our work, though inexact, is the first RNA folding algorithm to achieve linear runtime (and linear space) without imposing constraints on the output structure. Surprisingly, our approximate search results in even higher overall accuracy on a diverse database of sequences with known structures. More interestingly, it leads to significantly more accurate predictions on the longest sequence families in that database (16S and 23S Ribosomal RNAs), as well as improved accuracies for long-range base pairs (500+ nucleotides apart), both of which are well known to be challenging for the current models. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/LinearFold/LinearFold, and our webserver is at http://linearfold.org (sequence limit: 100 000nt). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liang Huang 0001, He Zhang 0025, Dezhong Deng, Kai Zhao 0003, Kaibo Liu, David A. Hendrix, David H. Mathews
Bioinform.5
2019 Structural Degradation Modeling Framework for Sparse Data Sets With an Application on Alzheimer's Disease
abstract
The rapid development of information technologies provided unprecedented big data environments for condition monitoring and degradation analyses. However, the available big data sets are often sparse with a limited number of observations per recorded unit. For example, in many healthcare systems, data are collected from a large number of patients, but the available observations from each patient are quite limited. Unfortunately, most of the existing approaches for data-driven degradation modeling may not work well in this scenario as they either pool the information from the population or require rich historical observations in each unit. To address the challenges in “sparse data environments,” this paper proposes a structural degradation modeling framework (SDM). The SDM is inspired by the recommender system, which provides recommendations about specific items for the user. In addition, it is also tailored to the needs of degradation modeling. In particular, the framework takes into consideration: 1) the available data from the unit of interest; 2) the population characteristics; 3) the relationship between the available units; and 4) the precision of the available units. Simulation studies and a case study that involves the Alzheimer's disease (AD) neuroimaging initiative data set are conducted, which shows satisfactory performance of the proposed method.
Abdallah A. Chehade, Kaibo Liu
IEEE Trans Autom. Sci. Eng.2
2019 Dynamic Inspection of Latent Variables in State-Space Systems
abstract
The state-space models (SSMs) are widely used in a variety of areas where a set of observable variables are used to track some latent variables. While most existing works focus on the statistical modeling of the relationship between the latent variables and observable variables or statistical inferences of the latent variables based on the observable variables, it comes to our awareness that an important problem has been largely neglected. In many applications, although the latent variables cannot be routinely acquired, they can be occasionally acquired to enhance the monitoring of the state-space system. Therefore, in this paper, novel dynamic inspection (DI) methods under a general framework of SSMs are developed to identify and inspect the latent variables that are most uncertain. Extensive numeric studies are conducted to demonstrate the effectiveness of the proposed methods.
Tianshu Feng, Xiaoning Qian, Kaibo Liu, Shuai Huang 0001
IEEE Trans Autom. Sci. Eng.3
2019 A Generic Health Index Approach for Multisensor Degradation Modeling and Sensor Selection
abstract
With recent development in sensor technology, multiple sensors have been widely adopted to monitor the degradation of a single unit simultaneously. The challenge of multisensor degradation modeling lies in that the sensor signals are often correlated and may contain only partial or even no information on the degradation status of a unit. To address these issues, this paper proposes a novel data fusion method that constructs a 1-D health index (HI) via automatically selecting and combining multiple sensor signals to better characterize the degradation process. In particular, this paper develops a new latent linear model that constructs the HI and selects informative sensors in a unified manner. Compared to the existing literature, the proposed method enjoys several unique advantages: 1) being able to derive the best linear unbiased estimator of the fusion coefficients; 2) offering high computational efficiency; 3) not requiring to know the exact value of the failure threshold; and 4) exhibiting general applicability in practice by not imposing restrictive assumptions on the degradation process. Simulation studies are presented to illustrate the effectiveness and evaluate the sensitivity of the proposed method. A case study on the degradation of aircraft gas turbine engines is also performed which shows a better prognostic performance of the proposed method compared with existing approaches.
Changyue Song, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2019 Causation-Based Monitoring and Diagnosis for Multivariate Categorical Processes With Ordinal Information
abstract
The monitoring and diagnosis of multivariate categorical processes (MCPs) have drawn increasing attention lately, as categorical variables have been frequently involved in modern quality control applications. In these applications, there may exist causal relationships among multiple categorical variables, where the attribute level of a cause variable influences that of its effect variable. In such a case, shifts occurring in a cause variable will propagate to its effect variable based on the causal structure. Furthermore, there usually exists natural order among the attribute levels of some categorical variables such as good, neutral, and bad for measuring the product quality. By assuming a latent continuous variable, the attribute levels of an ordinal categorical variable can be determined by classifying the value of the latent variable based on thresholds. In this paper, we leverage Bayesian networks (BNs) to characterize MCPs with a causal structure, where the categorical variables can be either nominal, ordinal or a combination of both. We develop one general control chart and one directional control chart, both of which fully exploit the causal relationships and the ordinal information for better process monitoring and diagnosis. Numerical simulations have demonstrated the superiority and robustness of our method in detecting and diagnosing the conditional probability shifts of nominal factors as well as the conditional latent location shifts of ordinal factors. Note to Practitioners-This paper aims at addressing the challenges in monitoring and diagnosing MCPs when there are causal relationships among the categorical variables. The developed method is greatly beneficial, especially when there are nominal and ordinal variables involved in the MCP. Specifically, a BN is employed to characterize the dependence structure of the variables involved in the process, and a latent continuous variable is utilized to model the orders of attribute levels of the ordinal variables. Then, a novel method is proposed to detect the probability shift in the nominal factors and the location shift on the latent variables of the ordinal variables based on the likelihood ratio test. A general monitoring control chart as well as a directional version which also facilitates diagnosis is proposed. In this paper, although our method is based on the assumption that the latent continuous variables of the ordinal variables follow logistic distributions, the proposed charts are demonstrated to perform efficiently and robustly in various cases as shown in the simulations and case studies.
Xiaochen Xian, Jian Li 0023, Kaibo Liu
IEEE Trans Autom. Sci. Eng.3
2018 A Collaborative Learning Framework for Estimating Many Individualized Regression Models in a Heterogeneous Population
abstract
Mixed-effect models (MEMs) have been found very useful for modeling complex dataset where many similar individualized regression models should be estimated. Like many statistical models, the success of these models builds on the assumption that a central tendency can effectively establish the population-level characteristics and covariates are sufficient to characterize the individual variation as derivation from the center. In many real-world problems, however, the dataset is collected from a rather heterogeneous population, where each individual has a distinct model. To fill in this gap, we propose a collaborative learning framework that provides a generic methodology for estimating a heterogeneous population of individualized regression models by exploiting the idea of “canonical models” and model regularization. By using a set of canonical models to represent the heterogeneous population characteristics, the canonical models span the modeling space for the individuals' models, e.g., although each individual model is distinct, its model parameter vector can be represented by the parameter vectors of the canonical models. Theoretical analysis is also conducted to reveal a connection between the proposed method and the MEMs. Both simulation studies and applications on Alzheimer's disease and degradation modeling of turbofan engines demonstrate the efficacy of the proposed method.
Kaibo Liu, Eunshin Byon, Xiaoning Qian, Shuai Huang 0001
IEEE Trans. Reliab.2
2018 Integration of Data-Level Fusion Model and Kernel Methods for Degradation Modeling and Prognostic Analysis
abstract
To prevent unexpected failures of complex engineering systems, multiple sensors have been widely used to simultaneously monitor the degradation process and make inference about the remaining useful life in real time. As each of the sensor signals often contains partial and dependent information, data-level fusion techniques have been developed that aim to construct a health index via the combination of multiple sensor signals. While the existing data-level fusion approaches have shown a promise for degradation modeling and prognostics, they are limited by only considering a linear fusion function. Such a linear assumption is usually insufficient to accurately characterize the complicated relations between multiple sensor signals and the underlying degradation process in practice, especially for complex engineering systems considered in this study. To address this issue, this study fills the literature gap by integrating kernel methods into the data-level fusion approaches to construct a health index for better characterizing the degradation process of the system. Through selecting a proper kernel function, the nonlinear relation between multiple sensor signals and the underlying degradation process can be captured. As a result, the constructed health index is expected to perform better in prognosis than existing data-level fusion methods that are based on the linear assumption. In fact, the existing data-level fusion models turn out to be only a special case of the proposed method. A case study based on the degradation signals of aircraft gas turbine engines is conducted and finally shows the developed health index by using the proposed method is insensitive for missing data and leads to an improved prognostic performance.
Changyue Song, Kaibo Liu, Xi Zhang 0006
IEEE Trans. Reliab.2
2017 Controlling the Residual Life Distribution of Parallel Unit Systems Through Workload Adjustment
abstract
Complex systems often consist of multiple units that are required to work together in parallel to satisfy a specific engineering objective. As an example, in manufacturing processes, several identical machines may need to operate together to simultaneously fabricate the same products in order to meet the high production demand. This parallel configuration is often designed with some level of redundancy to compensate for unexpected events. In this way, when only a small portion of units fail to operate due to either unexpected machine downtime or scheduled maintenance, the remaining units can still achieve the engineering objective by increasing their workloads up to the designed capacities. However, the workload of a unit apparently impacts the unit's degradation rate as well as its failure time. Specifically, this paper considers the case that a higher workload assignment accelerates the unit's degradation and vice versa. Based on this assumption, we develop a method to actively control the degradation as well as the predicted failure time of each unit by dynamically adjusting its workloads. Our goal is to prevent the overlap of unit failures within a certain time period through taking advantage of the natural redundancy of the parallel structure, which may potentially lead to a better utilization of maintenance resources as well as a consistently ensured system throughput. A numerical study is used to evaluate the performance of the proposed method under different scenarios.
Kaibo Liu, Nagi Gebraeel, Jianjun Shi 0001
IEEE Trans Autom. Sci. Eng.2
2017 Optimize the Signal Quality of the Composite Health Index via Data Fusion for Degradation Modeling and Prognostic Analysis
abstract
Due to the rapid development of sensing and computing technologies, multiple sensors have been widely used in a system to simultaneously monitor the health status of an operating unit. Such a data-rich environment creates an unprecedented opportunity to better understand the degradation behavior of the system and make accurate inferences about the remaining lifetime. Since data collected from multiple sensors are often correlated and each sensor data contains only partial information about the degraded unit, data fusion methodologies that integrate the data from multiple sensors provide an essential tool for degradation modeling and prognostics. To achieve this goal, a fundamental question needs to be answered first is how to measure the signal quality of a degradation signal. If such a question can be addressed, then the data fusion approach can be simplified as a mission-specific task: to construct a composite health index with the goal of optimizing its signal quality. In this paper, a new signal-to-noise ratio (SNR) metric that is tailored to the needs of degradation signals is proposed. Then, based on the new quality metric, we develop a data-level fusion model to construct a health index via fusion of multiple degradation-based sensor data. Our goal is that the developed health index provides a much better characterization of the health condition of the unit and thus leads to a better prediction of the remaining lifetime. A case study that involves the degradation dataset of aircraft gas turbine engines is conducted to numerically evaluate the performance of the developed health index regarding prognostics and further compare the result with existing literature.
Kaibo Liu, Abdallah A. Chehade, Changyue Song
IEEE Trans Autom. Sci. Eng.1
2017 Sensory-Based Failure Threshold Estimation for Remaining Useful Life Prediction
abstract
The rapid development of sensor and computing technology has created an unprecedented opportunity for condition monitoring and prognostic analysis in various manufacturing and healthcare industries. With the massive amount of sensor information available, important research efforts have been made in modeling the degradation signals of a unit and estimating its remaining useful life distribution. In particular, a unit is often considered to have failed when its degradation signal crosses a predefined failure threshold, which is assumed to be known a priori. Unfortunately, such a simplified assumption may not be valid in many applications given the stochastic nature of the underlying degradation mechanism. While there are some extended studies considering the variability in the estimated failure threshold via data-driven approaches, they focus on the failure threshold distribution of the population instead of that of an individual unit. Currently, the existing literature still lacks an effective approach to accurately estimate the failure threshold distribution of an operating unit based on its in-situ sensory data during condition monitoring. To fill this literature gap, this paper develops a convex quadratic formulation that combines the information from the degradation profiles of historical units and the in-situ sensory data from an operating unit to online estimate the failure threshold of this particular unit in the field. With a more accurate estimation of the failure threshold of the operating unit in real time, a better remaining useful life prediction is expected. Simulations as well as a case study involving a degradation dataset of aircraft turbine engines were used to numerically evaluate and compare the performance of the proposed methodology with the existing literature in the context of failure threshold estimation and remaining useful life prediction.
Abdallah A. Chehade, Scott Bonk, Kaibo Liu
IEEE Trans. Reliab.3
2016 Integration of Data Fusion Methodology and Degradation Modeling Process to Improve Prognostics
abstract
The rapid development of sensing and computing technologies has enabled multiple sensors embedded in a system to simultaneously monitor the degradation status of an operation unit. This creates a data-rich environment for degradation modeling and prognostics that could potentially lead to an accurate inference about the remaining lifetime of the degraded unit. However, as data collected from multiple sensors are often correlated and each sensor data contains only partial information about the same degradation process, there is a pressing need to develop data fusion methodologies that can integrate the data from multiple sensors for better characterizing the stochastic nature of the degradation process. Unlike other existing data fusion methodologies that treat the fusion procedure and the degradation modeling as two separate tasks, this paper aims at solving these two challenging problems in a unified manner. Specifically, we develop a methodology to construct a health index via fusion of multiple degradation-based sensor data. Our goal is that the developed health index provides a much better characterization of the condition of the unit and thus leads to a better prediction of the remaining lifetime. A case study that involves a degradation dataset of an aircraft gas turbine engine is implemented to numerically evaluate and compare the prognostic performance of the developed health index with existing literature.
Kaibo Liu, Shuai Huang 0001
IEEE Trans Autom. Sci. Eng.1
2016 An Automatic Process Monitoring Method Using Recurrence Plot in Progressive Stamping Processes
abstract
In progressive stamping processes, condition monitoring based on tonnage signals is of great practical significance. One typical fault in progressive stamping processes is a missing part in one of the die stations due to malfunction of part transfer in the press. One challenging question is how to detect the fault due to the missing part in certain die stations as such a fault often results in die or press damage, but only provides a small change in the tonnage signals. To address this issue, this article proposes a novel automatic process monitoring method using the recurrence plot (RP) method. Along with the developed method, we also provide a detailed interpretation of the representative patterns in the recurrence plot. Then, the corresponding relationship between the RPs and the tonnage signals under different process conditions is fully investigated. To differentiate the tonnage signals under normal and faulty conditions, we adopt the recurrence quantification analysis (RQA) to characterize the critical patterns in the RPs. A parameter learning algorithm is developed to set up the appropriate parameter of the RP method for progressive stamping processes. A real case study is provided to validate our approach, and the results are compared with the existing literature to demonstrate the outperformance of this proposed monitoring method.
Kaibo Liu, Xi Zhang 0006, Jianjun Shi 0001
IEEE Trans Autom. Sci. Eng.2
2016 Multiple Sensor Data Fusion for Degradation Modeling and Prognostics Under Multiple Operational Conditions
abstract
Due to the rapid advances in sensing and computing technology, multiple sensors have been widely used to simultaneously monitor the health status of an operation unit. This creates a data-rich environment, enabling an unprecedented opportunity to make better understanding and inference about the current and future behavior of the unit in real time. Depending on specific task requirements, a unit is often required to run under multiple operational conditions, each of which may affect the degradation path of the unit differently. Thus, two fundamental challenges remain to be solved for effective degradation modeling and prognostic analysis: 1) how to leverage the dependent information among multiple sensor signals to better understand the health condition of the unit; and 2) how to model the effects of multiple conditions on the degradation characteristics of the unit. To address these two issues, this paper develops a data fusion methodology that integrates the information from multiple sensors to construct a health index when the monitored unit runs under multiple operational conditions. Our goal is that the developed health index provides a much better characterization of the health condition of the degraded unit, and, thus, leads to a better prediction of the remaining lifetime. Unlike other existing approaches, the developed data fusion model combines the fusion procedure and the degradation modeling under different operational conditions in a unified manner. The effectiveness of the proposed method is demonstrated in a case study, which involves a degradation dataset of aircraft gas turbine engines collected from 21 sensors under six different operational conditions.
Kaibo Liu, Xi Zhang 0006, Jianjun Shi 0001
IEEE Trans. Reliab.2
2015 Domain-Knowledge Driven Cognitive Degradation Modeling for Alzheimer's Disease
abstract
Cognitive monitoring and screening holds great promises for early detection and intervention of Alzheimer's Disease (AD). A critical enabler is the personalized degradation model to predict the cognitive status over time. However, estimating such a model individual's data faces challenges due to the sparsity and fragmented nature of the cognitive data of each individual. To mitigate this problem, we propose novel methods, called the collaborative degradation model (CDM) together with its extended network regularized version, the NCDM, which can incorporate useful domain knowledge into the degradation modeling. While NCDM results in a difficult optimization problem, we are inspired by existing non-negative matrix factorization methods and develop an efficient algorithm to solve this problem and further provide theoretical results that ensure that the proposed algorithm can guarantee non-increasing property. Both simulation studies and the real-world application to AD are conducted across different degradation models and sampling schemes, which demonstrate the superiority of the proposed methods over existing methods.
Kaibo Liu, Eunshin Byon, Xiaoning Qian, Shuai Huang 0001
SDM2
2014 Adaptive Sensor Allocation Strategy for Process Monitoring and Diagnosis in a Bayesian Network
abstract
Multivariate process control in Distributed Sensor Networks (DSNs) is an important and challenging topic. Although a fully deployed sensor network will minimize information loss, the associated sensing cost can be overwhelming. Many efforts have been made to investigate the optimal sensor allocation strategy for different process control applications; however, most of them assume that the sensor layout is fixed once sensors are deployed in the system. This paper proposes a novel approach to adaptively reallocate sensor resources based on online observations, which can enhance both monitoring and diagnosis capabilities. The proposed adaptive sensor allocation strategy addresses two fundamental issues: when to reallocate sensors and how to update sensor layout. A max-min criterion is developed to manage sensor reallocation and process change detection in an integrated manner. To investigate the adaptive strategy, a Bayesian Network (BN) model is assumed available to represent the causal relationships among a set of variables. Case studies are performed on a hot forming process and a cap alignment process to illustrate the procedure and evaluate the performance of the proposed method under different fault scenarios.
Kaibo Liu, Xi Zhang 0006, Jianjun Shi 0001
IEEE Trans Autom. Sci. Eng.1
2013 A Data-Level Fusion Model for Developing Composite Health Indices for Degradation Modeling and Prognostic Analysis
abstract
Prognostics involves the effective utilization of condition or performance-based sensor signals to accurately estimate the remaining lifetime of partially degraded systems and components. The rapid development of sensor technology, has led to the use of multiple sensors to monitor the condition of an engineering system. It is therefore important to develop methodologies capable of integrating data from multiple sensors with the goal of improving the accuracy of predicting remaining lifetime. Although numerous efforts have focused on developing feature-level and decision-level fusion methodologies for prognostics, little research has targeted the development of “data-level” fusion models. In this paper, we present a methodology for constructing a composite health index for characterizing the performance of a system through the fusion of multiple degradation-based sensor data. This methodology includes data selection, data processing, and data fusion steps that lead to an improved degradation-based prognostic model. Our goal is that the composite health index provides a much better characterization of the condition of a system compared to relying solely on data from an individual sensor. Our methodology was evaluated through a case study involving a degradation dataset of an aircraft gas turbine engine that was generated by the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS).
Kaibo Liu, Nagi Gebraeel, Jianjun Shi 0001
IEEE Trans Autom. Sci. Eng.1