VLDB 2026 Research / reviewers in the wild / expert
Kwok-Leung Tsui
dblp:06/5615 · also Kwok Leung Tsui
· DBLP profile ↗
56ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0002-0558-2279ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 19 · 8 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 2Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIEF-Net: multimodal image-enhanced fusion network for intelligent fall risk prediction
Qizheng Zhao, Ruiyuan Wu, Manting Chen, Kwok-Leung Tsui, Yang Zhao 0009 |
Neural Networks | 4 |
| 2026 | New Look at Bayesian Prognostic MethodsabstractOnline remaining useful life (RUL) prediction is a core function of prognostics and health management (PHM), which provides solutions for comprehensive and personalised system management. RUL is realised by extrapolating timely updated prognostic models to reach a user-defined failure threshold. As of today, there are mainly two kinds of Bayesian prognostic methods. The first kind of Bayesian prognostic methods are Bayesian regression prognostic methods that directly use Bayes’ theorem to update degradation model parameters. The second kind of Bayesian prognostic methods is Bayesian state-space prognostic methods that firstly reformulate a degradation model by using state-space representation and consequently update state-space model parameters with Bayes’ theorem. However, comparisons of these two kinds of Bayesian prognostic methods have not been actively explored and discussed in a unified paper. In this study, similarities and differences between Bayesian regression prognostic methods and Bayesian state-space prognostic methods under the assumptions of additive Gaussian and Brownian motion errors were explored to enrich the PHM domain. A significant difference was observed between Bayesian regression prognostic methods and Bayesian state-space prognostic methods under different error assumptions. Experimental results showed that Bayesian state-space prognostic methods have greater RUL prediction uncertainties in two run-to-failure cases with different fluctuation strengths. However, under the assumption of Gaussian errors, because of the good ability of degradation tracking, Bayesian state-space prognostic methods may predict worse than Bayesian regression prognostic methods in a strong fluctuation dataset, which is not evident in the situation of Brownian motion errors.Note to Practitioners—The Bayesian update of model parameters considering online condition monitoring data is of great practical significance for describing individual degradation and RUL prediction. Practitioners need to know the similarities and differences between different Bayesian prognostic methods, as well as how to choose an appropriate Bayesian prognostic method for a specific predicted objective that practitioners care about. This paper introduces the similarities and differences between several classic Bayesian prognostic methods, and provides their performance comparisons in different RUL prediction scenarios, providing a reference for practitioners to use Bayesian model parameters updating to predict individual RUL. Jie Liu 0057, Dong Wang 0001, Jin-Zhen Kong, Naipeng Li, Zhike Peng, Kwok-Leung Tsui |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Sensor-Based Digital Biomarkers for Early Identification of Cognitive Frailty: A Systematic ReviewabstractCognitive Frailty (CF), characterized by the co-occurrence of Physical Frailty (PF) and Mild Cognitive Impairment (MCI), is increasingly recognized as a critical predictor of adverse health outcomes in aging populations. Despite its clinical significance, early identification of CF remains challenging due to heterogeneous diagnostic criteria and reliance on resource-intensive assessments. This systematic review synthesizes current evidence on sensor-based digital biomarkers for CF detection across eight databases following PRISMA 2020 guidelines. A combined qualitative and quantitative synthesis is conducted based on 20 eligible studies. Key findings reveal that Inertial Measurement Units (IMUs) are the most frequently utilized sensors (n = 6 studies), primarily capturing gait and motor function parameters under various functional tasks. Among digital biomarkers, gait features, particularly Stride Length under single-task conditions and Velocity under dual-task or single-task conditions, show significant associations with CF. Body composition metrics, particularly Appendicular Skeletal Muscle Mass Index (ASM), and physical activity parameters, such as Moderate-to-Vigorous-intensity Physical Activity (MVPA), show consistent negative associations with CF, while cardiovascular indices like Cardio-Ankle Vascular Index (CAVI) reveal positive correlations. Eighteen studies develop a total of 23 CF-related models, of which only four focus on predictive classification. These four studies report ten distinct models, with AUCs ranging from 0.696 to 0.911 and accuracy from 0.480 to 0.863. A quantitative synthesis of six models from three studies yields pooled estimates of sensitivity at 0.81 (95% CI: 0.55-0.94) and specificity at 0.80 (95% CI: 0.49-0.95), suggesting moderate-to-good discriminative performance despite notable heterogeneity in diagnostic definitions, sample sizes, sensor types, and modeling approaches. These results systematically map the evolving landscape of digital biomarkers for CF, emphasizing the need for standardized sensor protocols and rigorous validation to advance clinical translation. Ruiyuan Wu, Keyi Huang, Kwok-Leung Tsui, Jianbang Xiang, Yang Zhao 0009 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Lifeisgood: Learning Invariant Features via In-Label Swapping for Generalizing Out-of-Distribution in Machine Fault DiagnosisabstractIn machine fault diagnosis, conventional data-driven models trained by empirical risk minimization (ERM) often fail to generalize across domains with distinct data distributions caused by various machine operating conditions. One major reason is that ERM primarily focuses on informativeness of data labels and lacks sufficient attention on invariance of data features. To enable invariance on top of informativeness, a learning framework, learning invariant features via in-label swapping for generalizing out-of-distribution (Lifeisgood), is proposed in this study. Lifeisgood is inspired by a simple intuition that invariance can be assessed by checking changes in loss due to swapping certain entries of features with the same labels. Lifeisgood also enjoys a theoretical guarantee on improving testing domain performance under certain conditions based on a swapping 0-1 loss proposed in this work. To circumvent the training difficulties associated with the swapping 0-1 loss, a swapping cross-entropy loss is derived as a surrogate and theoretical justifications for such a relaxation are also provided. As a result, Lifeisgood can be employed conveniently to develop data-driven fault diagnosis models. In the experiments, Lifeisgood outperformed the majority of state-of-the-art methods in terms of average accuracy and exceeded the second-best by 25% in terms of the frequency of beating the generic ERM. The code is available at: https://github.com/mozhenling/doge-lifeisgood. Zhenling Mo, Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Trans. Cybern. | 3 |
| 2025 | Domain Generalization Study of Empirical Risk Minimization From Causal PerspectivesabstractEmpirical risk minimization (ERM) is a celebrated induction principle for developing data-driven models. However, ERM has received both pros and cons for its capability on domain generalization (DG). To this end, this paper attempts to study the success and failure of ERM at supervised DG classification tasks, both theoretically and empirically, with causal perspectives. In the theoretical aspect, we first explore different properties of a causal metric termed information flow, followed with discussing relationships between the information flow and the mutual information in the proposed causal graph. Next, we analyze the roles of the transformed causal feature and the transformed spurious feature on modeling performances. It reveals that the interaction between the spurious influencer and the transformed causal feature is the key determining the failure or success of ERM on DG. In the empirical study, we first simulate various DG settings based on the MNIST, Fashion MNIST, and CIFAR10 datasets. Next, we verify developed theories by testing three different neural network configurations in designed experiments. In addition, experiments based on real-world datasets are conducted to further consolidate key points of the proposed theories. To extend application benefits of the theoretical discoveries, a new risk minimization framework with a novel feature intervention for regulating ERM is proposed. It achieves DG improvements over ERM on real-world datasets of image segmentation, image classification, and text classification. Zhenling Mo, Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Trans. Multim. | 3 |
| 2025 | Extended Invariant Risk Minimization for Machine Fault Diagnosis With Label Noise and Data ShiftabstractIncorrect labels as well as the discrepancy between training and test domain data distributions can significantly affect the effectiveness of supervised data-driven models in machine fault diagnosis applications. Such a challenge can be characterized as the noisy label-domain generalization (NL-DG) problem. In this article, the extended invariant risk minimization (EIRM) is developed, which incorporates flat minima seeking to address the NL-DG challenge. The ability of handling NL-DG is realized by shifting the gradient penalty base from the dummy classifier to the entire model. EIRM is shown to be closely related to locating a flat minimum, which is crucial for label noise (LN) robustness and model generalization. Explorations on function smoothness and algorithm convergence are offered to understand EIRM from the theoretical aspect. An efficient implementation of EIRM is also developed to construct the fault diagnosis model. The EIRM-based fault diagnosis method is compared with strong benchmarks on multiple NL-DG tasks using actuator and gearbox fault datasets. Results indicate that the EIRM-based method on average is more effective than the benchmarks. The code is available at https://github.com/mozhenling/doge-eirm. Zhenling Mo, Zijun Zhang 0001, Qiang Miao, Kwok-Leung Tsui |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Distance-Aware Risk Minimization for Domain Generalization in Machine Fault DiagnosisabstractIndustrial Internet of Things (IIoT) connects machines, and it is important to build intelligent models to prevent machine failures by identifying incipient faults. To develop intelligent fault diagnosis models, empirical risk minimization (ERM)-based modeling paradigm has been prevalently applied. However, during model training, ERM primarily focuses on instance-to-prototype (ItP) distances from a prototypical perspective, which may limit its effectiveness in analyzing data of diverse distributions. To improve the ERM model, we propose considering additional instance-to-instance distances (ItI) and prototype-to-prototype (PtP) distances, leading to a new modeling framework—distance-aware risk minimization (DARM). To gain awareness of extra types of distances, two novel losses are proposed based on reformulations of soft-max cross entropy. Theoretical explorations are conducted to justify the significance of collectively considering ItP, ItI, and PtP distances. Methodologically, DARM can jointly minimize three types of distance-aware losses to train neural networks for fault diagnosis in the same fashion as ERM. In a comprehensive computational study, DARM consistently outperformed ERM in domain generalization (DG) tasks based on various machine fault diagnosis data sets. In addition, DARM has superior performance over several recent DG methods. The code is available athttps://github.com/mozhenling/doge-darm. Zhenling Mo, Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Internet Things J. | 3 |
| 2024 | PFFN: Periodic Feature-Folding Deep Neural Network for Traffic Condition ForecastingabstractAccurate forecasting of traffic conditions is critical for improving urban transportation safety, stability, and efficiency. It is challenging to produce explicit traffic forecasts due to complex and dynamic spatiotemporal contexts. Most existing works only capture partial characteristics and features of traffic data, and there still exists a great potential of uplifting the forecasting performances. In this paper, we propose a periodic feature-folding network (PFFN) that globally folds the spatial-temporal-geo-character features of collected data into pivotal levels to enhance the forecasting performance of traffic conditions. Meanwhile, the historical and recent traffic information within the subgraph is locally folded to capture similar traffic patterns and determine meaningful high-level features in latent space via a novel gated attention plug (GAP). During forecasting, the auxiliary road attributes are joint-wise folded with locally engineered features to realize the multi-step forecasting. Experimental results on two publicly accessible real-world urban traffic datasets show that the proposed PFFN can outperform the state-of-the-art benchmarks, significantly improving the performance of forecasting short-term traffic conditions. Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Internet Things J. | 3 |
| 2024 | A Hierarchical Transfer-Generative Framework for Automating Multianalytical Tasks in Rail Surface Defect InspectionabstractRail surface inspection is crucial for ensuring the safety and longevity of rail transport systems, grapples with the challenges posed by the scarcity of defective samples. Additionally, contemporary techniques in this domain typically fail to concurrently identify and localize defects at both image level and pixel levels. Addressing these intricacies, we present a hierarchical transfer-generative framework, the HTg-Net. This innovative framework is geared towards the automation and enhancement of multi-analytical tasks in rail surface inspection. The HTg-Net architecture synergistically melds two pivotal subnetworks: (1) the memory-guided generation subnetwork (MGN), which is endowed with a cutting-edge memory mechanism. This mechanism adeptly captures and recalls the typical patterns observed in rail images, facilitating the detection of anomalies or deviations; (2) the attention-focused segmentation subnetwork (ASN) is fortified with a gated attention mechanism and hierarchical weights transferred from MGN, enabling parallel feature extraction and enhancing defect localization. Rigorous evaluations of HTg-Net on three datasets elucidate its superior efficiency and performance over prevailing benchmarks, positioning it as an advanced solution for the comprehensive inspection of rail surface defects. Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Internet Things J. | 3 |
| 2024 | MSS-Former: Multiscale Skeletal Transformer for Intelligent Fall Risk Prediction in Older AdultsabstractFall, a leading cause of accidental death and injury in older adults aged 65 and above, has become a rapidly growing health concern in aging populations worldwide. Data-driven methods integrating depth imaging technology have received growing attention in automated fall risk assessment owing to their noninvasiveness and less dependence on healthcare professionals. However, most existing depth image data-based models neglect the inherent physiological and potential functional connections and lack sufficient real-world data validation. To fill the research gap, we developed a novel approach named multiscale skeletal transformer (MSS-Former), leveraging depth image technology and deep-learning models for effective fall risk prediction. Our contributions mainly consist of four parts. First, we introduced a multimodel output feature fusion transformer in fall risk prediction, enabling output merging and weighting from multiple model streams dynamically. Second, we developed an innovative scheme to construct interjoint skeletal topology, systematically focusing on joints’ intrinsic physiological and potential functional connections. Third, we constructed a ResNet-FPN, greatly enhancing multiscale feature extraction capabilities. Fourth, we conducted a field study in a local hospital and performed a comprehensive validation of our developed approach. The comparison results show that our approach achieved outstanding predictive performance, surpassing state-of-the-art methods on the real-world data set, with accuracy, precision, recall, and F1 scores of 97.84%, 97.33%, 96.97%, and 96.92%, respectively. In practice, the proposed approach would be of great value in the timely identification for individuals at high fall risk and facilitate decision making to take appropriate interventions. Qizheng Zhao, Xiaomao Fan, Manting Chen, Yutian Xiao, Eric Hiu Kwong Yeung, Kwok-Leung Tsui, Yang Zhao 0009 |
IEEE Internet Things J. | 7 |
| 2024 | Optimal Composite Likelihood Estimation and Prediction for Distributed Gaussian Process ModelingabstractLarge-scale Gaussian process (GP) modeling is becoming increasingly important in machine learning. However, the standard modeling method of GPs, which uses the maximum likelihood method and the best linear unbiased predictor, is designed to run on a single computer, which often has limited computing power. Therefore, there is a growing demand for approximate alternatives, such as composite likelihood methods, that can take advantage of the power of multiple computers. However, these alternative methods in the literature offer limited options for practitioners because most methods focus more on computational efficiency rather than statistical efficiency. Limited accurate solutions to the parameter estimation and prediction for fast GP modeling are available in the literature for supercomputing practitioners. Therefore, this study develops an optimal composite likelihood (OCL) scheme for distributed GP modeling that can minimize information loss in parameter estimation and model prediction. The proposed predictor, called the best linear unbiased block predictor (BLUBP), has the minimum prediction variance given the partitioned data. Numerical examples illustrate that both the proposed composite likelihood estimation and prediction methods provide more accurate performance than their traditional counterparts under various cases, and an extremely close approximation to the standard modeling method is observed. Qiang Zhou 0002, Kwok-Leung Tsui |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Sparsity-Constrained Invariant Risk Minimization for Domain Generalization With Application to Machinery Fault Diagnosis ModelingabstractMachine learning has been widely applied to study AI-informed machinery fault diagnosis. This work proposes a sparsity-constrained invariant risk minimization (SCIRM) framework, which develops machine-learning models with better generalization capacities for environmental disturbances in machinery fault diagnosis. The SCIRM is built by innovating the optimization formulation of the recently proposed invariant risk minimization (IRM) and its variants through the integration of sparsity constraints. We prove that if a sparsity measure is differentiable, scale invariant, and semistrictly quasi-convex, the SCIRM can be guaranteed to solve the domain generalization problem based on a few predefined problem settings. We mathematically derive a family of such sparsity measures. A practical process of implementing the SCIRM for machinery fault diagnosis tasks is offered. We first verify our theoretical exploration of the SCIRM by using simulation data. We further compare SCIRM with a set of state-of-the-art methods by using real machinery fault data collected under a variety of working conditions. The computational results confirm that the machinery fault diagnosis model developed by the SCIRM offers a higher generalization capacity and performs better than the other benchmarks across the different testing datasets. Zhenling Mo, Zijun Zhang 0001, Qiang Miao, Kwok-Leung Tsui |
IEEE Trans. Cybern. | 4 |
| 2024 | Sensor-Based Multifaceted Feature Extraction and Ensemble Elastic Net Approach for Assessing Fall Risk in Community-Dwelling Older AdultsabstractAccurate identification of community-dwelling older adults at high fall risk can facilitate timely intervention and significantly reduce fall incidents. Analyzing gait and balance capabilities via feature extraction and modeling through sensor-based motion data has emerged as a viable approach for fall risk assessment. However, the existing approaches for extracting key features related to fall risk lack inclusiveness, with limited consideration of the non-linear characteristics of sensor signals, such as signal complexity, self-similarity, and local stability. In this study, we developed a multifaceted feature extraction scheme employing diverse feature types, including demographic, descriptive statistical, non-linear, spatiotemporal and spectral features, derived from three-axis accelerometers and gyroscope data. This study is the first attempt to investigate non-linear features related to fall risk in multi-task scenarios from a dynamic system perspective. Based on the extracted multifaceted features, we propose an ensemble elastic net (E-E-N) approach for handling imbalanced data and offering high model interpretability. The E-E-N utilizes bootstrap sampling to construct base classifiers and employs a weighting mechanism to aggregate the base classifiers. We conducted a set of validation experiments using real-world data for comprehensive comparative analysis. The results demonstrate that the E-E-N approach exhibits superior predictive performance on fall risk classification. Our proposed approach offers a cost-effective tool for accurately assessing fall risk and alleviating the burden of continuous health monitoring in the long term. Lisha Yu, Kwok-Leung Tsui, Yang Zhao 0009 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | A Deep Generative Approach for Rail Foreign Object Detections via Semisupervised LearningabstractThe automated inspection and detection of foreign objects help prevent potential accidents and train derailments. Most existing approaches focus on the detection with prior labels, such as categories and locations of objects, and do not directly address detecting foreign objects of unknown categories, which can appear anytime on the rail track site. In this article, we develop a deep generative approach for detecting foreign objects without predefining the scope of objects. The detection procedure consists of the following three steps: first, the model composed of an autoencoder and a discriminator is developed via adversarial training based on normal rail images only; second, the detection of abnormal rail images is implemented based on the anomaly score obtained via the trained autoencoder; and finally, foreign objects are detected by filtering the subtle dissimilarity in normal areas and highlighting abnormal areas. The effectiveness of the proposed framework for the rail foreign object detection is validated with images collected by a train equipped with visual sensors. Computational results demonstrate that our proposal is capable to achieve an impressive performance on detecting numerous foreign objects. Moreover, two groups of benchmarking methods are employed to verify the superiority of the proposed framework. Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | An Indoor Fall Monitoring System: Robust, Multistatic Radar Sensing and Explainable, Feature-Resonated Deep Neural NetworkabstractIndoor fall monitoring is challenging for community-dwelling older adults due to the need for high accuracy and privacy concerns. Doppler radar is promising, given its low-cost and contactless sensing mechanism. However, the line-of-sight restriction limits the application of radar sensing in practice, as the Doppler signature will vary when the sensing angle changes, and signal strength will substantially degrade with large aspect angles. Additionally, the similarity of the Doppler signatures among different fall types makes classification challenging. To address these problems, we first present an experimental study to obtain Doppler signals under large and arbitrary aspect angles for diverse types of simulated activities. We then develop a novel, explainable, multi-stream, feature-resonated neural network (eMSFRNet) that achieves fall detection and a pioneering study of classifying seven fall types. eMSFRNet is robust to radar sensing angles and subjects, and is the first method that can resonate and enhance feature information from noisy/weak Doppler signatures. The multiple feature extractors - from ResNet, DenseNet, and VGGNet - extract diverse feature information with various spatial abstractions from a pair of Doppler signals. The resonated-fusion translates the multi-stream features to a single salient feature that is critical to fall detection and classification. eMSFRNet achieved 99.3% accuracy detecting falls and 76.8% accuracy classifying seven fall types. Our work is the firstmultistatic robust sensing system that overcomes the challenges associated with Doppler signatures under large and arbitrary aspect angles. Our work also demonstrates the potential to accommodate radar monitoring tasks that demand precise and robust sensing. Mengqi Shen, Kwok-Leung Tsui, Maury A. Nussbaum, Sunwook Kim, Fleming Lure |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | CAMV: A Crash Alarm Model for Vehicles Based on Internet of Vehicles DataabstractPredicting vehicle crashes is critical to improve urban transportation safety. It is however challenging to alarm impending crashes accurately as the volume of non-crash data dominates that of crash data. Most existing studies formulate the crash prediction as retrospective study and a binary classification problem, which offers the limited capacity of alarming crashes in advance. To fill such gap, this paper proposes a crash alarm model for vehicles (CAMV) developed via learning vehicle operational data. First, a non-crash learning block (NCLB) is developed to process time slices of operational data with a fixed length. It aims to model the inter-feature correlations and the temporal dependencies jointly. Next, a novel coarse-to-fine strategy is developed to selectively generate and calibrate alarms. Specifically, a statistical process control (SPC) module is developed at the coarse step to generate preliminary alarms with crash scores under an abnormal pattern. A self-calibration module (SCM) is next developed to eliminate false alarms and confirm the ultimate crash alarm. Experimental results verify the effectiveness of the proposed CAMV with crash data collected from Internet of Vehicles. Meanwhile, comparative experiments based on a set of benchmarking methods are conducted to verify the superiority of our proposal. Zijun Zhang 0001, Kwok-Leung Tsui |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Automatic fall risk assessment with Siamese network for stroke survivors using inertial sensor-based signalsabstractFall is a major threat to stroke survivors with the problems of gait and balance disorders in the rehabilitation phase following severe consequences on quality of life and a heavy burden to their families. Many solutions have been proposed to assess fall risk for elders based on inertial sensor-based signals, however, there still exists a great challenge of transferring them from elderly populations to the stroke-survivors populations as gait disorder patterns are significant difference between elders and stroke survivors. In this study, we conduct a pilot study to collect inertial sensor-based signals from stroke survivors when they performed the timed up and go test, and build an automatic fall risk assessment model with the architecture of Siamese network, with a merit of mitigating the problem of small sample size. Specifically, the proposed automatic fall risk assessment model consists of two parallel convolutional neural networks, each of which is composed of three convolutional layers, two max-pooling layers, and three fully connected layers. To utilize the space relation among accelerator-based and gyroscope-based signals, two-dimensional discrete wavelet transform extracts image-like features, wavelet coefficients, from inertial sensor-based signals as the input. Experimental results show that the proposed fall risk assessment model has achieved a promising results, which outperform cutting-edge methods with a big margin. The proposed fall risk assessment model with low computational complexity and limited memory consuming can be deployed on an embedded system to provide fall risk assessment service for stroke survivors in point-of-care environments or community settings. Xiaomao Fan, Yang Zhao 0009, Kuang-Hui Huang, Ya-Ting Wu, Tien-Lung Sun, Kwok-Leung Tsui |
Int. J. Intell. Syst. | 7 |
| 2022 | Nowcasting influenza-like illness (ILI) via a deep learning approach using google search data: An empirical study on Taiwan ILIabstractInfluenza outbreaks have brought increasing challenges to public health systems globally. The effective and efficient tracking of influenza can help authorities make informed and proactive decisions. In this study, we focus on nowcasting influenza epidemics at regional level. To alleviate the information lag between the release of the Centers for Disease Control's influenza-like illness (ILI) reports and real-time influenza activity, we incorporate Google search data and holiday effects to facilitate the ILI nowcasting. We develop a deep learning framework by extending the spatiotemporal residual network (ST-ResNet) to nowcast ILI rates in irregular-shaped region. We investigate the effect of temporal and spatial dependencies among irregular-shaped regions as well as external influence on ILI nowcasting. Various forecasting models, including time series models, penalized regression models, and other deep learning models, are employed for evaluating the performance of the proposed framework. Moreover, based on city-level ILI data in Taiwan, we conduct extensive experiments for methods validation. The results show that the extended ST-ResNet can effectively capture the complex spatiotemporal dependencies of ILI activity and the effects of external variables. Additionally, the strategy of incorporating Google search data and holiday effects improves the prediction. The findings may provide insights into the utility of various statistical and deep learning methods for effective ILI tracking at regional level and facilitate informed decision making for public health authorities. Yang Zhao 0009, Hsiang-Yu Yuan, Kwok-Leung Tsui |
Int. J. Intell. Syst. | 5 |
| 2022 | Multi-Graph Convolutional-Recurrent Neural Network (MGC-RNN) for Short-Term Forecasting of Transit Passenger FlowabstractShort-term forecasting of passenger flow is critical for transit management and crowd regulation. Spatial dependencies, temporal dependencies, inter-station correlations driven by other latent factors, and exogenous factors bring challenges to the short-term forecasts of passenger flow of urban rail transit networks. An innovative deep learning approach, Multi-Graph Convolutional-Recurrent Neural Network (MGC-RNN) is proposed to forecast passenger flow in urban rail transit systems to incorporate these complex factors. We propose to use multiple graphs to encode the spatial and other heterogenous inter-station correlations. The temporal dynamics of the inter-station correlations are also modeled via the proposed multi-graph convolutional-recurrent neural network structure. Inflow and outflow of all stations can be collectively predicted with multiple time steps ahead via a sequence to sequence(seq2seq) architecture. The proposed method is applied to the short-term forecasts of passenger flow in Shenzhen Metro, China. The experimental results show that MGC-RNN outperforms the benchmark algorithms in terms of forecasting accuracy. Besides, it is found that the inter-station driven by network distance, network structure, and recent flow patterns are significant factors for passenger flow forecasting. Moreover, the architecture of LSTM-encoder-decoder can capture the temporal dependencies well. In general, the proposed framework could provide multiple views of passenger flow dynamics for fine prediction and exhibit a possibility for multi-source heterogeneous data fusion in the spatiotemporal forecast tasks. Lishuai Li, Xinting Zhu, Kwok-Leung Tsui |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Forecasting influenza epidemics in Hong Kong using Google search queries data: A new integrated approach
Yunhao Liu 0002, Gengzhong Feng, Kwok-Leung Tsui, Shaolong Sun |
Expert Syst. Appl. | 3 |
| 2021 | Evaluating resampling methods and structured features to improve fall incident report identification by the severity levelabstractOBJECTIVE: This study aims to improve the classification of the fall incident severity level by considering data imbalance issues and structured features through machine learning. MATERIALS AND METHODS: We present an incident report classification (IRC) framework to classify the in-hospital fall incident severity level by addressing the imbalanced class problem and incorporating structured attributes. After text preprocessing, bag-of-words features, structured text features, and structured clinical features were extracted from the reports. Next, resampling techniques were incorporated into the training process. Machine learning algorithms were used to build classification models. IRC systems were trained, validated, and tested using a repeated and randomly stratified shuffle-split cross-validation method. Finally, we evaluated the system performance using the F1-measure, precision, and recall over 15 stratified test sets. RESULTS: The experimental results demonstrated that the classification system setting considering both data imbalance issues and structured features outperformed the other system settings (with a mean macro-averaged F1-measure of 0.733). Considering the structured features and resampling techniques, this classification system setting significantly improved the mean F1-measure for the rare class by 30.88% (P value < .001) and the mean macro-averaged F1-measure by 8.26% from the baseline system setting (P value < .001). In general, the classification system employing the random forest algorithm and random oversampling method outperformed the others. CONCLUSIONS: Structured features provide essential information for categorizing the fall incident severity level. Resampling methods help rebalance the class distribution of the original incident report data, which improves the performance of machine learning models. The IRC framework presented in this study effectively automates the identification of fall incident reports by the severity level. Jiaxing Liu 0002, Zoie Shui-Yee Wong, Hing-Yu So, Kwok-Leung Tsui |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | An Adaptive Gaussian Process-Based Search for Stochastically Constrained Optimization via SimulationabstractSimulation optimization (SO) techniques show a strong ability to solve large-scale problems. In this article, we concentrate on stochastically constrained SO. There are some challenges to tackle the problem: 1) the objective and constraints have no analytical forms and need to be evaluated via simulation; 2) we should make a tradeoff between exploiting around the best solution and exploring more unknown regions; and 3) both the objective value and feasibility determine the quality of a solution. Motivated by these issues, we propose an adaptive Gaussian process-based search (AGPS) to address stochastically constrained discrete SO problems. AGPS fast constructs the Gaussian process for each performance and then builds a new sampling distribution to adaptively balance exploration and exploitation considering the objective function and stochastic constraints. We show that AGPS converges to the set of globally optimal solutions with probability one. Numerical experiments demonstrate the superiority of our method compared with other advanced approaches.Note to Practitioners—Simulation is widely used to model complex and large-scale systems, such as healthcare, transportation, and supply chain logistics. When optimizing these systems, practitioners always assess overall performance by multiple indicators. Inspired by this issue, this article focuses on a general problem that aims to optimize the system’s primary performance while keeping the secondary performance within limits. We propose an adaptive Gaussian process-based method called AGPS to search for a high-quality solution. The bright spot of AGPS is that it can intelligently determine the quality of a solution considering all the stochastic performance and adaptively search the solution space. The merit lightens the burden of practitioners to design specific parameters for different performance indicators. Numerical experiments demonstrate that AGPS can apply to the large-scale discrete optimization problems with smooth function values and shows higher efficiency than existing methods. Hainan Guo, Kwok-Leung Tsui |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2021 | Active Balancing of Lithium-Ion Batteries Using Graph Theory and A-Star Search AlgorithmabstractThe heterogeneity of cells in a battery pack is inevitable but brings high risks of premature failure and even safety hazards. Accordingly, for safe and long-life operation, it is necessary to adjust the state of charge (SOC) of all in-pack cells to the same level. To address this problem, this article first proposes a battery SOC observer and analyzes its stability and convergence analysis using the Lyapunov direct method. Different to most available estimators is that the proposed method does not require the information of cell capacities. Then, after modeling the equalization system as a directed graph, the equalization problem is cast as a path searching problem. Finally, an A-star algorithm subject to balancing constraints is proposed to find the shortest path in this graph, corresponding to the most efficient SOC equalization. Experimental results show that the steady-state error of the proposed observer is less than $2\%$. It also demonstrates that the A-star algorithm can decrease the balancing time and energy loss during the balancing process by 9.59% and 19.5%, respectively, relative to the mean-difference-average method. Guangzhong Dong, Fangfang Yang, Kwok-Leung Tsui, Changfu Zou |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Signal-Disturbance Interfacing Elimination for Unbiased Model Parameter Identification of Lithium-Ion BatteryabstractA precisely parameterized battery model is the prerequisite of the model-based management of lithium-ion battery. However, the unexpected sensing of noises may discount the identification of model parameters in practical applications. This article focuses on the noise effect compensation and online parameter identification for the widely used equivalent circuit model. A novel degree of freedom (DOF) eliminator is proposed and combined with the Frisch scheme in a recursive fashion, for the first time, to coestimate the noise statistics and unbiased model parameters. A computationally tractable numerical solver is further proposed for the DOF eliminator to improve the real-time performance. Simulations and experiments are performed to validate the proposed method from theoretical to practical perspective. Results show that the proposed method can effectively mitigate the noise-induced identification biases and outperform the existing methods in terms of the accuracy and the robustness to noise corruption. Zhongbao Wei, Hongwen He, Josep Pou, Kwok-Leung Tsui, Zhongyi Quan, Yunwei Li 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Correction to "High-Speed Rail Suspension System Health Monitoring Using Multi-Location Vibration Data"abstractIn the above article[1],Table I,III, andIVshould show “N/m” instead of “kN/m” and they should also show “Ns/m” instead of “kNs/m.” The revised tables are shown below. Ning Hong, Lishuai Li, Weiran Yao, Yang Zhao 0009, Cai Yi, Jianhui Lin, Kwok-Leung Tsui |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2020 | Fault diagnosis of electrohydraulic actuator based on multiple source signals: An experimental investigation
Jianguo Miao, Jinglin Wang, Fangfang Yang, Kwok-Leung Tsui, Qiang Miao |
Neurocomputing | 5 |
| 2020 | Data-Driven Battery Health Prognosis Using Adaptive Brownian Motion ModelabstractDegradation dynamics modeling and health prognosis play extremely important roles in system prognostics and health management. Wiener process-based degradation models and remaining useful life (RUL) prediction methods have the advantage of high flexibility and efficiency, with features such as Brownian motion with drift and scale parameters. They can also quantify prediction uncertainty through inverse Gaussian distribution. However, prior studies use offline-identified model parameters, which can result in difficulties in both model adaptability and health prognosis. To improve the performance of Wiener process models, this article proposes a new data-driven Brownian motion model that utilizes the adaptive extended Kalman filter (AEKF) parameter identification method. The proposed model can update model parameters online and adapt to uncertain degradation operations. This data-driven method has the flexibility and efficiency of Brownian motion models but avoids their shortcomings in model adaptability and health prognosis. The model parameters and drift parameter are online estimated based on AEKF using limited historical system measurements. The effectiveness of the proposed data-driven framework in degradation modeling and RUL prediction is evaluated through simulations and experimental results on lithium-ion battery degradation data. The results show that the proposed approach has significant accuracy and robustness for both model adaptability and RUL prediction. Guangzhong Dong, Fangfang Yang, Zhongbao Wei, Jingwen Wei, Kwok-Leung Tsui |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Homecare-Oriented Intelligent Long-Term Monitoring of Blood Pressure Using Electrocardiogram SignalsabstractLong-term blood pressure (BP) monitoring is a widely used approach in a homecare intelligent system. However, BP is usually measured using cuff-based devices with tedious operation in practice, which may not be cost effective for continuous BP tracking. In this article, we propose a novel attention-based multitask network with a weighting scheme for BP estimation by analyzing and modeling single lead electrocardiogram (ECG) signals. Experimental results demonstrate that the proposed method could achieve mean error of systolic blood pressure, diastolic blood pressure, and mean arterial pressure estimation in levels of 0.18 ± 10.83, 1.24 ± 5.90, and 0.84 ± 6.47 mmHg, respectively. In comparison to other cutting-edge methods using ECG signals, the proposed method shows superior BP estimation performance. By integrating with a wearable/portable ECG monitoring device, the proposed model can be deployed to an embedded system or remote healthcare intelligent system to provide long-term BP monitoring service, which would help to reduce the incidence of malignant events happened in hypertensive population. Xiaomao Fan, Fan Xu 0003, Yang Zhao 0009, Kwok-Leung Tsui |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | High-Speed Rail Suspension System Health Monitoring Using Multi-Location Vibration DataabstractA novel data-driven framework to monitor the health status of high-speed rail suspension system by measuring train vibrations is proposed herein. Unlike existing methods, this framework does not rely on sophisticated dynamic models or high-fidelity simulations; it combines the power of data and domain knowledge to generate a model that can be trained quickly and adapted easily to different rail systems. In addition, the framework includes a module to generate a training dataset, tackling a typical challenge in real-world system monitoring, namely, the lack of labeled data due to practical limits. Based on the multi-output support vector regression (MSVR), the proposed framework can monitor the stiffness and damping coefficients of the suspension system using vibration signals measured on trains in real time. The framework comprises three modules. First, a simple suspension system dynamics model is built to generate a training dataset. Furthermore, key features are extracted from frequency response curves to reflect the impact of spring and damper degradation. Subsequently, a supervised learning model based on the MSVR is built to predict the stiffness and damping coefficients of suspension systems from features extracted in the second module. Once the model is built, real-time monitoring can be achieved by feeding the vibration signals as they are collected during operations. The proposed framework was evaluated on simulation data for its accuracy and tested on real-world operational data for its practicability. Ning Hong, Lishuai Li, Weiran Yao, Yang Zhao 0009, Cai Yi, Jianhui Lin, Kwok-Leung Tsui |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2019 | A Two-Stage HMM Model for Sleep/Wake Identification via Commercial Wearable DeviceabstractGood sleep habit is essential to maintain a quality life. Sensor-based wearable devices have been increasingly deployed to monitor activity and measure sleep duration in a non-intrusive, affordable, and portable way. Existing algorithms for detecting sleep and wake in wrist-worn devices are mainly based on activity count or accelerometer data inference. It is validated that sleep can be reflected by various vital signs, including physical activity, heart rates, and pulse oximetry. However, little attention has been paid to heart rates measurements for sleep/wake identification in commercial wearable devices together with activity information. Our study developed an unsupervised and personalized algorithm to infer sleep and wake states using heart rates and step counts based on hidden Markov models. The fusion of two HMMs successfully dealt with multi-granularity data and predicted sleep/wake states in the minimal granularity. The proposed algorithm was illustrated through a real-life case study. The agreement between our algorithm and Fitbit’s scoring was 89.35%. The proposed algorithm enabled identifying more afternoon naps, earlier sleep onset compared to Fitbit’s scoring. The results showed that heart rates were informative when distinguishing sleep and wake while compensating estimations driven from step counts. Jiaxing Liu 0002, Yang Zhao 0009, Boya Lai, Kwok-Leung Tsui |
SMC | 4 |
| 2019 | Process monitoring using variational autoencoder for high-dimensional nonlinear processes
Mingu Kwak, Kwok-Leung Tsui, Seoung Bum Kim |
Eng. Appl. Artif. Intell. | 3 |
| 2019 | Fostering linguistic decision-making under uncertainty: A proportional interval type-2 hesitant fuzzy TOPSIS approach based on Hamacher aggregation operators and andness optimization models
Zhen-Song Chen 0002, Yi Yang 0020, Xianjia Wang, Kwai-Sang Chin, Kwok-Leung Tsui |
Inf. Sci. | 5 |
| 2019 | Regularized Gaussian Mixture Model for High-Dimensional ClusteringabstractFinding low-dimensional representation of high-dimensional data sets is an important task in various applications. The fact that data sets often contain clusters embedded in different subspaces poses barrier to this task. Driven by the need in methods that enable clustering and finding each cluster's intrinsic subspace simultaneously, in this paper, we propose a regularized Gaussian mixture model (GMM) for clustering. Despite the advantages of GMM, such as its probabilistic interpretation and robustness against observation noise, traditional maximum-likelihood estimation for GMMs shows disappointing performance in high-dimensional setting. The proposed regularization method finds low-dimensional representations of the component covariance matrices, resulting in better estimation of local feature correlations. The regularization problem can be incorporated in the expectation maximization algorithm for maximizing the likelihood function of a GMM, with the M -step modified to incorporate the regularization. The M -step involves a determinant maximization problem, which can be solved efficiently. The performance of the proposed method is demonstrated using several simulated data sets. We also illustrate the potential value of the proposed method in applications using four real data sets. Yang Zhao 0009, Abhishek K. Shrivastava, Kwok-Leung Tsui |
IEEE Trans. Cybern. | 3 |
| 2019 | Forecasting Short-Term Passenger Flow: An Empirical Study on Shenzhen MetroabstractForecasting short-term traffic flow has been a critical topic in transportation research for decades, which aims to facilitate dynamic traffic control proactively by monitoring the present traffic and foreseeing its immediate future. In this paper, we focus on forecasting short-term passenger flow at subway stations by utilizing the data collected through an automatic fare collection (AFC) system along with various external factors, where passenger flow refers to the volume of arrivals at stations during a given period of time. Along this line, we propose a data-driven three-stage framework for short-term passenger flow forecasting, consisting of traffic data profiling, feature extraction, and predictive modeling. We investigate the effect of temporal and spatial features as well as external weather influence on passenger flow forecasting. Various forecasting models, including the time series model auto-regressive integrated moving average, linear regression, and support vector regression, are employed for evaluating the performance of the proposed framework. Moreover, using a real data set collected from the Shenzhen AFC system, we conduct extensive experiments for methods validation, feature evaluation, and data resolution demonstration. Liyang Tang, Yang Zhao 0009, Javier Cabrera, Jian Ma 0009, Kwok-Leung Tsui |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2018 | Two-stage aggregation paradigm for HFLTS possibility distributions: A hierarchical clustering perspective
Zhen-Song Chen 0002, Luis Martínez-López 0001, Kwai-Sang Chin, Kwok-Leung Tsui |
Expert Syst. Appl. | 4 |
| 2018 | An integrated machine learning framework for hospital readmission prediction
Shancheng Jiang, Kwai-Sang Chin, Gang Qu 0004, Kwok-Leung Tsui |
Knowl. Based Syst. | 4 |
| 2018 | Customizing Semantics for Individuals With Attitudinal HFLTS Possibility DistributionsabstractLinguistic computational techniques based on hesitant fuzzy linguistic term set (HFLTS) have been swiftly advanced on various fronts over the past five years. However, one critical issue in the existing theoretical development is that modeling possibility distribution based semantics involves a relatively strict constraint that linguistic terms are uniformly distributed across an HFLTS. Releasing the constraint of uniform HFLTS through which individual semantics could be customized is challenging yet intriguing for participants interested in this topic. Comparative linguistic expressions (CLEs) generated from context-free grammar facilitate flexible and accurate linguistic elicitation, and in consideration of computational simplicity, are transformed into HFLTSs that are machine manipulatable. It is imperative that the precision of customized individual semantics can be significantly improved with respect to different CLEs. This study proposes a novel possibility computation structure for HFLTS possibility distributions based on the linguistic terms similarity measure. The uniquely established linguistic terms in each and every CLE are initially treated as referential items for comparison. Then, possibilities of linguistic terms in a transformed HFLTS can be calculated as their similarity degrees to the predetermined referential item. Subsequently, the interweaving method in which a consistent inner interweaving matrix needs to be constructed is adopted for attitudinal characters to attain appealing degrees characterized in the unit interval. The generated attitudinal HFLTS possibility distributions provide a solution to the problem of modeling individually the semantic implications of CLEs. Several illustrative examples and comparative analyses further demonstrate that individual semantics endowed with attitudinal character model efficiently individual differences in cognitive styles. Zhen-Song Chen 0002, Kwai-Sang Chin, Luis Martínez-López 0001, Kwok-Leung Tsui |
IEEE Trans. Fuzzy Syst. | 4 |
| 2017 | Forecasting tourist arrivals with machine learning and internet search indexabstractThe queries entered into search engines register hundreds of millions of different searches by tourists, not only reflecting the trends of the searchers' preferences for travel products, but also offering a forecasting of their future travel behavior. This paper proposed a forecasting framework based on internet search index and machine learning to forecast tourist arrivals, and compared the forecasting performance of two different search engines data, Baidu and Google. The empirical results suggest that the proposed KELM models by fusing Baidu index and Google index can significantly improve the forecasting performance and outperform other benchmark models in terms of forecasting accuracy. Shaolong Sun, Shou-Yang Wang, Yunjie Wei, Xianduan Yang, Kwok-Leung Tsui |
IEEE BigData | 5 |
| 2017 | Modified genetic algorithm-based feature selection combined with pre-trained deep neural network for demand forecasting in outpatient department
Shancheng Jiang, Kwai-Sang Chin, Long Wang 0015, Gang Qu 0004, Kwok-Leung Tsui |
Expert Syst. Appl. | 5 |
| 2017 | Simulation Optimization for Medical Staff Configuration at Emergency Department in Hong KongabstractMedical staff configuration is a critical problem in the management of an emergency department (ED) in Hong Kong (HK). Given the service requirements by HK government, it is imperative for the hospital managers to develop medical staff configuration in a cost-and-time-effective way. In this paper, the medical staff configuration problem in ED is modeled as minimizing the total labor cost while satisfying the service quality requirements. To solve this issue, we propose a highly efficient search method, called random boundary generation with feasibility detection (RBG-FD). The random boundary generation (RBG) is applied to efficiently identify good-quality solutions based on the objective value. The feasibility detection (FD) procedure is used to retain the probability of correct feasibility detection of each sampled solution at the desired level, which intrinsically allocates a reasonable number of simulation replications. To estimate the performance measures of the ED, a discrete-event simulation model is developed to reflect the patient flow. Using these techniques, the efficiency of identifying the optimal staff configuration can be significantly improved. A case study is performed in a public hospital in HK. The numerical results indicate significantly higher practicability and efficiency of the proposed method with different patient arrival rates and service constraints. Note to Practitioners This paper seeks to solve the problem of minimizing the medical staff cost constrained by certain service requirements [i.e., patients' waiting times for treatment] at an emergency department in Hong Kong. In our formulation, these service requirements are characterized by some stochastic constraints. Most of the existing random search methods concentrate on the computing efforts in the neighborhood of the best-so-far solutions in order to obtain good-quality solutions. Due to the special structure of this problem and ease of computing the objective values, we proposed an efficient random search approach that iteratively identifies a solution with a better objective value than that of the current best solution. Experimental studies demonstrate the significantly higher efficiency of this method. In order to obtain the same solution quality, it is able to reduce the computational time by 90% compared with some existing approaches in the literature. Hainan Guo, Siyang Gao, Kwok-Leung Tsui, Tie Niu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2017 | Statistical Modeling of Bearing Degradation SignalsabstractBearings are the most common mechanical components used in machinery to support rotating shafts. Due to harsh working conditions, bearing performance deteriorates over time. To prevent any unexpected machinery breakdowns caused by bearing failures, statistical modeling of bearing degradation signals should be immediately conducted. In this paper, given observations of a health indicator, a statistical model of bearing degradation signals is proposed to describe two distinct stages existing in bearing degradation. More specifically, statistical modeling of Stage I aims to detect the first change point caused by an early bearing defect, and then statistical modeling of Stage II aims to predict bearing remaining useful life. More importantly, an underlying assumption used in the early work of Gebraeel et al. is discovered and reported in this paper. The work of Gebraeel et al. is extended to a more general prognostic method. Simulation and experimental case studies are investigated to illustrate how the proposed model works. Comparisons with the statistical model proposed by Gebraeel et al. for bearing remaining useful life prediction are conducted to highlight the superiority of the proposed statistical model. Dong Wang 0001, Kwok-Leung Tsui |
IEEE Trans. Reliab. | 2 |
| 2016 | The therapist assignment problem in home healthcare structures
Meiyan Lin, Kwai-Sang Chin, Xianjia Wang, Kwok-Leung Tsui |
Expert Syst. Appl. | 4 |
| 2016 | A Two-Stage Data-Driven-Based Prognostic Approach for Bearing Degradation ProblemabstractPrognostics of the remaining useful life (RUL) has emerged as a critical technique for ensuring the safety, availability, and efficiency of a complex system. To gain a better prognostic result, degradation information is quite useful because it can reflect the health status of a system. However, due to the lack of accurate information about the plants' degradation, the prognostic model is usually not well established. To solve this problem, this paper proposes a two-stage strategy that is in the context of data-driven modeling to predict the future health status of a bearing, where the degradation information was estimated by calculating the deviation of multiple statistics of vibration signals of a bearing from a known healthy state. Then, a prediction stage based on an enhanced Kalman filter and an expectation-maximization algorithm were used to estimate the RUL of the bearing adaptively. To verify the effectiveness of the proposed approach, a real-bearing degradation problem was implemented. Yu Wang 0043, Yizhen Peng, Yanyang Zi, Xiaohang Jin, Kwok-Leung Tsui |
IEEE Trans. Ind. Informatics | 5 |
| 2015 | An integrated Bayesian approach to prognositics of the remaining useful life and its application on bearing degradation problemabstractDegradation information of a complex mechanical system reflects the system's health status and is useful to predict the future progression of the fault or anomalous behaviors. This paper proposed a two-stage strategy to predict the future health status of a bearing by utilizing the bearing's degradation information. The first stage was implemented to monitor the bearing's health status until a degradation point was detected. When the bearing begins to degrade, a prediction stage based on Kalman filter was then used to estimate the remaining useful life (RUL) of the bearing. Finally, a real bearing degradation problem was used to verify our method. Yu Wang 0043, Yizhen Peng, Yanyang Zi, Xiaohang Jin, Kwok-Leung Tsui |
INDIN | 5 |
| 2015 | Accelerated Degradation Analysis for the Quality of a System Based on the Gamma ProcessabstractAs most systems these days are highly reliable with long lifetimes, failures of systems become rare; consequently, traditional failure time analysis may not be able to provide a precise assessment of the system reliability. In this regard, a degradation measure, as a percentage of the initial value, is an alternate way of describing the system health. This paper presents accelerated degradation analysis that characterizes the health and quality of systems with monotonic and bounded degradation. The maximum likelihood estimates (MLEs) of the model parameters are derived, based on a gamma process, time-scale transformation, and a power link function for associating the covariates. Then, methods of estimating the reliability, the mean and median lifetime, the conditional reliability, and the remaining useful life of systems under normal use conditions are all described. Moreover, approximate confidence intervals for the parameters of interest are developed based on the observed Fisher information matrix. A model validation metric with exact power is introduced. A Monte Carlo simulation study is carried out for evaluating the performance of the proposed methods. For an illustration of the proposed model, and the methods of inference developed here, a numerical example involving light intensity of light emitting diodes (LED) is analyzed. Man Ho Ling, Kwok-Leung Tsui, Narayanaswamy Balakrishnan 0001 |
IEEE Trans. Reliab. | 2 |
| 2014 | A Two-Step Parametric Method for Failure Prediction in Hard Disk DrivesabstractPredicting the impending failure of hard disk drives (HDDs) is crucial for preventing essential data from losing. In this paper, a two-step parametric method was developed to predict the impending failure of HDDs using the aggregate of statistical models. This method deals with the problem of failure prediction in two steps: anomaly detection and failure prediction. First, Mahalanobis distance was used for aggregating all the monitored variables into one index, which was then transformed into Gaussian variables by Box–Cox transformation. By defining an appropriate threshold, anomalies in HDDs were detected as a result. Second, a sliding-window-based generalized likelihood ratio test was proposed to track the anomaly progression in an HDD. When the occurrence of anomalies in a time interval is found to be statistically significant, indicating the HDD is approaching failure. In this work, we also derived a new cost function to adjust the prediction rate. This is important in a way to balance the failure detection rate and false alarm rate as well as to provide an advanced warning of HDD failures to the users, whereby the users can back up their data in time. Then the developed method was applied on a synthetic data set showing its effectiveness on predicting failures. To demonstrate the practical usefulness, this method was also applied on a real-life HDD data set. The result shows that our method could achieve 68% failure detection rate with 0% false alarm rate. This is much better than the results achieved by the state-of-the-art methods, such as support vector machine and hidden Markov models. Yu Wang 0043, Eden W. M. Ma, Tommy W. S. Chow, Kwok-Leung Tsui |
IEEE Trans. Ind. Informatics | 4 |
| 2013 | Calibration of Stochastic Computer Models Using Stochastic Approximation MethodsabstractComputer models are widely used to simulate real processes. Within the computer model, there always exist some parameters which are unobservable in the real process but need to be specified in the model. The procedure to adjust these unknown parameters in order to fit the model to observed data and improve predictive capability is known as calibration. Practically, calibration is typically done manually. In this paper, we propose an effective and efficient algorithm based on the stochastic approximation (SA) approach that can be easily automated. We first demonstrate the feasibility of applying stochastic approximation to stochastic computer model calibration and apply it to three stochastic simulation models. We compare our proposed SA approach with another direct calibration search method, the genetic algorithm. The results indicate that our proposed SA approach performs equally as well in terms of accuracy and significantly better in terms of computational search time. We further consider the calibration parameter uncertainty in the subsequent application of the calibrated model and propose an approach to quantify it using asymptotic approximations. Jun Yuan 0002, Szu Hui Ng, Kwok-Leung Tsui |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2013 | Online Anomaly Detection for Hard Disk Drives Based on Mahalanobis DistanceabstractA hard disk drive (HDD) failure may cause serious data loss and catastrophic consequences. Online health monitoring provides information about the degradation trend of the HDD, and hence the early warning of failures, which gives us a chance to save the data. This paper developed an approach for HDD anomaly detection using Mahalanobis distance (MD). Critical parameters were selected using failure modes, mechanisms, and effects analysis (FMMEA), and the minimum redundancy maximum relevance (mRMR) method. A self-monitoring, analysis, and reporting technology (SMART) data set is used to evaluate the performance of the developed approach. The result shows that about 67% of the anomalies of failed drives can be detected with zero false alarm rate, and most of them can provide users with at least 20 hours during which to backup the data. Yu Wang 0043, Qiang Miao, Eden W. M. Ma, Kwok-Leung Tsui, Michael G. Pecht |
IEEE Trans. Reliab. | 4 |
| 2013 | Degradation Data Analysis Using Wiener Processes With Measurement ErrorsabstractDegradation signals that reflect a system's health state are important for diagnostics and health management of complex systems. However, degradation signals are often compounded and contaminated by measurement errors, making data analysis a difficult task. Motivated by the wear problem of magnetic heads used in hard disk drives (HDDs), this paper investigates Wiener processes with measurement errors. We explore the traditional Wiener process with positive drifts compounded with i.i.d. Gaussian noises, and improve its estimation efficiency compared with the existing inference procedure. Furthermore, to capture the possible heterogeneity in a population, we develop a mixed effects model with measurement errors. Statistical inferences of this model are discussed. The mixed effects model subsumes several existing Wiener processes as its limiting cases, and thus it is useful for suggesting an appropriate Wiener process model for a specific dataset. The developed methodologies are then applied to the wear problem of magnetic heads of HDDs, and a light intensity degradation problem of light-emitting diodes. Zhisheng Ye 0001, Yu Wang 0043, Kwok-Leung Tsui, Michael G. Pecht |
IEEE Trans. Reliab. | 3 |
| 2012 | Ensemble-approaches for clustering health status of oil sand pumps
Francesco Di Maio, Peter W. Tse, Michael G. Pecht, Kwok-Leung Tsui, Enrico Zio |
Expert Syst. Appl. | 5 |
| 2011 | Logistic regression analysis for Predicting Methicillin-resistant Staphylococcus Aureus (MRSA) in-hospital mortalityabstractStatistical models have been widely used in public health and made a difference in a wide range of applications. For example, they provide new ideas for efficient feature selection. This paper attempts to demonstrate how to apply regression-based methods to accurately predict in-hospital mortality of Methicillin-resistant Staphylococcus Aureus (MRSA) patients. Logistic regression is used to predict the in-hospital death. It is found that admission age, residency, solid tumor, hemic malignancy, COAD, Dementia, PLT, Lymphocyte, Urea, and ALP are the significant prognostic factors (P<;0.1) for in-hospital survival. Using cross validation and random splitting and the prediction accuracy is around 85%. The future research direction is to strengthen the robustness of the predictive model. Possible direction is to make use of other data mining “blackbox” methods, such as k-NN and SVM. These models also need further validation on their performance and feature selection. Yizhen Hai, Vincent C. Cheng, Zoie Shui-Yee Wong, Kwok-Leung Tsui, Kwok-Yung Yuen |
ISI | 4 |
| 2011 | Prognostics and health monitoring for lithium-ion batteryabstractHealth monitoring is used to analyze and predict the battery health status. However, no matter what health monitoring methods and parameters are, a major aim is to improve the battery reliability through surveillance and prognostics. Hence, the latest known methods of state estimation and life prediction based on battery health monitoring are discussed in this paper. Through comparing their characteristics respectively, a prognostics-based fusion technique is proposed that combines physics-of-failure (PoF) with data-driven technology. The fusion approach not only investigates battery failure mechanism caused by environmental and internal characteristics, but also assesses parameters with aid of real-time health monitoring. The specific method is presented to realize the estimation on remaining useful life (RUL) of batteries. Yinjiao Xing, Qiang Miao, Kwok-Leung Tsui, Michael G. Pecht |
ISI | 3 |
| 2011 | Recent Research and Developments in Temporal and Spatiotemporal Surveillance for Public HealthabstractThe objective of public health surveillance is to systematically collect, analyze, and interpret public health data (about chronic or infectious diseases) to understand trends; detect changes in disease incidence and death rates; and plan, implement, and evaluate public health practices. Recently, studies have been conducted to develop methods and algorithms for health surveillance and disease detection. This paper attempts to review recent research on temporal and spatiotemporal surveillance methods. We have addressed specific challenges and research gaps in the relevant research. Lastly, we discuss a comparative example using a dataset of male thyroid cancer cases in New Mexico. Kwok-Leung Tsui, Zoie Shui-Yee Wong, Chen-ju Lin |
IEEE Trans. Reliab. | 1 |
| 2009 | Linear-mixed effects models for feature selection in high-dimensional NMR spectra
Yajun Mei, Seoung Bum Kim, Kwok-Leung Tsui |
Expert Syst. Appl. | 3 |
| 2008 | Invited Keynote Talk: Data Mining and Statistical Methods for Analyzing Microarray Experiments
Shin-Lian Lo, Kwok-Leung Tsui, Benjamin Barwick |
ISBRA | 2 |
| 2006 | FBP: A Frontier-Based Tree-Pruning AlgorithmabstractA frontier-based tree-pruning algorithm (FBP) is proposed. The new method has an order of computational complexity comparable to cost-complexity pruning (CCP). Regarding tree pruning, it provides a full spectrum of information: namely, (1) given the value of the penalization parameter λ, it gives the decision tree specified by the complexity-penalization approach; (2) given the size of a decision tree, it provides the range of the penalization parameter λ, within which the complexity-penalization approach renders this tree size; (3) it finds the tree sizes that are inadmissible—no matter what the value of the penalty parameter is, the resulting tree based on a complexity-penalization framework will never have these sizes. Simulations on real data sets reveal a “surprise:” in the complexity-penalization approach, most of the tree sizes are inadmissible. FBP facilitates a more faithful implementation of cross validation (CV), which is favored by simulations. Using FBP, a stability analysis of CV is proposed. Xiaoming Huo, Seoung Bum Kim, Kwok-Leung Tsui, Shuchun Wang |
INFORMS J. Comput. | 3 |