Zijun Zhang 0001

dblp:84/4245-1 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0002-2717-5033ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Computer networks · 9 · 9 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ComfortLLM: Compatible Multimodality Fusion-Oriented Large-Language Model for Industrial Fault Diagnosis With Diverse Data
Tao Hu 0015, Zhenling Mo, Zijun Zhang 0001
IEEE Internet Things J.4
2026 FedACT: Federated Agnostic Learning on Limited Decentralized CT Images With Knowledge Transferring Process
abstract
Existing federated learning primarily focuses on problem setups where servers and clients engage in model training for one or multiple specific tasks. However, in real-world scenarios, the required diagnoses at clinical sites can vary due to the diversity of conditions, differing not only from each other but also from those at the server. In this study, we concentrate on the practical yet challenging Federated Agnostic Learning (FAL), where client-side diagnostic tasks are agnostic. We introduce a novel FedACT method to address this issue, which is composed of two components. The first component extracts shared features among multiple agnostic tasks using an end-to-end similarity layer based on contrastive learning to enhance generalizability. In the second component, we design personalized task-specific branches for comprehensive tasks, including classification and segmentation. The branches can accurately accomplish the respective tasks through knowledge transfer and enhanced discriminative capabilities across various classes and tasks. Moreover, to better accommodate the potential heterogeneity of data and unseen tasks, specialized updating and aggregation methods are devised for FedACT. The experimental results demonstrate the effectiveness of FedACT in various scenarios under the FAL setting.
Liuyin Chen, Long Wang 0015, Guoyuan Liang, Zijun Zhang 0001
IEEE J. Biomed. Health Informatics4
2025 Neural ODE powered model for bearing remaining useful life predictions with intra- and inter-domain shifts
Tao Hu 0015, Zhenling Mo, Zijun Zhang 0001
Adv. Eng. Informatics3
2025 STRFormer: A spatial topological relationship-guided multi-modal variational fusion network for intelligent health state diagnosis of the manipulator
Bo Zhao 0026, Qiqiang Wu, Xianmin Zhang 0004, Zijun Zhang 0001
Adv. Eng. Informatics6
2025 Multitask Pointwise Mutual Information Learning for Bearing Remaining Useful Life Cross-Domain Imbalanced Regression
abstract
In modern industry, the Industrial Internet of Things (IIoT) has enabled system health analytics and monitoring through continuous data collection and networked connectivity. Bearing remaining useful life (RUL) prediction is one pivotal analytical task in preventing industrial system failures and optimizing maintenance schedules. Existing prediction methods using data face two critical engineering challenges: (1) performance degrades when deployed to unseen operational domains, and (2) imbalanced sensor data distributions cause biased predictions. Two challenges can even co-exist, which further limits the effectiveness of prediction methods using data. This study proposes a multi-task pointwise mutual information learning (MPML) based prediction model development method to tackle bearing RUL prediction under this compound challenge. MPML offers several key innovations through the following developments. First, an auxiliary task-assisted multi-task model learning scheme is devised to obtain task-wise generalizability for learning invariant latent features, and theoretical analysis is provided to explain the invariant feature learning mechanism. Furthermore, pointwise mutual information (PMI) modeling with a statistical explanation is proposed to impose adaptive biases, rectifying penalties for RUL misprediction in minority groups. Consequently, MPML addresses the unseen domain prediction through the multi-task learning scheme and effectively handles the data imbalance with the novel PMI-assisted loss. Extensive computational experiments are conducted to demonstrate the superiority of MPML, achieving a 33.12 square error compared to state-of-the-art methods.
Tao Hu 0015, Zhenling Mo, Zijun Zhang 0001
IEEE Internet Things J.3
2025 Semantics-Consistent Representation Learning for Industrial Fault Diagnosis in Unseen Domains
abstract
Domain generalizable fault diagnosis (DGFD) aims to develop robust and adaptable models for reliable fault diagnosis in unseen domains. Although many studies have been conducted to enhance model resistance against data distribution shifts, most existing models prioritize domain invariance and overlook class discriminability, which is crucial for DGFD. Furthermore, DGFD tasks often face significant challenges from the compound effect of class imbalance and subpopulation shifts. Therefore, this work proposes a novel Semantics-Consistent Representation Learning (SCRL) framework that enhances data-driven modeling for class-imbalanced DGFD (IDGFD) by more effectively learning domain-invariant but class-discriminative feature representations. To address class imbalance and prepare strategic data for robust model training, a multi-pipe interactive data processing scheme is designed to adaptively generate training samples. To improve the generalizability and discriminability of SCRL, modules for fault diagnosis, causal factorization, and affinity mining are jointly incorporated to address the data distribution and subpopulation shifts. This integration enables establishing domain-invariant and class-discriminative boundaries for effective IDGFD. Extensive experiments on four datasets demonstrate the superiority of SCRL over existing models in achieving accurate and robust fault diagnosis across unseen working conditions and systems. The code is available at https://github.com/ifuturekk/SCRL.
Jun Lin 0002, Zijun Zhang 0001
IEEE Internet Things J.3
2025 AUDAN: An adversarial unsupervised domain adaptation network informed decision-making in audit risk assessment
abstract
This paper develops an adversarial unsupervised domain adaptation network (AUDAN) for audit risk assessment from a data-driven perspective, which can offer a novel solution for the audit issue data distribution shift problem in audits. The proposed AUDAN consists of three neural network–based modules: the feature extractor, the domain classifier, and the risk-level classifier. A joint adversarial learning scheme based on these three modules is developed to enable learning discriminative and domain-invariant feature representations from audit issue data. In developing the AUDAN, the data is reorganized into source and target domains. A simple yet effective model selection technique, called latest held-out source risk validation, is proposed for the time-variant domain shift, where the data distribution of one period is closer to the data distribution of the adjacent period. The superiority of the latest held-out source risk validation technique has been theoretically justified. Computational experiments based on a real-world dataset were performed to verify the advantages and effectiveness of the AUDAN. The results showed that the AUDAN achieves superior testing performance in most cases. Additionally, the AUDAN yields robust classification results and outperforms the benchmarking models, including the train-on-target model, which is trained with the label information of target domain data revealed and can be regarded as a strong competitor. Ablation studies further showed the superiority of the developed latest held-out source risk validation method. © 2025 The Author(s).
Ziquan Ou, Zijun Zhang 0001
Knowl. Based Syst.2
2025 Distributed Secondary Control of DC Microgrids Under Unreliable Communication Networks
Haihua Guo, Xiaoran Dai, Siqi Bu, Zijun Zhang 0001
IEEE Trans Autom. Sci. Eng.4
2025 Privacy-Preserving Distributed Fault Diagnosis for Multiple Wind Farms Using a Federated Feature Fusion Method
Zijun Zhang 0001, Ershun Pan
IEEE Trans Autom. Sci. Eng.3
2025 Lifeisgood: Learning Invariant Features via In-Label Swapping for Generalizing Out-of-Distribution in Machine Fault Diagnosis
abstract
In machine fault diagnosis, conventional data-driven models trained by empirical risk minimization (ERM) often fail to generalize across domains with distinct data distributions caused by various machine operating conditions. One major reason is that ERM primarily focuses on informativeness of data labels and lacks sufficient attention on invariance of data features. To enable invariance on top of informativeness, a learning framework, learning invariant features via in-label swapping for generalizing out-of-distribution (Lifeisgood), is proposed in this study. Lifeisgood is inspired by a simple intuition that invariance can be assessed by checking changes in loss due to swapping certain entries of features with the same labels. Lifeisgood also enjoys a theoretical guarantee on improving testing domain performance under certain conditions based on a swapping 0-1 loss proposed in this work. To circumvent the training difficulties associated with the swapping 0-1 loss, a swapping cross-entropy loss is derived as a surrogate and theoretical justifications for such a relaxation are also provided. As a result, Lifeisgood can be employed conveniently to develop data-driven fault diagnosis models. In the experiments, Lifeisgood outperformed the majority of state-of-the-art methods in terms of average accuracy and exceeded the second-best by 25% in terms of the frequency of beating the generic ERM. The code is available at: https://github.com/mozhenling/doge-lifeisgood.
Zhenling Mo, Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Trans. Cybern.2
2025 Domain Generalization Study of Empirical Risk Minimization From Causal Perspectives
abstract
Empirical risk minimization (ERM) is a celebrated induction principle for developing data-driven models. However, ERM has received both pros and cons for its capability on domain generalization (DG). To this end, this paper attempts to study the success and failure of ERM at supervised DG classification tasks, both theoretically and empirically, with causal perspectives. In the theoretical aspect, we first explore different properties of a causal metric termed information flow, followed with discussing relationships between the information flow and the mutual information in the proposed causal graph. Next, we analyze the roles of the transformed causal feature and the transformed spurious feature on modeling performances. It reveals that the interaction between the spurious influencer and the transformed causal feature is the key determining the failure or success of ERM on DG. In the empirical study, we first simulate various DG settings based on the MNIST, Fashion MNIST, and CIFAR10 datasets. Next, we verify developed theories by testing three different neural network configurations in designed experiments. In addition, experiments based on real-world datasets are conducted to further consolidate key points of the proposed theories. To extend application benefits of the theoretical discoveries, a new risk minimization framework with a novel feature intervention for regulating ERM is proposed. It achieves DG improvements over ERM on real-world datasets of image segmentation, image classification, and text classification.
Zhenling Mo, Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Trans. Multim.2
2025 Extended Invariant Risk Minimization for Machine Fault Diagnosis With Label Noise and Data Shift
abstract
Incorrect labels as well as the discrepancy between training and test domain data distributions can significantly affect the effectiveness of supervised data-driven models in machine fault diagnosis applications. Such a challenge can be characterized as the noisy label-domain generalization (NL-DG) problem. In this article, the extended invariant risk minimization (EIRM) is developed, which incorporates flat minima seeking to address the NL-DG challenge. The ability of handling NL-DG is realized by shifting the gradient penalty base from the dummy classifier to the entire model. EIRM is shown to be closely related to locating a flat minimum, which is crucial for label noise (LN) robustness and model generalization. Explorations on function smoothness and algorithm convergence are offered to understand EIRM from the theoretical aspect. An efficient implementation of EIRM is also developed to construct the fault diagnosis model. The EIRM-based fault diagnosis method is compared with strong benchmarks on multiple NL-DG tasks using actuator and gearbox fault datasets. Results indicate that the EIRM-based method on average is more effective than the benchmarks. The code is available at https://github.com/mozhenling/doge-eirm.
Zhenling Mo, Zijun Zhang 0001, Qiang Miao, Kwok-Leung Tsui
IEEE Trans. Neural Networks Learn. Syst.2
2024 A Multi-task Learning Governed Temporal Convolutional Network for Predicting Rare HVAC Compressor Faults
abstract
Prediction of the compressor faults in heating, ventilation and air conditioning (HVAC) based on collected data is critical to its operational safety and energy efficiency. Main challenges include the unclear mechanism of utilizing data to effectively predict the compressor faults and the impact of fault scarcity. To make a response, this study develops a novel multitask learning governed temporal convolutional network (MTTCN) for the rare HVAC compressor fault prediction task. The MTTCN adopts a TCN-based backbone guided by three designed learning tasks to derive valuable latent features from raw long-range inputs for more robust rare fault predictions. The primary rare fault prediction task is assisted by focal loss to overcome the fault scarcity, while two auxiliary tasks, the contrastive learning and condition prediction, are employed to sharpen the feature discrimination and augment the temporal prediction capabilities of MTTCN. To ensure an efficient and intelligent optimization process, an improved uncertainty-based dynamic weighting mechanism is developed to adaptively balance the weights of task-specific losses during training. Results of experiments conducted reveal the superior performance of MTTCN over existing models in predicting rare HVAC compressor faults.
Zijun Zhang 0001
IJCNN2
2024 Distance-Aware Risk Minimization for Domain Generalization in Machine Fault Diagnosis
abstract
Industrial Internet of Things (IIoT) connects machines, and it is important to build intelligent models to prevent machine failures by identifying incipient faults. To develop intelligent fault diagnosis models, empirical risk minimization (ERM)-based modeling paradigm has been prevalently applied. However, during model training, ERM primarily focuses on instance-to-prototype (ItP) distances from a prototypical perspective, which may limit its effectiveness in analyzing data of diverse distributions. To improve the ERM model, we propose considering additional instance-to-instance distances (ItI) and prototype-to-prototype (PtP) distances, leading to a new modeling framework—distance-aware risk minimization (DARM). To gain awareness of extra types of distances, two novel losses are proposed based on reformulations of soft-max cross entropy. Theoretical explorations are conducted to justify the significance of collectively considering ItP, ItI, and PtP distances. Methodologically, DARM can jointly minimize three types of distance-aware losses to train neural networks for fault diagnosis in the same fashion as ERM. In a comprehensive computational study, DARM consistently outperformed ERM in domain generalization (DG) tasks based on various machine fault diagnosis data sets. In addition, DARM has superior performance over several recent DG methods. The code is available athttps://github.com/mozhenling/doge-darm.
Zhenling Mo, Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Internet Things J.2
2024 PFFN: Periodic Feature-Folding Deep Neural Network for Traffic Condition Forecasting
abstract
Accurate forecasting of traffic conditions is critical for improving urban transportation safety, stability, and efficiency. It is challenging to produce explicit traffic forecasts due to complex and dynamic spatiotemporal contexts. Most existing works only capture partial characteristics and features of traffic data, and there still exists a great potential of uplifting the forecasting performances. In this paper, we propose a periodic feature-folding network (PFFN) that globally folds the spatial-temporal-geo-character features of collected data into pivotal levels to enhance the forecasting performance of traffic conditions. Meanwhile, the historical and recent traffic information within the subgraph is locally folded to capture similar traffic patterns and determine meaningful high-level features in latent space via a novel gated attention plug (GAP). During forecasting, the auxiliary road attributes are joint-wise folded with locally engineered features to realize the multi-step forecasting. Experimental results on two publicly accessible real-world urban traffic datasets show that the proposed PFFN can outperform the state-of-the-art benchmarks, significantly improving the performance of forecasting short-term traffic conditions.
Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Internet Things J.2
2024 A Hierarchical Transfer-Generative Framework for Automating Multianalytical Tasks in Rail Surface Defect Inspection
abstract
Rail surface inspection is crucial for ensuring the safety and longevity of rail transport systems, grapples with the challenges posed by the scarcity of defective samples. Additionally, contemporary techniques in this domain typically fail to concurrently identify and localize defects at both image level and pixel levels. Addressing these intricacies, we present a hierarchical transfer-generative framework, the HTg-Net. This innovative framework is geared towards the automation and enhancement of multi-analytical tasks in rail surface inspection. The HTg-Net architecture synergistically melds two pivotal subnetworks: (1) the memory-guided generation subnetwork (MGN), which is endowed with a cutting-edge memory mechanism. This mechanism adeptly captures and recalls the typical patterns observed in rail images, facilitating the detection of anomalies or deviations; (2) the attention-focused segmentation subnetwork (ASN) is fortified with a gated attention mechanism and hierarchical weights transferred from MGN, enabling parallel feature extraction and enhancing defect localization. Rigorous evaluations of HTg-Net on three datasets elucidate its superior efficiency and performance over prevailing benchmarks, positioning it as an advanced solution for the comprehensive inspection of rail surface defects.
Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Internet Things J.2
2024 MNHP-GAE: A Novel Manipulator Intelligent Health State Diagnosis Method in Highly Imbalanced Scenarios
abstract
As a classical and crucial component in industrial systems, the manipulators are widely employed in precision manufacturing scenarios because of their advantages of high stiffness, large load support capability, and high precision. During their service, it is inevitable that they encounter data imbalance scenarios due to the occasional and low-frequency failure behaviors. But in order to address these issues, the majority of the approaches already in use need the assistance of extra tools. Thus, a novel intelligent health state diagnosis model, named multiple neighbor homogeneous property-embedded graph auto-encoder (MNHP-GAE), is developed to get around this restriction and apply it to the manipulators. Its core is to realize the expansion and enrichment of the feature space by mining effective complementary information from homogeneous property samples without the assistance of data augmentation and other technologies. Specifically, the wavelet decomposition reconstruction and dynamic time warping are integrated to promote the quantification of the sample similarity and enable the construction of homogeneous property graph samples. Following that, a unique graph auto-encoder module with the multi-head attention mechanism is constructed to extract complementary information from homogeneous property nodes and match it for diagnostic tasks. Finally, through a multi-case experimental validation scenario constructed by a 3-PRR planar parallel manipulator experimental platform, the superior performances of the proposed MNHP-GAE model in highly unbalanced scenarios are fully demonstrated.
Bo Zhao 0026, Qiqiang Wu, Zhenling Mo, Zijun Zhang 0001, Xianmin Zhang 0004
IEEE Internet Things J.5
2024 Sparsity-Constrained Invariant Risk Minimization for Domain Generalization With Application to Machinery Fault Diagnosis Modeling
abstract
Machine learning has been widely applied to study AI-informed machinery fault diagnosis. This work proposes a sparsity-constrained invariant risk minimization (SCIRM) framework, which develops machine-learning models with better generalization capacities for environmental disturbances in machinery fault diagnosis. The SCIRM is built by innovating the optimization formulation of the recently proposed invariant risk minimization (IRM) and its variants through the integration of sparsity constraints. We prove that if a sparsity measure is differentiable, scale invariant, and semistrictly quasi-convex, the SCIRM can be guaranteed to solve the domain generalization problem based on a few predefined problem settings. We mathematically derive a family of such sparsity measures. A practical process of implementing the SCIRM for machinery fault diagnosis tasks is offered. We first verify our theoretical exploration of the SCIRM by using simulation data. We further compare SCIRM with a set of state-of-the-art methods by using real machinery fault data collected under a variety of working conditions. The computational results confirm that the machinery fault diagnosis model developed by the SCIRM offers a higher generalization capacity and performs better than the other benchmarks across the different testing datasets.
Zhenling Mo, Zijun Zhang 0001, Qiang Miao, Kwok-Leung Tsui
IEEE Trans. Cybern.2
2024 An Anomaly-Free Representation Learning Approach for Efficient Railway Foreign Object Detection
abstract
Applying machine vision to facilitate railway anomaly detections faces a grand challenge in that anomalous samples for model training are insufficient due to their infrequent occurrence and wide diversity. An anomaly-free representation learning approach (ARLA) is developed in this article to realize a machine vision-powered railway foreign object detection (RFOD) that does not rely on anomalous samples. The ARLA consists of two components, a memory-suppress diffusion network module and a contrastive dissimilarity network. The former network module well considers the diversity of normal patterns and reconstructs high-fidelity normal images. The latter network module enables image-level and pixelwise foreign object detections based on well-defined dissimilarity scores and distance maps. The ARLA realizes an efficient RFOD, which leverages only normal images in training and does not compromise the detection performance at the inference stage. Improvements offered by ARLA in terms of pixelwise detection performance and model complexity against two groups of benchmarks have been consistently observed based on computational studies using the railway dataset.
Zijun Zhang 0001, Min Xie 0001
IEEE Trans. Ind. Informatics2
2024 A Stochastic Recurrent Encoder Decoder Network for Multistep Probabilistic Wind Power Predictions
abstract
In this article, a stochastic recurrent encoder decoder neural network (SREDNN), which considers latent random variables in its recurrent structures, is developed for the first time for the generative multistep probabilistic wind power predictions (MPWPPs). The SREDNN enables the stochastic recurrent model under the encoder-decoder framework to engage exogenous covariates to produce better MPWPP. The SREDNN consists of five components, the prior network, the inference network, the generative network, the encoder recurrent network, and the decoder recurrent network. The SREDNN is equipped with two critical advantages compared with conventional RNN-based methods. First, the integration over the latent random variable builds an infinite Gaussian mixture model (IGMM) as the observation model, which drastically increases the expressiveness of the wind power distribution. Secondly, hidden states of the SREDNN are updated in a stochastic way, which builds an infinite mixture of the IGMM for describing the ultimate wind power distribution and enables the SREDNN to model complex patterns across wind speed and wind power sequences. Computational experiments are conducted on a dataset of a commercial wind farm having 25 wind turbines (WTs) and two publicly assessable WT datasets to verify the advantages and effectiveness of the SREDNN for MPWPP. Experimental results show that the SREDNN achieves a lower negative form of the continuously ranked probability score (CRPS*) as well as a superior sharpness and comparable reliability of prediction intervals by comparing against considered benchmarking models. Results also show the clear benefit gained from considering latent random variables in SREDNN.
Zijun Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 A Deep Generative Approach for Rail Foreign Object Detections via Semisupervised Learning
abstract
The automated inspection and detection of foreign objects help prevent potential accidents and train derailments. Most existing approaches focus on the detection with prior labels, such as categories and locations of objects, and do not directly address detecting foreign objects of unknown categories, which can appear anytime on the rail track site. In this article, we develop a deep generative approach for detecting foreign objects without predefining the scope of objects. The detection procedure consists of the following three steps: first, the model composed of an autoencoder and a discriminator is developed via adversarial training based on normal rail images only; second, the detection of abnormal rail images is implemented based on the anomaly score obtained via the trained autoencoder; and finally, foreign objects are detected by filtering the subtle dissimilarity in normal areas and highlighting abnormal areas. The effectiveness of the proposed framework for the rail foreign object detection is validated with images collected by a train equipped with visual sensors. Computational results demonstrate that our proposal is capable to achieve an impressive performance on detecting numerous foreign objects. Moreover, two groups of benchmarking methods are employed to verify the superiority of the proposed framework.
Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Trans. Ind. Informatics2
2023 CAMV: A Crash Alarm Model for Vehicles Based on Internet of Vehicles Data
abstract
Predicting vehicle crashes is critical to improve urban transportation safety. It is however challenging to alarm impending crashes accurately as the volume of non-crash data dominates that of crash data. Most existing studies formulate the crash prediction as retrospective study and a binary classification problem, which offers the limited capacity of alarming crashes in advance. To fill such gap, this paper proposes a crash alarm model for vehicles (CAMV) developed via learning vehicle operational data. First, a non-crash learning block (NCLB) is developed to process time slices of operational data with a fixed length. It aims to model the inter-feature correlations and the temporal dependencies jointly. Next, a novel coarse-to-fine strategy is developed to selectively generate and calibrate alarms. Specifically, a statistical process control (SPC) module is developed at the coarse step to generate preliminary alarms with crash scores under an abnormal pattern. A self-calibration module (SCM) is next developed to eliminate false alarms and confirm the ultimate crash alarm. Experimental results verify the effectiveness of the proposed CAMV with crash data collected from Internet of Vehicles. Meanwhile, comparative experiments based on a set of benchmarking methods are conducted to verify the superiority of our proposal.
Zijun Zhang 0001, Kwok-Leung Tsui
IEEE Trans. Intell. Transp. Syst.2
2022 A Deep-Learning-Powered Near-Real-Time Detection of Railway Track Major Components: A Two-Stage Computer-Vision-Based Method
abstract
A deep-learning-powered two-stage method for automating the inspection of railway track major components is developed in this article. Rails and two types of fasteners: 1) bolts and 2) clippers, are considered as major targeted objects in this study. Based on railway images, the developed method realizes the accurate railway track inspection via two stages: 1) the initial detection and 2) the detection calibration. At stage I, a squeeze and excitation participated YOLOv3 model is developed to generate initial detection results. A domain-logic-based hybrid model (DLHM) developed with the domain knowledge is introduced to enhance the detection performance at stage II. The DLHM consists of two modules: 1) a module for the problematic region calibration and 2) another module for the symmetric region calibration. The developed DLHM offers a high probability on inspecting overlooked or misclassified interested objects generated from stage I. The effectiveness of the proposed method for detecting railway tracks is validated with field collected railway images. An overall 95.2% mAP can be achieved via the proposed method. Four state-of-the-art deep-learning-based methods are considered as benchmarks to verify advantages of the proposed method. Via a deep comparative analytics, we show that the proposed method offers a state-of-the-art performance in the railway track major component inspection task.
Li Zhuang, Haoyang Qi, Zijun Zhang 0001
IEEE Internet Things J.4
2022 Denoising temporal convolutional recurrent autoencoders for time series classification
Zijun Zhang 0001, Long Wang 0015, Xiong Luo
Inf. Sci.2
2022 A Continual Learning-Based Framework for Developing a Single Wind Turbine Cybertwin Adaptively Serving Multiple Modeling Tasks
abstract
This article proposes a generalized neural continual learning-based cybertwin (GNC) modeling framework to realize developing one wind turbine (WT) cybertwin serving multiple modeling tasks in the wind farm operations and maintenance (O&M). A generalized WT cybertwin modeling problem, which considers modeling one cybertwin for multiple tasks without additional computational burden, is studied for the first time. Fully connected neural networks are adopted as the backbone for developing the GNC model. The online elastic weight consolidation method is incorporated to mitigate the catastrophic forgetting phenomenon among different modeling tasks. Computational experiments are conducted to validate the effectiveness of the proposed GNC framework based on the supervisory control and data acquisition data. Modeling tasks in three important problems of the wind farm O&M, the WT gearbox failure detection, WT blade breakage detection, and wind power prediction, are considered in the experiment. Compared with other benchmarking models, such as the multiple neural cybertwins, neural cybertwin, and regularized neural cybertwin, the proposed GNC can achieve high accuracies on both new tasks and existing tasks, which further verifies the WT cybertwin generalization via the proposed GNC.
Luoxiao Yang, Long Wang 0015, Zijun Zhang 0001
IEEE Trans. Ind. Informatics4
2022 The Automatic Rail Surface Multi-Flaw Identification Based on a Deep Learning Powered Framework
abstract
Rails of unhealthy conditions are considered as major targets in the rail surface inspection and this study focuses on inspecting five types of major rail surface flaws, corrugations, defects, the shelling, squats, and grinding marks, via analyzing railway images. We propose a deep learning powered rail surface multi-flaw identification framework composed of two main components, a novel rail extractor for extracting rails from the background and a cascading rail surface flaw identifier for precisely identifying different flaws. The novelty of the cascading rail surface flaw identifier includes: 1) An unhealthy rail detector developed based on a DenseNet backbone for recognizing the healthy/unhealthy status on rail surfaces and 2) A rail flaw classifier for identifying flaw types on unhealthy rail surfaces. A new feature joint learning process integrating latent features derived from selected hierarchies of the DenseNet backbone as well as two traditional feature extractors, the local binary pattern and the gray level co-occurrence matrix, is developed to facilitate the rail flaw classifier to offer accurate identification results. The effectiveness of the proposed framework for rail surface multi-flaw identification is validated with datasets provided by an industrial partner in China and collected from online sources. Based on collected datasets, the proposed framework is capable to identify the rail with unhealthy conditions and its flaw type. The overall identification performance can achieve a 98.2% accuracy. Three groups of benchmarking methods are employed to verify advantages of the proposed framework. Computational results demonstrate the impressive performance of the proposed framework in the rail surface multi-flaw identification and its applicability on new datasets.
Li Zhuang, Haoyang Qi, Zijun Zhang 0001
IEEE Trans. Intell. Transp. Syst.3
2021 Soil-Moisture-Sensor-Based Automated Soil Water Content Cycle Classification With a Hybrid Symbolic Aggregate Approximation Algorithm
abstract
This article proposes a hybrid symbolic aggregate approximation and vector space model (SAX-VSM) method for automatically classifying soil water content cycles. In the proposed method, a novel similarity measure, the distance weighted cosine (DWC) similarity measure, is introduced to improve the classification performance of the SAX-VSM. The DWC similarity measure incorporates both direction and distance information of feature vectors. Meanwhile, a mixed-integer optimization problem is formulated to determine hyperparameters. An extended Rao-1 algorithm, I-Rao-1 algorithm, is developed to solve such optimization problems. To verify the feasibility and effectiveness of the proposed method, three soil moisture data sets collected from the Florida research trials are employed. Compared with state-of-the-art methods, the proposed method has achieved the best performance based on all data sets in terms of the highest accuracy, precision, and recall values. Therefore, it is promising to apply the proposed method into real applications in the smart irrigation system.
Zhongju Wang 0002, Long Wang 0015, Chao Huang 0002, Zijun Zhang 0001, Xiong Luo
IEEE Internet Things J.4
2021 A Conditional Convolutional Autoencoder-Based Method for Monitoring Wind Turbine Blade Breakages
abstract
The wind turbine blade breakage is a catastrophic failure to a wind farm. Its earlier detection is critical to prevent the unscheduled downtime and loss of whole assets. This article presents a conditional convolutional autoencoder-based monitoring method, which is of twofold, for identifying wind turbine blade breakages. First, a novel conditional convolutional autoencoder taking a multivariate set of data as input is developed to derive reconstruction errors, which reflect changes of system dynamics caused by impending blade breakages. Next, a statistical process control principle is applied to develop boundaries for triggering blade breakage alarms based on reconstruction errors. The effectiveness of the conditional convolutional autoencoder-based method is validated with datasets collected by supervisory control and data acquisition systems installed in multiple commercial wind farms. We also demonstrate advantages of the conditional convolutional autoencoder-based monitoring method by benchmarking against the classical autoencoder and conditional autoencoder-based monitoring methods.
Luoxiao Yang, Zijun Zhang 0001
IEEE Trans. Ind. Informatics2
2018 Short-Term Wind Speed Forecasting via Stacked Extreme Learning Machine With Generalized Correntropy
abstract
Recently, wind speed forecasting as an effective computing technique plays an important role in advancing industry informatics, while dealing with these issues of control and operation for renewable power systems. However, it is facing some increasing difficulties to handle the large-scale dataset generated in these forecasting applications, with the purpose of ensuring stable computing performance. In response to such limitation, this paper proposes a more practical approach through the combination of extreme-learning machine (ELM) method and deep-learning model. ELM is a novel computing paradigm that enables the neural network (NN) based learning to be achieved with fast training speed and good generalization performance. The stacked ELM (SELM) is an advanced ELM algorithm under deep-learning framework, which works efficiently on memory consumption decrease. In this paper, an enhanced SELM is accordingly developed via replacing the Euclidean norm of the mean square error (MSE) criterion in ELM with the generalized correntropy criterion to further improve the forecasting performance. The advantage of the enhanced SELM with generalized correntropy to achieve better forecasting performance mainly relies on the following aspect. Generalized correntropy is a stable and robust nonlinear similarity measure while employing machine learning method to forecast wind speed, where the outliers may exist in some industrially measured values. Specifically, the experimental results of short-term and ultra-short-term forecasting on real wind speed data show that the proposed approach can achieve better computing performance compared with other traditional and more recent methods.
Xiong Luo, Jiankun Sun, Long Wang 0015, Weiping Wang 0007, Wenbing Zhao 0001, Jinsong Wu 0001, Jenq-Haur Wang, Zijun Zhang 0001
IEEE Trans. Ind. Informatics8
2018 Differential Evolution With a New Encoding Mechanism for Optimizing Wind Farm Layout
abstract
This paper presents a differential evolution algorithm with a new encoding mechanism for efficiently solving the optimal layout of the wind farm, with the aim of maximizing the power output. In the modeling of the wind farm, the wake effects among different wind turbines are considered and the Weibull distribution is employed to estimate the wind speed distribution. In the process of evolution, a new encoding mechanism for the locations of wind turbines is designed based on the characteristics of the wind farm layout. This encoding mechanism is the first attempt to treat the location of each wind turbine as an individual. As a result, the whole population represents a layout. Compared with the traditional encoding, the advantages of this encoding mechanism are twofold: 1) the dimension of the search space is reduced to two, and 2) a crucial parameter (i.e., the population size) is eliminated. In addition, differential evolution serves as the search engine and the caching technique is adopted to enhance the computational efficiency. The comparative analysis between the proposed method and seven other state-of-the-art methods is conducted based on two wind scenarios. The experimental results indicate that the proposed method is able to obtain the best overall performance, in terms of the power output and execution time.
Yong Wang 0002, Hao Liu 0024, Huan Long, Zijun Zhang 0001, Shengxiang Yang
IEEE Trans. Ind. Informatics4
2017 Wind Turbine Gearbox Failure Identification With Deep Neural Networks
abstract
The feasibility of monitoring the health of wind turbine (WT) gearboxes based on the lubricant pressure data in the supervisory control and data acquisition system is investigated in this paper. A deep neural network (DNN)-based framework is developed to monitor conditions of WT gearboxes and identify their impending failures. Six data-mining algorithms, thek-nearest neighbors, least absolute shrinkage and selection operator, ridge regression (Ridge), support vector machines, shallow neural network, as well as DNN, are applied to model the lubricant pressure. A comparative analysis of developed data-driven models is conducted and the DNN model is the most accurate. To prevent the overfitting of the DNN model, a dropout algorithm is applied into the DNN training process. Computational results show that the prediction error will shift before the occurrences of gearbox failures. An exponentially weighted moving average control chart is deployed to derive criteria for detecting the shifts. The effectiveness of the proposed monitoring approach is demonstrated by examining real cases from wind farms in China and benchmarked against the gearbox monitoring based on the oil temperature data.
Long Wang 0015, Zijun Zhang 0001, Huan Long, Ruihua Liu
IEEE Trans. Ind. Informatics2
2016 Wind Turbine Modeling With Data-Driven Methods and Radially Uniform Designs
abstract
This paper proposes a radially uniform (RU) design to sample representative datasets from a large volume of wind turbine data to build accurate data-driven models. The sampling capability and computational complexity are theoretically analyzed. It is shown that the RU design is representative of the original dataset and has computational complexity that is of the same order as sorting algorithms. Five algorithms, the neural networks (NN), multivariate adaptive regression splines (MARS), support vector machines (SVM), k nearest neighbors (kNN), and linear regression (LR) are applied to model the wind turbine power output, drive-train vibratory acceleration, and tower vibratory acceleration based on the training dataset and sampled datasets. Extensive computational experiments are conducted to demonstrate advantages of the RU sampler over the random and maximin samplers. Results show that RU sampler outperforms the random sampler for building all five types of models and is more effective than the maximin sampler for building nonlinear models.
Matthias H. Y. Tan, Zijun Zhang 0001
IEEE Trans. Ind. Informatics2
2015 Data-driven minimization of pump operating and maintenance cost
Zijun Zhang 0001, Xiaofei He 0006, Andrew Kusiak
Eng. Appl. Artif. Intell.1
2013 Modeling and analysis of pumps in a wastewater treatment plant: A data-mining approach
Andrew Kusiak, Yaohui Zeng, Zijun Zhang 0001
Eng. Appl. Artif. Intell.3