Jiawei Xiao

dblp:314/0188 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 A prior knowledge-enhanced self-supervised learning framework using time-frequency invariance for machinery intelligent fault diagnosis with small samples
Jiawei Xiao, Chao Wei 0011, Xiaoxi Ding, Wenbing Huang 0001
Eng. Appl. Artif. Intell.2
2024 Online learning for data streams with bi-dynamic distributions
Huigui Yan, Jiawei Xiao, Shina Niu, Siqi Dong, Dianlong You
Inf. Sci.3
2024 HmmSeNet: A Novel Single Domain Generalization Equipment Fault Diagnosis Under Unknown Working Speed Using Histogram Matching Mixup
abstract
Equipments regularly change working speeds during real-time production owing to process requirements. Applying deep learning models trained in a single speed domain straightforwardly to other unknown speed domains is extremely challenging single-domain generalization problem. Therefore, this article proposes a histogram matching mixup based sequential embedding network (HmmSeNet) for single-domain generalization of intelligent fault diagnosis under unknown speeds. HmmSeNet consists of four components: histogram matching mixup (HMM); sequential embedding (SE); separable convolution; and decision making. First, inspired by histogram matching and Mixup, the HMM data augmentation method is proposed. HMM is capable of synthesizing data with the same semantic information, but different distributions from a single source domain data during the training process, thus augmenting the source speed domain to the unknown speed domains. Then, SE utilizes trainable linear dimensional boosting to approximate the distribution between samples, which reduces the effect of sample amplitude distribution shifts caused by speed changes and allows the model to learn domain-invariant features. Finally, three layers of separable convolution and global average pooling are used to accomplish an accurate and robust recognition task. Experimental results on three datasets show that the proposed approach is only trained on a single speed domain, while it has good diagnostic performance on other unknown speed domains, even varying speed domains. The powerful generality and flexible deployment capability of HmmSeNet for speed changes are also demonstrated by ablation experimental analysis and dimensional analysis.
Xiaoxi Ding, Chao Wei 0011, Jiawei Xiao, Rui Liu 0036, Wenbin Huang 0002
IEEE Trans. Ind. Informatics4
2024 Online Learning for Data Streams With Incomplete Features and Labels
abstract
Online learning is critical for handling complex data streams in Big Data-related applications. This study explores a new online learning problem where both the features and labels are incomplete. Such incompleteness poses a critical challenge in determining the latent relationship between incomplete features and labels. Unfortunately, existing online learning methods only consider a few cases of incomplete feature spaces, such as trapezoidal, evolvable, and capricious data streams, limiting their applicability to this problem. To bridge this gap, this study proposes a novel algorithm ofOnlineLearning for Data Streams withIncompleteFeatures andLabels (OLIFL). OLIFL imposes no constraints on changing patterns of feature space and does not require all instances to be labeled with two-fold ideas. First, OLIFL explores the informativeness of individual features to update the classifier by dynamically maintaining global feature space and updating the informativeness matrix. Second, it estimates the label confidence of unlabeled instances to control their negative effects by limiting the error upper bound. Extensive experiments on benchmark datasets are conducted in five scenarios: three incomplete feature (trapezoidal, evolvable, and capricious) spaces, and two incomplete labels (only missing labels and missing both features and labels). In addition, we explore the sensitivity of the model to parameters, and its usability and response efficiency in handling concept drifts. The results show that OLIFL significantly outperforms its rivals. Moreover, we use OLIFL to classify a movie review task as real application verification.
Dianlong You, Huigui Yan, Jiawei Xiao, Zhen Chen 0007, Di Wu 0056, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.3
2023 Online Multi-Label Streaming Feature Selection With Label Correlation
abstract
Multi-label streaming feature selection has attracted extensive attention in diverse big data applications. However, most existing works focused on the scenarios where labels are independent, while ignoring the real scenarios that they may be interdependent and correlated with each other. This paper aims to fill this gap by developing a novel online multi-label streaming feature selection scheme by taking into account the existence of label correlation, known as (OMSFSLC). In our design, we first calculate the correlation degree between labels to obtain the label weight. Then, we integrate the mutual information and the label weight to evaluate the correlation between features and labels. In particular, it consists of three stages: 1) online significance analysis, which can determine the significant features via the correlation degree between the newly arriving features and labels; 2) online relevance analysis, which can obtain relevant features via the mutual information; and 3) online redundancy analysis, which can filter the redundant features for removal via pairwise comparison. We implement our solution and conduct extensive experiments on benchmark datasets for performance evaluations. The experimental results exhibit that OMSFSLCsignificantly outperforms the state-of-the-art methods in terms of effectiveness and efficiency.
Dianlong You, Yang Wang 0164, Jiawei Xiao, Yaojin Lin, Maosheng Pan, Zhen Chen 0007, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.3
2023 Online Learning From Incomplete and Imbalanced Data Streams
abstract
Learning with streaming data has attracted extensive research interest in recent years. Existing online learning approaches have specific assumptions regarding data streams, such as requiring fixed or varying feature spaces with explicit patterns and balanced class distributions. While the data streams generated in many real scenarios commonly have arbitrarily incomplete feature spaces and dynamic imbalanced class distributions, making existing approaches be unsuitable for real applications. To address this issue, this paper proposes a novelOnlineLearning fromIncomplete andImbalancedDataStreams (OLI$^{2}$DS) algorithm. OLI$^{2}$DS has a two-fold main idea: 1) it follows the empirical risk minimization principle to identify the most informative features of incomplete feature spaces, and 2) it develops a dynamic cost strategy to handle imbalanced class distributions in real-time by transforming F-measure optimization into a weighted surrogate loss minimization. To evaluate OLI$^{2}$DS, we compare it with state-of-the-art related algorithms in three kinds of experiments. First, we adopt 14 real datasets to simulate three scenarios of incomplete feature spaces, i.e., trapezoidal, feature evolvable, and capricious data streams. Second, based on a benchmark online analyzer, we generate 13 datasets to simulate incomplete data streams with different imbalance ratios. Third, we analyze concept drift in two simulated scenes, i.e., online learning and data stream mining, and verify the adaption of OLI$^{2}$DS on repeated concept drifts and variable imbalance ratios. The results demonstrate that OLI$^{2}$DS achieves a significantly better performance than its rivals. Besides, a real-world case study on movie review classification is conducted to elaborate on our OLI$^{2}$DS algorithm's effectiveness. Code is released athttps://github.com/youdianlong/OLI2DS.
Dianlong You, Jiawei Xiao, Yang Wang 0164, Huigui Yan, Di Wu 0056, Zhen Chen 0007, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.2
2022 Online feature selection for multi-source streaming features
Dianlong You, Miaomiao Sun, Shunpan Liang, Yang Wang 0164, Jiawei Xiao, Fuyong Yuan, Xindong Wu 0001
Inf. Sci.6