Yanshan Xiao

dblp:65/4144 · DBLP profile ↗
← Back
45ranked-venue papers in the field
9as first author
29since 2021 · last 2026
0000-0002-5788-042XORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 28 (6 first)Data Mining & Knowledge Discovery · 12 (2 first)Database Systems & Data Management · 3Information Retrieval & Web Search · 2 (1 first)
YearPublicationVenuePosition
2026 DyFASA: Dynamic frequency-domain-aware spatial-channel attention for efficient lung disease detection from chest x-rays
Bo Liu 0002, Qingshan Tang, Yanshan Xiao, Weijie Zeng, Xinzhe Jiang, Yunlong Sun, Zongxiong Yang
Inf. Sci.3
2025 Boosting one-class transfer learning for multiple view uncertain data
Bo Liu 0002, Fan Cao, Yanshan Xiao
Inf. Sci.4
2025 Semi-supervised manifold regularized multi-task learning with privileged information
Bo Liu 0002, Baoqing Li, Yanshan Xiao, Zhitong Wang, Boxu Zhou, Shengxin He, Chenlong Ye, Fan Cao
Inf. Sci.3
2025 A multi-view forward positive and unlabeled graph learning method based on dictionary learning
Bo Liu 0002, Chenlong Ye, Yanshan Xiao, Baoqing Li, Zhitong Wang, Boxu Zhou, Shengxin He, Fan Cao
Inf. Sci.3
2025 Multi-view ordinal regression with feature augmentation and privileged information learning
Yanshan Xiao, Linbin Chen, Bo Liu 0002
Inf. Sci.1
2024 Privacy preservation-based federated learning with uncertain data
abstract
Federated learning (FL) belongs to distributed machine learning . It allows data information sharing between users while protecting their data privacy at the same time. However, in many real-world scenarios, the data collected by client devices may be affected by noise in the working environment, leading to the decreased accuracy and confidence. Therefore, it is necessary to take measures to reduce data uncertainty in order to enhance the performance of FL algorithms. Traditional FL methods encounter challenges in handling uncertain data, which motivates the introduction of multi-view learning in this paper which designs an FL model suitable for highly variable data characteristics . We first achieve information complementarity among data views while ensuring the consistency in data representation . Furthermore, we quantify data uncertainty using the reachable region of data noise, thereby improving model robustness. To maintain data privacy between clients, we design an adaptive Kalman filter-based differential protection security protocol . Clients use the protocol to process local data and upload it to the master server, which returns the updated model parameters to the clients. The experimental results demonstrate the effectiveness of the federated learning model proposed in this paper.
Fan Cao, Bo Liu 0002, Yanshan Xiao
Inf. Sci.5
2024 Dictionary-based multi-instance learning method with universum information
Fan Cao, Bo Liu 0002, Yanshan Xiao
Inf. Sci.4
2024 The multi-task transfer learning for multiple data streams with uncertain data
Bo Liu 0002, Yanshan Xiao, Zhiyu Zheng, Xiaokai Li 0002, Tiantian Peng
Inf. Sci.3
2024 Self-paced method for transfer partial label learning
Bo Liu 0002, Zhiyu Zheng, Yanshan Xiao, Xiaokai Li 0002, Tiantian Peng
Inf. Sci.3
2024 A new multi-view multi-label model with privileged information learning
Yanshan Xiao, Bo Liu 0002, Xiangjun Kong
Inf. Sci.1
2024 An Efficient Transfer Learning Method with Auxiliary Information
abstract
Transfer learning (TL) is an information reuse learning tool, which can help us learn better classification effect than traditional single task learning, because transfer learning can share information within the task-to-task model. Most TL algorithms are studied in the field of data improvement, doing some data extraction and transformation. However, it ignores that existing the additional information to improve the model’s accuracy, like Universum samples in the training data with privileged information. In this article, we focus on considering prior data to improve the TL algorithm, and the additional features also called privileged information are incorporated into the learning to improve the learning paradigm. In addition, we also carry out the Universum samples which do not belong to any indicated categories into the transfer learning paradigm to improve the utilization of prior knowledge. We propose a new TL Model (PU-TLSVM), in which each task with corresponding privileged features and Universum data is considered in the proposed model, so as to apply tasks with a priori data to the training stage. Then, we use Lagrange duality theorem to optimize our model to obtain the optimal discriminant for target task classification. Finally, we make a lot of predictions and tests to compare the actual effectiveness of the proposed method with the previous methods. The experiment results indicate that the proposed method is more effective and robust than other baselines.
Bo Liu 0002, Liangjiao Li, Yanshan Xiao, Ruiguang Huang
ACM Trans. Knowl. Discov. Data3
2023 Self-paced multi-view positive and unlabeled graph learning with auxiliary information
Bo Liu 0002, Tiantian Peng, Yanshan Xiao, Xiaokai Li 0002, Zhiyu Zheng
Inf. Sci.3
2023 Semi-supervised Multi-task Learning with Auxiliary data
Bo Liu 0002, Yanshan Xiao, Ruiguang Huang, Liangjiao Li
Inf. Sci.3
2023 Multi-view multi-label learning with high-order label correlation
Bo Liu 0002, Weibin Li 0003, Yanshan Xiao, Laiwang Liu, Changdong Liu
Inf. Sci.3
2023 Multi-task ordinal regression with labeled and unlabeled data
Yanshan Xiao, Liangwang Zhang, Bo Liu 0002, Ruichu Cai
Inf. Sci.1
2023 A Ranking Based Multi-View Method for Positive and Unlabeled Graph Classification
Bo Liu 0002, Zhiyong Che, Haowen Zhong, Yanshan Xiao
IEEE Trans. Knowl. Data Eng.4
2022 Dictionary-based transfer learning with Universum data
Zhiyong Che, Bo Liu 0002, Yanshan Xiao, Luyue Lin
Inf. Sci.3
2022 Adaptive robust Adaboost-based twin support vector machine with universum data
Bo Liu 0002, Ruiguang Huang, Yanshan Xiao, Liangjiao Li
Inf. Sci.3
2022 A new self-paced learning method for privilege-based positive and unlabeled learning
Bo Liu 0002, Yanshan Xiao, Ruiguang Huang, Liangjiao Li
Inf. Sci.3
2022 AdaBoost-based transfer learning with privileged information
Bo Liu 0002, Laiwang Liu, Yanshan Xiao, Changdong Liu, Weibin Li 0003
Inf. Sci.3
2022 Multi-view support vector ordinal regression with data uncertainty
Yanshan Xiao, Bo Liu 0002, Xiangjun Kong, Adi Alhudhaif, Fayadh Alenezi
Inf. Sci.1
2022 Multi-view partial label machine
Yanshan Xiao, Bo Liu 0002
Inf. Sci.2
2022 Multi-task manifold learning for partial label learning
Yanshan Xiao, Kairun Wen, Bo Liu 0002, Xiangjun Kong
Inf. Sci.2
2022 A Latent Variable Augmentation Method for Image Categorization with Insufficient Training Samples
abstract
Over the past few years, we have made great progress in image categorization based on convolutional neural networks (CNNs). These CNNs are always trained based on a large-scale image data set; however, people may only have limited training samples for training CNN in the real-world applications. To solve this problem, one intuition is augmenting training samples. In this article, we propose an algorithm called Lavagan ( La tent V ariables A ugmentation Method based on G enerative A dversarial N ets) to improve the performance of CNN with insufficient training samples. The proposed Lavagan method is mainly composed of two tasks. The first task is that we augment a number latent variables (LVs) from a set of adaptive and constrained LVs distributions. In the second task, we take the augmented LVs into the training procedure of the image classifier. By taking these two tasks into account, we propose a uniform objective function to incorporate the two tasks into the learning. We then put forward an alternative two-play minimization game to minimize this uniform loss function such that we can obtain the predictive classifier. Moreover, based on Hoeffding’s Inequality and Chernoff Bounding method, we analyze the feasibility and efficiency of the proposed Lavagan method, which manifests that the LV augmentation method is able to improve the performance of Lavagan with insufficient training samples. Finally, the experiment has shown that the proposed Lavagan method is able to deliver more accurate performance than the existing state-of-the-art methods.
Luyue Lin, Xin Zheng 0001, Bo Liu 0002, Wei Chen 0103, Yanshan Xiao
ACM Trans. Knowl. Discov. Data5
2022 New Multi-View Classification Method with Uncertain Data
abstract
Multi-view classification aims at designing a multi-view learning strategy to train a classifier from multi-view data, which are easily collected in practice. Most of the existing works focus on multi-view classification by assuming the multi-view data are collected with precise information. However, we always collect the uncertain multi-view data due to the collection process is corrupted with noise in real-life application. In this case, this article proposes a novel approach, called uncertain multi-view learning with support vector machine (UMV-SVM) to cope with the problem of multi-view learning with uncertain data. The method first enforces the agreement among all the views to seek complementary information of multi-view data and takes the uncertainty of the multi-view data into consideration by modeling reachability area of the noise. Then it proposes an iterative framework to solve the proposed UMV-SVM model such that we can obtain the multi-view classifier for prediction. Extensive experiments on real-life datasets have shown that the proposed UMV-SVM can achieve a better performance for uncertain multi-view classification in comparison to the state-of-the-art multi-view classification methods.
Bo Liu 0002, Haowen Zhong, Yanshan Xiao
ACM Trans. Knowl. Discov. Data3
2021 Twin support vector machines with privileged information
Zhiyong Che, Bo Liu 0002, Yanshan Xiao
Inf. Sci.3
2021 An efficient dictionary-based multi-view learning method
Bo Liu 0002, Yanshan Xiao, Weibin Li 0003, Laiwang Liu, Changdong Liu
Inf. Sci.3
2021 A nearest-neighbor search model for distance metric learning
Yibang Ruan, Yanshan Xiao, Bo Liu 0002
Inf. Sci.2
2021 Transfer learning-based one-class dictionary learning for recommendation data stream
Haoxin Xie, Bo Liu 0002, Yanshan Xiao
Inf. Sci.3
2020 Semi-Supervised Multi-view clustering based on orthonormality-constrained nonnegative matrix factorization
Bo Liu 0002, Yanshan Xiao, Luyue Lin
Inf. Sci.3
2020 A new transfer learning-based method for label proportions problem
Yanshan Xiao, HuaiPei Wang, Bo Liu 0002
Inf. Sci.1
2020 A new self-paced method for multiple instance boosting learning
Yanshan Xiao, Xiaozhou Yang, Bo Liu 0002
Inf. Sci.1
2015 A robust one-class transfer learning method with uncertain data
Yanshan Xiao, Bo Liu 0002, Philip S. Yu
Knowl. Inf. Syst.1
2014 An efficient orientation distance-based discriminative feature extraction method for multi-classification
Bo Liu 0002, Yanshan Xiao, Philip S. Yu, Longbing Cao
Knowl. Inf. Syst.2
2014 Uncertain One-Class Learning and Concept Summarization Learning on Uncertain Data Streams
abstract
This paper presents a novel framework to uncertain one-class learning and concept summarization learning on uncertain data streams. Our proposed framework consists of two parts. First, we put forward uncertain one-class learning to cope with data of uncertainty. We first propose a local kernel-density-based method to generate a bound score for each instance, which refines the location of the corresponding instance, and then construct an uncertain one-class classifier (UOCC) by incorporating the generated bound score into a one-class SVM-based learning phase. Second, we propose a support vectors (SVs)-based clustering technique to summarize the concept of the user from the history chunks by representing the chunk data using support vectors of the uncertain one-class classifier developed on each chunk, and then extend k-mean clustering method to cluster history chunks into clusters so that we can summarize concept from the history chunks. Our proposed framework explicitly addresses the problem of one-class learning and concept summarization learning on uncertain one-class data streams. Extensive experiments on uncertain data streams demonstrate that our proposed uncertain one-class learning method performs better than others, and our concept summarization method can summarize the evolving interests of the user from the history chunks.
Bo Liu 0002, Yanshan Xiao, Philip S. Yu, Longbing Cao, Yun Zhang 0001
IEEE Trans. Knowl. Data Eng.2
2014 An Efficient Approach for Outlier Detection with Imperfect Data Labels
abstract
The task of outlier detection is to identify data objects that are markedly different from or inconsistent with the normal set of data. Most existing solutions typically build a model using the normal data and identify outliers that do not fit the represented model very well. However, in addition to normal data, there also exist limited negative examples or outliers in many applications, and data may be corrupted such that the outlier detection data is imperfectly labeled. These make outlier detection far more difficult than the traditional ones. This paper presents a novel outlier detection approach to address data with imperfect labels and incorporate limited abnormal examples into learning. To deal with data with imperfect labels, we introduce likelihood values for each input data which denote the degree of membership of an example toward the normal and abnormal classes respectively. Our proposed approach works in two steps. In the first step, we generate a pseudo training dataset by computing likelihood values of each example based on its local behavior. We present kernel \(k\) -means clustering method and kernel LOF-based method to compute the likelihood values. In the second step, we incorporate the generated likelihood values and limited abnormal examples into SVDD-based learning framework to build a more accurate classifier for global outlier detection. By integrating local and global outlier detection, our proposed method explicitly handles data with imperfect labels and enhances the performance of outlier detection. Extensive experiments on real life datasets have demonstrated that our proposed approaches can achieve a better tradeoff between detection rate and false alarm rate as compared to state-of-the-art outlier detection approaches.
Bo Liu 0002, Yanshan Xiao, Philip S. Yu, Longbing Cao
IEEE Trans. Knowl. Data Eng.2
2013 One-Class Transfer Learning with Uncertain Data
Bo Liu 0002, Philip S. Yu, Yanshan Xiao
PAKDD (1)3
2013 Robust Textual Data Streams Mining Based on Continuous Transfer Learning
abstract
In textual data stream environment, concept drift can occur at any time, existing approaches partitioning streams into chunks can have problem if the chunk boundary does not coincide with the change point which is impossible to predict. Since concept drift can occur at any point of the streams, it will certainly occur within chunks, which is called random concept drift. The paper proposed an approach, which is called chunk level-based concept drift method (CLCD), that can overcome this chunking problem by continuously monitoring chunk characteristics to revise the classifier based on transfer learning in positive and unlabeled (PU) textual data stream environment. Our proposed approach works in three steps. In the first step, we propose core vocabulary-based criteria to justify and identify random concept drift. In the second step, we put forward the extension of LELC (PU learning by extracting likely positive and negative micro-clusters)[1], called soft-LELC, to extract representative examples from unlabeled data, and assign a confidence score to each extracted example. The assigned confidence score represents the degree of belongingness of an example towards its corresponding class. In the third step, we set up a transfer learning-based SVM to build an accurate classifier for the chunks where concept drift is identified in the first step. Extensive experiments have shown that CLCD can capture random concept drift, and outperforms state-of-the-art methods in positive and unlabeled textual data stream environments.
Longbing Cao, Bo Liu 0002, Yanshan Xiao, Philip S. Yu
SDM4
2013 MODS: Multiple One-class Data Streams Learning from Homogeneous Data
abstract
This paper presents a novel approach, called MODS, to build an accurate time evolving classifier from multiple one-class data streams learning time evolving classifier. Our proposed MODS approach works in two steps. In the first step, we first construct local one-class classifiers on the labeled positive examples from each sub-data stream respectively. We then collect the informative examples (support vectors) around each local one-class classifier, which can support the decision boundary of the classifier. This is called support vector preservation principle. In the second step, we construct a global one-class classifier on the collected informative examples. By using the support vector preservation principle, our proposed MODS explicitly addresses the problem of building accurate classifier from multiple one-class data streams. Extensive experiments on real life data streams have demonstrated that our MODS approach can achieve high performance and efficiency for the multiple one-class data streams learning in comparison with other approaches.
Bo Liu 0002, Yanshan Xiao, Philip S. Yu
SDM3
2013 SVDD-based outlier detection on uncertain data
Bo Liu 0002, Yanshan Xiao, Longbing Cao, Feiqi Deng
Knowl. Inf. Syst.2
2011 One-Class-Based Uncertain Data Stream Learning
abstract
This paper presents a novel approach to one-class-based uncertain data stream learning. Our proposed approach works in three steps. Firstly, we put forward a local kernel-density-based method to generate a bound score for each instance, which refines the location of the corresponding instance. Secondly, we construct an uncertain one-class classifier by incorporating the generated bound score into a one-class SVM-based learning phase. Thirdly, we devise an ensemble classifier, integrated from uncertain one-class classifiers built on the current and historical chunks, to cope with the concept drift involved in the uncertain data stream environment. Our proposed method explicitly handles the uncertainty of the input data and enhances the ability of one-class learning in reducing the sensitivity to noise. Extensive experiments on uncertain data streams demonstrate that our proposed approach can achieve better performance and is highly robust to noise in comparison with state-of-the-art one-class learning method.
Bo Liu 0002, Yanshan Xiao, Longbing Cao, Philip S. Yu
SDM2
2010 Orientation distance-based discriminative feature extraction for multi-class classification
abstract
Feature extraction is an effective step in data mining and machine learning. While many feature extraction methods have been proposed for clustering, classification and regression, very limited work has been done on multi-class classification problems. In fact, the accuracy of multi-class classification problems relies on well-extracted features, the modeling part aside. This paper proposes a new feature extraction method, namely extracting orientation distance-based discriminative (ODD) features, which is particularly designed for multi-class classification problems. The proposed method works in two steps. In the first step, we extend the Fisher Discriminant idea to determine more appropriate kernel function and map the input data with all classes into a feature space. In the second step, the ODD features are extracted based on the one-vs-all scheme to generate discriminative features between a pattern and each hyperplane. These newly extracted features are treated as the representative features and are further used in the subsequent classification procedure. Substantial experiments on both UCI and real-world datasets have been conducted to investigate the performance of ODD features based multi-class classification. The statistical results show that the classification accuracy based on ODD features outperforms that of the state-of-the-art feature extraction methods.
Bo Liu 0002, Yanshan Xiao, Longbing Cao, Philip S. Yu
CIKM2
2010 K-farthest-neighbors-based concept boundary determination for support vector data description
abstract
Support vector data description (SVDD) is very useful for one-class classification. However, it incurs high time complexity in handling large scale data. In this paper, we propose a novel and efficient method, named K-Farthest-Neighbors-based Concept Boundary Detection (KFN-CBD for short), to improve the SVDD learning efficiency on large datasets. This work is motivated by the observation that SVDD classifier is determined by support vectors (SVs), and removing the non-support vectors (non-SVs) will not change the classifier but will reduce computational costs. Our approach consists of two steps. In the first step, we propose the K-farthest-neighbors method to identify the samples around the hyper-sphere surface, which are more likely to be SVs. At the same time, a new tree search strategy of M-tree is presented to speed up the K-farthest neighbor query. In the second step, the non-SVs are eliminated from the training set, and only the identified boundary samples are used to train the SVDD classifier. By removing the non-SVs, the training time of SVDD can be substantially reduced.Extensive experiments have shown that KFN-CBD achieves around 6 times speedup compared to the standard SVDD, and obtains the comparable classification quality as the entire dataset used.
Yanshan Xiao, Bo Liu 0002, Longbing Cao
CIKM1
2010 Exploiting Local Data Uncertainty to Boost Global Outlier Detection
abstract
This paper presents a novel hybrid approach to outlier detection by incorporating local data uncertainty into the construction of a global classifier. To deal with local data uncertainty, we introduce a confidence value to each data example in the training data, which measures the strength of the corresponding class label. Our proposed method works in two steps. Firstly, we generate a pseudo training dataset by computing a confidence value of each input example on its class label. We present two different mechanisms: kernel k-means clustering algorithm and kernel LOF-based algorithm, to compute the confidence values based on the local data behavior. Secondly, we construct a global classifier for outlier detection by generalizing the SVDD-based learning framework to incorporate both positive and negative examples as well as their associated confidence values. By integrating local and global outlier detection, our proposed method explicitly handles the uncertainty of the input data and enhances the ability of SVDD in reducing the sensitivity to noise. Extensive experiments on real life datasets demonstrate that our proposed method can achieve a better tradeoff between detection rate and false alarm rate as compared to four state-of-the-art outlier detection algorithms.
Bo Liu 0002, Jie Yin 0001, Yanshan Xiao, Longbing Cao, Philip S. Yu
ICDM3
2010 SMILE: A Similarity-Based Approach for Multiple Instance Learning
abstract
Multiple instance learning (MIL) is a generalization of supervised learning which attempts to learn useful information from bags of instances. In MIL, the true labels of the instances in positive bags are not always available for training. This leads to a critical challenge, namely, handling the ambiguity of instance labels in positive bags. To address this issue, this paper proposes a novel MIL method named SMILE (Similarity-based Multiple Instance LEarning). It introduces a similarity weight to each instance in positive bag, which represents the instance similarity towards the positive and negative classes. The instances in positive bags, together with their similarity weights, are thereafter incorporated into the learning phase to build an extended SVM-based predictive classifier. Experiments on three real-world datasets consisting of 12 subsets show that SMILE achieves markedly better classification accuracy than state-of-the-art MIL methods.
Yanshan Xiao, Bo Liu 0002, Longbing Cao, Jie Yin 0001, Xindong Wu 0001
ICDM1