Xuegang Hu

dblp:38/5910 · DBLP profile ↗
← Back
119ranked-venue papers
13as first author
64since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 6 first-author · 43 since 2021Databases, data management, data science and information retrieval · 24 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 6 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Label-Aware Augmentation for Long-Tailed Cross-Network Node Classification
abstract
ABSTRACT Cross‐network node classification (CNC) aims to classify the nodes of an unlabeled graph by leveraging a graph with rich labelled nodes. The node representation and cross‐network discrepancy are two important issues in CNC. Most methods argue that the node representation is crucial to the performance of CNC and learn the representation based on the rich neighbourhood. However, in applications, the networks follow a long‐tailed distribution in their node degrees, that is, most nodes are tail nodes linked to a small amount of neighbours. The sparsity of neighbourhood challenges existing methods in representation and classification. To this end, we propose a label‐aware cross‐network node classification (LA‐CNC) method. First, a label‐consistent augmentation is designed for each network to enrich the representation by augmenting the neighbourhood of tail nodes. Second, a label contrast loss is introduced and combined with adversarial loss to enhance the distinguishability of cross‐network invariant representation. Extensive experiments demonstrate that our method outperforms state‐of‐the‐art methods on several datasets.
Yuhong Zhang 0002, Congmei Shi, Xuegang Hu
Expert Syst. J. Knowl. Eng.3
2026 Unified online weak multi-Label learning in drifting and imbalanced data streams
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001
Knowl. Based Syst.2
2026 Dual-module collaborative learning for open-set recognition with noisy labels
Yaojin Lin, Pei-Pei Li 0001, Xuegang Hu
Pattern Recognit.4
2026 Online multi-label classification under noisy and changing label distribution
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001, Jun Hu 0016
Pattern Recognit.2
2026 DSPL: Dual-Space Prompt Learning With Context Bias Decoupling for Open Set Recognition
abstract
Open Set Recognition (OSR) aims to accurately classify known classes and effectively reject unknown classes in open-world scenarios, which is essential for building safe and reliable intelligent systems. Existing studies have shown that auxiliary data-based methods—utilizing outlier exposure or prompt-based guidance to strengthen OSR models—have achieved significant improvements. However, these methods are highly sensitive to the selection of data. Considering the demand for domain expertise and the cost implications, synthesizing data from known samples becomes a preferable alternative. To achieve performance on par with methods that incorporate auxiliary data, it is essential to address the following key challenges: (1) obtaining meaningful pseudo-unknown samples and prompts without manual intervention, and (2) establishing a reliable recognition mechanism during the inference phase. In this paper, we propose a novel Dual-Space Prompt Learning with Context Bias Decoupling (DSPL) method to address the above issues. Specifically, we symmetrically model the same number of unknown classes as known classes to achieve a balanced class distribution in the embedding space. Learnable textual templates are designed for both types of data. Meanwhile, boundary samples are synthesized and treated as pseudo-unknown samples by decoupling discriminative and co-occurring features. Additionally, unknown sample detection is performed by integrating the maximum posterior over known and unknown classes, thereby extending the advantages of symmetric modeling from the training phase to inference. Extensive experiments demonstrate that the proposed method DSPL achieves state-of-the-art performance compared with the most advanced OSR methods.
Xuegang Hu, Yaojin Lin, Pei-Pei Li 0001
IEEE Trans. Knowl. Data Eng.2
2026 MGCD: Multiple-Granularity Cognitive Diagnosis in Intelligent Education Systems
abstract
Cognitive diagnosis (CD) is an important task in the field of intelligent education, aiming to discover the proficiency of students on knowledge concepts with response logs. In applications, different users of the tutoring system demand for a diagnosis of knowledge concepts at different granularities. However, recent methods assume that the concepts are of the same granularity and use explicit correlations between same-granularity concepts to improve the diagnosis performance. If required for diagnosing multi-granularity concepts, these methods will face diminished performance or partial invalidation. To this end, we make the first attempt for multiple-granularity cognitive diagnosis, i.e., diagnosis on coarse- and fine-grained concepts simultaneously. Specifically, in a skillful way, the same-granularity correlations are captured and embedded into concept representations in view of concept semantics and cross-granularity correlations to model the proficiency influence between concepts implicitly. Then, the specific loss for single-granularity diagnosis and the general loss for the consistency of multi-granularity are designed to train the model jointly, achieving multiple-granularity diagnosis. Extensive experiments demonstrate that our method can achieve state-of-the-art accuracy on both coarse- and fine-grained concepts.
Yuhong Zhang 0002, Tiancheng He, Chenyang Bu, Kui Yu, Xuegang Hu, Xindong Wu 0001
ACM Trans. Inf. Syst.5
2025 Multi-granularity decision information integration network for hierarchical classification via local and global constraints
Pei-Pei Li 0001, Xuegang Hu, Shengxing Bai, Yaojin Lin
Appl. Intell.3
2025 Entity alignment in noisy knowledge graph
Xuegang Hu
Appl. Intell.3
2025 A Plug-in for cognitive diagnosis method based on correlation representation under long-tailed distribution
Yuhong Zhang 0002, Tiancheng He, Shengyu Xu, Chenyang Bu, Xuegang Hu
Expert Syst. Appl.6
2025 Hypergraph representation learning for identifying circRNA-disease associations
Yang Li 0111, Xuegang Hu, Pei-Pei Li 0001, Lei Wang 0121, Zhu-Hong You
Pattern Recognit.2
2025 Weak Multi-Label Data Stream Classification Under Distribution Changes in Labels
abstract
Multi-label stream classification aims to address the challenge of dynamically assigning multiple labels to sequentially-arrived instances. In real situations, only partial labels of instances can be observed due to the expensive human annotations, and the problem of label distribution changes arises from multiple labels in a streaming mode, but few existing works jointly consider such challenges. Motivated by this, we propose the problem of weak multi-label stream classification (WMSC) and an online classification algorithm robust to weak labels. Specifically, we incrementally update the margin-based model using information from both the past model and the current incoming instance with partially observed labels. To increase the robustness to weak labels, we first adjust the classification margin of negative labels using the label causality matrix, which is constructed by the conditional probability of label pairs. Second, we introduce the label prototype matrix to regulate the margin by controlling the weighting parameter of the slack term. Additionally, to handle the potential distribution changes in labels, we utilize the instance-specific threshold via online thresholding to perform binary classification, which is formulated as a regression problem. Finally, theoretical analysis and empirical experimental results are presented to demonstrate the effectiveness of WMSC in classifying unobserved streaming instances.
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001, Jun Hu 0016
IEEE Trans. Big Data2
2025 A Cognitive Diagnosis Model With Nonlinear Dependence Between Students and Exercises
abstract
Cognitive diagnosis (CD) aims to discover students’ mastery of knowledge concepts through response logs and it is an important task in intelligence education. CD is generally performed with the advantage of students doing exercises, which is learned by linear interacting between student proficiency and exercise difficulty. However, existing methods represent the proficiency and difficulty in view of knowledge concepts independently, which is in coarse-granularity, leading to the indiscrimination for different combinations of students and exercises under the same concept. To this end, we propose a cognitive diagnosis method with nonlinear dependence between students and exercises (CDND), in which, more fine-grained information is captured to distinguish the different combinations of students and exercises. First, the nonlinear dependence is captured with a novel attention-like mechanism by interacting between students and exercises. Moreover, this nonlinear dependence is embedded as the interactive discrimination vector. Second, the interactive discrimination is used to adjust the advantage using a second-order interaction, which can strengthen the discrimination of the advantage under different combinations of students and exercises. Extensive experimental results on three real data sets validate our interactive discrimination has good compatibility with existing methods, and with it, our CDND achieves an obvious improvement in performance. Our code is available onhttps://github.com/joyce99/LinZhihao/tree/main/CDND-master.
Yuhong Zhang 0002, Chenyang Bu, Kui Yu, Xuegang Hu, Xindong Wu 0001
IEEE Trans. Comput. Soc. Syst.5
2025 Semi-Supervised Short Text Stream Classification Based on Drift-Aware Incremental Deep Learning
abstract
Real-world applications have produced massive short text streams. Contrary to the traditional normal texts, they present the characteristics such as short length, only having few labeled data, high-velocity, high-volume and dynamic data distributions, which deteriorate the issues of data sparseness, label missing and concept drift. Obviously, it is a huge challenge for existing short text (stream) classification algorithms due to the poor effectiveness, where they always assume all short texts are completely labeled and little attention is paid on the concept drift issue hidden in short text streams. Therefore, we propose a novel semi-supervised short text steam classification method based on the drift-aware incremental deep learning ensemble model. Specifically, with the sliding window mechanism, we firstly fuse three types of statistical, semantic and structure information to solve the data sparseness issue. Secondly, a semi-supervised incremental deep learning ensemble model based on GCN and the refined LSTM is developed to adapt to the high-volume, high-velocity and label missing short text streams. Thirdly, a label-probability distribution based concept drift detector is introduced to distinguish concept drifts. Finally, as compared with eleven well-known classification methods, extensive experiments demonstrate the effectiveness of the proposed method in the handling of short text streams with limited labeled data.
Pei-Pei Li 0001, Shiying Yu, Xuegang Hu
IEEE Trans. Knowl. Data Eng.4
2024 Class-Specific Semantic Generation and Reconstruction Learning for Open Set Recognition
Yaojin Lin, Pei-Pei Li 0001, Jun Hu 0016, Xuegang Hu
IJCAI5
2024 Knowledge and separating soft verbalizer based prompt-tuning for multi-label short text classification
Zhanwang Chen, Pei-Pei Li 0001, Xuegang Hu
Appl. Intell.3
2024 Prediction of Freezing of Gait in Parkinson's disease based on multi-channel time-series neural network
Xuegang Hu, Rongjun Ge, Chenchu Xu, Jinglin Zhang 0004, Zhifan Gao, Shu Zhao 0005, Kemal Polat
Artif. Intell. Medicine2
2024 Learning shared and non-redundant label-specific features for partial multi-label classification
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001, Yuhang Ge
Inf. Sci.2
2024 Lightweight multi-scale generative adversarial network with attention for image denoising
Xuegang Hu
Multim. Syst.1
2024 Multi-scale information fusion generative adversarial network for real-world noisy image denoising
Xuegang Hu
Mach. Vis. Appl.1
2024 LFFNet: lightweight feature-enhanced fusion network for real-time semantic segmentation of road scenes
Xuegang Hu, Juelin Gong
Pattern Anal. Appl.1
2024 A cross-network node classification method in open-set scenario
Yuhong Zhang 0002, Yunlong Ji, Kui Yu, Xuegang Hu, Xindong Wu 0001
Pattern Recognit.4
2024 Gradient-based multi-label feature selection considering three-way variable interaction
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001
Pattern Recognit.2
2024 Second-order Confidence Network for Early Classification of Time Series
abstract
Time series data are ubiquitous in a variety of disciplines. Early classification of time series, which aims to predict the class label of a time series as early and accurately as possible, is a significant but challenging task in many time-sensitive applications. Existing approaches mainly utilize heuristic stopping rules to capture stopping signals from the prediction results of time series classifiers. However, heuristic stopping rules can only capture obvious stopping signals, which makes these approaches give either correct but late predictions or early but incorrect predictions. To tackle the problem, we propose a novel second-order confidence network for early classification of time series, which can automatically learn to capture implicit stopping signals in early time series in a unified framework. The proposed model leverages deep neural models to capture temporal patterns and outputs second-order confidence to reflect the implicit stopping signals. Specifically, our model exploits the data not only from a time step but also from the probability sequence to capture stopping signals. By combining stopping signals from the classifier output and the second-order confidence, we design a more robust trigger to decide whether or not to request more observations from future time steps. Experimental results show that our approach can achieve superior results in early classification compared to state-of-the-art approaches.
Junwei Lv, Yuqi Chu, Jun Hu 0016, Pei-Pei Li 0001, Xuegang Hu
ACM Trans. Intell. Syst. Technol.5
2024 FDKT: Towards an Interpretable Deep Knowledge Tracing via Fuzzy Reasoning
abstract
In educational data mining, knowledge tracing (KT) aims to model learning performance based on student knowledge mastery. Deep-learning-based KT models perform remarkably better than traditional KT and have attracted considerable attention. However, most of them lack interpretability, making it challenging to explain why the model performed well in the prediction. In this paper, we propose an interpretable deep KT model, referred to as fuzzy deep knowledge tracing (FDKT) via fuzzy reasoning. Specifically, we formalize continuous scores into several fuzzy scores using the fuzzification module. Then, we input the fuzzy scores into the fuzzy reasoning module (FRM). FRM is designed to deduce the current cognitive ability, based on which the future performance was predicted. FDKT greatly enhanced the intrinsic interpretability of deep-learning-based KT through the interpretation of the deduction of student cognition. Furthermore, it broadened the application of KT to continuous scores. Improved performance with regard to both the advantages of FDKT was demonstrated through comparisons with the state-of-the-art models.
Fei Liu 0038, Chenyang Bu, Haotian Zhang 0007, Le Wu 0001, Kui Yu, Xuegang Hu
ACM Trans. Inf. Syst.6
2023 Global and Adaptive Local Label Correlation for Multi-label Learning with Missing Labels
abstract
Label missing is a major challenge in multi-label learning. Many existing methods try to use label correlation to recover ground-truth labels, but they only focus on the label correlation within the original label space, however, the label correlation learned in this way is incomplete. Thus, inspired$b$y the matrix adaptive column correlation, we propose a method to continuously adjust the label correlation matrix while the labels are filled in by adaptive column correlation learning method. Specifically, to reduce the impact of the missing labelson label correlation, the label space is firstly completed through manifold regularization while learning the local label information by adaptive column correlation learning in the complemented label space. Secondly, the global label correlation is utilized by adding a low-rank constraint to the entire label space. Finally, by jointly taking advantage of the global and adaptive local label correlation, our proposed approach achieves superior performance on both synthetic and real-world data sets from diverse domains compared to state-of-the art baselines.
Qingxia Jiang, Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu
IJCNN4
2023 Semi-supervised Short Text Classification Based On Dual-channel Data Augmentation
abstract
In practical applications, high-quality labeled data is critical for short text classification. But in many cases, it is expensive and time-consuming to obtain labeled information. Semi-supervised short text classification is hence attracting more attention. However, due to the sparsity of short texts, the performance of existing short text classification models always needs to be improved. Therefore, in this paper, we propose a semi-supervised Short text classification method based on Dual-Channel data Augmentation called SDCA. More specifically, in order to solve the sparsity of short texts, this model first adopts the multi-stage word-level TCN (Temporal Convolutional Network)-based attention to enhanced semantic features and an one-dimensional convolution-based attention mechanism to augment the relevance of surrounding short texts. Secondly, the unlabeled data are augmented by word embedding weighted augmentation and word replacement augmentation, so that the model can make full use of the unlabeled short texts and further enhance the network training. Finally, extensive experiments conducted on four benchmark datasets demonstrate the effectiveness of the proposed model on semi-supervised short text classification.
Pei-Pei Li 0001, Xuegang Hu
IJCNN3
2023 Multi-source Multi-label Feature Selection
abstract
Feature selection for multi-source multi-label data has attracted much attention, because there are a lot of scenarios that produce multiple source data with multi-labels in the real-world applications, which aggravates the problems of dimensional disaster and the label skewness. However, most of existing multi-source feature selection methods miss the multi-label issue while existing multi-label feature selection methods can not select the optimal feature set in the multi-source environment. Motivated by this, we propose a novel feature selection method, called MSMLFS. To be specific, the Inf-FS algorithm is firstly introduced to handle multi-label feature selection for each data source, which considers the label weight in the feature selection. Secondly, the over-sampling mechanism and the inter-source feature fusion method are used to handle the label skewness of multi-label data and the feature selection in multiple sources respectively. Finally, extensive experiments conducted on synthetic and real-world multi-source multi-label data sets demonstrate that the proposed method outperforms several state-of-the-art multi-source or multi-label feature selection methods.
Xiulan Yuan, Xuegang Hu, Pei-Pei Li 0001
IJCNN2
2023 Meta Multi-agent Exercise Recommendation: A Game Application Perspective
abstract
Exercise recommendation is a fundamental and important task in the E-learning system, facilitating students' personalized learning. Most existing exercise recommendation algorithms design a scoring criterion (e.g., weakest mastery, lowest historical correctness) in conjunction with experience, and then recommend the recommended knowledge concepts (KCs). These algorithms rely entirely on the scoring criteria by treating exercise recommendations as a centralized system. However, it is a complex problem for the centralized system to choose a limited number of exercises in a period of time to consolidate and learn the KCs efficiently. Moreover, different groups of students (e.g., different countries, schools, or classes) have different solutions for the same group of KCs according to their own situations, in the spirit of competency-based instructing. Therefore, we propose Meta Multi-Agent Exercise Recommendation (MMER). Specifically, we design the multi-agent exercise recommendation module, in which the KCs involved in exercises are considered agents with competition and cooperation among them. And the meta-training stage is designed to learn a robust recommendation module for new student groups. Extensive experiments on real-world datasets validate the satisfactory performance of the proposed model. Furthermore, the effectiveness of the multi-agent and meta-training part is demonstrated for the model in recommendation applications.
Fei Liu 0038, Xuegang Hu, Shuochen Liu, Chenyang Bu, Le Wu 0001
KDD2
2023 Cross-View Sample-Enriched Graph Contrastive Learning Network for Personalized Micro-video Recommendation
abstract
Micro-video recommendation has attracted extensive research attention with the increasing popularity of micro-video sharing platforms. Recently, graph contrastive learning (GCL) is adopted for enhancing the performance of graph neural network based micro-video recommendation. However, these GCL methods may suffer from the following problems: (1) they fail to fully exploit the potential of contrastive learning for ignoring or misjudging highly similar samples, and (2) the complementary recommendation effects between graph structure information and multi-modal feature information are not effectively utilized. In this paper, we propose a novel Cross-View Sample-Enriched Graph Contrastive Learning Network (CSGCL) for micro-video recommendation. Specifically, we build a collaborative learning view and a semantic learning view to learn node representations. For the collaborative learning view, we leverage similar nodes at the structure level to construct an effective collaborative contrastive objective. For the semantic learning view, we derive the k-nearest neighbor graph generated from multi-modal features as the semantic graphs and build a semantic contrastive objective for learning high-quality micro-video representations. Finally, a cross-view contrastive objective is designed to consider the mutually complementary recommendation effects by maximizing the agreement between the two above views. Extensive experiments on three real-world datasets demonstrate that the proposed model outperforms the baselines.
Ying He 0008, Gong-Qing Wu, Desheng Cai, Xuegang Hu
ICMR4
2023 LBARNet: Lightweight bilateral asymmetric residual network for real-time semantic segmentation
Xuegang Hu, Baoman Zhou
Comput. Graph.1
2023 Meta-path based graph contrastive learning for micro-video recommendation
Ying He 0008, Gong-Qing Wu, Desheng Cai, Xuegang Hu
Expert Syst. Appl.4
2023 Lightweight attention-guided redundancy-reuse network for real-time semantic segmentation
abstract
Abstract Semantic segmentation is a critical topic in computer vision, and it has numerous practical applications, including mobile devices, autonomous driving, and many other fields. However, in these application scenarios, it is often essential for the segmentation models to achieve a balance between efficiency and performance. A lightweight attention‐guided redundancy‐reuse network (LARNet) was proposed to address this challenge in this paper. Specifically, the multi‐scale asymmetric redundancy reuse (MAR) module was designed as the main component of the encoder for dense encoding of contextual semantic features. Furthermore, the efficient attention fusion (EAF) module was established for multi‐scale information fusion via the channel and spatial attention mechanisms in the decoder. A series of experiments were conducted to verify the proposed network. The results of tests on multiple datasets suggest that the network has higher accuracy and faster speed than the existing real‐time semantic segmentation methods.
Xuegang Hu, Shuhan Xu, Liyuan Jing
IET Image Process.1
2023 Intelligent Internet of Things in Mammography Screening Using Multicenter Transformation Between Unified Capsules
abstract
Mammography screening is one of the important applications for the intelligent Internet of Things (IoT). Due to the efficient and personalized cyber-medicine system, early diagnosis can successfully reduce the breast cancer mortality rate by AI-driven healthcare. However, it is a huge challenge to extend the conventional single-center into the multicenter mammography screening, thus improving the effectiveness and robustness of intelligent IoT-based devices. To address this problem, we utilize multicenter mammograms by the modified capsule neural network and propose a novel framework called multicenter transformation between unified capsules (MLT-UniCaps) in this article. The proposed MLT-UniCaps is composed of Attentional Pose Embedding, Dynamic Source Capsule Traversal, and Adaptive Target Capsule Fusion to realize an intelligent remote assistant diagnosis. Attentional Pose Embedding extracts feature vectors via variations in position, orientation, scale, and lighting as the poses through an adversarial convolutional neural network with an attention-based layer. Based on the pose presentation, Dynamic Source Capsule Traversal deploys a dynamic routing mechanism between neurons to build a source cancer classifier for single-center mammography screening. Using the source cancer classifier, Adaptive Target Capsule Fusion integrates various centers of mammograms as the universal cancer detectors and optimizes heterogeneous distribution among them by the transformation-likelihood maximization. Owing to the three components, MLT-UniCaps effectively improves the results of single-center mammography screening and works in the multicenter breast cancer diagnosis. By comprehensive experiments on 58 965 samples, the proposed MLT-UniCaps obtains 90.1% of overall classification accuracy on single-center trials and 73.8% of overall F1 score on multicenter trials. All the experimental results illustrated that our MLT-UniCaps, an intelligent IoT-based clinical tool, inures the benefit of mammography screening.
Xuegang Hu, Jinglin Zhang 0001, Chenchu Xu, Zhifan Gao
IEEE Internet Things J.2
2023 Lightweight multi-scale attention-guided network for real-time semantic segmentation
Xuegang Hu, Yuanjing Liu
Image Vis. Comput.1
2023 A Drift-Sensitive Distributed LSTM Method for Short Text Stream Classification
abstract
Real-world applications especially in the fields of social media have produced massive short text streams. Unlike traditional normal texts, these data present the characteristics of short length, high-volume, high-velocity and variable data distribution etc, which lead to the issues of data sparsity and concept drift. It is hence very challenging for existing short text classification algorithms. Therefore, we propose a flexible Long Short-Term Memory (LSTM) ensemble network based short text stream classification approach, which is implemented in a distributed mode while maintaining the high-accuracy advantage of deep learning models. More specifically, external resource based short text embedding using a pretrained embedding model and CNN is first proposed for the solution to the data sparsity of short texts. Second, to adapt to the high-volume and high-velocity short text streams, a flexible LSTM network is developed and implemented in a distributed mode for classifying short text data streams. Third, a concept drift factor is introduced for adapting to the concept drifts caused by the changing of data distributions. Finally, experiments conducted on three real short text data sets demonstrate that as compared with several state-of-the-art short text (stream) classification approaches, the proposed approach can classify short text streams effectively and efficiently while adapting to concept drifts.
Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu, Kui Yu
IEEE Trans. Big Data5
2023 Crowdsourcing Truth Inference via Reliability-Driven Multi-View Graph Embedding
abstract
Crowdsourcing truth inference aims to assign a correct answer to each task from candidate answers that are provided by crowdsourced workers. A common approach is to generate workers’ reliabilities to represent the quality of answers. Although crowdsourced triples can be converted into various crowdsourced relationships, the available related methods are not effective in capturing these relationships to alleviate the harm to inference that is caused by conflicting answers. In this research, we propose aReliability-drivenMulti-viewGraphEmbedding framework forTruthinference (TiReMGE), which explores multiple crowdsourced relationships by organically integrating worker reliabilities into a graph space that is constructed from crowdsourced triples. Specifically, to create an interactive environment, we propose a reliability-driven initialization criterion for initializing vectors of tasks and workers as interactive carriers of reliabilities. From the perspective of multiple crowdsourced relationships, a multi-view graph embedding framework is proposed for reliability information interaction on a task-worker graph, which encodes latent crowdsourced relationships into vectors of workers and tasks for reliability update and truth inference. A heritable reliability updating method based on the Lagrange multiplier method is proposed to obtain reliabilities that match the quality of workers for interaction by a novel constraint law. Our ultimate goal is to minimize the Euclidean distance between the encoded task vector and the answer that is provided by a worker with high reliability. Extensive experimental results on nine real-world datasets demonstrate that TiReMGE significantly outperforms the nine state-of-the-art baselines.
Gong-Qing Wu, Xingrui Zhuo, Xianyu Bao, Xuegang Hu, Richang Hong, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data4
2023 High-Dimensional Multi-Label Data Stream Classification With Concept Drifting Detection
abstract
Multi-label data streams such as Web texts and images have been popular on the Web. These data present the characteristics of multiple label, high dimensionality, high volume, high velocity and especial concept drift etc. Thus, multi-label data stream classification is a very challenging and significant task especially in the handling of high-dimensional data with concept drifts. However, this challenge has received little attention from the research community. Therefore, we propose the max-relevance and min-redundancy based algorithm adaptation approach for the efficient and effective classification on multi-label data streams with high-dimensional attributes and concept drifts .11.Source codes and data sets are available at below. https://github.com/peipeilihfut/MLStreamClassificationIn order to reduce the impact from the high-dimensional data with noisy attributes, we first refine the minimal-redundancy-maximal-relevance criterion based on mutual information to select qualified features in multi-label data streams. Secondly, we propose the data distribution based concept drifting detection approach to distinguish concept drifts hidden in data streams. Finally, we build an incremental ensemble classification model for efficiently classifying multi-label data streams. Extensive studies show that our approach can get optimal subsets of features while maintaining a good performance in the multi-label classification, as compared to several state-of-the-art multi-label feature selection algorithms using two efficient multi-label classification methods as base classifiers. Meanwhile, our approach is superior to three well-known multi-label data stream classification approaches in the effectiveness and efficiency.
Pei-Pei Li 0001, Xuegang Hu, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.3
2022 A Baseline for Early Classification of Time Series in An Open World
abstract
Early classification of time series aims to accurately predict the class label of a time series as early as possible, which is significant but challenging in many time-sensitive applications. Existing early classification methods hold a basic closed-world assumption that the classifier must have seen the classes of test samples. However, new samples that do not belong to any trained class may appear in the real world. In this paper, we first address the early classification in an open world and design two detectors to identify which known class or unknown class a sample belongs to. Specifically, based on the observed data, an early known-class detector is designed to determine the known-class confidence and an early unknown-class detector is designed to determine the unknown-class confidence according to the Minimum Reliable Length (MRL) and the Weibull distribution of each class. Experimental results evaluated on real-world datasets demonstrate that the proposed model can identify samples of unknown and known classes accurately and early.
Junwei Lv, Xuegang Hu
COMPSAC2
2022 Dual Confidence Learning Network for Open-World Time Series Classification
Junwei Lv, Ying He 0008, Xuegang Hu, Desheng Cai, Yuqi Chu, Jun Hu 0016
DASFAA (2)3
2022 A lightweight network with multi-scale information interaction attention for real-time semantic segmentation
abstract
Real-time semantic segmentation is an important field in computer vision. It is widely employed in real-world scenarios such as mobile devices and autonomous driving, requiring networks to achieve a trade-off between efficiency, performance, and model size. This paper proposes a lightweight network with multi-scale information interaction attention (MSIANet) to solve this issue. Specifically, we designed a multi-scale information interaction module (MSI) is the main component of the encoder and is used to densely encode contextual semantic features. Moreover, we designed the multi-channel attention fusion module (MAF) in the decoder part, thereby realizing multi-scale information fusion through channel attention mechanism and spatial attention mechanism. We verify our method through numerous experiments and prove that our network possesses fewer parameters and faster inference speed compared to most of the existing real-time semantic segmentation methods in multiple datasets.
Shuhan Xu, Xuegang Hu
ICMV2
2022 Multi-label Learning with Data Self-augmentation
Yuhang Ge, Xuegang Hu, Pei-Pei Li 0001, Haobo Wang 0001, Junbo Zhao 0002
ICONIP (4)2
2022 DAGKT: Difficulty and Attempts Boosted Graph-Based Knowledge Tracing
Fei Liu 0038, Wenhao Liang, Yuhong Zhang 0002, Chenyang Bu, Xuegang Hu
ICONIP (2)6
2022 Structure Similarity Graph for Cross-Network Node Classification
abstract
Cross-network node classification aims to use a labeled source network to classify nodes in an unlabeled target network. Most of the existing cross-network node classification methods learn the network representations by capturing the node neighborhood and train the classifier on these representations. The performance is highly dependent on the high-quality neighborhood in the network. However, in applications, the degree of nodes generally follows a long-tail distribution, i.e., a significant proportion of nodes are tail nodes with sparse neighborhood. It poses a challenge to existing methods. To this end, a structure similarity graph for cross-network node classitication method (SCNC) is proposed in this paper. Firstly, the potential links between nodes are predicted with the structural similarity metric to construct structure similarity graph, which can enrich the neighborhood of tail nodes. Then, the embedding representations of the structural similarity graph are learned to capture more neighborhood information. Finally, the adversarial is used to learn the domain invariant representations to address cross-network divergence. Extensive experimental results show that our SCNC outperforms the state-of-the-art methods.
Xinzheng Li, Yuhong Zhang 0002, Pei-Pei Li 0001, Xuegang Hu
ICTAI4
2022 A Zero-shot Learning Method with a Multi-Modal Knowledge Graph
abstract
Zero-shot learning aims to recognize unseen-classes using some seen-class samples as training set. It is challenging owing to that the feature representations of unseen-class samples are unavailable. Existing methods transfer the mapping from seen-classes to unseen-classes with the correlation as a bridge, in which, the semantic representations are used to discriminate the classes. However, the unavailability of visual representations for unseen-classes and the insufficient discrimination of semantic representations make the zero-shot learning challenging. Therefore, the visual representations are learned as complements to semantic representations to construct a multi-modal knowledge graph (KG), and a zero-shot learning method based on multi-modal KG is proposed in this paper. Specially, a semantic KG is introduced to capture the correlation of classes, and with the correlation, the visual feature representations of all classes are learned. Then, the discriminative visual representations and the semantic representations are used together to construct a multi-modal KG. With the multi-modal KG, the classifier for seen-classes is transferred to unseen classes. Extensive experimental results show the effectiveness of our method.
Yuhong Zhang 0002, Haitao Shu, Chenyang Bu, Xuegang Hu
ICTAI4
2022 Semi-Supervised Online Kernel Extreme Learning Machine for Multi-Label Data Stream Classification
abstract
Semi-supervised multi-label data stream classification serves a practical yet challenging task since only a small number of labeled instances are available in real streaming environments. However, the mainstream of existing semi-supervised multi-label classification technique is focused on the batch process. Meanwhile, many data stream classification approaches have been proposed, and one of popularly used base models is ELM (extreme learning machine). But only few ELM-based algorithms are proposed for the multi-label data stream classification. Therefore, in this paper, we present a novel Semi-supervised Online Extreme Learning Machine with Kernel function for multi-label data stream classification, called SSO-KELM. Specifically, we firstly introduce the kernel function to output the multi-dimensional vector, for adapting to the multi-label data. Secondly, to make full use of labeled and unlabeled data in a data stream, we derive a novel online semi-supervised ELM algorithm, which can adapt to the stream setting and achieve a higher classification accuracy. Finally, extensive experiments conducted on six benchmark multi-label data sets demonstrate the effectiveness of our approach compared to state-of-the-art approaches.
Shiyuan Qiu, Pei-Pei Li 0001, Xuegang Hu
IJCNN3
2022 An Online Dirichlet Model based on Sentence Embedding and DBSCAN for Noisy Short Text Stream Clustering
abstract
Short text stream clustering has received widespread attention due to the rise of various social medias. However, short text streams present the following characteristics such as infinite length, text sparsity and ambiguity, topic evolution and containing noisy data. Existing short text clustering methods do not make full use of the semantic information of short texts to solve the sparsity and ambiguity of short texts and few methods take the noise into account in short text stream. Therefore, in this paper, we propose an Online Dirichlet model based on Sentence Embedding and DBSCAN for noisy short text stream clustering, called ODSE. Firstly, to handle the text sparsity and ambiguity, we use Sentence-Bert to represent each short text for achieving the globally semantic information of each short text. Secondly, to handle the noisy data contained in short texts, we introduce the buffer mechanism and refine the Dirichlet process multinomial mixture model using the DBSCAN algorithm except the above sentence embedding input. This model can handel the short texts one by one. Besides, to adapt to the infinite length and topic evolution, we introduce the forgetting mechanism to update the clusters. Finally, extensive experiments demonstrate that as compared to several state-of-art algorithms, our proposed approach can achieve better performances on four benchmark short text datasets.
Xianliang Si, Pei-Pei Li 0001, Xuegang Hu, Yuhong Zhang 0002
IJCNN3
2022 Multi-Label Learning with Missing Labels via Common and Label-Specific Features
Mengxuan Sun, Pei-Pei Li 0001, Xuegang Hu
IJCNN4
2022 PU Matrix Completion Based Multi-label Classification with Missing Labels
Zhidong Huang, Pei-Pei Li 0001, Xuegang Hu
ISDA (2)3
2022 APGKT: Exploiting Associative Path on Skills Graph for Knowledge Tracing
Haotian Zhang 0007, Chenyang Bu, Fei Liu 0038, Shuochen Liu, Yuhong Zhang 0002, Xuegang Hu
PRICAI (1)6
2022 Multi-modal Component Representation for Multi-source Domain Adaptation Method
Yuhong Zhang 0002, Lin Qian, Xuegang Hu
PRICAI (1)4
2022 Joint pyramid attention network for real-time semantic segmentation of urban scenes
Xuegang Hu, Liyuan Jing, Uroosa Sehar
Appl. Intell.1
2022 Robust and accurate prediction of self-interacting proteins from protein sequence information by exploiting weighted sparse representation based classifier
abstract
BACKGROUND: Self-interacting proteins (SIPs), two or more copies of the protein that can interact with each other expressed by one gene, play a central role in the regulation of most living cells and cellular functions. Although numerous SIPs data can be provided by using high-throughput experimental techniques, there are still several shortcomings such as in time-consuming, costly, inefficient, and inherently high in false-positive rates, for the experimental identification of SIPs even nowadays. Therefore, it is more and more significant how to develop efficient and accurate automatic approaches as a supplement of experimental methods for assisting and accelerating the study of predicting SIPs from protein sequence information. RESULTS: In this paper, we present a novel framework, termed GLCM-WSRC (gray level co-occurrence matrix-weighted sparse representation based classification), for predicting SIPs automatically based on protein evolutionary information from protein primary sequences. More specifically, we firstly convert the protein sequence into Position Specific Scoring Matrix (PSSM) containing protein sequence evolutionary information, exploiting the Position Specific Iterated BLAST (PSI-BLAST) tool. Secondly, using an efficient feature extraction approach, i.e., GLCM, we extract abstract salient and invariant feature vectors from the PSSM, and then perform a pre-processing operation, the adaptive synthetic (ADASYN) technique, to balance the SIPs dataset to generate new feature vectors for classification. Finally, we employ an efficient and reliable WSRC model to identify SIPs according to the known information of self-interacting and non-interacting proteins. CONCLUSIONS: Extensive experimental results show that the proposed approach exhibits high prediction performance with 98.10% accuracy on the yeast dataset, and 91.51% accuracy on the human dataset, which further reveals that the proposed model could be a useful tool for large-scale self-interacting protein prediction and other bioinformatics tasks detection in the future.
Yang Li 0111, Xuegang Hu, Zhu-Hong You, Liping Li 0003, Pei-Pei Li 0001
BMC Bioinform.2
2022 LARFNet: Lightweight asymmetric refining fusion network for real-time semantic segmentation
Xuegang Hu, Juelin Gong
Comput. Graph.1
2022 Combining context-relevant features with multi-stage attention network for short text classification
Pei-Pei Li 0001, Xuegang Hu
Comput. Speech Lang.3
2022 Hierarchical GAN-Tree and Bi-Directional Capsules for multi-label image classification
Xuegang Hu, Pei-Pei Li 0001, Philip S. Yu
Knowl. Based Syst.2
2022 Learning common and label-specific features for multi-Label classification with correlation information
Pei-Pei Li 0001, Xuegang Hu, Kui Yu
Pattern Recognit.3
2022 Representation learning with deep sparse auto-encoder for multi-task learning
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001
Pattern Recognit.4
2022 Fuzzy Bayesian Knowledge Tracing
abstract
Online education promotes the sharing of learning resources. Knowledge tracing (KT) is aimed at tracking the cognition function of students according to their performance on various exercises at different times and has attracted considerable attention. Existing KT models primarily use bisection representations for the performance and cognitive states of students, thus limiting the application scope of these models and the accuracy of the evaluation of student cognitive performance in learning processes. Therefore, fuzzy Bayesian KT models (namely, FBKT and T2FBKT) are proposed to address continuous score scenarios (e.g., subjective examinations) so that the applicability of KT models may be broadened. Moreover, fine-grained cognitive states can be discerned. In particular, referring to type-2 fuzzy theory, T2FBKT mitigates the model uncertainty of FBKT induced by uncertain parameters. Finally, extensive experiments demonstrate the effectiveness of the proposed fuzzy KT models.
Fei Liu 0038, Xuegang Hu, Chenyang Bu, Kui Yu
IEEE Trans. Fuzzy Syst.2
2021 Multi-Label Learning with Missing Features
abstract
Multi-label learning deals with the problem that each example is associated with multiple class labels simultaneously. Existing multi-label learning approaches all assume the feature space is completed and construct classification models using examples with sufficient feature information. However, in real-world applications, it is difficult to get a fully completed feature matrix, that is, only partial feature information of each example can be obtained. In this paper, we formalize this problem as multi-label learning with missing features. To tackle this problem, we learn a feature correlation matrix and apply it to obtain a new supplementary feature matrix, which has richer feature information than the original missing feature matrix. After that, to improve the performance of multi-label classification, we constrain feature correlation on coefficient matrix by assuming that if two features are strongly correlated, the similarity between their corresponding parameter vector will be large. Besides, we also constrain label correlation on output of labels to capture more sufficient relationships between different labels. Extensive experiments show a competitive performance of our method against other state-of the-art multi-label learning approaches.
Pei-Pei Li 0001, Yizhang Zou, Xuegang Hu
IJCNN4
2021 Multi-Label Streaming Feature Selection via Class-Imbalance Aware Rough Set
abstract
Multi-label feature selection aims to select discriminative attributes in multi-label scenario, but most of existing multi-label feature selection methods fail to consider streaming features, i.e. features gradually flow one by one, which is more common in real-world applications. In addition, though there are already some representative works on multi-label streaming feature selection, they fail to tackle the class-imbalance problem, which exists widely in multi-label learning. In fact, class-imbalance will lead to the performance degradation of multi-label learning models. Thus considering class-imbalance problem in multi-label scenario is beneficial to multi-label feature selection because more precise feature evaluation is achieved. Motivated by this, we propose a new rough set named as class-imbalance aware rough set model which can fit class-imbalance problem well. To address streaming features, we construct a novel streaming feature selection framework called SFSCI(Streaming Feature Selection via Class-Imbalance aware rough set), which contains online irrelevancy discarding and online redundancy reduction. Finally, an empirical study on a series of benchmark data sets demonstrates that the proposed method is superior to other state-of-the-art multi-label feature selection methods, including several multi-label streaming feature selection methods.
Yizhang Zou, Xuegang Hu, Pei-Pei Li 0001
IJCNN2
2021 Adversarial training with Wasserstein distance for learning cross-lingual word embeddings
Yuling Li 0001, Yuhong Zhang 0002, Kui Yu, Xuegang Hu
Appl. Intell.4
2021 Attentive interaction-driven entity resolution over multi-source web information
Ying He 0008, Gong-Qing Wu, Desheng Cai, Shengjie Hu, Xianyu Bao, Xuegang Hu
Neurocomputing6
2021 Cognitive structure learning model for hierarchical multi-label text classification
Xuegang Hu, Pei-Pei Li 0001, Philip S. Yu
Knowl. Based Syst.2
2021 Semi-supervised classification on data streams with recurring concept drift and concept evolution
Xiulin Zheng, Pei-Pei Li 0001, Xuegang Hu, Kui Yu
Knowl. Based Syst.3
2020 Representation learning via serial robust autoencoder for domain adaptation
Shuai Yang 0003, Yuhong Zhang 0002, Hao Wang 0008, Pei-Pei Li 0001, Xuegang Hu
Expert Syst. Appl.5
2020 DRU-net: a novel U-net for biomedical image segmentation
abstract
With the wide applications of biomedical images in the medical field, the segmentation of biomedical images plays an important role in clinical diagnosis, pathological analysis, and medical intervention. Full convolutional neural networks, especially U‐net, have improved the performance of segmentation greatly in recent years. However, due to their regular geometric structure, the standard convolutions that they use are inherently limited in dealing with geometric transformations while biomedical objects have huge variations in shape and size. In this study, the authors propose the DRU‐net, which is a novel U‐net with deformable encoder and reshaping upsampling convolution decoder, for biomedical image segmentation. First, deformable convolutional networks are applied and improved to enhance the learning ability of the encoder for geometric transformations. Second, a novel upsampling method named reshape upsampling convolution is proposed for better‐restoring resolution and fusion features. Furthermore, focal loss is used to address class imbalance and model overwhelmed problems in biomedical image segmentation tasks. Theoretic analysis and experimental results have shown that the proposed algorithm not only reduces the number of parameters of U‐Net, but also achieves produces competitive results compared with the state‐of‐the‐art algorithms in terms of various quantitative measures on Drosophila electron microscopy dataset and Warwick‐QU dataset.
Xuegang Hu, Hongguang Yang
IET Image Process.1
2020 Semi-supervised representation learning via dual autoencoders for domain adaptation
Shuai Yang 0003, Hao Wang 0008, Yuhong Zhang 0002, Pei-Pei Li 0001, Yi Zhu 0006, Xuegang Hu
Knowl. Based Syst.6
2020 Wasserstein GAN based on Autoencoder with back-translation for cross-lingual embedding mappings
Yuhong Zhang 0002, Yuling Li 0001, Yi Zhu 0006, Xuegang Hu
Pattern Recognit. Lett.4
2019 A Multi-Grouped LS-SVM Method for Short-Term Urban Traffic Flow Prediction
abstract
Predicting short-term urban traffic flow is a non- trivial task, for an intelligent transportation system could greatly facilitate urban transportation infrastructure construction and enhances the efficiency of traffic control. Unfortunately, urban traffic flow is influenced by numerous factors, which increases the complexity of prediction. In this paper, Multi-Grouped Least Squares Support Vector Machine (MLS-SVM) is proposed for short-term urban traffic flow prediction. In MLS-SVM, spatiotemporal factors (e.g., time, geography, and environment) are divided into different groups. Correlations between each grouped factor are then recognized. Finally, the predicted effect is optimized by combining sub- models for each group. Real-world datasets are used in the experiments of traffic flow prediction. Comparing with the rival methods (i.e., LS-SVM, Wavelet Neural Network, Multi-Factor Pattern Recognition), the simulation results demonstrated the validity and stability of MLS-SVM.
Fei Liu 0038, Zhenchun Wei, Zhensheng Huang, Yang Lu 0015, Xuegang Hu, Lei Shi 0011
GLOBECOM5
2019 Representation learning via serial autoencoders for domain adaptation
Shuai Yang 0003, Yuhong Zhang 0002, Yi Zhu 0006, Pei-Pei Li 0001, Xuegang Hu
Neurocomputing5
2019 Transfer learning with deep manifold regularized auto-encoders
Yi Zhu 0006, Xindong Wu 0001, Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu
Neurocomputing5
2019 Online streaming feature selection using adapted Neighborhood Rough Set
Peng Zhou 0008, Xuegang Hu, Pei-Pei Li 0001, Xindong Wu 0001
Inf. Sci.2
2019 OFS-Density: A novel online streaming feature selection method
Peng Zhou 0008, Xuegang Hu, Pei-Pei Li 0001, Xindong Wu 0001
Pattern Recognit.2
2018 A survey on online feature selection with streaming features
Xuegang Hu, Peng Zhou 0008, Pei-Pei Li 0001, Jing Wang 0021, Xindong Wu 0001
Frontiers Comput. Sci.1
2018 Transfer learning with stacked reconstruction independent component analysis
Yi Zhu 0006, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001
Knowl. Based Syst.2
2018 Co-occurrence pattern mining based on a biological approximation scoring matrix
Dan Guo 0001, Ermao Yuan, Xuegang Hu, Xindong Wu 0001
Pattern Anal. Appl.3
2018 Online Biterm Topic Model based short text stream classification using short text expansion and concept drifting detection
Xuegang Hu, Pei-Pei Li 0001
Pattern Recognit. Lett.1
2018 Learning From Short Text Streams With Topic Drifts
abstract
Short text streams such as search snippets and micro blogs have been popular on the Web with the emergence of social media. Unlike traditional normal text streams, these data present the characteristics of short length, weak signal, high volume, high velocity, topic drift, etc. Short text stream classification is hence a very challenging and significant task. However, this challenge has received little attention from the research community. Therefore, a new feature extension approach is proposed for short text stream classification with the help of a large-scale semantic network obtained from a Web corpus. It is built on an incremental ensemble classification model for efficiency. First, more semantic contexts based on the senses of terms in short texts are introduced to make up of the data sparsity using the open semantic network, in which all terms are disambiguated by their semantics to reduce the noise impact. Second, a concept cluster-based topic drifting detection method is proposed to effectively track hidden topic drifts. Finally, extensive studies demonstrate that as compared to several well-known concept drifting detection methods in data stream, our approach can detect topic drifts effectively, and it enables handling short text streams effectively while maintaining the efficiency as compared to several state-of-the-art short text classification approaches.
Pei-Pei Li 0001, Xuegang Hu, Yuhong Zhang 0002, Lei Li 0002, Xindong Wu 0001
IEEE Trans. Cybern.4
2017 Three-layer concept drifting detection in text data streams
Yuhong Zhang 0002, Guang Chu, Pei-Pei Li 0001, Xuegang Hu, Xindong Wu 0001
Neurocomputing4
2017 Online feature selection for high-dimensional class-imbalanced data
Peng Zhou 0008, Xuegang Hu, Pei-Pei Li 0001, Xindong Wu 0001
Knowl. Based Syst.2
2016 Concept Based Short Text Stream Classification with Topic Drifting Detection
abstract
Short text stream classification is a challengingand significant task due to the characteristics of short length, weak signal, high velocity and especially topic drifting in short text stream. However, this challenge has received little attention from the research community. Motivated by this, we propose a new feature extension approach for short text stream classification using a large scale, general purpose semantic network obtained from a web corpus. Our approach is built on an incremental ensemble classification model. First, in terms of the open semantic network, we introduce more semantic contexts in short texts to make up of the data sparsity. Meanwhile, we disambiguate terms by their semantics to reduce the noise impact. Second, to effectively track hidden topic drifts, we propose a concept cluster based topic drifting detection method. Finally, extensive experiments demonstratethat our approach can detect topic drifts effectively compared to several well-known concept drifting detection methods in data streams. Meanwhile, our approach can perform best in the classification of text data streams compared to several stateof-the-art short text classification approaches.
Pei-Pei Li 0001, Xuegang Hu, Yuhong Zhang 0002, Lei Li 0002, Xindong Wu 0001
ICDM3
2016 A Label Correlation Based Weighting Feature Selection Approach for Multi-label Data
Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu
WAIM (2)5
2016 Domain adaptation via Multi-Layer Transfer Learning
Jianhan Pan, Xuegang Hu, Pei-Pei Li 0001, Huizong Li, Yuhong Zhang 0002, Yaojin Lin
Neurocomputing2
2016 A social tag clustering method based on common co-occurrence group similarity
abstract
Social tagging systems are widely applied in Web 2.0. Many users use these systems to create, organize, manage, and share Internet resources freely. However, many ambiguous and uncontrolled tags produced by social tagging systems not only worsen users’ experience, but also restrict resources’ retrieval efficiency. Tag clustering can aggregate tags with similar semantics together, and help mitigate the above problems. In this paper, we first present a common co-occurrence group similarity based approach, which employs the ternary relation among users, resources, and tags to measure the semantic relevance between tags. Then we propose a spectral clustering method to address the high dimensionality and sparsity of the annotating data. Finally, experimental results show that the proposed method is useful and efficient.
Huizong Li, Xuegang Hu, Yaojin Lin, Jianhan Pan
Frontiers Inf. Technol. Electron. Eng.2
2016 Multi-bridge transfer learning
Xuegang Hu, Jianhan Pan, Pei-Pei Li 0001, Huizong Li, Yuhong Zhang 0002
Knowl. Based Syst.1
2015 Learning concept-drifting data streams with random ensemble decision trees
Pei-Pei Li 0001, Xindong Wu 0001, Xuegang Hu, Hao Wang 0008
Neurocomputing3
2015 Quadruple Transfer Learning: Exploiting both shared and non-shared concepts for text classification
Jianhan Pan, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001, Yaojin Lin, Huizong Li, Lei Li 0002
Knowl. Based Syst.2
2015 Visual data denoising with a unified Schatten-p norm and ℓq norm regularized principal component pursuit
Jing Wang 0021, Meng Wang 0001, Xuegang Hu, Shuicheng Yan
Pattern Recognit.3
2015 Cross-domain sentiment classification-feature divergence, polarity divergence or both?
Yuhong Zhang 0002, Xuegang Hu, Pei-Pei Li 0001, Lei Li 0002, Xindong Wu 0001
Pattern Recognit. Lett.2
2015 A Large Probabilistic Semantic Network Based Approach to Compute Term Similarity
abstract
Measuring semantic similarity between two terms is essential for a variety of text analytics and understanding applications. Currently, there are two main approaches for this task, namely the knowledge based and the corpus based approaches. However, existing approaches are more suitable for semantic similarity between words rather than the more general multi-word expressions (MWEs), and they do not scale very well. Contrary to these existing techniques, we propose an efficient and effective approach for semantic similarity using a large scale semantic network. This semantic network is automatically acquired from billions of web documents. It consists of millions of concepts, which explicitly model the context of semantic relationships. In this paper, we first show how to map two terms into the concept space, and compare their similarity there. Then, we introduce a clustering approach to orthogonalize the concept space in order to improve the accuracy of the similarity measure. Finally, we conduct extensive studies to demonstrate that our approach can accurately compute the semantic similarity between terms of MWEs and with ambiguity, and significantly outperforms 12 competing methods under Pearson Correlation Coefficient. Meanwhile, our approach is much more efficient than all competing algorithms, and can be used to compute semantic similarity in a large scale.
Pei-Pei Li 0001, Haixun Wang, Kenny Q. Zhu, Zhongyuan Wang 0006, Xuegang Hu, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.5
2015 Online Feature Selection with Group Structure Analysis
abstract
Online selection of dynamic features has attracted intensive interest in recent years. However, existing online feature selection methods evaluate features individually and ignore the underlying structure of a feature stream. For instance, in image analysis, features are generated in groups which represent color, texture, and other visual information. Simply breaking the group structure in feature selection may degrade performance. Motivated by this observation, we formulate the problem as an online group feature selection. The problem assumes that features are generated individually but there are group structures in the feature stream. To the best of our knowledge, this is the first time that the correlation among streaming features has been considered in the online feature selection process. To solve this problem, we develop a novel online group feature selection method named OGFS. Our proposed approach consists of two stages: online intra-group selection and online inter-group selection. In the intra-group selection, we design a criterion based on spectral analysis to select discriminative features in each group. In the inter-group selection, we utilize a linear regression model to select an optimal subset. This two-stage procedure continues until there are no more features arriving or some predefined stopping conditions are met. Finally, we apply our method to multiple tasks including image classification and face verification. Extensive empirical studies performed on real-world and benchmark data sets demonstrate that our method outperforms other state-of-the-art online feature selection methods.
Jing Wang 0021, Meng Wang 0001, Pei-Pei Li 0001, Luoqi Liu, Zhong-Qiu Zhao, Xuegang Hu, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.6
2014 Extremal optimization-based semi-supervised algorithm with conflict pairwise constraints for community detection
abstract
The research on community structure is a key to analyze the network functionality and topology, and thus it is significant to detect and analysis the community structure. During the abstract process from an actual system to a network, especially for a large-scale network, it is inevitable to have mistaken connections between nodes or have connection missing. In addition, in real applications, from time to time we can obtain prior information in the form of pairwise constraints between nodes besides topology information, although they may be inaccurate or conflicted. These noises in the network-related information will dramatically reduce the accuracy of community detection. Hence, in this paper, we introduce a dissimilarity index to determine the trustworthiness of pairwise constraints and settle the conflict of pairwise constraints. Then, focusing on the community detection with false connections or conflicted connections, we propose a pairwise constrained structure-enhanced extremal optimization-based semi-supervised algorithm (PCSEO-SS algorithm). Compared with existing semi-supervised community detection approaches, the experimental results executed on real networks and synthetic networks, show that PCSEO-SS can solve the problem of false connections or conflicted connections to some extent and detect the community structure more precisely.
Lei Li 0002, Mei Du, Guanfeng Liu 0001, Xuegang Hu, Gong-Qing Wu
ASONAM4
2014 A Hybrid Feature Selection Approach by Correlation-Based Filters and SVM-RFE
abstract
Selecting a feature subset with strong discriminative power is a critical process for high dimensional data analysis, which has attracted much attention in many application domains, such as text categorization and genome projects. Since traditional feature selection methods provide limited contributions to classification, many researchers resort to hybrid or elaborate approaches to choose interesting features. In this paper, we propose a novel hybrid approach by correlation-based filters and Support Vector Machine-Recursive Feature Elimination (SVM-RFE) method for robust feature selection, which aims to yield robust results by aggregating multiple feature subsets (groups). Specifically, in the first stage, we incorporate correlation-based filters to identify Predominant Features and Complementary Features, and generate multiple groups for robustness, in the second stage, we aggregate multiple groups with SVM-RFE into a compact feature subset for high classification accuracy. Extensive experimental studies on both UCI data sets and microarray data sets have confirmed the effectiveness of our proposed approach.
Xuegang Hu, Pei-Pei Li 0001, Yuhong Zhang 0002, Huizong Li
ICPR2
2014 Searching for Recent Celebrity Images in Microblog Platform
abstract
With the explosive growth and widespread accessibility of image content in social media, many users are eagerly searching for most recent and relevant images on topics of their interests. However, most current microblog platforms merely make use of textual information, specifically, keywords, for image search which cannot achieve satisfactory results since in most cases image content is inconsistent with textual content. In this paper we tackle this problem under the application of searching for celebrity image. The proposed method is based on the idea of refining the initial text-based search results by utilizing multimedia plus social information. Given a text search query, we first obtain an initial text-based result. Next, we extract a seed tweet set whose images contain faces recognized as celebrities and texts contain the expanded keywords. Third, we extend the seed set based on visual and user information. Lastly, we employ a multi-modal graph based learning method to properly rank the obtained tweets by integrating social and visual information. Extensive experiments on data collected from Tencent Weibo demonstrate that our proposed method could approximately achieve 3-fold improvement in results as compared to the text baseline, typically used in microblog search service.
Na Zhao 0004, Richang Hong, Meng Wang 0001, Xuegang Hu, Tat-Seng Chua
ACM Multimedia4
2014 Ensemble learning from multiple information sources via label propagation and consensus
Yaojin Lin, Xuegang Hu, Xindong Wu 0001
Appl. Intell.2
2014 Quality of information-based source assessment and selection
Yaojin Lin, Xuegang Hu, Xindong Wu 0001
Neurocomputing2
2014 Knowledge reduction for decision tables with attribute value taxonomies
Mingquan Ye, Xindong Wu 0001, Xuegang Hu, Donghui Hu
Knowl. Based Syst.3
2014 Robust Face Recognition via Adaptive Sparse Representation
abstract
Sparse representation (or coding)-based classification (SRC) has gained great success in face recognition in recent years. However, SRC emphasizes the sparsity too much and overlooks the correlation information which has been demonstrated to be critical in real-world face recognition problems. Besides, some paper considers the correlation but overlooks the discriminative ability of sparsity. Different from these existing techniques, in this paper, we propose a framework called adaptive sparse representation-based classification (ASRC) in which sparsity and correlation are jointly considered. Specifically, when the samples are of low correlation, ASRC selects the most discriminative samples for representation, like SRC; when the training samples are highly correlated, ASRC selects most of the correlated and discriminative samples for representation, rather than choosing some related samples randomly. In general, the representation model is adaptive to the correlation structure that benefits from both l1-norm and l2-norm. Extensive experiments conducted on publicly available data sets verify the effectiveness and robustness of the proposed algorithm by comparing it with the state-of-the-art methods.
Jing Wang 0021, Canyi Lu, Meng Wang 0001, Pei-Pei Li 0001, Shuicheng Yan, Xuegang Hu
IEEE Trans. Cybern.6
2013 Subchloroplast Location Prediction via Homolog Knowledge Transfer and Feature Selection
abstract
The accuracy of subchloroplast location prediction algorithms often depends on predictive and succinct features derived from proteins. Thus, to improve the prediction accuracy, this paper proposes a novel SubChloroplast location prediction method, called SCHOTS, which integrates the HOmolog knowledge Transfer and feature Selection methods. SCHOTS contains two stages. First, discriminating features are generated by WS-LCHI, a Weighted Gene Ontology (GO) transfer model based on bit-Score of proteins and Logarithmic transformation of CHI-square. Second, the more informative GO terms are selected from the features. Extensive studies conducted on three real datasets demonstrate that SCHOTS outperforms three off-the-shelf subchloroplast prediction methods.
Xindong Wu 0001, Gong-Qing Wu, Xuegang Hu
AAAI4
2013 Web news extraction via path ratios
abstract
In addition to the news content, most web news pages also contain navigation panels, advertisements, related news links etc. These non-news items not only exist outside the news region, but are also present in the news content region. Effectively extracting the news content and filtering the noise have important effects on the follow-up activities of content management and analysis. Our extensive case studies have indicated that there exists potential relevance between web content layouts and their tag paths. Based on this observation, we design two tag path features to measure the importance of nodes: Text to tag Path Ratio (TPR) and Extended Text to tag Path Ratio (ETPR), and describe the calculation process of TPR by traversing the parsing tree of a web news page. In this paper, we present Content Extraction via Path Ratios (CEPR) - a fast, accurate and general on-line method for distinguishing news content from non-news content by the TPR/ETPR histogram effectively. In order to improve the ability of CEPR in extracting short texts, we propose a Gaussian smoothing method weighted by a tag path edit distance. This approach can enhance the importance of internal-link nodes but ignore noise nodes existing in news content. Experimental results on the CleanEval datasets and web news pages randomly selected from well-known websites show that CEPR can extract across multi-resources, multi-styles, and multi-languages. The average F and average score with CEPR is 8.69% and 14.25% higher than CETR, which demonstrates better web news extraction performance than most existing methods.
Gong-Qing Wu, Xuegang Hu, Xindong Wu 0001
CIKM3
2013 Flexible Pattern Matching with Gap-Length and One-Off Conditions
abstract
This paper focuses on pattern matching with wildcard, gap-length and one-off conditions. It is difficult to achieve optimal solutions. We propose an FNP algorithm based on Free-Node Optimum Pruning. Each Free-Node set is a set of nodes labeled by the same number which appear on different layers in a directed graph structure WON-Net. Compared on biological data and artificial data, experimental results show that (1) FNP has a significant advantage with its solutions, winning above 95% of biological data among similar algorithms. There are theorems on obtaining optimal solutions by FNP. (2)FNP demonstrates an evident advantage on running time, when |Net| is large and k is small. (|Net| denotes the number of independent substructures without losing solutions in WON-Net and k is the number of Free-Node sets.)
Dan Guo 0001, Taining Xiang, Xuegang Hu, Xindong Wu 0001
ICTAI3
2013 Online Group Feature Selection
Jing Wang 0021, Zhong-Qiu Zhao, Xuegang Hu, Yiu-Ming Cheung, Meng Wang 0001, Xindong Wu 0001
IJCAI3
2013 Pattern matching with wildcards and gap-length constraints based on a centrality-degree graph
Dan Guo 0001, Xuegang Hu, Fei Xie 0002, Xindong Wu 0001
Appl. Intell.2
2013 Multi-level rough set reduction for decision rule mining
Mingquan Ye, Xindong Wu 0001, Xuegang Hu, Donghui Hu
Appl. Intell.3
2013 Mining stable patterns in multiple correlated databases
Yaojin Lin, Xuegang Hu, Xindong Wu 0001
Decis. Support Syst.2
2013 Anonymizing classification data using rough set theory
Mingquan Ye, Xindong Wu 0001, Xuegang Hu, Donghui Hu
Knowl. Based Syst.3
2012 An Ensemble Method Based on Confidence Probability for Multi-domain Sentiment Classification
Yuhong Zhang 0002, Xuegang Hu
ICIC (1)3
2012 Learning from concept drifting data streams with unlabeled data
Xindong Wu 0001, Pei-Pei Li 0001, Xuegang Hu
Neurocomputing3
2012 Mining Recurring Concept Drifts with Limited Labeled Streaming Data
abstract
Tracking recurring concept drifts is a significant issue for machine learning and data mining that frequently appears in real-world stream classification problems. It is a challenge for many streaming classification algorithms to learn recurring concepts in a data stream environment with unlabeled data, and this challenge has received little attention from the research community. Motivated by this challenge, this article focuses on the problem of recurring contexts in streaming environments with limited labeled data. We propose a semi-supervised classification algorithm for data streams with REcurring concept Drifts and Limited LAbeled data, called REDLLA, in which a decision tree is adopted as the classification model. When growing a tree, a clustering algorithm based on k -means is installed to produce concept clusters and unlabeled data are labeled in the method of majority-class at leaves. In view of deviations between history and new concept clusters, potential concept drifts are distinguished and recurring concepts are maintained. Extensive studies on both synthetic and real-world data confirm the advantages of our REDLLA algorithm over three state-of-the-art online classification algorithms of CVFDT, DWCDS, and CDRDT and several known online semi-supervised algorithms, even in the case with more than 90% unlabeled data.
Pei-Pei Li 0001, Xindong Wu 0001, Xuegang Hu
ACM Trans. Intell. Syst. Technol.3
2011 An Efficient Ensemble Method for Classifying Skewed Data Streams
Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001
ICIC (3)2
2011 Random Ensemble Decision Trees for Learning Concept-Drifting Data Streams
Pei-Pei Li 0001, Xindong Wu 0001, Qianhui Althea Liang, Xuegang Hu, Yuhong Zhang 0002
PAKDD (1)4
2011 A Bit-Parallel Algorithm for Sequential Pattern Matching with Wildcards
abstract
Pattern matching with both gap constraints and the one-off condition is a challenging topic, especially in bioinformatics, information retrieval, and dictionary query. Among the algorithms to solve the problem, the most efficient one is SAIL, which is time consuming, especially when the pattern is long. In addition, existing algorithms based on bit-parallelism cannot handle a pattern that has only one pattern character between successive wildcards and the minimum local length constraints are zero. We propose an algorithm BPBM to handle online sequential pattern matching. In BPBM, an extended bit-parallelism operation is used to accelerate the matching process. An effective transition window mechanism with two nondeterministic finite state automatons (NFAs) is adopted to drop the useless scan window. It identifies gap constraints automatically and just scans once to export occurrences with exact match positions. Theoretical analysis and experimental results show that the BPBM algorithm is more competitive than other peers. It has an absolute advantage on search time complexity. It also has better stability that decreases operation costs with the increasing of the size of sequence alphabet or the length of the pattern. We also study off-line pattern matching. With twice pruning, left-most and right-most, we can increase the matching ratio about 2.08% on average.
Dan Guo 0001, Xiao-Li Hong, Xuegang Hu, Jun Gao 0006, Ying-Ling Liu, Gong-Qing Wu, Xindong Wu 0001
Cybern. Syst.3
2010 Learning from Concept Drifting Data Streams with Unlabeled Data
abstract
Contrary to the previous beliefs that all arrived streaming data are labeled and the class labels are immediately availa- ble, we propose a Semi-supervised classification algorithm for data streams with concept drifts and UNlabeled data, called SUN. SUN is based on an evolved decision tree. In terms of deviation between history concept clusters and new ones generated by a developed clustering algorithm of k-Modes, concept drifts are distinguished from noise at leaves. Extensive studies on both synthetic and real data demonstrate that SUN performs well compared to several known online algorithms on unlabeled data. A conclusion is hence drawn that a feasible reference framework is provided for tackling concept drifting data streams with unlabeled data.
Pei-Pei Li 0001, Xindong Wu 0001, Xuegang Hu
AAAI3
2010 Logistic Regression for Transductive Transfer Learning from Multiple Sources
Yuhong Zhang 0002, Xuegang Hu, Yucheng Fang
ADMA (2)2
2010 Sequential Pattern Mining with Wildcards
abstract
Sequential pattern mining is an important research task in many domains, such as biological science. In this paper, we study the problem of mining frequent patterns from sequences with wildcards. The user can specify the gap constraints with flexibility. Given a subject sequence, a minimal support threshold and a gap constraint, we aim to find frequent patterns whose supports in the sequence are no less than the given support threshold. We design an efficient mining algorithm MAIL that utilizes the candidate occurrences of the prefix to compute the support of a pattern that avoids the rescanning of the sequence. We present two pruning strategies to improve the completeness and the time efficiency of MAIL. Experiments show that MAIL mines 2 times more patterns than one of its peers and the time performance is 12 times faster on average than its another peer.
Fei Xie 0002, Xindong Wu 0001, Xuegang Hu, Jun Gao 0006, Dan Guo 0001, Yulian Fei, Ertian Hua
ICTAI (1)3
2009 Parameter Estimdation in Semi-Random Decision Tree Ensembling on Streaming Data
Pei-Pei Li 0001, Qianhui Althea Liang, Xindong Wu 0001, Xuegang Hu
PAKDD4
2008 Mining Concept-Drifting Data Streams with Multiple Semi-Random Decision Trees
Pei-Pei Li 0001, Xuegang Hu, Xindong Wu 0001
ADMA2
2007 Qualitative Simulation and Reasoning with Feature Reduction Based on Boundary Conditional Entropy of Knowledge
Yusheng Cheng, Yousheng Zhang, Xuegang Hu, Xiaoyao Jiang
PAKDD3
2007 A Semi-Random Multiple Decision-Tree Algorithm for Mining Data Streams
Xuegang Hu, Pei-Pei Li 0001, Xindong Wu 0001, Gong-Qing Wu
J. Comput. Sci. Technol.1