VLDB 2026 Research / reviewers in the wild / expert
Yuhong Zhang 0002
dblp:01/1504-2
· DBLP profile ↗
45ranked-venue papers
16as first author
26since 2021 · last 2026
0000-0001-7031-0889ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 12 first-author · 21 since 2021Databases, data management, data science and information retrieval · 8 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Label-Aware Augmentation for Long-Tailed Cross-Network Node ClassificationabstractABSTRACT Cross‐network node classification (CNC) aims to classify the nodes of an unlabeled graph by leveraging a graph with rich labelled nodes. The node representation and cross‐network discrepancy are two important issues in CNC. Most methods argue that the node representation is crucial to the performance of CNC and learn the representation based on the rich neighbourhood. However, in applications, the networks follow a long‐tailed distribution in their node degrees, that is, most nodes are tail nodes linked to a small amount of neighbours. The sparsity of neighbourhood challenges existing methods in representation and classification. To this end, we propose a label‐aware cross‐network node classification (LA‐CNC) method. First, a label‐consistent augmentation is designed for each network to enrich the representation by augmenting the neighbourhood of tail nodes. Second, a label contrast loss is introduced and combined with adversarial loss to enhance the distinguishability of cross‐network invariant representation. Extensive experiments demonstrate that our method outperforms state‐of‐the‐art methods on several datasets. Yuhong Zhang 0002, Congmei Shi, Xuegang Hu |
Expert Syst. J. Knowl. Eng. | 1 |
| 2026 | Cluster-guided and weight-decorrelated contrastive network for graph clustering
Shengxing Bai, Yuhong Zhang 0002, Peng Zhou 0008, Xindong Wu 0001 |
Knowl. Based Syst. | 2 |
| 2026 | MGCD: Multiple-Granularity Cognitive Diagnosis in Intelligent Education SystemsabstractCognitive diagnosis (CD) is an important task in the field of intelligent education, aiming to discover the proficiency of students on knowledge concepts with response logs. In applications, different users of the tutoring system demand for a diagnosis of knowledge concepts at different granularities. However, recent methods assume that the concepts are of the same granularity and use explicit correlations between same-granularity concepts to improve the diagnosis performance. If required for diagnosing multi-granularity concepts, these methods will face diminished performance or partial invalidation. To this end, we make the first attempt for multiple-granularity cognitive diagnosis, i.e., diagnosis on coarse- and fine-grained concepts simultaneously. Specifically, in a skillful way, the same-granularity correlations are captured and embedded into concept representations in view of concept semantics and cross-granularity correlations to model the proficiency influence between concepts implicitly. Then, the specific loss for single-granularity diagnosis and the general loss for the consistency of multi-granularity are designed to train the model jointly, achieving multiple-granularity diagnosis. Extensive experiments demonstrate that our method can achieve state-of-the-art accuracy on both coarse- and fine-grained concepts. Yuhong Zhang 0002, Tiancheng He, Chenyang Bu, Kui Yu, Xuegang Hu, Xindong Wu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2025 | An Robust Entity Alignment Method based on Knowledge Distillation with Noisy Aligned PairsabstractEntity alignment (EA) aims to find the same entities in different knowledge graphs. Existing EA methods assume the supervised aligned pairs without noise. In applications, noisy pairs lead to degradation of EA performance. To this end, a robust EA method based on knowledge distillation is proposed for noisy pairs. Firstly, the dual-teacher model with online distillation is designed, in which, noise discriminator is performed to improve the noise resistance of teacher models. Secondly, a student model is offline distilled from the dual-teacher model without using the noisy supervised pairs, further enhancing the robustness of student model. In addition, the entity structure is combined with entity representation for alignment inference to alleviate the bias of entity representation in noisy environment. Extensive experiments demonstrate the effectiveness of the proposed method. Yuhong Zhang 0002, Hangchi Song, Chenyang Bu, Kui Yu |
CIKM | 1 |
| 2025 | A Plug-in for cognitive diagnosis method based on correlation representation under long-tailed distribution
Yuhong Zhang 0002, Tiancheng He, Shengyu Xu, Chenyang Bu, Xuegang Hu |
Expert Syst. Appl. | 1 |
| 2025 | A Cognitive Diagnosis Model With Nonlinear Dependence Between Students and ExercisesabstractCognitive diagnosis (CD) aims to discover students’ mastery of knowledge concepts through response logs and it is an important task in intelligence education. CD is generally performed with the advantage of students doing exercises, which is learned by linear interacting between student proficiency and exercise difficulty. However, existing methods represent the proficiency and difficulty in view of knowledge concepts independently, which is in coarse-granularity, leading to the indiscrimination for different combinations of students and exercises under the same concept. To this end, we propose a cognitive diagnosis method with nonlinear dependence between students and exercises (CDND), in which, more fine-grained information is captured to distinguish the different combinations of students and exercises. First, the nonlinear dependence is captured with a novel attention-like mechanism by interacting between students and exercises. Moreover, this nonlinear dependence is embedded as the interactive discrimination vector. Second, the interactive discrimination is used to adjust the advantage using a second-order interaction, which can strengthen the discrimination of the advantage under different combinations of students and exercises. Extensive experimental results on three real data sets validate our interactive discrimination has good compatibility with existing methods, and with it, our CDND achieves an obvious improvement in performance. Our code is available onhttps://github.com/joyce99/LinZhihao/tree/main/CDND-master. Yuhong Zhang 0002, Chenyang Bu, Kui Yu, Xuegang Hu, Xindong Wu 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | A cross-network node classification method in open-set scenario
Yuhong Zhang 0002, Yunlong Ji, Kui Yu, Xuegang Hu, Xindong Wu 0001 |
Pattern Recognit. | 1 |
| 2024 | Diverse Structure-Aware Relation Representation in Cross-Lingual Entity AlignmentabstractCross-lingual entity alignment (CLEA) aims to find equivalent entity pairs between knowledge graphs (KGs) in different languages. It is an important way to connect heterogeneous KGs and facilitate knowledge completion. Existing methods have found that incorporating relations into entities can effectively improve KG representation and benefit entity alignment, and these methods learn relation representation depending on entities, which cannot capture the diverse structures of relations. However, multiple relations in KG form diverse structures, such as adjacency structure and ring structure. This diversity of relation structures makes the relation representation challenging. Therefore, we propose to construct the weighted line graphs to model the diverse structures of relations and learn relation representation independently from entities. Especially, owing to the diversity of adjacency structures and ring structures, we propose to construct adjacency line graph and ring line graph, respectively, to model the structures of relations and to further improve entity representation. In addition, to alleviate the hubness problem in alignment, we introduce the optimal transport into alignment and compute the distance matrix in a different way. From a global perspective, we calculate the optimal 1-to-1 alignment bi-directionally to improve the alignment accuracy. Experimental results on two benchmark datasets show that our proposed method significantly outperforms state-of-the-art CLEA methods in both supervised and unsupervised manners. Yuhong Zhang 0002, Kui Yu, Xindong Wu 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | Adaptive Prototype Interaction Network for Few-Shot Knowledge Graph CompletionabstractFew-shot knowledge graph completion (FKGC), which aims to infer new triples for a relation using only a few reference triples of the relation, has attracted much attention in recent years. Most existing FKGC methods learn a transferable embedding space, where entity pairs belonging to the same relations are close to each other. In real-world knowledge graphs (KGs), however, some relations may involve multiple semantics, and their entity pairs are not always close due to having different meanings. Hence, the existing FKGC methods may yield suboptimal performance when handling multiple semantic relations in the few-shot scenario. To solve this problem, we propose a new method named adaptive prototype interaction network (APINet) for FKGC. Our model consists of two major components: 1) an interaction attention encoder (InterAE) to capture the underlying relational semantics of entity pairs by modeling the interactive information between head and tail entities and 2) an adaptive prototype net (APNet) to generate relation prototypes adaptive to different query triples by extracting query-relevant reference pairs and reducing the data inconsistency between support and query sets. Experimental results on two public datasets demonstrate that APINet outperforms several state-of-the-art FKGC methods. The ablation study demonstrates the rationality and effectiveness of each component of APINet. Yuling Li 0001, Kui Yu, Yuhong Zhang 0002, Jiye Liang, Xindong Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Probabilistic model with evolutionary optimization for cognitive diagnosisabstractCognitive Diagnostic Models (CDMs) aim to analyze students' cognitive levels of each knowledge component (KC) by mining educational data. Existing CDMs can be mainly divided into two categories, i.e., traditional probability-based and neural-network-based. Most probabilistic models have the advantages of simplicity and good interpretability, but suffer from slow training time in the case of a large number of KCs. Neural-network-based methods are widely considered to be superior to probabilistic models due to their good performance. However, neural network methods are less interpretable than probabilistic models, thus limiting their usefulness in practice. Because most existing probabilistic models are optimized iteratively based on single-point-based search methods, they may be easily trapped in local optimum due to the influence of the initial points. And evolutionary algorithms (EAs) have good global search ability. Therefore, an interesting question is whether a simple probabilistic model based on evolutionary optimization can rival neural-network in limited optimization time. Thus, a hybrid EA with a customized local search is proposed. Experimental results on three real-world datasets show that our method outperforms the compared 7 models (including 2 state-of-the-art neural-network-based models); and the running time of our method is significantly less than the compared probabilistic models. Chenyang Bu, Zhiyong Cao, Chenlong He, Yuhong Zhang 0002 |
GECCO | 4 |
| 2023 | Global and Adaptive Local Label Correlation for Multi-label Learning with Missing LabelsabstractLabel missing is a major challenge in multi-label learning. Many existing methods try to use label correlation to recover ground-truth labels, but they only focus on the label correlation within the original label space, however, the label correlation learned in this way is incomplete. Thus, inspired$b$y the matrix adaptive column correlation, we propose a method to continuously adjust the label correlation matrix while the labels are filled in by adaptive column correlation learning method. Specifically, to reduce the impact of the missing labelson label correlation, the label space is firstly completed through manifold regularization while learning the local label information by adaptive column correlation learning in the complemented label space. Secondly, the global label correlation is utilized by adding a low-rank constraint to the entire label space. Finally, by jointly taking advantage of the global and adaptive local label correlation, our proposed approach achieves superior performance on both synthetic and real-world data sets from diverse domains compared to state-of-the art baselines. Qingxia Jiang, Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu |
IJCNN | 3 |
| 2023 | TransD-based Multi-hop Meta Learning for Few-shot Knowledge Graph CompletionabstractFew-shot knowledge graph completion (FKGC), which aims to infer missing facts about a relation from only a few reference triples, has recently attracted great attention. The core of solving the FKGC task is to learn a vector representation for each few-shot relation using the corresponding entity represen-tations. To this end, existing models generally enhance entity representations with their direct neighbors. However, a large number of entities have few direct neighbors. Hence, encoding only direct neighborhood is insufficient to obtain satisfactory en-tity representations. In addition, current models typically utilize static embeddings to represent entities, ignoring their diverse semantics, i.e., an entity may show distinct semantics within different few-shot relations. To address these issues, we propose a new FKGC framework, namely TransD-based Multi-hop Meta Learning (TDML). TDML consists of three main components: a multi-hop neighbor encoder to enhance entity representations by aggregating heterogeneous multi-hop neighbors, a transformer encoder to generate the relation meta representations, and a TransD-based relation representation updater that allows each entity to exhibit relation-specific semantics and tune the relation meta representations. Extensive experiments on two public datasets demonstrate that our model outperforms state-of-the-art FKGC methods. Jindi Li, Kui Yu, Yuling Li 0001, Yuhong Zhang 0002 |
IJCNN | 4 |
| 2023 | A Drift-Sensitive Distributed LSTM Method for Short Text Stream ClassificationabstractReal-world applications especially in the fields of social media have produced massive short text streams. Unlike traditional normal texts, these data present the characteristics of short length, high-volume, high-velocity and variable data distribution etc, which lead to the issues of data sparsity and concept drift. It is hence very challenging for existing short text classification algorithms. Therefore, we propose a flexible Long Short-Term Memory (LSTM) ensemble network based short text stream classification approach, which is implemented in a distributed mode while maintaining the high-accuracy advantage of deep learning models. More specifically, external resource based short text embedding using a pretrained embedding model and CNN is first proposed for the solution to the data sparsity of short texts. Second, to adapt to the high-volume and high-velocity short text streams, a flexible LSTM network is developed and implemented in a distributed mode for classifying short text data streams. Third, a concept drift factor is introduced for adapting to the concept drifts caused by the changing of data distributions. Finally, experiments conducted on three real short text data sets demonstrate that as compared with several state-of-the-art short text (stream) classification approaches, the proposed approach can classify short text streams effectively and efficiently while adapting to concept drifts. Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu, Kui Yu |
IEEE Trans. Big Data | 4 |
| 2023 | Independent Relation Representation With Line Graph for Cross-Lingual Entity AlignmentabstractCross-lingual entity alignment, which is an important task in the field of graph mining, aims to find equivalent entity pairs from two knowledge graphs. Recent methods show that relation representation can be used to improve entity representation and entity alignment. However, relation representation is learned dependently on entity representation, and both representations are learned from node-centered knowledge graphs. This dependency of relation on entity results in poor relation representation and leads to limited enhancement of entity representation. Therefore, to address this challenge, we propose novel relation-aware line graph neural networks for cross-lingual entity alignment (RALG). More specifically, we first propose to learn relation representation with heterogeneous line graphs independently from entities. The constructed heterogeneous line graphs can capture the correlation of relations explicitly. Secondly, we design a new way of aggregation in the form of triples to strengthen the relevance between entities and their corresponding relations. Experiments conducted on real-world datasets show that independent learning of relation representation with line graphs can represent relations better, and our method achieves better performance than the state-of-the-art methods for entity alignment. Yuhong Zhang 0002, Kui Yu, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Learning Inter-Entity-Interaction for Few-Shot Knowledge Graph CompletionabstractFew-shot knowledge graph completion (FKGC) aims to infer unknown fact triples of a relation using its few-shot reference entity pairs.Recent FKGC studies focus on learning semantic representations of entity pairs by separately encoding the neighborhoods of head and tail entities.Such practice, however, ignores the inter-entity interaction, resulting in low-discrimination representations for entity pairs, especially when these entity pairs are associated with 1-to-N, N-to-1, and N-to-N relations.To address this issue, this paper proposes a novel FKGC model, named Cross-Interaction Attention Network (CIAN) to investigate the inter-entity interaction between head and tail entities.Specifically, we first explore the interactions within entities by computing the attention between the task relation and each entity neighbor, and then model the interactions between head and tail entities by letting an entity to attend to the neighborhood of its paired entity.In this way, CIAN can figure out the relevant semantics between head and tail entities, thereby generating more discriminative representations for entity pairs.Extensive experiments on two public datasets show that CIAN outperforms several state-of-the-art methods.The source code is available at https://github.com/cjlyl/FKGC-CIAN. Yuling Li 0001, Kui Yu, Yuhong Zhang 0002 |
EMNLP | 4 |
| 2022 | DAGKT: Difficulty and Attempts Boosted Graph-Based Knowledge Tracing
Fei Liu 0038, Wenhao Liang, Yuhong Zhang 0002, Chenyang Bu, Xuegang Hu |
ICONIP (2) | 4 |
| 2022 | Structure Similarity Graph for Cross-Network Node ClassificationabstractCross-network node classification aims to use a labeled source network to classify nodes in an unlabeled target network. Most of the existing cross-network node classification methods learn the network representations by capturing the node neighborhood and train the classifier on these representations. The performance is highly dependent on the high-quality neighborhood in the network. However, in applications, the degree of nodes generally follows a long-tail distribution, i.e., a significant proportion of nodes are tail nodes with sparse neighborhood. It poses a challenge to existing methods. To this end, a structure similarity graph for cross-network node classitication method (SCNC) is proposed in this paper. Firstly, the potential links between nodes are predicted with the structural similarity metric to construct structure similarity graph, which can enrich the neighborhood of tail nodes. Then, the embedding representations of the structural similarity graph are learned to capture more neighborhood information. Finally, the adversarial is used to learn the domain invariant representations to address cross-network divergence. Extensive experimental results show that our SCNC outperforms the state-of-the-art methods. Xinzheng Li, Yuhong Zhang 0002, Pei-Pei Li 0001, Xuegang Hu |
ICTAI | 2 |
| 2022 | A Zero-shot Learning Method with a Multi-Modal Knowledge GraphabstractZero-shot learning aims to recognize unseen-classes using some seen-class samples as training set. It is challenging owing to that the feature representations of unseen-class samples are unavailable. Existing methods transfer the mapping from seen-classes to unseen-classes with the correlation as a bridge, in which, the semantic representations are used to discriminate the classes. However, the unavailability of visual representations for unseen-classes and the insufficient discrimination of semantic representations make the zero-shot learning challenging. Therefore, the visual representations are learned as complements to semantic representations to construct a multi-modal knowledge graph (KG), and a zero-shot learning method based on multi-modal KG is proposed in this paper. Specially, a semantic KG is introduced to capture the correlation of classes, and with the correlation, the visual feature representations of all classes are learned. Then, the discriminative visual representations and the semantic representations are used together to construct a multi-modal KG. With the multi-modal KG, the classifier for seen-classes is transferred to unseen classes. Extensive experimental results show the effectiveness of our method. Yuhong Zhang 0002, Haitao Shu, Chenyang Bu, Xuegang Hu |
ICTAI | 1 |
| 2022 | An Online Dirichlet Model based on Sentence Embedding and DBSCAN for Noisy Short Text Stream ClusteringabstractShort text stream clustering has received widespread attention due to the rise of various social medias. However, short text streams present the following characteristics such as infinite length, text sparsity and ambiguity, topic evolution and containing noisy data. Existing short text clustering methods do not make full use of the semantic information of short texts to solve the sparsity and ambiguity of short texts and few methods take the noise into account in short text stream. Therefore, in this paper, we propose an Online Dirichlet model based on Sentence Embedding and DBSCAN for noisy short text stream clustering, called ODSE. Firstly, to handle the text sparsity and ambiguity, we use Sentence-Bert to represent each short text for achieving the globally semantic information of each short text. Secondly, to handle the noisy data contained in short texts, we introduce the buffer mechanism and refine the Dirichlet process multinomial mixture model using the DBSCAN algorithm except the above sentence embedding input. This model can handel the short texts one by one. Besides, to adapt to the infinite length and topic evolution, we introduce the forgetting mechanism to update the clusters. Finally, extensive experiments demonstrate that as compared to several state-of-art algorithms, our proposed approach can achieve better performances on four benchmark short text datasets. Xianliang Si, Pei-Pei Li 0001, Xuegang Hu, Yuhong Zhang 0002 |
IJCNN | 4 |
| 2022 | APGKT: Exploiting Associative Path on Skills Graph for Knowledge Tracing
Haotian Zhang 0007, Chenyang Bu, Fei Liu 0038, Shuochen Liu, Yuhong Zhang 0002, Xuegang Hu |
PRICAI (1) | 5 |
| 2022 | Multi-modal Component Representation for Multi-source Domain Adaptation Method
Yuhong Zhang 0002, Lin Qian, Xuegang Hu |
PRICAI (1) | 1 |
| 2022 | A Structure-Aware Method for Cross-domain Text Classification
Yuhong Zhang 0002, Lin Qian, Pei-Pei Li 0001, Guocheng Liu |
PRICAI (2) | 1 |
| 2022 | Representation learning with deep sparse auto-encoder for multi-task learning
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001 |
Pattern Recognit. | 5 |
| 2021 | An Unsupervised Domain Adaptation Being Aware of Domain-specific and Label InformationabstractDomain adaptation aims to leverage the knowledge in a label-rich source domain to facilitate the learning task in an unlabeled target domain with a different distribution. Adversarial-based domain adaptation methods have attracted increasing attention due to their remarkable performance. However, most methods learn the invariant features by aligning distribution between domains, while ignoring the domain-specific and label information. It will lead to unsatisfying invariant features and cause improper label alignment when there is much specific information. Therefore, in this paper, we propose an unsupervised method to learn more robust and discriminative invariant features for domain adaptation by using the specific features and label information. Specially, a separate batch normalization layer is introduced to replace the completely shared layer to capture domain-specific information, which will benefit the learning of invariant features. Then, to make the adaptation more sufficiently, a symmetric design of classifier and the corresponding adversarial training loss are used to realize domain-wise and label-wise alignment. The two steps are optimized iteratively to improve the performance of the model. The extensive experiments on three benchmark datasets have demonstrated the effectiveness of our method. Yuhong Zhang 0002 |
IJCNN | 1 |
| 2021 | Adversarial training with Wasserstein distance for learning cross-lingual word embeddings
Yuling Li 0001, Yuhong Zhang 0002, Kui Yu, Xuegang Hu |
Appl. Intell. | 2 |
| 2021 | Learning Cross-Lingual Mappings in Imperfectly Isomorphic Embedding SpacesabstractOne mainstream method in cross-lingual word embeddings is to learn a linear mapping between two monolingual embedding spaces using a training dictionary. Successful linear mappings require isomorphic embedding spaces. However, monolingual embedding spaces are not perfectly isomorphic, and therefore, a linear mapping cannot align them accurately. In this study, we assume that two embedding spaces are composed of near-isomorphic translation pairs (NearITP) and non-isomorphic translation pairs. Owing to the nature of similar substructures, NearITP can make linear mapping work well. Motivated by this, we design a screening strategy to identify NearITP effectively. Based on this strategy, we find that the proportion of NearITP in the commonly used training dictionary is relatively low, leading to sub-optimal results. To address this problem, we propose a general framework that can be combined with any of the mapping methods, which further boosts subsequent mapping. Experimental results demonstrate that our framework is an improvement over existing mapping-based methods, and outperforms state-of-the-art models on two public data sets. Moreover, we show that our framework can be successfully generalized to contextual word embeddings such as multilingual BERT (mBERT), and further enhances the cross-lingual properties of mBERT. Yuling Li 0001, Kui Yu, Yuhong Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Representation learning via serial robust autoencoder for domain adaptation
Shuai Yang 0003, Yuhong Zhang 0002, Hao Wang 0008, Pei-Pei Li 0001, Xuegang Hu |
Expert Syst. Appl. | 2 |
| 2020 | Semi-supervised representation learning via dual autoencoders for domain adaptation
Shuai Yang 0003, Hao Wang 0008, Yuhong Zhang 0002, Pei-Pei Li 0001, Yi Zhu 0006, Xuegang Hu |
Knowl. Based Syst. | 3 |
| 2020 | Wasserstein GAN based on Autoencoder with back-translation for cross-lingual embedding mappings
Yuhong Zhang 0002, Yuling Li 0001, Yi Zhu 0006, Xuegang Hu |
Pattern Recognit. Lett. | 1 |
| 2019 | Representation learning via serial autoencoders for domain adaptation
Shuai Yang 0003, Yuhong Zhang 0002, Yi Zhu 0006, Pei-Pei Li 0001, Xuegang Hu |
Neurocomputing | 2 |
| 2019 | Transfer learning with deep manifold regularized auto-encoders
Yi Zhu 0006, Xindong Wu 0001, Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu |
Neurocomputing | 4 |
| 2018 | Transfer learning with stacked reconstruction independent component analysis
Yi Zhu 0006, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001 |
Knowl. Based Syst. | 3 |
| 2018 | Learning From Short Text Streams With Topic DriftsabstractShort text streams such as search snippets and micro blogs have been popular on the Web with the emergence of social media. Unlike traditional normal text streams, these data present the characteristics of short length, weak signal, high volume, high velocity, topic drift, etc. Short text stream classification is hence a very challenging and significant task. However, this challenge has received little attention from the research community. Therefore, a new feature extension approach is proposed for short text stream classification with the help of a large-scale semantic network obtained from a Web corpus. It is built on an incremental ensemble classification model for efficiency. First, more semantic contexts based on the senses of terms in short texts are introduced to make up of the data sparsity using the open semantic network, in which all terms are disambiguated by their semantics to reduce the noise impact. Second, a concept cluster-based topic drifting detection method is proposed to effectively track hidden topic drifts. Finally, extensive studies demonstrate that as compared to several well-known concept drifting detection methods in data stream, our approach can detect topic drifts effectively, and it enables handling short text streams effectively while maintaining the efficiency as compared to several state-of-the-art short text classification approaches. Pei-Pei Li 0001, Xuegang Hu, Yuhong Zhang 0002, Lei Li 0002, Xindong Wu 0001 |
IEEE Trans. Cybern. | 5 |
| 2017 | Three-layer concept drifting detection in text data streams
Yuhong Zhang 0002, Guang Chu, Pei-Pei Li 0001, Xuegang Hu, Xindong Wu 0001 |
Neurocomputing | 1 |
| 2016 | Concept Based Short Text Stream Classification with Topic Drifting DetectionabstractShort text stream classification is a challengingand significant task due to the characteristics of short length, weak signal, high velocity and especially topic drifting in short text stream. However, this challenge has received little attention from the research community. Motivated by this, we propose a new feature extension approach for short text stream classification using a large scale, general purpose semantic network obtained from a web corpus. Our approach is built on an incremental ensemble classification model. First, in terms of the open semantic network, we introduce more semantic contexts in short texts to make up of the data sparsity. Meanwhile, we disambiguate terms by their semantics to reduce the noise impact. Second, to effectively track hidden topic drifts, we propose a concept cluster based topic drifting detection method. Finally, extensive experiments demonstratethat our approach can detect topic drifts effectively compared to several well-known concept drifting detection methods in data streams. Meanwhile, our approach can perform best in the classification of text data streams compared to several stateof-the-art short text classification approaches. Pei-Pei Li 0001, Xuegang Hu, Yuhong Zhang 0002, Lei Li 0002, Xindong Wu 0001 |
ICDM | 4 |
| 2016 | A Label Correlation Based Weighting Feature Selection Approach for Multi-label Data
Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu |
WAIM (2) | 4 |
| 2016 | Domain adaptation via Multi-Layer Transfer Learning
Jianhan Pan, Xuegang Hu, Pei-Pei Li 0001, Huizong Li, Yuhong Zhang 0002, Yaojin Lin |
Neurocomputing | 6 |
| 2016 | Multi-bridge transfer learning
Xuegang Hu, Jianhan Pan, Pei-Pei Li 0001, Huizong Li, Yuhong Zhang 0002 |
Knowl. Based Syst. | 6 |
| 2015 | Quadruple Transfer Learning: Exploiting both shared and non-shared concepts for text classification
Jianhan Pan, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001, Yaojin Lin, Huizong Li, Lei Li 0002 |
Knowl. Based Syst. | 3 |
| 2015 | Cross-domain sentiment classification-feature divergence, polarity divergence or both?
Yuhong Zhang 0002, Xuegang Hu, Pei-Pei Li 0001, Lei Li 0002, Xindong Wu 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | A Hybrid Feature Selection Approach by Correlation-Based Filters and SVM-RFEabstractSelecting a feature subset with strong discriminative power is a critical process for high dimensional data analysis, which has attracted much attention in many application domains, such as text categorization and genome projects. Since traditional feature selection methods provide limited contributions to classification, many researchers resort to hybrid or elaborate approaches to choose interesting features. In this paper, we propose a novel hybrid approach by correlation-based filters and Support Vector Machine-Recursive Feature Elimination (SVM-RFE) method for robust feature selection, which aims to yield robust results by aggregating multiple feature subsets (groups). Specifically, in the first stage, we incorporate correlation-based filters to identify Predominant Features and Complementary Features, and generate multiple groups for robustness, in the second stage, we aggregate multiple groups with SVM-RFE into a compact feature subset for high classification accuracy. Extensive experimental studies on both UCI data sets and microarray data sets have confirmed the effectiveness of our proposed approach. Xuegang Hu, Pei-Pei Li 0001, Yuhong Zhang 0002, Huizong Li |
ICPR | 5 |
| 2012 | An Ensemble Method Based on Confidence Probability for Multi-domain Sentiment Classification
Yuhong Zhang 0002, Xuegang Hu |
ICIC (1) | 2 |
| 2011 | An Efficient Ensemble Method for Classifying Skewed Data Streams
Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001 |
ICIC (3) | 3 |
| 2011 | Random Ensemble Decision Trees for Learning Concept-Drifting Data Streams
Pei-Pei Li 0001, Xindong Wu 0001, Qianhui Althea Liang, Xuegang Hu, Yuhong Zhang 0002 |
PAKDD (1) | 5 |
| 2010 | Logistic Regression for Transductive Transfer Learning from Multiple Sources
Yuhong Zhang 0002, Xuegang Hu, Yucheng Fang |
ADMA (2) | 1 |