VLDB 2026 Research / reviewers in the wild / expert
Hanrui Wu
dblp:200/9625
· DBLP profile ↗
35ranked-venue papers
20as first author
28since 2021 · last 2026
0000-0003-3565-6635ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 15 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Local and High-Order Consistency Coding and Adaptation for Cross-Hypergraph Node ClassificationabstractNode classification is a fundamental task in hypergraph learning. Existing methods generally assume that there are a few labeled nodes given in advance. However, in a newly formed hypergraph, collecting label information is challenging and costly in practice. Besides, current approaches mainly exploit the local consistency relationship, i.e., direct neighborhood information, while ignoring the high-order consistency relationship, i.e., high-order proximity information, limiting the discrimination of the latent representations. To address these issues, we propose leveraging knowledge from an auxiliary well-labeled hypergraph (source hypergraph) to assist the learning tasks in the target hypergraph, thus studying the cross-hypergraph node classification problem. Specifically, we propose a model, namely Local and High-order Consistency Coding and Adaptation (LHCCA), which learns both discriminative and transferable node representations. On the one hand, for each hypergraph, by exploiting the local and high-order consistency relationships, LHCCA obtains two kinds of representations, which are then coded by an attention mechanism to achieve a unified representation. On the other hand, the coded source and target node representations are enforced adversarial domain adaptation and contrastive learning to discover transferable features for adaptation. Furthermore, we derive theoretical analyses to establish desirable properties of the proposed model. Extensive experiments on several real-world datasets are conducted, and the promising results demonstrate the effectiveness of the proposed model. Hanrui Wu, Yanxin Wu, Zhao-Rong Lai, Jinyi Long, Michael Kwok-Po Ng, C. L. Philip Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Consistent and specific multi-view multi-label learning with correlation information
Jia Zhang 0019, Hanrui Wu, Guodong Du 0002, Jinyi Long |
Inf. Sci. | 3 |
| 2025 | Towards spatio-temporal representation learning for EEG classification in motor imagery-based BCI system
Siwei Liu 0014, Jia Zhang 0019, Hanrui Wu, Guoxu Zhou, Qibin Zhao, Jinyi Long |
Knowl. Based Syst. | 3 |
| 2025 | EEG Feature Selection in Emotion Recognition Using a Fuzzy Information-Theoretic Based Optimization ApproachabstractFor electroencephalogram (EEG)-based emotion recognition, various EEG features are extracted from frequency, time, and time-frequency domains for modeling. Nevertheless, there is no a standard subset of EEG features widely accepted in this research field, giving rise to the challenge of curse dimensionality. To cope with the challenge, many EEG feature selection (FS) methods have been put forward based on information theory. Generally, these methods suffer from the issue of delivering a suboptimal result with heuristic search, and they are also inefficient in balancing the influence of different terms like feature relevance and feature redundancy. Based on this, we present a new fuzzy information-theoretic based optimization approach to attain the goal. To be specific, fuzzy mutual information is unitized to evaluate EEG features from the relevance and redundancy perspective, and data structure information is captured to exploit feature manifold simultaneously. Then, a unified optimization framework is designed to take all of them into consideration, thereby inducing a globally optimal result of EEG FS. Extensive empirical studies on three EEG emotional datasets reveal that our method is able to found out a discriminative and non-redundant feature subset from different domains, and therefore achieves the performance improvement of emotion recognition. Jia Zhang 0019, Siwei Liu 0014, Hanrui Wu, Jinyi Long |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | SMLE: Semi-Supervised Multi-Label Learning with Label EnhancementabstractSemi-supervised multi-label learning (SSMLL) involves learning a multi-label classifier from a small set of labeled data and a large set of unlabeled data. Label enhancement (LE), accounting for the relative importance of labels, has been effective in improving the performance of supervised multi-label learning models. Nevertheless, generating a robust SSMLL model with LE based on incomplete label information remains challenging. In this paper, we pioneer the idea of applying LE to SSMLL. First, we design a kNN aggregation-based method, aiming to assign pseudo-labels to unlabeled data and perform the LE process by aggregating label information from neighboring instances. Leveraging the topological structure of the feature space is an effective LE approach for training. However, LE, decoupled from the training process, lacks the dynamic feedback of the training model. To improve this, we incorporate a label propagation mechanism that iteratively optimizes the LE process with the guidance of the available label information. Moreover, we consider local label correlations according to local linear embedding to further enhance the generalization ability of the learning model. Extensive experiments demonstrate that the proposed approach can effectively recover latent label information, resulting in significant performance improvement in SSMLL. Qianzhi Ye, Jia Zhang 0019, Hanrui Wu, Tianlong Gu, C. L. Philip Chen, Jinyi Long |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Cold-start User Recommendation via Heterogeneous Domain AdaptationabstractIn recommendation systems, cold-start user recommendation is a challenging problem, where precise recommendations are required for users who have not appeared before. Several existing cold-start user recommendation models adopt domain adaptation to extract information from auxiliary source domains to assist the recommendations on the target domain. In this article, we propose that the cold-start user recommendation problem can be formulated by the heterogeneous domain adaption approach. We determine a transformation of user features, e.g., user social relations and historical interactions between warm users and their interested items, into a latent space so that the loss function is set by user feature reconstruction and by feature and distribution matching in the heterogeneous domains. The resulting optimization problem can be solved by matrix eigendecomposition, and the cold-start users’ preferences can thus be obtained. We also extend the proposed model using neural networks. We perform extensive experiments on several real-world datasets, and the results in terms of Precision, Recall, NDCG, and Hit Rate verify the effectiveness of the proposed model. Hanrui Wu, Yanxin Wu, Nuosi Li, Jia Zhang 0019, Michael Kwok-Po Ng, Jinyi Long |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Hypergraph Joint Representation Learning for Hypervertices and Hyperedges via Cross ExpansionabstractHypergraph captures high-order information in structured data and obtains much attention in machine learning and data mining. Existing approaches mainly learn representations for hypervertices by transforming a hypergraph to a standard graph, or learn representations for hypervertices and hyperedges in separate spaces. In this paper, we propose a hypergraph expansion method to transform a hypergraph to a standard graph while preserving high-order information. Different from previous hypergraph expansion approaches like clique expansion and star expansion, we transform both hypervertices and hyperedges in the hypergraph to vertices in the expanded graph, and construct connections between hypervertices or hyperedges, so that richer relationships can be used in graph learning. Based on the expanded graph, we propose a learning model to embed hypervertices and hyperedges in a joint representation space. Compared with the method of learning separate spaces for hypervertices and hyperedges, our method is able to capture common knowledge involved in hypervertices and hyperedges, and also improve the data efficiency and computational efficiency. To better leverage structure information, we minimize the graph reconstruction loss to preserve the structure information in the model. We perform experiments on both hypervertex classification and hyperedge classification tasks to demonstrate the effectiveness of our proposed method. Yuguang Yan, Hanrui Wu, Ruichu Cai |
AAAI | 4 |
| 2024 | High-order proximity and relation analysis for cross-network heterogeneous node classification
Hanrui Wu, Yanxin Wu, Nuosi Li, Min Yang 0007, Jia Zhang 0019, Michael Kwok-Po Ng, Jinyi Long |
Mach. Learn. | 1 |
| 2024 | Simplicial Complex Neural NetworksabstractGraph-structured data, where nodes exhibit either pair-wise or high-order relations, are ubiquitous and essential in graph learning. Despite the great achievement made by existing graph learning models, these models use the direct information (edges or hyperedges) from graphs and do not adopt the underlying indirect information (hidden pair-wise or high-order relations). To address this issue, in this paper, we propose a general framework named Simplicial Complex Neural (SCN) network, in which we construct a simplicial complex based on the direct and indirect graph information from a graph so that all information can be employed in the complex network learning. Specifically, we learn representations of simplices by aggregating and integrating information from all the simplices together via layer-by-layer simplicial complex propagation. In consequence, the representations of nodes, edges, and other high-order simplices are obtained simultaneously and can be used for learning purposes. By making use of block matrix properties, we derive the theoretical bound of the simplicial complex filter learnt by the propagation and establish the generalization error bound of the proposed simplicial complex network. We perform extensive experiments on node (0-simplex), edge (1-simplex), and triangle (2-simplex) classifications, and promising results demonstrate the performance of the proposed method is better than that of existing graph and hypergraph network approaches. Hanrui Wu, Andy M. Yip, Jinyi Long, Jia Zhang 0019, Michael Kwok-Po Ng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Semi-supervised imbalanced multi-label classification with label propagation
Guodong Du 0002, Jia Zhang 0019, Hanrui Wu, Peiliang Wu, Shaozi Li |
Pattern Recognit. | 4 |
| 2024 | Collaborative contrastive learning for hypergraph node classification
Hanrui Wu, Nuosi Li, Jia Zhang 0019, Sentao Chen, Michael Kwok-Po Ng, Jinyi Long |
Pattern Recognit. | 1 |
| 2024 | Transferable graph auto-encoders for cross-network node classification
Hanrui Wu, Yanxin Wu, Jia Zhang 0019, Michael Kwok-Po Ng, Jinyi Long |
Pattern Recognit. | 1 |
| 2024 | Feature Matching Machine for Cold-Start RecommendationabstractIn recommendation systems, the cold-start issue is a long-standing problem where no historical interaction records are given for certain users or items. Under this circumstance, recommendations for new users or new items become challenging. To address this problem, most existing approaches seek to discover a latent common space for users and items. However, these methods require a strong assumption that a shared space exists where the distributions of users and items are identical, which may limit the recommendation performance. In this article, we propose a novel model called Feature Matching Machine (FMM) to learn latent informative user and item representations. Different from previous methods, for warm users (or items), FMM learns two kinds of latent features, i.e., one is constructed by a hypergraph auto-encoder based on historical interactions between users and items, and the other is built by a multi-layer perceptron based on users (or items). Subsequently, FMM matches these two latent feature representations so as to discover the relationships across users (or items) and cold-start items (or users). We conduct extensive experiments on several real-world datasets and compare the proposed method with well-known baseline methods. Promising results demonstrate the effectiveness and efficiency of the proposed model. Hanrui Wu, Nuosi Li, Ka Ho Kwok, Xuheng Cai, Jia Zhang 0019, Jinyi Long, Michael Kwok-Po Ng |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Iterative Refinement for Multi-Source Visual Domain Adaptation (Extended abstract)abstractMulti-source domain adaptation (MSDA) aims to leverage the knowledge in multiple source domains to assist the prediction in a target domain, where the source and target domains have different data distributions. This paper presents a MSDA model to investigate both domain discrepancy and domain relevance, whose interactions are also exploited to gradually refine the learning performance. Particularly, the proposed model contains two components, i.e., feature spaces learning and transferred weights learning. The former one minimizes the domain discrepancy and the latter one evaluates the domain relevance. Experimental results on several real-world datasets demonstrate the effectiveness of the proposed model. Hanrui Wu, Yuguang Yan, Guosheng Lin, Min Yang 0007, Michael Kwok-Po Ng, Qingyao Wu |
ICDE | 1 |
| 2023 | Transferable Feature Selection for Unsupervised Domain Adaptation : Extended AbstractabstractDomain adaptation aims at extracting knowledge from auxiliary source domains to assist the learning task in a target domain. Since the distributions of the source and target domains are different, directly using source data to build a classifier for the target domain may hamper the classification performance on the target data. In this paper, we propose to find a feature subset that is both transferable and discriminative, so that both the domain discrepancy and the classification loss measured on the selected features can be reduced. To achieve this, we formulate a new sparse learning model that is able to jointly reduce the domain discrepancy and select informative features for classification. Extensive experiments on real-world data sets demonstrate the effectiveness of the proposed method. Yuguang Yan, Hanrui Wu, Yuzhong Ye, Chaoyang Bi, Qingyao Wu, Michael Kwok-Po Ng |
ICDE | 2 |
| 2023 | Group-preserving label-specific feature selection for multi-label learning
Jia Zhang 0019, Hanrui Wu, Min Jiang 0005, Shaozi Li, Yong Tang 0001, Jinyi Long |
Expert Syst. Appl. | 2 |
| 2023 | Hypergraph Collaborative Network on Vertices and HyperedgesabstractIn many practical datasets, such as co-citation and co-authorship, relationships across the samples are more complex than pair-wise. Hypergraphs provide a flexible and natural representation for such complex correlations and thus obtain increasing attention in the machine learning and data mining communities. Existing deep learning-based hypergraph approaches seek to learn the latent vertex representations based on either vertices or hyperedges from previous layers and focus on reducing the cross-entropy error over labeled vertices to obtain a classifier. In this paper, we propose a novel model called Hypergraph Collaborative Network (HCoN), which takes the information from both previous vertices and hyperedges into consideration to achieve informative latent representations and further introduces the hypergraph reconstruction error as a regularizer to learn an effective classifier. We evaluate the proposed method on two cases, i.e., semi-supervised vertex and hyperedge classifications. We carry out the experiments on several benchmark datasets and compare our method with several state-of-the-art approaches. Experimental results demonstrate that the performance of the proposed method is better than that of the baseline methods. Hanrui Wu, Yuguang Yan, Michael Kwok-Po Ng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Riemannian representation learning for multi-source domain adaptation
Sentao Chen, Lin Zheng 0003, Hanrui Wu |
Pattern Recognit. | 3 |
| 2023 | Adversarial Auto-encoder Domain Adaptation for Cold-start Recommendation with Positive and Negative HypergraphsabstractThis article presents a novel model named Adversarial Auto-encoder Domain Adaptation to handle the recommendation problem under cold-start settings. Specifically, we divide the hypergraph into two hypergraphs, i.e., a positive hypergraph and a negative one. Below, we adopt the cold-start user recommendation for illustration. After achieving positive and negative hypergraphs, we apply hypergraph auto-encoders to them to obtain positive and negative embeddings of warm users and items. Additionally, we employ a multi-layer perceptron to get warm and cold-start user embeddings called regular embeddings. Subsequently, for warm users, we assign positive and negative pseudo-labels to their positive and negative embeddings, respectively, and treat their positive and regular embeddings as the source and target domain data, respectively. Then, we develop a matching discriminator to jointly minimize the classification loss of the positive and negative warm user embeddings and the distribution gap between the positive and regular warm user embeddings. In this way, warm users’ positive and regular embeddings are connected. Since the positive hypergraph maintains the relations between positive warm user and item embeddings, and the regular warm and cold-start user embeddings follow a similar distribution, the regular cold-start user embedding and positive item embedding are bridged to discover their relationship. The proposed model can be easily extended to handle the cold-start item recommendation by changing inputs. We perform extensive experiments on real-world datasets for both cold-start user and cold-start item recommendations. Promising results in terms of precision, recall, normalized discounted cumulative gain, and hit rate verify the effectiveness of the proposed method. Hanrui Wu, Jinyi Long, Nuosi Li, Dahai Yu 0001, Michael Kwok-Po Ng |
ACM Trans. Inf. Syst. | 1 |
| 2023 | Cold-Start Next-Item Recommendation by User-Item Matching and Auto-EncodersabstractRecommendation systems provide personalized service to users and aim at suggesting to them items that they may prefer. There is an increasing requirement of next-item recommendation systems to infer a user's next favor item based on his/her historical selection of items. In this article, we study the next-item recommendation under the cold-start situation, where the users in the system share no interaction with the new items. Specifically, we seek to address the problem from the perspective of zero-shot learning (ZSL), which classifies samples whose classes are unseen during training. To this end, we crystallize the relationship and setting from ZSL to cold-start next-item recommendation, and further propose a novel model called User-Item Matching and Auto-encoders (UIMA) which learns the latent embeddings for both users and items by exploiting user historical preferences and item attributes. Concretely, UIMA consists of three components, i.e., two auto-encoders for learning user and item embeddings and a matching network to explore the relationship between the learned user and item embeddings. We perform experiments on several cold-start next-item recommendation datasets, including movies, music, and bookmarks. Promising results demonstrate the effectiveness of the proposed method for cold-start next-item recommendation. Hanrui Wu, Chung Wang Wong, Jia Zhang 0019, Yuguang Yan, Dahai Yu 0001, Jinyi Long, Michael Kwok-Po Ng |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | Multiple Graphs and Low-Rank Embedding for Multi-Source Heterogeneous Domain AdaptationabstractMulti-source domain adaptation is a challenging topic in transfer learning, especially when the data of each domain are represented by different kinds of features, i.e., Multi-source Heterogeneous Domain Adaptation (MHDA). It is important to take advantage of the knowledge extracted from multiple sources as well as bridge the heterogeneous spaces for handling the MHDA paradigm. This article proposes a novel method named Multiple Graphs and Low-rank Embedding (MGLE), which models the local structure information of multiple domains using multiple graphs and learns the low-rank embedding of the target domain. Then, MGLE augments the learned embedding with the original target data. Specifically, we introduce the modules of both domain discrepancy and domain relevance into the multiple graphs and low-rank embedding learning procedure. Subsequently, we develop an iterative optimization algorithm to solve the resulting problem. We evaluate the effectiveness of the proposed method on several real-world datasets. Promising results show that the performance of MGLE is better than that of the baseline methods in terms of several metrics, such as AUC, MAE, accuracy, precision, F1 score, and MCC, demonstrating the effectiveness of the proposed method. Hanrui Wu, Michael Kwok-Po Ng |
ACM Trans. Knowl. Discov. Data | 1 |
| 2022 | Hypergraph Convolution on Nodes-Hyperedges Network for Semi-Supervised Node ClassificationabstractHypergraphs have shown great power in representing high-order relations among entities, and lots of hypergraph-based deep learning methods have been proposed to learn informative data representations for the node classification problem. However, most of these deep learning approaches do not take full consideration of either the hyperedge information or the original relationships among nodes and hyperedges. In this article, we present a simple yet effective semi-supervised node classification method named Hypergraph Convolution on Nodes-Hyperedges network, which performs filtering on both nodes and hyperedges as well as recovers the original hypergraph with the least information loss. Instead of only reducing the cross-entropy loss over the labeled samples as most previous approaches do, we additionally consider the hypergraph reconstruction loss as prior information to improve prediction accuracy. As a result, by taking both the cross-entropy loss on the labeled samples and the hypergraph reconstruction loss into consideration, we are able to achieve discriminative latent data representations for training a classifier. We perform extensive experiments on the semi-supervised node classification problem and compare the proposed method with state-of-the-art algorithms. The promising results demonstrate the effectiveness of the proposed method. Hanrui Wu, Michael Kwok-Po Ng |
ACM Trans. Knowl. Discov. Data | 1 |
| 2022 | Iterative Refinement for Multi-Source Visual Domain AdaptationabstractOne of the main challenges in multi-source domain adaptation is how to reduce the domain discrepancy between each source domain and a target domain, and then evaluate the domain relevance to determine how much knowledge should be transferred from different source domains to the target domain. However, most prior approaches barely consider both discrepancies and relevance among domains. In this paper, we propose an algorithm, called Iterative Refinement based on Feature Selection and the Wasserstein distance (IRFSW), to solve semi-supervised domain adaptation with multiple sources. Specifically, IRFSW aims to explore both the discrepancies and relevance among domains in an iterative learning procedure, which gradually refines the learning performance until the algorithm stops. In each iteration, for each source domain and the target domain, we develop a sparse model to select features in which the domain discrepancy and training loss are reduced simultaneously. Then a classifier is constructed with the selected features of the source and labeled target data. After that, we exploit optimal transport over the selected features to calculate the transferred weights. The weight values are taken as the ensemble weights to combine the learned classifiers to control the amount of knowledge transferred from source domains to the target domain. Experimental results validate the effectiveness of the proposed method. Hanrui Wu, Yuguang Yan, Guosheng Lin, Min Yang 0007, Michael Kwok-Po Ng, Qingyao Wu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Transferable Feature Selection for Unsupervised Domain AdaptationabstractDomain adaptation aims at extracting knowledge from auxiliary source domains to assist the learning task in a target domain. In classification problems, since the distributions of the source and target domains are different, directly using source data to build a classifier for the target domain may hamper the classification performance on the target data. Fortunately, in many tasks, there can be some features that are transferable, i.e., the source and target domains share similar properties. On the other hand, it is common that the source data contain noisy features which may degrade the learning performance in the target domain. This issue, however, is barely studied in existing works. In this paper, we propose to find a feature subset that is transferable across the source and target domains. As a result, the domain discrepancy measured on the selected features can be reduced. Moreover, we seek to find the most discriminative features for classification. To achieve the above goals, we formulate a new sparse learning model that is able to jointly reduce the domain discrepancy and select informative features for classification. We develop two optimization algorithms to address the derived learning problem. Extensive experiments on real-world data sets demonstrate the effectiveness of the proposed method. Yuguang Yan, Hanrui Wu, Yuzhong Ye, Chaoyang Bi, Qingyao Wu, Michael Kwok-Po Ng |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Knowledge Preserving and Distribution Alignment for Heterogeneous Domain AdaptationabstractDomain adaptation aims at improving the performance of learning tasks in a target domain by leveraging the knowledge extracted from a source domain. To this end, one can perform knowledge transfer between these two domains. However, this problem becomes extremely challenging when the data of these two domains are characterized by different types of features, i.e., the feature spaces of the source and target domains are different, which is referred to as heterogeneous domain adaptation (HDA). To solve this problem, we propose a novel model called Knowledge Preserving and Distribution Alignment (KPDA), which learns an augmented target space by jointly minimizing information loss and maximizing domain distribution alignment. Specifically, we seek to discover a latent space, where the knowledge is preserved by exploiting the Laplacian graph terms and reconstruction regularizations. Moreover, we adopt the Maximum Mean Discrepancy to align the distributions of the source and target domains in the latent space. Mathematically, KPDA is formulated as a minimization problem with orthogonal constraints, which involves two projection variables. Then, we develop an algorithm based on the Gauss–Seidel iteration scheme and split the problem into two subproblems, which are solved by searching algorithms based on the Barzilai–Borwein (BB) stepsize. Promising results demonstrate the effectiveness of the proposed method. Hanrui Wu, Qingyao Wu, Michael Kwok-Po Ng |
ACM Trans. Inf. Syst. | 1 |
| 2021 | Domain Invariant and Agnostic Adaptation
Sentao Chen, Hanrui Wu, Cheng Liu 0001 |
Knowl. Based Syst. | 2 |
| 2021 | Joint Visual and Semantic Optimization for zero-shot learning
Hanrui Wu, Yuguang Yan, Sentao Chen, Xiangkang Huang, Qingyao Wu, Michael Kwok-Po Ng |
Knowl. Based Syst. | 1 |
| 2021 | Heterogeneous Domain Adaptation by Information Capturing and Distribution MatchingabstractHeterogeneous domain adaptation (HDA) is a challenging problem because of the different feature representations in the source and target domains. Most HDA methods search for mapping matrices from the source and target domains to discover latent features for learning. However, these methods barely consider the reconstruction error to measure the information loss during the mapping procedure. In this paper, we propose to jointly capture the information and match the source and target domain distributions in the latent feature space. In the learning model, we propose to minimize the reconstruction loss between the original and reconstructed representations to preserve information during transformation and reduce the Maximum Mean Discrepancy between the source and target domains to align their distributions. The resulting minimization problem involves two projection variables with orthogonal constraints that can be solved by the generalized gradient flow method, which can preserve orthogonal constraints in the computational procedure. We conduct extensive experiments on several image classification datasets to demonstrate that the effectiveness and efficiency of the proposed method are better than those of state-of-the-art HDA methods. Hanrui Wu, Hong Zhu 0012, Yuguang Yan, Jiaju Wu 0001, Yifan Zhang 0004, Michael Kwok-Po Ng |
IEEE Trans. Image Process. | 1 |
| 2020 | Geometric Knowledge Embedding for unsupervised domain adaptation
Hanrui Wu, Yuguang Yan, Yuzhong Ye, Michael Kwok-Po Ng, Qingyao Wu |
Knowl. Based Syst. | 1 |
| 2020 | Domain-attention Conditional Wasserstein Distance for Multi-source Domain AdaptationabstractMulti-source domain adaptation has received considerable attention due to its effectiveness of leveraging the knowledge from multiple related sources with different distributions to enhance the learning performance. One of the fundamental challenges in multi-source domain adaptation is how to determine the amount of knowledge transferred from each source domain to the target domain. To address this issue, we propose a new algorithm, called Domain-attention Conditional Wasserstein Distance (DCWD), to learn transferred weights for evaluating the relatedness across the source and target domains. In DCWD, we design a new conditional Wasserstein distance objective function by taking the label information into consideration to measure the distance between a given source domain and the target domain. We also develop an attention scheme to compute the transferred weights of different source domains based on their conditional Wasserstein distances to the target domain. After that, the transferred weights can be used to reweight the source data to determine their importance in knowledge transfer. We conduct comprehensive experiments on several real-world data sets, and the results demonstrate the effectiveness and efficiency of the proposed method. Hanrui Wu, Yuguang Yan, Michael Kwok-Po Ng, Qingyao Wu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2019 | Online Heterogeneous Transfer Learning by Knowledge TransitionabstractIn this article, we study the problem of online heterogeneous transfer learning, where the objective is to make predictions for a target data sequence arriving in an online fashion, and some offline labeled instances from a heterogeneous source domain are provided as auxiliary data. The feature spaces of the source and target domains are completely different, thus the source data cannot be used directly to assist the learning task in the target domain. To address this issue, we take advantage of unlabeled co-occurrence instances as intermediate supplementary data to connect the source and target domains, and perform knowledge transition from the source domain into the target domain. We propose a novel online heterogeneous transfer learning algorithm called O nline H eterogeneous K nowledge T ransition (OHKT) for this purpose. In OHKT, we first seek to generate pseudo labels for the co-occurrence data based on the labeled source data, and then develop an online learning algorithm to classify the target sequence by leveraging the co-occurrence data with pseudo labels. Experimental results on real-world data sets demonstrate the effectiveness and efficiency of the proposed algorithm. Hanrui Wu, Yuguang Yan, Yuzhong Ye, Huaqing Min, Michael Kwok-Po Ng, Qingyao Wu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2018 | Semi-Supervised Optimal Transport for Heterogeneous Domain AdaptationabstractHeterogeneous domain adaptation (HDA) aims to exploit knowledge from a heterogeneous source domain to improve the learning performance in a target domain. Since the feature spaces of the source and target domains are different, the transferring of knowledge is extremely difficult. In this paper, we propose a novel semi-supervised algorithm for HDA by exploiting the theory of optimal transport (OT), a powerful tool originally designed for aligning two different distributions. To match the samples between heterogeneous domains, we propose to preserve the semantic consistency between heterogeneous domains by incorporating label information into the entropic Gromov-Wasserstein discrepancy, which is a metric in OT for different metric spaces, resulting in a new semi-supervised scheme. Via the new scheme, the target and transported source samples with the same label are enforced to follow similar distributions. Lastly, based on the Kullback-Leibler metric, we develop an efficient algorithm to optimize the resultant problem. Comprehensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of our proposed method. Yuguang Yan, Wen Li 0001, Hanrui Wu, Huaqing Min, Mingkui Tan, Qingyao Wu |
IJCAI | 3 |
| 2017 | Learning Discriminative Correlation Subspace for Heterogeneous Domain AdaptationabstractDomain adaptation aims to reduce the effort on collecting and annotating target data by leveraging knowledge from a different source domain. The domain adaptation problem will become extremely challenging when the feature spaces of the source and target domains are different, which is also known as the heterogeneous domain adaptation (HDA) problem. In this paper, we propose a novel HDA method to find the optimal discriminative correlation subspace for the source and target data. The discriminative correlation subspace is inherited from the canonical correlation subspace between the source and target data, and is further optimized to maximize the discriminative ability for the target domain classifier. We formulate a joint objective in order to simultaneously learn the discriminative correlation subspace and the target domain classifier. We then apply an alternating direction method of multiplier (ADMM) algorithm to address the resulting non-convex optimization problem. Comprehensive experiments on two real-world data sets demonstrate the effectiveness of the proposed method compared to the state-of-the-art methods. Yuguang Yan, Wen Li 0001, Michael Kwok-Po Ng, Mingkui Tan, Hanrui Wu, Huaqing Min, Qingyao Wu |
IJCAI | 5 |
| 2017 | Online transfer learning by leveraging multiple source domains
Qingyao Wu, Xiaoming Zhou, Yuguang Yan, Hanrui Wu, Huaqing Min |
Knowl. Inf. Syst. | 4 |
| 2017 | Online Transfer Learning with Multiple Homogeneous or Heterogeneous SourcesabstractTransfer learning techniques have been broadly applied in applications where labeled data in a target domain are difficult to obtain while a lot of labeled data are available in related source domains. In practice, there can be multiple source domains that are related to the target domain, and how to combine them is still an open problem. In this paper, we seek to leverage labeled data from multiple source domains to enhance classification performance in a target domain where the target data are received in an online fashion. This problem is known as the online transfer learning problem. To achieve this, we propose novel online transfer learning paradigms in which the source and target domains are leveraged adaptively. We consider two different problem settings: homogeneous transfer learning and heterogeneous transfer learning. The proposed methods work in an online manner, where the weights of the source domains are adjusted dynamically. We provide the mistake bounds of the proposed methods and perform comprehensive experiments on real-world data sets to demonstrate the effectiveness of the proposed algorithms. Qingyao Wu, Hanrui Wu, Xiaoming Zhou, Mingkui Tan, Yuguang Yan, Tianyong Hao |
IEEE Trans. Knowl. Data Eng. | 2 |