Pei Yang 0001

dblp:51/6008-1 · DBLP profile ↗
← Back
39ranked-venue papers
18as first author
17since 2021 · last 2025
0000-0001-8926-9695ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 12 first-author · 7 since 2021Databases, data management, data science and information retrieval · 16 · 14 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Multi-label body constitution recognition via dual transform MLP-like architecture using tongue images
abstract
According to the traditional Chinese medicine composite constitution theory, body constitution recognition is modeled as a unique task of multi-label problems using tongue images. Although MLP-like architecture is one of the base models except for CNN-based, Transformer-based, and MLP-like-based for computer vision tasks, the performance of the MLP-like-based architecture for the multi-label BCR task needs to expand for multi-label BCR task with a specific module using tongue images. Thus, a novel dual transform MLP-like (DT-MLP) model is designed, in which discrete Fourier transform and chaos transform are used for feature extraction of their branches. Besides, a new multi-label tongue body constitution dataset is constructed. The results demonstrate that the proposed DT-MLP outperforms other state-of-the-art MLP-like models on mAP, AUC, and Acc, respectively.
Mengjian Zhang, Guihua Wen, Pei Yang 0001
ICASSP3
2025 Bimodal speech emotion recognition via contrastive self-alignment learning
Guihua Wen, Pei Yang 0001, Lianqi Liu
Expert Syst. Appl.3
2025 Multi-label body constitution recognition via HWmixer-MLP for facial and tongue images
Mengjian Zhang, Guihua Wen, Pei Yang 0001, Changjun Wang, Chuyun Chen
Expert Syst. Appl.3
2025 Chaos-MLP: Chaotic Transform MLP-Like Architecture for Medical Images Multi-Label Recognition Task
abstract
The theory of "three-stage prevention" in view of the body constitution is the key technology of modern Chinese medicine for "Preventive Treatment of Diseases". In particular, automated body constitution recognition (BCR) is an integral part of intelligent Traditional Chinese Medicine (TCM), which is extremely valuable for disease prevention and diagnosis. Actually, BCR is a challenging multi-label recognition task by the TCM composite constitution theory. First, two new databases are constructed, one is a multi-label facial body constitution (MFBC), and another is a multi-label tongue body constitution (MTBC). Second, a novel MLP-like architecture, named Chaos-MLP, is designed for the BCR task, which interacts with the channel chaotic features of extracted medical images and fuses them with the width and height channel direction features, respectively. Notably, the chaotic transform can enhance the distinguishability of extracted features from the medical images. Moreover, we propose a binary center cognitive gravity loss (BCCGL) to enhance the learning ability of the Chaos-MLP for unbalanced body constitution labels. Our proposed method shows superior performance on both MFBC and MTBC datasets than other state-of-the-art (SOTA) MLP-like networks and a vision graph-based neural network (VGNN), which include Wave-MLP, Cycle-MLP, Vip, and Active-MLP.
Mengjian Zhang, Guihua Wen, Pei Yang 0001, Changjun Wang, Xuhui Huang, Chuyun Chen
IEEE J. Biomed. Health Informatics3
2024 Speech Relationship Learning for Cross-Corpus Speech Emotion Recognition
abstract
Cross-Corpus Speech Emotion Recognition (SER) aims to identify human emotions from speech across different speakers and languages. Previous work engaged in extracting the domain-invariant features among individual samples that are most relevant to emotions, ignoring rich relationships between speech instances, which are also significant factors that strongly influence the sentiments. To explore those potential relationships across multiple corpora, we introduce a novel cross-corpus SER architecture with speech relationship learning. Specifically, during training, we employ the attention mechanism on the entire input batch, embedding the sample-level similar features in emotion space into new representations. Furthermore, a dual discriminator structure is proposed for improving the similarity calculation performance through adversarial training, and a domain-wise shared classifier with batch label smoothing strategy is proposed to enhance the network generalization ability. Experiments on the CASIA, EMODB and SAVEE datasets have demonstrated that the proposed method outperforms the state-of-the-art cross-corpus SER methods.
Yinru He, Guihua Wen, Pei Yang 0001
ICASSP3
2024 Multi-geometry embedded transformer for facial expression recognition in videos
Guihua Wen, Pei Yang 0001, Chuyun Chen
Expert Syst. Appl.4
2024 CDGT: Constructing diverse graph transformers for emotion recognition from facial videos
Guihua Wen, Pei Yang 0001, Chuyun Chen
Neural Networks4
2024 From patch, sample to domain: Capture geometric structures for few-shot learning
Qiaonan Li, Guihua Wen, Pei Yang 0001
Pattern Recognit.3
2024 CFAN-SDA: Coarse-Fine Aware Network With Static-Dynamic Adaptation for Facial Expression Recognition in Videos
abstract
Video-based facial expression recognition (FER) is a challenging task due to the dynamic emotional changes with variant frames in video sequences. This paper proposes a novel coarse-fine aware network with static-dynamic adaptation (CFAN-SDA) for in-the wild video-based FER. From coarse to fine, our method leverages cross-domain static FER database to boost video-based FER performance, and then explore hierarchical spatial-temporal feature learning. Specifically, different from existing methods, we design a static-dynamic adaptation learning to explore the knowledge transfer from labeled static images to unlabeled frames of video, which captures the features of coarse-grained emotion to find those important expression-related frames. Furthermore, we present hierarchical spatial-temporal transformers to better learn features of fine-grained expression, which consist of multi-view spatial transformer and frame-clip temporal transformer. The former captures multi-view spatial regions information from global to local, and the latter achieves cross-frame and cross-clip temporal interaction to select the key frame-level and clip-level multi-scale temporal information for fusing. Extensive experimental results on dynamic FER databases indicate that CFAN-SDA achieves superior performance compared to the state-of-the-art models.
Guihua Wen, Pei Yang 0001, Chuyun Chen
IEEE Trans. Circuits Syst. Video Technol.3
2024 Modeling Hierarchical Structural Distance for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to estimate a transferable model for unlabeled target domains by exploiting labeled source data. Optimal Transport (OT) based methods have recently been proven to be a promising solution for UDA with a solid theoretical foundation and competitive performance. However, most of these methods solely focus on domain-level OT alignment by leveraging the geometry of domains for domain-invariant features based on the global embeddings of images. However, global representations of images may destroy image structure, leading to the loss of local details that offer category-discriminative information. This study proposes an end-to-end Deep Hierarchical Optimal Transport method (DeepHOT), which aims to learn both domain-invariant and category-discriminative representations by mining hierarchical structural relations among domains. The main idea is to incorporate a domain-level OT and image-level OT into a unified OT framework, hierarchical optimal transport, to model the underlying geometry in both domain space and image space. In DeepHOT framework, an image-level OT serves as the ground distance metric for the domain-level OT, leading to the hierarchical structural distance. Compared with the ground distance of the conventional domain-level OT, the image-level OT captures structural associations among local regions of images that are beneficial to classification. In this way, DeepHOT, a unified OT framework, not only aligns domains by domain-level OT, but also enhances the discriminative power through image-level OT. Moreover, to overcome the limitation of high computational complexity, we propose a robust and efficient implementation of DeepHOT by approximating origin OT with sliced Wasserstein distance in image-level OT and accomplishing the mini-batch unbalanced domain-level OT. Extensive experiments show the superiority of DeepHOT in several benchmark datasets. The code will be released on GitHub.
Yingxue Xu, Guihua Wen, Pei Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Cross-Domain Sample Relationship Learning for Facial Expression Recognition
abstract
Cross-domain facial expression recognition is confronted by the problem of the large distribution discrepancy and samples inconsistencies between the source domain and target domain. To solve this problem, we propose a cross-domain sample relationship learning (CSRL) method that explores useful intrinsic sample relationships of two domains to narrow the domain discrepancy. Specifically, during the training stage, we first design inter-domain sample transformers to explore the sample similarity relationships between the source and target domains, and then deploy intra-domain sample transformers to capture the internal similar structure of the samples in each domain. Thus dual sample relationships can be learned to align the cross-domain similar samples and preserve the domain-specific information, which can facilitate both the inter-domain invariant features and intra-domain invariant features learning. Subsequently, we design a joint alignment strategy by simultaneously deploying the feature distribution alignment and cross-domain sample relationship learning. Thus, both local similar samples and global domain distribution of two domains can be well aligned to enhance the generalization ability of the model. Experimental results on several benchmark databases show the superiority of CSRL over some state-of-the-art methods.
Guihua Wen, Pengcheng Wen, Pei Yang 0001, Cheng Li 0025
IEEE Trans. Multim.4
2022 Deep non-negative tensor factorization with multi-way EMG data
Qi Tan 0001, Pei Yang 0001, Guihua Wen
Neural Comput. Appl.2
2022 Task-Coupling Elastic Learning for Physical Sign-Based Medical Image Classification
abstract
Physical signs of patients indicate crucial evidence for diagnosing both location and nature of the disease, where there is a sequential relationship between the two tasks. Thus their joint learning can utilize intrinsic association by transferring related knowledge across relevant tasks. Choosing the right time to transfer is a critical problem for joint learning. However, how to dynamically adjust when tasks interact to capture the right time for transferring related knowledge is still an open issue. To this end, we propose a Task-Coupling Elastic Learning (TCEL) framework to model the task relatedness for classifying disease-location and disease-nature based on physical sign images. The main idea is to dynamically transfer relevant knowledge by progressively shifting task-coupling from loose to tight during the multi-stage training. In the early stage of training, we relax the constraints of modeling relations to focus more in learning the generic task-common features. In the later stage, the semantic guidance will be strengthened to learn the task-specific features. Specifically, a dynamic sequential module (DSM) is proposed to explicitly model the sequential relationship and enable multi-stage training. Moreover, to address the side effect of DSM, a new loss regularization is proposed. The extensive experiments on these two clinical datasets show the superiority of the proposed method over the baselines, and demonstrate the effectiveness of the proposed task-coupling elastic mechanism.
Yingxue Xu, Guihua Wen, Pei Yang 0001, Baochao Fan, Mingnan Luo, Changjun Wang
IEEE J. Biomed. Health Informatics3
2022 Graph-Based Visual-Semantic Entanglement Network for Zero-Shot Image Recognition
abstract
Zero-shot learning uses semantic attributes to connect the search space of unseen objects. In recent years, although the deep convolutional network brings powerful visual modeling capabilities to the ZSL task, its visual features have severe pattern inertia and lack of representation of semantic relationships, which leads to severe bias and ambiguity. In response to this, we propose the Graph-based Visual-Semantic Entanglement Network to conduct graph modeling of visual features, which is mapped to semantic attributes by using a knowledge graph, it contains several novel designs: 1. it establishes a multi-path entangled network with the convolutional neural network (CNN) and the graph convolutional network (GCN), which input the visual features from CNN to GCN to model the implicit semantic relations, then GCN feedback the graph modeled information to CNN features; 2. it uses attribute word vectors as the target for the graph semantic modeling of GCN, which forms a self-consistent regression for graph modeling and supervise GCN to learn more personalized attribute relations; 3. it fuses and supplements the hierarchical visual-semantic features refined by graph modeling into visual embedding. Our method outperforms state-of-the-art approaches on multiple representative ZSL datasets: AwA2, CUB, and SUN by promoting the semantic linkage modelling of visual features.
Guihua Wen, Adriane Chapman, Pei Yang 0001, Mingnan Luo, Yingxue Xu, Dan Dai, Wendy Hall 0001
IEEE Trans. Multim.4
2021 Fully-channel regional attention network for disease-location recognition with tongue images
Guihua Wen, Mingnan Luo, Pei Yang 0001, Dan Dai, Zhiwen Yu 0002, Changjun Wang, Wendy Hall 0001
Artif. Intell. Medicine4
2021 Multi-source Seq2seq guided by knowledge for Chinese healthcare consultation
abstract
Online healthcare consultation offers people a convenient way to consult doctors. In this paper, we aim at building a generative dialog system for Chinese healthcare consultation. As the original Seq2seq architecture tends to suffer the issue of generating low-quality responses, the multi-source Seq2seq architecture generating more informative responses is much more preferred in this task. The multi-source Seq2seq architecture takes advantage of retrieval techniques to obtain responses from the database, and then takes these responses alongside the user-issued question as input. However, some of the retrieved responses might be not much related to the user-issued question, resulting in the generation of unsatisfying responses that are not correct in diagnosis or instead provide inappropriate advice on prevention or treatment. Therefore, this paper proposes multi-source Seq2seq guided by knowledge (MSSGK) to handle this problem. MSSGK differs from the multi-source Seq2seq architecture in that domain knowledge, including disease labels and topic labels about prevention and treatment, is introduced into the response generation via a multi-task learning framework. To better exploit the domain knowledge, we propose three attention mechanisms to provide more appropriate guidance for response generation. Experimental results on a dataset of real-world healthcare consultation show the effectiveness of the proposed method.
Yanghui Li, Guihua Wen, Mingnan Luo, Baochao Fan, Changjun Wang, Pei Yang 0001
J. Biomed. Informatics7
2021 MVANet: Multi-Task Guided Multi-View Attention Network for Chinese Food Recognition
abstract
Food recognition plays a much critical role in various health-care applications. However, it poses many challenges to current approaches due to the diverse appearances of food dishes and the non-uniform composition of ingredients for the foods in the same category. Current methods primarily focus on the appearance of foods without considering their semantic information, easily finding the wrong attention areas of food images. Second, these methods lack the dynamic weighting of multiple semantic features in the modeling process. Thus this paper proposes a novel Multi-View Attention Network within the multi-task learning framework that incorporates multiple semantic features into the food recognition task from both ingredient recognition and recipe modeling. It also utilizes the multi-view attention mechanism to automatically adjust the weights of different semantic features and enables different tasks to interact with each other so as to obtain a more comprehensive feature representation. The experiments conducted on both ChineseFoodNet and VIREO Food-172 benchmark databases validate the proposed method with the obvious improvement of the performance and the lower parameter size.
Haozan Liang, Guihua Wen, Mingnan Luo, Pei Yang 0001, Yingxue Xu
IEEE Trans. Multim.5
2020 Complex heterogeneity learning: A theoretical and empirical study
Pei Yang 0001, Qi Tan 0001, Jingrui He
Pattern Recognit.1
2019 Deep Multi-Task Learning with Adversarial-and-Cooperative Nets
abstract
In this paper, we propose a deep multi-Task learning model based on Adversarial-and-COoperative nets (TACO). The goal is to use an adversarial-and-cooperative strategy to decouple the task-common and task-specific knowledge, facilitating the fine-grained knowledge sharing among tasks. TACO accommodates multiple game players, i.e., feature extractors, domain discriminator, and tri-classifiers. They play the MinMax games adversarially and cooperatively to distill the task-common and task-specific features, while respecting their discriminative structures. Moreover, it adopts a divide-and-combine strategy to leverage the decoupled multi-view information to further improve the generalization performance of the model. The experimental results show that our proposed method significantly outperforms the state-of-the-art algorithms on the benchmark datasets in both multi-task learning and semi-supervised domain adaptation scenarios.
Pei Yang 0001, Qi Tan 0001, Jieping Ye, Hanghang Tong, Jingrui He
IJCAI1
2019 Task-Adversarial Co-Generative Nets
abstract
In this paper, we propose Task-Adversarial co-Generative Nets (TAGN) for learning from multiple tasks. It aims to address the two fundamental issues of multi-task learning, i.e., domain shift and limited labeled data, in a principled way. To this end, TAGN first learns the task-invariant representations of features to bridge the domain shift among tasks. Based on the task-invariant features, TAGN generates the plausible examples for each task to tackle the data scarcity issue. In TAGN, we leverage multiple game players to gradually improve the quality of the co-generation of features and examples by using an adversarial strategy. It simultaneously learns the marginal distribution of task-invariant features across different tasks and the joint distributions of examples with labels for each task. The theoretical study shows the desired results: at the equilibrium point of the multi-player game, the feature extractor exactly produces the task-invariant features for different tasks, while both the generator and the classifier perfectly replicate the joint distribution for each task. The experimental results on the benchmark data sets demonstrate the effectiveness of the proposed approach.
Pei Yang 0001, Qi Tan 0001, Hanghang Tong, Jingrui He
KDD1
2018 Heterogeneous representation learning with separable structured sparsity regularization
Pei Yang 0001, Qi Tan 0001, Yada Zhu, Jingrui He
Knowl. Inf. Syst.1
2018 Feature co-shrinking for co-clustering
Qi Tan 0001, Pei Yang 0001, Jingrui He
Pattern Recognit.2
2018 Function-on-Function Regression with Mode-Sparsity Regularization
abstract
Functional data is ubiquitous in many domains, such as healthcare, social media, manufacturing process, sensor networks, and so on. The goal of function-on-function regression is to build a mapping from functional predictors to functional response. In this article, we propose a novel function-on-function regression model based on mode-sparsity regularization. The main idea is to represent the regression coefficient function between predictor and response as the double expansion of basis functions, and then use a mode-sparsity regularization to automatically filter out irrelevant basis functions for both predictors and responses. The proposed approach is further extended to the tensor version to accommodate multiple functional predictors. While allowing the dimensionality of the regression weight matrix or tensor to be relatively large, the mode-sparsity regularized model facilitates the multi-way shrinking of basis functions for each mode. The proposed mode-sparsity regularization covers a wide spectrum of sparse models for function-on-function regression. The resulting optimization problem is challenging due to the non-smooth property of the mode-sparsity regularization. We develop an efficient algorithm to solve the problem, which works in an iterative update fashion, and converges to the global optimum. Furthermore, we analyze the generalization performance of the proposed method and derive an upper bound for the consistency between the recovered function and the underlying true function. The effectiveness of the proposed approach is verified on benchmark functional datasets in various domains.
Pei Yang 0001, Qi Tan 0001, Jingrui He
ACM Trans. Knowl. Discov. Data1
2017 Multi-task Function-on-function Regression with Co-grouping Structured Sparsity
abstract
The growing importance of functional data has fueled the rapid development of functional data analysis, which treats the infinite-dimensional data as continuous functions rather than discrete, finite-dimensional vectors. On the other hand, heterogeneity is an intrinsic property of functional data due to the variety of sources to collect the data. In this paper, we propose a novel multi-task function-on-function regression approach to model both the functionality and heterogeneity of data. The basic idea is to simultaneously model the relatedness among tasks and correlations among basis functions by using the co-grouping structured sparsity to encourage similar tasks to behave similarly in shrinking the basis functions. The resulting optimization problem is challenging due to the non-smoothness and non-separability of the co-grouping structured sparsity. We present an efficient algorithm to solve the problem, and prove its separability, convexity, and global convergence. The proposed algorithm is applicable to a wide spectrum of structured sparsity regularized techniques, such as structured ℓ2,p norm and structured Schatten p-norm. The effectiveness of the proposed approach is verified on benchmark functional data sets collected from various domains.
Pei Yang 0001, Qi Tan 0001, Jingrui He
KDD1
2016 Heterogeneous Representation Learning with Structured Sparsity Regularization
abstract
Motivated by real applications, heterogeneous learning has emerged as an important research area, which aims to model the co-existence of multiple types of heterogeneity. In this paper, we propose a HEterogeneous REpresentation learning model with structured Sparsity regularization (HERES) to learn from multiple types of heterogeneity. HERES aims to leverage two kinds of information to build a robust learning system. One is the rich correlations among heterogeneous data such as task relatedness, view consistency, and label correlation. The other is the prior knowledge of the data in the form of, e.g., the soft-clustering of the tasks. HERES is a generic framework for heterogeneous learning, which integrates multi-task, multi-view, and multi-label learning into a principled framework based on representation learning. The objective of HERES is to minimize the reconstruction loss of using the factor matrices to recover the input matrix for heterogeneous data, regularized by the structured sparsity constraint. The resulting optimization problem is challenging due to the non-smoothness and non-separability of structured sparsity. We develop an iterative updating method to solve the problem. Furthermore, we prove that the reformulation of structured sparsity is separable, which leads to a family of efficient and scalable algorithms for solving structured sparsity penalized problems. The experimental results in comparison with state-of-the-art methods demonstrate the effectiveness of the proposed approach.
Pei Yang 0001, Jingrui He
ICDM1
2016 Functional Regression with Mode-Sparsity Constraint
abstract
Functional data is ubiquitous in many domains such as healthcare, social media, manufacturing process, sensor networks, etc. Functional data analysis involves the analysis of data which is treated as infinite-dimensional continuous functions rather than discrete, finite-dimensional vectors. In this paper, we propose a novel function-on-function regression model based on mode-sparsity regularization. The main idea is to represent the regression coefficient function between predictor and response as the double expansion of basis functions, and then use mode-sparsity constraint to automatically filter out the irrelevant basis functions for both predictors and responses. The mode-sparsity regularization covers a wide spectrum of sparse models for function-on-function regression. The resulting optimization problem is challenging due to the non-smooth property of the mode-sparsity. We develop an efficient and convergence-guaranteed algorithm to solve the problem. The effectiveness of the proposed approach is verified on benchmark functional data sets in various domains.
Pei Yang 0001, Jingrui He
ICDM1
2016 Jointly Modeling Label and Feature Heterogeneity in Medical Informatics
abstract
Multiple types of heterogeneity including label heterogeneity and feature heterogeneity often co-exist in many real-world data mining applications, such as diabetes treatment classification, gene functionality prediction, and brain image analysis. To effectively leverage such heterogeneity, in this article, we propose a novel graph-based model for Learning with both Label and Feature heterogeneity, namely L 2 F . It models the label correlation by requiring that any two label-specific classifiers behave similarly on the same views if the associated labels are similar, and imposes the view consistency by requiring that view-based classifiers generate similar predictions on the same examples. The objective function for L 2 F is jointly convex. To solve the optimization problem, we propose an iterative algorithm, which is guaranteed to converge to the global optimum. One appealing feature of L 2 F is that it is capable of handling data with missing views and labels. Furthermore, we analyze its generalization performance based on Rademacher complexity, which sheds light on the benefits of jointly modeling the label and feature heterogeneity. Experimental results on various biomedical datasets show the effectiveness of the proposed approach.
Pei Yang 0001, Hongxia Yang, Haoda Fu, Dawei Zhou 0003, Jieping Ye, Theodoros Lappas, Jingrui He
ACM Trans. Knowl. Discov. Data1
2016 A Generalized Hierarchical Multi-Latent Space Model for Heterogeneous Learning
abstract
In many real world applications such as image annotation, gene function prediction, and insider threat detection, the data collected from heterogeneous sources often exhibit multiple types of heterogeneity, such as task heterogeneity, view heterogeneity, and label heterogeneity. To address this problem, we propose a Hierarchical Multi-Latent Space (HiMLS) learning framework to jointly model the triple types of heterogeneity. The basic idea is to learn a hierarchical multi-latent space by which we can simultaneously leverage the task relatedness, view consistency and the label correlations to improve the learning performance. We first propose a multi-latent space approach to model the complex heterogeneity, which is then used as a building block to stack up a multi-layer structure in order to learn the hierarchical multi-latent space. In such a way, we can gradually learn the more abstract concepts in the higher level. We present two instantiated models of the generalized framework using different divergence measures. The two-phase learning algorithms are used to train the multi-layer models. We drive the multiplicative update rules for pre-training and fine-tuning in each model, and prove the convergence and correctness of the update methods. The effectiveness of the proposed approach is verified on various data sets.
Pei Yang 0001, Hasan Davulcu, Yada Zhu, Jingrui He
IEEE Trans. Knowl. Data Eng.1
2015 A Graph-Based Hybrid Framework for Modeling Complex Heterogeneity
abstract
Data heterogeneity is an intrinsic property of many high impact applications, such as insider threat detection, traffic prediction, brain image analysis, quality control in manufacturing processes, etc. Furthermore, multiple types of heterogeneity (e.g., task/view/instance heterogeneity) often co-exist in these applications, thus pose new challenges to existing techniques, most of which are tailored for a single or dual types of heterogeneity. To address this problem, in this paper, we propose a novel graph-based hybrid approach to simultaneously model multiple types of heterogeneity in a principled framework. The objective is to maximize the smoothness consistency of the neighboring nodes, bag-instance correlation together with task relatedness on the hybrid graphs, and simultaneously minimize the empirical classification loss. Furthermore, we analyze its performance based on Rademacher complexity, which sheds light on the benefits of jointly modeling multiple types of heterogeneity. To solve the resulting non-convex non-smooth problem, we propose an iterative algorithm named M3 Learning, which combines block coordinate descent and the bundle method for optimization. Experimental results on various data sets show the effectiveness of the proposed algorithm.
Pei Yang 0001, Jingrui He
ICDM1
2015 Model Multiple Heterogeneity via Hierarchical Multi-Latent Space Learning
abstract
In many real world applications such as satellite image analysis, gene function prediction, and insider threat detection, the data collected from heterogeneous sources often exhibit multiple types of heterogeneity, such as task heterogeneity, view heterogeneity, and label heterogeneity. To address this problem, we propose a Hierarchical Multi-Latent Space (HiMLS) learning approach to jointly model the triple types of heterogeneity. The basic idea is to learn a hierarchical multi-latent space by which we can simultaneously leverage the task relatedness, view consistency and the label correlations to improve the learning performance. We first propose a multi-latent space framework to model the complex heterogeneity, which is used as a building block to stack up a multi-layer structure so as to learn the hierarchical multi-latent space. In such a way, we can gradually learn the more abstract concepts in the higher level. Then, a deep learning algorithm is proposed to solve the optimization problem. The experimental results on various data sets show the effectiveness of the proposed approach.
Pei Yang 0001, Jingrui He
KDD1
2015 Learning Complex Rare Categories with Dual Heterogeneity
abstract
In the era of big data, it is often the case that the self-similar rare categories in a large data set are of great importance, such as the malicious insiders in big organizations, and the IC devices with defects in semiconductor manufacturing. Furthermore, such rare categories often exhibit multiple types of heterogeneity, such as the task heterogeneity, which originates from data collected in multiple domains, and the view heterogeneity, which originates from multiple information sources. Existing methods for learning rare categories mainly focus on the homogeneous settings, i.e., a single task and a single view. In this paper, for the first time, we study complex rare categories with both task and view heterogeneity, and propose a novel optimization framework named M2LID. It introduces a boundary characterization metric to capture the sharp changes in density near the boundary of the rare categories in the feature space, and constructs a graph-based model to leverage both task and view heterogeneity. Furthermore, M2LID integrates them in a way of mutual benefit. We also present an effective algorithm to solve this framework, analyze its performance from various aspects, and demonstrate its effectiveness on both synthetic and real datasets.
Pei Yang 0001, Jingrui He, Jia-Yu Pan
SDM1
2014 Learning from Label and Feature Heterogeneity
abstract
Multiple types of heterogeneity, such as label heterogeneity and feature heterogeneity, often co-exist in many real-world data mining applications, such as news article categorization, gene functionality prediction. To effectively leverage such heterogeneity, in this paper, we propose a novel graph-based framework for Learning with both Label and Feature heterogeneities, namely L2F. It models the label correlation by requiring that any two label-specific classifiers behave similarly on the same views if the associated labels are similar, and imposes the view consistency by requiring that view-based classifiers generate similar predictions on the same examples. To solve the resulting optimization problem, we propose an iterative algorithm, which is guaranteed to converge to the global optimum. Furthermore, we analyze its generalization performance based on Rademacher complexity, which sheds light on the benefits of jointly modeling the label and feature heterogeneity. Experimental results on various data sets show the effectiveness of the proposed approach.
Pei Yang 0001, Jingrui He, Hongxia Yang, Haoda Fu
ICDM1
2014 Democracy is good for ranking: towards multi-view rank learning and adaptation in web search
abstract
Web search ranking models are learned from features originated from different views or perspectives of document relevancy, such as query dependent or independent features. This seems intuitively conformant to the principle of multi-view approach that leverages distinct complementary views to improve model learning. In this paper, we aim to obtain optimal separation of ranking features into non-overlapping subsets (i.e., views), and use such different views for rank learning and adaptation. We present a novel semi-supervised multi-view ranking model, which is then extended into an adaptive ranker for search domains where no training data exists. The core idea is to proactively strengthen view consistency (i.e., the consistency between different rankings each predicted by a distinct view-based ranker) especially when training and test data follow divergent distributions. For this purpose, we propose a unified framework based on listwise ranking scheme to mutually reinforce the view consistency of target queries and the appropriate weighting of source queries that act as prior knowledge. Based on LETOR and Yahoo Learning to Rank datasets, our method significantly outperforms some strong baselines including single-view ranking models commonly used and multi-view ranking models that do not impose view consistency on target data.
Wei Gao 0001, Pei Yang 0001
WSDM2
2014 Information-Theoretic Multi-view Domain Adaptation: A Theoretical and Empirical Study
abstract
Multi-view learning aims to improve classification performance by leveraging the consistency among different views of data. The incorporation of multiple views was paid little attention in the studies of domain adaptation, where the view consistency based on source data is largely violated in the target domain due to the distribution gap between different domain data. In this paper, we leverage multiple views for cross-domain document classification. The central idea is to strengthen the views' consistency on target data by identifying the associations of domain-specific features from different domains. We present an Information-theoretic Multi-view Adaptation Model (IMAM) using a multi-way clustering scheme, where word and link clusters can draw together seemingly unrelated features across domains, which boosts the consistency between document clusterings that are based on the respective word and link views. Moreover, we demonstrate that IMAM can always find the document clustering with the minimal disagreement rate to the overlap of view-based clusterings. We provide both theoretical and empirical justifications of the proposed method. Our experiments show that IMAM significantly outperforms traditional multi-view algorithm co-training, the co-training-based adaptation algorithm CODA, the single-view transfer model CoCC and the large-margin-based multi-view transfer model MVTL-LM.
Pei Yang 0001, Wei Gao 0001
J. Artif. Intell. Res.1
2014 Knowledge transfer across different domain data with multiple views
Qi Tan 0001, Huifang Deng, Pei Yang 0001
Neural Comput. Appl.3
2013 Multi-View Discriminant Transfer Learning
Pei Yang 0001, Wei Gao 0001
IJCAI1
2013 A link-bridged topic model for cross-domain document classification
Pei Yang 0001, Wei Gao 0001, Qi Tan 0001, Kam-Fai Wong
Inf. Process. Manag.1
2012 Kernel Mean Matching with a Large Margin
Qi Tan 0001, Huifang Deng, Pei Yang 0001
ADMA3
2009 An Information-Theoretic Approach for Multi-task Learning
Pei Yang 0001, Qi Tan 0001, Yehua Ding
ADMA1