Changying Du

dblp:33/10023 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
1since 2021 · last 2022
0000-0002-4489-0020ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Generative modeling · 28% Probabilistic and Bayesian machine learning · 22% Representation and self-supervised learning · 18%
Databases, data mining, and information retrieval
4 papers
Recommender systems · 75% Information retrieval · 25%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 50% Computational social science and digital humanities · 50%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
multi-view learning
0.522017
Nonlinear Maximum Margin Multi-View Learning with Adaptive Kernel · IJCAI 2017
Online Bayesian Max-Margin Subspace Multi-View Learning · IJCAI 2016
Computer vision › 3D vision
brain decoding
0.412020
Conditional Generative Neural Decoding with Structured CNN Feature Prediction · AAAI 2020
Machine learning › Generative modeling › diffusion model
conditional generation
0.412020
Conditional Generative Neural Decoding with Structured CNN Feature Prediction · AAAI 2020
Computer vision › 3D vision › brain decoding
visual decoding
0.412020
Conditional Generative Neural Decoding with Structured CNN Feature Prediction · AAAI 2020
Recommender systems
item tagging
0.412020
JIT2R: A Joint Framework for Item Tagging and Tag-based Recommendation · SIGIR 2020
Recommender systems › side information integration
tag-based recommendation
0.412020
JIT2R: A Joint Framework for Item Tagging and Tag-based Recommendation · SIGIR 2020
Machine learning › Generative modeling › adversarial inference
adversarially learned inference
0.312018
Multi-view Adversarially Learned Inference for Cross-domain Joint Distribution Matching · KDD 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
Multi-view Adversarially Learned Inference for Cross-domain Joint Distribution Matching · KDD 2018
Computational social science and digital humanities › computational linguistics
opinion mining
0.312018
StockAssIstant: A Stock AI Assistant for Reliability Modeling of Stock Comments · KDD 2018
Computational finance and economics › financial market analysis
stock market analysis
0.312018
StockAssIstant: A Stock AI Assistant for Reliability Modeling of Stock Comments · KDD 2018
Information retrieval › image retrieval
hashing-based image retrieval
0.312018
Redundancy-resistant Generative Hashing for Image Retrieval · IJCAI 2018
Information retrieval
image retrieval
0.312018
Redundancy-resistant Generative Hashing for Image Retrieval · IJCAI 2018
Recommender systems › context-aware recommendation
emotion-aware recommendation
0.312017
A Location-Sentiment-Aware Recommender System for Both Home-Town and Out-of-Town Users · KDD 2017
Recommender systems
point-of-interest recommendation
0.312017
A Location-Sentiment-Aware Recommender System for Both Home-Town and Out-of-Town Users · KDD 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › recursive bayesian estimation
bayesian online learning
0.212016
Online Bayesian Max-Margin Subspace Multi-View Learning · IJCAI 2016
Machine learning › Optimization for machine learning › regularized risk minimization
max-margin learning
0.212016
Online Bayesian Max-Margin Subspace Multi-View Learning · IJCAI 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212015
Bayesian Maximum Margin Principal Component Analysis · AAAI 2015
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.212015
Bayesian Maximum Margin Principal Component Analysis · AAAI 2015
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
principal component analysis
0.212015
Bayesian Maximum Margin Principal Component Analysis · AAAI 2015
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.212015
Bayesian Maximum Margin Principal Component Analysis · AAAI 2015
Natural language and speech › Information extraction and text analysis
text classification
0.212013
Triplex transfer learning: exploiting both shared and distinct concepts for text classification · WSDM 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent structure discovery
latent semantic learning
0.112012
Multi-task Semi-supervised Semantic Feature Learning for Classification · ICDM 2012
Recommender systems › collaborative filtering
matrix factorization
0.112012
Multi-task Semi-supervised Semantic Feature Learning for Classification · ICDM 2012
Recommender systems › collaborative filtering › matrix factorization
nonnegative matrix tri-factorization
0.112012
Multi-task Semi-supervised Semantic Feature Learning for Classification · ICDM 2012
Recommender systems
collaborative filtering
0.112020
JIT2R: A Joint Framework for Item Tagging and Tag-based Recommendation · SIGIR 2020
Machine learning › Transfer learning and domain adaptation › domain alignment
cross-domain feature alignment
0.112018
Multi-view Adversarially Learned Inference for Cross-domain Joint Distribution Matching · KDD 2018
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model
0.112018
Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional Data · ACM Multimedia 2018
Machine learning › Generative modeling
variational autoencoder
0.112018
Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional Data · ACM Multimedia 2018
Recommender systems
cold-start recommendation
0.112017
A Location-Sentiment-Aware Recommender System for Both Home-Town and Out-of-Town Users · KDD 2017
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
manifold regularization
0.012013
Triplex transfer learning: exploiting both shared and distinct concepts for text classification · WSDM 2013

Methods — techniques the papers use, named apart from their topics

variational autoencoder · 1.0semi-supervised learning · 0.9bayesian posterior regularization · 0.7adversarial training · 0.7data augmentation · 0.5structured multi-output regression · 0.4multi-task learning · 0.4convolutional neural network · 0.4bootstrapping · 0.4variational inference · 0.3time-series feature extraction · 0.3semantic analysis · 0.3gaussian mixture model · 0.3ensemble learning · 0.3adversarial learning · 0.3probabilistic generative model · 0.3latent variable model · 0.3alternating iterative optimization · 0.1
YearPublicationVenuePosition
2022 Structured Neural Decoding With Multitask Transfer Learning of Deep Neural Network Representations
abstract
The reconstruction of visual information from human brain activity is a very important research topic in brain decoding. Existing methods ignore the structural information underlying the brain activities and the visual features, which severely limits their performance and interpretability. Here, we propose a hierarchically structured neural decoding framework by using multitask transfer learning of deep neural network (DNN) representations and a matrix-variate Gaussian prior. Our framework consists of two stages, Voxel2Unit and Unit2Pixel. In Voxel2Unit, we decode the functional magnetic resonance imaging (fMRI) data to the intermediate features of a pretrained convolutional neural network (CNN). In Unit2Pixel, we further invert the predicted CNN features back to the visual images. Matrix-variate Gaussian prior allows us to take into account the structures between feature dimensions and between regression tasks, which are useful for improving decoding effectiveness and interpretability. This is in contrast with the existing single-output regression models that usually ignore these structures. We conduct extensive experiments on two real-world fMRI data sets, and the results show that our method can predict CNN features more accurately and reconstruct the perceived natural images and faces with higher quality.
Changde Du, Changying Du, Haibao Wang, Huiguang He
IEEE Trans. Neural Networks Learn. Syst.2
2020 Conditional Generative Neural Decoding with Structured CNN Feature Prediction
abstract
Decoding visual contents from human brain activity is a challenging task with great scientific value. Two main facts that hinder existing methods from producing satisfactory results are 1) typically small paired training data; 2) under-exploitation of the structural information underlying the data. In this paper, we present a novel conditional deep generative neural decoding approach with structured intermediate feature prediction. Specifically, our approach first decodes the brain activity to the multilayer intermediate features of a pretrained convolutional neural network (CNN) with a structured multi-output regression (SMR) model, and then inverts the decoded CNN features to the visual images with an introspective conditional generation (ICG) model. The proposed SMR model can simultaneously leverage the covariance structures underlying the brain activities, the CNN features and the prediction tasks to improve the decoding accuracy and interpretability. Further, our ICG model can 1) leverage abundant unpaired images to augment the training data; 2) self-evaluate the quality of its conditionally generated images; and 3) adversarially improve itself without extra discriminator. Experimental results show that our approach yields state-of-the-art visual reconstructions from brain activities.
Changde Du, Changying Du, Huiguang He
AAAI2
2020 JIT2R: A Joint Framework for Item Tagging and Tag-based Recommendation
abstract
Predicting tags for a given item and leveraging tags to assist item recommendation are two popular research topics in the field of recommender system. Previous studies mostly focus only one of them to make contributions. However, we believe that these tasks are inherently correlated with each other: tags can provide additional information to profile items for more accurate recommendation; user behaviors can help to infer item relationships to benefit the item tagging process. In order to take the advantages of such mutually influential signals, we propose to integrate item tagging and tag-based recommendation into a unified model. We firstly design a basic framework, where the user-item interaction signals are leveraged to supervise the item tagging process. Then we extend the basic model with a bootstrapping technique to circulate such mutual improvements between different tasks. We conduct extensive experiments based on real-word datasets to demonstrate our model's superiorities.
Xu Chen 0017, Changying Du, Xiuqiang He 0001, Jun Wang 0012
SIGIR2
2020 Online Bayesian max-margin subspace learning for multi-view classification and regression
Jia He 0001, Changying Du, Fuzhen Zhuang, Qing He 0003, Guoping Long
Mach. Learn.2
2019 Doubly Semi-Supervised Multimodal Adversarial Learning for Classification, Generation and Retrieval
abstract
Learning over incomplete multi-modality data is a challenging problem with strong practical applications. Most existing multi-modal data imputation approaches have two limitations: (1) they are unable to accurately control the semantics of imputed modalities; and (2) without a shared low-dimensional latent space, they do not scale well with multiple modalities. To overcome the limitations, we propose a novel doubly semi-supervised multi-modal learning framework (DSML) with a modality-shared latent space and modality-specific generators, encoders and classifiers. We design novel softmax-based discriminators to train all modules adversarially. As a unified framework, DSML can be applied in multi-modal semi-supervised classification, missing modality imputation and fast cross-modality retrieval tasks simultaneously. Experiments on multiple datasets demonstrate its advantages.
Changde Du, Changying Du, Huiguang He
ICME2
2019 Reconstructing Perceived Images From Human Brain Activities With Bayesian Deep Multiview Learning
abstract
Neural decoding, which aims to predict external visual stimuli information from evoked brain activities, plays an important role in understanding human visual system. Many existing methods are based on linear models, and most of them only focus on either the brain activity pattern classification or visual stimuli identification. Accurate reconstruction of the perceived images from the measured human brain activities still remains challenging. In this paper, we propose a novel deep generative multiview model for the accurate visual image reconstruction from the human brain activities measured by functional magnetic resonance imaging (fMRI). Specifically, we model the statistical relationships between the two views (i.e., the visual stimuli and the evoked fMRI) by using two view-specific generators with a shared latent space. On the one hand, we adopt a deep neural network architecture for visual image generation, which mimics the stages of human visual processing. On the other hand, we design a sparse Bayesian linear model for fMRI activity generation, which can effectively capture voxel correlations, suppress data noise, and avoid overfitting. Furthermore, we devise an efficient mean-field variational inference method to train the proposed model. The proposed method can accurately reconstruct visual images via Bayesian inference. In particular, we exploit a posterior regularization technique in the Bayesian inference to regularize the model posterior. The quantitative and qualitative evaluations conducted on multiple fMRI data sets demonstrate the proposed method can reconstruct visual images more accurately than the state of the art.
Changde Du, Changying Du, Huiguang He
IEEE Trans. Neural Networks Learn. Syst.2
2018 Redundancy-resistant Generative Hashing for Image Retrieval
abstract
By optimizing probability distributions over discrete latent codes, Stochastic Generative Hashing (SGH) bypasses the critical and intractable binary constraints on hash codes. While encouraging results were reported, SGH still suffers from the deficient usage of latent codes, i.e., there often exist many uninformative latent dimensions in the code space, a disadvantage inherited from its auto-encoding variational framework. Motivated by the fact that code redundancy usually is severer when more complex decoder network is used, in this paper, we propose a constrained deep generative architecture to simplify the decoder for data reconstruction. Specifically, our new framework forces the latent hashing codes to not only reconstruct data through the generative network but also retain minimal squared L2 difference to the last real-valued network hidden layer. Furthermore, during posterior inference, we propose to regularize the standard auto-encoding objective with an additional term that explicitly accounts for the negative redundancy degree of latent code dimensions. We interpret such modifications as Bayesian posterior regularization and design an adversarial strategy to optimize the generative, the variational, and the redundancy-resistanting parameters. Empirical results show that our new method can significantly boost the quality of learned codes and achieve state-of-the-art performance for image retrieval.
Changying Du, Xingyu Xie, Changde Du, Hao Wang 0005
IJCAI1
2018 Multi-view Adversarially Learned Inference for Cross-domain Joint Distribution Matching
abstract
Many important data mining problems can be modeled as learning a (bidirectional) multidimensional mapping between two data domains. Based on the generative adversarial networks (GANs), particularly conditional ones, cross-domain joint distribution matching is an increasingly popular kind of methods addressing such problems. Though significant advances have been achieved, there are still two main disadvantages of existing models, i.e., the requirement of large amount of paired training samples and the notorious instability of training. In this paper, we propose a multi-view adversarially learned inference (ALI) model, termed as MALI, to address these issues. Unlike the common practice of learning direct domain mappings, our model relies on shared latent representations of both domains and can generate arbitrary number of paired faking samples, benefiting from which usually very few paired samples (together with sufficient unpaired ones) is enough for learning good mappings. Extending the vanilla ALI model, we design novel discriminators to judge the quality of generated samples (both paired and unpaired), and provide theoretical analysis of our new formulation. Experiments on image-to-image translation, image-to-attribute generation (multi-label classification), attribute-to-image generation tasks demonstrate that our semi-supervised learning framework yields significant performance improvements over existing ones. Results on cross-modality retrieval show that our latent space based method can achieve competitive similarity search performance in relative fast speed, compared to those methods that compute similarities in the high-dimensional data space.
Changying Du, Changde Du, Xingyu Xie, Chen Zhang 0003, Hao Wang 0005
KDD1
2018 StockAssIstant: A Stock AI Assistant for Reliability Modeling of Stock Comments
abstract
Stock comments from analysts contain important consulting information for investors to foresee stock volatility and market trends. Existing studies on stock comments usually focused on capturing coarse-grained opinion polarities or understanding market fundamentals. However, investors are often overwhelmed and confused by massive comments with huge noises and ambiguous opinions. Therefore, it is an emerging need to have a fine-grained stock comment analysis tool to identify more reliable stock comments. To this end, this paper provides a solution called StockAssIstant for modeling the reliability of stock comments by considering multiple factors, such as stock price trends, comment content, and the performances of analysts, in a holistic manner. Specifically, we first analyze the pattern of analysts' opinion dynamics from historical comments. Then, we extract key features from the time-series constructed by using the semantic information in comment text, stock prices and the historical behaviors of analysts. Based on these features, we propose an ensemble learning based approach for measuring the reliability of comments. Finally, we conduct extensive experiments and provide a trading simulation on real-world stock data. The experimental results and the profit achieved by the simulated trading in 12-month period clearly validate the effectiveness of our approach for modeling the reliability of stock comments.
Chen Zhang 0003, Changying Du, Hongzhi Yin, Hao Wang 0005
KDD4
2018 Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional Data
abstract
There are threefold challenges in emotion recognition. First, it is difficult to recognize human's emotional states only considering a single modality. Second, it is expensive to manually annotate the emotional data. Third, emotional data often suffers from missing modalities due to unforeseeable sensor malfunction or configuration issues. In this paper, we address all these problems under a novel multi-view deep generative framework. Specifically, we propose to model the statistical relationships of multi-modality emotional data using multiple modality-specific generative networks with a shared latent space. By imposing a Gaussian mixture assumption on the posterior approximation of the shared latent variables, our framework can learn the joint deep representation from multiple modalities and evaluate the importance of each modality simultaneously. To solve the labeled-data-scarcity problem, we extend our multi-view model to semi-supervised learning scenario by casting the semi-supervised classification problem as a specialized missing data imputation task. To address the missing-modality problem, we further extend our semi-supervised multi-view model to deal with incomplete data, where a missing view is treated as a latent variable and integrated out during inference. This way, the proposed overall framework can utilize all available (both labeled and unlabeled, as well as both complete and incomplete) data to improve its generalization ability. The experiments conducted on two real multi-modal emotion datasets demonstrated the superiority of our framework.
Changde Du, Changying Du, Hao Wang 0005, Jinpeng Li 0002, Wei-Long Zheng, Bao-Liang Lu, Huiguang He
ACM Multimedia2
2017 Nonlinear Maximum Margin Multi-View Learning with Adaptive Kernel
abstract
Existing multi-view learning methods based on kernel function either require the user to select and tune a single predefined kernel or have to compute and store many Gram matrices to perform multiple kernel learning. Apart from the huge consumption of manpower, computation and memory resources, most of these models seek point estimation of their parameters, and are prone to overfitting to small training data. This paper presents an adaptive kernel nonlinear max-margin multi-view learning model under the Bayesian framework. Specifically, we regularize the posterior of an efficient multi-view latent variable model by explicitly mapping the latent representations extracted from multiple data views to a random Fourier feature space where max-margin classification constraints are imposed. Assuming these random features are drawn from Dirichlet process Gaussian mixtures, we can adaptively learn shift-invariant kernels from data according to Bochners theorem. For inference, we employ the data augmentation idea for hinge loss, and design an efficient gradient-based MCMC sampler in the augmented space. Having no need to compute the Gram matrix, our algorithm scales linearly with the size of training set. Extensive experiments on real-world datasets demonstrate that our method has superior performance.
Jia He 0001, Changying Du, Changde Du, Fuzhen Zhuang, Qing He 0003, Guoping Long
IJCAI2
2017 Sharing deep generative representation for perceived image reconstruction from human brain activity
abstract
Decoding human brain activities via functional magnetic resonance imaging (fMRI) has gained increasing attention in recent years. While encouraging results have been reported in brain states classification tasks, reconstructing the details of human visual experience still remains difficult. Two main challenges that hinder the development of effective models are the perplexing fMRI measurement noise and the high dimensionality of limited data instances. Existing methods generally suffer from one or both of these issues and yield dissatisfactory results. In this paper, we tackle this problem by casting the reconstruction of visual stimulus as the Bayesian inference of missing view in a multiview latent variable model. Sharing a common latent representation, our joint generative model of external stimulus and brain response is not only “deep” in extracting nonlinear features from visual images, but also powerful in capturing correlations among voxel activities of fMRI recordings. The nonlinearity and deep structure endow our model with strong representation ability, while the correlations of voxel activities are critical for suppressing noise and improving prediction. We devise an efficient variational Bayesian method to infer the latent variables and the model parameters. To further improve the reconstruction accuracy, the latent representations of testing instances are enforced to be close to that of their neighbours from the training set via posterior regularization. Experiments on three fMRI recording datasets demonstrate that our approach can more accurately reconstruct visual stimuli.
Changde Du, Changying Du, Huiguang He
IJCNN2
2017 A Location-Sentiment-Aware Recommender System for Both Home-Town and Out-of-Town Users
abstract
Spatial item recommendation has become an important means to help people discover interesting locations, especially when people pay a visit to unfamiliar regions. Some current researches are focusing on modelling individual and collective geographical preferences for spatial item recommendation based on users' check-in records, but they fail to explore the phenomenon of user interest drift across geographical regions, i.e., users would show different interests when they travel to different regions. Besides, they ignore the influence of public comments for subsequent users' check-in behaviors. Specifically, it is intuitive that users would refuse to check in to a spatial item whose historical reviews seem negative overall, even though it might fit their interests. Therefore, it is necessary to recommend the right item to the right user at the right location. In this paper, we propose a latent probabilistic generative model called LSARS to mimic the decision-making process of users' check-in activities both in home-town and out-of-town scenarios by adapting to user interest drift and crowd sentiments, which can learn location-aware and sentiment-aware individual interests from the contents of spatial items and user reviews. Due to the sparsity of user activities in out-of-town regions, LSARS is further designed to incorporate the public preferences learned from local users' check-in behaviors. Finally, we deploy LSARS into two practical application scenes: spatial item recommendation and target user discovery. Extensive experiments on two large-scale location-based social networks (LBSNs) datasets show that LSARS achieves better performance than existing state-of-the-art methods.
Hao Wang 0005, Yanmei Fu, Qinyong Wang, Hongzhi Yin, Changying Du, Hui Xiong 0001
KDD5
2016 Online Bayesian Max-Margin Subspace Multi-View Learning
Jia He 0001, Changying Du, Fuzhen Zhuang, Qing He 0003, Guoping Long
IJCAI2
2016 Online variational Bayesian Support Vector Regression
abstract
Traditional Support Vector Regression (SVR) solvers require user pre-specified penalty (regularization) parameter as input and typically model the training data with maximum a posterior (MAP) principle. The resultant point estimates can be affected seriously by inappropriate regularization, outliers and noise, especially when training online. In this paper, we address the aforementioned problems by developing a Bayesian SVR model with the pseudo-likelihood and data augmentation idea. Then we perform variational posterior inference in an augmented variable space and the approximate posterior of model weights, rather than point estimates as in traditional SVR, are used to make robust predictions. Besides, once the approximate posterior is obtained from a given set of data, we can regard it as model prior when dealing with new arrival data, which leads to a natural way to extend our batch model to the online scenario. Experiments on several benchmark regression problems as well as a real vehicle accident rate prediction task show that our models have superior performance while inferring penalty parameter automatically.
Siqi Deng, Kan Gao, Changying Du, Wenjing Ma, Guoping Long, Yucheng Li 0002
IJCNN3
2016 Bayesian Group Feature Selection for Support Vector Learning Machines
Changde Du, Changying Du, Shandian Zhe, A-Li Luo, Qing He 0003, Guoping Long
PAKDD (1)2
2016 Efficient Bayesian Maximum Margin Multiple Kernel Learning
Changying Du, Changde Du, Guoping Long, Xin Jin 0004, Yucheng Li 0002
ECML/PKDD (1)1
2016 Learning Beyond Predefined Label Space via Bayesian Nonparametric Topic Modelling
Changying Du, Fuzhen Zhuang, Jia He 0001, Qing He 0003, Guoping Long
ECML/PKDD (1)1
2016 Online Bayesian Multiple Kernel Bipartite Ranking
Changying Du, Changde Du, Guoping Long, Qing He 0003, Yucheng Li 0002
UAI1
2015 Bayesian Maximum Margin Principal Component Analysis
abstract
Supervised dimensionality reduction has shown great advantages in finding predictive subspaces. Previous methods rarely consider the popular maximum margin principle and are prone to overfitting to usually small training data, especially for those under the maximum likelihood framework. In this paper, we present a posterior-regularized Bayesian approach to combine Principal Component Analysis (PCA) with the max-margin learning. Based on the data augmentation idea for max-margin learning and the probabilistic interpretation of PCA, our method can automatically infer the weight and penalty parameter of max-margin learning machine, while finding the most appropriate PCA subspace simultaneously under the Bayesian framework. We develop a fast mean-field variational inference algorithm to approximate the posterior. Experimental results on various classification tasks show that our method outperforms a number of competitors.
Changying Du, Shandian Zhe, Fuzhen Zhuang, Yuan Qi 0001, Qing He 0003, Zhongzhi Shi
AAAI1
2015 Heterogeneous Multi-task Semantic Feature Learning for Classification
abstract
Multi-task Learning (MTL) aims to learn multiple related tasks simultaneously instead of separately to improve generalization performance of each task. Most existing MTL methods assumed that the multiple tasks to be learned have the same feature representation. However, this assumption may not hold for many real-world applications. In this paper, we study the problem of MTL with heterogeneous features for each task. To address this problem, we first construct an integrated graph of a set of bipartite graphs to build a connection among different tasks. We then propose a multi-task nonnegative matrix factorization (MTNMF) method to learn a common semantic feature space underlying different heterogeneous feature spaces of each task. Finally, based on the common semantic features and original heterogeneous features, we model the heterogenous MTL problem as a multi-task multi-view learning (MTMVL) problem. In this way, a number of existing MTMVL methods can be applied to solve the problem effectively. Extensive experiments on three real-world problems demonstrate the effectiveness of our proposed method.
Xin Jin 0004, Fuzhen Zhuang, Sinno Jialin Pan, Changying Du, Ping Luo 0001, Qing He 0003
CIKM4
2014 Multi-task Multi-view Learning for Heterogeneous Tasks
abstract
Multi-task multi-view learning deals with the learning scenarios where multiple tasks are associated with each other through multiple shared feature views. All previous works for this problem assume that the tasks use the same set of class labels. However, in real world there exist quite a few applications where the tasks with several views correspond to different set of class labels. This new learning scenario is called Multi-task Multi-view Learning for Heterogeneous Tasks in this study. Then, we propose a Multi-tAsk MUlti-view Discriminant Analysis (MAMUDA) method to solve this problem. Specifically, this method collaboratively learns the feature transformations for different views in different tasks by exploring the shared task-specific and problem intrinsic structures. Additionally, MAMUDA method is convenient to solve the multi-class classification problems. Finally, the experiments on two real-world problems demonstrate the effectiveness of MAMUDA for heterogeneous tasks.
Xin Jin 0004, Fuzhen Zhuang, Hui Xiong 0001, Changying Du, Ping Luo 0001, Qing He 0003
CIKM4
2014 Nonparametric Bayesian Multi-Task Large-margin Classification
abstract
In this paper, we present a nonparametric Bayesian multi-task large-margin classification model which can cluster tasks into the most appropriate number of groups and induce flexible model sharing within each task group simultaneously. Specifically, we first show a very simple method to integrate large margin learning with hierarchical Bayesian models by employing an important variant of the standard SVMi.e.proximal SVM (PSVM)whose loss function is used to define a novel likelihood function. And then we assume that the model parameter of each task consists of two parts: one is shared within each task group (group-level parameter) while the other is specific to each distinct task (task rescaling parameter). A Dirichlet process prior is imposed on the group-level parameter while the task rescaling parameter is assigned a one-mean Laplace prior. Finally the parameter of a task is the corresponding group parameter times its specific rescaling parameter. We give efficient Markov chain Monte Calo (MCMC) algorithm to conduct model inference. Experiments on the Landmine detection data and the UCI Yeast data demonstrate the effectiveness of our method.
Changying Du, Jia He 0001, Fuzhen Zhuang, Yuan Qi 0001, Qing He 0003
ECAI1
2014 Clustering in extreme learning machine feature space
Qing He 0003, Xin Jin 0004, Changying Du, Fuzhen Zhuang, Zhongzhi Shi
Neurocomputing3
2014 Triplex Transfer Learning: Exploiting Both Shared and Distinct Concepts for Text Classification
abstract
Transfer learning focuses on the learning scenarios when the test data from target domains and the training data from source domains are drawn from similar but different data distributions with respect to the raw features. Along this line, some recent studies revealed that the high-level concepts, such as word clusters, could help model the differences of data distributions, and thus are more appropriate for classification. In other words, these methods assume that all the data domains have the same set of shared concepts, which are used as the bridge for knowledge transfer. However, in addition to these shared concepts, each domain may have its own distinct concepts. In light of this, we systemically analyze the high-level concepts, and propose a general transfer learning framework based on nonnegative matrix trifactorization, which allows to explore both shared and distinct concepts among all the domains simultaneously. Since this model provides more flexibility in fitting the data, it can lead to better classification accuracy. Moreover, we propose to regularize the manifold structure in the target domains to improve the prediction performances. To solve the proposed optimization problem, we also develop an iterative algorithm and theoretically analyze its convergence properties. Finally, extensive experiments show that the proposed model can outperform the baseline methods with a significant margin. In particular, we show that our method works much better for the more challenging tasks when there are distinct concepts in the data.
Fuzhen Zhuang, Ping Luo 0001, Changying Du, Qing He 0003, Zhongzhi Shi, Hui Xiong 0001
IEEE Trans. Cybern.3
2013 Triplex transfer learning: exploiting both shared and distinct concepts for text classification
abstract
Transfer learning focuses on the learning scenarios when the test data from target domains and the training data from source domains are drawn from similar but different data distributions with respect to the raw features. Along this line, some recent studies revealed that the high-level concepts, such as word clusters, could help model the differences of data distributions, and thus are more appropriate for classification. In other words, these methods assume that all the data domains have the same set of shared concepts, which are used as the bridge for knowledge transfer. However, in addition to these shared concepts, each domain may have its own distinct concepts. In light of this, we systemically analyze the high-level concepts, and propose a general transfer learning framework based on nonnegative matrix trifactorization, which allows to explore both shared and distinct concepts among all the domains simultaneously. Since this model provides more flexibility in fitting the data, it can lead to better classification accuracy. Moreover, we propose to regularize the manifold structure in the target domains to improve the prediction performances. To solve the proposed optimization problem, we also develop an iterative algorithm and theoretically analyze its convergence properties. Finally, extensive experiments show that the proposed model can outperform the baseline methods with a significant margin. In particular, we show that our method works much better for the more challenging tasks when there are distinct concepts in the data.
Fuzhen Zhuang, Ping Luo 0001, Changying Du, Qing He 0003, Zhongzhi Shi
WSDM3
2012 Multi-task Semi-supervised Semantic Feature Learning for Classification
abstract
Multi-task learning has proven to be useful to boost the learning of multiple related but different tasks. Meanwhile, latent semantic models such as LSA and LDA are popular and effective methods to extract discriminative semantic features of high dimensional dyadic data. In this paper, we present a method to combine these two techniques together by introducing a new matrix tri-factorization based formulation for semi-supervised latent semantic learning, which can incorporate labeled information into traditional unsupervised learning of latent semantics. Our inspiration for multi-task semantic feature learning comes from two facts, i.e., 1) multiple tasks generally share a set of common latent semantics, and 2) a semantic usually has a stable indication of categories no matter which task it is from. Thus to make multiple tasks learn from each other we wish to share the associations between categories and those common semantics among tasks. Along this line, we propose a novel joint Nonnegative matrix tri-factorization framework with the aforesaid associations shared among tasks in the form of a semantic-category relation matrix. Our new formulation for multi-task learning can simultaneously learn (1) discriminative semantic features of each task, (2) predictive structure and categories of unlabeled data in each task, (3) common semantics shared among tasks and specific semantics exclusive to each task. We give alternating iterative algorithm to optimize our objective and theoretically show its convergence. Finally extensive experiments on text data along with the comparison with various baselines and three state-of-the-art multi-task learning algorithms demonstrate the effectiveness of our method.
Changying Du, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi
ICDM1
2011 A parallel incremental extreme SVM classifier
Qing He 0003, Changying Du, Fuzhen Zhuang, Zhongzhi Shi
Neurocomputing2