Pengfei Wei 0001

dblp:29/11273-1 · DBLP profile ↗
← Back
42ranked-venue papers
18as first author
26since 2021 · last 2024
0000-0001-8093-0803ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 15 first-author · 19 since 2021Databases, data management, data science and information retrieval · 13 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 9 since 2021
YearPublicationVenuePosition
2024 Prior and Prediction Inverse Kernel Transformer for Single Image Defocus Deblurring
abstract
Defocus blur, due to spatially-varying sizes and shapes, is hard to remove. Existing methods either are unable to effectively handle irregular defocus blur or fail to generalize well on other datasets. In this work, we propose a divide-and-conquer approach to tackling this issue, which gives rise to a novel end-to-end deep learning method, called prior-and-prediction inverse kernel transformer (P2IKT), for single image defocus deblurring. Since most defocus blur can be approximated as Gaussian blur or its variants, we construct an inverse Gaussian kernel module in our method to enhance its generalization ability. At the same time, an inverse kernel prediction module is introduced in order to flexibly address the irregular blur that cannot be approximated by Gaussian blur. We further design a scale recurrent transformer, which estimates mixing coefficients for adaptively combining the results from the two modules and runs the scale recurrent ``coarse-to-fine" procedure for progressive defocus deblurring. Extensive experimental results demonstrate that our P2IKT outperforms previous methods in terms of PSNR on multiple defocus deblurring datasets.
Peng Tang 0004, Zhiqiang Xu 0003, Chunlai Zhou, Pengfei Wei 0001, Peng Han 0005, Xin Cao 0001, Tobias Lasser
AAAI4
2024 Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
abstract
Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspects: 1) previous works of zero-shot TTS are typically trained with single-sentence prompts, which significantly restricts their performance when the data is relatively sufficient during the inference stage. 2) The prosodic information in prompts is highly coupled with timbre, making it untransferable to each other. This paper introduces Mega-TTS 2, a generic prompting mechanism for zero-shot TTS, to tackle the aforementioned challenges. Specifically, we design a powerful acoustic autoencoder that separately encodes the prosody and timbre information into the compressed latent space while providing high-quality reconstructions. Then, we propose a multi-reference timbre encoder and a prosody latent language model (P-LLM) to extract useful information from multi-sentence prompts. We further leverage the probabilities derived from multiple P-LLM outputs to produce transferable and controllable prosody. Experimental results demonstrate that Mega-TTS 2 could not only synthesize identity-preserving speech with a short prompt of an unseen speaker from arbitrary sources but consistently outperform the fine-tuning method when the volume of data ranges from 10 seconds to 5 minutes. Furthermore, our method enables to transfer various speaking styles to the target timbre in a fine-grained and controlled manner. Audio samples can be found in https://boostprompt.github.io/boostprompt/.
Ziyue Jiang 0001, Jinglin Liu, Yi Ren 0006, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang 0006, Chen Zhang 0020, Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001
ICLR9
2024 Graph Domain Adaptation: A Generative View
abstract
Recent years have witnessed tremendous interest in deep learning on graph-structured data. Due to the high cost of collecting labeled graph-structured data, domain adaptation is important to supervised graph learning tasks with limited samples. However, current graph domain adaptation methods are generally adopted from traditional domain adaptation tasks, and the properties of graph-structured data are not well utilized. For example, the observed social networks on different platforms are controlled not only by the different crowds or communities but also by domain-specific policies and background noise. Based on these properties in graph-structured data, we first assume that the graph-structured data generation process is controlled by three independent types of latent variables, i.e., the semantic latent variables, the domain latent variables, and the random latent variables. Based on this assumption, we propose a disentanglement-based unsupervised domain adaptation method for the graph-structured data, which applies variational graph auto-encoders to recover these latent variables and disentangles them via three supervised learning modules. Extensive experimental results on two real-world datasets in the graph classification task reveal that our method not only significantly outperforms the traditional domain adaptation methods and the disentangled-based domain adaptation methods but also outperforms the state-of-the-art graph domain adaptation algorithms. The code is available at https://github.com/rynewu224/GraphDA .
Ruichu Cai, Fengzhu Wu, Zijian Li 0001, Pengfei Wei 0001, Lingling Yi, Kun Zhang 0001
ACM Trans. Knowl. Discov. Data4
2023 Adaptive Policy Learning for Offline-to-Online Reinforcement Learning
abstract
Conventional reinforcement learning (RL) needs an environment to collect fresh data, which is impractical when online interactions are costly. Offline RL provides an alternative solution by directly learning from the previously collected dataset. However, it will yield unsatisfactory performance if the quality of the offline datasets is poor. In this paper, we consider an offline-to-online setting where the agent is first learned from the offline dataset and then trained online, and propose a framework called Adaptive Policy Learning for effectively taking advantage of offline and online data. Specifically, we explicitly consider the difference between the online and offline data and apply an adaptive update scheme accordingly, that is, a pessimistic update strategy for the offline dataset and an optimistic/greedy update scheme for the online dataset. Such a simple and effective method provides a way to mix the offline and online RL and achieve the best of both worlds. We further provide two detailed algorithms for implementing the framework through embedding value or policy-based RL algorithms into it. Finally, we conduct extensive experiments on popular continuous control tasks, and results show that our algorithm can learn the expert policy with high sample efficiency even when the quality of offline dataset is poor, e.g., random dataset.
Xufang Luo, Pengfei Wei 0001, Xuan Song 0001, Dongsheng Li 0002, Jing Jiang 0002
AAAI3
2023 Virtual Try-On with Pose-Garment Keypoints Guided Inpainting
abstract
Virtual try-on is an important technology supporting on-line apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existing virtual try-on methods usually present distortions in garment shape and lose pattern details. In this paper, we propose a pose-garment keypoints guided inpainting method for the image-based virtual try-on task, which produces high-fidelity try-on images and well preserves the shapes and patterns of the garments. In our method, human pose and garment keypoints are extracted from source images and constructed as graphs to predict the garment keypoints at the target pose. After which, the predicted key-points are used as guide information to predict the target segmentation map and warp the garment image. The try-on image is finally generated with a semantic-conditioned inpainting scheme using the segmentation map and recomposed person image as conditions. To verify the effectiveness of our proposed method, we conduct extensive experiments on the VITON-HD dataset under both paired and unpaired experimental settings. The qualitative and quantitative results show that our method significantly outperforms prior methods at different image resolutions. The codes repository link is https://github.com/lizhi-ntu/KGI.
Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Alex Chichung Kot
ICCV2
2023 Dynamic Job Shop Scheduling via Deep Reinforcement Learning
abstract
Recently, deep reinforcement learning (DRL) is shown to be promising in learning dispatching rules end-to-end for complex scheduling problems. However, most research is limited to deterministic problems. In this paper, we focus on the dynamic job-shop scheduling problem (DJSP), which is a complex dynamic optimization problem under uncertainty. We propose a DRL based method to learn dispatching policies for DJSP. Unlike existing DRL based dynamic scheduling methods that use a fixed number of dispatching rules as actions, our decision-making framework directly selects legitimate jobs, which is able to break the limitations imposed by priority dispatching rules. We design two training methods, including a gradient based algorithm with dense rewards, and an evolutionary strategy with sparse rewards. Extensive experiments show that our DRL method can learn high-quality DJSP dispatching policies, and can significantly outperform a state-of-the-art Genetic Programming (GP) based dispatching rule learning method.
Xinjie Liang, Wen Song 0004, Pengfei Wei 0001
ICTAI3
2023 AudioQR: Deep Neural Audio Watermarks For QR Code
abstract
Image-based quick response (QR) code is frequently used, but creates barriers for the visual impaired people. With the goal of ``AI for good", this paper proposes the AudioQR, a barrier-free QR coding mechanism for the visually impaired population via deep neural audio watermarks. Previous audio watermarking approaches are mainly based on handcrafted pipelines, which is less secure and difficult to apply in large-scale scenarios. In contrast, AudioQR is the first comprehensive end-to-end pipeline that hides watermarks in audio imperceptibly and robustly. To achieve this, we jointly train an encoder and decoder, where the encoder is structured as a concatenation of transposed convolutions and multi-receptive field fusion modules. Moreover, we customize the decoder training with a stochastic data augmentation chain to make the watermarked audio robust towards different audio distortions, such as environment background, room impulse response when playing through the air, music surrounding, and Gaussian noise. Experiment results indicate that AudioQR can efficiently hide arbitrary information into audio without introducing significant perceptible difference. Our code is available at https://github.com/xinghua-qu/AudioQR.
Xinghua Qu, Xiang Yin 0006, Pengfei Wei 0001, Lu Lu 0015, Zejun Ma 0001
IJCAI3
2023 S2CD: Self-heuristic Speaker Content Disentanglement for Any-to-Any Voice Conversion
Pengfei Wei 0001, Xiang Yin 0006, Xinghua Qu, Zhiqiang Xu 0003, Zejun Ma 0001
INTERSPEECH1
2023 Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement Perspective
abstract
Unsupervised video domain adaptation is a practical yet challenging task. In this work, for the first time, we tackle it from a disentanglement view. Our key idea is to handle the spatial and temporal domain divergence separately through disentanglement. Specifically, we consider the generation of cross-domain videos from two sets of latent factors, one encoding the static information and another encoding the dynamic information. A Transfer Sequential VAE (TranSVAE) framework is then developed to model such generation. To better serve for adaptation, we propose several objectives to constrain the latent factors. With these constraints, the spatial divergence can be readily removed by disentangling the static domain-specific information out, and the temporal divergence is further reduced from both frame- and video-levels through adversarial learning. Extensive experiments on the UCF-HMDB, Jester, and Epic-Kitchens datasets verify the effectiveness and superiority of TranSVAE compared with several state-of-the-art approaches.
Pengfei Wei 0001, Lingdong Kong, Xinghua Qu, Yi Ren 0006, Zhiqiang Xu 0003, Jing Jiang 0002, Xiang Yin 0006
NeurIPS1
2023 Adaptive Transfer Kernel Learning for Transfer Gaussian Process Regression
abstract
Transfer regression is a practical and challenging problem with important applications in various domains, such as engineering design and localization. Capturing the relatedness of different domains is the key of adaptive knowledge transfer. In this paper, we investigate an effective way of explicitly modelling domain relatedness through transfer kernel, a transfer-specified kernel that considers domain information in the covariance calculation. Specifically, we first give the formal definition of transfer kernel, and introduce three basic general forms that well cover existing related works. To cope with the limitations of the basic forms in handling complex real-world data, we further propose two advanced forms. Corresponding instantiations of the two forms are developed, namely${Trk}_{\alpha \beta }$and${Trk}_{\omega }$based on multiple kernel learning and neural networks, respectively. For each instantiation, we present a condition with which the positive semi-definiteness is guaranteed and a semantic meaning is interpreted to the learned domain relatedness. Moreover, the condition can be easily used in the learning ofTrGP$_{\alpha \beta }$andTrGP$_{\omega }$that are the Gaussian process models with the transfer kernels${Trk}_{\alpha \beta }$and${Trk}_{\omega }$respectively. Extensive empirical studies show the effectiveness ofTrGP$_{\alpha \beta }$andTrGP$_{\omega }$on domain relatedness modelling and transfer adaptiveness.
Pengfei Wei 0001, Yiping Ke, Yew-Soon Ong, Zejun Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Transfer Kernel Learning for Multi-Source Transfer Gaussian Process Regression
abstract
Multi-source transfer regression is a practical and challenging problem where capturing the diverse relatedness of different domains is the key of adaptive knowledge transfer. In this paper, we propose an effective way of explicitly modeling the domain relatedness of each domain pair through transfer kernel learning. Specifically, we first discuss the advantages and disadvantages of existing transfer kernels in handling the multi-source transfer regression problem. To cope with the limitations of the existing transfer kernels, we further propose a novel multi-source transfer kernel$k_{ms}$. The proposed$k_{ms}$assigns a learnable parametric coefficient to model the relatedness of each inter-domain pair, and simultaneously regulates the relatedness of the intra-domain pair to be 1. Moreover, to capture the heterogeneous data characteristics of multiple domains,$k_{ms}$exploits different standard kernels for different domain pairs. We further provide a theorem that not only guarantees the positive semi-definiteness of$k_{ms}$but also conveys a semantic interpretation to the learned domain relatedness. Moreover, the theorem can be easily used in the learning of the corresponding transfer Gaussian process model with$k_{ms}$. Extensive empirical studies show the effectiveness of our proposed method on domain relatedness modelling and transfer performance.
Pengfei Wei 0001, Thanh Vinh Vo, Xinghua Qu, Yew-Soon Ong, Zejun Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Adaptive Multi-Source Causal Inference from Observational Data
abstract
We propose a new approach to estimate causal effects from observational data. We leverage multiple data sources which share similar causal mechanisms with the scarce target observations to help infer causal effects in the target domain. The data sources may be available in sequence or some unplanned order. Causal inference can be carried out without prior knowledge of the data discrepancy between the source and target observations. We introduce three levels of knowledge transfer through modelling the outcomes, treatments, and confounders to achieve consistent positive transfer. We incorporate parametric transfer factors to adaptively control the transfer strength, thus achieving a fair and balanced knowledge transfer between the sources and the target. We also empirically show the effectiveness of the proposed method as compared with recent baselines.
Thanh Vinh Vo, Pengfei Wei 0001, Trong Nghia Hoang, Tze-Yun Leong
CIKM2
2022 Dynamic Transfer Gaussian Process Regression
abstract
In this paper, we work on a challenging dynamic transfer regression problem where domains come in a streaming manner. At each time stage, a new domain emerges and is taken as the target domain while all the domains in previous time stages are taken as source domains. We propose a transfer Gaussian process model GPdk with a novel dynamic transfer kernel DyTK to handle the dynamic transfer regression problem. Specifically, DyTK is with a sequential form to fit the domain stream. To adaptively control the knowledge transfer strength, DyTK is designed to be capable of modeling the inter-domain relatedness of every inter-domain pair. A theorem that ensures DyTK to be positive semi-definite is then proposed. We also theoretically analyze the transfer performance of GPdk by deriving its generalization error bounds. The error bounds further motivate us to propose a parameter reuse strategy to alleviate the scalability issue of GPdk along time. Extensive experiments on both synthetic and real-world datasets show the effectiveness of GPdk in handling dynamic transfer regression problems.
Pengfei Wei 0001, Xinghua Qu, Wen Song 0004, Zejun Ma 0001
CIKM1
2022 Language Adaptive Cross-Lingual Speech Representation Learning with Sparse Sharing Sub-Networks
abstract
Unsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across multiple languages. However, standard XLSR model suffers from language interference problem due to the lack of language specific modeling ability. In this work, we investigate language adaptive training on XLSR models. More importantly, we propose a novel language adaptive pretraining approach based on sparse sharing sub-networks. It makes room for language specific modeling by pruning out unimportant parameters for each language, without requiring any manually designed language specific component. After pruning, each language only maintains a sparse sub-network, while the sub-networks are partially shared with each other. Experimental results on a downstream multilingual speech recognition task show that our proposed method significantly outperforms baseline XLSR models on both high resource and low resource languages. Besides, our proposed method consistently outperforms other adaptation methods and requires fewer parameters.
Yizhou Lu, Mingkun Huang, Xinghua Qu, Pengfei Wei 0001, Zejun Ma 0001
ICASSP4
2022 Importance Prioritized Policy Distillation
abstract
Policy distillation (PD) has been widely studied in deep reinforcement learning (RL), while existing PD approaches assume that the demonstration data (i.e., state-action pairs in frames) in a decision making sequence is uniformly distributed. This may bring in unwanted bias since RL is a reward maximizing process instead of simple label matching. Given such an issue, we denote the frame importance as its contribution to the expected reward on a particular frame, and hypothesize that adapting such frame importance could benefit the performance of the distilled student policy. To verify our hypothesis, we analyze why and how frame importance matters in RL settings. Based on the analysis, we propose an importance prioritized PD framework that highlights the training on important frames, so as to learn efficiently. Particularly, the frame importance is measured by the reciprocal of weighted Shannon entropy from a teacher policy's action prescriptions. Experiments on Atari games and policy compression tasks show that capturing the frame importance significantly boosts the performance of the distilled policies.
Xinghua Qu, Yew-Soon Ong, Abhishek Gupta 0001, Pengfei Wei 0001, Zhu Sun 0001, Zejun Ma 0001
KDD4
2022 Synthesising Audio Adversarial Examples for Automatic Speech Recognition
abstract
Adversarial examples in automatic speech recognition (ASR) are naturally sounded by humans yet capable of fooling well trained ASR models to transcribe incorrectly. Existing audio adversarial examples are typically constructed by adding constrained perturbations on benign audio inputs. Such attacks are therefore generated with an audio dependent assumption. For the first time, we propose the Speech Synthesising based Attack (SSA), a novel threat model that constructs audio adversarial examples entirely from scratch, i.e., without depending on any existing audio to fool cutting-edge ASR models. To this end, we introduce a conditional variational auto-encoder (CVAE) as the speech synthesiser. Meanwhile, an adaptive sign gradient descent algorithm is proposed to solve the adversarial audio synthesis task. Experiments on three datasets (i.e., Audio Mnist, Common Voice, and Librispeech) show that our method could synthesise naturally sounded audio adversarial examples to mislead the start-of-the-art ASR models. Our web-page containing generated audio demos is at https://sites.google.com/view/ssa-asr/home.
Xinghua Qu, Pengfei Wei 0001, Mingyong Gao, Zhu Sun 0001, Yew-Soon Ong, Zejun Ma 0001
KDD2
2022 Taxi demand forecasting based on the temporal multimodal information fusion graph neural network
Wenxiong Liao, Bi Zeng, Jianqi Liu, Pengfei Wei 0001, Xiaochun Cheng
Appl. Intell.4
2022 A novel concavity based method for automatic segmentation of touching cells in microfluidic chips
Qiqiang Li, Wen Song 0004, Pengfei Wei 0001, Jing Guo 0007
Expert Syst. Appl.4
2022 Semi-supervised multi-view graph convolutional networks with application to webpage classification
Fei Wu 0004, Xiaoyuan Jing, Pengfei Wei 0001, Chao Lan, Yimu Ji 0001, Guoping Jiang, Qinghua Huang
Inf. Sci.3
2022 Dual-aligned unsupervised domain adaptation with graph convolutional networks
Fei Wu 0004, Pengfei Wei 0001, Guangwei Gao, Changhui Hu 0001, Qi Ge, Xiaoyuan Jing
Multim. Tools Appl.2
2022 Modality and Event Adversarial Networks for Multi-Modal Fake News Detection
abstract
With the popularity of news on social media, fake news has become an important issue for the public and government. There exist some fake news detection methods that focus on information exploration and utilization from multiple modalities, e.g., text and image. However, how to effectively learn both modality-invariant and event-invariant discriminant features is still a challenge. In this paper, we propose a novel approach named Modality and Event Adversarial Networks (MEAN) for fake news detection. It contains two parts: a multi-modal generator and a dual discriminator. The multi-modal generator extracts latent discriminant feature representations of text and image modalities. A decoder is adopted to reduce information loss in the generation process for each modality. The dual discriminator includes a modality discriminator and an event discriminator. The discriminator learns to classify the event or the modality of features, and network training is guided by the adversarial scheme. Experiments on two widely used datasets show that MEAN can perform better than state-of-the-art related multi-modal fake news detection methods.
Pengfei Wei 0001, Fei Wu 0004, Ying Sun 0023, Xiaoyuan Jing
IEEE Signal Process. Lett.1
2022 Subdomain Adaptation With Manifolds Discrepancy Alignment
abstract
Reducing domain divergence is a key step in transfer learning. Existing works focus on the minimization of global domain divergence. However, two domains may consist of several shared subdomains, and differ from each other in each subdomain. In this article, we take the local divergence of subdomains into account in transfer. Specifically, we propose to use the low-dimensional manifold to represent the subdomain, and align the local data distribution discrepancy in each manifold across domains. A manifold maximum mean discrepancy (M3D) is developed to measure the local distribution discrepancy in each manifold. We then propose a general framework, called transfer with manifolds discrepancy alignment (TMDA), to couple the discovery of data manifolds with the minimization of M3D. We instantiate TMDA in the subspace learning case considering both the linear and nonlinear mappings. We also instantiate TMDA in the deep learning framework. Experimental studies show that TMDA is a promising method for various transfer learning tasks.
Pengfei Wei 0001, Yiping Ke, Xinghua Qu, Tze-Yun Leong
IEEE Trans. Cybern.1
2022 Easy-But-Effective Domain Sub-Similarity Learning for Transfer Regression
abstract
Transfer covariance function, which can model domain similarity and adaptively control the knowledge transfer across domains, is widely used in transfer learning. In this paper, we concentrate on Gaussian process (GP) models using a transfer covariance function for regression problems in a black-box learning scenario. Precisely, we investigate a family of rather general transfer covariance functions,${T}_{*}$, that can model the heterogeneous sub-similarities of domains through multiple kernel learning. A necessary and sufficient condition to obtain validGPs using${T}_{*}$($GP_{T_{*}}$) for any data is given. This condition becomes specially handy for practical applications as (i) it enables semantic interpretations of the sub-similarities and (ii) it can readily be used for model learning. In particular, we propose a computationally inexpensive model learning rule that can explicitly capture different sub-similarities of domains. We propose two instantiations of$GP_{T_{*}}$, one with a set of predefined constant base kernels and one with a set of learnable parametric base kernels. Extensive experiments on 36 synthetic transfer tasks and 12 real-world transfer tasks demonstrate the effectiveness of$GP_{T_{*}}$on the sub-similarity capture and the transfer performance.
Pengfei Wei 0001, Ramón Sagarna, Yiping Ke, Yew-Soon Ong
IEEE Trans. Knowl. Data Eng.1
2021 Causal Modeling with Stochastic Confounders
abstract
This work extends causal inference in temporal models with stochastic confounders. We propose a new approach to variational estimation of causal inference based on a representer theorem with a random input space. We estimate causal effects involving latent confounders that may be interdependent and time-varying from sequential, repeated measurements in an observational study. Our approach extends current work that assumes independent, non-temporal latent confounders with potentially biased estimators. We introduce a simple yet elegant algorithm without parametric specification on model components. Our method avoids the need for expensive and careful parameterization in deploying complex models, such as deep neural networks in existing approaches, for causal inference and analysis. We demonstrate the effectiveness of our approach on various benchmark temporal datasets.
Thanh Vinh Vo, Pengfei Wei 0001, Wicher Bergsma, Tze-Yun Leong
AISTATS2
2021 Semantic Preserving Generative Adversarial Network For Cross-Modal Hashing
abstract
Cross-modal hashing has achieved significant progress in recent years. However, how to effectively learn more discriminative hash codes of each modality and simultaneous alleviate the loss of modality information is still a challenging problem. Focusing on this problem, in this paper, we propose a novel cross-modal hashing approach named Semantic Preserving Generative Adversarial Network (SPGAN). The overall network architecture consists of two sub-networks, i.e., a semantic preserving generative adversarial network module and a discriminative hashing module. The generator maps text features into the image feature space. And the discriminator judges whether the feature representations are real image features or generated image features. The adversarial learning process can effectively reduce modality difference and preserve information of the image modality as much as possible. The discriminative hashing module projects the real and generated image features into a Hamming space to obtain hash codes, and explores semantic similarities for enhancing the discriminant ability of hash codes. Experiments on two widely used datasets demonstrate that SPGAN can outperform state-of-the-art related works.
Fei Wu 0004, Xiaokai Luo, Qinghua Huang, Pengfei Wei 0001, Ying Sun 0023, Xiwei Dong, Zhiyong Wu 0006
ICIP4
2021 Practical Multisource Transfer Regression With Source-Target Similarity Captures
abstract
A key challenge in many applications of multisource transfer learning is to explicitly capture the diverse source-target similarities. In this article, we are concerned with stretching the set of practical approaches based on Gaussian process (GP) models to solve multisource transfer regression problems. Precisely, we first investigate the feasibility and performance of a family of transfer covariance functions that represent the pairwise similarity of each source and the target domain. We theoretically show that using such a transfer covariance function for general GP modeling can only capture the same similarity coefficient for all the sources, and thus may result in unsatisfactory transfer performance. This outcome, together with the scalability issues of a single GP based approach, leads us to propose TCMSStack, an integrated framework incorporating a separate transfer covariance function for each source and stacking. Contrary to typical stacking approaches, TCMSStack learns the source-target similarity in each base GP model by considering the dependencies of the other sources along the process. We introduce two instances of the proposed TCMSStack. Extensive experiments on one synthetic and two real-world data sets, with learning settings up to 11 sources for the latter, demonstrate the effectiveness of our approach.
Pengfei Wei 0001, Ramón Sagarna, Yiping Ke, Yew-Soon Ong
IEEE Trans. Neural Networks Learn. Syst.1
2020 Succinct Adaptive Manifold Transfer
abstract
Capturing the relatedness of different domains is a key challenge in transferring knowledge across domains. In this paper, we propose an effective and efficient Gaussian process (GP) modelling framework, mTGPmk, that can explicitly model domain relatedness and adaptively control the space as well as the strength of knowledge transfer. mTGPmk takes both the discrepancy of input feature space and the discrepancy of predictive function into account in the transfer procedure. Specifically, mTGPmk adaptively selects a good latent manifold shared by different domains, and utilizes a parametric similarity coefficient to measure the predictive function covariance of different domains in this manifold. The latent shared manifold and the similarity coefficient are jointly learned in a coupled manner. By doing so, mTGPmk maximizes the strength of the shared knowledge transfer by choosing the transfer space with the best transfer capacity. More importantly, mTGPmk exploits a succinct and computationally efficient manifold learning approach so that it can be well trained with scarce target training data. Extensive experimental studies using 36 synthetic transfer tasks and 10 real-world transfer tasks show the effectiveness of mTGPmk on capturing the relatedness and the transfer adaptiveness.
Pengfei Wei 0001, Yiping Ke, Zhiqiang Xu 0003, Tze-Yun Leong
CIKM1
2020 Randomized Transferable Machine
abstract
Feature-based transfer is one of the most effective methodologies for transfer learning. Existing studies usually assume that the learned new feature representation is truly domain-invariant, and thus directly train a transfer model M on source domain. In this paper, we consider a more realistic scenario where the new feature representation is suboptimal and small divergence still exists across domains. We propose a new learning strategy with a transfer model called Randomized Transferable Machine (RTM). More specifically, we work on source data with the new feature representation learned from existing feature-based transfer methods. The key idea is to enlarge source training data populations by randomly corrupting source data using some noises, and then train a transfer model ~M that performs well on all the corrupted source data populations. In principle, the more corruptions are made, the higher the probability of the target data can be covered by the constructed source populations. and thus better transfer performance can be achieved by ~M An ideal case is with infinite corruptions, which however is infeasible in reality. We develop a marginalized solution with linear regression model and dropout noise. With a marginalization trick, we can train an RTM that is equivalently to training using infinite source noisy populations without truly conducting any corruption. More importantly, such an RTM has a closed-form solution, which enables very fast and efficient training. Extensive experiments on various real-world transfer tasks show that RTM is a promising transfer model.
Pengfei Wei 0001, Tze-Yun Leong
ICPR1
2020 MESA: Boost Ensemble Imbalanced Learning with MEta-SAmpler
abstract
Imbalanced learning (IL), i.e., learning unbiased models from class-imbalanced data, is a challenging problem. Typical IL methods including resampling and reweighting were designed based on some heuristic assumptions. They often suffer from unstable performance, poor applicability, and high computational cost in complex tasks where their assumptions do not hold. In this paper, we introduce a novel ensemble IL framework named MESA. It adaptively resamples the training set in iterations to get multiple classifiers and forms a cascade ensemble model. MESA directly learns the sampling strategy from data to optimize the final metric beyond following random heuristics. Moreover, unlike prevailing meta-learning-based IL solutions, we decouple the model-training and meta-training in MESA by independently train the meta-sampler over task-agnostic meta-data. This makes MESA generally applicable to most of the existing learning models and the meta-sampler can be efficiently applied to new tasks. Extensive experiments on both synthetic and real-world tasks demonstrate the effectiveness, robustness, and transferability of MESA. Our code is available at https://github.com/ZhiningLiu1998/mesa.
Zhining Liu 0002, Pengfei Wei 0001, Jing Jiang 0002, Wei Cao 0007, Jiang Bian 0002, Yi Chang 0001
NeurIPS2
2020 Cooperative Heterogeneous Deep Reinforcement Learning
abstract
Numerous deep reinforcement learning agents have been proposed, and each of them has its strengths and flaws. In this work, we present a Cooperative Heterogeneous Deep Reinforcement Learning (CHDRL) framework that can learn a policy by integrating the advantages of heterogeneous agents. Specifically, we propose a cooperative learning framework that classifies heterogeneous agents into two classes: global agents and local agents. Global agents are off-policy agents that can utilize experiences from the other agents. Local agents are either on-policy agents or population-based evolutionary algorithms (EAs) agents that can explore the local area effectively. We employ global agents, which are sample-efficient, to guide the learning of local agents so that local agents can benefit from the sample-efficient agents and simultaneously maintain their advantages, e.g., stability. Global agents also benefit from effective local searches. Experimental studies on a range of continuous control tasks from the Mujoco benchmark show that CHDRL achieves better performance compared with state-of-the-art baselines.
Pengfei Wei 0001, Jing Jiang 0002, Guodong Long, Qinghua Lu 0001, Chengqi Zhang
NeurIPS2
2019 Knowledge Transfer based on Multiple Manifolds Assumption
abstract
Unsupervised domain adaptation is a popular but challenging problem setting. Existing unsupervised domain adaptation methods are based on the single manifold assumption, i.e., data are sampled from a single low-dimensional manifold, and thus may not well capture the complex characteristic of the real-world data. In this paper, we propose to transfer knowledge across domains under the multiple manifolds assumption that assumes the data are sampled from multiple low-dimensional manifolds. Specifically, we develop a multiple manifolds information transfer framework (MMIT). The proposed MMIT aims to transfer the multiple manifolds information, which is represented by the data manifold neighborhood structure, with the the best adaptation capacity. To do so, we propose to couple the multiple manifolds information transfer with the domain distribution discrepancy minimization in the adaptation procedure. Experimental studies demonstrate that MMIT achieves the promising adaptation performance on various real-world adaptation tasks.
Pengfei Wei 0001, Yiping Ke
CIKM1
2019 Learning Disentangled Semantic Representation for Domain Adaptation
abstract
Domain adaptation is an important but challenging task. Most of the existing domain adaptation methods struggle to extract the domain-invariant representation on the feature space with entangling domain information and semantic information. Different from previous efforts on the entangled feature space, we aim to extract the domain invariant semantic information in the latent disentangled semantic representation (DSR) of the data. In DSR, we assume the data generation process is controlled by two independent sets of variables, i.e., the semantic latent variables and the domain latent variables. Under the above assumption, we employ a variational auto-encoder to reconstruct the semantic latent variables and domain latent variables behind the data. We further devise a dual adversarial network to disentangle these two sets of reconstructed latent variables. The disentangled semantic latent variables are finally adapted across the domains. Experimental studies testify that our model yields state-of-the-art performance on several domain adaptation benchmark datasets.
Ruichu Cai, Zijian Li 0001, Pengfei Wei 0001, Jie Qiao, Kun Zhang 0001, Zhifeng Hao 0004
IJCAI3
2019 Multi-source sequential knowledge regression by using transfer RNN units
Xiurui Xie, Guisong Liu, Pengfei Wei 0001, Hong Qu 0002
Neural Networks4
2019 A General Domain Specific Feature Transfer Framework for Hybrid Domain Adaptation
abstract
Heterogeneous domain adaptation needs supplementary information to link up different domains. However, such supplementary information may not always be available in real cases. In this paper, a new problem setting called hybrid domain adaptation is investigated. It is a special case of heterogeneous domain adaptation, in which different domains share some common features, but also have their own domain specific features. We leverage upon common features instead of supplementary information to achieve effective adaptation. We propose a general domain specific feature transfer framework, which can link up different domains using common features and simultaneously reduce domain divergences. Specifically, we learn the translations between common features and domain specific features. Then, we cross-use the learned translations to transfer the domain specific features of one domain to another domain. Finally, we compose a homogeneous space in which the domain divergences are minimized. We instantiate the general framework to a linear case and a nonlinear case. Extensive experiments verify the effectiveness of the two cases.
Pengfei Wei 0001, Yiping Ke, Chi Keong Goh
IEEE Trans. Knowl. Data Eng.1
2019 Feature Analysis of Marginalized Stacked Denoising Autoenconder for Unsupervised Domain Adaptation
abstract
Marginalized stacked denoising autoencoder (mSDA), has recently emerged with demonstrated effectiveness in domain adaptation. In this paper, we investigate the rationale for why mSDA benefits domain adaptation tasks from the perspective of adaptive regularization. Our investigations focus on two types of feature corruption noise: Gaussian noise (mSDAg) and Bernoulli dropout noise (mSDAbd). Both theoretical and empirical results demonstrate that mSDAbd successfully boosts the adaptation performance but mSDAgfails to do so. We then propose a new mSDA with data-dependent multinomial dropout noise (mSDAmd) that overcomes the limitations of mSDAbdand further improves the adaptation performance. Our mSDAmdis based on a more realistic assumption: different features are correlated and, thus, should be corrupted with different probabilities. Experimental results demonstrate the superiority of mSDAmdto mSDAbdon the adaptation performance and the convergence speed. Finally, we propose a deep transferable feature coding (DTFC) framework for unsupervised domain adaptation. The motivation of DTFC is that mSDA fails to consider the distribution discrepancy across different domains in the feature learning process. We introduce a new element to mSDA: domain divergence minimization by maximum mean discrepancy. This element is essential for domain adaptation as it ensures the extracted deep features to have a small distribution discrepancy. The effectiveness of DTFC is verified by extensive experiments on three benchmark data sets for both Bernoulli dropout noise and multinomial dropout noise.
Pengfei Wei 0001, Yiping Ke, Chi Keong Goh
IEEE Trans. Neural Networks Learn. Syst.1
2018 Transfer Hawkes Processes with Content Information
abstract
Hawkes processes are widely used for modeling event cascades. However, content and cross-domain information which is also instrumental in modeling is usually neglected. In this paper, we propose a novel model called transfer Hybrid Least Square for Hawkes (trHLSH) that incorporates Hawkes processes with content and cross-domain information. We also present the effective learning algorithm for the model. Evaluation on both synthetic and real-world datasets demonstrates that the proposed model can jointly learn knowledge from temporal, content and cross-domain information, and has better performance in terms of network recovery and prediction.
Tianbo Li, Pengfei Wei 0001, Yiping Ke
ICDM2
2018 Uncluttered Domain Sub-Similarity Modeling for Transfer Regression
abstract
Transfer covariance functions, which can model domain similarities and adaptively control the knowledge transfer across domains, are widely used in Gaussian process (GP) based transfer learning. We focus on regression problems in a black-box learning scenario, and study a family of rather general transfer covariance functions, T_*, that can model the similarity heterogeneity of domains through multiple kernel learning. A necessary and sufficient condition that (i) validates GPs using T_* for any data and (ii) provides semantic interpretations is given. Moreover, building on this condition, we propose a computationally inexpensive model learning rule that can explicitly capture different sub-similarities of domains. Extensive experiments on one synthetic dataset and four real-world datasets demonstrate the effectiveness of the learned GP on the sub-similarity capture and the transfer performance.
Pengfei Wei 0001, Ramón Sagarna, Yiping Ke, Yew-Soon Ong
ICDM1
2017 Domain Specific Feature Transfer for Hybrid Domain Adaptation
abstract
Heterogeneous domain adaptation needs supplementary information to link up domains. However, this supplementary information is unavailable in many real cases. In this paper, a new problem setting called hybrid domain adaptation is investigated. It is a special case of heterogeneous domain adaptation in which different domains share some common features, but also have their own domain specific features. In this case, it can be efficiently solved without any supplementary information by using the common features to link up the domains in adaptation. We propose a domain specific feature transfer (DSFT) method, which can link up different domains using the common features and simultaneously reduce domain divergences. Specifically, we first learn the translations between the common features and the domain specific features. Then we cross-use the learned translations to transfer the domain specific features of one domain to another domain. Finally, we compose a homogeneous space in which the domain divergences are minimized. Extensive experiments verify the effectiveness of our proposed method.
Pengfei Wei 0001, Yiping Ke, Chi Keong Goh
ICDM1
2017 Source-Target Similarity Modelings for Multi-Source Transfer Gaussian Process Regression
abstract
A key challenge in multi-source transfer learning is to capture the diverse inter-domain similarities. In this paper, we study different approaches based on Gaussian process models to solve the multi-source transfer regression problem. Precisely, we first investigate the feasibility and performance of a family of transfer covariance functions that represent the pairwise similarity of each source and the target domain. We theoretically show that using such a transfer covariance function for general Gaussian process modelling can only capture the same similarity coefficient for all the sources, and thus may result in unsatisfactory transfer performance. This leads us to propose TC$_{MS}$Stack, an integrated strategy incorporating the benefits of the transfer covariance function and stacking. Extensive experiments on one synthetic and two real-world datasets, with learning settings of up to 11 sources for the latter, demonstrate the effectiveness of our proposed TC$_{MS}$Stack.
Pengfei Wei 0001, Ramón Sagarna, Yiping Ke, Yew-Soon Ong, Chi Keong Goh
ICML1
2016 Deep Nonlinear Feature Coding for Unsupervised Domain Adaptation
Pengfei Wei 0001, Yiping Ke, Chi Keong Goh
IJCAI1
2013 Speech Emotional Features Extraction Based on Electroglottograph
abstract
This study proposes two classes of speech emotional features extracted from electroglottography (EGG) and speech signal. The power-law distribution coefficients (PLDC) of voiced segments duration, pitch rise duration, and pitch down duration are obtained to reflect the information of vocal folds excitation. The real discrete cosine transform coefficients of the normalized spectrum of EGG and speech signal are calculated to reflect the information of vocal tract modulation. Two experiments are carried out. One is of proposed features and traditional features based on sequential forward floating search and sequential backward floating search. The other is the comparative emotion recognition based on support vector machine. The results show that proposed features are better than those commonly used in the case of speaker-independent and content-independent speech emotion recognition.
Lijiang Chen, Xia Mao, Pengfei Wei 0001, Angelo Compare
Neural Comput.3
2012 Mandarin emotion recognition combining acoustic and emotional point information
Lijiang Chen, Xia Mao, Pengfei Wei 0001, Yu-Li Xue, Mitsuru Ishizuka
Appl. Intell.3