Xudong Mao

dblp:55/11145 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-1952-4176ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Aspect-Oriented Prompt with Adaptive Cross-Modal Fusion for Multimodal Sentiment Analysis
Xudong Mao, Fuqiang Yu, Lap-Kei Lee, Fu Lee Wang, Zhenguo Yang
DASFAA (3)4
2026 Training for Identity, Inference for Controllability: A Unified Approach to Tuning-Free Face Personalization
abstract
Tuning-free face personalization methods have developed along two distinct paradigms: text embedding approaches that map facial features into the text embedding space, and adapter-based methods that inject features through auxiliary cross-attention layers. While both paradigms have shown promise, existing methods struggle to simultaneously achieve high identity fidelity and flexible text controllability. We introduce UniID, a unified tuning-free framework that synergistically integrates both paradigms. Our key insight is that when merging these approaches, they should mutually reinforce only identity-relevant information while preserving the original diffusion prior for non-identity attributes. We realize this through a principled training-inference strategy: during training, we employ an identity-focused learning scheme that guides both branches to capture identity features exclusively; at inference, we introduce a normalized rescaling mechanism that recovers the text controllability of the base diffusion model while enabling complementary identity signals to enhance each other. This principled design enables UniID to achieve high-fidelity face personalization with flexible text controllability. Extensive experiments against six state-of-the-art methods demonstrate that UniID achieves superior performance in both identity preservation and text controllability.
Lianyu Pang, Qiping Wang 0004, Baoquan Zhao, Zhenguo Yang, Qing Li 0001, Xudong Mao
ICMR7
2026 Expressive Human Volumetric Video Generation With Rich Text
abstract
Plain text has become the dominant interactive interface for text-driven human volumetric video generation. However, its limited customization options hinder users from expressing motion effects with accuracy. For example, plain text struggles to specify continuous variables such as motion amplitude, speed, and joint trajectories with precision, and it fails to convey stylized motion characteristics. Additionally, crafting detailed textual prompts for complex motion sequences is cumbersome, while excessively long prompts strain text encoders. To address these limitations, we propose a rich text-based framework that supports font styles, sizes, and trajectory sketching. By extracting motion-related attributes from rich text, our method enables fine-grained control over motion styles, precise speed regulation, and accurate joint trajectory manipulation. These capabilities are realized through gradient-guided noise editing and ControlNet-based motion optimization, which operate within the latent motion diffusion process. Specifically, we design a unified gradient-guided adaptation mechanism to ensure that the generated motion video adheres strictly to the specified constraints. Furthermore, we introduce realism-oriented optimization for stylistic and joint-level control, refining motion synthesis at a granular level to produce smoother, more natural movements. We present multiple comparative evaluations showcasing volumetric video generation from both rich text and plain text. Through quantitative analysis, we demonstrate that our method surpasses strong plain-text baselines, producing expressive, customizable human volumetric motion videos.
Guanghui Yue 0001, Wei Zhou 0021, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2025 CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
abstract
Recent advances in text-to-image personalization have enabled high-quality and controllable image synthesis for user-provided concepts. However, existing methods still struggle to balance identity preservation with text alignment. Our approach is based on the fact that generating prompt-aligned images requires a precise semantic understanding of the prompt, which involves accurately processing the interactions between the new concept and its surrounding context tokens within the CLIP text encoder. To address this, we aim to embed the new concept properly into the input embedding space of the text encoder, allowing for seamless integration with existing tokens. We introduce Context Regularization (CoRe), which enhances the learning of the new concept's text embedding by regularizing its context tokens in the prompt. This is based on the insight that appropriate output vectors of the text encoder for the context tokens can only be achieved if the new concept's text embedding is correctly learned. CoRe can be applied to arbitrary prompts without requiring the generation of corresponding images, thus improving the generalization of the learned text embedding. Additionally, CoRe can serve as a test-time optimization technique to further enhance the generations for specific prompts. Comprehensive experiments demonstrate that our method outperforms several baseline methods in both identity preservation and text alignment.
Feize Wu, Yun Pang, Lianyu Pang, Jian Yin 0001, Baoquan Zhao, Qing Li 0001, Xudong Mao
AAAI8
2025 MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal Learning
abstract
Recent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints.
Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao
ICME7
2025 DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
abstract
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets.
Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
ACM Multimedia4
2025 VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
abstract
Cross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex queries. To evaluate our method, we develop CrossVideoQA, a comprehensive benchmark specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning.
Yiran Meng, Junhong Ye, Wei Zhou 0021, Guanghui Yue 0001, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao
ACM Multimedia5
2025 PairEdit: Learning Semantic Variations for Exemplar-based Image Editing
abstract
Recent advancements in text-guided image editing have achieved notable success by leveraging natural language prompts for fine-grained semantic control. However, certain editing semantics are challenging to specify precisely using textual descriptions alone. A practical alternative involves learning editing semantics from paired source-target examples. Existing exemplar-based editing methods still rely on text prompts describing the change within paired examples or learning implicit text-based editing instructions. In this paper, we introduce PairEdit, a novel visual editing method designed to effectively learn complex editing semantics from a limited number of image pairs or even a single image pair, without using any textual guidance. We propose a target noise prediction that explicitly models semantic variations within paired images through a guidance direction term. Moreover, we introduce a content-preserving noise schedule to facilitate more effective semantic learning. We also propose optimizing distinct LoRAs to disentangle the learning of semantic variations from content. Extensive qualitative and quantitative evaluations demonstrate that PairEdit successfully learns intricate semantics while significantly improving content consistency compared to baseline methods. Code is available at https://github.com/xudonmao/PairEdit.
Haoguang Lu, Zhenguo Yang, Aurele Tohokantche Gnanha, Fu Lee Wang, Qing Li 0001, Xudong Mao
NeurIPS7
2025 EviD-GAN: Improving GAN With an Infinite Set of Discriminators at Negligible Cost
abstract
Ensemble learning improves the capability of convolutional neural network (CNN)-based discriminators, whose performance is crucial to the quality of generated samples in generative adversarial network (GAN). However, this learning strategy results in a significant increase in the number of parameters along with computational overhead. Meanwhile, the suitable number of discriminators required to enhance GAN performance is still being investigated. To mitigate these issues, we propose an evidential discriminator for GAN (EviD-GAN)-code is available at https://github.com/Tohokantche/EviD-GAN-to learn both the model (epistemic) and data (aleatoric) uncertainties. Specifically, by analyzing three GAN models, the relation between the distribution of discriminator's output and the generator performance has been discovered yielding a general formulation of GAN framework. With the above analysis, the evidential discriminator learns the degree of aleatoric and epistemic uncertainties via imposing a higher order distribution constraint over the likelihood as expressed in the discriminator's output. This constraint can learn an ensemble of likelihood functions corresponding to an infinite set of discriminators. Thus, EviD-GAN aggregates knowledge through the ensemble learning of discriminator that allows the generator to benefit from an informative gradient flow at a negligible computational cost. Furthermore, inspired by the gradient direction in maximum mean discrepancy (MMD)-repulsive GAN, we design an asymmetric regularization scheme for EviD-GAN. Unlike MMD-repulsive GAN that performs at the distribution level, our regularization scheme is based on a pairwise loss function, performs at the sample level, and is characterized by an asymmetric behavior during the training of generator and discriminator. Experimental results show that the proposed evidential discriminator is cost-effective, consistently improves GAN in terms of Frechet inception distance (FID) and inception score (IS), and performs better than other competing models that use multiple discriminators.
Aurele Tohokantche Gnanha, Wenming Cao 0002, Xudong Mao, Si Wu 0002, Hau-San Wong, Qing Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Cross Initialization for Face Personalization of Text-to-Image Models
abstract
Recently, there has been a surge in face personalization techniques, benefiting from the advanced capabilities of pretrained text-to-image diffusion models. Among these, a notable method is Textual Inversion, which generates personalized images by inverting given images into textual embeddings. However, methods based on Textual Inversion still struggle with balancing the tradeoff between reconstruction quality and editability. In this study, we examine this issue through the lens of initialization. Upon closely examining traditional initialization methods, we identified a significant disparity between the initial and learned embeddings in terms of both scale and orientation. The scale of the learned embedding can be up to 100 times greater than that of the initial embedding. Such a significant change in the embedding could increase the risk of overfitting, thereby compromising the editability. Driven by this observation, we introduce a novel initialization method, termed Cross Initialization, that significantly narrows the gap between the initial and learned embeddings. This method not only improves both reconstruction and editability but also reduces the optimization steps from 5,000 to 320. Furthermore, we apply a regularization term to keep the learned embedding close to the initial embedding. We show that when combined with Cross Initialization, this regularization term can effectively improve editability. We provide comprehensive empirical evidence to demonstrate the superior performance of our method compared to the baseline methods. Notably, in our experiments, Cross Initialization is the only method that successfully edits an individual's facial expression. Additionally, a fast version of our method allows for capturing an input image in roughly 26 seconds, while surpassing the baseline methods in terms of both reconstruction and editability. Code is available at https://github.com/lyuPang/CrossInitialization.
Lianyu Pang, Jian Yin 0001, Haoran Xie 0001, Qiping Wang 0002, Qing Li 0001, Xudong Mao
CVPR6
2024 AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
abstract
Recent advances in text-to-image models have enabled high-quality personalized image synthesis based on user-provided concepts with flexible textual control. In this work, we analyze the limitations of two primary techniques in text-to-image personalization: Textual Inversion and DreamBooth. When integrating the learned concept into new prompts, Textual Inversion tends to overfit the concept, while DreamBooth often overlooks it. We attribute these issues to the incorrect learning of the embedding alignment for the concept. To address this, we introduce AttnDreamBooth, a novel approach that separately learns the embedding alignment, the attention map, and the subject identity across different training stages. We also introduce a cross-attention map regularization term to enhance the learning of the attention map. Our method demonstrates significant improvements in identity preservation and text alignment compared to the baseline methods.
Lianyu Pang, Jian Yin 0001, Baoquan Zhao, Feize Wu, Fu Lee Wang, Qing Li 0001, Xudong Mao
NeurIPS7
2022 Cycle Encoding of a StyleGAN Encoder for Improved Reconstruction and Editability
abstract
GAN inversion aims to invert an input image into the latent space of a pre-trained GAN. Despite the recent advances in GAN inversion, there remain challenges to mitigate the tradeoff between distortion and editability, i.e. reconstructing the input image accurately and editing the inverted image with a small visual quality drop. The recently proposed pivotal tuning model makes significant progress towards reconstruction and editability, by using a two-step approach that first inverts the input image into a latent code, called pivot code, and then alters the generator so that the input image can be accurately mapped into the pivot code. Here, we show that both reconstruction and editability can be improved by a proper design of the pivot code. We present a simple yet effective method, named cycle encoding, for a high-quality pivot code. The key idea of our method is to progressively train an encoder in varying spaces according to a cycle scheme: W->W+->W. This training methodology preserves the properties of both W and W+ spaces, i.e. high editability of W and low distortion of W+. To further decrease the distortion, we also propose to refine the pivot code with an optimization-based method, where a regularization term is introduced to reduce the degradation in editability. Qualitative and quantitative comparisons to several state-of-the-art methods demonstrate the superiority of our approach.
Xudong Mao, Liujuan Cao, Aurele Tohokantche Gnanha, Zhenguo Yang, Qing Li 0001, Rongrong Ji
ACM Multimedia1
2022 αβ-GAN: Robust generative adversarial networks
Aurele Tohokantche Gnanha, Wenming Cao 0002, Xudong Mao, Si Wu 0002, Hau-San Wong, Qing Li 0001
Inf. Sci.3
2022 The residual generator: An improved divergence minimization framework for GAN
Aurele Tohokantche Gnanha, Wenming Cao 0002, Xudong Mao, Si Wu 0002, Hau-San Wong, Qing Li 0001
Pattern Recognit.3
2022 Learning semantic alignment from image for text-guided image inpainting
Yucheng Xie, Zehang Lin, Zhenguo Yang, Xingcai Wu, Xudong Mao, Qing Li 0001, Wenyin Liu
Vis. Comput.6
2021 Generative Semi-supervised Learning for Multivariate Time Series Imputation
abstract
The missing values, widely existed in multivariate time series data, hinder the effective data analysis. Existing time series imputation methods do not make full use of the label information in real-life time series data. In this paper, we propose a novel semi-supervised generative adversarial network model, named SSGAN, for missing value imputation in multivariate time series data. It consists of three players, i.e., a generator, a discriminator, and a classifier. The classifier predicts labels of time series data, and thus it drives the generator to estimate the missing values (or components), conditioned on observed components and data labels at the same time. We introduce a temporal reminder matrix to help the discriminator better distinguish the observed components from the imputed ones. Moreover, we theoretically prove that, SSGAN using the temporal reminder matrix and the classifier does learn to estimate missing values converging to the true data distribution when the Nash equilibrium is achieved. Extensive experiments on three public real-world datasets demonstrate that, SSGAN yields a more than 15% gain in performance, compared with the state-of-the-art methods.
Xiaoye Miao, Jun Wang 0184, Yunjun Gao, Xudong Mao, Jianwei Yin
AAAI5
2021 Image-to-Image Translation via Hierarchical Style Disentanglement
abstract
Recently, image-to-image translation has made significant progress in achieving both multi-label (i.e., translation conditioned on different labels) and multi-style (i.e., generation with diverse styles) tasks. However, due to the unexplored independence and exclusiveness in the labels, existing endeavors are defeated by involving uncontrolled manipulations to the translation results. In this paper, we propose Hierarchical Style Disentanglement (HiSD) to address this issue. Specifically, we organize the labels into a hierarchical tree structure, in which independent tags, exclusive attributes, and disentangled styles are allocated from top to bottom. Correspondingly, a new translation process is designed to adapt the above structure, in which the styles are identified for controllable translations. Both qualitative and quantitative results on the CelebA-HQ dataset verify the ability of the proposed HiSD. The code has been released at https://github.com/imlixinyang/HiSD.
Shengchuan Zhang, Jie Hu 0018, Liujuan Cao, Xiaopeng Hong, Xudong Mao, Feiyue Huang, Yongjian Wu 0001, Rongrong Ji
CVPR6
2021 Collaborative Learning of Bidirectional Decoders for Unsupervised Text Style Transfer
abstract
Unsupervised text style transfer aims to alter the underlying style of the text to a desired value while keeping its style-independent semantics, without the support of parallel training corpora.Existing methods struggle to achieve both high style conversion rate and low content loss, exhibiting the over-transfer and undertransfer problems.We attribute these problems to the conflicting driving forces of the style conversion goal and content preservation goal.In this paper, we propose a collaborative learning framework for unsupervised text style transfer using a pair of bidirectional decoders, one decoding from left to right while the other decoding from right to left.In our collaborative learning mechanism, each decoder is regularized by knowledge from its peer which has a different knowledge acquisition process.The difference is guaranteed by their opposite decoding directions and a distinguishability constraint.As a result, mutual knowledge distillation drives both decoders to a better optimum and alleviates the over-transfer and undertransfer problems.Experimental results on two benchmark datasets show that our framework achieves strong empirical results on both style compatibility and content preservation.
Yun Ma 0001, Yangbin Chen, Xudong Mao, Qing Li 0001
EMNLP (1)3
2021 Multi-branch Semantic Learning Network for Text-to-Image Synthesis
abstract
In this paper, we propose a multi-branch semantic learning network (MSLN) to generate image according to textual description by taking into account global and local textual semantics, which consists of two stages. The first stage generates a coarse-grained image based on the sentence features. In the second stage, a multi-branch fine-grained generation model is constructed to inject the sentence-level and word-level semantics into two coarse-grained images by global and local attention modules, which generate global and local fine-grained image textures, respectively. In particular, we devise a channel fusion module (CFM) to fuse the global and local fine-grained features in the multi-branch fine-grained stage and generate the output image. Extensive experiments conducted on the CUB-200 dataset and Oxford-102 dataset demonstrate the superior performance of the proposed method. (e.g., FID is reduced from 16.09 to 14.43 on CUB-200).
Jiading Ling, Xingcai Wu, Zhenguo Yang, Xudong Mao, Qing Li 0001, Wenyin Liu
MMAsia4
2021 Revisiting Discriminator in GAN Compression: A Generator-discriminator Cooperative Compression Scheme
abstract
Recently, a series of algorithms have been explored for GAN compression, which aims to reduce tremendous computational overhead and memory usages when deploying GANs on resource-constrained edge devices. However, most of the existing GAN compression work only focuses on how to compress the generator, while fails to take the discriminator into account. In this work, we revisit the role of discriminator in GAN compression and design a novel generator-discriminator cooperative compression scheme for GAN compression, termed GCC. Within GCC, a selective activation discriminator automatically selects and activates convolutional channels according to a local capacity constraint and a global coordination constraint, which help maintain the Nash equilibrium with the lightweight generator during the adversarial training and avoid mode collapse. The original generator and discriminator are also optimized from scratch, to play as a teacher model to progressively refine the pruned generator and the selective activation discriminator. A novel online collaborative distillation scheme is designed to take full advantage of the intermediate feature of the teacher generator and discriminator to further boost the performance of the lightweight generator. Extensive experiments on various GAN-based generation tasks demonstrate the effectiveness and generalization of GCC. Among them, GCC contributes to reducing 80% computational costs while maintains comparable performance in image translation tasks.
Jie Wu 0032, Xuefeng Xiao 0001, Fei Chao 0001, Xudong Mao, Rongrong Ji
NeurIPS5
2020 Mixing Up Real Samples and Adversarial Samples for Semi-Supervised Learning
abstract
Consistency regularization methods have shown great success in semi-supervised learning tasks. Most existing methods focus on either the local neighborhood or in-between neighborhood of training samples to enforce the consistency constraint. In this paper, we propose a novel generalized framework called Adversarial Mixup (AdvMixup), which unifies the local and in-between neighborhood approaches by defining a virtual data distribution along the paths between the training samples and adversarial samples. Experimental results on both synthetic data and benchmark datasets exhibit that our AdvMixup can achieve better performance and robustness than state-of-the-art methods for semi-supervised learning.
Yun Ma 0001, Xudong Mao, Yangbin Chen, Qing Li 0001
IJCNN2
2019 GAN with Pixel and Perceptual Regularizations for Photo-Realistic Joint Deblurring and Super-Resolution
Zhenguo Yang, Xudong Mao, Qing Li 0001, Wenyin Liu
CGI3
2019 Self-attention StarGAN for Multi-domain Image-to-Image Translation
Ziliang He, Zhenguo Yang, Xudong Mao, Jianming Lv, Qing Li 0001, Wenyin Liu
ICANN (3)3
2019 Generating Adversarial Examples by Adversarial Networks for Semi-supervised Learning
Yun Ma 0001, Xudong Mao, Yangbin Chen, Qing Li 0001
WISE2
2019 Multi-Task Learning for Abstractive and Extractive Summarization
abstract
The abstractive method and extractive method are two main approaches for automatic document summarization. In this paper, to fully integrate the relatedness and advantages of both approaches, we propose a general unified framework for abstractive summarization which incorporates extractive summarization as an auxiliary task. In particular, our framework is composed of a shared hierarchical document encoder, a hierarchical attention mechanism-based decoder, and an extractor. We adopt multi-task learning method to train these two tasks jointly, which enables the shared encoder to better capture the semantics of the document. Moreover, as our main task is abstractive summarization, we constrain the attention learned in the abstractive task with the labels of the extractive task to strengthen the consistency between the two tasks. Experiments on the CNN/DailyMail dataset demonstrate that both the auxiliary task and the attention constraint contribute to improve the performance significantly, and our model is comparable to the state-of-the-art abstractive models. In addition, we cut half number of labels of the extractive task, pretrain the extractor, and jointly train the two tasks using the estimated sentence salience of the extractive task to constrain the attention of the abstractive task. The results do not decrease much compared with using full-labeled data of the auxiliary task.
Yangbin Chen, Yun Ma 0001, Xudong Mao, Qing Li 0001
Data Sci. Eng.3
2019 On the Effectiveness of Least Squares Generative Adversarial Networks
abstract
Unsupervised learning with generative adversarial networks (GANs) has proven to be hugely successful. Regular GANs hypothesize the discriminator as a classifier with the sigmoid cross entropy loss function. However, we found that this loss function may lead to the vanishing gradients problem during the learning process. To overcome such a problem, we propose in this paper the Least Squares Generative Adversarial Networks (LSGANs) which adopt the least squares loss for both the discriminator and the generator. We show that minimizing the objective function of LSGAN yields minimizing the Pearson$\chi ^2$divergence. We also show that the derived objective function that yields minimizing the Pearson$\chi ^2$divergence performs better than the classical one of using least squares for classification. There are two benefits of LSGANs over regular GANs. First, LSGANs are able to generate higher quality images than regular GANs. Second, LSGANs perform more stably during the learning process. For evaluating the image quality, we conduct both qualitative and quantitative experiments, and the experimental results show that LSGANs can generate higher quality images than regular GANs. Furthermore, we evaluate the stability of LSGANs in two groups. One is to compare between LSGANs and regular GANs without gradient penalty. We conduct three experiments, including Gaussian mixture distribution, difficult architectures, and a newly proposed method — datasets with small variability, to illustrate the stability of LSGANs. The other one is to compare between LSGANs with gradient penalty (LSGANs-GP) and WGANs with gradient penalty (WGANs-GP). The experimental results show that LSGANs-GP succeed in training for all the difficult architectures used in WGANs-GP, including 101-layer ResNet.
Xudong Mao, Qing Li 0001, Haoran Xie 0001, Raymond Y. K. Lau, Zhen Wang 0004, Stephen Paul Smolley
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Unpaired Multi-Domain Image Generation via Regularized Conditional GANs
abstract
In this paper, we study the problem of multi-domain image generation, the goal of which is to generate pairs of corresponding images from different domains. With the recent development in generative models, image generation has achieved great progress and has been applied to various computer vision tasks. However, multi-domain image generation may not achieve the desired performance due to the difficulty of learning the correspondence of different domain images, especially when the information of paired samples is not given. To tackle this problem, we propose Regularized Conditional GAN (RegCGAN) which is capable of learning to generate corresponding images in the absence of paired training data. RegCGAN is based on the conditional GAN, and we introduce two regularizers to guide the model to learn the corresponding semantics of different domains. We evaluate the proposed model on several tasks for which paired training data is not given, including the generation of edges and photos, the generation of faces with different attributes, etc. The experimental results show that our model can successfully generate corresponding images for all these tasks, while outperforms the baseline methods. We also introduce an approach of applying RegCGAN to unsupervised domain adaptation.
Xudong Mao, Qing Li 0001
IJCAI1
2017 Least Squares Generative Adversarial Networks
abstract
Unsupervised learning with generative adversarial networks (GANs) has proven hugely successful. Regular GANs hypothesize the discriminator as a classifier with the sigmoid cross entropy loss function. However, we found that this loss function may lead to the vanishing gradients problem during the learning process. To overcome such a problem, we propose in this paper the Least Squares Generative Adversarial Networks (LSGANs) which adopt the least squares loss function for the discriminator. We show that minimizing the objective function of LSGAN yields minimizing the Pearson X2 divergence. There are two benefits of LSGANs over regular GANs. First, LSGANs are able to generate higher quality images than regular GANs. Second, LSGANs perform more stable during the learning process. We evaluate LSGANs on LSUN and CIFAR-10 datasets and the experimental results show that the images generated by LSGANs are of better quality than the ones generated by regular GANs. We also conduct two comparison experiments between LSGANs and regular GANs to illustrate the stability of LSGANs.
Xudong Mao, Qing Li 0001, Haoran Xie 0001, Raymond Y. K. Lau, Zhen Wang 0004, Stephen Paul Smolley
ICCV1
2017 Recent advances in semantic computing and personalization
Haoran Xie 0001, Fu Lee Wang, Xudong Mao, Ke Li 0001, Qing Li 0001, Handing Wang
Neurocomputing3
2014 Popularity Tendency Analysis of Ranking-Oriented Collaborative Filtering from the Perspective of Loss Function
Xudong Mao, Qing Li 0001, Haoran Xie 0001, Yanghui Rao
DASFAA (1)1
2014 Mining Latent User Community for Tag-Based and Content-Based Search in Social Media
abstract
In recent years, there has been a proliferation of collaborative tagging systems in Web 2.0 communities. With the increasingly large amount of social data, how to manage and organize them becomes an important and crucial problem for folksonomy applications. To better understand and meet users’ needs, multimedia resources can be organized or indexed from these user perspectives; it is thus important to find latent user communities for social media applications. In this paper, we propose the mechanism of augmented folksonomy graph (AFG) to incorporate multi-faceted relations in social media, along with a novel density-based clustering method to discover latent user community from AFG by combining contents and tags of multimedia resources. To evaluate the proposed method, we conduct experiments on a public dataset, the empirical results of which show that our approach outperforms baseline ones in terms of tag-based and content-based personalized search.
Haoran Xie 0001, Qing Li 0001, Xudong Mao, Xiaodong Li 0007, Yi Cai 0001, Qianru Zheng
Comput. J.3
2014 Sentiment topic models for social emotion mining
Yanghui Rao, Qing Li 0001, Xudong Mao, Wenyin Liu
Inf. Sci.3
2014 Community-aware user profile enrichment in folksonomy
Haoran Xie 0001, Qing Li 0001, Xudong Mao, Xiaodong Li 0007, Yi Cai 0001, Yanghui Rao
Neural Networks3
2013 Finding Dominating Set from Verbal Contextual Graph for Personalized Search in Folksonomy
abstract
With the development of the Internet, user-generated data has been growing tremendously in Web 2.0 era. Facing such a big volume of resources in folksonomy, people need a method of fast exploration and indexing to find their demanded data. To achieve this goal, contextual information is indispensable and valuable to understand user preference and purpose. In sociolinguistics, context can be mainly categorized as verbal context and social context. Comparing with verbal context, social context not only requires domain knowledge to pre-define contextual attributes but also acquires additional data from users. However, there is no research of addressing irrelevant contextual factors for verbal context model so far. The dominating set from verbal context proposed in this paper is to fill this blank. We present the verbal context in folksonomy to capture the user intention, and propose a dominating set discovering method for this verbal context model to prune the irrelevant contextual factors and keep the major characteristics at the same time. Furthermore, the experiments, which are conducted on a public data set, show that the proposed method gives convincing results.
Haoran Xie 0001, Jingsheng Lei, Qing Li 0001, Xiaodong Li 0007, Xudong Mao, Yanghui Rao
Web Intelligence6
2012 Context-Aware Personalized Search Based on User and Resource Profiles in Folksonomies
Haoran Xie 0001, Qing Li 0001, Xudong Mao
APWeb3