Liqun Chen 0001

dblp:22/150-1 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
8since 2021 · last 2024
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
22 papers
Representation and self-supervised learning · 20% Vision and language · 15% Generative modeling · 14%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
2.342023
Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning · CVPR 2023
Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective · NeurIPS 2022
Vision-Language Pre-Training with Triple Contrastive Learning · CVPR 2022
Computer vision › Vision and language
cross-modal alignment
1.632022
Vision-Language Pre-Training with Triple Contrastive Learning · CVPR 2022
Multi-modal Alignment using Representation Codebook · CVPR 2022
Graph Optimal Transport for Cross-Domain Alignment · ICML 2020
Machine learning › Generative modeling
variational autoencoder
1.542020
Graph-Driven Generative Models for Heterogeneous Multi-Task Learning · AAAI 2020
Improving Textual Network Learning with Variational Homophilic Embeddings · NeurIPS 2019
Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models · ACL (1) 2019
Machine learning › Generative modeling
generative adversarial network
1.342019
Variational Annealing of GANs: A Langevin Perspective · ICML 2019
Adversarial Text Generation via Feature-Mover's Distance · NeurIPS 2018
Chi-square Generative Adversarial Network · ICML 2018
Machine learning › Graph learning
network embedding
1.232020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Improving Textual Network Learning with Variational Homophilic Embeddings · NeurIPS 2019
Improving Textual Network Embedding with Global Attention via Optimal Transport · ACL (1) 2019
Computer vision › Vision and language › multimodal representation
vision-language representation learning
1.122022
Vision-Language Pre-Training with Triple Contrastive Learning · CVPR 2022
Multi-modal Alignment using Representation Codebook · CVPR 2022
Natural language and speech › Language models and text generation
text generation
1.132020
Improving Text Generation with Student-Forcing Optimal Transport · EMNLP (1) 2020
Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models · ACL (1) 2019
Adversarial Text Generation via Feature-Mover's Distance · NeurIPS 2018
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.822023
Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning · CVPR 2023
Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective · NeurIPS 2022
Mathematical optimization
optimal transport
0.822020
Graph Optimal Transport for Cross-Domain Alignment · ICML 2020
Improving Sequence-to-Sequence Learning via Optimal Transport · ICLR (Poster) 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.722019
Variational Annealing of GANs: A Langevin Perspective · ICML 2019
Variational Inference and Model Selection with Generalized Evidence Bounds · ICML 2018
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
modality alignment
0.712023
Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning · CVPR 2023
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.712023
Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning · CVPR 2023
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.622017
Adversarial Symmetric Variational Autoencoder · NIPS 2017
ALICE: Towards Understanding Adversarial Learning for Joint Distribution Matching · NIPS 2017
Machine learning › Representation and self-supervised learning › vector quantization
representation codebook
0.612022
Multi-modal Alignment using Representation Codebook · CVPR 2022
Machine learning › Optimization for machine learning
stochastic optimization
0.612022
Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective · NeurIPS 2022
Computer vision › Vision and language
vision-language pretraining
0.612022
Vision-Language Pre-Training with Triple Contrastive Learning · CVPR 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.512021
Wasserstein Contrastive Representation Distillation · CVPR 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Wasserstein Contrastive Representation Distillation · CVPR 2021
Machine learning › Transfer learning and domain adaptation › knowledge transfer › representation transfer
representation distillation
0.512021
Wasserstein Contrastive Representation Distillation · CVPR 2021
Machine learning › Graph learning
dynamic graph learning
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Graph learning › graph neural network
graph convolutional network
0.412020
Graph-Driven Generative Models for Heterogeneous Multi-Task Learning · AAAI 2020
Machine learning › Graph learning
graph neural network
0.412020
Graph-Driven Generative Models for Heterogeneous Multi-Task Learning · AAAI 2020
Machine learning › Graph learning
graph representation learning
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Generative modeling › generative model
multi-task generative modeling
0.412020
Graph-Driven Generative Models for Heterogeneous Multi-Task Learning · AAAI 2020
Machine learning › Deep learning architectures and training › regularization
optimal transport regularization
0.412020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020
Natural language and speech › Language models and text generation › text generation
reinforcement learning for text generation
0.412020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.412020
Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning · AAAI 2020
Machine learning › Transfer learning and domain adaptation
optimal transport alignment
0.412019
Improving Textual Network Embedding with Global Attention via Optimal Transport · ACL (1) 2019
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning
0.412019
Improving Sequence-to-Sequence Learning via Optimal Transport · ICLR (Poster) 2019
Natural language and speech › Language models and text generation › text generation › synthetic text generation
adversarial text generation
0.312018
Adversarial Text Generation via Feature-Mover's Distance · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

optimal transport · 2.3wasserstein distance · 1.4contrastive learning · 1.1variational inference · 1.1mutual information maximization · 1.1importance weighting · 0.7geometric consistency loss · 0.7deep feature separation loss · 0.7contrastive loss · 0.7brownian-bridge loss · 0.7variational autoencoder · 0.4gromov-wasserstein distance · 0.4graph convolutional network · 0.4attention mechanism · 0.4sequence-to-sequence model · 0.4homophilic prior · 0.4generative model · 0.4
YearPublicationVenuePosition
2024 Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2
abstract
Text generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations.
Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin
IEEE Trans. Neural Networks Learn. Syst.6
2023 Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning
abstract
Contrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly match each other in the latent space. Yet it remains an open question how the modality alignment affects the downstream task performance. In this paper, based on an information-theoretic argument, we first prove that exact modality alignment is sub-optimal in general for down-stream prediction tasks. Hence we advocate that the key of better performance lies in meaningful latent modality structures instead of perfect modality alignment. To this end, we propose three general approaches to construct latent modality structures. Specifically, we design 1) a deep feature separation loss for intra-modality regularization; 2) a Brownian-bridge loss for inter-modality regularization; and 3) a geometric consistency loss for both intra- and intermodality regularization. Extensive experiments are conducted on two popular multi-modal representation learning frameworks: the CLIP-based two-tower model and the ALBEF-based fusion model. We test our model on a variety of tasks including zero/few-shot image classification, image-text retrieval, visual question answering, visual reasoning, and visual entailment. Our method achieves consistent improvements over existing methods, demonstrating the effectiveness and generalizability of our proposed approach on latent modality structure regularization.
Changyou Chen, Han Zhao 0002, Liqun Chen 0001, Qing Ping, Son Dinh Tran, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi
CVPR4
2022 Multi-modal Alignment using Representation Codebook
abstract
Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusion. Since image and text typically reside in different regions of the feature space, directly aligning them at instance level is challenging especially when features are still evolving during training. In this paper, we propose to align at a higher and more stable level using cluster representation. Specifically, we treat image and text as two “views” of the same entity, and encode them into a joint vision-language coding space spanned by a dictionary of cluster centers (codebook). We contrast positive and negative samples via their cluster assignments while simultaneously optimizing the cluster centers. To further smooth out the learning process, we adopt a teacher-student distillation paradigm, where the momentum teacher of one view guides the student learning of the other. We evaluated our approach on common vision language benchmarks and obtain new SoTA on zero-shot cross modality retrieval while being competitive on various other transfer tasks.
Jiali Duan, Liqun Chen 0001, Son Tran, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi
CVPR2
2022 Vision-Language Pre-Training with Triple Contrastive Learning
abstract
Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply performing cross-modal alignment (CMA) ignores data potential within each modality, which may result in degraded representations. For instance, although CMA-based models are able to map image-text pairs close together in the embedding space, they fail to ensure that similar inputs from the same modality stay close by. This problem can get even worse when the pre-training data is noisy. In this paper, we propose triple contrastive learning (TCL) for vision-language pre-training by leveraging both cross-modal and intra-modal self-supervision. Besides CMA, TCL introduces an intra-modal contrastive objective to provide complementary benefits in representation learning. To take advantage of localized and structural information from image and text input, TCL further maximizes the average MI between local regions of image/text and their global summary. To the best of our knowledge, ours is the first work that takes into account local structure information for multi-modality representation learning. Experimental evaluations show that our approach is competitive and achieves the new state of the art on various common downstream vision-language tasks such as image-text retrieval and visual question answering.
Jiali Duan, Son Tran, Yi Xu 0011, Sampath Chanda, Liqun Chen 0001, Belinda Zeng, Trishul Chilimbi, Junzhou Huang
CVPR6
2022 Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective
abstract
Contrastive learning (CL) has been the de facto technique for self-supervised representation learning (SSL), with impressive empirical success such as multi-modal representation learning. However, traditional CL loss only considers negative samples from a minibatch, which could cause biased gradients due to the non-decomposibility of the loss. For the first time, we consider optimizing a more generalized contrastive loss, where each data sample is associated with an infinite number of negative samples. We show that directly using minibatch stochastic optimization could lead to gradient bias. To remedy this, we propose an efficient Bayesian data augmentation technique to augment the contrastive loss into a decomposable one, where standard stochastic optimization can be directly applied without gradient bias. Specifically, our augmented loss defines a joint distribution over the model parameters and the augmented parameters, which can be conveniently optimized by a proposed stochastic expectation-maximization algorithm. Our framework is more general and is related to several popular SSL algorithms. We verify our framework on both small scale models and several large foundation models, including SSL of ImageNet and SSL for vision-language representation learning. Experiment results indicate the existence of gradient bias in all cases, and demonstrate the effectiveness of the proposed method on improving previous state of the arts. Remarkably, our method can outperform the strong MoCo-v3 under the same hyper-parameter setting with only around half of the minibatch size; and also obtains strong results in the recent public benchmark ELEVATER for few-shot image classification.
Changyou Chen, Yi Xu 0011, Liqun Chen 0001, Jiali Duan, Yiran Chen 0001, Son Tran, Belinda Zeng, Trishul Chilimbi
NeurIPS4
2021 Wasserstein Contrastive Representation Distillation
abstract
The primary goal of knowledge distillation (KD) is to encapsulate the information of a model learned from a teacher network into a student network, with the latter being more compact than the former. Existing work, e.g., using Kullback-Leibler divergence for distillation, may fail to capture important structural knowledge in the teacher network and often lacks the ability for feature generalization, particularly in situations when teacher and student are built to address different classification tasks. We propose Wasserstein Contrastive Representation Distillation (WCoRD), which leverages both primal and dual forms of Wasserstein distance for KD. The dual form is used for global knowledge transfer, yielding a contrastive learning objective that maximizes the lower bound of mutual information between the teacher and the student networks. The primal form is used for local contrastive knowledge transfer within a mini-batch, effectively matching the distributions of features between the teacher and the student networks. Experiments demonstrate that the proposed WCoRD method outperforms state-of-the-art approaches on privileged information distillation, model compression and cross-modal transfer.
Liqun Chen 0001, Dong Wang 0037, Zhe Gan, Jingjing Liu 0001, Ricardo Henao, Lawrence Carin
CVPR1
2021 Contextualized Perturbation for Textual Adversarial Attack
abstract
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dianqi Li, Yizhe Zhang 0002, Hao Peng 0009, Liqun Chen 0001, Chris Brockett, Ming-Ting Sun, William B. Dolan
NAACL-HLT4
2021 SpanPredict: Extraction of Predictive Document Spans with Neural Attention
abstract
Vivek Subramanian, Matthew Engelhard, Sam Berchuck, Liqun Chen, Ricardo Henao, Lawrence Carin. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Vivek Subramanian, Matthew Engelhard, Samuel Berchuck, Liqun Chen 0001, Ricardo Henao, Lawrence Carin
NAACL-HLT4
2020 Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning
abstract
Reinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions.
Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin
AAAI1
2020 Dynamic Embedding on Textual Networks via a Gaussian Process
abstract
Textual network embedding aims to learn low-dimensional representations of text-annotated nodes in a graph. Prior work in this area has typically focused on fixed graph structures; however, real-world networks are often dynamic. We address this challenge with a novel end-to-end node-embedding model, called Dynamic Embedding for Textual Networks with a Gaussian Process (DetGP). After training, DetGP can be applied efficiently to dynamic graphs without re-training or backpropagation. The learned representation of each node is a combination of textual and structural embeddings. Because the structure is allowed to be dynamic, our method uses the Gaussian process to take advantage of its non-parametric properties. To use both local and global graph structures, diffusion is used to model multiple hops between neighbors. The relative importance of global versus local structure for the embeddings is learned automatically. With the non-parametric nature of the Gaussian process, updating the embeddings for a changed graph structure requires only a forward pass through the learned model. Considering link prediction and node classification, experiments demonstrate the empirical effectiveness of our method compared to baseline approaches. We further show that DetGP can be straightforwardly and efficiently applied to dynamic textual networks.
Pengyu Cheng, Yitong Li 0001, Xinyuan Zhang 0001, Liqun Chen 0001, David E. Carlson, Lawrence Carin
AAAI4
2020 Graph-Driven Generative Models for Heterogeneous Multi-Task Learning
abstract
We propose a novel graph-driven generative model, that unifies multiple heterogeneous learning tasks into the same framework. The proposed model is based on the fact that heterogeneous learning tasks, which correspond to different generative processes, often rely on data with a shared graph structure. Accordingly, our model combines a graph convolutional network (GCN) with multiple variational autoencoders, thus embedding the nodes of the graph (i.e., samples for the tasks) in a uniform manner, while specializing their organization and usage to different tasks. With a focus on healthcare applications (tasks), including clinical topic modeling, procedure recommendation and admission-type prediction, we demonstrate that our method successfully leverages information across different tasks, boosting performance in all tasks and outperforming existing state-of-the-art approaches.
Wenlin Wang, Hongteng Xu, Zhe Gan, Bai Li 0001, Guoyin Wang 0002, Liqun Chen 0001, Qian Yang 0003, Wenqi Wang 0001, Lawrence Carin
AAAI6
2020 Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin
BMVC3
2020 Improving Text Generation with Student-Forcing Optimal Transport
abstract
Jianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuhchen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jianqiao Li, Chunyuan Li, Guoyin Wang 0002, Hao Fu 0002, Yuh-Chen Lin, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Ruiyi Zhang 0002, Wenlin Wang, Dinghan Shen, Qian Yang 0003, Lawrence Carin
EMNLP (1)6
2020 Graph Optimal Transport for Cross-Domain Alignment
abstract
Cross-domain alignment between two sets of entities (e.g., objects in an image, words in a sentence) is fundamental to both computer vision and natural language processing. Existing methods mainly focus on designing advanced attention mechanisms to simulate soft alignment, where no training signals are provided to explicitly encourage alignment. Plus, the learned attention matrices are often dense and difficult to interpret. We propose Graph Optimal Transport (GOT), a principled framework that builds upon recent advances in Optimal Transport (OT). In GOT, cross-domain alignment is formulated as a graph matching problem, by representing entities as a dynamically-constructed graph. Two types of OT distances are considered: (i) Wasserstein distance (WD) for node (entity) matching; and (ii) Gromov-Wasserstein distance (GWD) for edge (structure) matching. Both WD and GWD can be incorporated into existing neural network models, effectively acting as a drop-in regularizer. The inferred transport plan also yields sparse and self-normalized alignment, enhancing the interpretability of the learned model. Experiments show consistent outperformance of GOT over baselines across a wide range of tasks, including image-text retrieval, visual question answering, image captioning, machine translation, and text summarization.
Liqun Chen 0001, Zhe Gan, Yu Cheng 0001, Lawrence Carin, Jingjing Liu 0001
ICML1
2019 Improving Textual Network Embedding with Global Attention via Optimal Transport
abstract
Liqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Liqun Chen 0001, Guoyin Wang 0002, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang 0001, Wenlin Wang, Yizhe Zhang 0002, Lawrence Carin
ACL (1)1
2019 Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models
abstract
Variational autoencoders (VAEs) have received much attention recently as an end-toend architecture for text generation with latent variables.However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the posterior collapse issue has been widely identified in text-VAEs.In this paper, we propose to leverage several multi-level structures to learn a VAE model for generating long, and coherent text.In particular, a hierarchy of stochastic layers between the encoder and decoder networks is employed to abstract more informative and semantic-rich latent codes.Besides, we utilize a multi-level decoder structure to capture the coherent long-term structure inherent in long-form texts, by generating intermediate sentence representations as highlevel plan vectors.Extensive experimental results demonstrate that the proposed multi-level VAE model produces more coherent and less repetitive long text compared to baselines as well as can mitigate the posterior-collapse issue.
Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang 0002, Liqun Chen 0001, Xin Wang 0061, Jianfeng Gao 0001, Lawrence Carin
ACL (1)4
2019 Improving Sequence-to-Sequence Learning via Optimal Transport
Liqun Chen 0001, Yizhe Zhang 0002, Ruiyi Zhang 0002, Chenyang Tao, Zhe Gan, Bai Li 0001, Dinghan Shen, Changyou Chen, Lawrence Carin
ICLR (Poster)1
2019 Variational Annealing of GANs: A Langevin Perspective
abstract
The generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to improve GAN training. To enrich the understanding of this fast-growing yet almost exclusively heuristic-driven subject, we elucidate the theoretical roots of some of the empirical attempts to stabilize and improve GAN training with the introduction of likelihoods. We highlight new insights from variational theory of diffusion processes to derive a likelihood-based regularizing scheme for GAN training, and present a novel approach to train GANs with an unnormalized distribution instead of empirical samples. To substantiate our claims, we provide experimental evidence on how our theoretically-inspired new algorithms improve upon current practice.
Chenyang Tao, Shuyang Dai, Liqun Chen 0001, Ke Bai 0001, Junya Chen, Chang Liu 0030, Ruiyi Zhang 0002, Georgiy V. Bobashev, Lawrence Carin
ICML3
2019 On Fenchel Mini-Max Learning
abstract
Inference, estimation, sampling and likelihood evaluation are four primary goals of probabilistic modeling. Practical considerations often force modeling approaches to make compromises between these objectives. We present a novel probabilistic learning framework, called Fenchel Mini-Max Learning (FML), that accommodates all four desiderata in a flexible and scalable manner. Our derivation is rooted in classical maximum likelihood estimation, and it overcomes a longstanding challenge that prevents unbiased estimation of unnormalized statistical models. By reformulating MLE as a mini-max game, FML enjoys an unbiased training objective that (i) does not explicitly involve the intractable normalizing constant and (ii) is directly amendable to stochastic gradient descent optimization. To demonstrate the utility of the proposed approach, we consider learning unnormalized statistical models, nonparametric density estimation and training generative models, with encouraging empirical results presented.
Chenyang Tao, Liqun Chen 0001, Shuyang Dai, Junya Chen, Ke Bai 0001, Dong Wang 0037, Jianfeng Feng, Wenlian Lu, Georgiy V. Bobashev, Lawrence Carin
NeurIPS2
2019 Improving Textual Network Learning with Variational Homophilic Embeddings
abstract
The performance of many network learning applications crucially hinges on the success of network embedding algorithms, which aim to encode rich network information into low-dimensional vertex-based vector representations. This paper considers a novel variational formulation of network embeddings, with special focus on textual networks. Different from most existing methods that optimize a discriminative objective, we introduce Variational Homophilic Embedding (VHE), a fully generative model that learns network embeddings by modeling the semantic (textual) information with a variational autoencoder, while accounting for the structural (topology) information through a novel homophilic prior design. Homophilic vertex embeddings encourage similar embedding vectors for related (connected) vertices. The VHE encourages better generalization for downstream tasks, robustness to incomplete observations, and the ability to generalize to unseen vertices. Extensive experiments on real-world networks, for multiple tasks, demonstrate that the proposed method achieves consistently superior performance relative to competing state-of-the-art approaches.
Wenlin Wang, Chenyang Tao, Zhe Gan, Guoyin Wang 0002, Liqun Chen 0001, Xinyuan Zhang 0001, Ruiyi Zhang 0002, Qian Yang 0003, Ricardo Henao, Lawrence Carin
NeurIPS5
2018 Symmetric Variational Autoencoder and Connections to Adversarial Learning
abstract
A new form of the variational autoencoder (VAE) is proposed, based on the symmetric Kullback- Leibler divergence. It is demonstrated that learn- ing of the resulting symmetric VAE (sVAE) has close connections to previously developed adversarial-learning methods. This relationship helps unify the previously distinct techniques of VAE and adversarially learning, and provides insights that allow us to ameliorate shortcomings with some previously developed adversarial methods. In addition to an analysis that motivates and explains the sVAE, an extensive set of experiments validate the utility of the approach.
Liqun Chen 0001, Shuyang Dai, Yunchen Pu, Erjin Zhou, Chunyuan Li, Qinliang Su, Changyou Chen, Lawrence Carin
AISTATS1
2018 Variational Inference and Model Selection with Generalized Evidence Bounds
abstract
Recent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the conventional variational bound. It naturally subsumes the importance-weighted and Renyi bounds as special cases, and it is provably sharper than these counterparts. We also present an improved estimator for variational learning, and advocate a novel high signal-to-variance ratio update rule for the variational parameters. We discuss model-selection issues associated with existing evidence-lower-bound-based variational inference procedures, and show how to leverage the flexibility of our new formulation to address them. Empirical evidence is provided to validate our claims.
Liqun Chen 0001, Chenyang Tao, Ruiyi Zhang 0002, Ricardo Henao, Lawrence Carin
ICML1
2018 Chi-square Generative Adversarial Network
abstract
To assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and Hilbert space discrepancy metrics. We elucidate the theoretical connections between these three popular GAN training criteria and propose a novel procedure, called $\chi^2$ (Chi-square) GAN, that is conceptually simple, stable at training and resistant to mode collapse. Our procedure naturally generalizes to address the problem of simultaneous matching of multiple distributions. Further, we propose a resampling strategy that significantly improves sample quality, by repurposing the trained critic function via an importance weighting mechanism. Experiments show that the proposed procedure improves stability and convergence, and yields state-of-art results on a wide range of generative modeling tasks.
Chenyang Tao, Liqun Chen 0001, Ricardo Henao, Jianfeng Feng, Lawrence Carin
ICML2
2018 Adversarial Text Generation via Feature-Mover's Distance
abstract
Generative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose to improve text-generation GAN via a novel approach inspired by optimal transport. Specifically, we consider matching the latent feature distributions of real and synthetic sentences using a novel metric, termed the feature-mover's distance (FMD). This formulation leads to a highly discriminative critic and easy-to-optimize objective, overcoming the mode-collapsing and brittle-training problems in existing methods. Extensive experiments are conducted on a variety of tasks to evaluate the proposed model empirically, including unconditional text generation, style transfer from non-parallel text, and unsupervised cipher cracking. The proposed model yields superior performance, demonstrating wide applicability and effectiveness.
Liqun Chen 0001, Shuyang Dai, Chenyang Tao, Zhe Gan, Dinghan Shen, Yizhe Zhang 0002, Guoyin Wang 0002, Ruiyi Zhang 0002, Lawrence Carin
NeurIPS1
2018 A Unified Particle-Optimization Framework for Scalable Bayesian Sampling
Changyou Chen, Ruiyi Zhang 0002, Wenlin Wang, Bai Li 0001, Liqun Chen 0001
UAI5
2017 Triangle Generative Adversarial Networks
abstract
A Triangle Generative Adversarial Network ($\Delta$-GAN) is developed for semi-supervised cross-domain joint distribution matching, where the training data consists of samples from each domain, and supervision of domain correspondence is provided by only a few paired samples. $\Delta$-GAN consists of four neural networks, two generators and two discriminators. The generators are designed to learn the two-way conditional distributions between the two domains, while the discriminators implicitly define a ternary discriminative function, which is trained to distinguish real data pairs and two kinds of fake data pairs. The generators and discriminators are trained together using adversarial learning. Under mild assumptions, in theory the joint distributions characterized by the two generators concentrate to the data distribution. In experiments, three different kinds of domain pairs are considered, image-label, image-image and image-attribute pairs. Experiments on semi-supervised image classification, image-to-image translation and attribute-based image generation demonstrate the superiority of the proposed approach.
Zhe Gan, Liqun Chen 0001, Weiyao Wang 0002, Yunchen Pu, Yizhe Zhang 0002, Hao Liu 0015, Chunyuan Li, Lawrence Carin
NIPS2
2017 ALICE: Towards Understanding Adversarial Learning for Joint Distribution Matching
abstract
We investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable matched joint distributions for unsupervised and supervised tasks. We unify a broad family of adversarial models as joint distribution matching problems. Our approach stabilizes learning of unsupervised bidirectional adversarial learning methods. Further, we introduce an extension for semi-supervised learning tasks. Theoretical results are validated in synthetic data and real-world applications.
Chunyuan Li, Hao Liu 0015, Changyou Chen, Yunchen Pu, Liqun Chen 0001, Ricardo Henao, Lawrence Carin
NIPS5
2017 Adversarial Symmetric Variational Autoencoder
abstract
A new form of variational autoencoder (VAE) is developed, in which the joint distribution of data and codes is considered in two (symmetric) forms: (i) from observed data fed through the encoder to yield codes, and (ii) from latent codes drawn from a simple prior and propagated through the decoder to manifest data. Lower bounds are learned for marginal log-likelihood fits observed data and latent codes. When learning with the variational bound, one seeks to minimize the symmetric Kullback-Leibler divergence of joint density functions from (i) and (ii), while simultaneously seeking to maximize the two marginal log-likelihoods. To facilitate learning, a new form of adversarial training is developed. An extensive set of experiments is performed, in which we demonstrate state-of-the-art data reconstruction and generation on several image benchmarks datasets.
Yunchen Pu, Weiyao Wang 0002, Ricardo Henao, Liqun Chen 0001, Zhe Gan, Chunyuan Li, Lawrence Carin
NIPS4