Dinghan Shen

dblp:202/2287 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
4since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
27 papers
Language models and text generation · 23% Generative modeling · 16% Representation and self-supervised learning · 11%
Databases, data mining, and information retrieval
4 papers
Information retrieval · 100%

Topics — the 30 heaviest of 60, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
text generation
2.062020
Improving Text Generation with Student-Forcing Optimal Transport · EMNLP (1) 2020
Improving Adversarial Text Generation by Modeling the Distant Future · ACL 2020
Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models · ACL (1) 2019
Machine learning › Generative modeling
variational autoencoder
1.952020
Generative Semantic Hashing Enhanced via Boltzmann Machines · ACL 2020
Syntax-Infused Variational Autoencoder for Text Generation · ACL (1) 2019
Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models · ACL (1) 2019
Machine learning › Graph learning
network embedding
1.032019
Improving Textual Network Embedding with Global Attention via Optimal Transport · ACL (1) 2019
Diffusion Maps for Textual Network Embedding · NeurIPS 2018
Improved Semantic-Aware Network Embedding with Fine-Grained Word Alignment · EMNLP 2018
Machine learning › Deep learning architectures and training
data augmentation
1.022021
CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding · ICLR 2021
HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
text classification
1.032018
Learning Context-Aware Convolutional Filters for Text Processing · EMNLP 2018
Joint Embedding of Words and Labels for Text Classification · ACL (1) 2018
Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms · ACL (1) 2018
Machine learning › Generative modeling
generative adversarial network
0.932018
Adversarial Text Generation via Feature-Mover's Distance · NeurIPS 2018
Video Generation From Text · AAAI 2018
Adversarial Feature Matching for Text Generation · ICML 2017
Computer vision › Vision and language
cross-modal grounding
0.922021
Vision-Language Navigation Policy Learning and Adaptation · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation · CVPR 2019
Machine learning › Reinforcement learning
imitation learning
0.922021
Vision-Language Navigation Policy Learning and Adaptation · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation · CVPR 2019
Computer vision › Vision and language
vision-and-language navigation
0.922021
Vision-Language Navigation Policy Learning and Adaptation · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation · CVPR 2019
Information retrieval › similarity search
semantic hashing
0.822020
Generative Semantic Hashing Enhanced via Boltzmann Machines · ACL 2020
NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing · ACL (1) 2018
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.722019
Learning Compressed Sentence Representations for On-Device Text Processing · ACL (1) 2019
Deconvolutional Latent-Variable Model for Text Sequence Matching · AAAI 2018
Natural language and speech › Language models and text generation
natural language understanding
0.722021
CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding · ICLR 2021
HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text generation › synthetic text generation
adversarial text generation
0.622018
Adversarial Text Generation via Feature-Mover's Distance · NeurIPS 2018
Adversarial Feature Matching for Text Generation · ICML 2017
Machine learning › Representation and self-supervised learning › contrastive learning
contrastive data augmentation
0.512021
CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding · ICLR 2021
Robotics › Robot navigation and mapping
embodied navigation
0.512021
Vision-Language Navigation Policy Learning and Adaptation · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Natural language and speech › Language models and text generation
instruction following
0.512021
Vision-Language Navigation Policy Learning and Adaptation · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.512021
MixKD: Towards Efficient Distillation of Large-scale Language Models · ICLR 2021
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
LLM distillation
0.512021
MixKD: Towards Efficient Distillation of Large-scale Language Models · ICLR 2021
Machine learning › Reinforcement learning
policy learning
0.512021
Vision-Language Navigation Policy Learning and Adaptation · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Machine learning › Generative modeling › synthetic data generation
text data augmentation
0.512021
HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text generation
content planning
0.412020
Improving Adversarial Text Generation by Modeling the Distant Future · ACL 2020
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.412020
Improving Disentangled Text Representation Learning with Information-Theoretic Guidance · ACL 2020
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
0.422020
Deconvolutional Paragraph Representation Learning · NIPS 2017
Improving Disentangled Text Representation Learning with Information-Theoretic Guidance · ACL 2020
Machine learning › Transfer learning and domain adaptation
optimal transport alignment
0.412019
Improving Textual Network Embedding with Global Attention via Optimal Transport · ACL (1) 2019
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.412019
An End-to-End Generative Architecture for Paraphrase Generation · EMNLP/IJCNLP (1) 2019
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning
0.412019
Improving Sequence-to-Sequence Learning via Optimal Transport · ICLR (Poster) 2019
Natural language and speech › Language models and text generation › text generation › constrained text generation
syntax-guided text generation
0.412019
Syntax-Infused Variational Autoencoder for Text Generation · ACL (1) 2019
Mathematical optimization
optimal transport
0.412019
Improving Sequence-to-Sequence Learning via Optimal Transport · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.312018
Learning Context-Aware Convolutional Filters for Text Processing · EMNLP 2018
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
deconvolutional network
0.312018
Deconvolutional Latent-Variable Model for Text Sequence Matching · AAAI 2018

Methods — techniques the papers use, named apart from their topics

optimal transport · 1.9mixup · 1.0self-supervised imitation learning · 0.9matching critic · 0.9variational autoencoder · 0.7word embeddings · 0.7deconvolutional network · 0.6knowledge distillation · 0.5hidden-cut sampling · 0.5contrastive learning · 0.5variational inference · 0.4reparameterization · 0.4boltzmann machine · 0.4sequence-to-sequence model · 0.4neural architecture search · 0.3generative modeling · 0.3diffusion map · 0.3diffusion convolution · 0.3
YearPublicationVenuePosition
2021 HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability
abstract
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang
ACL/IJCNLP (1)2
2021 MixKD: Towards Efficient Distillation of Large-scale Language Models
Kevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 0001, Weizhu Chen, Changyou Chen, Lawrence Carin
ICLR3
2021 CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding
Yanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev, Weizhu Chen, Jiawei Han 0001
ICLR2
2021 Vision-Language Navigation Policy Learning and Adaptation
abstract
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms baseline methods by 10 percent on Success Rate weighted by Path Length (SPL) and achieves the state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore and adapt to unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7 to 11.7 percent).
Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Improving Disentangled Text Representation Learning with Information-Theoretic Guidance
abstract
Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang, Yitong Li, Lawrence Carin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang 0002, Yitong Li 0001, Lawrence Carin
ACL3
2020 Improving Adversarial Text Generation by Modeling the Distant Future
abstract
Auto-regressive text generation models usually focus on local fluency, and may cause inconsistent semantic meaning in long text generation.Further, automatically generating words with similar semantics is challenging, and hand-crafted linguistic rules are difficult to apply.We consider a text planning scheme and present a model-based imitation-learning approach to alleviate the aforementioned issues.Specifically, we propose a novel guider network to focus on the generative process over a longer horizon, which can assist next-word prediction and provide intermediate rewards for generator optimization.Extensive experiments demonstrate that the proposed method leads to improved performance.
Ruiyi Zhang 0002, Changyou Chen, Zhe Gan, Wenlin Wang, Dinghan Shen, Guoyin Wang 0002, Lawrence Carin
ACL5
2020 Generative Semantic Hashing Enhanced via Boltzmann Machines
abstract
Generative semantic hashing is a promising technique for large-scale information retrieval thanks to its fast retrieval speed and small memory footprint.For the tractability of training, existing generative-hashing methods mostly assume a factorized form for the posterior distribution, enforcing independence among the bits of hash codes.From the perspectives of both model representation and code space size, independence is always not the best assumption.In this paper, to introduce correlations among the bits of hash codes, we propose to employ the distribution of Boltzmann machine as the variational posterior.To address the intractability issue of training, we first develop an approximate method to reparameterize the distribution of a Boltzmann machine by augmenting it as a hierarchical concatenation of a Gaussian-like distribution and a Bernoulli distribution.Based on that, an asymptotically-exact lower bound is further derived for the evidence lower bound (ELBO).With these novel techniques, the entire model can be optimized efficiently.Extensive experimental results demonstrate that by effectively modeling correlations among different bits within a hash code, our model can achieve significant performance gains.
Qinliang Su, Dinghan Shen, Changyou Chen
ACL3
2020 Improving Text Generation with Student-Forcing Optimal Transport
abstract
Jianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuhchen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jianqiao Li, Chunyuan Li, Guoyin Wang 0002, Hao Fu 0002, Yuh-Chen Lin, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Ruiyi Zhang 0002, Wenlin Wang, Dinghan Shen, Qian Yang 0003, Lawrence Carin
EMNLP (1)11
2019 Improving Textual Network Embedding with Global Attention via Optimal Transport
abstract
Liqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Liqun Chen 0001, Guoyin Wang 0002, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang 0001, Wenlin Wang, Yizhe Zhang 0002, Lawrence Carin
ACL (1)4
2019 Learning Compressed Sentence Representations for On-Device Text Processing
abstract
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang, Qian Yang, Meng Tang, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang 0001, Qian Yang 0003, Asli Celikyilmaz, Lawrence Carin
ACL (1)1
2019 Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models
abstract
Variational autoencoders (VAEs) have received much attention recently as an end-toend architecture for text generation with latent variables.However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the posterior collapse issue has been widely identified in text-VAEs.In this paper, we propose to leverage several multi-level structures to learn a VAE model for generating long, and coherent text.In particular, a hierarchy of stochastic layers between the encoder and decoder networks is employed to abstract more informative and semantic-rich latent codes.Besides, we utilize a multi-level decoder structure to capture the coherent long-term structure inherent in long-form texts, by generating intermediate sentence representations as highlevel plan vectors.Extensive experimental results demonstrate that the proposed multi-level VAE model produces more coherent and less repetitive long text compared to baselines as well as can mitigate the posterior-collapse issue.
Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang 0002, Liqun Chen 0001, Xin Wang 0061, Jianfeng Gao 0001, Lawrence Carin
ACL (1)1
2019 Syntax-Infused Variational Autoencoder for Text Generation
abstract
We present a syntax-infused variational autoencoder (SIVAE), that integrates sentences with their syntactic trees to improve the grammar of generated sentences.Distinct from existing VAE-based text generative models, SIVAE contains two separate latent spaces, for sentences and syntactic trees.The evidence lower bound objective is redesigned correspondingly, by optimizing a joint distribution that accommodates two encoders and two decoders.SIVAE works with long shortterm memory architectures to simultaneously generate sentences and syntactic trees.Two versions of SIVAE are proposed: one captures the dependencies between the latent variables through a conditional prior network, and the other treats the latent variables independently such that syntactically-controlled sentence generation can be performed.Experimental results demonstrate the generative superiority of SIVAE on both reconstruction and targeted syntactic evaluations.Finally, we show that the proposed models can be used for unsupervised paraphrasing given different syntactic tree templates.
Xinyuan Zhang 0001, Yi Yang 0038, Siyang Yuan, Dinghan Shen, Lawrence Carin
ACL (1)4
2019 Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
abstract
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieves the new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%).
Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001
CVPR5
2019 Document Hashing with Mixture-Prior Generative Models
abstract
Wei Dong, Qinliang Su, Dinghan Shen, Changyou Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qinliang Su, Dinghan Shen, Changyou Chen
EMNLP/IJCNLP (1)3
2019 An End-to-End Generative Architecture for Paraphrase Generation
abstract
Qian Yang, Zhouyuan Huo, Dinghan Shen, Yong Cheng, Wenlin Wang, Guoyin Wang, Lawrence Carin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qian Yang 0003, Zhouyuan Huo, Dinghan Shen, Yong Cheng 0003, Wenlin Wang, Guoyin Wang 0002, Lawrence Carin
EMNLP/IJCNLP (1)3
2019 Improving Sequence-to-Sequence Learning via Optimal Transport
Liqun Chen 0001, Yizhe Zhang 0002, Ruiyi Zhang 0002, Chenyang Tao, Zhe Gan, Bai Li 0001, Dinghan Shen, Changyou Chen, Lawrence Carin
ICLR (Poster)8
2018 Video Generation From Text
abstract
Generating videos from text has proven to be a significant challenge for existing generative models. We tackle this problem by training a conditional generative model to extract both static and dynamic information from text. This is manifested in a hybrid framework, employing a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN). The static features, called "gist," are used to sketch text-conditioned background color and object layout structure. Dynamic features are considered by transforming input text into an image filter. To obtain a large amount of data for training the deep-learning model, we develop a method to automatically create a matched text-video corpus from publicly available online videos. Experimental results show that the proposed framework generates plausible and diverse short-duration smooth videos, while accurately reflecting the input text information. It significantly outperforms baseline models that directly adapt text-to-image generation procedures to produce videos. Performance is evaluated both visually and by adapting the inception score used to evaluate image generation in GANs.
Yitong Li 0001, Martin Renqiang Min, Dinghan Shen, David E. Carlson, Lawrence Carin
AAAI3
2018 Deconvolutional Latent-Variable Model for Text Sequence Matching
abstract
A latent-variable model is introduced for text matching, inferring sentence representations by jointly optimizing generative and discriminative objectives. To alleviate typical optimization challenges in latent-variable models for text, we employ deconvolutional networks as the sequence decoder (generator), providing learned latent codes with more semantic information and better generalization. Our model, trained in an unsupervised manner, yields stronger empirical predictive performance than a decoder based on Long Short-Term Memory (LSTM), with less parameters and considerably faster training. Further, we apply it to text sequence-matching problems. The proposed model significantly outperforms several strong sentence-encoding baselines, especially in the semi-supervised setting.
Dinghan Shen, Yizhe Zhang 0002, Ricardo Henao, Qinliang Su, Lawrence Carin
AAAI1
2018 NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing
abstract
Dinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Dinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin
ACL (1)1
2018 Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms
abstract
Dinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyuan Li, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Dinghan Shen, Guoyin Wang 0002, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang 0002, Chunyuan Li, Ricardo Henao, Lawrence Carin
ACL (1)1
2018 Joint Embedding of Words and Labels for Text Classification
abstract
Guoyin Wang, Chunyuan Li, Wenlin Wang, Yizhe Zhang, Dinghan Shen, Xinyuan Zhang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Guoyin Wang 0002, Chunyuan Li, Wenlin Wang, Yizhe Zhang 0002, Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin
ACL (1)5
2018 Topic Compositional Neural Language Model
abstract
We propose a Topic Compositional Neural Language Model (TCNLM), a novel method designed to simultaneously capture both the global semantic meaning and the local word-ordering structure in a document. The TCNLM learns the global semantic coherence of a document via a neural topic model, and the probability of each learned latent topic is further used to build a Mixture-of-Experts (MoE) language model, where each expert (corresponding to one topic) is a recurrent neural network (RNN) that accounts for learning the local structure of a word sequence. In order to train the MoE model efficiently, a matrix factorization method is applied, by extending each weight matrix of the RNN to be an ensemble of topic-dependent weight matrices. The degree to which each member of the ensemble is used is tied to the document-dependent probability of the corresponding topics. Experimental results on several corpora show that the proposed approach outperforms both a pure RNN-based model and other topic-guided language models. Further, our model yields sensible topics, and also has the capacity to generate meaningful sentences conditioned on given topics.
Wenlin Wang, Zhe Gan, Wenqi Wang 0001, Dinghan Shen, Jiaji Huang, Wei Ping, Sanjeev Satheesh, Lawrence Carin
AISTATS4
2018 Learning Context-Aware Convolutional Filters for Text Processing
abstract
Convolutional neural networks (CNNs) have recently emerged as a popular building block for natural language processing (NLP).Despite their success, most existing CNN models employed in NLP share the same learned (and static) set of filters for all input sentences.In this paper, we consider an approach of using a small meta network to learn contextaware convolutional filters for text processing.The role of meta network is to abstract the contextual information of a sentence or document into a set of input-aware filters.We further generalize this framework to model sentence pairs, where a bidirectional filter generation mechanism is introduced to encapsulate co-dependent sentence representations.In our benchmarks on four different tasks, including ontology classification, sentiment analysis, answer sentence selection, and paraphrase identification, our proposed model, a modified CNN with context-aware filters, consistently outperforms the standard CNN and attentionbased CNN baselines.By visualizing the learned context-aware filters, we further validate and rationalize the effectiveness of proposed framework.
Dinghan Shen, Martin Renqiang Min, Yitong Li 0001, Lawrence Carin
EMNLP1
2018 Improved Semantic-Aware Network Embedding with Fine-Grained Word Alignment
abstract
Network embeddings, which learn lowdimensional representations for each vertex in a large-scale network, have received considerable attention in recent years.For a wide range of applications, vertices in a network are typically accompanied by rich textual information such as user profiles, paper abstracts, etc.We propose to incorporate semantic features into network embeddings by matching important words between text sequences for all pairs of vertices.We introduce a word-by-word alignment framework that measures the compatibility of embeddings between word pairs, and then adaptively accumulates these alignment features with a simple yet effective aggregation function.In experiments, we evaluate the proposed framework on three real-world benchmarks for downstream tasks, including link prediction and multi-label vertex classification.Results demonstrate that our model outperforms state-of-the-art network embedding methods by a large margin.
Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin
EMNLP1
2018 Adversarial Text Generation via Feature-Mover's Distance
abstract
Generative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose to improve text-generation GAN via a novel approach inspired by optimal transport. Specifically, we consider matching the latent feature distributions of real and synthetic sentences using a novel metric, termed the feature-mover's distance (FMD). This formulation leads to a highly discriminative critic and easy-to-optimize objective, overcoming the mode-collapsing and brittle-training problems in existing methods. Extensive experiments are conducted on a variety of tasks to evaluate the proposed model empirically, including unconditional text generation, style transfer from non-parallel text, and unsupervised cipher cracking. The proposed model yields superior performance, demonstrating wide applicability and effectiveness.
Liqun Chen 0001, Shuyang Dai, Chenyang Tao, Zhe Gan, Dinghan Shen, Yizhe Zhang 0002, Guoyin Wang 0002, Ruiyi Zhang 0002, Lawrence Carin
NeurIPS6
2018 Diffusion Maps for Textual Network Embedding
abstract
Textual network embedding leverages rich text information associated with the network to learn low-dimensional vectorial representations of vertices. Rather than using typical natural language processing (NLP) approaches, recent research exploits the relationship of texts on the same edge to graphically embed text. However, these models neglect to measure the complete level of connectivity between any two texts in the graph. We present diffusion maps for textual network embedding (DMTE), integrating global structural information of the graph to capture the semantic relatedness between texts, with a diffusion-convolution operation applied on the text inputs. In addition, a new objective function is designed to efficiently preserve the high-order proximity using the graph diffusion. Experimental results show that the proposed approach outperforms state-of-the-art methods on the vertex-classification and link-prediction tasks.
Xinyuan Zhang 0001, Yitong Li 0001, Dinghan Shen, Lawrence Carin
NeurIPS3
2018 Parametric t-Distributed Stochastic Exemplar-Centered Embedding
Martin Renqiang Min, Dinghan Shen
ECML/PKDD (1)3
2017 Adversarial Feature Matching for Text Generation
abstract
The Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We employ a long short-term memory network as generator, and a convolutional network as discriminator. Instead of using the standard objective of GAN, we propose matching the high-dimensional latent feature distributions of real and synthetic sentences, via a kernelized discrepancy metric. This eases adversarial training by alleviating the mode-collapsing problem. Our experiments show superior performance in quantitative evaluation, and demonstrate that our model can generate realistic-looking sentences.
Yizhe Zhang 0002, Zhe Gan, Kai Fan 0002, Zhi Chen 0009, Ricardo Henao, Dinghan Shen, Lawrence Carin
ICML6
2017 Deconvolutional Paragraph Representation Learning
abstract
Learning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the length of the text. We propose a sequence-to-sequence, purely convolutional and deconvolutional autoencoding framework that is free of the above issue, while also being computationally efficient. The proposed method is simple, easy to implement and can be leveraged as a building block for many applications. We show empirically that compared to RNNs, our framework is better at reconstructing and correcting long paragraphs. Quantitative evaluation on semi-supervised text classification and summarization tasks demonstrate the potential for better utilization of long unlabeled text data.
Yizhe Zhang 0002, Dinghan Shen, Guoyin Wang 0002, Zhe Gan, Ricardo Henao, Lawrence Carin
NIPS2