Kevin Gimpel

dblp:47/1252 · DBLP profile ↗
← Back
61ranked-venue papers
10as first author
10since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021
YearPublicationVenuePosition
2024 Structured Tree Alignment for Evaluation of (Speech) Constituency Parsing
abstract
We present the structured average intersectionover-union ratio (STRUCT-IOU), a similarity metric between constituency parse trees motivated by the problem of evaluating speech parsers.STRUCT-IOU enables comparison between a constituency parse tree (over automatically recognized spoken word boundaries) with the ground-truth parse (over written words).To compute the metric, we project the groundtruth parse tree to the speech domain by forced alignment, align the projected ground-truth constituents with the predicted ones under certain structured constraints, and calculate the average IOU score across all aligned constituent pairs.STRUCT-IOU takes word boundaries into account and overcomes the challenge that the predicted words and ground truth may not have perfect one-to-one correspondence.Extending to the evaluation of text constituency parsing, we demonstrate that STRUCT-IOU can address token-mismatch issues, and shows higher tolerance to syntactically plausible parses than PARSEVAL (Black et al., 1991). 1
Freda Shi, Kevin Gimpel, Karen Livescu
ACL (1)2
2024 MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
abstract
It has been widely observed that exact or approximate MAP (mode-seeking) decoding from natural language generation (NLG) models consistently leads to degenerate outputs (Holtzman et al., 2019;Stahlberg and Byrne, 2019).Prior work has attributed this behavior to either a fundamental and unavoidable inadequacy of modes in probabilistic models or weaknesses in language modeling.Contrastingly, we argue that degenerate modes can even occur in the absence of any modeling error, due to contamination of the training data.Specifically, we argue that mixing even a tiny amount of low-entropy noise with a population text distribution can cause the data distribution's mode to become degenerate.We therefore propose to apply MAP decoding to the model's true conditional distribution where the conditioning variable explicitly avoids specific degenerate behavior.Using exact search, we empirically verify that the length-conditional modes of machine translation models and language models are indeed more fluent and topical than their unconditional modes.For the first time, we also share many examples of exact modal sequences from these models, and from several variants of the LLaMA-7B model.Notably, we observe that various kinds of degenerate modes persist, even at the scale of LLaMA-7B.Although we cannot tractably address these degeneracies with exact search, we perform a classifier-based approximate search on LLaMA-7B, a model which was not trained for instruction following, and find that we are able to elicit reasonable outputs without any finetuning.
Davis Yoshida, Kartik Goyal, Kevin Gimpel
ACL (1)3
2023 Audio-Visual Neural Syntax Acquisition
abstract
We study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce phrase structure using the inferred segment-level continuous representations. We present the Audio-Visual Neural Syntax Learner (AV-NSL) that learns phrase structure by listening to audio and looking at images, without ever being exposed to text. By training on paired images and spoken captions, AV-NSL exhibits the capability to infer meaningful phrase structures that are comparable to those derived by naturally-supervised text parsers, for both English and German. Our findings extend prior work in unsupervised language acquisition from speech and grounded grammar induction, and present one approach to bridge the gap between the two topics.
Cheng-I Lai, Freda Shi, Puyuan Peng, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David D. Cox, David F. Harwath, Yang Zhang 0001, Karen Livescu, James R. Glass
ASRU5
2023 The Benefits of Label-Description Training for Zero-Shot Text Classification
abstract
Pretrained language models have improved zero-shot text classification by allowing the transfer of semantic knowledge from the training data in order to classify among specific label sets in downstream tasks.We propose a simple way to further improve zero-shot accuracies with minimal effort.We curate small finetuning datasets intended to describe the labels for a task.Unlike typical finetuning data, which has texts annotated with labels, our data simply describes the labels in language, e.g., using a few related terms, dictionary/encyclopedia entries, and short templates.Across a range of topic and sentiment datasets, our method is more accurate than zero-shot by 17-19% absolute.It is also more robust to choices required for zero-shot classification, such as patterns for prompting the model to classify and mappings from labels to tokens in the model's vocabulary.Furthermore, since our data merely describes the labels but does not use input texts, finetuning on it yields a model that performs strongly on multiple text domains for a given label set, even improving over few-shot out-of-domain classification in multiple settings.
Lingyu Gao 0001, Debanjan Ghosh, Kevin Gimpel
EMNLP3
2022 Deep Clustering of Text Representations for Supervision-Free Probing of Syntax
abstract
We explore deep clustering of multilingual text representations for unsupervised model interpretation and induction of syntax. As these representations are high-dimensional, out-of-the-box methods like K-means do not work well. Thus, our approach jointly transforms the representations into a lower-dimensional cluster-friendly space and clusters them. We consider two notions of syntax: Part of Speech Induction (POSI) and Constituency Labelling (CoLab) in this work. Interestingly, we find that Multilingual BERT (mBERT) contains surprising amount of syntactic knowledge of English; possibly even as much as English BERT (E-BERT). Our model can be used as a supervision-free probe which is arguably a less-biased way of probing. We find that unsupervised probes show benefits from higher layers as compared to supervised probes. We further note that our unsupervised probe utilizes E-BERT and mBERT representations differently, especially for POSI. We validate the efficacy of our probe by demonstrating its capabilities as a unsupervised syntax induction technique. Our probe works well for both syntactic formalisms by simply adapting the input representations. We report competitive performance of our probe on 45-tag English POSI, state-of-the-art performance on 12-tag POSI across 10 languages, and competitive results on CoLab. We also perform zero-shot syntax induction on resource impoverished languages and report strong results.
Vikram Gupta, Freda Shi, Kevin Gimpel, Mrinmaya Sachan
AAAI3
2022 Chess as a Testbed for Language Model State Tracking
abstract
Transformer language models have made tremendous strides in natural language understanding tasks. However, the complexity of natural language makes it challenging to ascertain how accurately these models are tracking the world state underlying the text. Motivated by this issue, we consider the task of language modeling for the game of chess. Unlike natural language, chess notations describe a simple, constrained, and deterministic domain. Moreover, we observe that the appropriate choice of chess notation allows for directly probing the world state, without requiring any additional probing-related machinery. We find that: (a) With enough training data, transformer language models can learn to track pieces and predict legal moves with high accuracy when trained solely on move sequences. (b) For small training sets providing access to board state information during training can yield significant improvements. (c) The success of transformer language models is dependent on access to the entire game history i.e. “full attention”. Approximating this full attention results in a significant performance drop. We propose this testbed as a benchmark for future work on the development and analysis of transformer language models.
Shubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin Gimpel
AAAI4
2022 SummScreen: A Dataset for Abstractive Screenplay Summarization
abstract
We introduce SUMMSCREEN, a summarization dataset comprised of pairs of TV series transcripts and human written recaps.The dataset provides a challenging testbed for abstractive summarization for several reasons.Plot details are often expressed indirectly in character dialogues and may be scattered across the entirety of the transcript.These details must be found and integrated to form the succinct plot descriptions in the recaps.Also, TV scripts contain content that does not directly pertain to the central plot but rather serves to develop characters or provide comic relief.This information is rarely contained in recaps.Since characters are fundamental to TV series, we also propose two entity-centric evaluation metrics.Empirically, we characterize the dataset by evaluating several methods, including neural models and those based on nearest neighbors.An oracle extractive approach outperforms all benchmarked models according to automatic metrics, showing that the neural models are unable to fully exploit the input transcripts.Human evaluation and qualitative analysis reveal that our nonoracle models are competitive with their oracle counterparts in terms of generating faithful plot events and can benefit from better content selectors.Both oracle and non-oracle models generate unfaithful facts, suggesting future research directions.
Mingda Chen, Zewei Chu, Sam Wiseman, Kevin Gimpel
ACL (1)4
2022 Substructure Distribution Projection for Zero-Shot Cross-Lingual Dependency Parsing
abstract
We present substructure distribution projection (SUBDP), a technique that projects a distribution over structures in one domain to another, by projecting substructure distributions separately.Models for the target domain can then be trained, using the projected distributions as soft silver labels.We evaluate SUBDP on zeroshot cross-lingual dependency parsing, taking dependency arcs as substructures: we project the predicted dependency arc distributions in the source language(s) to target language(s), and train a target language parser on the resulting distributions.Given an English treebank as the only source of human supervision, SUBDP achieves better unlabeled attachment score than all prior work on the Universal Dependencies v2.2 (Nivre et al., 2020) test set across eight diverse target languages, as well as the best labeled attachment score on six languages.In addition, SUBDP improves zeroshot cross-lingual dependency parsing with very few (e.g., 50) supervised bitext pairs, across a broader range of target languages.
Freda Shi, Kevin Gimpel, Karen Livescu
ACL (1)2
2022 Moment Distributionally Robust Tree Structured Prediction
abstract
Structured prediction of tree-shaped objects is heavily studied under the name of syntactic dependency parsing. Current practice based on maximum likelihood or margin is either agnostic to or inconsistent with the evaluation loss. Risk minimization alleviates the discrepancy between training and test objectives but typically induces a non-convex problem. These approaches adopt explicit regularization to combat overfitting without probabilistic interpretation. We propose a moment-based distributionally robust optimization approach for tree structured prediction, where the worst-case expected loss over a set of distributions within bounded moment divergence from the empirical distribution is minimized. We develop efficient algorithms for arborescences and other variants of trees. We derive Fisher consistency, convergence rates and generalization bounds for our proposed method. We evaluate its empirical effectiveness on dependency parsing benchmarks.
Yeshu Li, Danyal Saeed, Brian D. Ziebart, Kevin Gimpel
NeurIPS5
2021 FlowPrior: Learning Expressive Priors for Latent Variable Sentence Models
abstract
Variational autoencoders (VAEs) are widely used for latent variable modeling of text.We focus on variations that learn expressive prior distributions over the latent variable.We find that existing training strategies are not effective for learning rich priors, so we add the importance-sampled log marginal likelihood as a second term to the standard VAE objective to help when learning the prior.Doing so improves results for all priors evaluated, including a novel choice for sentence VAEs based on normalizing flows (NF).Priors parameterized with NF are no longer constrained to a specific distribution family, allowing a more flexible way to encode the data distribution.Our model, which we call FlowPrior, shows a substantial improvement in language modeling tasks compared to strong baselines.We demonstrate that FlowPrior learns an expressive prior with analysis and several forms of evaluation involving generation.
Xiaoan Ding, Kevin Gimpel
NAACL-HLT2
2020 How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed Questions
abstract
We present a large-scale dataset for the task of rewriting an ill-formed natural language question to a well-formed one. Our multi-domain question rewriting (MQR) dataset is constructed from human contributed Stack Exchange question edit histories. The dataset contains 427,719 question pairs which come from 303 domains. We provide human annotations for a subset of the dataset as a quality estimate. When moving from ill-formed to well-formed questions, the question quality improves by an average of 45 points across three aspects. We train sequence-to-sequence neural models on the constructed dataset and obtain an improvement of 13.2% in BLEU-4 over baseline methods built from other data resources. We release the MQR dataset to encourage research on the problem of question rewriting.1
Zewei Chu, Mingda Chen, Miaosen Wang, Kevin Gimpel, Manaal Faruqui, Xiance Si
AAAI5
2020 PeTra: A Sparsely Supervised Memory Model for People Tracking
abstract
We propose PeTra, a memory-augmented neural network designed to track entities in its memory slots.PeTra is trained using sparse annotation from the GAP pronoun resolution dataset and outperforms a prior memory model on the task while using a simpler architecture.We empirically compare key modeling choices, finding that we can simplify several aspects of the design of the memory module while retaining strong performance.To measure the people tracking capability of memory models, we (a) propose a new diagnostic evaluation based on counting the number of unique entities in text, and (b) conduct a small scale human evaluation to compare evidence of people tracking in the memory logs of PeTra relative to a previous approach.PeTra is highly effective in both evaluations, demonstrating its ability to track people in its memory despite being trained with limited annotation.
Shubham Toshniwal, Allyson Ettinger, Kevin Gimpel, Karen Livescu
ACL3
2020 ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation
abstract
We propose to train a non-autoregressive machine translation model to minimize the energy defined by a pretrained autoregressive model.In particular, we view our non-autoregressive translation system as an inference network (Tu and Gimpel, 2018) trained to minimize the autoregressive teacher energy.This contrasts with the popular approach of training a non-autoregressive model on a distilled corpus consisting of the beam-searched outputs of such a teacher model.Our approach, which we call ENGINE (ENerGy-based Inference NEtworks), achieves state-of-the-art non-autoregressive results on the IWSLT 2014 DE-EN and WMT 2016 RO-EN datasets, approaching the performance of autoregressive models.
Lifu Tu, Richard Yuanzhe Pang, Sam Wiseman, Kevin Gimpel
ACL4
2020 Discriminatively-Tuned Generative Classifiers for Robust Natural Language Inference
abstract
While discriminative neural network classifiers are generally preferred, recent work has shown advantages of generative classifiers in term of data efficiency and robustness.In this paper, we focus on natural language inference (NLI).We propose GenNLI, a generative classifier for NLI tasks, and empirically characterize its performance by comparing it to five baselines, including discriminative models and large-scale pretrained language representation models like BERT.We explore training objectives for discriminative fine-tuning of our generative classifiers, showing improvements over log loss fine-tuning from prior work (Lewis and Fan, 2019).In particular, we find strong results with a simple unbounded modification to log loss, which we call the "infinilog loss".Our experiments show that GenNLI outperforms both discriminative and pretrained baselines across several challenging NLI experimental settings, including small training sets, imbalanced label distributions, and label noise.
Xiaoan Ding, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Kevin Gimpel
EMNLP (1)5
2020 On the Role of Supervision in Unsupervised Constituency Parsing
abstract
We analyze several recent unsupervised constituency parsing models, which are tuned with respect to the parsing F 1 score on the Wall Street Journal (WSJ) development set (1,700 sentences).We introduce strong baselines for them, by training an existing supervised parsing model (Kitaev and Klein, 2018) on the same labeled examples they access.When training on the 1,700 examples, or even when using only 50 examples for training and 5 for development, such a few-shot parsing approach can outperform all the unsupervised parsing methods by a significant margin.Fewshot parsing can be further improved by a simple data augmentation method and selftraining.This suggests that, in order to arrive at fair conclusions, we should carefully consider the amount of labeled data used for model development.We propose two protocols for future work on unsupervised parsing: (i) use fully unsupervised criteria for hyperparameter tuning and model selection; (ii) use as few labeled examples as possible for model development, and compare to few-shot parsing trained on the same labeled examples.1
Freda Shi, Karen Livescu, Kevin Gimpel
EMNLP (1)3
2020 Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks
abstract
Long document coreference resolution remains a challenging task due to the large memory and runtime requirements of current models.Recent work doing incremental coreference resolution using just the global representation of entities shows practical benefits but requires keeping all entities in memory, which can be impractical for long documents.We argue that keeping all entities in memory is unnecessary, and we propose a memoryaugmented neural network that tracks only a small bounded number of entities at a time, thus guaranteeing a linear runtime in length of document.We show that (a) the model remains competitive with models with high memory and computational requirements on OntoNotes and LitBank, and (b) the model learns an efficient memory management strategy easily outperforming a rule-based strategy.
Shubham Toshniwal, Sam Wiseman, Allyson Ettinger, Karen Livescu, Kevin Gimpel
EMNLP (1)5
2020 An Exploration of Arbitrary-Order Sequence Labeling via Energy-Based Inference Networks
abstract
Many tasks in natural language processing involve predicting structured outputs, e.g., sequence labeling, semantic role labeling, parsing, and machine translation.Researchers are increasingly applying deep representation learning to these problems, but the structured component of these approaches is usually quite simplistic.In this work, we propose several high-order energy terms to capture complex dependencies among labels in sequence labeling, including several that consider the entire label sequence.We use neural parameterizations for these energy terms, drawing from convolutional, recurrent, and selfattention networks.We use the framework of learning energy-based inference networks (Tu and Gimpel, 2018) for dealing with the difficulties of training and inference with such models.We empirically demonstrate that this approach achieves substantial improvement using a variety of high-order energy terms on four sequence labeling tasks, while having the same decoding speed as simple, local classifiers.We also find high-order energies to help in noisy data conditions. 1
Lifu Tu, Tianyu Liu 0001, Kevin Gimpel
EMNLP (1)3
2020 ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut
ICLR4
2019 Controllable Paraphrase Generation with a Syntactic Exemplar
abstract
Prior work on controllable text generation usually assumes that the controlled attribute can take on one of a small set of values known a priori.In this work, we propose a novel task, where the syntax of a generated sentence is controlled rather by a sentential exemplar.To evaluate quantitatively with standard metrics, we create a novel dataset with human annotations.We also develop a variational model with a neural module specifically designed for capturing syntactic knowledge and several multitask training objectives to promote disentangled representation learning.Empirically, the proposed model is observed to achieve improvements over baselines and learn to capture desirable characteristics.
Mingda Chen, Qingming Tang, Sam Wiseman, Kevin Gimpel
ACL (1)4
2019 Visually Grounded Neural Syntax Acquisition
abstract
We present the Visually Grounded Neural Syntax Learner (VG-NSL), an approach for learning syntactic representations and structures without explicit supervision.The model learns by looking at natural images and reading paired captions.VG-NSL generates constituency parse trees of texts, recursively composes representations for constituents, and matches them with images.We define the concreteness of constituents by their matching scores with images, and use it to guide the parsing of text.Experiments on the MSCOCO data set show that VG-NSL outperforms various unsupervised parsing approaches that do not use visual grounding, in terms of F 1 scores against gold parse trees.We find that VG-NSL is much more stable with respect to the choice of random initialization and the amount of training data.We also find that the concreteness acquired by VG-NSL correlates well with a similar measure defined by linguists.Finally, we also apply VG-NSL to multiple languages in the Multi30K data set, showing that our model consistently outperforms prior unsupervised approaches. 1
Freda Shi, Jiayuan Mao, Kevin Gimpel, Karen Livescu
ACL (1)3
2019 Beyond BLEU: Training Neural Machine Translation with Semantic Similarity
abstract
While most neural machine translation (NMT) systems are still trained using maximum likelihood estimation, recent work has demonstrated that optimizing systems to directly improve evaluation metrics such as BLEU can substantially improve final translation accuracy.However, training with BLEU has some limitations: it doesn't assign partial credit, it has a limited range of output values, and it can penalize semantically correct hypotheses if they differ lexically from the reference.In this paper, we introduce an alternative reward function for optimizing NMT systems that is based on recent work in semantic similarity.We evaluate on four disparate languages translated to English, and find that training with our proposed metric results in better translations as evaluated by BLEU, semantic similarity, and human evaluation, and also that the optimization procedure converges faster.Analysis suggests that this is because the proposed metric is more conducive to optimization, assigning partial credit and providing more diversity in scores than BLEU. 1
John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, Graham Neubig
ACL (1)3
2019 Simple and Effective Paraphrastic Similarity from Parallel Translations
abstract
We present a model and methodology for learning paraphrastic sentence embeddings directly from bitext, removing the timeconsuming intermediate step of creating paraphrase corpora.Further, we show that the resulting model can be applied to cross-lingual tasks where it both outperforms and is orders of magnitude faster than more complex stateof-the-art baselines.1
John Wieting, Kevin Gimpel, Graham Neubig, Taylor Berg-Kirkpatrick
ACL (1)2
2019 EntEval: A Holistic Evaluation Benchmark for Entity Representations
abstract
Mingda Chen, Zewei Chu, Yang Chen, Karl Stratos, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingda Chen, Zewei Chu, Karl Stratos, Kevin Gimpel
EMNLP/IJCNLP (1)5
2019 Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations
abstract
Mingda Chen, Zewei Chu, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingda Chen, Zewei Chu, Kevin Gimpel
EMNLP/IJCNLP (1)3
2019 Latent-Variable Generative Models for Data-Efficient Text Classification
abstract
Xiaoan Ding, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiaoan Ding, Kevin Gimpel
EMNLP/IJCNLP (1)2
2018 ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations
abstract
We describe PARANMT-50M, a dataset of more than 50 million English-English sentential paraphrase pairs.We generated the pairs automatically by using neural machine translation to translate the non-English side of a large parallel corpus, following Wieting et al. (2017).Our hope is that PARANMT-50M can be a valuable resource for paraphrase generation and can provide a rich source of semantic knowledge to improve downstream natural language understanding tasks.To show its utility, we use PARANMT-50M to train paraphrastic sentence embeddings that outperform all supervised systems on every SemEval semantic textual similarity competition, in addition to showing how it can be used for paraphrase generation. 1
John Wieting, Kevin Gimpel
ACL (1)2
2018 Variational Sequential Labelers for Semi-Supervised Learning
abstract
We introduce a family of multitask variational methods for semi-supervised sequence labeling.Our model family consists of a latentvariable generative model and a discriminative labeler.The generative models use latent variables to define the conditional probability of a word given its context, drawing inspiration from word prediction objectives commonly used in learning word embeddings.The labeler helps inject discriminative information into the latent space.We explore several latent variable configurations, including ones with hierarchical structure, which enables the model to account for both label-specific and word-specific information.Our models consistently outperform standard sequential baselines on 8 sequence labeling datasets, and improve further with unlabeled data.
Mingda Chen, Qingming Tang, Karen Livescu, Kevin Gimpel
EMNLP4
2018 A Study of All-Convolutional Encoders for Connectionist Temporal Classification
abstract
Connectionist temporal classification (CTC) is a popular sequence prediction approach for automatic speech recognition that is typically used with models based on recurrent neural networks (RNNs). We explore whether deep convolutional neural networks (CNNs) can be used effectively instead of RNNs as the “encoder” in CTC. CNNs lack an explicit representation of the entire sequence, but have the advantage that they are much faster to train. We present an exploration of CNN s as encoders for CTC models, in the context of character-based (lexicon-free) automatic speech recognition. In particular, we explore a range of one-dimensional convolutionallayers, which are particularly efficient. We compare the performance of our CNN-based models against typical RNN-based models in terms of training time, decoding time, model size and word error rate (WER) on the Switchboard Eva12000 corpus. We find that our CNN-based models are close in performance to LSTMs, while not matching them, and are much faster to train and decode.
Kalpesh Krishna, Liang Lu 0001, Kevin Gimpel, Karen Livescu
ICASSP3
2018 Learning Approximate Inference Networks for Structured Prediction
Lifu Tu, Kevin Gimpel
ICLR (Poster)2
2018 Adversarial Example Generation with Syntactically Controlled Paraphrase Networks
abstract
Mohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Mohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer
NAACL-HLT3
2018 Parsing Speech: a Neural Approach to Integrating Lexical and Acoustic-Prosodic Information
abstract
Trang Tran, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, Mari Ostendorf. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Trang Tran 0001, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, Mari Ostendorf
NAACL-HLT4
2018 Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
abstract
The growing importance of massive datasets with the advent of deep learning makes robustness to label noise a critical property for classifiers to have. Sources of label noise include automatic labeling for large datasets, non-expert labeling, and label corruption by data poisoning adversaries. In the latter case, corruptions may be arbitrarily bad, even so bad that a classifier predicts the wrong labels with high confidence. To protect against such sources of noise, we leverage the fact that a small set of clean labels is often easy to procure. We demonstrate that robustness to label noise up to severe strengths can be achieved by using a set of trusted data with clean labels, and propose a loss correction that utilizes trusted examples in a data-efficient manner to mitigate the effects of label noise on deep neural network classifiers. Across vision and natural language processing tasks, we experiment with various label noises at several strengths, and show that our method significantly outperforms existing methods.
Dan Hendrycks, Mantas Mazeika, Duncan Wilson, Kevin Gimpel
NeurIPS4
2017 Revisiting Recurrent Networks for Paraphrastic Sentence Embeddings
abstract
We consider the problem of learning general-purpose, paraphrastic sentence embeddings, revisiting the setting of Wieting et al. (2016b).While they found LSTM recurrent networks to underperform word averaging, we present several developments that together produce the opposite conclusion.These include training on sentence pairs rather than phrase pairs, averaging states to represent sequences, and regularizing aggressively.These improve LSTMs in both transfer learning and supervised settings.We also introduce a new recurrent architecture, the GATED RECURRENT AVER-AGING NETWORK, that is inspired by averaging and LSTMs while outperforming them both.We analyze our learned models, finding evidence of preferences for particular parts of speech and dependency relations.1
John Wieting, Kevin Gimpel
ACL (1)2
2017 Learning Paraphrastic Sentence Embeddings from Back-Translated Bitext
abstract
We consider the problem of learning general-purpose, paraphrastic sentence embeddings in the setting of Wieting et al. (2016b).We use neural machine translation to generate sentential paraphrases via back-translation of bilingual sentence pairs.We evaluate the paraphrase pairs by their ability to serve as training data for learning paraphrastic sentence embeddings.We find that the data quality is stronger than prior work based on bitext and on par with manually-written English paraphrase pairs, with the advantage that our approach can scale up to generate large training sets for many languages and domains.We experiment with several language pairs and data sources, and develop a variety of data filtering techniques.In the process, we explore how neural machine translation output differs from humanwritten sentences, finding clear differences in length, the amount of repetition, and the use of rare words. 1 1 Generated paraphrases and code are available at http: //ttic.uchicago.edu/˜wieting.R: We understand that has already commenced, but there is a long way to go.T: This situation has already commenced, but much still needs to be done.R: The restaurant is closed on Sundays.No breakfast is available on Sunday mornings
John Wieting, Jonathan Mallinson, Kevin Gimpel
EMNLP3
2017 A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
Dan Hendrycks, Kevin Gimpel
ICLR (Poster)2
2016 Commonsense Knowledge Base Completion
abstract
We enrich a curated resource of commonsense knowledge by formulating the problem as one of knowledge base completion (KBC). Most work in KBC focuses on knowledge bases like Freebase that relate entities drawn from a fixed set. However, the tuples in ConceptNet (Speer and Havasi, 2012) define relations between an unbounded set of phrases. We develop neural network models for scoring tuples on arbitrary phrases and evaluate them by their ability to distinguish true held-out tuples from false ones. We find strong performance from a bilinear model using a simple additive architecture to model phrases. We manually evaluate our trained model’s ability to assign quality scores to novel tuples, finding that it can propose tuples at the same quality level as mediumconfidence tuples from ConceptNet.
Xiang Li 0069, Aynaz Taheri, Lifu Tu, Kevin Gimpel
ACL (1)4
2016 Who did What: A Large-Scale Person-Centered Cloze Dataset
abstract
We have constructed a new "Who-did-What" dataset of over 200,000 fill-in-the-gap (cloze) multiple choice reading comprehension problems constructed from the LDC English Gigaword newswire corpus.The WDW dataset has a variety of novel features.First, in contrast with the CNN and Daily Mail datasets (Hermann et al., 2015) we avoid using article summaries for question formation.Instead, each problem is formed from two independent articles -an article given as the passage to be read and a separate article on the same events used to form the question.Second, we avoid anonymization -each choice is a person named entity.Third, the problems have been filtered to remove a fraction that are easily solved by simple baselines, while remaining 84% solvable by humans.We report performance benchmarks of standard systems and propose the WDW dataset as a challenge task for the community.1
Hai Wang 0013, Mohit Bansal, Kevin Gimpel, David A. McAllester
EMNLP4
2016 Charagram: Embedding Words and Sentences via Character n-grams
abstract
We present CHARAGRAM embeddings, a simple approach for learning character-based compositional models to embed textual sequences.A word or sentence is represented using a character n-gram count vector, followed by a single nonlinear transformation to yield a low-dimensional embedding.We use three tasks for evaluation: word similarity, sentence similarity, and part-of-speech tagging.We demonstrate that CHARAGRAM embeddings outperform more complex architectures based on character-level recurrent and convolutional neural networks, achieving new state-of-the-art performance on several similarity tasks. 1
John Wieting, Mohit Bansal, Kevin Gimpel, Karen Livescu
EMNLP3
2016 Efficient Segmental Cascades for Speech Recognition
abstract
Discriminative segmental models offer a way to incorporate flexible feature functions into speech recognition.However, their appeal has been limited by their computational requirements, due to the large number of possible segments to consider.Multi-pass cascades of segmental models introduce features of increasing complexity in different passes, where in each pass a segmental model rescores lattices produced by a previous (simpler) segmental model.In this paper, we explore several ways of making segmental cascades efficient and practical: reducing the feature set in the first pass, frame subsampling, and various pruning approaches.In experiments on phonetic recognition, we find that with a combination of such techniques, it is possible to maintain competitive performance while greatly reducing decoding, pruning, and training time.
Hao Tang 0002, Kevin Gimpel, Karen Livescu
INTERSPEECH3
2016 Constraints Based Convex Belief Propagation
abstract
Inference in Markov random fields subject to consistency structure is a fundamental problem that arises in many real-life applications. In order to enforce consistency, classical approaches utilize consistency potentials or encode constraints over feasible instances. Unfortunately this comes at the price of a serious computational bottleneck. In this paper we suggest to tackle consistency by incorporating constraints on beliefs. This permits derivation of a closed-form message-passing algorithm which we refer to as the Constraints Based Convex Belief Propagation (CBCBP). Experiments show that CBCBP outperforms the standard approach while being at least an order of magnitude faster.
Yaniv Tenzer, Alexander G. Schwing, Kevin Gimpel, Tamir Hazan
NIPS3
2016 End-to-end training approaches for discriminative segmental models
abstract
Recent work on discriminative segmental models has shown that they can achieve competitive speech recognition performance, using features based on deep neural frame classifiers. However, segmental models can be more challenging to train than standard frame-based approaches. While some segmental models have been successfully trained end to end, there is a lack of understanding of their training under different settings and with different losses.
Hao Tang 0002, Kevin Gimpel, Karen Livescu
SLT3
2015 Discriminative segmental cascades for feature-rich phone recognition
abstract
Discriminative segmental models, such as segmental conditional random fields (SCRFs) and segmental structured support vector machines (SSVMs), have had success in speech recognition via both lattice rescoring and first-pass decoding. However, such models suffer from slow decoding, hampering the use of computationally expensive features, such as segment neural networks or other high-order features. A typical solution is to use approximate decoding, either by beam pruning in a single pass or by beam pruning to generate a lattice followed by a second pass. In this work, we study discriminative segmental models trained with a hinge loss (i.e., segmental structured SVMs). We show that beam search is not suitable for learning rescoring models in this approach, though it gives good approximate decoding performance when the model is already well-trained. Instead, we consider an approach inspired by structured prediction cascades, which use max-marginal pruning to generate lattices. We obtain a high-accuracy phonetic recognition system with several expensive feature types: a segment neural network, a second-order language model, and second-order phone boundary features.
Hao Tang 0002, Kevin Gimpel, Karen Livescu
ASRU3
2015 Multi-Perspective Sentence Similarity Modeling with Convolutional Neural Networks
abstract
Modeling sentence similarity is complicated by the ambiguity and variability of linguistic expression.To cope with these challenges, we propose a model for comparing sentences that uses a multiplicity of perspectives.We first model each sentence using a convolutional neural network that extracts features at multiple levels of granularity and uses multiple types of pooling.We then compare our sentence representations at several granularities using multiple similarity metrics.We apply our model to three tasks, including the Microsoft Research paraphrase identification task and two SemEval semantic textual similarity tasks.We obtain strong performance on all tasks, rivaling or exceeding the state of the art without using external resources such as WordNet or parsers.
Kevin Gimpel, Jimmy Lin
EMNLP2
2015 Deep Multilingual Correlation for Improved Word Embeddings
abstract
Ang Lu, Weiran Wang, Mohit Bansal, Kevin Gimpel, Karen Livescu. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Ang Lu, Mohit Bansal, Kevin Gimpel, Karen Livescu
HLT-NAACL4
2015 A Sense-Topic Model for Word Sense Induction with Unsupervised Data Enrichment
abstract
Word sense induction (WSI) seeks to automatically discover the senses of a word in a corpus via unsupervised methods. We propose a sense-topic model for WSI, which treats sense and topic as two separate latent variables to be inferred jointly. Topics are informed by the entire document, while senses are informed by the local context surrounding the ambiguous word. We also discuss unsupervised ways of enriching the original corpus in order to improve model performance, including using neural word embeddings and external corpora to expand the context of each data instance. We demonstrate significant improvements over the previous state-of-the-art, achieving the best results reported to date on the SemEval-2013 WSI task.
Jing Wang 0102, Mohit Bansal, Kevin Gimpel, Brian D. Ziebart, Clement T. Yu
Trans. Assoc. Comput. Linguistics3
2015 From Paraphrase Database to Compositional Paraphrase Model and Back
abstract
The Paraphrase Database (PPDB; Ganitkevitch et al., 2013) is an extensive semantic resource, consisting of a list of phrase pairs with (heuristic) confidence estimates. However, it is still unclear how it can best be used, due to the heuristic nature of the confidences and its necessarily incomplete coverage. We propose models to leverage the phrase pairs from the PPDB to build parametric paraphrase models that score paraphrase pairs more accurately than the PPDB’s internal scores while simultaneously improving its coverage. They allow for learning phrase embeddings as well as improved word embeddings. Moreover, we introduce two new, manually annotated datasets to evaluate short-phrase paraphrasing models. Using our paraphrase model trained using PPDB, we achieve state-of-the-art results on standard word and bigram similarity tasks and beat strong baselines on our new short phrase paraphrase tasks.
John Wieting, Mohit Bansal, Kevin Gimpel, Karen Livescu
Trans. Assoc. Comput. Linguistics3
2014 Weakly-Supervised Learning with Cost-Augmented Contrastive Estimation
abstract
We generalize contrastive estimation in two ways that permit adding more knowledge to unsupervised learning.The first allows the modeler to specify not only the set of corrupted inputs for each observation, but also how bad each one is.The second allows specifying structural preferences on the latent variable used to explain the observations.They require setting additional hyperparameters, which can be problematic in unsupervised learning, so we investigate new methods for unsupervised model selection and system combination.We instantiate these ideas for part-of-speech induction without tag dictionaries, improving over contrastive estimation as well as strong benchmarks from the PASCAL 2012 shared task.
Kevin Gimpel, Mohit Bansal
EMNLP1
2014 A comparison of training approaches for discriminative segmental models
abstract
Segmental models such as segmental conditional random fields have had some recent success in lattice rescoring for speech recognition. They provide a flexible framework for incorpo-rating a wide range of features across different levels of units, such as phones and words. However, such models have mainly been trained by maximizing conditional likelihood, which may not be the best proxy for the task loss of speech recognition. In addition, there has been little work on designing cost func-tions as surrogates for the word error rate. In this paper, we investigate various losses and introduce a new cost function for training segmental models. We compare lattice rescoring results for multiple tasks and also study the impact of several choices required when optimizing these losses. Index Terms: speech recognition, segmental conditional ran-dom fields, empirical Bayes risk, large-margin training
Hao Tang 0002, Kevin Gimpel, Karen Livescu
INTERSPEECH2
2014 Phrase Dependency Machine Translation with Quasi-Synchronous Tree-to-Tree Features
abstract
Recent research has shown clear improvement in translation quality by exploiting linguistic syntax for either the source or target language. However, when using syntax for both languages (“tree-to-tree” translation), there is evidence that syntactic divergence can hamper the extraction of useful rules (Ding and Palmer 2005 ). Smith and Eisner ( 2006 ) introduced quasi-synchronous grammar, a formalism that treats non-isomorphic structure softly using features rather than hard constraints. Although a natural fit for translation modeling, its flexibility has proved challenging for building real-world systems. In this article, we present a tree-to-tree machine translation system inspired by quasi-synchronous grammar. The core of our approach is a new model that combines phrases and dependency syntax, integrating the advantages of phrase-based and syntax-based translation. We report statistically significant improvements over a phrase-based baseline on five of seven test sets across four language pairs. We also present encouraging preliminary results on the use of unsupervised dependency parsing for syntax-based machine translation.
Kevin Gimpel, Noah A. Smith
Comput. Linguistics1
2013 A Systematic Exploration of Diversity in Machine Translation
abstract
This paper addresses the problem of producing a diverse set of plausible translations.We present a simple procedure that can be used with any statistical machine translation (MT) system.We explore three ways of using diverse translations: (1) system combination, (2) discriminative reranking with rich features, and (3) a novel post-editing scenario in which multiple translations are presented to users.We find that diversity can improve performance on these tasks, especially for sentences that are difficult for MT.
Kevin Gimpel, Dhruv Batra, Chris Dyer, Gregory Shakhnarovich
EMNLP1
2013 Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
Olutobi Owoputi, Brendan T. O'Connor 0001, Chris Dyer, Kevin Gimpel, Nathan Schneider 0001, Noah A. Smith
HLT-NAACL4
2012 Word Salad: Relating Food Prices and Descriptions
Victor Chahuneau, Kevin Gimpel, Bryan R. Routledge, Lily Scherlis, Noah A. Smith
EMNLP-CoNLL2
2012 Structured Ramp Loss Minimization for Machine Translation
Kevin Gimpel, Noah A. Smith
HLT-NAACL1
2012 Concavity and Initialization for Unsupervised Dependency Parsing
Kevin Gimpel, Noah A. Smith
HLT-NAACL1
2011 Quasi-Synchronous Phrase Dependency Grammars for Machine Translation
Kevin Gimpel, Noah A. Smith
EMNLP1
2010 Distributed Asynchronous Online Learning for Natural Language Processing
Kevin Gimpel, Dipanjan Das 0001, Noah A. Smith
CoNLL1
2010 Softmax-Margin CRFs: Training Log-Linear Models with Cost Functions
Kevin Gimpel, Noah A. Smith
HLT-NAACL1
2010 Movie Reviews and Revenues: An Experiment in Text Regression
Mahesh Joshi, Dipanjan Das 0001, Kevin Gimpel, Noah A. Smith
HLT-NAACL3
2009 Cube Summing, Approximate Inference with Non-Local Features, and Dynamic Programming without Semirings
Kevin Gimpel, Noah A. Smith
EACL1
2009 Feature-Rich Translation by Quasi-Synchronous Lattice Parsing
Kevin Gimpel, Noah A. Smith
EMNLP1
2008 Logistic Normal Priors for Unsupervised Probabilistic Grammar Induction
abstract
We explore a new Bayesian model for probabilistic grammars, a family of distributions over discrete structures that includes hidden Markov models and probabilistic context-free grammars. Our model extends the correlated topic model framework to probabilistic grammars, exploiting the logistic normal distribution as a prior over the grammar parameters. We derive a variational EM algorithm for that model, and then experiment with the task of unsupervised grammar induction for natural language dependency parsing. We show that our model achieves superior results over previous models that use different priors.
Shay B. Cohen, Kevin Gimpel, Noah A. Smith
NIPS2