Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chris Dyer

dblp:41/6895 · also Christopher Dyer · DBLP profile ↗
← Back
124ranked-venue papers
10as first author
7since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 124 · 10 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
66 papers
Language models and text generation · 27% Information extraction and text analysis · 22% Machine translation · 10%
Theoretical computer science
2 papers
Automata and formal languages · 100%

Topics — the 30 heaviest of 134, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
syntactic parsing
1.242019
Text Genre and Training Data Size in Human-like Parsing · EMNLP/IJCNLP (1) 2019
Finding syntax in human encephalography with beam search · ACL (1) 2018
Ontology-Aware Token Embeddings for Prepositional Phrase Attachment · ACL (1) 2017
Natural language and speech › Language models and text generation
neural language model
1.032019
Learning to Discover, Ground and Use Words with Segmental Neural Language Models · ACL (1) 2019
On the State of the Art of Evaluation in Neural Language Models · ICLR (Poster) 2018
Learning to Create and Reuse Words in Open-Vocabulary Neural Language Modeling · ACL (1) 2017
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.942016
Distilling an Ensemble of Greedy Dependency Parsers into One MST Parser · EMNLP 2016
Training with Exploration Improves a Greedy Stack LSTM Parser · EMNLP 2016
Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs · EMNLP 2015
Machine learning › Representation and self-supervised learning
word representation
0.732015
Not All Contexts Are Created Equal: Better Word Representations with Variable Attention · EMNLP 2015
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation · EMNLP 2015
Sparse Overcomplete Word Vector Representations · ACL (1) 2015
Natural language and speech › Information extraction and text analysis
coreference resolution
0.622018
Syntactic Scaffolds for Semantic Structures · EMNLP 2018
Reference-Aware Language Models · EMNLP 2017
Natural language and speech › Language models and text generation
language modeling
0.622018
Fast Parametric Learning with Activation Memorization · ICML 2018
Reference-Aware Language Models · EMNLP 2017
Natural language and speech › Language models and text generation
decoding
0.622022
Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022
Efficient Minimum Error Rate Training and Minimum Bayes-Risk Decoding for Translation Hypergraphs and Lattices · ACL/IJCNLP 2009
Natural language and speech › Language models and text generation › decoding
adaptive decoding
0.612022
Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022
Machine learning › Generative modeling
energy-based model
0.612022
Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings · ICLR 2022
Natural language and speech › Language models and text generation
masked language modeling
0.612022
Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings · ICLR 2022
Natural language and speech › Language models and text generation
text generation
0.612022
Enabling Arbitrary Translation Objectives with Adaptive Tree Search · ICLR 2022
Computer vision › Vision and language
open-vocabulary models
0.522017
Learning to Create and Reuse Words in Open-Vocabulary Neural Language Modeling · ACL (1) 2017
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation · EMNLP 2015
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.532015
Learning Word Representations with Hierarchical Sparse Coding · ICML 2015
Evaluation of Word Vector Representations by Subspace Alignment · EMNLP 2015
Not All Contexts Are Created Equal: Better Word Representations with Variable Attention · EMNLP 2015
Natural language and speech › Machine translation
document-level machine translation
0.512021
Diverse Pretrained Context Encodings Improve Document Translation · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.512021
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering · NeurIPS 2021
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.512021
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering · NeurIPS 2021
Natural language and speech › Machine translation
statistical machine translation
0.542013
Translating into Morphologically Rich Languages with Synthetic Phrases · EMNLP 2013
Joint Feature Selection in Distributed Stochastic Learning for Large-Scale Discriminative Training in SMT · ACL (1) 2012
Discriminative Word Alignment with a Function Word Reordering Model · EMNLP 2010
Machine learning › Representation and self-supervised learning › word representation
word representation learning
0.522016
Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning · ACL (1) 2016
Learning Word Representations with Hierarchical Sparse Coding · ICML 2015
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.422016
Phonologically Aware Neural Model for Named Entity Recognition in Low Resource Transfer Settings · EMNLP 2016
Metaphor Detection with Cross-Lingual Model Transfer · ACL (1) 2014
Natural language and speech › Information extraction and text analysis
semantic parsing
0.422016
Semantic Parsing with Semi-Supervised Sequential Autoencoders · EMNLP 2016
A Discriminative Graph-Based Parser for the Abstract Meaning Representation · ACL (1) 2014
Computer vision › Video understanding and tracking
action segmentation
0.412020
Learning to Segment Actions from Observation and Narration · ACL 2020
Machine learning › Generative modeling › generative model
probabilistic generative model
0.412020
A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing · ACL 2020
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
transition-based dependency parsing
0.422015
Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs · EMNLP 2015
Transition-Based Dependency Parsing with Stack Long Short-Term Memory · ACL (1) 2015
Computer vision › Video understanding and tracking › action segmentation
weakly supervised action segmentation
0.412020
Learning to Segment Actions from Observation and Narration · ACL 2020
Multimedia analysis and retrieval › image analysis › image understanding
historical document analysis
0.412020
A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing · ACL 2020
Machine learning › Trustworthy machine learning › robustness
certified robustness
0.412019
Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation · EMNLP/IJCNLP (1) 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
collapsed variational inference
0.412019
Compound Probabilistic Context-Free Grammars for Grammar Induction · ACL (1) 2019
Machine learning › Trustworthy machine learning › robustness › certified robustness
interval bound propagation
0.412019
Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation · EMNLP/IJCNLP (1) 2019
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.412019
Scalable Syntax-Aware Language Models Using Knowledge Distillation · ACL (1) 2019
Machine learning › Trustworthy machine learning
robustness
0.412019
Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

expectation-maximization · 1.2latent variable model · 1.2LSTM · 0.8word embeddings · 0.8semi-supervised learning · 0.8backpropagation · 0.7tree search · 0.6metropolis-hastings · 0.6pre-trained language model · 0.5end-to-end differentiable training · 0.5context encoder · 0.5variational autoencoder · 0.4neural editor model · 0.4dynamic programming · 0.4collapsed variational inference · 0.4surprisal · 0.3sub-differentiable surrogate · 0.3parser action count · 0.3
YearPublicationVenuePosition
2023 Machine Learning for Ancient Languages: A Survey
abstract
Abstract Ancient languages preserve the cultures and histories of the past. However, their study is fraught with difficulties, and experts must tackle a range of challenging text-based tasks, from deciphering lost languages to restoring damaged inscriptions, to determining the authorship of works of literature. Technological aids have long supported the study of ancient texts, but in recent years advances in artificial intelligence and machine learning have enabled analyses on a scale and in a detail that are reshaping the field of humanities, similarly to how microscopes and telescopes have contributed to the realm of science. This article aims to provide a comprehensive survey of published research using machine learning for the study of ancient texts written in any language, script, and medium, spanning over three and a half millennia of civilizations around the ancient world. To analyze the relevant literature, we introduce a taxonomy of tasks inspired by the steps involved in the study of ancient documents: digitization, restoration, attribution, linguistic analysis, textual criticism, translation, and decipherment. This work offers three major contributions: first, mapping the interdisciplinary field carved out by the synergy between the humanities and machine learning; second, highlighting how active collaboration between specialists from both fields is key to producing impactful and compelling scholarship; third, highlighting promising directions for future work in this field. Thus, this work promotes and supports the continued collaborative impetus between the humanities and machine learning.
Thea Sommerschield, Yannis M. Assael, John Pavlopoulos, Vanessa Stefanak, Andrew W. Senior, Chris Dyer, John Bodel, Jonathan Prag, Ion Androutsopoulos, Nando de Freitas
Comput. Linguistics6
2022 Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings
Kartik Goyal, Chris Dyer, Taylor Berg-Kirkpatrick
ICLR2
2022 Enabling Arbitrary Translation Objectives with Adaptive Tree Search
Wang Ling, Wojciech Stokowiec, Domenic Donato, Chris Dyer, Lei Yu 0008, Laurent Sartran, Austin Matthews
ICLR4
2022 Transformer Grammars: Augmenting Transformer Language Models with Syntactic Inductive Biases at Scale
abstract
Abstract We introduce Transformer Grammars (TGs), a novel class of Transformer language models that combine (i) the expressive power, scalability, and strong performance of Transformers and (ii) recursive syntactic compositions, which here are implemented through a special attention mask and deterministic transformation of the linearized tree. We find that TGs outperform various strong baselines on sentence-level language modeling perplexity, as well as on multiple syntax-sensitive language modeling evaluation metrics. Additionally, we find that the recursive syntactic composition bottleneck which represents each sentence as a single vector harms perplexity on document-level language modeling, providing evidence that a different kind of memory mechanism—one that is independent of composed syntactic representations—plays an important role in current successful models of long text.
Laurent Sartran, Samuel Barrett, Adhiguna Kuncoro, Milos Stanojevic, Phil Blunsom, Chris Dyer
Trans. Assoc. Comput. Linguistics6
2021 Diverse Pretrained Context Encodings Improve Document Translation
abstract
Domenic Donato, Lei Yu, Chris Dyer. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Domenic Donato, Lei Yu 0008, Chris Dyer
ACL/IJCNLP (1)3
2021 Game-theoretic Vocabulary Selection via the Shapley Value and Banzhaf Index
abstract
Roma Patel, Marta Garnelo, Ian Gemp, Chris Dyer, Yoram Bachrach. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Roma Patel, Marta Garnelo, Ian Gemp, Chris Dyer, Yoram Bachrach
NAACL-HLT4
2021 End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering
abstract
We present an end-to-end differentiable training method for retrieval-augmented open-domain question answering systems that combine information from multiple retrieved documents when generating answers. We model retrieval decisions as latent variables over sets of relevant documents. Since marginalizing over sets of retrieved documents is computationally hard, we approximate this using an expectation-maximization algorithm. We iteratively estimate the value of our latent variable (the set of relevant documents for a given question) and then use this estimate to update the retriever and reader parameters. We hypothesize that such end-to-end training allows training signals to flow to the reader and then to the retriever better than staged-wise training. This results in a retriever that is able to select more relevant documents for a question and a reader that is trained on more accurate documents to generate an answer. Experiments on three benchmark datasets demonstrate that our proposed method outperforms all existing approaches of comparable size by 2-3% absolute exact match points, achieving new state-of-the-art results. Our results also demonstrate the feasibility of learning to retrieve to improve answer generation without explicit supervision of retrieval decisions.
Devendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer, Dani Yogatama
NeurIPS4
2020 Learning to Segment Actions from Observation and Narration
abstract
We apply a generative segmental model of task structure, guided by narration, to action segmentation in video.We focus on unsupervised and weakly-supervised settings where no action labels are known during training.Despite its simplicity, our model performs competitively with previous work on a dataset of naturalistic instructional videos.Our model allows us to vary the sources of supervision used in training, and we find that both task structure and narrative language provide large benefits in segmentation quality.
Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, Aida Nematzadeh
ACL4
2020 A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing
abstract
We propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents.We focus on clustering extracted glyph images into underlying templates in the presence of multiple confounding sources of variance.Our approach introduces a neural editor model that first generates well-understood printing phenomena like spatial perturbations from template parameters via interpertable latent variables, and then modifies the result by generating a non-interpretable latent vector responsible for inking variations, jitter, noise from the archiving process, and other unforeseen phenomena associated with Early Modern printing.Critically, by introducing an inference network whose input is restricted to the visual residual between the observation and the interpretably-modified template, we are able to control and isolate what the vector-valued latent variable captures.We show that our approach outperforms rigid interpretable clustering baselines (Ocular) and overly-flexible deep generative models (VAE) alike on the task of completely unsupervised discovery of typefaces in mixed-font documents.
Kartik Goyal, Chris Dyer, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick
ACL2
2020 Syntactic Structure Distillation Pretraining for Bidirectional Encoders
abstract
Textual representation learners trained on large amounts of data have achieved notable success on downstream tasks; intriguingly, they have also performed well on challenging tests of syntactic competence. Hence, it remains an open question whether scalable learners like BERT can become fully proficient in the syntax of natural language by virtue of data scale alone, or whether they still benefit from more explicit syntactic biases. To answer this question, we introduce a knowledge distillation strategy for injecting syntactic biases into BERT pretraining, by distilling the syntactically informative predictions of a hierarchical—albeit harder to scale—syntactic language model. Since BERT models masked words in bidirectional context, we propose to distill the approximate marginal distribution over words in context from the syntactic LM. Our approach reduces relative error by 2–21% on a diverse set of structured prediction tasks, although we obtain mixed results on the GLUE benchmark. Our findings demonstrate the benefits of syntactic biases, even for representation learners that exploit large amounts of data, and contribute to a better understanding of where syntactic biases are helpful in benchmarks of natural language understanding.
Adhiguna Kuncoro, Lingpeng Kong, Daniel Fried, Dani Yogatama, Laura Rimell, Chris Dyer, Phil Blunsom
Trans. Assoc. Comput. Linguistics6
2020 Better Document-Level Machine Translation with Bayes' Rule
abstract
We show that Bayes’ rule provides an effective mechanism for creating document translation models that can be learned from only parallel sentences and monolingual documents a compelling benefit because parallel documents are not always available. In our formulation, the posterior probability of a candidate translation is the product of the unconditional (prior) probability of the candidate output document and the “reverse translation probability” of translating the candidate output back into the source language. Our proposed model uses a powerful autoregressive language model as the prior on target language documents, but it assumes that each sentence is translated independently from the target to the source language. Crucially, at test time, when a source document is observed, the document language model prior induces dependencies between the translations of the source sentences in the posterior. The model’s independence assumption not only enables efficient use of available data, but it additionally admits a practical left-to-right beam-search algorithm for carrying out inference. Experiments show that our model benefits from using cross-sentence context in the language model, and it outperforms existing document translation approaches.
Lei Yu 0008, Laurent Sartran, Wojciech Stokowiec, Wang Ling, Lingpeng Kong, Phil Blunsom, Chris Dyer
Trans. Assoc. Comput. Linguistics7
2019 Learning to Discover, Ground and Use Words with Segmental Neural Language Models
abstract
We propose a segmental neural language model that combines the generalization power of neural networks with the ability to discover word-like units that are latent in unsegmented character sequences.In contrast to previous segmentation models that treat word segmentation as an isolated task, our model unifies word discovery, learning how words fit together to form sentences, and, by conditioning the model on visual context, how words' meanings ground in representations of nonlinguistic modalities.Experiments show that the unconditional model learns predictive distributions better than character LSTM models, discovers words competitively with nonparametric Bayesian word segmentation models, and that modeling language conditional on visual context improves performance on both.
Kazuya Kawakami, Chris Dyer, Phil Blunsom
ACL (1)2
2019 Compound Probabilistic Context-Free Grammars for Grammar Induction
abstract
We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar.In contrast to traditional formulations which learn a single stochastic grammar, our context-free rule probabilities are modulated by a per-sentence continuous latent variable, which induces marginal dependencies beyond the traditional context-free assumptions.Inference in this grammar is performed by collapsed variational inference, in which an amortized variational posterior is placed on the continuous variable, and the latent trees are marginalized with dynamic programming.Experiments on English and Chinese show the effectiveness of our approach compared to recent state-of-the-art methods for grammar induction from words with neural language models.
Chris Dyer, Alexander M. Rush
ACL (1)2
2019 Scalable Syntax-Aware Language Models Using Knowledge Distillation
abstract
Prior work has shown that, on small amounts of training data, syntactic neural language models learn structurally sensitive generalisations more successfully than sequential language models. However, their computational complexity renders scaling difficult, and it remains an open question whether structural biases are still necessary when sequential models have access to ever larger amounts of training data. To answer this question, we introduce an efficient knowledge distillation (KD) technique that transfers knowledge from a syntactic language model trained on a small corpus to an LSTM language model, hence enabling the LSTM to develop a more structurally sensitive representation of the larger training data it learns from. On targeted syntactic evaluations, we find that, while sequential LSTMs perform much better than previously reported, our proposed technique substantially improves on this baseline, yielding a new state of the art. Our findings and analysis affirm the importance of structural biases, even in models that learn from large amounts of data.
Adhiguna Kuncoro, Chris Dyer, Laura Rimell, Stephen Clark, Phil Blunsom
ACL (1)2
2019 Comparing Top-Down and Bottom-Up Neural Generative Dependency Models
abstract
Recurrent neural network grammars (RNNGs) generate sentences using phrase-structure syntax and perform very well in terms of both language modeling and parsing performance.However, since dependency annotations are much more readily available than phrase structure annotations, we propose two new generative models of projective dependency syntax, so as to explore whether generative dependency models are similarly effective.Both models use RNNs to represent the derivation history with making any explicit independence assumptions, but they differ in how they construct the trees: one builds the tree bottom up and the other top down, which profoundly changes the estimation problem faced by the learner.We evaluate the two models on three typologically different languages: English, Arabic, and Japanese.We find that both generative models improve parsing performance over a discriminative baseline, but, in contrast to RNNGs, they are significantly less effective than non-syntactic LSTM language models.Little difference between the tree construction orders is observed for either parsing or language modeling.
Austin Matthews, Graham Neubig, Chris Dyer
CoNLL3
2019 Text Genre and Training Data Size in Human-like Parsing
abstract
John Hale, Adhiguna Kuncoro, Keith Hall, Chris Dyer, Jonathan Brennan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
John T. Hale, Adhiguna Kuncoro, Keith Hall, Chris Dyer, Jonathan Brennan
EMNLP/IJCNLP (1)4
2019 Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation
abstract
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, Pushmeet Kohli. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, Pushmeet Kohli
EMNLP/IJCNLP (1)4
2018 A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models
abstract
Beam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do not directly consider the behaviour of the final decoding method. As a result, for cross-entropy trained models, beam decoding can sometimes yield reduced test performance when compared with greedy decoding. In order to train models that can more effectively make use of beam search, we propose a new training procedure that focuses on the final loss metric (e.g. Hamming loss) evaluated on the output of beam search. While well-defined, this "direct loss" objective is itself discontinuous and thus difficult to optimize. Hence, in our approach, we form a sub-differentiable surrogate objective by introducing a novel continuous approximation of the beam search decoding procedure.In experiments, we show that optimizing this new training objective yields substantially better results on two sequence tasks (Named Entity Recognition and CCG Supertagging) when compared with both cross entropy trained greedy decoding and cross entropy trained beam decoding baselines.
Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick
AAAI3
2018 LSTMs Can Learn Syntax-Sensitive Dependencies Well, But Modeling Structure Makes Them Better
abstract
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, Phil Blunsom. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, Phil Blunsom
ACL (1)2
2018 Finding syntax in human encephalography with beam search
abstract
Recurrent neural network grammars (RNNGs) are generative models of (tree, string) pairs that rely on neural networks to evaluate derivational choices.Parsing with them using beam search yields a variety of incremental complexity metrics such as word surprisal and parser action count.When used as regressors against human electrophysiological responses to naturalistic text, they derive two amplitude effects: an early peak and a P600-like later peak.By contrast, a non-syntactic neural language model yields no reliable effects.Model comparisons attribute the early peak to syntactic composition within the RNNG.This pattern of results recommends the RNNG+beam search combination as a mechanistic model of the syntactic processing that occurs during normal human language comprehension.
John T. Hale, Chris Dyer, Adhiguna Kuncoro, Jonathan Brennan
ACL (1)2
2018 Syntactic Scaffolds for Semantic Structures
abstract
We introduce the syntactic scaffold, an approach to incorporating syntactic information into semantic tasks.Syntactic scaffolds avoid expensive syntactic processing at runtime, only making use of a treebank during training, through a multitask objective.We improve over strong baselines on PropBank semantics, frame semantics, and coreference resolution, achieving competitive performance on all three tasks.
Swabha Swayamdipta, Sam Thomson, Kenton Lee, Luke Zettlemoyer, Chris Dyer, Noah A. Smith
EMNLP5
2018 On the State of the Art of Evaluation in Neural Language Models
Gábor Melis, Chris Dyer, Phil Blunsom
ICLR (Poster)2
2018 Memory Architectures in Recurrent Neural Network Language Models
Dani Yogatama, Yishu Miao, Gábor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, Phil Blunsom
ICLR (Poster)6
2018 Fast Parametric Learning with Activation Memorization
abstract
Neural networks trained with backpropagation often struggle to identify classes that have been observed a small number of times. In applications where most class labels are rare, such as language modelling, this can become a performance bottleneck. One potential remedy is to augment the network with a fast-learning non-parametric model which stores recent activations and class labels into an external memory. We explore a simplified architecture where we treat a subset of the model parameters as fast memory stores. This can help retain information over longer time intervals than a traditional memory, and does not require additional space or compute. In the case of image classification, we display faster binding of novel classes on an Omniglot image curriculum task. We also show improved performance for word-based language models on news reports (GigaWord), books (Project Gutenberg) and Wikipedia articles (WikiText-103) - the latter achieving a state-of-the-art perplexity of 29.2.
Jack W. Rae, Chris Dyer, Peter Dayan, Timothy P. Lillicrap
ICML2
2018 Using Morphological Knowledge in Open-Vocabulary Neural Language Models
abstract
Austin Matthews, Graham Neubig, Chris Dyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Austin Matthews, Graham Neubig, Chris Dyer
NAACL-HLT3
2018 Neural Arithmetic Logic Units
abstract
Neural networks can learn to represent and manipulate numerical information, but they seldom generalize well outside of the range of numerical values encountered during training. To encourage more systematic numerical extrapolation, we propose an architecture that represents numerical quantities as linear activations which are manipulated using primitive arithmetic operators, controlled by learned gates. We call this module a neural arithmetic logic unit (NALU), by analogy to the arithmetic logic unit in traditional processors. Experiments show that NALU-enhanced neural networks can learn to track time, perform arithmetic over images of numbers, translate numerical language into real-valued scalars, execute computer code, and count objects in images. In contrast to conventional architectures, we obtain substantially better generalization both inside and outside of the range of numerical values encountered during training, often extrapolating orders of magnitude beyond trained numerical ranges.
Andrew Trask, Felix Hill, Scott E. Reed, Jack W. Rae, Chris Dyer, Phil Blunsom
NeurIPS5
2018 Unsupervised Text Style Transfer using Language Models as Discriminators
abstract
Binary classifiers are employed as discriminators in GAN-based unsupervised style transfer models to ensure that transferred sentences are similar to sentences in the target domain. One difficulty with the binary discriminator is that error signal is sometimes insufficient to train the model to produce rich-structured language. In this paper, we propose a technique of using a target domain language model as the discriminator to provide richer, token-level feedback during the learning process. Because our language model scores sentences directly using a product of locally normalized probabilities, it offers more stable and more useful training signal to the generator. We train the generator to minimize the negative log likelihood (NLL) of generated sentences evaluated by a language model. By using continuous approximation of the discrete samples, our model can be trained using back-propagation in an end-to-end way. Moreover, we find empirically with a language model as a structured discriminator, it is possible to eliminate the adversarial training steps using negative samples, thus making training more stable. We compare our model with previous work using convolutional neural networks (CNNs) as discriminators and show our model outperforms them significantly in three tasks including word substitution decipherment, sentiment modification and related language translation.
Zhiting Hu, Chris Dyer, Eric P. Xing, Taylor Berg-Kirkpatrick
NeurIPS3
2018 The NarrativeQA Reading Comprehension Challenge
abstract
Reading comprehension (RC)—in contrast to information retrieval—requires integrating information and reasoning about events, entities, and their relations across a full document. Question answering is conventionally used to assess RC ability, in both artificial agents and children learning to read. However, existing RC datasets and tasks are dominated by questions that can be solved by selecting answers using superficial information (e.g., local context similarity or global term frequency); they thus fail to test for the essential integrative aspect of RC. To encourage progress on deeper comprehension of language, we present a new dataset and set of tasks in which the reader must answer questions about stories by reading entire books or movie scripts. These tasks are designed so that successfully answering their questions requires understanding the underlying narrative rather than relying on shallow pattern matching or salience. We show that although humans solve the tasks easily, standard RC models struggle on the tasks presented here. We provide an analysis of the dataset and the challenges it presents.
Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, Edward Grefenstette
Trans. Assoc. Comput. Linguistics4
2017 Ontology-Aware Token Embeddings for Prepositional Phrase Attachment
abstract
Type-level word embeddings use the same set of parameters to represent all instances of a word regardless of its context, ignoring the inherent lexical ambiguity in language.Instead, we embed semantic concepts (or synsets) as defined in WordNet and represent a word token in a particular context by estimating a distribution over relevant semantic concepts.We use the new, context-sensitive embeddings in a model for predicting prepositional phrase (PP) attachments and jointly learn the concept embeddings and model parameters.We show that using context-sensitive embeddings improves the accuracy of the PP attachment model by 5.4% absolute points, which amounts to a 34.4% relative reduction in errors.
Pradeep Dasigi, Waleed Ammar, Chris Dyer, Eduard H. Hovy
ACL (1)3
2017 Learning to Create and Reuse Words in Open-Vocabulary Neural Language Modeling
abstract
Fixed-vocabulary language models fail to account for one of the most characteristic statistical facts of natural language: the frequent creation and reuse of new word types.Although character-level language models offer a partial solution in that they can create word types not attested in the training corpus, they do not capture the "bursty" distribution of such words.In this paper, we augment a hierarchical LSTM language model that generates sequences of word tokens character by character with a caching mechanism that learns to reuse previously generated words.To validate our model we construct a new open-vocabulary language modeling corpus (the Multilingual Wikipedia Corpus; MWC) from comparable Wikipedia articles in 7 typologically diverse languages and demonstrate the effectiveness of our model across this range of languages.
Kazuya Kawakami, Chris Dyer, Phil Blunsom
ACL (1)2
2017 Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
abstract
Solving algebraic word problems requires executing a series of arithmetic operations-a program-to obtain a final answer.However, since programs can be arbitrarily complicated, inducing them directly from question-answer pairs is a formidable challenge.To make this task more feasible, we solve these problems by generating answer rationales, sequences of natural language and human-readable mathematical expressions that derive the final answer through a series of small steps.Although rationales do not explicitly specify programs, they provide a scaffolding for their structure via intermediate milestones.To evaluate our approach, we have created a new 100,000-sample dataset of questions, answers and rationales.Experimental results show that indirect supervision of program learning via answer rationales is a promising strategy for inducing arithmetic programs.
Wang Ling, Dani Yogatama, Chris Dyer, Phil Blunsom
ACL (1)3
2017 Should Neural Network Architecture Reflect Linguistic Structure?
abstract
I explore the hypothesis that conventional neural network models (e.g., recurrent neural networks) are incorrectly biased for making linguistically sensible generalizations when learning, and that a better class of models is based on architectures that reflect hierarchical structures for which considerable behavioral evidence exists. I focus on the problem of modeling and representing the meanings of sentences. On the generation front, I introduce recurrent neural network grammars (RNNGs), a joint, generative model of phrase-structure trees and sentences. RNNGs operate via a recursive syntactic process reminiscent of probabilistic context-free grammar generation, but decisions are parameterized using RNNs that condition on the entire (top-down, left-to-right) syntactic derivation history, thus relaxing context-free independence assumptions, while retaining a bias toward explaining decisions via "syntactically local" conditioning contexts. Experiments show that RNNGs obtain better results in generating language than models that don't exploit linguistic structure. On the representation front, I explore unsupervised learning of syntactic structures based on distant semantic supervision using a reinforcement-learning algorithm. The learner seeks a syntactic structure that provides a compositional architecture that produces a good representation for a downstream semantic task. Although the inferred structures are quite different from traditional syntactic analyses, the performance on the downstream tasks surpasses that of systems that use sequential RNNs and tree-structured RNNs based on treebank dependencies.
Chris Dyer
CoNLL1
2017 What Do Recurrent Neural Network Grammars Learn About Syntax?
abstract
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith
EACL (1)4
2017 Reference-Aware Language Models
abstract
We propose a general class of language models that treat reference as discrete stochastic latent variables.This decision allows for the creation of entity mentions by accessing external databases of referents (required by, e.g., dialogue generation) or past internal state (required to explicitly model coreferentiality).Beyond simple copying, our coreference model can additionally refer to a referent using varied mention forms (e.g., a reference to "Jane" can be realized as "she"), a characteristic feature of reference in natural languages.Experiments on three representative applications show our model variants outperform models based on deterministic attention and standard language modeling baselines.
Phil Blunsom, Chris Dyer, Wang Ling
EMNLP3
2017 Learning to Compose Words into Sentences with Reinforcement Learning
Dani Yogatama, Phil Blunsom, Chris Dyer, Edward Grefenstette, Wang Ling
ICLR (Poster)3
2017 The Neural Noisy Channel
Lei Yu 0008, Phil Blunsom, Chris Dyer, Edward Grefenstette, Tomás Kociský
ICLR (Poster)3
2017 Multitask Learning with CTC and Segmental CRF for Speech Recognition
abstract
Segmental conditional random fields (SCRFs) and connectionist temporal classification (CTC) are two sequence labeling methods used for end-to-end training of speech recognition models. Both models define a transcription probability by marginalizing decisions about latent segmentation alternatives to derive a sequence probability: the former uses a globally normalized joint model of segment labels and durations, and the latter classifies each frame as either an output symbol or a "continuation" of the previous label. In this paper, we train a recognition model by optimizing an interpolation between the SCRF and CTC losses, where the same recurrent neural network (RNN) encoder is used for feature extraction for both outputs. We find that this multitask objective improves recognition accuracy when decoding with either the SCRF or CTC models. Additionally, we show that CTC can also be used to pretrain the RNN encoder, which improves the convergence rate when learning the joint model.
Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith
INTERSPEECH3
2017 On-the-fly Operation Batching in Dynamic Computation Graphs
abstract
Dynamic neural networks toolkits such as PyTorch, DyNet, and Chainer offer more flexibility for implementing models that cope with data of varying dimensions and structure, relative to toolkits that operate on statically declared computations (e.g., TensorFlow, CNTK, and Theano). However, existing toolkits - both static and dynamic - require that the developer organize the computations into the batches necessary for exploiting high-performance data-parallel algorithms and hardware. This batching task is generally difficult, but it becomes a major hurdle as architectures become complex. In this paper, we present an algorithm, and its implementation in the DyNet toolkit, for automatically batching operations. Developers simply write minibatch computations as aggregations of single instance computations, and the batching algorithm seamlessly executes them, on the fly, in computationally efficient batches. On a variety of tasks, we obtain throughput similar to manual batches, as well as comparable speedups over single-instance learning on architectures that are impractical to batch manually.
Graham Neubig, Yoav Goldberg, Chris Dyer
NIPS3
2017 Greedy Transition-Based Dependency Parsing with Stack LSTMs
abstract
We introduce a greedy transition-based parser that learns to represent parser states using recurrent neural networks. Our primary innovation that enables us to do this efficiently is a new control structure for sequential neural networks—the stack long short-term memory unit (LSTM). Like the conventional stack data structures used in transition-based parsers, elements can be pushed to or popped from the top of the stack in constant time, but, in addition, an LSTM maintains a continuous space embedding of the stack contents. Our model captures three facets of the parser's state: (i) unbounded look-ahead into the buffer of incoming words, (ii) the complete history of transition actions taken by the parser, and (iii) the complete contents of the stack of partially built tree fragments, including their internal structures. In addition, we compare two different word representations: (i) standard word vectors based on look-up tables and (ii) character-based models of words. Although standard word embedding models work well in all languages, the character-based models improve the handling of out-of-vocabulary words, particularly in morphologically rich languages. Finally, we discuss the use of dynamic oracles in training the parser. During training, dynamic oracles alternate between sampling parser states from the training data and from the model as it is being learned, making the model more robust to the kinds of errors that will be made at test time. Training our model with dynamic oracles yields a linear-time greedy parser with very competitive performance.
Miguel Ballesteros, Chris Dyer, Yoav Goldberg, Noah A. Smith
Comput. Linguistics2
2016 Modeling Evolving Relationships Between Characters in Literary Novels
abstract
Studying characters plays a vital role in computationally representing and interpreting narratives. Unlike previous work, which has focused on inferring character roles, we focus on the problem of modeling their relationships. Rather than assuming a fixed relationship for a character pair, we hypothesize that relationships temporally evolve with the progress of the narrative, and formulate the problem of relationship modeling as a structured prediction problem. We propose a semi-supervised framework to learn relationship sequences from fully as well as partially labeled data. We present a Markovian model capable of accumulating historical beliefs about the relationship and status changes. We use a set of rich linguistic and semantically motivated features that incorporate world knowledge to investigate the textual content of narrative. We empirically demonstrate that such a framework outperforms competitive baselines.
Snigdha Chaturvedi, Hal Daumé III, Chris Dyer
AAAI4
2016 Synthesizing Compound Words for Machine Translation
abstract
Most machine translation systems construct translations from a closed vocabulary of target word forms, posing problems for translating into languages that have productive compounding processes.We present a simple and effective approach that deals with this problem in two phases.First, we build a classifier that identifies spans of the input text that can be translated into a single compound word in the target language.Then, for each identified span, we generate a pool of possible compounds which are added to the translation model as "synthetic" phrase translations.Experiments reveal that (i) we can effectively predict what spans can be compounded; (ii) our compound generation model produces good compounds; and (iii) modest improvements are possible in end-to-end English-German and English-Finnish translation tasks.We additionally introduce KomposEval, a new multi-reference dataset of English phrases and their translations into German compounds.
Austin Matthews, Eva Schlinger, Alon Lavie, Chris Dyer
ACL (1)4
2016 Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning
abstract
We use Bayesian optimization to learn curricula for word representation learning, optimizing performance on downstream tasks that depend on the learned representations as features.The curricula are modeled by a linear ranking function which is the scalar product of a learned weight vector and an engineered feature vector that characterizes the different aspects of the complexity of each instance in the training corpus.We show that learning the curriculum improves performance on a variety of downstream tasks over random orders and in comparison to the natural corpus order.
Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Brian MacWhinney, Chris Dyer
ACL (1)5
2016 Cross-lingual Models of Word Embeddings: An Empirical Comparison
abstract
Despite interest in using cross-lingual knowledge to learn word embeddings for various tasks, a systematic comparison of the possible approaches is lacking in the literature.We perform an extensive evaluation of four popular approaches of inducing cross-lingual embeddings, each requiring a different form of supervision, on four typologically different language pairs.Our evaluation setup spans four different tasks, including intrinsic evaluation on mono-lingual and cross-lingual similarity, and extrinsic evaluation on downstream semantic and syntactic applications.We show that models which require expensive cross-lingual knowledge almost always perform better, but cheaply supervised models often prove competitive on certain tasks.
Shyam Upadhyay, Manaal Faruqui, Chris Dyer, Dan Roth 0001
ACL (1)3
2016 Named Entity Recognition for Linguistic Rapid Response in Low-Resource Languages: Sorani Kurdish and Tajik
abstract
This paper describes our construction of named-entity recognition (NER) systems in two Western Iranian languages, Sorani Kurdish and Tajik, as a part of a pilot study of “Linguistic Rapid Response” to potential emergency humanitarian relief situations. In the absence of large annotated corpora, parallel corpora, treebanks, bilingual lexica, etc., we found the following to be effective: exploiting distributional regularities in monolingual data, projecting information across closely related languages, and utilizing human linguist judgments. We show promising results on both a four-month exercise in Sorani and a two-day exercise in Tajik, achieved with minimal annotation costs.
Patrick Littell, Kartik Goyal, David R. Mortensen, Alexa Little, Chris Dyer, Lori S. Levin
COLING5
2016 PanPhon: A Resource for Mapping IPA Segments to Articulatory Feature Vectors
abstract
This paper contributes to a growing body of evidence that—when coupled with appropriate machine-learning techniques–linguistically motivated, information-rich representations can outperform one-hot encodings of linguistic data. In particular, we show that phonological features outperform character-based models. PanPhon is a database relating over 5,000 IPA segments to 21 subsegmental articulatory features. We show that this database boosts performance in various NER-related tasks. Phonologically aware, neural CRF models built on PanPhon features are able to perform better on monolingual Spanish and Turkish NER tasks that character-based models. They have also been shown to work well in transfer models (as between Uzbek and Turkish). PanPhon features also contribute measurably to Orthography-to-IPA conversion tasks.
David R. Mortensen, Patrick Littell, Akash Bharadwaj, Kartik Goyal, Chris Dyer, Lori S. Levin
COLING5
2016 The Role of Context in Neural Morphological Disambiguation
abstract
Languages with rich morphology often introduce sparsity in language processing tasks. While morphological analyzers can reduce this sparsity by providing morpheme-level analyses for words, they will often introduce ambiguity by returning multiple analyses for the same surface form. The problem of disambiguating between these morphological parses is further complicated by the fact that a correct parse for a word is not only be dependent on the surface form but also on other words in its context. In this paper, we present a language-agnostic approach to morphological disambiguation. We address the problem of using context in morphological disambiguation by presenting several LSTM-based neural architectures that encode long-range surface-level and analysis-level contextual dependencies. We applied our approach to Turkish, Russian, and Arabic to compare effectiveness across languages, matching state-of-the-art results in two of the three languages. Our results also demonstrate that while context plays a role in learning how to disambiguate, the type and amount of context needed varies between languages.
Qinlan Shen, Daniel Clothiaux, Emily Tagtow, Patrick Littell, Chris Dyer
COLING5
2016 Greedy, Joint Syntactic-Semantic Parsing with Stack LSTMs
abstract
We present a transition-based parser that jointly produces syntactic and semantic dependencies. It learns a representation of the entire algorithm state, using stack long short-term memories. Our greedy inference algorithm has linear time, including feature extraction. On the CoNLL 2008--9 English shared tasks, we obtain the best published parsing performance among models that jointly learn syntax and semantics.
Swabha Swayamdipta, Miguel Ballesteros, Chris Dyer, Noah A. Smith
CoNLL3
2016 Training with Exploration Improves a Greedy Stack LSTM Parser
abstract
We adapt the greedy Stack-LSTM dependency parser of Dyer et al. (2015) to support a training-with-exploration procedure using dynamic oracles(Goldberg and Nivre, 2013) instead of cross-entropy minimization. This form of training, which accounts for model predictions at training time rather than assuming an error-free action history, improves parsing accuracies for both English and Chinese, obtaining very strong results for both languages. We discuss some modifications needed in order to get training with exploration to work well for a probabilistic neural-network.
Miguel Ballesteros, Yoav Goldberg, Chris Dyer, Noah A. Smith
EMNLP3
2016 Phonologically Aware Neural Model for Named Entity Recognition in Low Resource Transfer Settings
abstract
Named Entity Recognition is a well established information extraction task with many state of the art systems existing for a variety of languages.Most systems rely on language specific resources, large annotated corpora, gazetteers and feature engineering to perform well monolingually.In this paper, we introduce an attentional neural model which only uses language universal phonological character representations with word embeddings to achieve state of the art performance in a monolingual setting using supervision and which can quickly adapt to a new language with minimal or no data.We demonstrate that phonological character representations facilitate cross-lingual transfer, outperform orthographic representations and incorporating both attention and phonological features improves statistical efficiency of the model in 0-shot and low data transfer settings with no task specific feature engineering in the source or target language.
Akash Bharadwaj, David R. Mortensen, Chris Dyer, Jaime G. Carbonell
EMNLP3
2016 Transition-Based Dependency Parsing with Heuristic Backtracking
abstract
Comunicació presentada a Conference on Empirical Methods in Natural Language Processing
Jacob Buckman, Miguel Ballesteros, Chris Dyer
EMNLP3
2016 Character Sequence Models for Colorful Words
abstract
We present a neural network architecture to predict a point in space from the sequence of characters in the color's name. Using large scale color--name pairs obtained from an online design forum, we evaluate our model on a color Turing test and find that, given a name, the colors predicted by our model are preferred by annotators to names created by humans. Our datasets and demo system are available online at this http URL.
Kazuya Kawakami, Chris Dyer, Bryan R. Routledge, Noah A. Smith
EMNLP2
2016 Semantic Parsing with Semi-Supervised Sequential Autoencoders
abstract
We present a novel semi-supervised approach for sequence transduction and apply it to semantic parsing.The unsupervised component is based on a generative model in which latent sentences generate the unpaired logical forms.We apply this method to a number of semantic parsing tasks focusing on domains with limited access to labelled training data and extend those datasets with synthetically generated logical forms.
Tomás Kociský, Gábor Melis, Edward Grefenstette, Chris Dyer, Wang Ling, Phil Blunsom, Karl Moritz Hermann
EMNLP4
2016 Distilling an Ensemble of Greedy Dependency Parsers into One MST Parser
abstract
We introduce two first-order graph-based dependency parsers achieving a new state of the art.The first is a consensus parser built from an ensemble of independently trained greedy LSTM transition-based parsers with different random initializations.We cast this approach as minimum Bayes risk decoding (under the Hamming cost) and argue that weaker consensus within the ensemble is a useful signal of difficulty or ambiguity.The second parser is a "distillation" of the ensemble into a single model.We train the distillation parser using a structured hinge loss objective with a novel cost that incorporates ensemble uncertainty estimates for each possible attachment, thereby avoiding the intractable crossentropy computations required by applying standard distillation objectives to problems with structured outputs.The first-order distillation parser matches or surpasses the state of the art on English, Chinese, and German.
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Noah A. Smith
EMNLP4
2016 Generalizing and Hybridizing Count-based and Neural Language Models
abstract
Language models (LMs) are statistical models that calculate probabilities over sequences of words or other discrete symbols.Currently two major paradigms for language modeling exist: count-based n-gram models, which have advantages of scalability and test-time speed, and neural LMs, which often achieve superior modeling performance.We demonstrate how both varieties of models can be unified in a single modeling framework that defines a set of probability distributions over the vocabulary of words, and then dynamically calculates mixture weights over these distributions.This formulation allows us to create novel hybrid models that combine the desirable features of count-based and neural LMs, and experiments demonstrate the advantages of these approaches. 1
Graham Neubig, Chris Dyer
EMNLP2
2016 Segmental Recurrent Neural Networks for End-to-End Speech Recognition
abstract
We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based acoustic models, it does not rely on an external system to provide features or segmentation boundaries. Instead, this model marginalises out all the possible segmentations, and features are extracted from the RNN trained together with the segmental CRF. In essence, this model is self-contained and can be trained end-to-end. In this paper, we discuss practical training and decoding issues as well as the method to speed up the training in the context of speech recognition. We performed experiments on the TIMIT dataset. We achieved 17.3 phone error rate (PER) from the first-pass decoding --- the best reported result using CRFs, despite the fact that we only used a zeroth-order CRF and without using any language model.
Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith, Steve Renals
INTERSPEECH3
2016 Bridge-Language Capitalization Inference in Western Iranian: Sorani, Kurmanji, Zazaki, and Tajik
Patrick Littell, David R. Mortensen, Kartik Goyal, Chris Dyer, Lori S. Levin
LREC4
2016 Incorporating Structural Alignment Biases into an Attentional Neural Translation Model
abstract
Neural encoder-decoder models of machine translation have achieved impressive results, rivalling traditional translation models. However their modelling formulation is overly simplistic, and omits several key inductive biases built into traditional models. In this paper we extend the attentional neural translation model to include structural biases from word based alignment models, including positional bias, Markov conditioning, fertility and agreement over translation directions. We show improvements over a baseline attentional model and standard phrase-based model over several language pairs, evaluating on difficult languages in a low resource setting.
Trevor Cohn, Cong Duy Vu Hoang, Ekaterina Vymolova, Kaisheng Yao, Chris Dyer, Gholamreza Haffari
HLT-NAACL5
2016 Recurrent Neural Network Grammars
abstract
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith
HLT-NAACL1
2016 Morphological Inflection Generation Using Character Sequence to Sequence Learning
abstract
Manaal Faruqui, Yulia Tsvetkov, Graham Neubig, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Manaal Faruqui, Yulia Tsvetkov, Graham Neubig, Chris Dyer
HLT-NAACL4
2016 Generation from Abstract Meaning Representation using Tree Transducers
abstract
Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime Carbonell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime G. Carbonell
HLT-NAACL2
2016 Neural Architectures for Named Entity Recognition
abstract
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, Chris Dyer
HLT-NAACL5
2016 Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation Learning
abstract
Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David R. Mortensen, Alan W. Black, Lori S. Levin, Chris Dyer
HLT-NAACL9
2016 Hierarchical Attention Networks for Document Classification
abstract
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, Eduard Hovy. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Diyi Yang, Chris Dyer, Xiaodong He 0001, Alexander J. Smola, Eduard H. Hovy
HLT-NAACL3
2016 Mining Parallel Corpora from Sina Weibo and Twitter
abstract
Microblogs such as Twitter, Facebook, and Sina Weibo (China's equivalent of Twitter) are a remarkable linguistic resource. In contrast to content from edited genres such as newswire, microblogs contain discussions of virtually every topic by numerous individuals in different languages and dialects and in different styles. In this work, we show that some microblog users post “self-translated” messages targeting audiences who speak different languages, either by writing the same message in multiple languages or by retweeting translations of their original posts in a second language. We introduce a method for finding and extracting this naturally occurring parallel data. Identifying the parallel content requires solving an alignment problem, and we give an optimally efficient dynamic programming algorithm for this. Using our method, we extract nearly 3M Chinese–English parallel segments from Sina Weibo using a targeted crawl of Weibo users who post in multiple languages. Additionally, from a random sample of Twitter, we obtain substantial amounts of parallel data in multiple language pairs. Evaluation is performed by assessing the accuracy of our extraction approach relative to a manual annotation as well as in terms of utility as training data for a Chinese–English machine translation system. Relative to traditional parallel data resources, the automatically extracted parallel data yield substantial translation quality improvements in translating microblog text and modest improvements in translating edited news content.
Wang Ling, Luís Marujo, Chris Dyer, Alan W. Black, Isabel Trancoso
Comput. Linguistics3
2016 Cross-Lingual Bridges with Models of Lexical Borrowing
abstract
Linguistic borrowing is the phenomenon of transferring linguistic constructions (lexical, phonological, morphological, and syntactic) from a “donor” language to a “recipient” language as a result of contacts between communities speaking different languages. Borrowed words are found in all languages, and—in contrast to cognate relationships—borrowing relationships may exist across unrelated languages (for example, about 40% of Swahili’s vocabulary is borrowed from the unrelated language Arabic). In this work, we develop a model of morpho-phonological transformations across languages. Its features are based on universal constraints from Optimality Theory (OT), and we show that compared to several standard—but linguistically more naïve—baselines, our OT-inspired model obtains good performance at predicting donor forms from borrowed forms with only a few dozen training examples, making this a cost-effective strategy for sharing lexical information across languages. We demonstrate applications of the lexical borrowing model in machine translation, using resource-rich donor language to obtain translations of out-of-vocabulary loanwords in a lower resource language. Our framework obtains substantial improvements (up to 1.6 BLEU) over standard baselines.
Yulia Tsvetkov, Chris Dyer
J. Artif. Intell. Res.2
2016 Many Languages, One Parser
abstract
We train one multilingual model for dependency parsing and use it to parse sentences in several languages. The parsing model uses (i) multilingual word clusters and embeddings; (ii) token-level language information; and (iii) language-specific features (fine-grained POS tags). This input representation enables the parser not only to parse effectively in multiple languages, but also to generalize across languages based on linguistic universals and typological similarities, making it more effective to learn from limited annotations. Our parser’s performance compares favorably to strong baselines in a range of data scenarios, including when the target language has a large treebank, a small treebank, or no treebank for training.
Waleed Ammar, George Mulcaire, Miguel Ballesteros, Chris Dyer, Noah A. Smith
Trans. Assoc. Comput. Linguistics4
2015 Weakly-Supervised Grammar-Informed Bayesian CCG Parser Learning
abstract
Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that, in combination with a small universal set of rules, specify the syntactic configurations in which they may occur. Categories are selected from a large, recursively-defined set; this leads to high word-to-category ambiguity, which is one of the primary factors that make learning CCG parsers difficult, especially in the face of little data. Previous work has shown that learning sequence models for CCG tagging can be improved by using linguistically-motivated prior probability distributions over potential categories. We extend this approach to the task of learning a CCG parser from weak supervision. We present a Bayesian formulation for CCG parser induction that assumes only supervision in the form of an incomplete tag dictionary mapping some word types to sets of potential categories. Our approach outperforms a baseline model trained with uniform priors by exploiting universal, intrinsic properties of the CCG formalism to bias the model toward simpler, more cross-linguistically common categories.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
AAAI2
2015 Gaussian LDA for Topic Models with Word Embeddings
abstract
Rajarshi Das, Manzil Zaheer, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Rajarshi Das, Manzil Zaheer, Chris Dyer
ACL (1)3
2015 Unifying Bayesian Inference and Vector Space Models for Improved Decipherment
abstract
Qing Dou, Ashish Vaswani, Kevin Knight, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Qing Dou, Ashish Vaswani, Kevin Knight, Chris Dyer
ACL (1)4
2015 Transition-Based Dependency Parsing with Stack Long Short-Term Memory
abstract
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith
ACL (1)1
2015 Sparse Overcomplete Word Vector Representations
abstract
Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, Noah A. Smith
ACL (1)4
2015 A Supertag-Context Model for Weakly-Supervised CCG Parser Learning
abstract
Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that specify the syntactic configurations in which they may occur.We present a novel parsing model with the capacity to capture the associative adjacent-category relationships intrinsic to CCG by parameterizing the relationships between each constituent label and the preterminal categories directly to its left and right, biasing the model toward constituent categories that can combine with their contexts.This builds on the intuitions of Klein and Manning's (2002) "constituentcontext" model, which demonstrated the value of modeling context, but has the advantage of being able to exploit the properties of CCG.Our experiments show that our model outperforms a baseline in which this context information is not captured.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
CoNLL2
2015 Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs
abstract
We present extensions to a continuousstate dependency parsing method that makes it applicable to morphologically rich languages. Starting with a highperformance transition-based parser that uses long short-term memory (LSTM) recurrent neural networks to learn representations of the parser state, we replace lookup-based word representations with representations constructed from the orthographic representations of the words, also using LSTMs. This allows statistical sharing across word forms that are similar on the surface. Experiments for morphologically rich languages show that the parsing model benefits from incorporating the character-based encodings of words.
Miguel Ballesteros, Chris Dyer, Noah A. Smith
EMNLP2
2015 Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
abstract
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, Tiago Luís
EMNLP2
2015 Not All Contexts Are Created Equal: Better Word Representations with Variable Attention
abstract
Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramón Fermandez, Chris Dyer, Alan W Black, Isabel Trancoso, Chu-Cheng Lin. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramon Fermandez, Chris Dyer, Alan W. Black, Isabel Trancoso, Chu-Cheng Lin
EMNLP5
2015 Evaluation of Word Vector Representations by Subspace Alignment
abstract
Unsupervisedly learned word vectors have proven to provide exceptionally effective features in many NLP tasks.Most common intrinsic evaluations of vector quality measure correlation with similarity judgments.However, these often correlate poorly with how well the learned representations perform as features in downstream evaluation tasks.We present QVEC-a computationally inexpensive intrinsic evaluation measure of the quality of word embeddings based on alignment to a matrix of features extracted from manually crafted lexical resources-that obtains strong correlation with performance of the vectors in a battery of downstream semantic evaluation tasks. 1
Yulia Tsvetkov, Manaal Faruqui, Wang Ling, Guillaume Lample, Chris Dyer
EMNLP5
2015 Humor Recognition and Humor Anchor Extraction
abstract
Humor is an essential component in personal communication. How to create computational models to discover the structures behind humor, recognize humor and even extract humor anchors remains a challenge. In this work, we first identify several semantic structures behind humor and design sets of features for each structure, and next employ a com-putational approach to recognize humor. Furthermore, we develop a simple and effective method to extract anchors that enable humor in a sentence. Experiments conducted on two datasets demonstrate that our humor recognizer is effective in automatically distinguishing between humorous and non-humorous texts and our extracted humor anchors correlate quite well with human annotations. 1
Diyi Yang, Alon Lavie, Chris Dyer, Eduard H. Hovy
EMNLP3
2015 Learning Word Representations with Hierarchical Sparse Coding
abstract
We propose a new method for learning word representations using hierarchical regularization in sparse coding inspired by the linguistic study of word meanings. We show an efficient learning algorithm based on stochastic proximal methods that is significantly faster than previous approaches, making it possible to perform hierarchical sparse coding on a corpus of billions of word tokens. Experiments on various benchmark tasks—word similarity ranking, syntactic and semantic analogies, sentence completion, and sentiment analysis—demonstrate that the method outperforms or is competitive with state-of-the-art methods.
Dani Yogatama, Manaal Faruqui, Chris Dyer, Noah A. Smith
ICML3
2015 Retrofitting Word Vectors to Semantic Lexicons
abstract
Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard H. Hovy, Noah A. Smith
HLT-NAACL4
2015 Ontologically Grounded Multi-sense Representation Learning for Semantic Vector Space Models
abstract
Words are polysemous. However, most approaches to representation learning for lexical semantics assign a single vector to every surface word type. Meanwhile, lexical ontologies such as WordNet provide a source of complementary knowledge to distributional information, including a word sense inventory. In this paper we propose two novel and general approaches for generating sense-specific word embeddings that are grounded in an ontology. The first applies graph smoothing as a postprocessing step to tease the vectors of different senses apart, and is applicable to any vector space model. The second adapts predictive maximum likelihood models that learn word embeddings with latent variables representing senses grounded in an specified ontology. Empirical results on lexical semantic tasks show that our approaches effectively captures information from both the ontology and distributional statistics. Moreover, in most cases our sense-specific models outperform other models we compare against.
Sujay Kumar Jauhar, Chris Dyer, Eduard H. Hovy
HLT-NAACL2
2015 Unsupervised POS Induction with Word Embeddings
abstract
Chu-Cheng Lin, Waleed Ammar, Chris Dyer, Lori Levin. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Chu-Cheng Lin, Waleed Ammar, Chris Dyer, Lori S. Levin
HLT-NAACL3
2015 Two/Too Simple Adaptations of Word2Vec for Syntax Problems
abstract
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso
HLT-NAACL2
2015 Constraint-Based Models of Lexical Borrowing
abstract
Linguistic borrowing is the phenomenon of transferring linguistic constructions (lexical, phonological, morphological, and syntactic) from a "donor" language to a "recipient" language as a result of contacts between communities speaking different languages.Borrowed words are found in all languages, and-in contrast to cognate relationships-borrowing relationships may exist across unrelated languages (for example, about 40% of Swahili's vocabulary is borrowed from Arabic).In this paper, we develop a model of morpho-phonological transformations across languages with features based on universal constraints from Optimality Theory (OT).Compared to several standardbut linguistically naïve-baselines, our OTinspired model obtains good performance with only a few dozen training examples, making this a cost-effective strategy for sharing lexical information across languages.
Yulia Tsvetkov, Waleed Ammar, Chris Dyer
HLT-NAACL3
2015 Linguistic Fundamentals for Natural Language Processing: 100 Essentials from Morphology and Syntax Emily M. Bender (University of Washington) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 20), 2013, xvii+166 pp; paperbound, ISBN 978-1-62705-011-1, $40.00; e-book, ISBN 978-1-62705-012-8, $30.00 or by subscription
abstract
The phenomenal success of machine learning in engineering natural language applications has led to a curious situation: Natural language processing practitioners who were trained in the last 15 to 20 years may have established a quite successful career in this area with only a haphazard knowledge of the science of natural languages.The premise of the new volume by Emily M. Bender is that greater awareness of linguistics will enable continued technical progress, particularly as language applications are required to perform more intelligent processing in more languages.
Chris Dyer
Comput. Linguistics1
2014 A Discriminative Graph-Based Parser for the Abstract Meaning Representation
abstract
Meaning Representation (AMR) is a semantic formalism for which a growing set of annotated examples is available.We introduce the first approach to parse sentences into this representation, providing a strong baseline for future improvement.The method is based on a novel algorithm for finding a maximum spanning, connected subgraph, embedded within a Lagrangian relaxation of an optimization problem that imposes linguistically inspired constraints.Our approach is described in the general framework of structured prediction, allowing future incorporation of additional features and constraints, and may extend to other formalisms as well.Our open-source system, JAMR, is available at:
Jeffrey Flanigan, Sam Thomson, Jaime G. Carbonell, Chris Dyer, Noah A. Smith
ACL (1)4
2014 Metaphor Detection with Cross-Lingual Model Transfer
abstract
We show that it is possible to reliably dis-criminate whether a syntactic construction is meant literally or metaphorically using lexical semantic features of the words that participate in the construction. Our model is constructed using English resources, and we obtain state-of-the-art performance relative to previous work in this language. Using a model transfer approach by piv-oting through a bilingual dictionary, we show our model can identify metaphoric expressions in other languages. We pro-vide results on three new test sets in Span-ish, Farsi, and Russian. The results sup-port the hypothesis that metaphors are conceptual, rather than lexical, in nature. 1
Yulia Tsvetkov, Leonid Boytsov, Anatole Gershman, Eric Nyberg, Chris Dyer
ACL (1)5
2014 Automatic Classification of Communicative Functions of Definiteness
Archna Bhatia, Chu-Cheng Lin, Nathan Schneider 0001, Yulia Tsvetkov, Fatima Talib Al-Raisi, Laleh Roostapour, Jordan Bender, Abhimanu Kumar, Lori S. Levin, Mandy Simons, Chris Dyer
COLING11
2014 Weakly-Supervised Bayesian Learning of a CCG Supertagger
abstract
We present a Bayesian formulation for weakly-supervised learning of a Combinatory Categorial Grammar (CCG) supertagger with an HMM.We assume supervision in the form of a tag dictionary, and our prior encourages the use of crosslinguistically common category structures as well as transitions between tags that can combine locally according to CCG's combinators.Our prior is theoretically appealing since it is motivated by languageindependent, universal properties of the CCG formalism.Empirically, we show that it yields substantial improvements over previous work that used similar biases to initialize an EM-based learner.Additional gains are obtained by further shaping the prior with corpus-specific information that is extracted automatically from raw text and a tag dictionary.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
CoNLL2
2014 Learning from Post-Editing: Online Model Adaptation for Statistical Machine Translation
abstract
Using machine translation output as a starting point for human translation has become an increasingly common application of MT.We propose and evaluate three computationally efficient online methods for updating statistical MT systems in a scenario where post-edited MT output is constantly being returned to the system: (1) adding new rules to the translation model from the post-edited content, (2) updating a Bayesian language model of the target language that is used by the MT system, and (3) updating the MT system's discriminative parameters with a MIRA step.Individually, these techniques can substantially improve MT quality, even over strong baselines.Moreover, we see super-additive improvements when all three techniques are used in tandem.
Michael J. Denkowski, Chris Dyer, Alon Lavie
EACL2
2014 Improving Vector Space Word Representations Using Multilingual Correlation
abstract
The distributional hypothesis of Harris (1954), according to which the meaning of words is evidenced by the contexts they occur in, has motivated several effective techniques for obtaining vector space semantic representations of words using unannotated text corpora. This paper argues that lexico-semantic content should additionally be invariant across languages and proposes a simple technique based on canonical correlation analysis (CCA) for incorporating multilingual evidence into vectors generated monolingually. We evaluate the resulting word representations on standard lexical semantic evaluation tasks and show that our method produces substantially better semantic representations than monolingual techniques.
Manaal Faruqui, Chris Dyer
EACL2
2014 Augmenting Translation Models with Simulated Acoustic Confusions for Improved Spoken Language Translation
abstract
We propose a novel technique for adapting text-based statistical machine translation to deal with input from automatic speech recognition in spoken language translation tasks.We simulate likely misrecognition errors using only a source language pronunciation dictionary and language model (i.e., without an acoustic model), and use these to augment the phrase table of a standard MT system.The augmented system can thus recover from recognition errors during decoding using synthesized phrases.Using the outputs of five different English ASR systems as input, we find consistent and significant improvements in translation quality.Our proposed technique can also be used in conjunction with lattices as ASR output, leading to further improvements.
Yulia Tsvetkov, Florian Metze, Chris Dyer
EACL3
2014 A Dependency Parser for Tweets
abstract
We describe a new dependency parser for English tweets, TWEEBOPARSER. The parser builds on several contributions: new syntactic annotations for a corpus of tweets (TWEEBANK), with conventions informed by the domain; adaptations to a statistical parsing algorithm; and a new approach to exploiting out-of-domain Penn Treebank data. Our experiments show that the parser achieves over 80% unlabeled attachment accuracy on our new, high-quality test set and measure the benefit of our contributions. Our dataset and parser can be found at http://www.ark.cs.cmu.edu/TweetNLP.
Lingpeng Kong, Nathan Schneider 0001, Swabha Swayamdipta, Archna Bhatia, Chris Dyer, Noah A. Smith
EMNLP5
2014 Language Modeling with Power Low Rank Ensembles
abstract
We present power low rank ensembles (PLRE), a flexible framework for n-gram language modeling where ensembles of low rank matrices and tensors are used to obtain smoothed probability estimates of words in context.Our method can be understood as a generalization of ngram modeling to non-integer n, and includes standard techniques such as absolute discounting and Kneser-Ney smoothing as special cases.PLRE training is efficient and our approach outperforms stateof-the-art modified Kneser Ney baselines in terms of perplexity on large corpora as well as on BLEU score in a downstream machine translation task.
Ankur P. Parikh, Avneesh Saluja, Chris Dyer, Eric P. Xing
EMNLP3
2014 Latent-Variable Synchronous CFGs for Hierarchical Translation
abstract
Data-driven refinement of non-terminal categories has been demonstrated to be a reliable technique for improving mono-lingual parsing with PCFGs. In this pa-per, we extend these techniques to learn latent refinements of single-category syn-chronous grammars, so as to improve translation performance. We compare two estimators for this latent-variable model: one based on EM and the other is a spec-tral algorithm based on the method of mo-ments. We evaluate their performance on a Chinese–English translation task. The re-sults indicate that we can achieve signifi-cant gains over the baseline with both ap-proaches, but in particular the moments-based estimator is both faster and performs better than EM. 1
Avneesh Saluja, Chris Dyer, Shay B. Cohen
EMNLP2
2014 A Unified Annotation Scheme for the Semantic/Pragmatic Components of Definiteness
Archna Bhatia, Mandy Simons, Lori S. Levin, Yulia Tsvetkov, Chris Dyer, Jordan Bender
LREC5
2014 Augmenting English Adjective Senses with Supersenses
Yulia Tsvetkov, Nathan Schneider 0001, Dirk Hovy, Archna Bhatia, Manaal Faruqui, Chris Dyer
LREC6
2014 Dual Subtitles as Parallel Corpora
Shikun Zhang, Wang Ling, Chris Dyer
LREC3
2014 Conditional Random Field Autoencoders for Unsupervised Structured Prediction
Waleed Ammar, Chris Dyer, Noah A. Smith
NIPS2
2014 Locally Non-Linear Learning for Statistical Machine Translation via Discretization and Structured Regularization
abstract
Linear models, which support efficient learning and inference, are the workhorses of statistical machine translation; however, linear decision rules are less attractive from a modeling perspective. In this work, we introduce a technique for learning arbitrary, rule-local, non-linear feature transforms that improve model expressivity, but do not sacrifice the efficient inference and learning associated with linear models. To demonstrate the value of our technique, we discard the customary log transform of lexical probabilities and drop the phrasal translation probability in favor of raw counts. We observe that our algorithm learns a variation of a log transform that leads to better translation quality compared to the explicit log transform. We conclude that non-linear responses play an important role in SMT, an observation that we hope will inform the efforts of feature engineers.
Jonathan H. Clark, Chris Dyer, Alon Lavie
Trans. Assoc. Comput. Linguistics2
2014 Discriminative Lexical Semantic Segmentation with Gaps: Running the MWE Gamut
abstract
We present a novel representation, evaluation measure, and supervised models for the task of identifying the multiword expressions (MWEs) in a sentence, resulting in a lexical semantic segmentation. Our approach generalizes a standard chunking representation to encode MWEs containing gaps, thereby enabling efficient sequence tagging algorithms for feature-rich discriminative models. Experiments on a new dataset of English web text offer the first linguistically-driven evaluation of MWE identification with truly heterogeneous expression types. Our statistical sequence model greatly outperforms a lookup-based segmentation procedure, achieving nearly 60% F1 for MWE identification.
Nathan Schneider 0001, Emily Danchik, Chris Dyer, Noah A. Smith
Trans. Assoc. Comput. Linguistics3
2013 Microblogs as Parallel Corpora
Wang Ling, Guang Xiang, Chris Dyer, Alan W. Black, Isabel Trancoso
ACL (1)3
2013 Translating into Morphologically Rich Languages with Synthetic Phrases
abstract
Translation into morphologically rich languages is an important but recalcitrant problem in MT.We present a simple and effective approach that deals with the problem in two phases.First, a discriminative model is learned to predict inflections of target words from rich source-side annotations.Then, this model is used to create additional sentencespecific word-and phrase-level translations that are added to a standard translation model as "synthetic" phrases.Our approach relies on morphological analysis of the target language, but we show that an unsupervised Bayesian model of morphology can successfully be used in place of a supervised analyzer.We report significant improvements in translation quality when translating from English to Russian, Hebrew and Swahili.
Victor Chahuneau, Eva Schlinger, Noah A. Smith, Chris Dyer
EMNLP4
2013 A Systematic Exploration of Diversity in Machine Translation
abstract
This paper addresses the problem of producing a diverse set of plausible translations.We present a simple procedure that can be used with any statistical machine translation (MT) system.We explore three ways of using diverse translations: (1) system combination, (2) discriminative reranking with rich features, and (3) a novel post-editing scenario in which multiple translations are presented to users.We find that diversity can improve performance on these tasks, especially for sentences that are difficult for MT.
Kevin Gimpel, Dhruv Batra, Chris Dyer, Gregory Shakhnarovich
EMNLP3
2013 Paraphrasing 4 Microblog Normalization
abstract
Compared to the edited genres that have played a central role in NLP research, microblog texts use a more informal register with nonstandard lexical items, abbreviations, and free orthographic variation.When confronted with such input, conventional text analysis tools often perform poorly.Normalization -replacing orthographically or lexically idiosyncratic forms with more standard variants -can improve performance.We propose a method for learning normalization rules from machine translations of a parallel corpus of microblog messages.To validate the utility of our approach, we evaluate extrinsically, showing that normalizing English tweets and then translating improves translation quality (compared to translating unnormalized text) using three standard web translation services as well as a phrase-based translation system trained on parallel microblog data.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso
EMNLP2
2013 Knowledge-Rich Morphological Priors for Bayesian Language Models
Victor Chahuneau, Noah A. Smith, Chris Dyer
HLT-NAACL3
2013 A Simple, Fast, and Effective Reparameterization of IBM Model 2
Chris Dyer, Victor Chahuneau, Noah A. Smith
HLT-NAACL1
2013 Large-Scale Discriminative Training for Statistical Machine Translation Using Held-Out Line Search
Jeffrey Flanigan, Chris Dyer, Jaime G. Carbonell
HLT-NAACL2
2013 Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
Olutobi Owoputi, Brendan T. O'Connor 0001, Chris Dyer, Kevin Gimpel, Nathan Schneider 0001, Noah A. Smith
HLT-NAACL3
2013 Supersense Tagging for Arabic: the MT-in-the-Middle Attack
Nathan Schneider 0001, Behrang Mohit, Chris Dyer, Kemal Oflazer, Noah A. Smith
HLT-NAACL3
2012 Joint Feature Selection in Distributed Stochastic Learning for Large-Scale Discriminative Training in SMT
Patrick Simianer, Stefan Riezler, Chris Dyer
ACL (1)3
2012 Bayesian Language Modelling of German Compounds
Jan A. Botha, Chris Dyer, Phil Blunsom
COLING2
2012 A Bayesian Model for Learning SCFGs with Discontiguous Rules
Abby D. Levenberg, Chris Dyer, Phil Blunsom
EMNLP-CoNLL2
2011 Unsupervised Word Alignment with Arbitrary Features
Chris Dyer, Jonathan H. Clark, Alon Lavie, Noah A. Smith
ACL1
2011 Predicting a Scientific Community's Response to an Article
Dani Yogatama, Michael Heilman, Brendan T. O'Connor 0001, Chris Dyer, Bryan R. Routledge, Noah A. Smith
EMNLP4
2010 Discriminative Word Alignment with a Function Word Reordering Model
Hendra Setiawan, Chris Dyer, Philip Resnik
EMNLP2
2010 Two monolingual parses are better than one (synchronous parse)
Chris Dyer
HLT-NAACL1
2010 Context-free reordering, finite-state translation
Chris Dyer, Philip Resnik
HLT-NAACL1
2010 Monte Carlo techniques for phrase-based translation
Abhishek Arun, Barry Haddow, Philipp Koehn, Adam Lopez, Chris Dyer, Phil Blunsom
Mach. Transl.5
2009 A Gibbs Sampler for Phrasal Synchronous Grammar Induction
Phil Blunsom, Trevor Cohn, Chris Dyer, Miles Osborne
ACL/IJCNLP3
2009 Efficient Minimum Error Rate Training and Minimum Bayes-Risk Decoding for Translation Hypergraphs and Lattices
Shankar Kumar, Wolfgang Macherey, Chris Dyer, Franz Josef Och
ACL/IJCNLP3
2009 Monte Carlo inference and maximization for phrase-based translation
Abhishek Arun, Chris Dyer, Barry Haddow, Phil Blunsom, Adam Lopez, Philipp Koehn
CoNLL2
2009 Using a maximum entropy model to build segmentation lattices for MT
Chris Dyer
HLT-NAACL1
2008 Generalizing Word Lattice Translation
Chris Dyer, Smaranda Muresan, Philip Resnik
ACL1
2007 Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst
ACL11