VLDB 2026 Research / reviewers in the wild / expert
Matt Post
dblp:51/8151 · also Matthew Post
· DBLP profile ↗
32ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-1297-6794ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Machine translation · 58% Language models and text generation · 14% Information extraction and text analysis · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 25 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
machine translation evaluation |
2.2 | 3 | 2026 | PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation · ACL (1) 2026 Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies · ACL (1) 2024 Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing · EMNLP (1) 2020 |
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation |
1.9 | 3 | 2026 | PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation · ACL (1) 2026 Levenshtein Training for Word-level Quality Estimation · EMNLP (1) 2021 Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.7 | 1 | 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer · EMNLP 2023 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.7 | 1 | 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer · EMNLP 2023 |
Machine learning › Representation and self-supervised learning › visual representation
pixel representation |
0.7 | 1 | 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer · EMNLP 2023 |
Natural language and speech › Machine translation › neural machine translation
open vocabulary translation |
0.5 | 1 | 2021 | Robust Open-Vocabulary Translation from Visual Text Representations · EMNLP (1) 2021 |
Natural language and speech › Machine translation
robust machine translation |
0.5 | 1 | 2021 | Robust Open-Vocabulary Translation from Visual Text Representations · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › text segmentation
sentence segmentation |
0.5 | 1 | 2021 | A unified approach to sentence segmentation of punctuated text in many languages · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
text segmentation |
0.5 | 1 | 2021 | A unified approach to sentence segmentation of punctuated text in many languages · ACL/IJCNLP (1) 2021 |
Natural language and speech › Machine translation › machine translation evaluation › translation quality estimation
word-level quality estimation |
0.5 | 1 | 2021 | Levenshtein Training for Word-level Quality Estimation · EMNLP (1) 2021 |
Natural language and speech › Machine translation › machine translation evaluation
automatic evaluation metrics |
0.4 | 1 | 2020 | Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing · EMNLP (1) 2020 |
Natural language and speech › Machine translation
low-resource machine translation |
0.4 | 1 | 2020 | Simulated multiple reference training improves low-resource machine translation · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual word alignment |
0.4 | 1 | 2019 | A Discriminative Neural Model for Cross-Lingual Word Alignment · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › decoding › constrained decoding
lexically constrained decoding |
0.4 | 1 | 2019 | PARABANK: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-Constrained Neural Machine Translation · AAAI 2019 |
Natural language and speech › Machine translation
neural machine translation |
0.4 | 1 | 2019 | PARABANK: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-Constrained Neural Machine Translation · AAAI 2019 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.4 | 1 | 2019 | PARABANK: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-Constrained Neural Machine Translation · AAAI 2019 |
Natural language and speech › Language models and text generation
decoding |
0.3 | 1 | 2026 | PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation · ACL (1) 2026 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
0.3 | 1 | 2026 | PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation · ACL (1) 2026 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural Network · AAAI 2017 |
Computer vision › Image recognition and object detection › text recognition
robust word recognition |
0.3 | 1 | 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural Network · AAAI 2017 |
Natural language and speech › Language models and text generation › text correction
spelling correction |
0.3 | 1 | 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural Network · AAAI 2017 |
Machine learning › Representation and self-supervised learning › text embedding
cross-lingual representation |
0.2 | 1 | 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer · EMNLP 2023 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
levenshtein transformer |
0.1 | 1 | 2021 | Levenshtein Training for Word-level Quality Estimation · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
text representation |
0.1 | 1 | 2021 | Robust Open-Vocabulary Translation from Visual Text Representations · EMNLP (1) 2021 |
Computational social science and digital humanities
psycholinguistics |
0.1 | 1 | 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural Network · AAAI 2017 |
Methods — techniques the papers use, named apart from their topics
subword segmentation · 1.2regularization · 1.0pairwise supervision · 1.0human judgment differences · 1.0vocabulary expansion · 0.7transfer learning · 0.5sliding window · 0.5multilingual modeling · 0.5levenshtein transformer · 0.5iterative decoding · 0.5lexically-constrained neural machine translation · 0.4semi-character recurrent neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine TranslationabstractWe present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-free machine translation (MT) evaluation as a graded pairwise comparison.Given a source segment and two candidate translations, PEAR predicts the direction and magnitude of their quality difference.The metrics are trained using pairwise supervision derived from differences in human judgments, with an additional regularization term that encourages sign inversion under candidate order reversal.On the WMT24 meta-evaluation benchmark, PEAR outperforms strictly matched single-candidate QE baselines trained with the same data and backbones, isolating the benefit of the proposed pairwise formulation.Despite using substantially fewer parameters than recent large metrics, PEAR surpasses far larger QE models and reference-based metrics.Our analysis further indicates that PEAR yields a less redundant evaluation signal relative to other top metrics.Finally, we show that PEAR is an effective utility function for minimum Bayes risk (MBR) decoding, reducing pairwise scoring cost at negligible impact. Lorenzo Proietti 0002, Roman Grundkiewicz, Matt Post |
ACL (1) | 3 |
| 2024 | Navigating the Metrics Maze: Reconciling Score Magnitudes and AccuraciesabstractTen years ago, a single metric, BLEU, governed progress in machine translation research.For better or worse, there is no such consensus today, and consequently it is difficult for researchers to develop and retain intuitions about metric deltas that drove earlier research and deployment decisions.This paper investigates the "dynamic range" of a number of modern metrics in an effort to provide a collective understanding of the meaning of differences in scores both within and among metrics; in other words, we ask what point difference x in metric y is required between two systems for humans to notice?We conduct our evaluation on a new large dataset, ToShip23, using it to discover deltas at which metrics achieve system-level differences that are meaningful to humans, which we measure by pairwise system accuracy.We additionally show that this method of establishing delta-accuracy is more stable than the standard use of statistical p-values in regards to testset size.Where data size permits, we also explore the effect of metric deltas and accuracy across finer-grained features such as translation direction, domain, and system closeness. Tom Kocmi, Vilém Zouhar, Christian Federmann, Matt Post |
ACL (1) | 4 |
| 2024 | CTC-GMM: CTC Guided Modality Matching For Fast and Accurate Streaming Speech TranslationabstractModels for streaming speech translation (ST) can achieve high accuracy and low latency if they’re developed with vast amounts of paired audio in the source language and written text in the target language. Yet, these text labels for the target language are often pseudo labels due to the prohibitive cost of manual ST data labeling. In this paper, we introduce a methodology named Connectionist Temporal Classification guided modality matching (CTC-GMM) that enhances the streaming ST model by leveraging extensive machine translation (MT) text data. This technique employs CTC to compress the speech sequence into a compact embedding sequence that matches the corresponding text sequence, allowing us to utilize matched source-target language text pairs from the MT corpora to refine the streaming ST model further. Our evaluations with FLEURS and CoVoST2 show that the CTC-GMM approach can increase translation accuracy relatively by 13.9% and 6.4% respectively, while also boosting decoding speed by 59.7% on GPU. Rui Zhao 0017, Jinyu Li 0001, Ruchao Fan, Matt Post |
SLT | 4 |
| 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferabstractWe introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations.We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings.We explore various properties of pixel representations such as parameter sharing within and across scripts to better understand where they lead to positive transfer.We observe that these properties not only enable seamless cross-lingual transfer to unseen scripts, but make pixel representations more data-efficient than alternatives such as vocabulary expansion.We hope this work contributes to more extensible multilingual models for all languages and scripts. Elizabeth Salesky, Neha Verma 0001, Philipp Koehn, Matt Post |
EMNLP | 4 |
| 2022 | Large-Scale Streaming End-to-End Speech Translation with Neural TransducersabstractNeural transducers have been widely used in automatic speech recognition (ASR). In this paper, we introduce it to streaming end-to-end speech translation (ST), which aims to convert audio signals to texts in other languages directly. Compared with cascaded ST that performs ASR followed by text-based machine translation (MT), the proposed Transformer transducer (TT)-based ST model drastically reduces inference latency, exploits speech information, and avoids error propagation from ASR to MT. To improve the modeling capacity, we propose attention pooling for the joint network in TT. In addition, we extend TT-based ST to multilingual ST, which generates texts of multiple languages at the same time. Experimental results on a large-scale 50 thousand (K) hours pseudo-labeled training set show that TT-based ST not only significantly reduces inference time but also outperforms non-streaming cascaded ST for English-German translation. Jinyu Li 0001, Matt Post, Yashesh Gaur |
INTERSPEECH | 4 |
| 2021 | A unified approach to sentence segmentation of punctuated text in many languagesabstractRachel Wicks, Matt Post. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rachel Wicks, Matt Post |
ACL/IJCNLP (1) | 2 |
| 2021 | Levenshtein Training for Word-level Quality EstimationabstractWe propose a novel scheme to use the Levenshtein Transformer to perform the task of word-level quality estimation.A Levenshtein Transformer is a natural fit for this task: trained to perform decoding in an iterative manner, a Levenshtein Transformer can learn to post-edit without explicit supervision.To further minimize the mismatch between the translation task and the word-level QE task, we propose a two-stage transfer learning procedure on both augmented data and human postediting data.We also propose heuristics to construct reference labels that are compatible with subword-level finetuning and inference.Results on WMT 2020 QE shared task dataset show that our proposed method has superior data efficiency under the data-constrained setting and competitive performance under the unconstrained setting.* Shuoyang Ding had a part-time affiliation with Microsoft at the time of this work. Shuoyang Ding, Marcin Junczys-Dowmunt, Matt Post, Philipp Koehn |
EMNLP (1) | 3 |
| 2021 | Robust Open-Vocabulary Translation from Visual Text RepresentationsabstractMachine translation models have discrete vo cabularies and commonly use subword seg mentation techniques to achieve an 'open vo cabulary.'This approach relies on consis tent and correct underlying unicode sequences, and makes models susceptible to degrada tion from common types of noise and vari ation.Motivated by the robustness of hu man language processing, we propose the use of visual text representations, which dispense with a finite set of text embeddings in favor of continuous vocabularies created by process ing visually rendered text with sliding win dows.We show that models using visual text representations approach or match per formance of traditional text models on small and larger datasets.More importantly, mod els with visual embeddings demonstrate sig nificant robustness to varied types of noise, achieving e.g., 25.9 BLEU on a character per muted German-English task where subword models degrade to 1.9. Elizabeth Salesky, David Etter, Matt Post |
EMNLP (1) | 3 |
| 2021 | The Multilingual TEDx Corpus for Speech Recognition and TranslationabstractWe present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source languages. We segment transcripts into sentences and align them to the source-language audio and target-language translations. The corpus is released along with open-sourced code enabling extension to new talks and languages as they become available. Our corpus creation methodology can be applied to more languages than previous work, and creates multi-way parallel evaluation sets. We provide baselines in multiple ASR and ST settings, including multilingual models to improve translation performance for low-resource language pairs. Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni, Matteo Negri, Marco Turchi, Douglas W. Oard, Matt Post |
Interspeech | 8 |
| 2020 | Simulated multiple reference training improves low-resource machine translationabstractMany valid translations exist for a given sentence, yet machine translation (MT) is trained with a single reference translation, exacerbating data sparsity in low-resource settings. We introduce Simulated Multiple Reference Training (SMRT), a novel MT training method that approximates the full space of possible translations by sampling a paraphrase of the reference sentence from a paraphraser and training the MT model to predict the paraphraser's distribution over possible tokens. We demonstrate the effectiveness of SMRT in low-resource settings when translating to English, with improvements of 1.2 to 7.0 BLEU. We also find SMRT is complementary to back-translation. Huda Khayrallah, Brian Thompson 0001, Matt Post, Philipp Koehn |
EMNLP (1) | 3 |
| 2020 | Automatic Machine Translation Evaluation in Many Languages via Zero-Shot ParaphrasingabstractWe frame the task of machine translation evaluation as one of scoring machine translation output with a sequence-to-sequence paraphraser, conditioned on a human reference.We propose training the paraphraser as a multilingual NMT system, treating paraphrasing as a zero-shot translation task (e.g., Czech to Czech).This results in the paraphraser's output mode being centered around a copy of the input sequence, which represents the best case scenario where the MT system output matches a human reference.Our method is simple and intuitive, and does not require human judgements for training.Our single model (trained in 39 languages) outperforms or statistically ties with all prior metrics on the WMT 2019 segment-level shared metrics task in all languages (excluding Gujarati where the model had no training data).We also explore using our model for the task of quality estimation as a metric-conditioning on the source instead of the reference-and find that it significantly outperforms every submission to the WMT 2019 shared task on quality estimation in every language pair. Word-level paraphraser log probabilities H(out|in) sBLEU LASERCopy Jason went to school at the University of Madrid . -0.08 -0.26 -0.16 -0.16 -0.12 -0.11 -0.14 -0.10 -0.10 -0.11 -0.10 -0.13 100.0 1.000 Disfluent Jason went school at Brian Thompson 0001, Matt Post |
EMNLP (1) | 2 |
| 2020 | Benchmarking Neural and Statistical Machine Translation on Low-Resource African LanguagesabstractResearch in machine translation (MT) is developing at a rapid pace. However, most work in the community has focused on languages where large amounts of digital resources are available. In this study, we benchmark state of the art statistical and neural machine translation systems on two African languages which do not have large amounts of resources: Somali and Swahili. These languages are of social importance and serve as test-beds for developing technologies that perform reasonably well despite the low-resource constraint. Our findings suggest that statistical machine translation (SMT) and neural machine translation (NMT) can perform similarly in low-resource scenarios, but neural systems require more careful tuning to match performance. We also investigate how to exploit additional data, such as bilingual text harvested from the web, or user dictionaries; we find that NMT can significantly improve in performance with the use of these additional data. Finally, we survey the landscape of machine translation resources for the languages of Africa and provide some suggestions for promising future research directions. Kevin Duh, Paul McNamee, Matt Post, Brian Thompson 0001 |
LREC | 3 |
| 2020 | The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological ExplorationabstractWe present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible and counting. Our data is derived from scraping several online resources and merging them with existing corpora, combining them under a common scheme that is verse-parallel across all translations. We detail our effort to scrape, clean, align, and utilize this ripe multilingual dataset. The corpus captures the great typological variety of the world’s languages. We catalog this by showing highly similar proportions of representation of Ethnologue’s typological features in our corpus. We also give an example application: projecting pronoun features like clusivity across alignments to richly annotate languages which do not mark the distinction. Arya McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky |
LREC | 8 |
| 2020 | Membership Inference Attacks on Sequence-to-Sequence Models: Is My Data In Your Machine Translation System?abstractData privacy is an important issue for “machine learning as a service” providers. We focus on the problem of membership inference attacks: Given a data sample and black-box access to a model’s API, determine whether the sample existed in the model’s training data. Our contribution is an investigation of this problem in the context of sequence-to-sequence models, which are important in applications such as machine translation and video captioning. We define the membership inference problem for sequence generation, provide an open dataset based on state-of-the-art machine translation models, and report initial results on whether these models leak private information against several kinds of membership inference attacks. Sorami Hisamoto, Matt Post, Kevin Duh |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | PARABANK: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-Constrained Neural Machine TranslationabstractWe present PARABANK, a large-scale English paraphrase dataset that surpasses prior work in both quantity and quality. Following the approach of PARANMT (Wieting and Gimpel, 2018), we train a Czech-English neural machine translation (NMT) system to generate novel paraphrases of English reference sentences. By adding lexical constraints to the NMT decoding procedure, however, we are able to produce multiple high-quality sentential paraphrases per source sentence, yielding an English paraphrase resource with more than 4 billion generated tokens and exhibiting greater lexical diversity. Using human judgments, we also demonstrate that PARABANK’s paraphrases improve over PARANMT on both semantic similarity and fluency. Finally, we use PARABANK to train a monolingual NMT model with the same support for lexically-constrained decoding for sentence rewriting tasks. Edward J. Hu, Rachel Rudinger, Matt Post, Benjamin Van Durme |
AAAI | 3 |
| 2019 | Large-Scale, Diverse, Paraphrastic Bitexts via Sampling and ClusteringabstractProducing diverse paraphrases of a sentence is a challenging task.Natural paraphrase corpora are scarce and limited, while existing large-scale resources are automatically generated via back-translation and rely on beam search, which tends to lack diversity.We describe PARABANK 2, a new resource that contains multiple diverse sentential paraphrases, produced from a bilingual corpus using negative constraints, inference sampling, and clustering.We show that PARABANK 2 significantly surpasses prior work in both lexical and syntactic diversity while being meaningpreserving, as measured by human judgments and standardized metrics.Further, we illustrate how such paraphrastic resources may be used to refine contextualized encoders, leading to improvements in downstream tasks. Edward J. Hu, Nils Holzenberger, Matt Post, Benjamin Van Durme |
CoNLL | 4 |
| 2019 | A Discriminative Neural Model for Cross-Lingual Word AlignmentabstractElias Stengel-Eskin, Tzu-ray Su, Matt Post, Benjamin Van Durme. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Elias Stengel-Eskin, Tzu-Ray Su, Matt Post, Benjamin Van Durme |
EMNLP/IJCNLP (1) | 3 |
| 2019 | An Exploration of Placeholding in Neural Machine Translation
Matt Post, Shuoyang Ding, Marianna J. Martindale, Winston Wu |
MTSummit (1) | 1 |
| 2019 | An Interactive Teaching Tool for Introducing Novices to Machine TranslationabstractThe first step in the research process is developing an understanding of the problem at hand. Novices may be interested in learning about machine translation (MT), but often lack experience and intuition about the task of translation (either by human or machine) and its challenges. The goal of this work is to allow students to interactively discover why MT is an open problem, and encourage them to ask questions, propose solutions, and test intuitions. We present a hands-on activity in which students build and evaluate their own MT systems using curated parallel texts. By having students hand-engineer MT system rules in a simple user interface, which they can then run on real data, they gain intuition about why early MT research took this approach, where it fails, and what features of language make MT a challenging problem even today. Developing translation rules typically strikes novices as an obvious approach that should succeed, but the idea quickly struggles in the face of natural language complexity. This interactive, intuition-building exercise can be augmented by a discussion of state-of-the-art MT techniques and challenges, focusing on areas or aspects of linguistic complexity that the students found difficult. We envision this lesson plan being used in the framework of a larger AI or natural language processing course (where only a small amount of time can be dedicated to MT) or as a standalone activity. We describe and release the tool that supports this lesson, as well as accompanying data. Huda Khayrallah, Rebecca Knowles, Kevin Duh, Matt Post |
SIGCSE | 4 |
| 2018 | Fast Lexically Constrained Decoding with Dynamic Beam Allocation for Neural Machine TranslationabstractThe end-to-end nature of neural machine translation (NMT) removes many ways of manually guiding the translation process that were available in older paradigms.Recent work, however, has introduced a new capability: lexically constrained or guided decoding, a modification to beam search that forces the inclusion of pre-specified words and phrases in the output.However, while theoretically sound, existing approaches have computational complexities that are either linear (Hokamp and Liu, 2017) or exponential (Anderson et al., 2017) in the number of constraints.We present an algorithm for lexically constrained decoding with a complexity of O(1) in the number of constraints.We demonstrate the algorithm's remarkable ability to properly place these constraints, and use it to explore the shaky relationship between model and BLEU scores.Our implementation is available as part of SOCKEYE. Matt Post, David Vilar |
NAACL-HLT | 1 |
| 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural NetworkabstractLanguage processing mechanism by humans is generally more robust than computers. The Cmabrigde Uinervtisy (Cambridge University) effect from the psycholinguistics literature has demonstrated such a robust word processing mechanism, where jumbled words (e.g. Cmabrigde / Cambridge) are recognized with little cost. On the other hand, computational models for word recognition (e.g. spelling checkers) perform poorly on data with such noise. Inspired by the findings from the Cmabrigde Uinervtisy effect, we propose a word recognition model based on a semi-character level recurrent neural network (scRNN). In our experiments, we demonstrate that scRNN has significantly more robust performance in word spelling correction (i.e. word recognition) compared to existing spelling checkers and character-based convolutional neural network. Furthermore, we demonstrate that the model is cognitively plausible by replicating a psycholinguistics experiment about human reading difficulty using our model. Keisuke Sakaguchi, Kevin Duh, Matt Post, Benjamin Van Durme |
AAAI | 3 |
| 2017 | Syntax-Based Statistical Machine Translation Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn (University of Edinburgh, University of Edinburgh, Johns Hopkins University, Johns Hopkins University), edited by Graeme Hirst, volume 33), 2016, xvii+190 pp; paperback, ISBN 978-1-62705-900-8; ebook, ISBN 978-1-62705-502-4; doi: 10.2200/S00716ED1V04Y201604HLT033, $70abstractIn its early development, machine translation adopted rule-based approaches, which can include the use of language syntax. The late 1980s and early 1990s saw the inception of the statistical machine translation (SMT) approach, where translation models can be learned automatically from a parallel corpus rather than created manually by humans. Initial SMT models were word-based and phrase-based, without the use of syntactic knowledge. In phrase-based SMT, a source sentence is first segmented into phrases and then translated phrase-by-phrase with some reordering of the translated phrases in the target sentence. This has posed challenges when translating between two syntactically different languages. Syntax-based SMT approaches take advantage of syntactic knowledge within the framework of SMT. This book provides an introduction to syntax-based SMT approaches. It is a valuable resource for those who are interested in syntax-based SMT.The book consists of seven chapters. There is not an introduction chapter in this book, aside from the preface, which can be considered as a brief introduction. Readers are referred to Koehn (2010) for background knowledge. I think an introduction chapter categorized into sections would have been useful, before proceeding to describe the various models. The first two chapters provide principles applicable across various syntax-based SMT approaches. The next three chapters describe syntax-based SMT decoding in detail; this constitutes half of the book. Selected extended topics are provided in the next chapter, which is followed by a concluding chapter.Chapter 1 describes the models and formalisms applicable to syntax-based SMT. The first section describes the phrasal translation units in phrase-based SMT, its limitations, and how tree structures address the limitations of the phrase-based approach. This explanation is useful as translation units are the key difference between the phrase-based and syntax-based SMT approaches. The next two sections describe the grammar formalisms and the statistical models that define syntax-based SMT. The section that covers the grammar formalisms (i.e., synchronous context-free grammar [SCFG] and synchronous tree-substitution grammar [STSG]), would have been clearer if their differences were presented in a side-by-side illustrating example. The remainder of the chapter discusses different categories of syntax-based SMT approaches and the history of these approaches, which include string-to-string, string-to-tree, tree-to-string, and tree-to-tree SMT approaches. Although the syntax-based translation model in Galley et al. (2006) falls under the string-to-tree category, I wonder why hierarchical phrase-based SMT, or Hiero (Chiang, 2007), is not explicitly put under the string-to-string category, since Hiero also uses “unlabeled hierarchical phrases where there is no representation of linguistic categories.”Chapter 2 focuses on how the statistical framework of a syntax-based SMT approach learns its model from a word-aligned and parsed parallel text. The first section explains how phrase pairs are extracted as translation rules from a word-aligned sentence pair in phrase-based SMT (Koehn, Och, and Marcu, 2003), highlighting the definition of a phrase as a sequence of words and the alignment-consistency property of a phrase pair as defined in Och and Ney (2004). The remainder of the chapter introduces three predominant instantiations of syntax-based models: hierarchical phrase-based SMT (Hiero) (Chiang, 2007), which is a non-labeled syntax-based SMT approach arising from the phrase-based approach; syntax-augmented machine translation (SAMT), which introduces the notion of soft labels while keeping the nonlinguistic phrase notion; and GHKM (Galley et al., 2004), which only extracts translation rules consistent with constituency parse subtrees. This chapter is nicely organized and it is easy to follow the gradual evolution from phrase-based SMT to GHKM.Chapter 3 introduces the decoding formalism in the form of a directed hypergraph, defined as a set of vertices and a set of directed hyperedges. The first section introduces the notion of a weighted parse forest represented in a weighted hypergraph, representing alternative parse trees of a sentence. I found it important to pay careful attention to this section, in order to understand the next section and the following chapters. The next section presents various algorithms on a hypergraph to translate a sentence in a hypergraph representation of possible tree derivations. Overall, I found this chapter to contain many technical details. The last section of this chapter provides historical notes on the sources of these concepts. This chapter needs to be read before the next chapter, which assumes understanding of the concepts introduced in Chapter 3.Chapter 4 describes tree decoding—that is, decoding with the constituency parse tree of a source sentence as its input, focusing on the tree-to-string approach. The first two sections highlight decoding with local and non-local features, where non-local features accommodate n-gram language models and are more complex than local features. The next section is devoted to an in-depth description of a beam search algorithm on the parse tree of a source sentence. The description could have been improved if the running example showed the decoding steps. The next two sections present extensions to the concepts introduced in the earlier part of this chapter, by providing references to more efficient hypergraph operations. The content of this section requires readers who are interested in implementing an efficient tree-based algorithm to go through the cited references. Brief historical notes conclude this chapter nicely, by pointing to relevant materials for further reading.Chapter 5 describes string decoding with a source sentence string as its input. The first two sections describe beam search decoding algorithms in a binary SCFG, namely, a maximum of two non-terminal symbols on the right-hand side of each rule, adopted in Hiero and SAMT. The algorithms covered are a basic algorithm and an optimized algorithm. The complexity comparison between the two is nicely presented here, emphasizing the complexity reduction achieved by algorithm optimization. The handling of non-binary rules is described in the following section, illustrated by GHKM rule extraction. A mid-chapter summary section divides this chapter into two parts: beam search decoding and parsing. The second part describes parsing algorithms in the context of shared-category SCFG, assuming the same set of non-terminal symbols for the left-hand and right-hand sides of a rule, followed by a section extending the algorithm to STSG and distinct-category SCFG. The organization of this chapter is excellent. However, I feel that the inclusion of distinct-category SCFG decoding does not fit well into this chapter, as string decoding in string-to-tree SMT requires no knowledge of the source syntax. The historical notes also do not provide any references of prior work on string decoding using distinct-category SCFG.Chapter 6 contains various selected topics on syntax-based SMT. The first section discusses tree transformations, which make translation rule learning more effective. The description of non-context-free models serves as a prelude to the next section on dependency-based SMT, which covers dependency treelet (equivalent to the tree-to-string approach) and string-to-dependency (equivalent to the string-to-tree approach). The next section focuses on the ability of syntax-based SMT to have a more grammatical output compared with phrase-based SMT, although there is still room for improvement, including the use of unification grammars and semantic properties. Finally, the last section of this chapter explains how MT evaluation benefits from syntax-based SMT principles. Overall, this chapter enriches readers' knowledge beyond basic syntax-based SMT in the earlier chapters. I would also suggest the inclusion of phrase-based decoding approaches that use syntax-based features (Cherry, 2008; Chang et al., 2009).Chapter 7 nicely concludes this book by discussing the comparison between phrase-based and syntax-based SMT approaches and proposing possible future developments of syntax-based SMT. The chapter also highlights that syntax-driven MT predates statistical MT, as I mentioned at the beginning of this review.Overall, I found this book to be a useful reference book for those interested in syntax-based SMT. The book is well organized, which makes it easy for readers to refer to specific aspects of syntax-based SMT. An improvement can be made to the presentation of ideas in this book. Throughout the book, there are many technical keywords, resulting from the complexity of syntax-based SMT. It would be useful to highlight these keywords in a side bar to remind readers that they are important keywords. In addition, although examples are given throughout the book, it would be even more useful to use these examples to illustrate how the algorithms work, so that readers can gain a better understanding of the algorithms. Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn, Graeme Hirst, Christian Hadiwinoto |
Comput. Linguistics | 3 |
| 2016 | Reassessing the Goals of Grammatical Error Correction: Fluency Instead of GrammaticalityabstractThe field of grammatical error correction (GEC) has grown substantially in recent years, with research directed at both evaluation metrics and improved system performance against those metrics. One unvisited assumption, however, is the reliance of GEC evaluation on error-coded corpora, which contain specific labeled corrections. We examine current practices and show that GEC’s reliance on such corpora unnaturally constrains annotation and automatic evaluation, resulting in (a) sentences that do not sound acceptable to native speakers and (b) system rankings that do not correlate with human judgments. In light of this, we propose an alternate approach that jettisons costly error coding in favor of unannotated, whole-sentence rewrites. We compare the performance of existing metrics over different gold-standard annotations, and show that automatic evaluation with our new annotation scheme has very strong correlation with expert rankings (ρ = 0.82). As a result, we advocate for a fundamental and necessary shift in the goal of GEC, from correcting small, labeled error types, to producing text that has native fluency. Keisuke Sakaguchi, Courtney Napoles, Matt Post, Joel R. Tetreault |
Trans. Assoc. Comput. Linguistics | 3 |
| 2014 | Some insights from translating conversational telephone speechabstractWe report insights from translating Spanish conversational telephone speech into English text by cascading an automatic speech recognition (ASR) system with a statistical machine translation (SMT) system. The key new insight is that the informal register of conversational speech is a greater challenge for ASR than for SMT: the BLEU score for translating the reference transcript is 64%, but drops to 32% for translating automatic transcripts, whose word error rate (WER) is 40%. Several strategies are examined to mitigate the impact of ASR errors on the SMT output: (i) providing the ASR lattice, instead of the 1-best output, as input to the SMT system, (ii) training the SMT system on Spanish ASR output paired with English text, instead of Spanish reference transcripts, and (iii) improving the core ASR system. Each leads to consistent and complementary improvements in the SMT output. Compared to translating the 1-best output of an ASR system with 40% WER using an SMT system trained on Spanish reference transcripts, translating the output lattice of a better ASR system with 35% WER using an SMT system trained on ASR output improves BLEU from 32% to 38%. Matt Post, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 2 |
| 2014 | A Wikipedia-based Corpus for Contextualized Machine Translation
Jennifer Drexler Fox, Pushpendre Rastogi, Jacqueline Aguilar, Benjamin Van Durme, Matt Post |
LREC | 5 |
| 2014 | The Language Demographics of Amazon Mechanical TurkabstractWe present a large scale study of the languages spoken by bilingual workers on Mechanical Turk (MTurk). We establish a methodology for determining the language skills of anonymous crowd workers that is more robust than simple surveying. We validate workers’ self-reported language skill claims by measuring their ability to correctly translate words, and by geolocating workers to see if they reside in countries where the languages are likely to be spoken. Rather than posting a one-off survey, we posted paid tasks consisting of 1,000 assignments to translate a total of 10,000 words in each of 100 languages. Our study ran for several months, and was highly visible on the MTurk crowdsourcing platform, increasing the chances that bilingual workers would complete it. Our study was useful both to create bilingual dictionaries and to act as census of the bilingual speakers on MTurk. We use this data to recommend languages with the largest speaker populations as good candidates for other researchers who want to develop crowdsourced, multilingual technologies. To further demonstrate the value of creating data via crowdsourcing, we hire workers to create bilingual parallel corpora in six Indian languages, and use them to train statistical machine translation systems. Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, Chris Callison-Burch |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Learning to translate with products of novices: a suite of open-ended challenge problems for teaching MTabstractMachine translation (MT) draws from several different disciplines, making it a complex subject to teach. There are excellent pedagogical texts, but problems in MT and current algorithms for solving them are best learned by doing. As a centerpiece of our MT course, we devised a series of open-ended challenges for students in which the goal was to improve performance on carefully constrained instances of four key MT tasks: alignment, decoding, evaluation, and reranking. Students brought a diverse set of techniques to the problems, including some novel solutions which performed remarkably well. A surprising and exciting outcome was that student solutions or their combinations fared competitively on some tasks, demonstrating that even newcomers to the field can help improve the state-of-the-art on hard NLP problems while simultaneously learning a great deal. The problems, baseline code, and results are freely available. Adam Lopez, Matt Post, Chris Callison-Burch, Jonathan Weese, Juri Ganitkevitch, Narges Ahmidi, Olivia Buzek, Leah Hanson, Beaniesh Jamil, Matthias A. Lee, Ya-Ting Lin, Henry Pao, Fatima Rivera, Leili Shahriyari, Debu Sinha, Adam R. Teichert, Stephen Wampler, Michael Weinberger, Daguang Xu, Lin Yang 0002, Shang Zhao 0002 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 21 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 18 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 18 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 19 |
| 2012 | Stylometric Analysis of Scientific Articles
Shane Bergsma, Matt Post, David Yarowsky |
HLT-NAACL | 2 |