EDBT 2026 Demo / reviewers in the wild / expert
Chris Quirk
dblp:63/7183 · also Christopher Quirk
· DBLP profile ↗
50ranked-venue papers
11as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 11 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
24 papers |
Question answering and dialogue systems · 26% Machine translation · 17% Language models and text generation · 14% | |
| Databases, data mining, and information retrieval
4 papers |
Knowledge graphs · 72% Information retrieval · 21% Query processing and optimization · 8% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 100% |
Topics — the 30 heaviest of 48, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
user simulation |
0.9 | 1 | 2025 | SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants? · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.6 | 3 | 2018 | Confidence Modeling for Neural Semantic Parsing · ACL (1) 2018 Improved Semantic Parsers For If-Then Statements · ACL (1) 2016 Language to Code: Learning Semantic Parsers for If-This-Then-That Recipes · ACL (1) 2015 |
Knowledge graphs
knowledge graph embedding |
0.6 | 2 | 2019 | Embedding Edge-attributed Relational Hierarchies · SIGIR 2019 Compositional Learning of Embeddings for Relation Paths in Knowledge Base and Text · ACL (1) 2016 |
Natural language and speech › Language models and text generation
controllable text generation |
0.5 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
knowledge-grounded response generation |
0.5 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Machine translation
statistical machine translation |
0.4 | 4 | 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014 Bayesian Learning of Non-Compositional Phrases with Synchronous Parsing · ACL 2008 The impact of parse quality on syntactically-informed statistical machine translation · EMNLP 2006 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.4 | 1 | 2019 | Context-Aware Intent Identification in Email Conversations · SIGIR 2019 |
Machine learning › Deep learning architectures and training
positional encoding |
0.4 | 1 | 2019 | Novel positional encodings to enable tree-based transformers · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
transformer |
0.4 | 1 | 2019 | Novel positional encodings to enable tree-based transformers · NeurIPS 2019 |
Program synthesis and code generation
code translation |
0.4 | 1 | 2019 | Novel positional encodings to enable tree-based transformers · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
0.3 | 1 | 2018 | Confidence Modeling for Neural Semantic Parsing · ACL (1) 2018 |
Machine learning › Trustworthy machine learning › uncertainty estimation
confidence estimation |
0.3 | 1 | 2018 | Confidence Modeling for Neural Semantic Parsing · ACL (1) 2018 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.3 | 1 | 2018 | Confidence Modeling for Neural Semantic Parsing · ACL (1) 2018 |
Machine learning › Optimization for machine learning › training criteria
minimum error rate training |
0.3 | 2 | 2013 | Regularized Minimum Error Rate Training · EMNLP 2013 Optimal Search for Minimum Error Rate Training · EMNLP 2011 |
Knowledge graphs
link prediction |
0.2 | 1 | 2016 | Compositional Learning of Embeddings for Relation Paths in Knowledge Base and Text · ACL (1) 2016 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.2 | 1 | 2015 | Pre-Computable Multi-Layer Neural Network Language Models · EMNLP 2015 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2015 | Pre-Computable Multi-Layer Neural Network Language Models · EMNLP 2015 |
Natural language and speech › Language models and text generation
neural language model |
0.2 | 1 | 2015 | Pre-Computable Multi-Layer Neural Network Language Models · EMNLP 2015 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.2 | 3 | 2008 | Syntactic Models for Structural Word Insertion and Deletion during Translation · EMNLP 2008 The impact of parse quality on syntactically-informed statistical machine translation · EMNLP 2006 Dependency Treelet Translation: Syntactically Informed Phrasal SMT · ACL 2005 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.2 | 1 | 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014 |
Natural language and speech › Machine translation › statistical machine translation
translation rule extraction |
0.2 | 1 | 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014 |
Natural language and speech › Question answering and dialogue systems
dialogue system training |
0.2 | 1 | 2013 | Lightly Supervised Learning of Procedural Dialog Systems · ACL (1) 2013 |
Natural language and speech › Machine translation
neural machine translation |
0.2 | 1 | 2013 | Joint Language and Translation Modeling with Recurrent Neural Networks · EMNLP 2013 |
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model |
0.2 | 1 | 2013 | Joint Language and Translation Modeling with Recurrent Neural Networks · EMNLP 2013 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.2 | 1 | 2013 | Lightly Supervised Learning of Procedural Dialog Systems · ACL (1) 2013 |
Mathematical optimization
regularization |
0.2 | 1 | 2013 | Regularized Minimum Error Rate Training · EMNLP 2013 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › grounding
knowledge grounding |
0.1 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.1 | 1 | 2011 | Gappy Phrasal Alignment By Agreement · ACL 2011 |
Mathematical optimization
line search |
0.1 | 1 | 2011 | Optimal Search for Minimum Error Rate Training · EMNLP 2011 |
Machine learning › Graph learning › relation modeling
relational inference |
0.1 | 1 | 2019 | Embedding Edge-attributed Relational Hierarchies · SIGIR 2019 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.3user simulator · 0.9translation-based embedding · 0.8positional encoding · 0.8hyperbolic embedding · 0.8inductive attention · 0.5control phrase prediction · 0.5deep learning · 0.4context modeling · 0.4classical machine learning · 0.4dynamic programming · 0.2compositional path representation · 0.2text mining · 0.2automatic curation · 0.2l2 regularization · 0.2l0 regularization · 0.2gradient of expected BLEU · 0.2exact search · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?abstractYao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu 0004, Jianfeng Gao 0001 |
EMNLP | 7 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 6 |
| 2021 | Text Editing by CommandabstractFelix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao 0001, William B. Dolan |
NAACL-HLT | 5 |
| 2019 | Novel positional encodings to enable tree-based transformersabstractNeural models optimized for tree-based problems are of great value in tasks like SQL query extraction and program synthesis. On sequence-structured data, transformers have been shown to learn relationships across arbitrary pairs of positions more reliably than recurrent models. Motivated by this property, we propose a method to extend transformers to tree-structured data, enabling sequence-to-tree, tree-to-sequence, and tree-to-tree mappings. Our approach abstracts the transformer's sinusoidal positional encodings, allowing us to instead use a novel positional encoding scheme to represent node positions within trees. We evaluated our model in tree-to-tree program translation and sequence-to-tree semantic parsing settings, achieving superior performance over both sequence-to-sequence transformers and state-of-the-art tree-based LSTMs on several datasets. In particular, our results include a 22% absolute increase in accuracy on a JavaScript to CoffeeScript translation dataset. Vighnesh Leonardo Shiv, Chris Quirk |
NeurIPS | 2 |
| 2019 | Embedding Edge-attributed Relational HierarchiesabstractRelational embedding methods encode objects and their relations as low-dimensional vectors. While achieving competitive performance on a variety of relational inference tasks, these methods fall short of preserving the hierarchies that are often formed in existing graph data, and ignore the rich edge attributes that describe the relation facts. In this paper, we propose a novel embedding method that simultaneously preserve the hierarchical property and the edge information in the edge-attributed relational hierarchies. The proposed method preserves the hierarchical relations by leveraging the non-linearity of hyperbolic vector translations, for which the edge attributes are exploited to capture the importance of each relation fact. Our experiment is conducted on the well-known Enron organizational chart, where the supervision relations between employees of the Enron company are accompanied with email-based attributes. We show that our method produces relational embeddings of higher quality than state-of-the-art methods, and outperforms a variety of strong baselines in reconstructing the organizational chart. Muhao Chen 0001, Chris Quirk |
SIGIR | 2 |
| 2019 | Context-Aware Intent Identification in Email ConversationsabstractEmail continues to be one of the most important means of online communication. People spend a significant amount of time sending, reading, searching and responding to email in order to manage tasks, exchange information, etc. In this paper, we study intent identification in workplace email. We use a large scale publicly available email dataset to characterize intents in enterprise email and propose methods for improving intent identification in email conversations. Previous work focused on classifying email messages into broad topical categories or detecting sentences that contain action items or follow certain speech acts. In this work, we focus on sentence-level intent identification and study how incorporating more context (such as the full message body and other metadata) could improve the performance of the intent identification models. We experiment with several models for leveraging context including both classical machine learning and deep learning approaches. We show that modeling the interaction between sentence and context can significantly improve the performance. Wei Wang 0238, Saghar Hosseini, Ahmed Awadallah 0001, Paul N. Bennett, Chris Quirk |
SIGIR | 5 |
| 2018 | Confidence Modeling for Neural Semantic ParsingabstractIn this work we focus on confidence modeling for neural semantic parsers which are built upon sequence-to-sequence models.We outline three major causes of uncertainty, and design various metrics to quantify these factors.These metrics are then used to estimate confidence scores that indicate whether model predictions are likely to be correct.Beyond confidence estimation, we identify which parts of the input contribute to uncertain predictions allowing users to interpret their model, and verify or refine its input.Experimental results show that our confidence model significantly outperforms a widely used method that relies on posterior probability, and improves the quality of interpretation compared to simply relying on attention scores. Li Dong 0004, Chris Quirk, Mirella Lapata |
ACL (1) | 2 |
| 2017 | Distant Supervision for Relation Extraction beyond the Sentence BoundaryabstractThe growing demand for structured knowledge has led to great interest in relation extraction, especially in cases with limited supervision.However, existing distance supervision approaches only extract relations expressed in single sentences.In general, cross-sentence relation extraction is under-explored, even in the supervised-learning setting.In this paper, we propose the first approach for applying distant supervision to crosssentence relation extraction.At the core of our approach is a graph representation that can incorporate both standard dependencies and discourse relations, thus providing a unifying way to model relations within and across sentences.We extract features from multiple paths in this graph, increasing accuracy and robustness when confronted with linguistic variation and analysis error.Experiments on an important extraction task for precision medicine show that our approach can learn an accurate cross-sentence extractor, using only a small existing knowledge base and unlabeled text from biomedical research articles.Compared to the existing distant supervision paradigm, our approach extracted twice as many relations at similar precision, thus demonstrating the prevalence of cross-sentence relations and the promise of our approach. Chris Quirk, Hoifung Poon |
EACL (1) | 1 |
| 2017 | Cross-Sentence N-ary Relation Extraction with Graph LSTMsabstractPast work in relation extraction has focused on binary relations in single sentences. Recent NLP inroads in high-value domains have sparked interest in the more general setting of extracting n-ary relations that span multiple sentences. In this paper, we explore a general relation extraction framework based on graph long short-term memory networks (graph LSTMs) that can be easily extended to cross-sentence n-ary relation extraction. The graph formulation provides a unified way of exploring different LSTM approaches and incorporating various intra-sentential and inter-sentential dependencies, such as sequential, syntactic, and discourse relations. A robust contextual representation is learned for the entities, which serves as input to the relation classifier. This simplifies handling of relations with arbitrary arity, and enables multi-task learning with related relations. We evaluate this framework in two important precision medicine settings, demonstrating its effectiveness with both conventional supervised learning and distant supervision. Cross-sentence extraction produced larger knowledge bases. and multi-task learning significantly improved extraction accuracy. A thorough analysis of various LSTM approaches yielded useful insight the impact of linguistic analysis on extraction accuracy. Nanyun Peng 0001, Hoifung Poon, Chris Quirk, Kristina Toutanova, Scott Yih |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | Improved Semantic Parsers For If-Then StatementsabstractDigital personal assistants are becoming both more common and more useful.The major NLP challenge for personal assistants is machine understanding: translating natural language user commands into an executable representation.This paper focuses on understanding rules written as If-Then statements, though the techniques should be portable to other semantic parsing tasks.We view understanding as structure prediction and show improved models using both conventional techniques and neural network models.We also discuss various ways to improve generalization and reduce overfitting: synthetic training data from paraphrase, grammar combinations, feature selection and ensembles of multiple systems.An ensemble of these techniques achieves a new state of the art result with 8% accuracy improvement. Islam Beltagy, Chris Quirk |
ACL (1) | 2 |
| 2016 | Compositional Learning of Embeddings for Relation Paths in Knowledge Base and TextabstractModeling relation paths has offered significant gains in embedding models for knowledge base (KB) completion.However, enumerating paths between two entities is very expensive, and existing approaches typically resort to approximation with a sampled subset.This problem is particularly acute when text is jointly modeled with KB relations and used to provide direct evidence for facts mentioned in it.In this paper, we propose the first exact dynamic programming algorithm which enables efficient incorporation of all relation paths of bounded length, while modeling both relation types and intermediate nodes in the compositional path representations.We conduct a theoretical analysis of the efficiency gain from the approach.Experiments on two datasets show that it addresses representational limitations in prior approaches and improves accuracy in KB completion. Kristina Toutanova, Xi Victoria Lin, Scott Yih, Hoifung Poon, Chris Quirk |
ACL (1) | 5 |
| 2015 | Language to Code: Learning Semantic Parsers for If-This-Then-That RecipesabstractChris Quirk, Raymond Mooney, Michel Galley. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Chris Quirk, Raymond J. Mooney, Michel Galley |
ACL (1) | 1 |
| 2015 | Pre-Computable Multi-Layer Neural Network Language ModelsabstractIn the last several years, neural network models have significantly improved accuracy in a number of NLP tasks.However, one serious drawback that has impeded their adoption in production systems is the slow runtime speed of neural network models compared to alternate models, such as maximum entropy classifiers.In Devlin et al. (2014), the authors presented a simple technique for speeding up feed-forward embedding-based neural network models, where the dot product between each word embedding and part of the first hidden layer are pre-computed offline.However, this technique cannot be used for hidden layers beyond the first.In this paper, we explore a neural network architecture where the embedding layer feeds into multiple hidden layers that are placed "next to" one another so that each can be pre-computed independently.On a large scale language modeling task, this architecture achieves a 10x speedup at runtime and a significant reduction in perplexity when compared to a standard multilayer network. Jacob Devlin, Chris Quirk, Arul Menezes |
EMNLP | 2 |
| 2015 | An AMR parser for English, French, German, Spanish and Japanese and a new AMR-annotated corpusabstractIn this demonstration, we will present our online parser 1 that allows users to submit any sentence and obtain an analysis following the specification of AMR (Banarescu et al., 2014) to a large extent.This AMR analysis is generated by a small set of rules that convert a native Logical Form analysis provided by a preexisting parser (see Vanderwende, 2015) into the AMR format.While we demonstrate the performance of our AMR parser on data sets annotated by the LDC, we will focus attention in the demo on the following two areas: 1) we will make available AMR annotations for the data sets that were used to develop our parser, to serve as a supplement to the LDC data sets, and 2) we will demonstrate AMR parsers for German, French, Spanish and Japanese that make use of the same small set of LF-to-AMR conversion rules. Lucy Vanderwende, Arul Menezes, Chris Quirk |
HLT-NAACL | 3 |
| 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual DataabstractStatistical phrase-based translation learns translation rules from bilingual corpora, and has traditionally only used monolingual evidence to construct features that rescore existing translation candidates.In this work, we present a semi-supervised graph-based approach for generating new translation rules that leverages bilingual and monolingual data.The proposed technique first constructs phrase graphs using both source and target language monolingual corpora.Next, graph propagation identifies translations of phrases that were not observed in the bilingual corpus, assuming that similar phrases have similar translations.We report results on a large Arabic-English system and a medium-sized Urdu-English system.Our proposed approach significantly improves the performance of competitive phrasebased systems, leading to consistent improvements between 1 and 4 BLEU points on standard evaluation sets.Source!Target! el gato! los gatos!un gato! cat! the cat! the cats! a cat! Target!Prob.! the cat! 0.7! cat! 0.15! …! …! felino!canino!el perro!Target!Prob.! canine!0.6!dog! 0.3!…! …! Target!Prob.! the cats!0.8! cats! 0.1!…! … Avneesh Saluja, Hany Hassan, Kristina Toutanova, Chris Quirk |
ACL (1) | 4 |
| 2014 | Distributed open-domain conversational understanding framework with domain independent extractorsabstractTraditional spoken dialog systems are usually based on a centralized architecture, in which the number of domains is predefined, and the provider is fixed for a given domain and intent. The spoken language understanding (SLU) component is responsible for detecting domain and intents, and filling domain-specific slots. It is expensive and time-consuming in this architecture to add new and/or competing domains, intents, or providers. The rapid growth of service providers in the mobile computing market calls for an extensible dialog system framework. This paper presents a distributed dialog infrastructure where each domain or provider is agnostic of others, and processes the user utterances independently using their own knowledge or models, so that a new domain and new provider can be easily incorporated in. In addition, to facilitate each service provider building their own SLU models or algorithms, we introduce a new component, extractors, to provide intermediate semantic annotations such as entity mention tags, which can be plugged in arbitrarily as well. Each service provider can then rapidly develop their SLU parser with minimum efforts by providing some example sentences with intents and slots if needed. Our preliminary experimental results demonstrate the power of this new framework compared to a centralized architecture. Qi Li 0014, Gökhan Tür, Dilek Hakkani-Tür, Xiang Li 0066, Tim Paek, Asela Gunawardana, Chris Quirk |
SLT | 7 |
| 2014 | Literome: PubMed-scale genomic knowledge base in the cloudabstractMOTIVATION: Advances in sequencing technology have led to an exponential growth of genomics data, yet it remains a formidable challenge to interpret such data for identifying disease genes and drug targets. There has been increasing interest in adopting a systems approach that incorporates prior knowledge such as gene networks and genotype-phenotype associations. The majority of such knowledge resides in text such as journal publications, which has been undergoing its own exponential growth. It has thus become a significant bottleneck to identify relevant knowledge for genomic interpretation as well as to keep up with new genomics findings. RESULTS: In the Literome project, we have developed an automatic curation system to extract genomic knowledge from PubMed articles and made this knowledge available in the cloud with a Web site to facilitate browsing, searching and reasoning. Currently, Literome focuses on two types of knowledge most pertinent to genomic medicine: directed genic interactions such as pathways and genotype-phenotype associations. Users can search for interacting genes and the nature of the interactions, as well as diseases and drugs associated with a single nucleotide polymorphism or gene. Users can also search for indirect connections between two entities, e.g. a gene and a disease might be linked because an interacting gene is associated with a related disease. AVAILABILITY AND IMPLEMENTATION: Literome is freely available at literome.azurewebsites.net. Download for non-commercial use is available via Web services. Hoifung Poon, Chris Quirk, Charlie DeZiel, David Heckerman |
Bioinform. | 2 |
| 2013 | Lightly Supervised Learning of Procedural Dialog Systems
Svitlana Volkova, Pallavi Choudhury, Chris Quirk, William B. Dolan, Luke Zettlemoyer |
ACL (1) | 3 |
| 2013 | Joint Language and Translation Modeling with Recurrent Neural NetworksabstractWe present a joint language and translation model based on a recurrent neural network which predicts target words based on an unbounded history of both source and target words.The weaker independence assumptions of this model result in a vastly larger search space compared to related feedforward-based language or translation models.We tackle this issue with a new lattice rescoring algorithm and demonstrate its effectiveness empirically.Our joint model builds on a well known recurrent neural network language model (Mikolov, 2012) augmented by a layer of additional inputs from the source language.We show competitive accuracy compared to the traditional channel model features.Our best results improve the output of a system trained on WMT 2012 French-English data by up to 1.5 BLEU, and by 1.1 BLEU on average across several test sets. Michael Auli, Michel Galley, Chris Quirk, Geoffrey Zweig |
EMNLP | 3 |
| 2013 | Regularized Minimum Error Rate TrainingabstractMinimum Error Rate Training (MERT) remains one of the preferred methods for tuning linear parameters in machine translation systems, yet it faces significant issues.First, MERT is an unregularized learner and is therefore prone to overfitting.Second, it is commonly used on a noisy, non-convex loss function that becomes more difficult to optimize as the number of parameters increases.To address these issues, we study the addition of a regularization term to the MERT objective function.Since standard regularizers such as ℓ 2 are inapplicable to MERT due to the scale invariance of its objective function, we turn to two regularizers-ℓ 0 and a modification of ℓ 2and present methods for efficiently integrating them during search.To improve search in large parameter spaces, we also present a new direction finding algorithm that uses the gradient of expected BLEU to orient MERT's exact line searches.Experiments with up to 3600 features show that these extensions of MERT yield results comparable to PRO, a learner often used with large feature sets. Michel Galley, Chris Quirk, Colin Cherry, Kristina Toutanova |
EMNLP | 2 |
| 2013 | Monolingual Marginal Matching for Translation Model AdaptationabstractWhen using a machine translation (MT) model trained on OLD-domain parallel data to translate NEW-domain text, one major challenge is the large number of out-of-vocabulary (OOV) and new-translation-sense words.We present a method to identify new translations of both known and unknown source language words that uses NEW-domain comparable document pairs.Starting with a joint distribution of source-target word pairs derived from the OLD-domain parallel corpus, our method recovers a new joint distribution that matches the marginal distributions of the NEW-domain comparable document pairs, while minimizing the divergence from the OLD-domain distribution.Adding learned translations to our French-English MT model results in gains of about 2 BLEU points over strong baselines. Ann Irvine, Chris Quirk, Hal Daumé III |
EMNLP | 2 |
| 2013 | Paraphrase features to improve natural language understanding
Ruhi Sarikaya, Chris Brockett, Chris Quirk, William B. Dolan |
INTERSPEECH | 4 |
| 2013 | Morphological, Syntactical and Semantic Knowledge in Statistical Machine Translation
Marta R. Costa-jussà, Chris Quirk |
HLT-NAACL | 2 |
| 2013 | Beyond Left-to-Right: Multiple Decomposition Structures for SMT
Kristina Toutanova, Chris Quirk, Jianfeng Gao 0001 |
HLT-NAACL | 3 |
| 2012 | MSR SPLAT, a language analysis toolkit
Chris Quirk, Pallavi Choudhury, Jianfeng Gao 0001, Hisami Suzuki, Kristina Toutanova, Michael Gamon, Scott Yih, Colin Cherry, Lucy Vanderwende |
HLT-NAACL | 1 |
| 2012 | Linguistic Structure Prediction Noah A. Smith Carnegie Mellon University Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 13), 2011, xx+248 pp; paperbound, ISBN 978-1-60845-405-1, $60.00; ebook, ISBN 978-1-60845-406-8, $30.00 or by subscriptionabstractNoah Smith's ambitious new monograph, Linguistic Structure Prediction, “aims to bridge the gap between natural language processing and machine learning.” Given that current natural language processing (NLP) research makes heavy demands on machine-learning techniques, and a sizeable fraction of modern machine learning (ML) research focuses on structure prediction, this is clearly a timely and important topic. To address the gaps and overlaps between these two large and well-developed fields in five brief chapters is a difficult feat. The text, though not without its flaws, does an admirable job of building this bridge.An introductory first chapter surveys current research areas in statistical NLP, cataloging and defining many common linguistic structure prediction tasks. Machine learning students new to the area are likely to find this helpful albeit a bit terse; NLP students will likely consider this section primarily a review. The subsequent chapters change character abruptly, delving into mathematical details and heavy formalism.Chapter 2 introduces the concept of decoding, presenting five distinct viewpoints on the search for the highest scoring structure. The reader is quickly ushered through graphical models, polytopes, grammars, hypergraphs, and weighted deduction systems, with descriptions based on an example in sequence tagging. The broad coverage, multi-viewpoint discussion encourages the reader to make connections between many distinct approaches, and provides solid formalism for reasoning about decoding problems. It is a comprehensive introduction to the most common and effective decoding approaches, with one significant exception: the recent advances in dual decomposition and Lagrangian relaxation methods. Timing is likely the culprit. This book was developed mainly from 2006 to 2009, whereas dual decomposition did not attain notoriety in our community until a few years later (Rush et al. 2010). Relaxation approaches, though potentially a passing phase, have successfully broadened the reach of simpler decoding techniques into more complicated domains such as structured event extraction. They would have made a nice addition. Regardless, this second chapter equips the reader with sufficient machinery to solve a number of structured prediction problems.Chapter 3 applies the machinery described in the prior chapter to the problem of supervised structure induction. Probabilistic generative and conditional models are introduced in some detail, followed by a discussion of margin-based methods. Hidden Markov models (HMMs) and probabilistic context-free grammars are introduced in detail, followed by solid descriptions of maximum likelihood estimation and smoothing. The section on conditional models is well written and crucial, because so many commonly used tasks can be treated as sequence modeling using techniques such as conditional random fields. Much of the subject matter introduced abstractly in Chapter 2 is presented in this chapter using specific algorithms. For instance, sequence modeling is discussed broadly in Chapter 2; the specific algorithms for HMMs are fully defined in Chapter 3. This coarse-to-fine introduction of material may challenge readers who are accustomed to more practical descriptions of material. Were I to teach a course based on this book, I would be tempted to present the third chapter before the second.Chapter 4 focuses on semisupervised, unsupervised, and hidden variable learning. With a good mix of theory and practical examples, Expectation-Maximization (EM) is introduced and grounded in several problems, then generalized with log-linear models and approximated with contrastive estimation. Hard EM is mentioned in the context of several examples, though a more detailed description of this potentially important technique (cf. Spitkovsky et al., 2010) would bridge the material of Chapters 2 and 4. The chapter then describes Bayesian approaches to NLP, working from theory into specific techniques and landing in models. Finally, a brief section is devoted to the related area of hidden variable learning.Chapter 5 begins by describing the partition function, as well as inference techniques for the partition function and decoding methods. I found it strange that this important section was postponed so late in the book; much of the material was forward-referenced throughout Chapter 4. Regardless, the techniques are described in a unifying, generic manner. The book concludes with a discussion of minimum Bayes risk decoding, and a few other variants.Four appendices are devoted to optimization, experimental techniques, maximum entropy, and locally normalized conditional models. All of these sections provide some useful background. The section on hypothesis testing in Appendix B would be especially useful to students new to the area. It can be difficult to pick the correct hypothesis testing method in general, and this problem is exacerbated in structure prediction. This material serves as a good guide for a researcher hoping to evaluate how effective these methods are.I have some concerns about the intended audience. Descriptions quickly descend into heavy notation and require knowledge of a broad range of mathematical concepts, from marginal polytopes to semirings. I suspect the average NLP graduate student would find it difficult to approach much of the material without a series of courses in probability, statistics, and machine learning. The book is also very theoretical: Few concrete algorithms are provided. Instead, the concepts are introduced using only mathematics and formalism. For readers already conversant in the mapping from mathematical descriptions into concrete algorithms and implementations, this will not be a significant barrier. From the other direction, the structures used in NLP (e.g., dependency trees) are relatively well motivated, though a machine-learning researcher new to the area might benefit from a fuller introduction to NLP. However, the text serves as an effective guide for introducing the machine-learning community into the NLP community, but I feel it would be challenging to use in the other direction.This may be a personal bias, but I was surprised by the avoidance of machine-translation–related techniques, despite obvious influences. Why resort to the term “decoding” if not because of decipherment and translation? One of the most effective uses of hidden variable learning is in word alignment; it seems like a personal example. Of course, building an effective machine-translation system requires a huge amount of engineering in addition to the underlying theory, but I felt some discussion of the problem and effective techniques would be pertinent.Despite my struggles with the book's organization and a few important omissions, I must admit that I want a copy for my bookshelf. The author covers a huge amount of material in just under 200 pages, touching on some of the most important algorithmic techniques and viewpoints in modern statistical NLP. At times the text reads like a survey, touching very briefly on a huge range of topics. Yet the survey is comprehensive and enlightening, tying together a broad range of topics and viewpoints. Younger graduate students may require serious effort to comprehend the full text, but the modern NLP researcher looking to advance the state of the art in structured prediction must truly understand the concepts presented here. Chris Quirk |
Comput. Linguistics | 1 |
| 2011 | Gappy Phrasal Alignment By Agreement
Mohit Bansal, Chris Quirk, Robert C. Moore |
ACL | 2 |
| 2011 | Optimal Search for Minimum Error Rate Training
Michel Galley, Chris Quirk |
EMNLP | 2 |
| 2011 | Incremental Training and Intentional Over-fitting of Word Alignment
Qin Gao, William Lewis, Chris Quirk, Mei-Yuh Hwang |
MTSummit | 3 |
| 2011 | MT Detection in Web-Scraped Parallel Corpora
Spencer Rarrick, Chris Quirk, William Lewis |
MTSummit | 2 |
| 2011 | On the Expressivity of Linear Transductions
Markus Saers, Dekai Wu, Chris Quirk |
MTSummit | 3 |
| 2010 | Learning Phrase-Based Spelling Error Models from Clickthrough Data
Xu Sun 0001, Jianfeng Gao 0001, Daniel Micol, Chris Quirk |
ACL | 4 |
| 2010 | A Large Scale Ranker-Based System for Search Query Spelling Correction
Jianfeng Gao 0001, Daniel Micol, Chris Quirk, Xu Sun 0001 |
COLING | 4 |
| 2010 | Extracting Parallel Sentences from Comparable Corpora using Document Level Alignment
Jason Smith 0006, Chris Quirk, Kristina Toutanova |
HLT-NAACL | 2 |
| 2009 | Improving search engines using human computation gamesabstractWork on evaluating and improving the relevance of web search engines typically use human relevance judgments or clickthrough data. Both these methods look at the problem of learning the mapping from queries to web pages. In this paper, we identify some issues with this approach, and suggest an alternative approach, namely, learning a mapping from web pages to queries. In particular, we use human computation games to elicit data about web pages from players that can be used to improve search. We describe three human computation games that we developed, with a focus on Page Hunt, a single-player game. We describe experiments we conducted with several hundred game players, highlight some interesting aspects of the data obtained and define the 'findability' metric. We also show how we automatically extract query alterations for use in query refinement using techniques from bitext matching. The data that we elicit from players has several other applications including providing metadata for pages and identifying ranking issues. Raman Chandrasekar, Chris Quirk |
CIKM | 3 |
| 2009 | Less is More: Significance-Based N-gram Selection for Smaller, Better Language Models
Robert C. Moore, Chris Quirk |
EMNLP | 2 |
| 2009 | Page hunt: improving search engines using human computation gamesabstractThere has been a lot of work on evaluating and improving the relevance of web search engines. In this paper, we suggest using human computation games to elicit data from players that can be used to improve search. We describe Page Hunt, a single-player game. The data elicited using Page Hunt has several applications including providing metadata for pages, providing query alterations for use in query refinement, and identifying ranking issues. We describe an experiment with over 340 game players, and highlight some interesting aspects of the data obtained. Raman Chandrasekar, Chris Quirk |
SIGIR | 3 |
| 2008 | Bayesian Learning of Non-Compositional Phrases with Synchronous Parsing
Hao Zhang 0010, Chris Quirk, Robert C. Moore, Daniel Gildea |
ACL | 2 |
| 2008 | Random Restarts in Minimum Error Rate Training for Statistical Machine Translation
Robert C. Moore, Chris Quirk |
COLING | 2 |
| 2008 | Syntactic Models for Structural Word Insertion and Deletion during Translation
Arul Menezes, Chris Quirk |
EMNLP | 2 |
| 2007 | Faster beam-search decoding for phrasal statistical machine translation
Robert C. Moore, Chris Quirk |
MTSummit | 2 |
| 2007 | Generative models of noisy translations with applications to parallel fragment extraction
Chris Quirk, Raghavendra Udupa, Arul Menezes |
MTSummit | 1 |
| 2006 | The impact of parse quality on syntactically-informed statistical machine translation
Chris Quirk, Simon Corston-Oliver |
EMNLP | 1 |
| 2006 | Do we need phrases? Challenging the conventional wisdom in Statistical Machine Translation
Chris Quirk, Arul Menezes |
HLT-NAACL | 1 |
| 2006 | Dependency treelet translation: the convergence of statistical and example-based machine-translation?
Chris Quirk, Arul Menezes |
Mach. Transl. | 1 |
| 2005 | Dependency Treelet Translation: Syntactically Informed Phrasal SMTabstractWe describe a novel approach to statistical machine translation that combines syntactic information in the source language with recent advances in phrasal translation. This method requires a source-language dependency parser, target language word segmentation and an unsupervised word alignment component. We align a parallel corpus, project the source dependency parse onto the target sentence, extract dependency treelet translation pairs, and train a tree-based ordering model. We describe an efficient decoder and show that using these tree-based models in combination with conventional SMT models provides a promising approach that incorporates the power of phrasal SMT with the linguistic generality available in a parser. Chris Quirk, Arul Menezes, Colin Cherry |
ACL | 1 |
| 2004 | Unsupervised Construction of Large Paraphrase Corpora: Exploiting Massively Parallel News Sources
William B. Dolan, Chris Quirk, Chris Brockett |
COLING | 2 |
| 2004 | Monolingual Machine Translation for Paraphrase Generation
Chris Quirk, Chris Brockett, William B. Dolan |
EMNLP | 1 |
| 2004 | Training a Sentence-Level Machine Translation Confidence Measure
Chris Quirk |
LREC | 1 |
| 2003 | Disambiguation of English PP attachment using multilingual aligned dataabstractPrepositional phrase attachment (PP attachment) is a major source of ambiguity in English. It poses a substantial challenge to Machine Translation (MT) between English and languages that are not characterized by PP attachment ambiguity. In this paper we present an unsupervised, bilingual, corpus-based approach to the resolution of English PP attachment ambiguity. As data we use aligned linguistic representations of the English and Japanese sentences from a large parallel corpus of technical texts. The premise of our approach is that with large aligned, parsed, bilingual (or multilingual) corpora, languages can learn non-trivial linguistic information from one another with high accuracy. We contend that our approach can be extended to linguistic phenomena other than PP attachment. Lee Schwartz, Takako Aikawa, Chris Quirk |
MTSummit | 3 |