EDBT 2026 Demo / reviewers in the wild / expert
Kristina Toutanova
dblp:25/1520
· DBLP profile ↗
65ranked-venue papers
19as first author
14since 2021 · last 2025
0009-0005-3458-9049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 18 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert ModelsabstractWe propose a method to optimize language model pre-training data mixtures through efficient approximation of the cross-entropy loss corresponding to each candidate mixture via a Mixture of Data Experts (MDE).We use this approximation as additional features in a regression model, trained from observations of model loss for a small number of mixtures.Experiments with Transformer language models between 70M and 10B parameters on the SlimPajama dataset show that our method achieves significantly better performance than approaches that train regression models using only the mixture rates as input features.Combining our method with an objective that takes into account cross-entropy on end task data leads to superior performance on few-shot downstream evaluations.We also provide theoretical insights on why aggregation of data expert predictions can provide good approximations to model losses for data mixtures. Lior Belenki, Alekh Agarwal, Tianze Shi, Kristina Toutanova |
ACL (1) | 4 |
| 2025 | Understanding Museum Exhibits using Vision-Language Reasoning
Ada-Astrid Balauca, Sanjana Garai, Stefan Balauca, Rasesh Udayakumar Shetty, Naitik Agrawal, Dhwanil Subhashbhai Shah, Yuqian Fu, Xi Wang 0021, Kristina Toutanova, Danda Pani Paudel, Luc Van Gool |
ICCV | 9 |
| 2024 | Taming CLIP for Fine-Grained and Structured Visual Understanding of Museum Exhibits
Ada-Astrid Balauca, Danda Pani Paudel, Kristina Toutanova, Luc Van Gool |
ECCV (76) | 3 |
| 2024 | Efficient End-to-End Visual Document Understanding with Rationale DistillationabstractWang Zhu, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wang Zhu 0001, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova |
NAACL-HLT | 6 |
| 2023 | QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set OperationsabstractFormulating selective information needs results in queries that implicitly specify set operations, such as intersection, union, and difference.For instance, one might search for "shorebirds that are not sandpipers" or "science-fiction films shot in England".To study the ability of retrieval systems to meet such information needs, we construct QUEST, a dataset of 3357 natural language queries with implicit set operations, that map to a set of entities corresponding to Wikipedia documents.The dataset challenges models to match multiple constraints mentioned in queries with corresponding evidence in documents and correctly perform various set operations.The dataset is constructed semi-automatically using Wikipedia category names.Queries are automatically composed from individual categories, then paraphrased and further validated for naturalness and fluency by crowdworkers.Crowdworkers also assess the relevance of entities based on their documents and highlight attribution of query constraints to spans of document text.We analyze several modern retrieval systems, finding that they often struggle on such queries.Queries involving negation and conjunction are particularly challenging and systems are further challenged with combinations of these operations. 1 * Work done during an internship at Google.retrieving an exhaustive document set, instead lim-65 iting annotation to the top few results of a baseline 66 information retrieval system.67 To analyze how well retrieval systems handle 68 such queries, we present QUEST, a dataset with 69 natural language queries from four domains, that 70 are mapped to relatively comprehensive sets of en-71 tities corresponding to Wikipedia pages.We use 72 Wikipedia categories and their mapping to entities 73 in Wikipedia as a building block for our dataset 74 construction approach, but do not allow access to 75 this semi-structured data source at inference time, 76 to simulate text-based retrieval.Wikipedia cate-77 gories represent a broad set of natural language 78 descriptions of entity properties and often corre-79 spond to selective information need queries that 80 could be plausibly issued by a search engine user 81 ([At least 90% of the time based on our filtering?]).82 The correspondence between property names and 83 document text is also often subtle and requires so-84 phisticated reasoning to determine relevance, rep-85 resenting the natural language inference challenge 86 inherent in the task, while the knowledge of cate-87 gory membership allows us to construct relatively 88 comprehensive sets of candidate entities for atomic 89 categories and their combinations.90 Our dataset construction process is outlined in 91 Figure 1.The base queries in our dataset are 92 semi-automatically generated using Wikipedia cat-93 egory names.To construct queries, we sample 94 category names and compose them into complex 95 queries by using pre-defined templates (for exam-96 ple, A \ B \ C).Next, we ask crowdworkers to 97 paraphrase these automatically generated queries, 98 while ensuring that the paraphrased queries are 99 fluent and clearly describe what a user could be 00 looking for.These are then validated for natural-01 ness and fluency by a different set of crowdworkers, 02 and filtered according to those criteria.Finally, for 03 a large subset of our dataset, we collect scalar rel-04 evance labels based on the entity documents, and 05 textual attributions mapping query constraints to 06 spans of document text, to aid the development of 07 systems that can make precise inferences based on 08 trusted sources.09 Performing well on this dataset requires sys-10 tems that can match query constraints with cor-11 Chaitanya Malaviya, Peter Shaw 0004, Ming-Wei Chang, Kenton Lee, Kristina Toutanova |
ACL (1) | 5 |
| 2023 | Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesabstractLarge-scale multi-modal pre-training models such as CLIP [30] and PaLI [8] exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (Oven), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct Oven-Wiki‡by repurposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. Oven-Wiki challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-ofthe-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities. Hexiang Hu, Yi Luan, Yang Chen 0065, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, Ming-Wei Chang |
ICCV | 7 |
| 2023 | Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingabstractVisually-situated language is ubiquitous---sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images. Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu 0001, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw 0004, Ming-Wei Chang, Kristina Toutanova |
ICML | 10 |
| 2023 | From Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesabstractMuch of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use — via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks. Peter Shaw 0004, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, Kristina Toutanova |
NeurIPS | 9 |
| 2022 | Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingabstractLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Linlu Qiu, Peter Shaw 0004, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova |
EMNLP | 8 |
| 2022 | Improving Compositional Generalization with Latent Structure and Data AugmentationabstractLinlu Qiu, Peter Shaw, Panupong Pasupat, Pawel Nowak, Tal Linzen, Fei Sha, Kristina Toutanova. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Linlu Qiu, Peter Shaw 0004, Panupong Pasupat, Pawel Krzysztof Nowak, Tal Linzen, Fei Sha, Kristina Toutanova |
NAACL-HLT | 7 |
| 2021 | Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?abstractPeter Shaw, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Peter Shaw 0004, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova |
ACL/IJCNLP (1) | 4 |
| 2021 | Representations for Question Answering from Documents with Tables and TextabstractTables in Web documents are pervasive and can be directly used to answer many of the queries searched on the Web, motivating their integration in question answering. Very often information presented in tables is succinct and hard to interpret with standard language representations. On the other hand, tables often appear within textual context, such as an article describing the table. Using the information from an article as additional context can potentially enrich table representations. In this work we aim to improve question answering from tables by refining table representations based on information from surrounding text. We also present an effective method to combine text and table-based predictions for question answering from full documents, obtaining significant improvements on the Natural Questions dataset. Victoria Zayats, Kristina Toutanova, Mari Ostendorf |
EACL | 2 |
| 2021 | Joint Passage Ranking for Diverse Multi-Answer RetrievalabstractWe study multi-answer retrieval, an underexplored problem that requires retrieving passages to cover multiple distinct answers for a given question.This task requires joint modeling of retrieved passages, as models should not repeatedly retrieve passages containing the same answer at the cost of missing a different valid answer.In this paper, we introduce JPR, the first joint passage retrieval model for multi-answer retrieval.JPR makes use of an autoregressive reranker that selects a sequence of passages, each conditioned on previously selected passages.JPR is trained to select passages that cover new answers at each timestep and uses a tree-decoding algorithm to enable flexibility in the degree of diversity.Compared to prior approaches, JPR achieves significantly better answer coverage on three multianswer datasets.When combined with downstream question answering, the improved retrieval enables larger answer generation models since they need to consider fewer passages, establishing a new state-of-the-art. Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, Hannaneh Hajishirzi |
EMNLP (1) | 4 |
| 2021 | Sparse, Dense, and Attentional Representations for Text RetrievalabstractAbstract Dual encoders perform retrieval by encoding documents and queries into dense low-dimensional vectors, scoring each document by its inner product with the query. We investigate the capacity of this architecture relative to sparse bag-of-words models and attentional neural networks. Using both theoretical and empirical analysis, we establish connections between the encoding dimension, the margin between gold and lower-ranked documents, and the document length, suggesting limitations in the capacity of fixed-length encodings to support precise retrieval of long documents. Building on these insights, we propose a simple neural model that combines the efficiency of dual encoders with some of the expressiveness of more costly attentional architectures, and explore sparse-dense hybrids to capitalize on the precision of sparse retrieval. These models outperform strong alternatives in large-scale retrieval. Yi Luan, Jacob Eisenstein, Kristina Toutanova, Michael Collins 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Probabilistic Assumptions Matter: Improved Models for Distantly-Supervised Document-Level Question AnsweringabstractWe address the problem of extractive question answering using document-level distant supervision, pairing questions and relevant documents with answer strings.We compare previously used probability space and distant supervision assumptions (assumptions on the correspondence between the weak answer string labels and possible answer mention spans).We show that these assumptions interact, and that different configurations provide complementary benefits.We demonstrate that a multiobjective model can efficiently combine the advantages of multiple assumptions and outperform the best individual formulation.Our approach outperforms previous state-of-the-art models by 4.3 points in F1 on TriviaQA-Wiki and 1.7 points in Rouge-L on NarrativeQA summaries.1 Hao Cheng 0002, Ming-Wei Chang, Kenton Lee, Kristina Toutanova |
ACL | 4 |
| 2019 | Latent Retrieval for Weakly Supervised Open Domain Question AnsweringabstractRecent work on open domain question answering (QA) assumes strong supervision of the supporting evidence and/or assumes a blackbox information retrieval (IR) system to retrieve evidence candidates.We argue that both are suboptimal, since gold evidence is not always available, and QA is fundamentally different from IR.We show for the first time that it is possible to jointly learn the retriever and reader from question-answer string pairs and without any IR system.In this setting, evidence retrieval from all of Wikipedia is treated as a latent variable.Since this is impractical to learn from scratch, we pre-train the retriever with an Inverse Cloze Task.We evaluate on open versions of five QA datasets.On datasets where the questioner already knows the answer, a traditional IR system such as BM25 is sufficient.On datasets where a user is genuinely seeking an answer, we show that learned retrieval is crucial, outperforming BM25 by up to 19 points in exact match. Kenton Lee, Ming-Wei Chang, Kristina Toutanova |
ACL (1) | 3 |
| 2019 | Zero-Shot Entity Linking by Reading Entity DescriptionsabstractWe present the zero-shot entity linking task, where mentions must be linked to unseen entities without in-domain labeled data.The goal is to enable robust transfer to highly specialized domains, and so no metadata or alias tables are assumed.In this setting, entities are only identified by text descriptions, and models must rely strictly on language understanding to resolve the new entities.First, we show that strong reading comprehension models pre-trained on large unlabeled data can be used to generalize to unseen entities.Second, we propose a simple and effective adaptive pre-training strategy, which we term domainadaptive pre-training (DAP), to address the domain shift problem associated with linking unseen entities in a new domain.We present experiments on a new dataset that we construct for this task and show that DAP improves over strong pre-training baselines, including BERT. Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, Honglak Lee |
ACL (1) | 4 |
| 2019 | Natural Questions: a Benchmark for Question Answering ResearchabstractWe present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia page from the top 5 search results, and annotates a long answer (typically a paragraph) and a short answer (one or more entities) if present on the page, or marks null if no long/short answer is present. The public release consists of 307,373 training examples with single annotations; 7,830 examples with 5-way annotations for development data; and a further 7,842 examples with 5-way annotated sequestered as test data. We present experiments validating quality of the data. We also describe analysis of 25-way annotations on 302 examples, giving insights into human variability on the annotation task. We introduce robust metrics for the purposes of evaluating question answering systems; demonstrate high human upper bounds on these metrics; and establish baseline results using competitive methods drawn from related literature. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins 0001, Ankur P. Parikh, Christopher Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V. Le, Slav Petrov |
Trans. Assoc. Comput. Linguistics | 11 |
| 2017 | A Nested Attention Neural Hybrid Model for Grammatical Error CorrectionabstractJianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Jianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao 0001 |
ACL (1) | 3 |
| 2017 | Cross-Sentence N-ary Relation Extraction with Graph LSTMsabstractPast work in relation extraction has focused on binary relations in single sentences. Recent NLP inroads in high-value domains have sparked interest in the more general setting of extracting n-ary relations that span multiple sentences. In this paper, we explore a general relation extraction framework based on graph long short-term memory networks (graph LSTMs) that can be easily extended to cross-sentence n-ary relation extraction. The graph formulation provides a unified way of exploring different LSTM approaches and incorporating various intra-sentential and inter-sentential dependencies, such as sequential, syntactic, and discourse relations. A robust contextual representation is learned for the entities, which serves as input to the relation classifier. This simplifies handling of relations with arbitrary arity, and enables multi-task learning with related relations. We evaluate this framework in two important precision medicine settings, demonstrating its effectiveness with both conventional supervised learning and distant supervision. Cross-sentence extraction produced larger knowledge bases. and multi-task learning significantly improved extraction accuracy. A thorough analysis of various LSTM approaches yielded useful insight the impact of linguistic analysis on extraction accuracy. Nanyun Peng 0001, Hoifung Poon, Chris Quirk, Kristina Toutanova, Scott Yih |
Trans. Assoc. Comput. Linguistics | 4 |
| 2016 | Microsummarization of Online Reviews: An Experimental StudyabstractMobile and location-based social media applications provide platforms for users to share brief opinions about products, venues, and services. These quickly typed opinions, or microreviews, are a valuable source of current sentiment on a wide variety of subjects. However, there is currently little research on how to mine this information to present it back to users in easily consumable way. In this paper, we introduce the task of microsummarization, which combines sentiment analysis, summarization, and entity recognition in order to surface key content to users. We explore unsupervised and supervised methods for this task, and find we can reliably extract relevant entities and the sentiment targeted towards them using crowdsourced labels as supervision. In an end-to-end evaluation, we find our best-performing system is vastly preferred by judges over a traditional extractive summarization approach. This work motivates an entirely new approach to summarization, incorporating both sentiment analysis and item extraction for modernized, at-a-glance presentation of public opinion. Rebecca Mason, Benjamin Gaska, Benjamin Van Durme, Pallavi Choudhury, Ted Hart, William B. Dolan, Kristina Toutanova, Margaret Mitchell |
AAAI | 7 |
| 2016 | Compositional Learning of Embeddings for Relation Paths in Knowledge Base and TextabstractModeling relation paths has offered significant gains in embedding models for knowledge base (KB) completion.However, enumerating paths between two entities is very expensive, and existing approaches typically resort to approximation with a sampled subset.This problem is particularly acute when text is jointly modeled with KB relations and used to provide direct evidence for facts mentioned in it.In this paper, we propose the first exact dynamic programming algorithm which enables efficient incorporation of all relation paths of bounded length, while modeling both relation types and intermediate nodes in the compositional path representations.We conduct a theoretical analysis of the efficiency gain from the approach.Experiments on two datasets show that it addresses representational limitations in prior approaches and improves accuracy in KB completion. Kristina Toutanova, Xi Victoria Lin, Scott Yih, Hoifung Poon, Chris Quirk |
ACL (1) | 1 |
| 2016 | A Dataset and Evaluation Metrics for Abstractive Compression of Sentences and Short ParagraphsabstractWe introduce a manually-created, multireference dataset for abstractive sentence and short paragraph compression.First, we examine the impact of single-and multi-sentence level editing operations on human compression quality as found in this corpus.We observe that substitution and rephrasing operations are more meaning preserving than other operations, and that compressing in context improves quality.Second, we systematically explore the correlations between automatic evaluation metrics and human judgments of meaning preservation and grammaticality in the compression task, and analyze the impact of the linguistic units used and precision versus recall measures on the quality of the metrics.Multi-reference evaluation metrics are shown to offer significant advantage over single reference-based metrics. Kristina Toutanova, Chris Brockett, Ke M. Tran, Saleema Amershi |
EMNLP | 1 |
| 2016 | E-TIPSY: Search Query Corpus Annotated with Entities, Term Importance, POS Tags, and Syntactic Parses
Yuval Marton, Kristina Toutanova |
LREC | 2 |
| 2015 | Model Selection for Type-Supervised Learning with Application to POS TaggingabstractModel selection (picking, for example, the feature set and the regularization strength) is crucial for building high-accuracy NLP models.In supervised learning, we can estimate the accuracy of a model on a subset of the labeled data and choose the model with the highest accuracy.In contrast, here we focus on type-supervised learning, which uses constraints over the possible labels for word types for supervision, and labeled data is either not available or very small.For the setting where no labeled data is available, we perform a comparative study of previously proposed and one novel model selection criterion on type-supervised POS-tagging in nine languages.For the setting where a small labeled set is available, we show that the set should be used for semi-supervised learning rather than for model selection onlyusing it for model selection reduces the error by less than 5%, whereas using it for semi-supervised learning reduces the error by 44%. Kristina Toutanova, Waleed Ammar, Pallavi Choudhury, Hoifung Poon |
CoNLL | 1 |
| 2015 | Representing Text for Joint Embedding of Text and Knowledge BasesabstractModels that learn to represent textual and knowledge base relations in the same continuous latent space are able to perform joint inferences among the two kinds of relations and obtain high accuracy on knowledge base completion (Riedel et al., 2013).In this paper we propose a model that captures the compositional structure of textual relations, and jointly optimizes entity, knowledge base, and textual relation representations.The proposed model significantly improves performance over a model that does not share parameters among textual relations with common sub-structure. Kristina Toutanova, Danqi Chen 0001, Patrick Pantel, Hoifung Poon, Pallavi Choudhury, Michael Gamon |
EMNLP | 1 |
| 2015 | Detecting Translation Direction: A Cross-Domain StudyabstractParallel corpora are constructed by taking a document authored in one language and translating it into another language.However, the information about the authored and translated sides of the corpus is usually not preserved.When available, this information can be used to improve statistical machine translation.Existing statistical methods for translation direction detection have low accuracy when applied to the realistic out-of-domain setting, especially when the input texts are short.Our contributions in this work are threefold: 1) We develop a multi-corpus parallel dataset with translation direction labels at the sentence level, 2) we perform a comparative evaluation of previously introduced features for translation direction detection in a cross-domain setting and 3) we generalize a previously introduced type of features to outperform the best previously proposed features in detecting translation direction and achieve 0.80 precision with 0.85 recall. Sauleh Eetemadi, Kristina Toutanova |
HLT-NAACL | 2 |
| 2015 | Grounded Semantic Parsing for Complex Knowledge ExtractionabstractRecently, there has been increasing interest in learning semantic parsers with indirect supervision, but existing work focuses almost exclusively on question answering.Separately, there have been active pursuits in leveraging databases for distant supervision in information extraction, yet such methods are often limited to binary relations and none can handle nested events.In this paper, we generalize distant supervision to complex knowledge extraction, by proposing the first approach to learn a semantic parser for extracting nested event structures without annotated examples, using only a database of such complex events and unannotated text.The key idea is to model the annotations as latent variables, and incorporate a prior that favors semantic parses containing known events.Experiments on the GENIA event extraction dataset show that our approach can learn from and extract complex biological pathway events.Moreover, when supplied with just five example words per event type, it becomes competitive even among supervised systems, outperforming 19 out of 24 teams that participated in the original shared task. Ankur P. Parikh, Hoifung Poon, Kristina Toutanova |
HLT-NAACL | 3 |
| 2015 | Survey of data-selection methods in statistical machine translation
Sauleh Eetemadi, William Lewis, Kristina Toutanova, Hayder Radha |
Mach. Transl. | 3 |
| 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual DataabstractStatistical phrase-based translation learns translation rules from bilingual corpora, and has traditionally only used monolingual evidence to construct features that rescore existing translation candidates.In this work, we present a semi-supervised graph-based approach for generating new translation rules that leverages bilingual and monolingual data.The proposed technique first constructs phrase graphs using both source and target language monolingual corpora.Next, graph propagation identifies translations of phrases that were not observed in the bilingual corpus, assuming that similar phrases have similar translations.We report results on a large Arabic-English system and a medium-sized Urdu-English system.Our proposed approach significantly improves the performance of competitive phrasebased systems, leading to consistent improvements between 1 and 4 BLEU points on standard evaluation sets.Source!Target! el gato! los gatos!un gato! cat! the cat! the cats! a cat! Target!Prob.! the cat! 0.7! cat! 0.15! …! …! felino!canino!el perro!Target!Prob.! canine!0.6!dog! 0.3!…! …! Target!Prob.! the cats!0.8! cats! 0.1!…! … Avneesh Saluja, Hany Hassan, Kristina Toutanova, Chris Quirk |
ACL (1) | 3 |
| 2014 | Asymmetric Features Of Human Generated TranslationabstractDistinct properties of translated text have been the subject of research in linguistics for many year (Baker, 1993).In recent years computational methods have been developed to empirically verify the linguistic theories about translated text (Baroni and Bernardini, 2006).While many characteristics of translated text are more apparent in comparison to the original text, most of the prior research has focused on monolingual features of translated and original text.The contribution of this work is introducing bilingual features that are capable of explaining differences in translation direction using localized linguistic phenomena at the phrase or sentence level, rather than using monolingual statistics at the document level.We show that these bilingual features outperform the monolingual features used in prior work (Kurokawa et al., 2009) for the task of classifying translation direction. Sauleh Eetemadi, Kristina Toutanova |
EMNLP | 2 |
| 2013 | Regularized Minimum Error Rate TrainingabstractMinimum Error Rate Training (MERT) remains one of the preferred methods for tuning linear parameters in machine translation systems, yet it faces significant issues.First, MERT is an unregularized learner and is therefore prone to overfitting.Second, it is commonly used on a noisy, non-convex loss function that becomes more difficult to optimize as the number of parameters increases.To address these issues, we study the addition of a regularization term to the MERT objective function.Since standard regularizers such as ℓ 2 are inapplicable to MERT due to the scale invariance of its objective function, we turn to two regularizers-ℓ 0 and a modification of ℓ 2and present methods for efficiently integrating them during search.To improve search in large parameter spaces, we also present a new direction finding algorithm that uses the gradient of expected BLEU to orient MERT's exact line searches.Experiments with up to 3600 features show that these extensions of MERT yield results comparable to PRO, a learner often used with large feature sets. Michel Galley, Chris Quirk, Colin Cherry, Kristina Toutanova |
EMNLP | 4 |
| 2013 | Beyond Left-to-Right: Multiple Decomposition Structures for SMT
Kristina Toutanova, Chris Quirk, Jianfeng Gao 0001 |
HLT-NAACL | 2 |
| 2012 | Multilingual Named Entity Recognition using Parallel Data and Metadata from Wikipedia
Sungchul Kim, Kristina Toutanova, Hwanjo Yu |
ACL (1) | 2 |
| 2012 | MSR SPLAT, a language analysis toolkit
Chris Quirk, Pallavi Choudhury, Jianfeng Gao 0001, Hisami Suzuki, Kristina Toutanova, Michael Gamon, Scott Yih, Colin Cherry, Lucy Vanderwende |
HLT-NAACL | 5 |
| 2011 | Unsupervised Bilingual Morpheme Segmentation and Alignment with Context-rich Hidden Semi-Markov Models
Jason Naradowsky, Kristina Toutanova |
ACL | 2 |
| 2011 | Learning Discriminative Projections for Text Similarity Measures
Scott Yih, Kristina Toutanova, John C. Platt, Christopher Meek |
CoNLL | 2 |
| 2011 | Clickthrough-based latent semantic models for web searchabstractThis paper presents two new document ranking models for Web search based upon the methods of semantic representation and the statistical translation-based approach to information retrieval (IR). Assuming that a query is parallel to the titles of the documents clicked on for that query, large amounts of query-title pairs are constructed from clickthrough data; two latent semantic models are learned from this data. One is a bilingual topic model within the language modeling framework. It ranks documents for a query by the likelihood of the query being a semantics-based translation of the documents. The semantic representation is language independent and learned from query-title pairs, with the assumption that a query and its paired titles share the same distribution over semantic topics. The other is a discriminative projection model within the vector space modeling framework. Unlike Latent Semantic Analysis and its variants, the projection matrix in our model, which is used to map from term vectors into sematic space, is learned discriminatively such that the distance between a query and its paired title, both represented as vectors in the projected semantic space, is smaller than that between the query and the titles of other documents which have no clicks for that query. These models are evaluated on the Web search task using a real world data set. Results show that they significantly outperform their corresponding baseline models, which are state-of-the-art. Jianfeng Gao 0001, Kristina Toutanova, Scott Yih |
SIGIR | 2 |
| 2010 | Translingual Document Representations from Discriminative Projections
John C. Platt, Kristina Toutanova, Scott Yih |
EMNLP | 2 |
| 2010 | Extracting Parallel Sentences from Comparable Corpora using Document Level Alignment
Jason Smith 0006, Chris Quirk, Kristina Toutanova |
HLT-NAACL | 3 |
| 2009 | A global model for joint lemmatization and part-of-speech prediction
Kristina Toutanova, Colin Cherry |
ACL/IJCNLP | 1 |
| 2009 | Joint Optimization for Machine Translation System Combination
Xiaodong He 0001, Kristina Toutanova |
EMNLP | 2 |
| 2009 | Unsupervised Morphological Segmentation with Log-Linear Models
Hoifung Poon, Colin Cherry, Kristina Toutanova |
HLT-NAACL | 3 |
| 2008 | Applying Morphology Generation Models to Machine Translation
Kristina Toutanova, Hisami Suzuki, Achim Ruopp |
ACL | 1 |
| 2008 | Bayesian Semi-Supervised Chinese Word Segmentation for Statistical Machine Translation
Jia Xu 0004, Jianfeng Gao 0001, Kristina Toutanova, Hermann Ney |
COLING | 3 |
| 2008 | A Global Joint Model for Semantic Role LabelingabstractWe present a model for semantic role labeling that effectively captures the linguistic intuition that a semantic argument frame is a joint structure, with strong dependencies among the arguments. We show how to incorporate these strong dependencies in a statistical joint model with a rich set of features over multiple argument phrases. The proposed model substantially outperforms a similar state-of-the-art local model that does not include dependencies among different arguments. We evaluate the gains from incorporating this joint information on the Propbank corpus, when using correct syntactic parse trees as input, and when using automatically derived parse trees. The gains amount to 24.1% error reduction on all arguments and 36.8% on core arguments for gold-standard parse trees on Propbank. For automatic parse trees, the error reductions are 8.3% and 10.3% on all and core arguments, respectively. We also present results on the CoNLL 2005 shared task data set. Additionally, we explore considering multiple syntactic analyses to cope with parser noise and uncertainty. Kristina Toutanova, Aria Haghighi, Christopher D. Manning |
Comput. Linguistics | 1 |
| 2007 | A Discriminative Syntactic Word Order Model for Machine Translation
Pi-Chuan Chang, Kristina Toutanova |
ACL | 2 |
| 2007 | A Comparative Study of Parameter Estimation Methods for Statistical Natural Language Processing
Jianfeng Gao 0001, Galen Andrew, Kristina Toutanova |
ACL | 4 |
| 2007 | Generating Complex Morphology for Machine Translation
Einat Minkov, Kristina Toutanova, Hisami Suzuki |
ACL | 2 |
| 2007 | Generating Case Markers in Machine Translation
Kristina Toutanova, Hisami Suzuki |
HLT-NAACL | 1 |
| 2007 | A Bayesian LDA-based model for semi-supervised part-of-speech taggingabstractWe present a novel Bayesian model for semi-supervised part-of-speech tagging. Our model extends the Latent Dirichlet Allocation model and incorporates the intuition that words’ distributions over tags, p(t|w), are sparse. In addition we in- troduce a model for determining the set of possible tags of a word which captures important dependencies in the ambiguity classes of words. Our model outper- forms the best previously proposed model for this task on a standard dataset. Kristina Toutanova |
NIPS | 1 |
| 2006 | Learning to Predict Case Markers in JapaneseabstractJapanese case markers, which indicate the grammatical relation of the complement NP to the predicate, often pose challenges to the generation of Japanese text, be it done by a foreign language learner, or by a machine translation (MT) system. In this paper, we describe the task of predicting Japanese case markers and propose machine learning methods for solving it in two settings: (i) monolingual, when given information only from the Japanese sentence; and (ii) bilingual, when also given information from a corresponding English source sentence in an MT context. We formulate the task after the well-studied task of English semantic role labelling, and explore features from a syntactic dependency structure of the sentence. For the monolingual task, we evaluated our models on the Kyoto Corpus and achieved over 84% accuracy in assigning correct case markers for each phrase. For the bilingual task, we achieved an accuracy of 92% per phrase using a bilingual dataset from a technical domain. We show that in both settings, features that exploit dependency information, whether derived from gold-standard annotations or automatically assigned, contribute significantly to the prediction of case markers. Hisami Suzuki, Kristina Toutanova |
ACL | 2 |
| 2006 | Competitive generative models with structure learning for NLP classification tasks
Kristina Toutanova |
EMNLP | 1 |
| 2006 | Automatic Semantic Role Labeling
Scott Yih, Kristina Toutanova |
HLT-NAACL | 2 |
| 2005 | Joint Learning Improves Semantic Role LabelingabstractDespite much recent progress on accurate semantic role labeling, previous work has largely used independent classifiers, possibly combined with separate label sequence models via Viterbi decoding. This stands in stark contrast to the linguistic observation that a core argument frame is a joint structure, with strong dependencies between arguments. We show how to build a joint model of argument frames, incorporating novel features that model these interactions into discriminative log-linear models. This system achieves an error reduction of 22% on all arguments and 32% on core arguments over a state-of-the art independent classifier for gold-standard parse trees on PropBank. Kristina Toutanova, Aria Haghighi, Christopher D. Manning |
ACL | 1 |
| 2005 | A Joint Model for Semantic Role Labeling
Aria Haghighi, Kristina Toutanova, Christopher D. Manning |
CoNLL | 2 |
| 2004 | The Leaf Path Projection View of Parse Trees: Exploring String Kernels for HPSG Parse Selection
Kristina Toutanova, Penka Markova, Christopher D. Manning |
EMNLP | 1 |
| 2004 | Learning random walk models for inducing word dependency distributionsabstractMany NLP tasks rely on accurately estimating word dependency probabilities P(ω1|ω2), where the words w1 and w2 have a particular relationship (such as verb-object). Because of the sparseness of counts of such dependencies, smoothing and the ability to use multiple sources of knowledge are important challenges. For example, if the probability P(N|V) of noun N being the subject of verb V is high, and V takes similar objects to V', and V' is synonymous to V", then we want to conclude that P(N|V") should also be reasonably high---even when those words did not cooccur in the training data.To capture these higher order relationships, we propose a Markov chain model, whose stationary distribution is used to give word probability estimates. Unlike the manually defined random walks used in some link analysis algorithms, we show how to automatically learn a rich set of parameters for the Markov chain's transition probabilities. We apply this model to the task of prepositional phrase attachment, obtaining an accuracy of 87.54%. Kristina Toutanova, Christopher D. Manning, Andrew Y. Ng |
ICML | 1 |
| 2003 | Optimizing Local Probability Models for Statistical Parsing
Kristina Toutanova, Mark Mitchell, Christopher D. Manning |
ECML | 1 |
| 2003 | Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network
Kristina Toutanova, Daniel Klein 0001, Christopher D. Manning, Yoram Singer |
HLT-NAACL | 1 |
| 2002 | Pronunciation Modeling for Improved Spelling CorrectionabstractThis paper presents a method for incorporating word pronunciation information in a noisy channel model for spelling correction. The proposed method builds an explicit error model for word pronunciations. By modeling pronunciation similarities between words we achieve a substantial performance improvement over the previous best performing models for spelling correction. Kristina Toutanova, Robert C. Moore |
ACL | 1 |
| 2002 | The LinGO Redwoods Treebank: Motivation and Preliminary Applications
Stephan Oepen, Kristina Toutanova, Stuart M. Shieber, Christopher D. Manning, Dan Flickinger, Thorsten Brants |
COLING | 2 |
| 2002 | Feature Selection for a Rich HPSG Grammar Using Decision Trees
Kristina Toutanova, Christopher D. Manning |
CoNLL | 1 |
| 2002 | Extentions to HMM-based Statistical Word Alignment ModelsabstractThis paper describes improved HMM-based word level alignment models for statistical machine translation. We present a method for using part of speech tag information to improve alignment accuracy, and an approach to modeling fertility and correspondence to the empty word in an HMM alignment model. We present accuracy results from evaluating Viterbi alignments against human-judged alignments on the Canadian Hansards corpus, as compared to a bigram HMM, and IBM model 4. The results show up to 16% alignment error reduction. Kristina Toutanova, H. Tolga Ilhan, Christopher D. Manning |
EMNLP | 1 |
| 2001 | Text Classification in a Hierarchical Mixture Model for Small Training SetsabstractDocuments are commonly categorized into hierarchies of topics, such as the ones maintained by Yahoo! and the Open Directory project, in order to facilitate browsing and other interactive forms of information retrieval. In addition, topic hierarchies can be utilized to overcome the sparseness problem in text categorization with a large number of categories, which is the main focus of this paper. This paper presents a hierarchical mixture model which extends the standard naive Bayes classifier and previous hierarchical approaches. Improved estimates of the term distributions are made by differentiation of words in the hierarchy according to their level of generality/specificity. Experiments on the Newsgroups and the Reuters-21578 dataset indicate improved performance of the proposed classifier in comparison to other state-of-the-art methods on datasets with a small number of positive examples. Kristina Toutanova, Francine Chen 0001, Kris Popat, Thomas Hofmann 0001 |
CIKM | 1 |