VLDB 2026 Research / reviewers in the wild / expert
Scott Yih
dblp:07/7129 · also Scott Wen-tau Yih, Wen-Tau Yih, Wen-tau Yih
· DBLP profile ↗
107ranked-venue papers
13as first author
42since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 11 first-author · 41 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Factuality with Explicit Working MemoryabstractMingda Chen, Yang Li, Karthik Padthe, Rulin Shao, Alicia Yi Sun, Luke Zettlemoyer, Gargi Ghosh, Wen-tau Yih. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mingda Chen, Karthik Padthe, Rulin Shao, Alicia Sun, Luke Zettlemoyer, Gargi Ghosh, Scott Yih |
ACL (1) | 8 |
| 2025 | DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversabstractLarge language models (LLMs) have demonstrated strong effectiveness and robustness when fine-tuned as dense retrievers.However, their large parameter size presents significant computational challenges at inference time.While smaller retrievers offer better efficiency, they often fail to generalize effectively with limited supervised fine-tuning data.In this work, we introduce DRAMA, a training framework that leverages LLMs to train smaller generalizable dense retrievers.In particular, we adopt pruned LLMs as the backbone and train on diverse LLM-augmented data in a single-stage contrastive learning setup.Experiments show that DRAMA offers better multilingual and long-context capabilities than traditional encoder-based retrievers, and achieves strong performance across multiple tasks and languages.1 * Equal contribution.† Work done while at Meta. 1 Code and checkpoints will be available at https://github. com/facebookresearch/dpr-scale/tree/main/drama. Xueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin, Scott Yih, Xilun Chen 0002 |
ACL (1) | 5 |
| 2025 | Memory Layers at ScaleabstractMemory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsely activated memory layers complement compute-heavy dense feed-forward layers, providing dedicated capacity to store and retrieve information cheaply. This work takes memory layers beyond proof-of-concept, proving their utility at contemporary scale. On downstream tasks, language models augmented with our improved memory layer outperform dense models with more than twice the computation budget, as well as mixture-of-expert models when matched for both compute and parameters. We find gains are especially pronounced for factual tasks. We provide a fully parallelizable memory layer implementation, demonstrating scaling laws with up to 128B memory parameters, pretrained to 1 trillion tokens, comparing to base models with up to 8B parameters. Vincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Scott Yih, Luke Zettlemoyer, Gargi Ghosh |
ICML | 4 |
| 2025 | SelfCite: Self-Supervised Alignment for Context Attribution in Large Language ModelsabstractWe introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses. Instead of only relying on costly and labor-intensive annotations, SelfCite leverages a reward signal provided by the LLM itself through context ablation: If a citation is necessary, removing the cited text from the context should prevent the same response; if sufficient, retaining the cited text alone should preserve the same response. This reward can guide the inference-time best-of-N sampling strategy to improve citation quality significantly, as well as be used in preference optimization to directly fine-tune the models for generating better citations. The effectiveness of SelfCite is demonstrated by increasing citation F1 up to 5.3 points on the LongBench-Cite benchmark across five long-form question answering tasks. The source code is available at https://github.com/facebookresearch/SelfCite. Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Shen 0001, Zhaofeng Wu, Hu Xu 0001, Xi Victoria Lin, James R. Glass, Shang-Wen Li 0001, Scott Yih |
ICML | 9 |
| 2025 | Meta CLIP 2: A Worldwide Scaling RecipeabstractContrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data.
In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP. Yung-Sung Chuang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, James R. Glass, Lifei Huang, Jason Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu 0003, Saining Xie, Scott Yih, Shang-Wen Li 0001, Hu Xu 0001 |
NeurIPS | 14 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 14 |
| 2025 | Group-Level Data Selection for Efficient PretrainingabstractThe efficiency and quality of language model pretraining are largely determined by the way pretraining data are selected. In this paper, we introduce *Group-MATES*, an efficient group-level data selection approach to optimize the speed-quality frontier of language model pretraining. Specifically, Group-MATES parameterizes costly group-level selection with a relational data influence model. To train this model, we sample training trajectories of the language model and collect oracle data influences alongside. The relational data influence model approximates the oracle data influence by weighting individual influence with relationships among training data. To enable efficient selection with our relational data influence model, we partition the dataset into small clusters using relationship weights and select data within each cluster independently. Experiments on DCLM 400M-4x, 1B-1x, and 3B-1x show that Group-MATES achieves 3.5\%-9.4\% relative performance gains over random selection across 22 downstream tasks, nearly doubling the improvements achieved by state-of-the-art individual data selection baselines. Furthermore, Group-MATES reduces the number of tokens required to reach a certain downstream performance by up to 1.75x, substantially elevating the speed-quality frontier. Further analyses highlight the critical role of relationship weights in the relational data influence model and the effectiveness of our cluster-based inference. Our code is open-sourced at https://github.com/facebookresearch/Group-MATES. Zichun Yu, Arnold Overwijk, Scott Yih, Chenyan Xiong |
NeurIPS | 5 |
| 2024 | Instruction-tuned Language Models are Better Knowledge LearnersabstractZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, Srini Iyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhengbao Jiang, Zhiqing Sun, Pedro Rodríguez 0001, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Scott Yih, Srinivasan Iyer 0001 |
ACL (1) | 8 |
| 2024 | MoDE: CLIP Data Experts via ClusteringabstractThe success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web- crawled data. We present Mixture of Data Experts (MoDE) and learn a system of CLIP data experts via clustering. Each data expert is trained on one data cluster, being less sensitive to false negative noises in other clusters. At inference time, we ensemble their outputs by applying weights determined through the correlation between task metadata and cluster conditions. To estimate the correlation pre-cisely, the samples in one cluster should be semantically similar, but the number of data experts should still be rea-sonable for training and inference. As such, we consider the ontology in human language and propose to use fine- grained cluster centers to represent each data expert at a coarse-grained level. Experimental studies show that four CLIP data experts on ViT-B/16 outperform the ViT-L/14 by OpenAI CLIP and OpenCLIP on zero-shot image classification but with less (<35%) training cost. Meanwhile, MoDE can train all data expert asynchronously and can flexibly include new data experts. The code is available here. Jiawei Ma, Po-Yao Huang 0001, Saining Xie, Shang-Wen Li 0001, Luke Zettlemoyer, Shih-Fu Chang, Scott Yih, Hu Xu 0001 |
CVPR | 7 |
| 2024 | Few-Shot Data Synthesis for Open Domain Multi-Hop Question AnsweringabstractFew-shot learning for open domain multi-hop question answering typically relies on the incontext learning capability of large language models (LLMs).While powerful, these LLMs usually contain tens or hundreds of billions of parameters, making them rather inefficient at inference time.To improve performance of smaller language models, we propose a data synthesis framework for multi-hop question answering that requires less than 10 humanannotated question answer pairs.Our framework depends only on rich, naturally-occurring relationships among documents and is built upon the data generation functions parameterized by LLMs and prompts.We synthesize millions of multi-hop questions and claims to finetune language models, evaluated on popular benchmarks for multi-hop question answering and fact verification.Empirically, our approach improves model performance significantly, allowing the finetuned models to be competitive with GPT-3.5 based approaches while being almost one-third the size in parameter count.What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?Query1: the eastern section of the Colorado orogeny Query2: the elevation range for the High Plains Unanswerable Question generation Query generation Query verification Question: What is the elevation range for … Query: the eastern section of the Colorado … Observation: 1: The Colorado orogeny, or Colorado orogen, was an orogeny … … Answer: 1,800 to 7,000 ft Question answering 1,800 to 7,000 ft Randomly sample document pairsThe causes of World War II are debated, but contributing factors included the Second Italo-Ethiopian War, Spanish Civil War, Second Sino-Japanese War, Soviet-Japanese border conflicts, the rise of fascism in Europe, and European tensions in the aftermath of World War I.The Second Italo-Ethiopian War, also referred to as the Second Italo-Abyssinian War, was a war of aggression which was fought between Italy and Ethiopia from October 1935 to February 1937. Events occurred in sequenceThe Colorado orogeny, or Colorado orogen, was an orogeny in Colorado and surrounding areas which was a part of the development of the ancestral Rockies.The eastern sector extends into the High Plains and is called the Central Plains orogeny.The High Plains are a subregion of the Great Plains.From east to west, the High Plains rise in elevation from around 1,800 to 7,000 ft (550 to 2,130 m). Extra geographical informationNew York, often called New York City or NYC, is the most populous city in the United States.With a 2020 population of 8,804,190 distributed over 300.46 square miles (778.2 km 2 ), the city is the most densely populated major city in the United States.NYC is more than twice as populous as Los Angeles, the nation's second-largest city.Los Angeles, often referred to by its initials L.A., officially the City of Los Angeles, is the most populous city in the U.S. state of California. Extra demographic information Mingda Chen, Xilun Chen 0002, Scott Yih |
EACL (1) | 3 |
| 2024 | Altogether: Image Captioning via Re-aligning Alt-textabstractHu Xu, Po-Yao Huang, Xiaoqing Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy, Luke Zettlemoyer, Wen-tau Yih, Shang-Wen Li, Saining Xie, Christoph Feichtenhofer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hu Xu 0001, Po-Yao Huang 0001, Xiaoqing Ellen Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy, Luke Zettlemoyer, Scott Yih, Shang-Wen Li 0001, Saining Xie, Christoph Feichtenhofer |
EMNLP | 10 |
| 2024 | RA-DIT: Retrieval-Augmented Dual Instruction TuningabstractRetrieval-augmented language models (RALMs) improve performance by accessing long-tail and up-to-date knowledge from external data stores, but are challenging to build. Existing approaches require either expensive retrieval-specific modifications to LM pre-training or use post-hoc integration of the data store that leads to suboptimal performance. We introduce Retrieval-Augmented Dual Instruction Tuning (RA-DIT), a lightweight fine-tuning methodology that provides a third option by retrofitting any LLM with retrieval capabilities. Our approach operates in two distinct fine-tuning steps: (1) one updates a pre-trained LM to better use retrieved information, while (2) the other updates the retriever to return more relevant results, as preferred by the LM. By fine-tuning over tasks that require both knowledge utilization and contextual awareness, we demonstrate that each stage yields significant performance improvements, and using both leads to additional gains. Our best model, RA-DIT 65B, achieves state-of-the-art performance across a range of knowledge-intensive zero- and few-shot learning benchmarks, significantly outperforming existing in-context RALM approaches by up to +8.9% in 0-shot setting and +1.4% in 5-shot setting on average. Xi Victoria Lin, Xilun Chen 0002, Mingda Chen, Maria Lomeli, Richard James 0001, Pedro Rodríguez 0001, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, Scott Yih |
ICLR | 12 |
| 2024 | In-Context Pretraining: Language Modeling Beyond Document BoundariesabstractLanguage models are currently trained to predict tokens given document prefixes, enabling them to zero shot long form generation and prompting-style tasks which can be reduced to document completion. We instead present IN-CONTEXT PRETRAINING, a new approach where language models are trained on a sequence of related documents, thereby explicitly encouraging them to read and reason across document boundaries. Our approach builds on the fact that current pipelines train by concatenating random sets of shorter documents to create longer context windows; this improves efficiency even though the prior documents provide no signal for predicting the next document. Given this fact, we can do IN-CONTEXT PRETRAINING by simply changing the document ordering so that each context contains related documents, and directly applying existing pretraining pipelines. However, this document sorting problem is challenging. There are billions of documents and we would like the sort to maximize contextual similarity for every document without repeating any data. To do this, we introduce approximate algorithms for finding related documents with efficient nearest neighbor search and constructing coherent batches with a graph cover algorithm. Our experiments show IN-CONTEXT PRETRAINING offers a scalable and simple approach to significantly enhance LM performance: we see notable improvements in tasks that require more complex contextual reasoning, including in-context learning (+8%), reading comprehension (+15%), faithfulness to previous contexts (+16%), long-context reasoning (+5%), and retrieval augmentation (+9%). Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Scott Yih, Mike Lewis |
ICLR | 9 |
| 2024 | REPLUG: Retrieval-Augmented Black-Box Language ModelsabstractWeijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, Wen-tau Yih. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James 0001, Mike Lewis, Luke Zettlemoyer, Scott Yih |
NAACL-HLT | 8 |
| 2024 | Nearest Neighbor Speculative Decoding for LLM Generation and AttributionabstractLarge language models (LLMs) often hallucinate and lack the ability to provide attribution for their generations. Semi-parametric LMs, such as kNN-LM, approach these limitations by refining the output of an LM for a given prompt using its nearest neighbor matches in a non-parametric data store. However, these models often exhibit slow inference speeds and produce non-fluent texts. In this paper, we introduce Nearest Neighbor Speculative Decoding (NEST), a novel semi-parametric language modeling approach that is capable of incorporating real-world text spans of arbitrary length into the LM generations and providing attribution to their sources. NEST performs token-level retrieval at each inference step to compute a semi-parametric mixture distribution and identify promising span continuations in a corpus. It then uses an approximate speculative decoding procedure that accepts a prefix of the retrieved span or generates a new token. NEST significantly enhances the generation quality and attribution rate of the base LM across a variety of knowledge-intensive tasks, surpassing the conventional kNN-LM method and performing competitively with in-context retrieval augmentation. In addition, NEST substantially improves the generation speed, achieving a 1.8x speedup in inference time when applied to Llama-2-Chat 70B. Code will be released at https://github.com/facebookresearch/NEST/tree/main. Minghan Li 0002, Xilun Chen 0002, Ari Holtzman, Beidi Chen, Jimmy Lin, Scott Yih, Xi Victoria Lin |
NeurIPS | 6 |
| 2024 | FLAME : Factuality-Aware Alignment for Large Language ModelsabstractAlignment is a procedure to fine-tune pre-trained large language models (LLMs) to follow natural language instructions and serve as helpful AI assistants.
We have observed, however, that the conventional alignment process fails to enhance the factual accuracy of LLMs, and often leads to the generation of more false facts (i.e., *hallucination*).
In this paper, we study how to make the LLM alignment process more factual, by first identifying factors that lead to hallucination in both alignment steps: supervised fine-tuning (SFT) and reinforcement learning (RL).
In particular, we find that training the LLM on new or unfamiliar knowledge can encourage hallucination.
This makes SFT less factual as it trains on human-labeled data that may be novel to the LLM.
Furthermore, reward functions used in standard RL often inadequately capture factuality and favor longer and more detailed responses, which inadvertently promote hallucination.
Based on these observations, we propose *FactuaLity-aware AlignMEnt*, comprised of *factuality-aware SFT* and *factuality-aware RL* through direct preference optimization.
Experiments show that our proposed *FLAME* guides LLMs to output more factual responses while maintaining their instruction-following capability. Sheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong, Jimmy Lin, Scott Yih, Xilun Chen 0002 |
NeurIPS | 6 |
| 2024 | CRAG - Comprehensive RAG BenchmarkabstractRetrieval-Augmented Generation (RAG) has recently emerged as a promising solution to alleviate Large Language Model (LLM)’s deficiency in lack of knowledge. Existing RAG datasets, however, do not adequately represent the diverse and dynamic nature of real-world Question Answering (QA) tasks. To bridge this gap, we introduce the Comprehensive RAG Benchmark (CRAG), a factual question answering benchmark of 4,409 question-answer pairs and mock APIs to simulate web and Knowledge Graph (KG) search. CRAG is designed to encapsulate a diverse array of questions across five domains and eight question categories, reflecting varied entity popularity from popular to long-tail, and temporal dynamisms ranging from years to seconds. Our evaluation on this benchmark highlights the gap to fully trustworthy QA. Whereas most advanced LLMs achieve $\le 34\%$ accuracy on CRAG, adding RAG in a straightforward manner improves the accuracy only to 44%. State-of-the-art industry RAG solutions only answer 63% questions without any hallucination. CRAG also reveals much lower accuracy in answering questions regarding facts with higher dynamism, lower popularity, or higher complexity, suggesting future research directions. The CRAG benchmark laid the groundwork for a KDD Cup 2024 challenge, attracted thousands of participants and submissions. We commit to maintaining CRAG to serve research communities in advancing RAG solutions and general QA solutions. CRAG is available at https://github.com/facebookresearch/CRAG/. Kai Sun 0006, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Daniel Gui, Ziran Will Jiang, Ziyu Jiang, Lingkun Kong, Brian Moran, Eting Yuan, Hanwen Zha, Nan Tang 0001, Lei Chen 0002, Nicolas Scheffer, Rakesh Wanga, Scott Yih, Xin Dong 0001 |
NeurIPS | 26 |
| 2023 | CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector RetrievalabstractMinghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Minghan Li 0002, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Scott Yih, Xilun Chen 0002 |
ACL (1) | 7 |
| 2023 | Learning to Simulate Natural Language Feedback for Interactive Semantic ParsingabstractInteractive semantic parsing based on natural language (NL) feedback, where users provide feedback to correct the parser mistakes, has emerged as a more practical scenario than the traditional one-shot semantic parsing.However, prior work has heavily relied on humanannotated feedback data to train the interactive semantic parser, which is prohibitively expensive and not scalable.In this work, we propose a new task of simulating NL feedback for interactive semantic parsing.We accompany the task with a novel feedback evaluator.The evaluator is specifically designed to assess the quality of the simulated feedback, based on which we decide the best feedback simulator from our proposed variants.On a text-to-SQL dataset, we show that our feedback simulator can generate high-quality NL feedback to boost the error correction ability of a specific parser.In low-data settings, our feedback simulator can help achieve comparable error correction performance as trained using the costly, full set of human annotations.1 Yintao Tai, Sida I. Wang, Scott Yih, Ziyu Yao 0002 |
ACL (1) | 5 |
| 2023 | FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationabstractSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Scott Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 5 |
| 2023 | InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, Mike Lewis |
ICLR | 8 |
| 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationabstractWe introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io. Yuhang Lai, Chengxi Li 0011, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Sida I. Wang, Tao Yu 0009 |
ICML | 7 |
| 2023 | LEVER: Learning to Verify Language-to-Code Generation with ExecutionabstractThe advent of large language models trained on code (code LLMs) has led to significant progress in language-to-code generation. State-of-the-art approaches in this area combine LLM decoding with sample pruning and reranking using test cases or heuristics based on the execution results. However, it is challenging to obtain test cases for many real-world language-to-code applications, and heuristics cannot well capture the semantic features of the execution results, such as data type and value range, which often indicates the correctness of the program. In this work, we propose LEVER, a simple approach to improve language-to-code generation by learning to verify the generated programs with their execution results. Specifically, we train verifiers to determine whether a program sampled from the LLMs is correct or not based on the natural language input, the program itself and its execution results. The sampled programs are reranked by combining the verification score with the LLM generation probability, and marginalizing over programs with the same execution results. On four datasets across the domains of table QA, math QA and basic Python programming, LEVER consistently improves over the base code LLMs (4.6% to 10.9% with code-davinci-002) and achieves new state-of-the-art results on all of them. Ansong Ni, Srinivasan Iyer 0001, Dragomir R. Radev, Veselin Stoyanov, Scott Yih, Sida I. Wang, Xi Victoria Lin |
ICML | 5 |
| 2023 | Retrieval-Augmented Multimodal Language ModelingabstractRecent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all their knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). Specifically, for the retriever, we use a pretrained CLIP, and for the generator, we train a CM3 Transformer on the LAION dataset. Our resulting model, named Retrieval-Augmented CM3 (RA-CM3), is the first multimodal model that can retrieve and generate both text and images. We show that RA-CM3 significantly outperforms baseline multimodal models such as DALL-E and CM3 on both image and caption generation tasks (12 FID and 17 CIDEr improvements on MS-COCO), while requiring much less compute for training (<30% of DALL-E). Moreover, we show that RA-CM3 exhibits novel capabilities such as faithful image generation and multimodal in-context learning (e.g., image generation from demonstrations). Michihiro Yasunaga, Armen Aghajanyan, Richard James 0001, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, Scott Yih |
ICML | 9 |
| 2023 | Coder Reviewer Reranking for Code GenerationabstractSampling diverse programs from a code language model and reranking with model likelihood is a popular method for code generation but it is prone to preferring degenerate solutions. Inspired by collaborative programming, we propose Coder-Reviewer reranking. We augment Coder language models from past work, which generate programs given language instructions, with Reviewer models, which evaluate the likelihood of the instruction given the generated programs. We perform an extensive study across six datasets with eight models from three model families. Experimental results show that Coder-Reviewer reranking leads to consistent and significant improvement (up to 17% absolute accuracy gain) over reranking with the Coder model only. When combined with executability filtering, Coder-Reviewer reranking can often outperform the minimum Bayes risk method. Coder-Reviewer reranking is easy to implement by prompting, can generalize to different programming languages, and works well with off-the-shelf hyperparameters. Tao Yu 0009, Tatsunori B. Hashimoto, Mike Lewis, Scott Yih, Daniel Fried, Sida I. Wang |
ICML | 5 |
| 2022 | On Continual Model Refinement in Out-of-Distribution Data StreamsabstractReal-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrophic forgetting.However, existing continual learning (CL) problem setups cannot cover such a realistic and complex scenario.In response to this, we propose a new CL problem formulation dubbed continual model refinement (CMR).Compared to prior CL settings, CMR is more practical and introduces unique challenges (boundary-agnostic and non-stationary distribution shift, diverse mixtures of multiple OOD data clusters, error-centric streams, etc.).We extend several existing CL approaches to the CMR setting and evaluate them extensively.For benchmarking and analysis, we propose a general sampling algorithm to obtain dynamic OOD data streams with controllable nonstationarity, as well as a suite of metrics measuring various aspects of online performance.Our experiments and detailed analysis reveal the promise and challenges of the CMR problem, supporting that studying CMR in dynamic OOD streams can benefit the longevity of deployed NLP models in production. 1 Bill Y. Lin, Sida I. Wang, Xi Victoria Lin, Robin Jia, Xiang Ren 0001, Scott Yih |
ACL (1) | 7 |
| 2022 | UniPELT: A Unified Framework for Parameter-Efficient Language Model TuningabstractYuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Scott Yih, Madian Khabsa. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yuning Mao, Lambert Mathias, Amjad Almahairi, Hao Ma 0001, Jiawei Han 0001, Scott Yih, Madian Khabsa |
ACL (1) | 7 |
| 2022 | Improving Passage Retrieval with Zero-Shot Question GenerationabstractWe propose a simple and effective re-ranking method for improving passage retrieval in open question answering.The re-ranker re-scores retrieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage.This approach can be applied on top of any retrieval method (e.g.neural or keywordbased), does not require any domain-or taskspecific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question).When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy.We also obtain new stateof-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.1 Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Scott Yih, Joelle Pineau, Luke Zettlemoyer |
EMNLP | 5 |
| 2022 | DiffCSE: Difference-based Contrastive Learning for Sentence EmbeddingsabstractYung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang 0001, Shiyu Chang, Marin Soljacic, Shang-Wen Li 0001, Scott Yih, James R. Glass |
NAACL-HLT | 8 |
| 2022 | Boosted Dense RetrieverabstractPatrick Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Patrick S. H. Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel 0001 |
NAACL-HLT | 5 |
| 2022 | Simple Local Attentions Remain Competitive for Long-Context TasksabstractWenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen 0002, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad |
NAACL-HLT | 7 |
| 2022 | Autoregressive Search Engines: Generating Substrings as Document IdentifiersabstractKnowledge-intensive language tasks require NLP systems to both provide the correct answer and retrieve supporting evidence for it in a given corpus. Autoregressive language models are emerging as the de-facto standard for generating answers, with newer and more powerful systems emerging at an astonishing pace. In this paper we argue that all this (and future) progress can be directly applied to the retrieval problem with minimal intervention to the models' architecture. Previous work has explored ways to partition the search space into hierarchical structures and retrieve documents by autoregressively generating their unique identifier. In this work we propose an alternative that doesn't force any structure in the search space: using all ngrams in a passage as its possible identifiers. This setup allows us to use an autoregressive model to generate and score distinctive ngrams, that are then mapped to full passages through an efficient data structure. Empirically, we show this not only outperforms prior autoregressive approaches but also leads to an average improvement of at least 10 points over more established retrieval solutions for passage-level retrieval on the KILT benchmark, establishing new state-of-the-art downstream performance on some datasets, while using a considerably lighter memory footprint than competing systems. Code available in the supplementary materials. Pre-trained models will be made available. Michele Bevilacqua, Giuseppe Ottaviano, Patrick S. H. Lewis, Scott Yih, Sebastian Riedel 0001, Fabio Petroni |
NeurIPS | 4 |
| 2022 | BiT: Robustly Binarized Multi-distilled TransformerabstractModern pre-trained transformers have rapidly advanced the state-of-the-art in machine learning, but have also grown in parameters and computational complexity, making them increasingly difficult to deploy in resource-constrained environments. Binarization of the weights and activations of the network can significantly alleviate these issues, however, is technically challenging from an optimization perspective. In this work, we identify a series of improvements that enables binary transformers at a much higher accuracy than what was possible previously. These include a two-set binarization scheme, a novel elastic binary activation function with learned parameters, and a method to quantize a network to its limit by successively distilling higher precision models into lower precision students. These approaches allow for the first time, fully binarized transformer models that are at a practical level of accuracy, approaching a full-precision BERT baseline on the GLUE language understanding benchmark within as little as 5.9%. Code and models are available at:https://github.com/facebookresearch/bit. Zechun Liu, Barlas Oguz, Aasish Pappu, Scott Yih, Meng Li 0004, Raghuraman Krishnamoorthi, Yashar Mehdad |
NeurIPS | 5 |
| 2022 | QUASER: Question Answering with Scalable Extractive RationalizationabstractDesigning natural language processing (NLP) models that produce predictions by first extracting a set of relevant input sentences, i.e., rationales, is gaining importance for improving model interpretability and producing supporting evidence for users. Current unsupervised approaches are designed to extract rationales that maximize prediction accuracy, which is invariably obtained by exploiting spurious correlations in datasets, and leads to unconvincing rationales. In this paper, we introduce unsupervised generative models to extract dual-purpose rationales, which must not only be able to support a subsequent answer prediction, but also support a reproduction of the input query. We show that such models can produce more meaningful rationales, that are less influenced by dataset artifacts, and as a result, also achieve the state-of-the-art on rationale extraction metrics on four datasets from the ERASER benchmark, significantly improving upon previous unsupervised methods. Our multi-task model is scalable and enables using state-of-the-art pretrained language models to design explainable question answering systems. Asish Ghoshal, Srinivasan Iyer 0001, Bhargavi Paranjape, Kushal Lakhotia, Scott Yih, Yashar Mehdad |
SIGIR | 5 |
| 2021 | On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized StudyabstractDivyansh Kaushik, Douwe Kiela, Zachary C. Lipton, Wen-tau Yih. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Divyansh Kaushik, Douwe Kiela, Zachary C. Lipton, Scott Yih |
ACL/IJCNLP (1) | 4 |
| 2021 | Multi-Task Retrieval for Knowledge-Intensive TasksabstractJean Maillard, Vladimir Karpukhin, Fabio Petroni, Wen-tau Yih, Barlas Oguz, Veselin Stoyanov, Gargi Ghosh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jean Maillard, Vladimir Karpukhin, Fabio Petroni, Scott Yih, Barlas Oguz, Veselin Stoyanov, Gargi Ghosh |
ACL/IJCNLP (1) | 4 |
| 2021 | Joint Verification and Reranking for Open Fact Checking Over TablesabstractMichael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, Sebastian Riedel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Scott Yih, Sebastian Riedel 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | FiD-Ex: Improving Sequence-to-Sequence Models for Extractive Rationale GenerationabstractNatural language (NL) explanations of model predictions are gaining popularity as a means to understand and verify decisions made by large black-box pre-trained models, for tasks such as Question Answering (QA) and Fact Verification.Recently, pre-trained sequence to sequence (seq2seq) models have proven to be very effective in jointly making predictions, as well as generating NL explanations.However, these models have many shortcomings; they can fabricate explanations even for incorrect predictions, they are difficult to adapt to long input documents, and their training requires a large amount of labeled data.In this paper, we develop FiD-Ex 1 , which addresses these shortcomings for seq2seq models by: 1) introducing sentence markers to eliminate explanation fabrication by encouraging extractive generation, 2) using the fusion-in-decoder architecture to handle long input contexts, and 3) intermediate fine-tuning on re-structured open domain QA datasets to improve few-shot performance.FiD-Ex significantly improves over prior work in terms of explanation metrics and task accuracy on five tasks from the ERASER explainability benchmark in both fully supervised and few-shot settings. Kushal Lakhotia, Bhargavi Paranjape, Asish Ghoshal, Scott Yih, Yashar Mehdad, Srinivasan Iyer 0001 |
EMNLP (1) | 4 |
| 2021 | On the Influence of Masking Policies in Intermediate Pre-trainingabstractCurrent NLP models are predominantly trained through a two-stage "pre-train then fine-tune" pipeline.Prior work has shown that inserting an intermediate pre-training stage, using heuristic masking policies for masked language modeling (MLM), can significantly improve final performance.However, it is still unclear (1) in what cases such intermediate pre-training is helpful, (2) whether hand-crafted heuristic objectives are optimal for a given task, and (3) whether a masking policy designed for one task is generalizable beyond that task.In this paper, we perform a large-scale empirical study to investigate the effect of various masking policies in intermediate pre-training with nine selected tasks across three categories.Crucially, we introduce methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.We conclude that the success of intermediate pre-training is dependent on appropriate pre-train corpus, selection of output format (i.e., masked spans or full sentence), and clear understanding of the role that MLM plays for the downstream task.In addition, we find our learned masking policies outperform the heuristic of masking named entities on TriviaQA, and policies learned from one task can positively transfer to other tasks in certain cases, inviting future research in this direction. Qinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte, Hao Ma 0001, Scott Yih, Xiang Ren 0001, Madian Khabsa |
EMNLP (1) | 6 |
| 2021 | Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz |
ICLR | 8 |
| 2021 | RECONSIDER: Improved Re-Ranking using Span-Focused Cross-Attention for Open Domain Question AnsweringabstractSrinivasan Iyer, Sewon Min, Yashar Mehdad, Wen-tau Yih. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Srinivasan Iyer 0001, Sewon Min, Yashar Mehdad, Scott Yih |
NAACL-HLT | 4 |
| 2021 | On Unifying Misinformation DetectionabstractNayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma, Wen-tau Yih, Madian Khabsa. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma 0001, Scott Yih, Madian Khabsa |
NAACL-HLT | 6 |
| 2020 | TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataabstractRecent years have witnessed the burgeoning of pretrained language models (LMs) for textbased natural language (NL) understanding tasks.Such models are typically trained on free-form NL text, hence may not be suitable for tasks like semantic parsing over structured data, which require reasoning over both free-form NL questions and structured tabular data (e.g., database tables).In this paper we present TABERT, a pretrained LM that jointly learns representations for NL sentences and (semi-)structured tables.TABERT is trained on a large corpus of 26 million tables and their English contexts.In experiments, neural semantic parsers using TABERT as feature representation layers achieve new best results on the challenging weakly-supervised semantic parsing benchmark WIKITABLEQUESTIONS, while performing competitively on the text-to-SQL dataset SPIDER. 1 Graham Neubig, Scott Yih, Sebastian Riedel 0001 |
ACL | 3 |
| 2020 | Dense Passage Retrieval for Open-Domain Question AnsweringabstractVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen 0001, Scott Yih |
EMNLP (1) | 8 |
| 2020 | Efficient One-Pass End-to-End Entity Linking for QuestionsabstractWe present ELQ, a fast end-to-end entity linking model for questions, which uses a biencoder to jointly perform mention detection and linking in one pass.Evaluated on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question, ELQ outperforms the previous state of the art by a large margin of +12.7% and +19.6% F1, respectively.With a very fast inference time (1.57examples/s on a single CPU), ELQ can be useful for downstream question answering systems.In a proof-of-concept experiment, we demonstrate that using ELQ significantly improves the downstream QA performance of GraphRetriever (Min et al., 2019). 1 Belinda Z. Li, Sewon Min, Srinivasan Iyer 0001, Yashar Mehdad, Scott Yih |
EMNLP (1) | 5 |
| 2020 | Unsupervised Question Decomposition for Question AnsweringabstractWe aim to improve question answering (QA) by decomposing hard questions into simpler sub-questions that existing QA systems are capable of answering.Since labeling questions with decompositions is cumbersome, we take an unsupervised approach to produce sub-questions, also enabling us to leverage millions of questions from the internet.Specifically, we propose an algorithm for One-to-N Unsupervised Sequence transduction (ONUS) that learns to map one hard, multi-hop question to many simpler, singlehop sub-questions.We answer sub-questions with an off-the-shelf QA model and give the resulting answers to a recomposition model that combines them into a final answer.We show large QA improvements on HOTPOTQA over a strong baseline on the original, out-ofdomain, and multi-hop dev sets.ONUS automatically learns to decompose different kinds of questions, while matching the utility of supervised and heuristic decomposition methods for QA and exceeding those methods in fluency.Qualitatively, we find that using subquestions is promising for shedding light on why a QA system makes a prediction. 1 * KC was a part-time research scientist at Facebook AI Research while working on this paper.1 Our code, data, and pretrained models are available at https://github.com/facebookresearch/ UnsupervisedDecomposition. What profession do H. L. Mencken and Albert Camus have in common? Ethan Perez, Patrick S. H. Lewis, Scott Yih, Kyunghyun Cho, Douwe Kiela |
EMNLP (1) | 3 |
| 2020 | An Imitation Game for Learning Semantic Parsers from User InteractionabstractDespite the widely successful applications, building a semantic parser is still a tedious process in practice with challenges from costly data annotation and privacy risks.We suggest an alternative, human-in-the-loop methodology for learning semantic parsers directly from users.A semantic parser should be introspective of its uncertainties and prompt for user demonstrations when uncertain.In doing so it also gets to imitate the user behavior and continue improving itself autonomously with the hope that eventually it may become as good as the user in interpreting their questions.To combat the sparsity of demonstrations, we propose a novel annotation-efficient imitation learning algorithm, which iteratively collects new datasets by mixing demonstrated states and confident predictions and retrains the semantic parser in a Dataset Aggregation fashion (Ross et al., 2011).We provide a theoretical analysis of its cost bound and also empirically demonstrate its promising performance on the text-to-SQL problem. 1 Ziyu Yao 0002, Yiqi Tang, Scott Yih, Huan Sun 0001, Yu Su 0001 |
EMNLP (1) | 3 |
| 2020 | Abductive Commonsense Reasoning
Chandra Bhagavatula, Ronan Le Bras 0001, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, Yejin Choi 0001 |
ICLR | 8 |
| 2020 | Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksabstractLarge pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline. Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal 0001, Heinrich Küttler, Mike Lewis, Scott Yih, Tim Rocktäschel, Sebastian Riedel 0001, Douwe Kiela |
NeurIPS | 9 |
| 2019 | QUAREL: A Dataset and Models for Answering Questions about Qualitative RelationshipsabstractMany natural la guage questions require recognizing and reasoning with qualitative relationships (e.g., in science, economics, and medicine), but are challenging to answer with corpus-based methods. Qualitative modeling provides tools that support such reasoning, but the semantic parsing task of mapping questions into those models has formidable challenges. We present QUAREL, a dataset of diverse story questions involving qualitative relationships that characterize these challenges, and techniques that begin to address them. The dataset has 2771 questions relating 19 different types of quantities. For example, “Jenny observes that the robot vacuum cleaner moves slower on the living room carpet than on the bedroom carpet. Which carpet has more friction?” We contribute (1) a simple and flexible conceptual framework for representing these kinds of questions; (2) the QUAREL dataset, including logical forms, exemplifying the parsing challenges; and (3) two novel models for this task, built as extensions of type-constrained semantic parsing. The first of these models (called QUASP+) significantly outperforms off-the-shelf tools on QUAREL. The second (QUASP+ZERO) demonstrates zero-shot capability, i.e., the ability to handle new qualitative relationships without requiring additional training data, something not possible with previous models. This work thus makes inroads into answering complex, qualitative questions that require reasoning, and scaling to new relationships at low cost. The dataset and models are available at http://data.allenai.org/quarel. Oyvind Tafjord, Peter Clark, Matt Gardner 0001, Scott Yih, Ashish Sabharwal |
AAAI | 4 |
| 2019 | Everything Happens for a Reason: Discovering the Purpose of Actions in Procedural TextabstractBhavana Dalvi, Niket Tandon, Antoine Bosselut, Wen-tau Yih, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bhavana Dalvi, Niket Tandon, Antoine Bosselut, Scott Yih, Peter Clark |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Model-based Interactive Semantic Parsing: A Unified Framework and A Text-to-SQL Case StudyabstractZiyu Yao, Yu Su, Huan Sun, Wen-tau Yih. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ziyu Yao 0002, Yu Su 0001, Huan Sun 0001, Scott Yih |
EMNLP/IJCNLP (1) | 4 |
| 2019 | FlowQA: Grasping Flow in History for Conversational Machine Comprehension
Hsin-Yuan Huang, Eunsol Choi, Scott Yih |
ICLR (Poster) | 3 |
| 2018 | A Knowledge-Grounded Neural Conversation ModelabstractNeural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more useful conversational applications. This paper presents a novel, fully data-driven, and knowledge-grounded neural conversation model aimed at producing more contentful responses. We generalize the widely-used Sequence-to-Sequence (Seq2Seq) approach by conditioning responses on both conversation history and external “facts”, allowing the model to be versatile and applicable in an open-domain setting. Our approach yields significant improvements over a competitive Seq2Seq baseline. Human judges found that our outputs are significantly more informative. Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, William B. Dolan, Jianfeng Gao 0001, Scott Yih, Michel Galley |
AAAI | 6 |
| 2018 | QuAC: Question Answering in ContextabstractWe present QuAC, a dataset for Question Answering in Context that contains 14K information-seeking QA dialogs (100K questions in total).The dialogs involve two crowd workers: (1) a student who poses a sequence of freeform questions to learn as much as possible about a hidden Wikipedia text, and (2) a teacher who answers the questions by providing short excerpts from the text.QuAC introduces challenges not found in existing machine comprehension datasets: its questions are often more open-ended, unanswerable, or only meaningful within the dialog context, as we show in a detailed qualitative evaluation.We also report results for a number of reference models, including a recently state-ofthe-art reading comprehension architecture extended to model dialog context.Our best model underperforms humans by 20 F1, suggesting that there is significant room for future work on this data.Dataset, baseline, and leaderboard available at http://quac.ai.How was perversion handled?How long was he there?How popular did she become?How did Mark Felt contact Woodword?How did the meeting go?How did it do on the charts?When was she born?When was it founded?When was the breakup? Eunsol Choi, He He 0001, Mohit Iyyer, Mark Yatskar, Scott Yih, Yejin Choi 0001, Percy Liang, Luke Zettlemoyer |
EMNLP | 5 |
| 2018 | Policy Shaping and Generalized Update Equations for Semantic Parsing from DenotationsabstractSemantic parsing from denotations faces two key challenges in model training: (1) given only the denotations (e.g., answers), search for good candidate semantic parses, and (2) choose the best model update algorithm.We propose effective and general solutions to each of them.Using policy shaping, we bias the search procedure towards semantic parses that are more compatible to the text, which provide better supervision signals for training.In addition, we propose an update equation that generalizes three different families of learning algorithms, which enables fast model exploration.When experimented on a recently proposed sequential question answering dataset, our framework leads to a new state-of-theart model that outperforms previous work by 5.0% absolute on exact match accuracy.Question: what nation scored the most points Dipendra Misra, Ming-Wei Chang, Xiaodong He 0001, Scott Yih |
EMNLP | 4 |
| 2018 | Dissecting Contextual Word Embeddings: Architecture and RepresentationabstractContextual word representations derived from pre-trained bidirectional language models (biLMs) have recently been shown to provide significant improvements to the state of the art for a wide range of NLP tasks.However, many questions remain as to how and why these models are so effective.In this paper, we present a detailed empirical study of how the choice of neural architecture (e.g.LSTM, CNN, or self attention) influences both end task accuracy and qualitative properties of the representations that are learned.We show there is a tradeoff between speed and accuracy, but all architectures learn high quality contextual representations that outperform word embeddings for four challenging NLP tasks.Additionally, all architectures learn representations that vary with network depth, from exclusively morphological based at the word embedding layer through local syntax based in the lower contextual layers to longer range semantics such coreference at the upper layers.Together, these results suggest that unsupervised biLMs, independent of architecture, are learning much more about the structure of language than previously appreciated. Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, Scott Yih |
EMNLP | 4 |
| 2018 | Reasoning about Actions and State Changes by Injecting Commonsense KnowledgeabstractComprehending procedural text, e.g., a paragraph describing photosynthesis, requires modeling actions and the state changes they produce, so that questions about entities at different timepoints can be answered.Although several recent systems have shown impressive progress in this task, their predictions can be globally inconsistent or highly improbable.In this paper, we show how the predicted effects of actions in the context of a paragraph can be improved in two ways: (1) by incorporating global, commonsense constraints (e.g., a non-existent entity cannot be destroyed), and (2) by biasing reading with preferences from large-scale corpora (e.g., trees rarely move).Unlike earlier methods, we treat the problem as a neural structured prediction task, allowing hard and soft constraints to steer the model away from unlikely predictions.We show that the new model significantly outperforms earlier systems on a benchmark dataset for procedural text comprehension (+8% relative gain), and that it also avoids some of the nonsensical predictions that earlier systems make. Niket Tandon, Bhavana Dalvi, Joel Grus, Scott Yih, Antoine Bosselut, Peter Clark |
EMNLP | 4 |
| 2018 | Tracking State Changes in Procedural Text: a Challenge Dataset and Models for Process Paragraph ComprehensionabstractBhavana Dalvi, Lifu Huang, Niket Tandon, Wen-tau Yih, Peter Clark. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Bhavana Dalvi, Lifu Huang, Niket Tandon, Scott Yih, Peter Clark |
NAACL-HLT | 4 |
| 2017 | Search-based Neural Structured Learning for Sequential Question AnsweringabstractRecent work in semantic parsing for question answering has focused on long and complicated questions, many of which would seem unnatural if asked in a normal conversation between two humans.In an effort to explore a conversational QA setting, we present a more realistic task: answering sequences of simple but inter-related questions.We collect a dataset of 6,066 question sequences that inquire about semistructured tables from Wikipedia, with 17,553 question-answer pairs in total.To solve this sequential question answering task, we propose a novel dynamic neural semantic parsing framework trained using a weakly supervised reward-guided search.Our model effectively leverages the sequential context to outperform state-of-the-art QA systems that are designed to answer highly complex questions. Mohit Iyyer, Scott Yih, Ming-Wei Chang |
ACL (1) | 2 |
| 2017 | Maximum Margin Reward Networks for Learning from Explicit and Implicit SupervisionabstractNeural networks have achieved state-ofthe-art performance on several structuredoutput prediction tasks, trained in a fully supervised fashion.However, annotated examples in structured domains are often costly to obtain, which thus limits the applications of neural networks.In this work, we propose Maximum Margin Reward Networks, a neural networkbased framework that aims to learn from both explicit (full structures) and implicit supervision signals (delayed feedback on the correctness of the predicted structure).On named entity recognition and semantic parsing, our model outperforms previous systems on the benchmark datasets, CoNLL-2003 and WebQuestionsSP. Haoruo Peng, Ming-Wei Chang, Scott Yih |
EMNLP | 3 |
| 2017 | Cross-Sentence N-ary Relation Extraction with Graph LSTMsabstractPast work in relation extraction has focused on binary relations in single sentences. Recent NLP inroads in high-value domains have sparked interest in the more general setting of extracting n-ary relations that span multiple sentences. In this paper, we explore a general relation extraction framework based on graph long short-term memory networks (graph LSTMs) that can be easily extended to cross-sentence n-ary relation extraction. The graph formulation provides a unified way of exploring different LSTM approaches and incorporating various intra-sentential and inter-sentential dependencies, such as sequential, syntactic, and discourse relations. A robust contextual representation is learned for the entities, which serves as input to the relation classifier. This simplifies handling of relations with arbitrary arity, and enables multi-task learning with related relations. We evaluate this framework in two important precision medicine settings, demonstrating its effectiveness with both conventional supervised learning and distant supervision. Cross-sentence extraction produced larger knowledge bases. and multi-task learning significantly improved extraction accuracy. A thorough analysis of various LSTM approaches yielded useful insight the impact of linguistic analysis on extraction accuracy. Nanyun Peng 0001, Hoifung Poon, Chris Quirk, Kristina Toutanova, Scott Yih |
Trans. Assoc. Comput. Linguistics | 5 |
| 2016 | Compositional Learning of Embeddings for Relation Paths in Knowledge Base and TextabstractModeling relation paths has offered significant gains in embedding models for knowledge base (KB) completion.However, enumerating paths between two entities is very expensive, and existing approaches typically resort to approximation with a sampled subset.This problem is particularly acute when text is jointly modeled with KB relations and used to provide direct evidence for facts mentioned in it.In this paper, we propose the first exact dynamic programming algorithm which enables efficient incorporation of all relation paths of bounded length, while modeling both relation types and intermediate nodes in the compositional path representations.We conduct a theoretical analysis of the efficiency gain from the approach.Experiments on two datasets show that it addresses representational limitations in prior approaches and improves accuracy in KB completion. Kristina Toutanova, Xi Victoria Lin, Scott Yih, Hoifung Poon, Chris Quirk |
ACL (1) | 3 |
| 2016 | Learning from Explicit and Implicit Supervision Jointly For Algebra Word ProblemsabstractAutomatically solving algebra word problems has raised considerable interest recently.Existing state-of-the-art approaches mainly rely on learning from human annotated equations.In this paper, we demonstrate that it is possible to efficiently mine algebra problems and their numerical solutions with little to no manual effort.To leverage the mined dataset, we propose a novel structured-output learning algorithm that aims to learn from both explicit (e.g., equations) and implicit (e.g., solutions) supervision signals jointly.Enabled by this new algorithm, our model gains 4.6% absolute improvement in accuracy on the ALG-514 benchmark compared to the one without using implicit supervision.The final model also outperforms the current state-of-the-art approach by 3%. Shyam Upadhyay, Ming-Wei Chang, Kai-Wei Chang 0001, Scott Yih |
EMNLP | 4 |
| 2016 | Question Answering with Knowledge Base, Web and BeyondabstractIn this tutorial, we give the audience a coherent overview of the research of question answering (QA). We first introduce a variety of QA problems proposed by pioneer researchers and briefly describe the early efforts. By contrasting with the current research trend in this domain, the audience can easily comprehend what technical problems remain challenging and what the main breakthroughs and opportunities are during the past half century. For the rest of the tutorial, we select three categories of the QA problems that have recently attracted a great deal of attention in the research community, and present the tasks with the latest technical survey. We conclude the tutorial by discussing the new opportunities and future directions of QA research. Scott Yih, Hao Ma 0001 |
SIGIR | 1 |
| 2016 | Table Cell Search for Question AnsweringabstractTables are pervasive on the Web. Informative web tables range across a large variety of topics, which can naturally serve as a significant resource to satisfy user information needs. Driven by such observations, in this paper, we investigate an important yet largely under-addressed problem: Given millions of tables, how to precisely retrieve table cells to answer a user question. This work proposes a novel table cell search framework to attack this problem. We first formulate the concept of a relational chain which connects two cells in a table and represents the semantic relation between them. With the help of search engine snippets, our framework generates a set of relational chains pointing to potentially correct answer cells. We further employ deep neural networks to conduct more fine-grained inference on which relational chains best match the input question and finally extract the corresponding answer cells. Based on millions of tables crawled from the Web, we evaluate our framework in the open-domain question answering (QA) setting, using both the well-known WebQuestions dataset and user queries mined from Bing search engine logs. On WebQuestions, our framework is comparable to state-of-the-art QA systems based on knowledge bases (KBs), while on Bing queries, it outperforms other systems with a 56.7% relative gain. Moreover, when combined with results from our framework, KB-based QA performance can obtain a relative improvement of 28.1% to 66.7%, demonstrating that web tables supply rich knowledge that might not exist or is difficult to be identified in existing KBs. Huan Sun 0001, Hao Ma 0001, Xiaodong He 0001, Scott Yih, Yu Su 0001, Xifeng Yan |
WWW | 4 |
| 2015 | Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge BaseabstractWen-tau Yih, Ming-Wei Chang, Xiaodong He, Jianfeng Gao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Scott Yih, Ming-Wei Chang, Xiaodong He 0001, Jianfeng Gao 0001 |
ACL (1) | 1 |
| 2015 | WikiQA: A Challenge Dataset for Open-Domain Question AnsweringabstractWe describe the WIKIQA dataset, a new publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.Most previous work on answer sentence selection focuses on a dataset created using the TREC-QA data, which includes editor-generated questions and candidate answer sentences selected by matching content words in the question.WIKIQA is constructed using a more natural process and is more than an order of magnitude larger than the previous dataset.In addition, the WIKIQA dataset also includes questions for which there are no correct sentences, enabling researchers to work on answer triggering, a critical component in any QA system.We compare several systems on the task of answer sentence selection on both datasets and also describe the performance of a system on the problem of answer triggering using the WIKIQA dataset. Yi Yang 0038, Scott Yih, Christopher Meek |
EMNLP | 2 |
| 2015 | Deep Learning and Continuous Representations for Natural Language ProcessingabstractDeep learning techniques have demonstrated tremendous success in the speech and language processing community in recent years, establishing new state-ofthe-art performance in speech recognition, language modeling, and have shown great potential for many other natural language processing tasks. The focus of this tutorial is to provide an extensive overview on recent deep learning approaches to problems in language or text processing, with particular emphasis on important real-world applications including language understanding, semantic representation modeling, question answering and semantic parsing, etc. Scott Yih, Xiaodong He 0001, Jianfeng Gao 0001 |
HLT-NAACL | 1 |
| 2015 | Open Domain Question Answering via Semantic EnrichmentabstractMost recent question answering (QA) systems query large-scale knowledge bases (KBs) to answer a question, after parsing and transforming natural language questions to KBs-executable forms (e.g., logical forms). As a well-known fact, KBs are far from complete, so that information required to answer questions may not always exist in KBs. In this paper, we develop a new QA system that mines answers directly from the Web, and meanwhile employs KBs as a significant auxiliary to further boost the QA performance. Specifically, to the best of our knowledge, we make the first attempt to link answer candidates to entities in Freebase, during answer candidate generation. Several remarkable advantages follow: (1) Redundancy among answer candidates is automatically reduced. (2) The types of an answer candidate can be effortlessly determined by those of its corresponding entity in Freebase. (3) Capitalizing on the rich information about entities in Freebase, we can develop semantic features for each answer candidate after linking them to Freebase. Particularly, we construct answer-type related features with two novel probabilistic models, which directly evaluate the appropriateness of an answer candidate's types under a given question. Overall, such semantic features turn out to play significant roles in determining the true answers from the large answer candidate pool. The experimental results show that across two testing datasets, our QA system achieves an 18%~54% improvement under F_1 metric, compared with various existing QA systems. Huan Sun 0001, Hao Ma 0001, Scott Yih, Chen-Tse Tsai, Jingjing Liu 0001, Ming-Wei Chang |
WWW | 3 |
| 2014 | Learning Continuous Phrase Representations for Translation ModelingabstractThis paper tackles the sparsity problem in estimating phrase translation probabilities by learning continuous phrase representations, whose distributed nature enables the sharing of related phrases in their representations.A pair of source and target phrases are projected into continuous-valued vector representations in a low-dimensional latent space, where their translation score is computed by the distance between the pair in this new space.The projection is performed by a neural network whose weights are learned on parallel training data.Experimental evaluation has been performed on two WMT translation tasks.Our best result improves the performance of a state-of-the-art phrase-based statistical machine translation system trained on WMT 2012 French-English data by up to 1.3 BLEU points. Jianfeng Gao 0001, Xiaodong He 0001, Scott Yih, Li Deng 0001 |
ACL (1) | 3 |
| 2014 | Typed Tensor Decomposition of Knowledge Bases for Relation ExtractionabstractWhile relation extraction has traditionally been viewed as a task relying solely on textual data, recent work has shown that by taking as input existing facts in the form of entity-relation triples from both knowl-edge bases and textual data, the perfor-mance of relation extraction can be im-proved significantly. Following this new paradigm, we propose a tensor decompo-sition approach for knowledge base em-bedding that is highly scalable, and is es-pecially suitable for relation extraction. By leveraging relational domain knowl-edge about entity type information, our learning algorithm is significantly faster than previous approaches and is better able to discover new relations missing from the database. In addition, when ap-plied to a relation extraction task, our ap-proach alone is comparable to several ex-isting systems, and improves the weighted mean average precision of a state-of-the-art method by 10 points when used as a subcomponent. 1 Kai-Wei Chang 0001, Scott Yih, Bishan Yang, Christopher Meek |
EMNLP | 2 |
| 2014 | Joint semantic utterance classification and slot filling with recursive neural networksabstractIn recent years, continuous space models have proven to be highly effective at language processing tasks ranging from paraphrase detection to language modeling. These models are distinctive in their ability to achieve generalization through continuous space representations, and compositionality through arithmetic operations on those representations. Examples of such models include feed-forward and recurrent neural network language models. Recursive neural networks (RecNNs) extend this framework by providing an elegant mechanism for incorporating both discrete syntactic structure and continuous-space word and phrase representations into a powerful compositional model. In this paper, we show that RecNNs can be used to perform the core spoken language understanding (SLU) tasks in a spoken dialog system, more specifically domain and intent determination, concurrently with slot filling, in one jointly trained model. We find that a very simple RecNN model achieves competitive performance on the benchmark ATIS task, as well as on a Microsoft Cortana conversational understanding task. Zhaohan Guo, Gökhan Tür, Scott Yih, Geoffrey Zweig |
SLT | 3 |
| 2013 | Question Answering Using Enhanced Lexical Semantic Models
Scott Yih, Ming-Wei Chang, Christopher Meek, Andrzej Pastusiak |
ACL (1) | 1 |
| 2013 | Multi-Relational Latent Semantic AnalysisabstractWe present Multi-Relational Latent Semantic Analysis (MRLSA) which generalizes Latent Semantic Analysis (LSA).MRLSA provides an elegant approach to combining multiple relations between words by constructing a 3-way tensor.Similar to LSA, a lowrank approximation of the tensor is derived using a tensor decomposition.Each word in the vocabulary is thus represented by a vector in the latent semantic space and each relation is captured by a latent square matrix.The degree of two words having a specific relation can then be measured through simple linear algebraic operations.We demonstrate that by integrating multiple relations from both homogeneous and heterogeneous information sources, MRLSA achieves stateof-the-art performance on existing benchmark datasets for two relations, antonymy and is-a. Kai-Wei Chang 0001, Scott Yih, Christopher Meek |
EMNLP | 2 |
| 2013 | Animacy Detection with Voting ModelsabstractAnimacy detection is a problem whose solution has been shown to be beneficial for a number of syntactic and semantic tasks.We present a state-of-the-art system for this task which uses a number of simple classifiers with heterogeneous data sources in a voting scheme.We show how this framework can give us direct insight into the behavior of the system, allowing us to more easily diagnose sources of error. Joshua L. Moore, Christopher J. C. Burges, Erin Renshaw, Scott Yih |
EMNLP | 4 |
| 2013 | Linguistic Regularities in Continuous Space Word Representations
Tomás Mikolov, Scott Yih, Geoffrey Zweig |
HLT-NAACL | 2 |
| 2013 | Combining Heterogeneous Models for Measuring Relational Similarity
Alisa Zhila, Scott Yih, Christopher Meek, Geoffrey Zweig, Tomás Mikolov |
HLT-NAACL | 2 |
| 2013 | Dual Coordinate Descent Algorithms for Efficient Large Margin Structured PredictionabstractDue to the nature of complex NLP problems, structured prediction algorithms have been important modeling tools for a wide range of tasks. While there exists evidence showing that linear Structural Support Vector Machine (SSVM) algorithm performs better than structured Perceptron, the SSVM algorithm is still less frequently chosen in the NLP community because of its relatively slow training speed. In this paper, we propose a fast and easy-to-implement dual coordinate descent algorithm for SSVMs. Unlike algorithms such as Perceptron and stochastic gradient descent, our method keeps track of dual variables and updates the weight vector more aggressively. As a result, this training process is as efficient as existing online learning methods, and yet derives consistently better models, as evaluated on four benchmark NLP datasets for part-of-speech tagging, named-entity recognition and dependency parsing. Ming-Wei Chang, Scott Yih |
Trans. Assoc. Comput. Linguistics | 2 |
| 2012 | Polarity Inducing Latent Semantic Analysis
Scott Yih, Geoffrey Zweig, John C. Platt |
EMNLP-CoNLL | 1 |
| 2012 | MSR SPLAT, a language analysis toolkit
Chris Quirk, Pallavi Choudhury, Jianfeng Gao 0001, Hisami Suzuki, Kristina Toutanova, Michael Gamon, Scott Yih, Colin Cherry, Lucy Vanderwende |
HLT-NAACL | 7 |
| 2012 | Measuring Word Relatedness Using Heterogeneous Vector Space Models
Scott Yih, Vahed Qazvinian |
HLT-NAACL | 1 |
| 2011 | Learning Discriminative Projections for Text Similarity Measures
Scott Yih, Kristina Toutanova, John C. Platt, Christopher Meek |
CoNLL | 1 |
| 2011 | Domain Adaptation with Ensemble of Feature Groups
Rajhans Samdani, Scott Yih |
IJCAI | 2 |
| 2011 | Clickthrough-based latent semantic models for web searchabstractThis paper presents two new document ranking models for Web search based upon the methods of semantic representation and the statistical translation-based approach to information retrieval (IR). Assuming that a query is parallel to the titles of the documents clicked on for that query, large amounts of query-title pairs are constructed from clickthrough data; two latent semantic models are learned from this data. One is a bilingual topic model within the language modeling framework. It ranks documents for a query by the likelihood of the query being a semantics-based translation of the documents. The semantic representation is language independent and learned from query-title pairs, with the assumption that a query and its paired titles share the same distribution over semantic topics. The other is a discriminative projection model within the vector space modeling framework. Unlike Latent Semantic Analysis and its variants, the projection matrix in our model, which is used to map from term vectors into sematic space, is learned discriminatively such that the distance between a query and its paired title, both represented as vectors in the projected semantic space, is smaller than that between the query and the titles of other documents which have no clicks for that query. These models are evaluated on the Web search task using a real world data set. Results show that they significantly outperform their corresponding baseline models, which are state-of-the-art. Jianfeng Gao 0001, Kristina Toutanova, Scott Yih |
SIGIR | 3 |
| 2010 | Translingual Document Representations from Discriminative Projections
John C. Platt, Kristina Toutanova, Scott Yih |
EMNLP | 3 |
| 2010 | Adaptive near-duplicate detection via similarity learningabstractIn this paper, we present a novel near-duplicate document detection method that can easily be tuned for a particular domain. Our method represents each document as a real-valued sparse k-gram vector, where the weights are learned to optimize for a specified similarity function, such as the cosine similarity or the Jaccard coefficient. Near-duplicate documents can be reliably detected through this improved similarity measure. In addition, these vectors can be mapped to a small number of hash-values as document signatures through the locality sensitive hashing scheme for efficient similarity computation. We demonstrate our approach in two target domains: Web news articles and email messages. Our method is not only more accurate than the commonly used methods such as Shingles and I-Match, but also shows consistent improvement across the domains, which is a desired property lacked by existing methods. Hannaneh Hajishirzi, Scott Yih, Alek Kolcz |
SIGIR | 2 |
| 2009 | Learning Term-weighting Functions for Similarity Measures
Scott Yih |
EMNLP | 1 |
| 2008 | Partitioned logistic regression for spam filteringabstractNaive Bayes and logistic regression perform well in different regimes. While the former is a very simple generative model which is efficient to train and performs well empirically in many applications,the latter is a discriminative model which often achieves better accuracy and can be shown to outperform naive Bayes asymptotically. In this paper, we propose a novel hybrid model, partitioned logistic regression, which has several advantages over both naive Bayes and logistic regression. This model separates the original feature space into several disjoint feature groups. Individual models on these groups of features are learned using logistic regression and their predictions are combined using the naive Bayes principle to produce a robust final estimation. We show that our model is better both theoretically and empirically. In addition, when applying it in a practical application, email spam filtering, it improves the normalized AUC score at 10% false-positive rate by 28.8% and 23.6% compared to naive Bayes and logistic regression, when using the exact same training examples. Ming-Wei Chang, Scott Yih, Christopher Meek |
KDD | 2 |
| 2008 | The Importance of Syntactic Parsing and Inference in Semantic Role LabelingabstractWe present a general framework for semantic role labeling. The framework combines a machine-learning technique with an integer linear programming-based inference procedure, which incorporates linguistic and structural constraints into a global decision process. Within this framework, we study the role of syntactic parsing information in semantic role labeling. We show that full syntactic parsing information is, by far, most relevant in identifying the argument, especially, in the very first stage—the pruning stage. Surprisingly, the quality of the pruning stage cannot be solely determined based on its recall and precision. Instead, it depends on the characteristics of the output candidates that determine the difficulty of the downstream problems. Motivated by this observation, we propose an effective and simple approach of combining different semantic role labeling systems through joint inference, which significantly improves its performance. Our system has been evaluated in the CoNLL-2005 shared task on semantic role labeling, and achieves the highest F1 score among 19 participants. Vasin Punyakanok, Dan Roth 0001, Scott Yih |
Comput. Linguistics | 3 |
| 2007 | Improving Similarity Measures for Short Segments of Text
Scott Yih, Christopher Meek |
AAAI | 1 |
| 2007 | Multi-Document Summarization by Maximizing Informative Content-Words
Scott Yih, Joshua Goodman 0001, Lucy Vanderwende, Hisami Suzuki |
IJCAI | 1 |
| 2007 | Raising the baseline for high-precision text classifiersabstractMany important application areas of text classifiers demand high precision andit is common to compare prospective solutions to the performance of Naive Bayes. This baseline is usually easy to improve upon, but in this work we demonstrate that appropriate document representation can make out performing this classifier much more challenging. Most importantly, we provide a link between Naive Bayes and the logarithmic opinion pooling of the mixture-of-experts framework, which dictates a particular type of document length normalization. Motivated by document-specific feature selection we propose monotonic constraints on document term weighting, which is shown as an effective method of fine-tuning document representation. The discussion is supported by experiments using three large email corpora corresponding to the problem of spam detection, where high precision is of particular importance. Alek Kolcz, Scott Yih |
KDD | 2 |
| 2007 | Site-Independent Template-Block Detection
Alek Kolcz, Scott Yih |
PKDD | 2 |
| 2006 | Improved Discriminative Bilingual Word AlignmentabstractFor many years, statistical machine translation relied on generative models to provide bilingual word alignments. In 2005, several independent efforts showed that discriminative models could be used to enhance or replace the standard generative approach. Building on this work, we demonstrate substantial improvement in word-alignment accuracy, partly though improved training methods, but predominantly through selection of more and better features. Our best model produces the lowest alignment error rate yet reported on Canadian Hansards bilingual data. Robert C. Moore, Scott Yih, Andreas Bode |
ACL | 2 |
| 2006 | Automatic Semantic Role Labeling
Scott Yih, Kristina Toutanova |
HLT-NAACL | 1 |
| 2006 | Finding advertising keywords on web pagesabstractA large and growing number of web pages display contextual advertising based on keywords automatically extracted from the text of the page, and this is a substantial source of revenue supporting the web today. Despite the importance of this area, little formal, published research exists. We describe a system that learns how to extract keywords from web pages for advertisement targeting. The system uses a number of features, such as term frequency of each potential keyword, inverse document frequency, presence in meta-data, and how often the term occurs in search query logs. The system is trained with a set of example pages that have been hand-labeled with "relevant" keywords. Based on this training, it can then extract new keywords from previously unseen pages. Accuracy is substantially better than several baseline systems. Scott Yih, Joshua Goodman 0001, Vitor R. Carvalho |
WWW | 1 |
| 2005 | Generalized Inference with Multiple Semantic Role Labeling Systems
Peter Koomen, Vasin Punyakanok, Dan Roth 0001, Scott Yih |
CoNLL | 4 |
| 2005 | Integer linear programming inference for conditional random fieldsabstractInference in Conditional Random Fields and Hidden Markov Models is done using the Viterbi algorithm, an efficient dynamic programming algorithm. In many cases, general (non-local and non-sequential) constraints may exist over the output sequence, but cannot be incorporated and exploited in a natural way by this inference procedure. This paper proposes a novel inference procedure based on integer linear programming (ILP) and extends CRF models to naturally and efficiently support general constraint structures. For sequential constraints, this procedure reduces to simple linear programming as the inference process. Experimental evidence is supplied in the context of an important NLP problem, semantic role labeling. Dan Roth 0001, Scott Yih |
ICML | 2 |
| 2005 | The Necessity of Syntactic Parsing for Semantic Role Labeling
Vasin Punyakanok, Dan Roth 0001, Scott Yih |
IJCAI | 3 |
| 2005 | Learning and Inference over Constrained Output
Vasin Punyakanok, Dan Roth 0001, Scott Yih, Dav Zimak |
IJCAI | 3 |
| 2004 | Semantic Role Labeling Via Integer Linear Programming Inference
Vasin Punyakanok, Dan Roth 0001, Scott Yih, Dav Zimak |
COLING | 3 |
| 2004 | Semantic Role Labeling Via Generalized Inference Over Classifiers
Vasin Punyakanok, Dan Roth 0001, Scott Yih, Dav Zimak, Yuancheng Tu |
CoNLL | 3 |
| 2004 | A Linear Programming Formulation for Global Inference in Natural Language Tasks
Dan Roth 0001, Scott Yih |
CoNLL | 2 |
| 2004 | Mining Online Deal Forums for Hot DealsabstractOnline deal forums are public places where participants share with each other news and information regarding "deals" such as sales promotion events by online stores. The large number of messages in the forums and their inherent uncertainty make it difficult for even seasoned users to identify useful deal information from the forums. We develop an intelligent deal alert service which assists ordinary Web surfers to find useful deals by mining online deal forums. It periodically crawls relevant deal forums to collect fresh message posts and responses, and evaluate them using a form of probabilistic text classification. Users may be notified of new, "potentially" useful deal messages via emails or they may browse them using their favorite Web browser. We train and evaluate the service using deal posts and responses collected from actual deal forums in the Web. The preliminary evaluation results show that the service is quite effective in reducing the time to find useful deals. Scott Yih, Po-Hao Chang, WooYoung Kim |
Web Intelligence | 1 |
| 2002 | Probabilistic Reasoning for Entity & Relation Recognition
Dan Roth 0001, Scott Yih |
COLING | 2 |
| 2001 | Relational Learning via Propositional Algorithms: An Information Extraction Case Study
Dan Roth 0001, Scott Yih |
IJCAI | 2 |