VLDB 2026 Research / reviewers in the wild / expert
Sewon Min
dblp:203/9401
· DBLP profile ↗
31ranked-venue papers
11as first author
22since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 11 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DS SERVE: A Framework for Efficient and Scalable Neural RetrievalabstractWe present DS SERVE, a framework that transforms large-scale text datasets—comprising half a trillion tokens—into a high-performance neural retrieval system. DS SERVE offers both a web interface and API endpoints, achieving low latency with modest memory overhead on a single node. The framework also supports inference-time tradeoffs between latency, accuracy, and result diversity. We anticipate that DS SERVE will be broadly useful for a range of applications such as large-scale retrieval-augmented generation (RAG), training data attribution, training a search agent, and beyond. Jinjian Liu, Xinxi Lyu, Rulin Shao, Joseph Gonzalez 0001, Matei Zaharia, Sewon Min |
AAAI | 7 |
| 2025 | OLMoE: Open Mixture-of-Experts Language ModelsabstractWe introduce OLMoE, a fully open, state-of-the-art language model leveraging sparse Mixture-of-Experts (MoE). OLMoE-1B-7B has 7 billion (B) parameters but uses only 1B per input token. We pretrain it on 5 trillion tokens and further adapt it to create OLMoE-1B-7B-Instruct. Our models outperform all available models with similar active parameters, even surpassing larger ones like Llama2-13B-Chat and DeepSeekMoE-16B. We present novel findings on MoE training, define and analyze new routing properties showing high specialization in our model, and open-source all our work: model weights, training data, code, and logs. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Pete Walsh 0001, Oyvind Tafjord, Nathan Lambert 0001, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, Dave Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi |
ICLR | 6 |
| 2025 | Organize the Web: Constructing Domains Enhances Pre-Training Data CurationabstractModern language models are trained on large, unstructured datasets consisting of trillions of tokens and obtained by crawling the web. The unstructured nature makes it difficult to reason about their contents and develop systematic approaches to data curation. In this paper, we unpack monolithic web corpora by developing taxonomies of their contents and organizing them into domains. We introduce WebOrganizer, a framework for organizing web pages in terms of both their topic and format. Using these two complementary notions of domains, we automatically annotate pre-training data by distilling annotations from a large language model into efficient classifiers. This allows us to study how data from different domains should be mixed to improve models on downstream tasks, and we show that we can combine insights about effective topics and formats to further boost performance. We demonstrate that our domain mixing also improves existing methods that select data based on quality. Furthermore, we study and compare how quality-based methods will implicitly change the domain mixture. Overall, our work demonstrates that constructing and mixing domains provides a valuable complement to quality-based data curation methods, opening new avenues for effective and insightful pre-training data curation. Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen 0001, Luca Soldaini |
ICML | 3 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 23 |
| 2024 | CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model GenerationabstractTong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tong Chen 0005, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi 0001, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh |
EMNLP | 4 |
| 2024 | BTR: Binary Token Representations for Efficient Retrieval Augmented Language ModelsabstractRetrieval augmentation addresses many critical problems in large language models such as hallucination, staleness, and privacy leaks.
However, running retrieval-augmented language models (LMs) is slow and difficult to scale due to processing large amounts of retrieved text.
We introduce binary token representations (BTR), which use 1-bit vectors to precompute every token in passages, significantly reducing computation during inference.
Despite the potential loss of accuracy, our new calibration techniques and training objectives restore performance. Combined with offline and runtime compression, this only requires 127GB of disk space for encoding 3 billion tokens in Wikipedia.
Our experiments show that on five knowledge-intensive NLP tasks, BTR accelerates state-of-the-art inference by up to 4x and reduces storage by over 100x while maintaining over 95% task performance. Our code is publicly available at https://github.com/csarron/BTR. Sewon Min, Yizhong Wang, Hannaneh Hajishirzi |
ICLR | 2 |
| 2024 | SILO Language Models: Isolating Legal Risk In a Nonparametric DatastoreabstractThe legality of training language models (LMs) on copyrighted or otherwise restricted data is under intense debate. However, as we show, model performance significantly degrades if trained only on low-risk text (e.g., out-of-copyright books or government documents), due to its limited size and domain coverage. We present SILO, a new language model that manages this risk-performance tradeoff during inference. SILO is built by (1) training a parametric LM on the Open License Corpus (OLC), a new corpus we curate with 228B tokens of public domain and permissively licensed text and (2) augmenting it with a more general and easily modifiable nonparametric datastore (e.g., containing copyrighted books or news) that is only queried during inference. The datastore allows use of high-risk data without training on it, supports sentence-level data attribution, and enables data producers to opt out from the model by removing content from the store. These capabilities can foster compliance with data-use regulations such as the fair use doctrine in the United States and the GDPR in the European Union. Our experiments show that the parametric LM struggles on its own with domains not covered by OLC. However, access to the datastore greatly improves out of domain performance, closing 90% of the performance gap with an LM trained on the Pile, a more diverse corpus with mostly high-risk text. We also analyze which nonparametric approach works best, where the remaining errors lie, and how performance scales with datastore size. Our results suggest that it is possible to build high quality language models while mitigating legal risk. Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer |
ICLR | 1 |
| 2024 | In-Context Pretraining: Language Modeling Beyond Document BoundariesabstractLanguage models are currently trained to predict tokens given document prefixes, enabling them to zero shot long form generation and prompting-style tasks which can be reduced to document completion. We instead present IN-CONTEXT PRETRAINING, a new approach where language models are trained on a sequence of related documents, thereby explicitly encouraging them to read and reason across document boundaries. Our approach builds on the fact that current pipelines train by concatenating random sets of shorter documents to create longer context windows; this improves efficiency even though the prior documents provide no signal for predicting the next document. Given this fact, we can do IN-CONTEXT PRETRAINING by simply changing the document ordering so that each context contains related documents, and directly applying existing pretraining pipelines. However, this document sorting problem is challenging. There are billions of documents and we would like the sort to maximize contextual similarity for every document without repeating any data. To do this, we introduce approximate algorithms for finding related documents with efficient nearest neighbor search and constructing coherent batches with a graph cover algorithm. Our experiments show IN-CONTEXT PRETRAINING offers a scalable and simple approach to significantly enhance LM performance: we see notable improvements in tasks that require more complex contextual reasoning, including in-context learning (+8%), reading comprehension (+15%), faithfulness to previous contexts (+16%), long-context reasoning (+5%), and retrieval augmentation (+9%). Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Scott Yih, Mike Lewis |
ICLR | 2 |
| 2024 | REPLUG: Retrieval-Augmented Black-Box Language ModelsabstractWeijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, Wen-tau Yih. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James 0001, Mike Lewis, Luke Zettlemoyer, Scott Yih |
NAACL-HLT | 2 |
| 2024 | Scaling Retrieval-Based Language Models with a Trillion-Token DatastoreabstractScaling laws with respect to the amount of training data and the number of parameters allow us to predict the cost-benefit trade-offs of pretraining language models (LMs) in different configurations. In this paper, we consider another dimension of scaling: the amount of data available at inference time. Specifically, we find that increasing the size of the datastore used by a retrieval-based LM monotonically improves language modeling and several downstream tasks without obvious saturation, such that a smaller model augmented with a large datastore outperforms a larger LM-only model on knowledge-intensive tasks. By plotting compute-optimal scaling curves with varied datastore, model, and pretraining data sizes, we show that using larger datastores can significantly improve model performance for the same training compute budget. We carry out our study by constructing a 1.4 trillion-token datastore named MassiveDS, which is the largest and the most diverse open-sourced datastore for retrieval-based LMs to date, and designing an efficient pipeline for studying datastore scaling in an accessible manner. Finally, we analyze the effect of improving the retriever, datastore quality filtering, and other design choices on our observed scaling trends. Overall, our results show that datastore size should be considered as an integral part of LM efficiency and performance trade-offs. To facilitate future research, we open-source our datastore and code at https://github.com/RulinShao/retrieval-scaling. Rulin Shao, Jacqueline He, Akari Asai, Tim Dettmers, Sewon Min, Luke Zettlemoyer, Pang Wei Koh |
NeurIPS | 6 |
| 2023 | CREPE: Open-Domain Question Answering with False PresuppositionsabstractWhen asking about unfamiliar topics, information seeking users often pose questions with false presuppositions.Most existing question answering (QA) datasets, in contrast, assume all questions have well defined answers.We introduce CREPE, a QA dataset containing a natural distribution of presupposition failures from online information-seeking forums.We find that 25% of questions contain false presuppositions, and provide annotations for these presuppositions and their corrections.Through extensive baseline experiments, we show that adaptations of existing open-domain QA models can find presuppositions moderately well, but struggle when predicting whether a presupposition is factually correct.This is in large part due to difficulty in retrieving relevant evidence passages from a large text corpus.CREPE provides a benchmark to study question answering in the wild, and our analyses provide avenues for future work in better modeling and further studying the task. 1Question: If there's an equal and opposite reaction for everything, how does any action happen?Isn't it balanced out by the opposite reaction?False presupposition: The equal and opposite reaction applies to the same object.Correction: Based on Newton's Law of Motion, the equal and opposite reaction applies to the other object.Only forces that are applied to the same object would be cancelled out. Newton's laws of motionFrom Wikipedia, the free encyclopedia Inputs given to the human raters Question: Why do prosecuters/courts seek/sentence prison time greater than the expected lifespan of the offender (i.e. 150 years in prison)?Why not simply sentence those criminals to 'life' in prison instead?Comment: Sentencing options are written into state laws.Life in prison is different in state laws than 150 years.Some of it comes into play with the "cruel and unusual punishment" clause in the Constitution too.Life in prison may not be "cruel and unusual" for a murder sentence, but it might be for, say, child sex trafficking.But if you trafficked 10 kids and the sentence is 15 years for each one, you get an effective life sentence that will also stand up, Constitutionally, against a "cruel and unusual punishment" defense. Outputs human raters rateReference Presupposition: It does not make sense to sentence a person to 150 years in prison if they can't live that long anyways, prosecutors should use the life in prison sentence instead.Correction: The defendant can argue the life in prison sentence as cruel and unusual, so the actual year sentence is better to give than the alternative.GOLD-COMMENT track, Dedicated Presupposition: Penalties should be able to be sentenced to life in prison.Correction: Life in prison is different in state laws than 150 years in prison.GOLD-COMMENT track, Unified Presupposition: If a criminal is sentenced to life in prison, they should be sentenced to life in prison.Correction: It is not the case that if a criminal is sentenced to life in prison, they should be sentenced to life in prison.Main, Dedicated Presupposition: Penalties should be able to be imposed on criminals for life.Correction: The longer the sentence, the more likely the prosecution will seek to sentence the offender to life in prison.Main, Unified Presupposition: Prosecutor's should seek prison time greater than the expected lifespan of the offender.Correction: It is not the case that prosecutor's should seek prison time greater than the expected lifespan of the offender.Table 13: An example of the input and the output human raters are given for the human evaluation of the writing subtask.Note that human raters are not given which output is a reference or from which system. Inputs given to the human ratersQuestion: Why did scientists in the 1970s think that there was going to be a new ice age soon?Comment: They didn't.Between 1965 and 1979, there was 7 papers talking about global cooling (not ice age and not necessarily soon).During the same period there was 44 papers about global warming.The media just liked the sensationalism, so there was some news article and a front page on the Times Magazine.They started with a minority of scientist talking about global cooling in a time period when there was still a lot of unknown in climate science and changed that to Scientific consensus that an Ice Age is coming soon.The 7 papers were the following : McComick and Ludwig 1967, Barrett 1971, Rasool and Xinyan Yu 0001, Sewon Min, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 2 |
| 2023 | Z-ICL: Zero-Shot In-Context Learning with Pseudo-DemonstrationsabstractAlthough large language models can be prompted for both zero-and few-shot learning, performance drops significantly when no demonstrations are available.In this paper, we introduce Z-ICL, a new zero-shot method that closes the gap by constructing pseudo-demonstrations for a given test input using a raw text corpus.Concretely, pseudodemonstrations are constructed by (1) finding the nearest neighbors to the test input from the corpus and pairing them with random task labels, and (2) applying a set of techniques to reduce the amount of direct copying the model does from the resulting demonstrations.Evaluation on nine classification datasets shows that Z-ICL outperforms previous zero-shot methods by a significant margin, and is on par with incontext learning with few-shot labeled training data.Overall, Z-ICL provides a significantly higher estimate of the zero-shot performance levels of a model, and supports future efforts to develop better pseudo-demonstrations that further improve zero-shot results. 1 Xinxi Lyu, Sewon Min, Iz Beltagy, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 2 |
| 2023 | Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What MattersabstractBoshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, Huan Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Boshi Wang, Sewon Min, Xiang Deng 0001, You Wu 0001, Luke Zettlemoyer, Huan Sun 0001 |
ACL (1) | 2 |
| 2023 | FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationabstractSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Scott Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 1 |
| 2023 | InSCIt: Information-Seeking Conversations with Mixed-Initiative InteractionsabstractAbstract In an information-seeking conversation, a user may ask questions that are under-specified or unanswerable. An ideal agent would interact by initiating different response types according to the available knowledge sources. However, most current studies either fail to or artificially incorporate such agent-side initiative. This work presents InSCIt, a dataset for Information-Seeking Conversations with mixed-initiative Interactions. It contains 4.7K user-agent turns from 805 human-human conversations where the agent searches over Wikipedia and either directly answers, asks for clarification, or provides relevant information to address user queries. The data supports two subtasks, evidence passage identification and response generation, as well as a human evaluation protocol to assess model performance. We report results of two systems based on state-of-the-art models of conversational knowledge identification and open-domain question answering. Both systems significantly underperform humans, suggesting ample room for improvement in future studies.1 Zeqiu Wu, Ryu Parish, Hao Cheng 0002, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, Hannaneh Hajishirzi |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | Noisy Channel Language Model Prompting for Few-Shot Text ClassificationabstractWe introduce a noisy channel approach for language model prompting in few-shot text classification.Instead of computing the likelihood of the label given the input (referred as direct models), channel models compute the conditional probability of the input given the label, and are thereby required to explain every word in the input.We use channel models for recently proposed few-shot learning methods with no or very limited updates to the language model parameters, via either in-context demonstration or prompt tuning.Our experiments show that, for both methods, channel models significantly outperform their direct counterparts, which we attribute to their stability, i.e., lower variance and higher worstcase accuracy.We also present extensive ablations that provide recommendations for when to use channel prompt tuning instead of other competitive methods (e.g., direct head tuning): channel prompt tuning is preferred when the number of training examples is small, labels in the training data are imbalanced, or generalization to unseen labels is required. Sewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
ACL (1) | 1 |
| 2022 | FaVIQ: FAct Verification from Information-seeking QuestionsabstractDespite significant interest in developing general purpose fact checking models, it is challenging to construct a large-scale fact verification dataset with realistic real-world claims.Existing claims are either authored by crowdworkers, thereby introducing subtle biases that are difficult to control for, or manually verified by professional fact checkers, causing them to be expensive and limited in scale.In this paper, we construct a large-scale challenging fact verification dataset called FAVIQ, consisting of 188k claims derived from an existing corpus of ambiguous information-seeking questions.The ambiguities in the questions enable automatically constructing true and false claims that reflect user confusions (e.g., the year of the movie being filmed vs. being released).Claims in FAVIQ are verified to be natural, contain little lexical bias, and require a complete understanding of the evidence for verification.Our experiments show that the stateof-the-art models are far from solving our new task.Moreover, training on our data helps in professional fact-checking, outperforming models trained on the widely used dataset FEVER or in-domain data by up to 17% absolute.Altogether, our data will serve as a challenging benchmark for natural language understanding and support future progress in professional fact checking.1 Jungsoo Park, Sewon Min, Jaewoo Kang, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 2 |
| 2022 | Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?abstractLarge language models (LMs) are able to incontext learn-perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs.However, there has been little understanding of how the model learns and which aspects of the demonstrations contribute to end task performance.In this paper, we show that ground truth demonstrations are in fact not required-randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks, consistently over 12 different models including GPT-3.Instead, we find that other aspects of the demonstrations are the key drivers of end task performance, including the fact that they provide a few examples of (1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence.Together, our analysis provides a new way of understanding how and why in-context learning works, while opening up new questions about how much can be learned from large language models through inference alone. Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP | 1 |
| 2022 | Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous PromptsabstractDaniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson 0001, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh 0001, Yejin Choi 0001 |
NAACL-HLT | 3 |
| 2022 | MetaICL: Learning to Learn In ContextabstractSewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Sewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi |
NAACL-HLT | 1 |
| 2021 | Joint Passage Ranking for Diverse Multi-Answer RetrievalabstractWe study multi-answer retrieval, an underexplored problem that requires retrieving passages to cover multiple distinct answers for a given question.This task requires joint modeling of retrieved passages, as models should not repeatedly retrieve passages containing the same answer at the cost of missing a different valid answer.In this paper, we introduce JPR, the first joint passage retrieval model for multi-answer retrieval.JPR makes use of an autoregressive reranker that selects a sequence of passages, each conditioned on previously selected passages.JPR is trained to select passages that cover new answers at each timestep and uses a tree-decoding algorithm to enable flexibility in the degree of diversity.Compared to prior approaches, JPR achieves significantly better answer coverage on three multianswer datasets.When combined with downstream question answering, the improved retrieval enables larger answer generation models since they need to consider fewer passages, establishing a new state-of-the-art. Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, Hannaneh Hajishirzi |
EMNLP (1) | 1 |
| 2021 | RECONSIDER: Improved Re-Ranking using Span-Focused Cross-Attention for Open Domain Question AnsweringabstractSrinivasan Iyer, Sewon Min, Yashar Mehdad, Wen-tau Yih. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Srinivasan Iyer 0001, Sewon Min, Yashar Mehdad, Scott Yih |
NAACL-HLT | 2 |
| 2020 | Dense Passage Retrieval for Open-Domain Question AnsweringabstractVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen 0001, Scott Yih |
EMNLP (1) | 3 |
| 2020 | Efficient One-Pass End-to-End Entity Linking for QuestionsabstractWe present ELQ, a fast end-to-end entity linking model for questions, which uses a biencoder to jointly perform mention detection and linking in one pass.Evaluated on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question, ELQ outperforms the previous state of the art by a large margin of +12.7% and +19.6% F1, respectively.With a very fast inference time (1.57examples/s on a single CPU), ELQ can be useful for downstream question answering systems.In a proof-of-concept experiment, we demonstrate that using ELQ significantly improves the downstream QA performance of GraphRetriever (Min et al., 2019). 1 Belinda Z. Li, Sewon Min, Srinivasan Iyer 0001, Yashar Mehdad, Scott Yih |
EMNLP (1) | 2 |
| 2020 | AmbigQA: Answering Ambiguous Open-domain QuestionsabstractAmbiguity is inherent to open-domain question answering; especially when exploring new topics, it can be difficult to ask questions that have a single, unambiguous answer.In this paper, we introduce AMBIGQA, a new open-domain question answering task which involves finding every plausible answer, and then rewriting the question for each one to resolve the ambiguity.To study this task, we construct AMBIGNQ, a dataset covering 14,042 questions from NQ-OPEN, an existing opendomain QA benchmark.We find that over half of the questions in NQ-OPEN are ambiguous, with diverse sources of ambiguity such as event and entity references.We also present strong baseline models for AMBIGQA which we show benefit from weakly supervised learning that incorporates NQ-OPEN, strongly suggesting our new task and data will support significant future research effort.Our data and baselines are available at https://nlp.cs. washington.edu/ambigqa.Type Example Event references (39%) What season does meredith and derek get married in grey's anatomy?Q: In what season do Meredith and Derek get informally married in Grey's Anatomy? / A: Season 5 Q: In what season do Meredith and Derek get legally married in Grey's Anatomy? / A: Season 7 Properties (27%) How many episode in seven deadly sins season 2? Q: How many episodes were there in seven deadly sins season 2, not including the OVA episode?/ A: 25 Q: How many episodes were there in seven deadly sins season 2, including the OVA episode?/ A: 26 Entity references (23%) How many sacks does clay matthews have in his career?Q: How many sacks does Clay Matthews Jr. have in his career?/ A: 69.5 Q: How many sacks does Clay Matthews III have in his career?/ A: 91.5 Answer types (16%) Who sings the song what a beautiful name it is?Q: Which group sings the song what a beautiful name it is?/ A: Hillsong Live Q: Who is the lead singer of the song what a beautiful name it is?/ A: Brooke Ligertwood Sewon Min, Julian Michael, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP (1) | 1 |
| 2019 | Compositional Questions Do Not Necessitate Multi-hop ReasoningabstractMulti-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs.We argue that it can be difficult to construct large multi-hop RC datasets.For example, even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant.Our analysis is centered on HOTPOTQA, where we show that single-hop reasoning can solve much more of the dataset than previously thought.We introduce a single-hop BERT-based RC model that achieves 67 F1-comparable to state-of-theart multi-hop models.We also design an evaluation setting where humans are not shown all of the necessary paragraphs for the intended multi-hop reasoning but can still answer over 80% of questions.Together with detailed error analysis, these results suggest there should be an increasing focus on the role of evidence in multi-hop reasoning and possibly even a shift towards information retrieval style evaluations with large and diverse evidence collections. Sewon Min, Eric Wallace, Sameer Singh 0001, Matt Gardner 0001, Hannaneh Hajishirzi, Luke Zettlemoyer |
ACL (1) | 1 |
| 2019 | Multi-hop Reading Comprehension through Question Decomposition and RescoringabstractMulti-hop Reading Comprehension (RC) requires reasoning and aggregation across several paragraphs.We propose a system for multi-hop RC that decomposes a compositional question into simpler sub-questions that can be answered by off-the-shelf single-hop RC models.Since annotations for such decomposition are expensive, we recast subquestion generation as a span prediction problem and show that our method, trained using only 400 labeled examples, generates sub-questions that are as effective as humanauthored sub-questions.We also introduce a new global rescoring approach that considers each decomposition (i.e. the sub-questions and their answers) to select the best final answer, greatly improving overall performance.Our experiments on HOTPOTQA show that this approach achieves the state-of-the-art results, while providing explainable evidence for its decision making in the form of sub-questions. Sewon Min, Victor Zhong, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 1 |
| 2019 | A Discrete Hard EM Approach for Weakly Supervised Question AnsweringabstractSewon Min, Danqi Chen, Hannaneh Hajishirzi, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sewon Min, Danqi Chen 0001, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Efficient and Robust Question Answering from Minimal Context over DocumentsabstractNeural models for question answering (QA) over documents have achieved significant performance improvements.Although effective, these models do not scale to large corpora due to their complex modeling of interactions between the document and the question.Moreover, recent work has shown that such models are sensitive to adversarial inputs.In this paper, we study the minimal context required to answer the question, and find that most questions in existing datasets can be answered with a small set of sentences.Inspired by this observation, we propose a simple sentence selector to select the minimal set of sentences to feed into the QA model.Our overall system achieves significant reductions in training (up to 15 times) and inference times (up to 13 times), with accuracy comparable to or better than the state-of-the-art on SQuAD, NewsQA, TriviaQA and SQuAD-Open.Furthermore, our experimental results and analyses show that our approach is more robust to adversarial inputs. Sewon Min, Victor Zhong, Richard Socher, Caiming Xiong |
ACL (1) | 1 |
| 2018 | Neural Speed Reading via Skim-RNN
Minjoon Seo, Sewon Min, Ali Farhadi, Hannaneh Hajishirzi |
ICLR (Poster) | 2 |
| 2017 | Query-Reduction Networks for Question Answering
Minjoon Seo, Sewon Min, Ali Farhadi, Hannaneh Hajishirzi |
ICLR (Poster) | 2 |