EDBT 2026 Demo / reviewers in the wild / expert
Hannaneh Hajishirzi
dblp:52/1296 · also Hanna Hajishirzi
· DBLP profile ↗
144ranked-venue papers
9as first author
85since 2021 · last 2025
0000-0002-1055-6657ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 140 · 8 first-author · 83 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackabstractLester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lester James V. Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar 0009, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi |
ACL (1) | 8 |
| 2025 | Steering off Course: Reliability Challenges in Steering Language ModelsabstractPatrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal, Hannaneh Hajishirzi, Sachin Kumar. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Patrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal, Hannaneh Hajishirzi, Sachin Kumar 0009 |
ACL (1) | 4 |
| 2025 | Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsabstractToday’s most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well- tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog. Matt Deitke, Sangho Lee 0008, Rohun Tripathi, Yue Yang 0006, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, Yen-Sung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert 0001, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz 0002, Aaron Sarnat, Byron Bischoff, Pete Walsh 0001, Chris Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hannaneh Hajishirzi, Ross B. Girshick, Ali Farhadi, Aniruddha Kembhavi |
CVPR | 47 |
| 2025 | s1: Simple test-time scalingabstractNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candes, Tatsunori Hashimoto. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Niklas Muennighoff, Zitong Yang, Xiang Li 0063, Li Fei-Fei 0001, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel J. Candès, Tatsunori B. Hashimoto |
EMNLP | 6 |
| 2025 | SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureabstractDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Zejiang Shen, Doug Downey, Hannaneh Hajishirzi, Arman Cohan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Dave Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh 0001, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Shen 0001, Doug Downey, Hannaneh Hajishirzi, Arman Cohan |
EMNLP | 13 |
| 2025 | Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-IndexabstractLanguage models are trained mainly on massive text data from the Internet, and it becomes increasingly important to understand this data source. Exact-match search engines enable searching in large text corpora – counting string appearances and retrieving the enclosing documents – yet the high storage overhead hinders their application on Internet-scale data. We present Infini-gram mini, an efficient and scalable system that can make petabyte-level text corpora searchable. Based on the FM-index data structure (Ferragina and Manzini, 2000), which simultaneously indexes and compresses text, our system creates indexes with size only 44% of the corpus. Infini-gram mini greatly improves upon the best existing implementation of FM-index in terms of indexing speed (18\times) and memory use during both indexing (3.2\times reduction) and querying (down to a negligible amount). We index 83TB of Internet text in 99 days with a single 128-core CPU node (or 19 hours if using 137 such nodes). We show one important use case of Infini-gram mini in a large-scale analysis of benchmark contamination. We find several core LM evaluation benchmarks to be heavily contaminated in Internet crawls (up to 74.2% in GSM8K), which could lead to overestimating the capabilities of language models if trained on such data. We host a benchmark contamination bulletin to share the contamination rate of many core and community-contributed benchmarks. We also release a web interface and an API endpoint to serve general search queries on Infini-gram mini indexes. Jiacheng Liu 0010, Yejin Choi 0001, Noah A. Smith, Hannaneh Hajishirzi |
EMNLP | 5 |
| 2025 | Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice QuestionsabstractMultiple-choice question answering (MCQA) is a key competence of performant transformer language models that is tested by mainstream benchmarks. However, recent evidence shows that models can have quite a range of performance, particularly when the task format is diversified slightly (such as by shuffling answer choice order). In this work we ask: how do successful models perform formatted MCQA? We employ vocabulary projection and activation patching methods to localize key hidden states that encode relevant information for predicting the correct answer. We find that prediction of a specific answer symbol is causally attributed to a few middle layers, and specifically their multi-head self-attention mechanisms. We show that subsequent layers increase the probability of the predicted answer symbol in vocabulary space, and that this probability increase is associated with a sparse set of attention heads with unique roles. We additionally uncover differences in how different models adjust to alternative symbols. Finally, we demonstrate that a synthetic task can disentangle sources of model error to pinpoint when a model has learned formatted MCQA, and show that logit differences between answer choice tokens continue to grow over the course of training. Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi, Ashish Sabharwal |
ICLR | 4 |
| 2025 | Organize the Web: Constructing Domains Enhances Pre-Training Data CurationabstractModern language models are trained on large, unstructured datasets consisting of trillions of tokens and obtained by crawling the web. The unstructured nature makes it difficult to reason about their contents and develop systematic approaches to data curation. In this paper, we unpack monolithic web corpora by developing taxonomies of their contents and organizing them into domains. We introduce WebOrganizer, a framework for organizing web pages in terms of both their topic and format. Using these two complementary notions of domains, we automatically annotate pre-training data by distilling annotations from a large language model into efficient classifiers. This allows us to study how data from different domains should be mixed to improve models on downstream tasks, and we show that we can combine insights about effective topics and formats to further boost performance. We demonstrate that our domain mixing also improves existing methods that select data based on quality. Furthermore, we study and compare how quality-based methods will implicitly change the domain mixture. Overall, our work demonstrates that constructing and mixing domains provides a valuable complement to quality-based data curation methods, opening new avenues for effective and insightful pre-training data curation. Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen 0001, Luca Soldaini |
ICML | 4 |
| 2025 | A Systematic Examination of Preference Learning through the Lens of Instruction-FollowingabstractJoongwon Kim, Anirudh Goyal, Aston Zhang, Bo Xiong, Rui Hou, Melanie Kambadur, Dhruv Mahajan, Hannaneh Hajishirzi, Liang Tan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Joongwon Kim, Anirudh Goyal, Aston Zhang, Melanie Kambadur, Dhruv Mahajan 0001, Hannaneh Hajishirzi |
NAACL (Long Papers) | 8 |
| 2025 | ComPO: Community Preferences for Language Model PersonalizationabstractSachin Kumar, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sachin Kumar 0009, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi |
NAACL (Long Papers) | 5 |
| 2025 | Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationabstractDeveloping large language models is expensive and often involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more reliable and useful for such decisions, and interventions to design higher-quality evaluation benchmarks. We introduce two key metrics that show differences in current benchmarks: signal, a benchmark’s ability to separate better models from worse models, and noise, a benchmark’s sensitivity to random variability between training steps. We demonstrate that benchmarks with a better signal-to-noise ratio are more reliable when making decisions at small scale, and those with less noise have lower scaling law prediction error. These results suggest that improving signal or noise will lead to more useful benchmarks, so we introduce four interventions designed to directly affect signal or noise. For example, we propose that switching to a metric that has better signal and noise (e.g., perplexity rather than accuracy) leads to better reliability and scaling law error. We also find that filtering noisy benchmarks such that they have better signal-to-noise ratio leads to more reliable evaluations. We also find that averaging the output of a model's checkpoints to reduce noise leads to consistent improvements. We conclude by recommending that those creating new benchmarks, or selecting which existing benchmarks to use, aim for high signal and low noise. We use 30 benchmarks for these experiments, and 465 open-weight language models from 60M to 32B parameters, resulting in a new, publicly available dataset of 50K evaluation benchmark results, totaling 200M instances. David Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu, Noah A. Smith, Hannaneh Hajishirzi, Kyle Lo, Jesse Dodge |
NeurIPS | 6 |
| 2025 | Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model TrainingabstractThe right batch size is important when training language models at scale: a large batch size is necessary for fast training, but a batch size that is *too large* will harm token efficiency. To navigate this tradeoff, McCandlish et al. (2018) suggest that a *critical batch size* (CBS), below which training will not substantially degrade loss, can be estimated based on the gradient noise scale during training. While their method has been adopted in practice, e.g., when training GPT-3, strong assumptions are required to justify gradient noise as a proxy for the CBS, which makes it unclear whether their approach should be trusted in practice, limiting its applicability. In this paper, we introduce a simple, empirical approach to *directly* measure the CBS and show how the CBS evolves over training. Applying our approach to the OLMo models, we find that CBS is near 0 at initialization, increases rapidly at first, and then plateaus as training progresses. Furthermore, we find that this trend holds across different model sizes (1B and 7B), suggesting CBS from small training runs can inform larger-scale training runs. Our findings about how the CBS changes over training motivate *batch size warmup* as a natural way to reliably train language models at large batch size: start the batch size small and increase it as the CBS grows. To validate this claim, we use batch size warmup to train OLMo 1B to slightly better loss than the original training run with 43% fewer gradient steps. This shows how our framework can be applied to reliably train language models at larger batch sizes, increasing data parallelism without compromising performance. William Merrill, Shane Arora, Dirk Groeneveld, Hannaneh Hajishirzi |
NeurIPS | 4 |
| 2025 | Generalizing Verifiable Instruction FollowingabstractA crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like only answer with yes or no" ormention the word `abracadabra' at least 3 times" that the user adds to craft a more useful answer.Even today's strongest models struggle with fulfilling such constraints. We find that most models strongly overfit on a small set of verifiable constraints from the benchmarks that test these abilities, a skill called precise instruction following, and are not able to generalize well to unseen output constraints. We introduce a new benchmark, IFBench, to evaluate precise instruction following generalization on 58 new, diverse, and challenging verifiable out-of-domain constraints. In addition, we perform an extensive analysis of how and on what data models can be trained to improve precise instruction following generalization. Specifically, we carefully design constraint verification modules and show that reinforcement learning with verifiable rewards (RLVR) significantly improves instruction following. In addition to IFBench, we release 29 additional new hand-annotated training constraints and verification functions, RLVR training prompts, and code. Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert 0001, Hannaneh Hajishirzi |
NeurIPS | 8 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 21 |
| 2025 | OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative GeneralizationabstractRecent large language models (LLMs) with long-chain-of-thought reasoning—such as DeepSeek-R1—have achieved impressive results on Olympiad-level mathematics benchmarks. However, they often rely on a narrow set of strategies and struggle with problems that require a novel way of thinking. To systematically investigate these limitations, we introduce OMEGA—Out-of-distribution Math Problems Evaluation with 3 Generalization Axes—a controlled yet diverse bench- mark designed to evaluate three axes of out-of-distribution generalization, inspired by Boden’s typology of creativity: (1) Exploratory—applying known problem- solving skills to more complex instances within the same problem domain; (2) Com- positional—combining distinct reasoning skills, previously learned in isolation, to solve novel problems that require integrating these skills in new and coherent ways; and (3) Transformative—adopting novel, often unconventional strategies by moving beyond familiar approaches to solve problems more effectively. OMEGA consists of programmatically generated training–test pairs derived from templated problem generators across geometry, number theory, algebra, combinatorics, logic, and puzzles, with solutions verified using symbolic, numerical, or graphical methods. We evaluate frontier (or top-tier) LLMs and observe sharp performance degradation as problem complexity increases. Moreover, we fine-tune the Qwen-series models across all generalization settings and observe notable improvements in exploratory generalization, while compositional generalization remains limited, and transformative reasoning shows little to no improvement. By isolating and quantifying these fine-grained failures, OMEGA lays the groundwork for advancing LLMs toward genuine mathematical creativity beyond mechanical proficiency. Our code and dataset are available at https://github.com/sunblaze-ucb/omega. Yiyou Sun, Shawn Hu, Georgia Zhou, Ken Zheng, Hannaneh Hajishirzi, Nouha Dziri, Dawn Song |
NeurIPS | 5 |
| 2025 | SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded TasksabstractWe present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods. Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan |
NeurIPS | 12 |
| 2025 | A Large-Scale Study of Reranker Relevance Feedback at InferenceabstractNeural IR systems often employ a retrieve-and-rerank framework: a bi-encoder retrieves a fixed number of candidates (e.g., 𝐾=100), which a cross-encoder then reranks.Recent studies have indicated that relevance feedback from the reranker at inference time can improve the recall of the retriever.The approach works by updating the retriever's query representations via a distillation process that aligns it with the reranker's predictions.While a powerful idea, the arguably narrow scope of past studies focusing on a small number of specific domains such as english question answering and entity retrieval has left a gap in our understanding of how well it generalizes.In this paper, we study inference-time reranker relevance feedback extensively across multiple retrieval domains, languages, and modalities, while also investigating aspects such as the performance and latency implications of the number of distillation updates and feedback candidates. Revanth Gangi Reddy, Pradeep Dasigi, Md. Arafat Sultan, Arman Cohan, Avirup Sil, Heng Ji 0001, Hannaneh Hajishirzi |
SIGIR | 7 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 43 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 32 |
| 2024 | Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, Yezhou Yang |
ECCV (22) | 8 |
| 2024 | CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model GenerationabstractTong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tong Chen 0005, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi 0001, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh |
EMNLP | 7 |
| 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionabstractDespite their remarkable capabilities, large language models (LLMs) often produce responses containing factual inaccuracies due to their sole reliance on the parametric knowledge they encapsulate. Retrieval-Augmented Generation (RAG), an ad hoc approach that augments LMs with retrieval of relevant knowledge, decreases such issues. However, indiscriminately retrieving and incorporating a fixed number of retrieved passages, regardless of whether retrieval is necessary, or passages are relevant, diminishes LM versatility or can lead to unhelpful response generation. We introduce a new framework called **Self-Reflective Retrieval-Augmented Generation (Self-RAG)** that enhances an LM's quality and factuality through retrieval and self-reflection.
Our framework trains a single arbitrary LM that adaptively retrieves passages on-demand, and generates and reflects on retrieved passages and its generations using special tokens, called {\it reflection} tokens. Generating reflection tokens makes the LM controllable during the inference phase, enabling it to tailor its behavior to diverse task requirements.
Experiments show that Self-RAG (7B and 13B parameters) significantly outperforms state-of-the-art LLMs and retrieval-augmented models on a diverse set of tasks.
Specifically, Self-RAG outperforms ChatGPT and retrieval-augmented Llama2-chat on Open-domain QA, reasoning, and fact verification tasks, and it shows significant gains in improving factuality and citation accuracy for long-form generations relative to these models. Our code and trained models are available at https://selfrag.github.io/ Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, Hannaneh Hajishirzi |
ICLR | 5 |
| 2024 | BTR: Binary Token Representations for Efficient Retrieval Augmented Language ModelsabstractRetrieval augmentation addresses many critical problems in large language models such as hallucination, staleness, and privacy leaks.
However, running retrieval-augmented language models (LMs) is slow and difficult to scale due to processing large amounts of retrieved text.
We introduce binary token representations (BTR), which use 1-bit vectors to precompute every token in passages, significantly reducing computation during inference.
Despite the potential loss of accuracy, our new calibration techniques and training objectives restore performance. Combined with offline and runtime compression, this only requires 127GB of disk space for encoding 3 billion tokens in Wikipedia.
Our experiments show that on five knowledge-intensive NLP tasks, BTR accelerates state-of-the-art inference by up to 4x and reduces storage by over 100x while maintaining over 95% task performance. Our code is publicly available at https://github.com/csarron/BTR. Sewon Min, Yizhong Wang, Hannaneh Hajishirzi |
ICLR | 4 |
| 2024 | What's In My Big Data?abstractLarge text corpora are the backbone of language models.
However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination).
In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora. WIMBD builds on two basic capabilities---count and search---*at scale*, which allows us to analyze more than 35 terabytes on a standard compute node.
We apply WIMBD to ten different corpora used to train popular language models, including *C4*, *The Pile*, and *RedPajama*.
Our analysis uncovers several surprising and previously undocumented findings about these corpora, including the high prevalence of duplicate, synthetic, and low-quality content, personally identifiable information, toxic language, and benchmark contamination.
For instance, we find that about 50% of the documents in *RedPajama* and *LAION-2B-en* are duplicates. In addition, several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE.
We open-source WIMBD's code and artifacts to provide a standard set of evaluations for new text-based corpora and to encourage more analyses and transparency around them. Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh 0001, Dirk Groeneveld, Luca Soldaini, Sameer Singh 0001, Hannaneh Hajishirzi, Noah A. Smith, Jesse Dodge |
ICLR | 11 |
| 2024 | MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsabstractLarge Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/. Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 0010, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng 0002, Kai-Wei Chang 0001, Michel Galley, Jianfeng Gao 0001 |
ICLR | 6 |
| 2024 | SILO Language Models: Isolating Legal Risk In a Nonparametric DatastoreabstractThe legality of training language models (LMs) on copyrighted or otherwise restricted data is under intense debate. However, as we show, model performance significantly degrades if trained only on low-risk text (e.g., out-of-copyright books or government documents), due to its limited size and domain coverage. We present SILO, a new language model that manages this risk-performance tradeoff during inference. SILO is built by (1) training a parametric LM on the Open License Corpus (OLC), a new corpus we curate with 228B tokens of public domain and permissively licensed text and (2) augmenting it with a more general and easily modifiable nonparametric datastore (e.g., containing copyrighted books or news) that is only queried during inference. The datastore allows use of high-risk data without training on it, supports sentence-level data attribution, and enables data producers to opt out from the model by removing content from the store. These capabilities can foster compliance with data-use regulations such as the fair use doctrine in the United States and the GDPR in the European Union. Our experiments show that the parametric LM struggles on its own with domains not covered by OLC. However, access to the datastore greatly improves out of domain performance, closing 90% of the performance gap with an LM trained on the Pile, a more diverse corpus with mostly high-risk text. We also analyze which nonparametric approach works best, where the remaining errors lie, and how performance scales with datastore size. Our results suggest that it is possible to build high quality language models while mitigating legal risk. Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer |
ICLR | 5 |
| 2024 | Data Engineering for Scaling Language Models to 128K ContextabstractWe study continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in particular *the ability to utilize information at arbitrary input locations*, is a capability that is mostly already acquired through large-scale pretraining, and that this capability can be readily extended to contexts substantially longer than seen during training (e.g., 4K to 128K) through lightweight continual pretraining on appropriate data mixture. We investigate the *quantity* and *quality* of the data for continual pretraining: (1) for quantity, we show that 500 million to 5 billion tokens are enough to enable the model to retrieve information anywhere within the 128K context; (2) for quality, our results equally emphasize *domain balance* and *length upsampling*. Concretely, naïvely upsampling longer data on certain domains like books, a common practice of existing work, gives suboptimal performance; a balanced domain mixture is equally important. We demonstrate that continual pretraining of the full model on 1B-5B tokens of such data is an effective and affordable strategy for scaling the context length of language models to 128K. Our recipe outperforms strong open-source long-context models and closes the gap to frontier models like GPT-4 128K. Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Hao Peng 0018 |
ICML | 5 |
| 2024 | APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and InferenceabstractFine-tuning and inference with large Language Models (LM) are generally known to be expensive. Parameter-efficient fine-tuning over pretrained LMs reduces training memory by updating a small number of LM parameters but does not improve inference efficiency. Structured pruning improves LM inference efficiency by removing consistent parameter blocks, yet often increases training memory and time. To improve both training and inference efficiency, we introduce APT that adaptively prunes and tunes parameters for the LMs. At the early stage of fine-tuning, APT dynamically adds salient tuning parameters for fast and accurate convergence while discarding unimportant parameters for efficiency. Compared to baselines, our experiments show that APT maintains up to 98% task performance when pruning RoBERTa and T5 models with 40% parameters left while keeping 86.4% LLaMA models’ performance with 70% parameters remaining. Furthermore, APT speeds up LMs’ fine-tuning by up to 8$\times$ and reduces large LMs’ memory training footprint by up to 70%. Our code and models are publicly available at https://github.com/ROIM1998/APT. Bowen Zhao 0004, Hannaneh Hajishirzi |
ICML | 2 |
| 2024 | BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferabstractAkari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Akari Asai, Sneha Reddy Kudugunta, Xinyan Yu 0001, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi |
NAACL-HLT | 9 |
| 2024 | The Art of Saying No: Contextual Noncompliance in Language ModelsabstractChat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of ``unsafe'' queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30\% of requests.To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities. Faeze Brahman, Sachin Kumar 0009, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 14 |
| 2024 | MatFormer: Nested Transformer for Elastic InferenceabstractFoundation models are applied in a broad spectrum of settings with different inference constraints, from massive multi-accelerator clusters to resource-constrained standalone mobile devices. However, the substantial costs associated with training these models often limit the number of unique model sizes that can be offered. Consequently, practitioners are compelled to select a model that may not be optimally aligned with their specific latency and cost requirements. We present MatFormer, a novel Transformer architecture designed to provide elastic inference across diverse deployment constraints. MatFormer achieves this by incorporating a nested Feed Forward Network (FFN) block structure within a standard Transformer model. During training, we optimize the parameters of multiple nested FFN blocks with varying sizes, enabling the extraction of hundreds of accurate smaller models without incurring additional computational costs. We empirically validate the efficacy of MatFormer across different model classes (decoders and encoders) and modalities (language and vision), demonstrating its potential for real-world deployment. We show that a 850M decoder-only MatFormer language model (MatLM) allows us to extract multiple smaller models spanning from 582M to 850M parameters, each exhibiting better validation loss and one-shot downstream evaluations than independently trained counterparts. Furthermore, we observe that smaller encoders extracted from a universal MatFormer-based ViT (MatViT) encoder preserve the metric-space structure for adaptive large-scale retrieval. Finally, we showcase that speculative decoding with the accurate and consistent submodels extracted from MatFormer can lead to significant reduction in inference latency. Devvrit, Sneha Reddy Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit S. Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham M. Kakade, Ali Farhadi, Prateek Jain 0002 |
NeurIPS | 8 |
| 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackabstractLearning from preference feedback has emerged as an essential step for improving the generation quality and performance of modern language models (LMs). Despite its widespread use, the way preference-based learning is applied varies wildly, with differing data, learning algorithms, and evaluations used, making disentangling the impact of each aspect difficult. In this work, we identify four core aspects of preference-based learning: preference data, learning algorithm, reward model, and policy training prompts, systematically investigate the impact of these components on downstream model performance, and suggest a recipe for strong learning for preference feedback. Our findings indicate that all aspects are important for performance, with better preference data leading to the largest improvements, followed by the choice of learning algorithm, the use of improved reward models, and finally the use of additional unlabeled prompts for policy training. Notably, PPO outperforms DPO by up to 2.5% in math and 1.2% in general domains. High-quality preference data leads to improvements of up to 8% in instruction following and truthfulness. Despite significant gains of up to 5% in mathematical evaluation when scaling up reward models, we surprisingly observe marginal improvements in other categories. Hamish Ivison, Yizhong Wang, Jiacheng Liu 0010, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert 0001, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 9 |
| 2024 | Paloma: A Benchmark for Evaluating Language Model FitabstractEvaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Pete Walsh 0001, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson 0001, Jesse Dodge |
NeurIPS | 13 |
| 2024 | ActionAtlas: A VideoQA Benchmark for Domain-specialized Action RecognitionabstractOur world is full of varied actions and moves in specialized fields that we, as humans, seek to identify and learn about. To evaluate the effectiveness of multi-modal models in helping us recognize such fine-grained actions, we introduce ActionAtlas, a video question answering (VideoQA) benchmark on fine-grained action recognition with short videos across various sports. ActionAtlas contains 554 videos spanning 284 actions across 42 sports with 1161 actions as total potential choices. Unlike most existing action recognition benchmarks that focus on simplistic actions, often identifiable from a single frame, ActionAtlas focuses on intricate movements and tests the models' ability to discern subtle differences. Additionally, each video in ActionAtlas also includes a question, which helps to more accurately pinpoint the action's performer in scenarios where multiple individuals are involved in different activities. We evaluate proprietary and open models on this benchmark and show that the state-of-the-art models only perform at most 48.73% accurately where random chance is 20%. Furthermore, our results show that a high frame sampling rate is essential for recognizing actions in ActionAtlas, a feature that current top proprietary models like Gemini lack in their default settings. Mohammadreza Salehi, Aditya Kusupati, Ranjay Krishna, Yejin Choi 0001, Hannaneh Hajishirzi, Ali Farhadi |
NeurIPS | 6 |
| 2024 | Decoding-Time Language Model Alignment with Multiple ObjectivesabstractAligning language models (LMs) to human preferences has emerged as a critical pursuit, enabling these models to better serve diverse user needs. Existing methods primarily focus on optimizing LMs for a single reward function, limiting their adaptability to varied objectives.
Here, we propose $\textbf{multi-objective decoding~(MOD)}$, a decoding-time algorithm that outputs the next token from a linear combination of predictions of all base models, for any given weighting over different objectives.
We exploit a common form among a family of $f$-divergence regularized alignment approaches (such as PPO, DPO, and their variants) to identify a closed-form solution by Legendre transform, and derive an efficient decoding strategy.
Theoretically, we show why existing approaches can be sub-optimal even in natural settings and obtain optimality guarantees for our method.
Empirical results demonstrate the effectiveness of the algorithm. For example, compared to a parameter-merging baseline, MOD achieves 12.8\% overall reward improvement when equally optimizing towards $3$ objectives. Moreover, we experiment with MOD on combining three fully-finetuned
LMs of different model sizes, each aimed at different objectives such as safety, coding, and general user preference. Unlike traditional methods that require careful curation of a mixture of datasets to achieve comprehensive improvement, we can quickly experiment with preference weightings using MOD to find the best combination of models. Our best combination reduces toxicity on Toxigen to nearly 0\% and achieves 7.9--33.3\% improvement across three other metrics ($\textit{i.e.}$, Codex@1, GSM-COT, BBH-COT). Ruizhe Shi, Yifang Chen 0001, Yushi Hu, Alisa Liu, Hannaneh Hajishirzi, Noah A. Smith, Simon S. Du |
NeurIPS | 5 |
| 2023 | CREPE: Open-Domain Question Answering with False PresuppositionsabstractWhen asking about unfamiliar topics, information seeking users often pose questions with false presuppositions.Most existing question answering (QA) datasets, in contrast, assume all questions have well defined answers.We introduce CREPE, a QA dataset containing a natural distribution of presupposition failures from online information-seeking forums.We find that 25% of questions contain false presuppositions, and provide annotations for these presuppositions and their corrections.Through extensive baseline experiments, we show that adaptations of existing open-domain QA models can find presuppositions moderately well, but struggle when predicting whether a presupposition is factually correct.This is in large part due to difficulty in retrieving relevant evidence passages from a large text corpus.CREPE provides a benchmark to study question answering in the wild, and our analyses provide avenues for future work in better modeling and further studying the task. 1Question: If there's an equal and opposite reaction for everything, how does any action happen?Isn't it balanced out by the opposite reaction?False presupposition: The equal and opposite reaction applies to the same object.Correction: Based on Newton's Law of Motion, the equal and opposite reaction applies to the other object.Only forces that are applied to the same object would be cancelled out. Newton's laws of motionFrom Wikipedia, the free encyclopedia Inputs given to the human raters Question: Why do prosecuters/courts seek/sentence prison time greater than the expected lifespan of the offender (i.e. 150 years in prison)?Why not simply sentence those criminals to 'life' in prison instead?Comment: Sentencing options are written into state laws.Life in prison is different in state laws than 150 years.Some of it comes into play with the "cruel and unusual punishment" clause in the Constitution too.Life in prison may not be "cruel and unusual" for a murder sentence, but it might be for, say, child sex trafficking.But if you trafficked 10 kids and the sentence is 15 years for each one, you get an effective life sentence that will also stand up, Constitutionally, against a "cruel and unusual punishment" defense. Outputs human raters rateReference Presupposition: It does not make sense to sentence a person to 150 years in prison if they can't live that long anyways, prosecutors should use the life in prison sentence instead.Correction: The defendant can argue the life in prison sentence as cruel and unusual, so the actual year sentence is better to give than the alternative.GOLD-COMMENT track, Dedicated Presupposition: Penalties should be able to be sentenced to life in prison.Correction: Life in prison is different in state laws than 150 years in prison.GOLD-COMMENT track, Unified Presupposition: If a criminal is sentenced to life in prison, they should be sentenced to life in prison.Correction: It is not the case that if a criminal is sentenced to life in prison, they should be sentenced to life in prison.Main, Dedicated Presupposition: Penalties should be able to be imposed on criminals for life.Correction: The longer the sentence, the more likely the prosecution will seek to sentence the offender to life in prison.Main, Unified Presupposition: Prosecutor's should seek prison time greater than the expected lifespan of the offender.Correction: It is not the case that prosecutor's should seek prison time greater than the expected lifespan of the offender.Table 13: An example of the input and the output human raters are given for the human evaluation of the writing subtask.Note that human raters are not given which output is a reference or from which system. Inputs given to the human ratersQuestion: Why did scientists in the 1970s think that there was going to be a new ice age soon?Comment: They didn't.Between 1965 and 1979, there was 7 papers talking about global cooling (not ice age and not necessarily soon).During the same period there was 44 papers about global warming.The media just liked the sensationalism, so there was some news article and a front page on the Times Magazine.They started with a minority of scientist talking about global cooling in a time period when there was still a lot of unknown in climate science and changed that to Scientific consensus that an Ice Age is coming soon.The 7 papers were the following : McComick and Ludwig 1967, Barrett 1971, Rasool and Xinyan Yu 0001, Sewon Min, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 4 |
| 2023 | PuMer: Pruning and Merging Tokens for Efficient Vision Language ModelsabstractLarge-scale vision language (VL) models use Transformers to perform cross-modal interactions between the input text and image.These cross-modal interactions are computationally expensive and memory-intensive due to the quadratic complexity of processing the input image and text.We present PuMer 1 : a token reduction framework that uses text-informed Pruning and modality-aware Merging strategies to progressively reduce the tokens of input image and text, improving model inference speed and reducing memory footprint.PuMer learns to keep salient image tokens related to the input text and merges similar textual and visual tokens by adding lightweight token reducer modules at several cross-modal layers in the VL model.Training PuMer is mostly the same as finetuning the original VL model but faster.Our evaluation for two vision language models on four downstream VL tasks shows PuMer increases inference throughput by up to 2x and reduces memory footprint by over 50% while incurring less than a 1% accuracy drop. 2 Bhargavi Paranjape, Hannaneh Hajishirzi |
ACL (1) | 3 |
| 2023 | HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot GeneralisationabstractRecent NLP models have shown the remarkable ability to effectively generalise 'zero-shot' to new tasks using only natural language instructions as guidance.However, many of these approaches suffer from high computational costs due to their reliance on concatenating lengthy instructions with every input example, resulting in costly reprocessing of the instruction.To avoid this, we introduce Hypernetworks for INstruction Tuning (HINT), which convert task instructions and examples into parameter-efficient modules inserted into an underlying model using a pretrained text encoder, eliminating the need to include instructions in the model input.The hypernetwork in HINT also produces an encoded instruction, which we concatenate with encoded inputs during decoding to further improve performance.HINT models outperform strong state-of-theart baselines by over 10% when controlling for compute (measured in FLOPs).By converting instructions into modules, HINT models can effectively disregard the length of instructions and few-shot example inputs in terms of compute usage.As a result, HINT can enhance its performance by up to 25% by incorporating additional few-shot data, while utilizing only up to 5% more compute.This combines the strengths of parameter-efficient fine-tuning and in-context learning.We release our code publicly 1 . Hamish Ivison, Akshita Bhagia, Yizhong Wang, Hannaneh Hajishirzi, Matthew E. Peters |
ACL (1) | 4 |
| 2023 | Z-ICL: Zero-Shot In-Context Learning with Pseudo-DemonstrationsabstractAlthough large language models can be prompted for both zero-and few-shot learning, performance drops significantly when no demonstrations are available.In this paper, we introduce Z-ICL, a new zero-shot method that closes the gap by constructing pseudo-demonstrations for a given test input using a raw text corpus.Concretely, pseudodemonstrations are constructed by (1) finding the nearest neighbors to the test input from the corpus and pairing them with random task labels, and (2) applying a set of techniques to reduce the amount of direct copying the model does from the resulting demonstrations.Evaluation on nine classification datasets shows that Z-ICL outperforms previous zero-shot methods by a significant margin, and is on par with incontext learning with few-shot labeled training data.Overall, Z-ICL provides a significantly higher estimate of the zero-shot performance levels of a model, and supports future efforts to develop better pseudo-demonstrations that further improve zero-shot results. 1 Xinxi Lyu, Sewon Min, Iz Beltagy, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 5 |
| 2023 | When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesabstractAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 6 |
| 2023 | Self-Instruct: Aligning Language Models with Self-Generated InstructionsabstractYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 7 |
| 2023 | Elaboration-Generating Commonsense Question Answering at ScaleabstractIn question answering requiring common sense, language models (e.g., GPT-3) have been used to generate text expressing background knowledge that helps improve performance.Yet the cost of working with such models is very high; in this work, we finetune smaller language models to generate useful intermediate context, referred to here as elaborations.Our framework alternates between updating two language models-an elaboration generator and an answer predictor-allowing each to influence the other.Using less than 0.5% of the parameters of GPT-3, our model outperforms alternatives with similar sizes and closes the gap with GPT-3 on four commonsense question answering benchmarks.Human evaluations show that the quality of the generated elaborations is high. 1 Wenya Wang 0001, Vivek Srikumar, Hannaneh Hajishirzi, Noah A. Smith |
ACL (1) | 3 |
| 2023 | FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context LearningabstractLarge pre-trained models are capable of fewshot in-context learning (ICL), i.e., performing a new task by prepending a few demonstrations before the test input.However, the concatenated demonstrations are often excessively long and induce additional computation.Inspired by fusion-in-decoder (FiD) models which efficiently aggregate more passages and thus outperforms concatenation-based models in opendomain QA, we hypothesize that similar techniques can be applied to improve the efficiency and end-task performance of ICL.To verify this, we present a comprehensive study on applying three fusion methods-concatenationbased (early fusion), FiD (intermediate), and ensemble-based (late)-to ICL.We adopt a meta-learning setup where a model is first trained to perform ICL on a mixture of tasks using one selected fusion method, then evaluated on held-out tasks for ICL.Results on 11 heldout tasks show that FiD-ICL matches or outperforms the other two fusion methods.Additionally, we show that FiD-ICL (1) is 10x faster at inference time compared to concat-based and ensemble-based ICL, as we can easily precompute the representations of in-context examples and reuse them; (2) enables scaling up to meta-training 3B-sized models, which would fail for concat-based ICL.1 Qinyuan Ye, Iz Beltagy, Matthew E. Peters, Xiang Ren 0001, Hannaneh Hajishirzi |
ACL (1) | 5 |
| 2023 | Crystal: Introspective Reasoners Reinforced with Self-FeedbackabstractExtensive work has shown that the performance and interpretability of commonsense reasoning can be improved via knowledge-augmented reasoning methods, where the knowledge that underpins the reasoning process is explicitly verbalized and utilized.However, existing implementations, including "chain-of-thought" and its variants, fall short in capturing the introspective nature of knowledge required in commonsense reasoning, and in accounting for the mutual adaptation between the generation and utilization of knowledge.We propose a novel method to develop an introspective commonsense reasoner, CRYSTAL.To tackle commonsense problems, it first introspects for knowledge statements related to the given question, and subsequently makes an informed prediction that is grounded in the previously introspected knowledge.The knowledge introspection and knowledge-grounded reasoning modes of the model are tuned via reinforcement learning to mutually adapt, where the reward derives from the feedback given by the model itself.Experiments show that CRYSTAL significantly outperforms both the standard supervised finetuning and chain-of-thought distilled methods, and enhances the transparency of the commonsense reasoning process.Our work ultimately validates the feasibility and potential of reinforcing a neural model with self-feedback. 1 Jiacheng Liu 0010, Ramakanth Pasunuru, Hannaneh Hajishirzi, Yejin Choi 0001, Asli Celikyilmaz |
EMNLP | 3 |
| 2023 | Vera: A General-Purpose Plausibility Estimation Model for Commonsense StatementsabstractToday's language models can be remarkably intelligent yet still produce text that contains trivial commonsense errors.Therefore, we seek a retrospective verification approach that can reflect on the commonsense plausibility of the machine text, and introduce VERA, a general-purpose model that learns to estimate the commonsense plausibility of declarative statements.To support diverse commonsense domains, VERA is trained on ∼7M commonsense statements that are automatically converted from 19 QA datasets and two commonsense knowledge bases, and using a combination of three training objectives.When applied to solving commonsense problems in the verification format, VERA substantially outperforms existing models that can be repurposed for commonsense verification, even including GPT-3.5/ChatGPT/GPT-4, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs.We find that VERA excels at filtering machinegenerated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings. Jiacheng Liu 0010, Wenya Wang 0001, Dianzhuo Wang, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
EMNLP | 6 |
| 2023 | TaskWeb: Selecting Better Source Tasks for Multi-task NLPabstractRecent work in NLP has shown promising results in training models on large amounts of tasks to achieve better generalization.However, it is not well-understood how tasks are related, and how helpful training tasks can be chosen for a new task.In this work, we investigate whether knowing task relationships via pairwise task transfer improves choosing one or more source tasks that help to learn a new target task.We provide TASKWEB, a largescale benchmark of pairwise task transfers for 22 NLP tasks using three different model types, sizes, and adaptation methods, spanning about 25,000 experiments.Then, we design a new method TASKSHOP based on our analysis of TASKWEB.TASKSHOP uses TASKWEB to estimate the benefit of using a source task for learning a new target task, and to choose a subset of helpful training tasks for multi-task training.Our method improves overall rankings and top-k precision of source tasks by 10% and 38%, respectively.We also use TASKSHOP to build much smaller multi-task training sets that improve zero-shot performances across 11 different target tasks by at least 4.3%. Joongwon Kim, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2023 | FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationabstractSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Scott Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 9 |
| 2023 | Editing models with task arithmetic
Gabriel Ilharco, Marco Túlio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi |
ICLR | 5 |
| 2023 | AGRO: Adversarial discovery of error-prone Groups for Robust Optimization
Bhargavi Paranjape, Pradeep Dasigi, Vivek Srikumar, Luke Zettlemoyer, Hannaneh Hajishirzi |
ICLR | 5 |
| 2023 | Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, Yejin Choi 0001 |
ICLR | 7 |
| 2023 | Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World ModellingabstractReinforcement learning (RL) agents typically learn tabula rasa, without prior knowledge of the world. However, if initialized with knowledge of high-level subgoals and transitions between subgoals, RL agents could utilize this Abstract World Model (AWM) for planning and exploration. We propose using few-shot large language models (LLMs) to hypothesize an AWM, that will be verified through world experience, to improve sample efficiency of RL agents. Our DECKARD agent applies LLM-guided exploration to item crafting in Minecraft in two phases: (1) the Dream phase where the agent uses an LLM to decompose a task into a sequence of subgoals, the hypothesized AWM; and (2) the Wake phase where the agent learns a modular policy for each subgoal and verifies or corrects the hypothesized AWM. Our method of hypothesizing an AWM with LLMs and then verifying the AWM based on agent experience not only increases sample efficiency over contemporary methods by an order of magnitude but is also robust to and corrects errors in the LLM, successfully blending noisy internet-scale information from LLMs with knowledge grounded in environment dynamics. Kolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi 0001, Hannaneh Hajishirzi, Sameer Singh 0001, Roy Fox |
ICML | 5 |
| 2023 | DataComp: In search of the next generation of multimodal datasetsabstractMultimodal datasets are a critical component in recent breakthroughs such as CLIP, Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the machine learning ecosystem, we introduce DataComp, a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl. Participants in our benchmark design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets. Our benchmark consists of multiple compute scales spanning four orders of magnitude, which enables the study of scaling trends and makes the benchmark accessible to researchers with varying resources. Our baseline experiments show that the DataComp workflow leads to better training sets. Our best baseline, DataComp-1B, enables training a CLIP ViT-L/14 from scratch to 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI's CLIP ViT-L/14 by 3.7 percentage points while using the same training procedure and compute. We release \datanet and all accompanying code at www.datacomp.ai. Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang 0001, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah M. Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alexandros G. Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, Ludwig Schmidt |
NeurIPS | 26 |
| 2023 | GenEval: An object-focused framework for evaluating text-to-image alignmentabstractRecent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automated methods are critical for evaluating the increasingly large number of new models. However, most current automated evaluation metrics like FID or CLIPScore only offer a distribution-level measure of image quality or image-text alignment, and are unsuited for fine-grained or instance-level analysis. In this paper, we introduce GenEval, an object-focused framework to evaluate compositional image properties such as object co-occurrence, position, count, and color. We show that current object detection models can be leveraged to evaluate text-to-image models on a variety of generation tasks with strong human agreement, and that other discriminative vision models can be linked to this pipeline to further verify properties like object color. We then evaluate several open-source text-to-image models and analyze their relative reasoning capabilities on our benchmark. We find that recent models demonstrate significant improvement on these tasks, though they are still lacking in complex capabilities such as spatial relations and attribute binding. Finally, we demonstrate how GenEval might be used to help discover existing failure modes, in order to inform development of the next generation of text-to-image models. Our code to run the GenEval framework will be made publicly available at https://github.com/djghosh13/geneval. Dhruba Ghosh, Hannaneh Hajishirzi, Ludwig Schmidt |
NeurIPS | 2 |
| 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open ResourcesabstractIn this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu, our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.Our experiments show that different instruction-tuning datasets can uncover or enhance specific skills, while no single dataset (or combination) provides the best performance across all evaluations. Interestingly, we find that model and human preference-based evaluations fail to reflect differences in model capabilities exposed by benchmark-based evaluations, suggesting the need for the type of systemic evaluation performed in this work. Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and instruction-tuning data is required to close the gap. We release our instruction-tuned models, including a fully finetuned 65B Tülu, along with our code, data, and evaluation framework to facilitate future research. Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, Dave Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, Hannaneh Hajishirzi |
NeurIPS | 11 |
| 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingabstractLanguage models (LMs) often exhibit undesirable text generation behaviors, including generating false, toxic, or irrelevant outputs.
Reinforcement learning from human feedback (RLHF)---where human preference judgments on LM outputs are transformed into a learning signal---has recently shown promise in addressing these issues. However, such holistic feedback conveys limited information on long text outputs; it does not indicate which aspects of the outputs influenced user preference; e.g., which parts contain what type(s) of errors. In this paper, we use fine-grained human feedback (e.g., which sentence is false, which sub-sentence is irrelevant) as an explicit training signal. We introduce Fine-Grained RLHF, a framework that enables training and learning from reward functions that are fine-grained in two respects: (1) density, providing a reward after every segment (e.g., a sentence) is generated; and (2) incorporating multiple reward models associated with different feedback types (e.g., factual incorrectness, irrelevance, and information incompleteness). We conduct experiments on detoxification and long-form question answering to illustrate how learning with this reward function leads to improved performance, supported by both automatic and human evaluation. Additionally, we show that LM behaviors can be customized using different combinations of fine-grained reward models. We release all data, collected human feedback, and codes at https://FineGrainedRLHF.github.io. Zeqiu Wu, Yushi Hu, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, Hannaneh Hajishirzi |
NeurIPS | 9 |
| 2023 | InSCIt: Information-Seeking Conversations with Mixed-Initiative InteractionsabstractAbstract In an information-seeking conversation, a user may ask questions that are under-specified or unanswerable. An ideal agent would interact by initiating different response types according to the available knowledge sources. However, most current studies either fail to or artificially incorporate such agent-side initiative. This work presents InSCIt, a dataset for Information-Seeking Conversations with mixed-initiative Interactions. It contains 4.7K user-agent turns from 805 human-human conversations where the agent searches over Wikipedia and either directly answers, asks for clarification, or provides relevant information to address user queries. The data supports two subtasks, evidence passage identification and response generation, as well as a human evaluation protocol to assess model performance. We report results of two systems based on state-of-the-art models of conversational knowledge identification and open-domain question answering. Both systems significantly underperform humans, suggesting ample room for improvement in future studies.1 Zeqiu Wu, Ryu Parish, Hao Cheng 0002, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, Hannaneh Hajishirzi |
Trans. Assoc. Comput. Linguistics | 7 |
| 2022 | Generated Knowledge Prompting for Commonsense ReasoningabstractJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, Hannaneh Hajishirzi. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Jiacheng Liu 0010, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras 0001, Yejin Choi 0001, Hannaneh Hajishirzi |
ACL (1) | 8 |
| 2022 | Noisy Channel Language Model Prompting for Few-Shot Text ClassificationabstractWe introduce a noisy channel approach for language model prompting in few-shot text classification.Instead of computing the likelihood of the label given the input (referred as direct models), channel models compute the conditional probability of the input given the label, and are thereby required to explain every word in the input.We use channel models for recently proposed few-shot learning methods with no or very limited updates to the language model parameters, via either in-context demonstration or prompt tuning.Our experiments show that, for both methods, channel models significantly outperform their direct counterparts, which we attribute to their stability, i.e., lower variance and higher worstcase accuracy.We also present extensive ablations that provide recommendations for when to use channel prompt tuning instead of other competitive methods (e.g., direct head tuning): channel prompt tuning is preferred when the number of training examples is small, labels in the training data are imbalanced, or generalization to unseen labels is required. Sewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
ACL (1) | 3 |
| 2022 | Cross-Task Generalization via Natural Language Crowdsourcing InstructionsabstractHumans (e.g., crowdworkers) have a remarkable ability in solving different tasks, by simply reading textual instructions that define them and looking at a few examples.Despite the success of the conventional supervised learning on individual datasets, such models often struggle with generalization across tasks (e.g., a question-answering system cannot solve classification tasks).A long-standing challenge in AI is to build a model that learns a new task by understanding the humanreadable instructions that define it.To study this, we introduce NATURAL INSTRUCTIONS, a dataset of 61 distinct tasks, their humanauthored instructions, and 193k task instances (input-output pairs).The instructions are obtained from crowdsourcing instructions used to create existing NLP datasets and mapped to a unified schema.Using this meta-dataset, we measure cross-task generalization by training models on seen tasks and measuring generalization to the remaining unseen ones.We adopt generative pre-trained language models to encode task-specific instructions along with input and generate task output.Our results indicate that models benefit from instructions when evaluated in terms of generalization to unseen tasks (19% better for models utilizing instructions).These models, however, are far behind an estimated performance upperbound, indicating significant room for more progress in this direction.1 Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi |
ACL (1) | 4 |
| 2022 | FaVIQ: FAct Verification from Information-seeking QuestionsabstractDespite significant interest in developing general purpose fact checking models, it is challenging to construct a large-scale fact verification dataset with realistic real-world claims.Existing claims are either authored by crowdworkers, thereby introducing subtle biases that are difficult to control for, or manually verified by professional fact checkers, causing them to be expensive and limited in scale.In this paper, we construct a large-scale challenging fact verification dataset called FAVIQ, consisting of 188k claims derived from an existing corpus of ambiguous information-seeking questions.The ambiguities in the questions enable automatically constructing true and false claims that reflect user confusions (e.g., the year of the movie being filmed vs. being released).Claims in FAVIQ are verified to be natural, contain little lexical bias, and require a complete understanding of the evidence for verification.Our experiments show that the stateof-the-art models are far from solving our new task.Moreover, training on our data helps in professional fact-checking, outperforming models trained on the widely used dataset FEVER or in-domain data by up to 17% absolute.Altogether, our data will serve as a challenging benchmark for natural language understanding and support future progress in professional fact checking.1 Jungsoo Park, Sewon Min, Jaewoo Kang, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 5 |
| 2022 | Robust fine-tuning of zero-shot modelsabstractLarge pre-trained models such as CLIP or ALIGN offer consistent accuracy across a range of data distributions when performing zero-shot inference (i.e., without fine-tuning on a specific dataset). Although existing fine-tuning methods substantially improve accuracy on a given target distribution, they often reduce robustness to distribution shifts. We address this tension by introducing a simple and effective method for improving robustness while fine-tuning: ensembling the weights of the zero-shot and fine-tuned models (WiSE-FT). Compared to standard fine-tuning, WiSE-FT provides large accuracy improvements under distribution shift, while preserving high accuracy on the target distribution. On ImageNet and five derived distribution shifts, WiSE-FT improves accuracy under distribution shift by 4 to 6 percentage points (pp) over prior work while increasing ImageNet accuracy by 1.6 pp. WiSE-FT achieves similarly large robustness gains (2 to 23 pp) on a diverse set of six further distribution shifts, and accuracy gains of 0.8 to 3.3 pp compared to standard fine-tuning on commonly used transfer learning datasets. These improvements come at no additional computational cost during fine-tuning or inference. Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, Ludwig Schmidt |
CVPR | 8 |
| 2022 | Rainier: Reinforced Knowledge Introspector for Commonsense Question AnsweringabstractKnowledge underpins reasoning.Recent research demonstrates that when relevant knowledge is provided as additional context to commonsense question answering (QA), it can substantially enhance the performance even on top of state-of-the-art.The fundamental challenge is where and how to find such knowledge that is high quality and on point with respect to the question; knowledge retrieved from knowledge bases are incomplete and knowledge generated from language models are inconsistent.We present RAINIER 1 , or Reinforced Knowledge Introspector, that learns to generate contextually relevant knowledge in response to given questions.Our approach starts by imitating knowledge generated by GPT-3, then learns to generate its own knowledge via reinforcement learning where rewards are shaped based on the increased performance on the resulting question answering.RAINIER demonstrates substantial and consistent performance gains when tested over 9 different commonsense benchmarks: including 5 datasets that are seen during model training, as well as 4 datasets that are kept unseen.Our work is the first to report that knowledge generated by models that are orders of magnitude smaller than GPT-3, even without direct supervision on the knowledge itself, can exceed the quality of commonsense knowledge elicited from GPT-3. Jiacheng Liu 0010, Skyler Hallinan, Ximing Lu, Sean Welleck, Hannaneh Hajishirzi, Yejin Choi 0001 |
EMNLP | 6 |
| 2022 | ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft PromptsabstractThis work introduces a new multi-task, parameter-efficient language model (LM) tuning method that learns to transfer knowledge across different tasks via a mixture of soft prompts-small prefix embedding vectors pretrained for different tasks.Our method, called ATTEMPT (ATTEntional Mixtures of Prompt Tuning), obtains source prompts as encodings of large-scale source tasks into a small number of parameters and trains an attention module to interpolate the source prompts and a newly initialized target prompt for every instance in the target task.During training, only the target task prompt and the attention weights, which are shared between tasks in multi-task training, are updated, while the original LM and source prompts are intact.ATTEMPT is highly parameter-efficient (e.g., updates 2,300 times fewer parameters than full fine-tuning), while achieving high task performance using knowledge from high-resource tasks.Moreover, it is modular using pre-trained soft prompts and can flexibly add or remove source prompts for effective knowledge transfer.Our experimental results across 21 diverse NLP datasets show that ATTEMPT significantly outperforms prompt tuning and outperforms or matches fully finetuned or other parameter-efficient tuning approaches that use over ten times more parameters.Finally, ATTEMPT outperforms previous work in few-shot learning settings. Akari Asai, Mohammadreza Salehi, Matthew E. Peters, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2022 | Correcting Diverse Factual Errors in Abstractive Summarization via Post-Editing and Language Model InfillingabstractAbstractive summarization models often generate inconsistent summaries containing factual errors or hallucinated content.Recent works focus on correcting factual errors in generated summaries via post-editing.Such correction models are trained using adversarial nonfactual summaries constructed using heuristic rules for injecting errors.However, generating non-factual summaries using heuristics often does not generalize well to actual model errors.In this work, we propose to generate hard, representative synthetic examples of nonfactual summaries through infilling language models.With this data, we train a more robust fact-correction model to post-edit the summaries to improve factual consistency.Through quantitative and qualitative experiments on two popular summarization datasets-CNN/DM and XSum-we show that our approach vastly outperforms prior methods in correcting erroneous summaries.Our model-FACTEDITimproves factuality scores by over ∼11 points on CNN/DM and over ∼31 points on XSum on average across multiple summarization models, producing more factual summaries while maintaining competitive summarization quality. 1 The first vaccine for Ebola was approved by the FDA in 2019 in the US, five years after the initial outbreak in 2014.To produce the vaccine, scientists had to sequence the DNA of Ebola, then identify possible vaccines, and finally show successful clinical trials.Scientists say a vaccine for COVID-19 is unlikely to be ready this year, although clinical trials have already started.Scientists believe a vaccine for Covid-19 might not be ready this year.The first vaccine for Ebola took 5 years to be approved by the FDA.Scientists believe a vaccine for Ebola might not be ready this year.The first vaccine for Ebola took 5 years to be produced by the CBP. Vidhisha Balachandran, Hannaneh Hajishirzi, William W. Cohen, Yulia Tsvetkov |
EMNLP | 2 |
| 2022 | Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?abstractLarge language models (LMs) are able to incontext learn-perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs.However, there has been little understanding of how the model learns and which aspects of the demonstrations contribute to end task performance.In this paper, we show that ground truth demonstrations are in fact not required-randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks, consistently over 12 different models including GPT-3.Instead, we find that other aspects of the demonstrations are the key drivers of end task performance, including the fact that they provide a few examples of (1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence.Together, our analysis provides a new way of understanding how and why in-context learning works, while opening up new questions about how much can be learned from large language models through inference alone. Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP | 6 |
| 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningabstractCompared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dialogue context.Moreover, it can be expensive to retrain well-established retrievers such as search engines that are originally developed for nonconversational queries.To facilitate their use, we develop a query rewriting model CONQRR that rewrites a conversational question in the context into a standalone question.It is trained with a novel reward function to directly optimize towards retrieval using reinforcement learning and can be adapted to any off-theshelf retriever.CONQRR achieves state-ofthe-art results on a recent open-domain CQA dataset containing conversations from three different sources, and is effective for two different off-the-shelf retrievers.Our extensive analysis also shows the robustness of CON-QRR to out-of-domain dialogues as well as to zero query rewriting supervision. Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, Gaurav Tomar |
EMNLP | 5 |
| 2022 | Knowledge Base Question Answering by Case-based Reasoning over SubgraphsabstractQuestion answering (QA) over knowledge bases (KBs) is challenging because of the diverse, essentially unbounded, types of reasoning patterns needed. However, we hypothesize in a large KB, reasoning patterns required to answer a query type reoccur for various entities in their respective subgraph neighborhoods. Leveraging this structural similarity between local neighborhoods of different subgraphs, we introduce a semiparametric model (CBR-SUBG) with (i) a nonparametric component that for each query, dynamically retrieves other similar $k$-nearest neighbor (KNN) training queries along with query-specific subgraphs and (ii) a parametric component that is trained to identify the (latent) reasoning patterns from the subgraphs of KNN queries and then apply them to the subgraph of the target query. We also propose an adaptive subgraph collection strategy to select a query-specific compact subgraph, allowing us to scale to full Freebase KB containing billions of facts. We show that CBR-SUBG can answer queries requiring subgraph reasoning patterns and performs competitively with the best models on several KBQA benchmarks. Our subgraph collection strategy also produces more compact subgraphs (e.g. 55% reduction in size for WebQSP while increasing answer recall by 4.85%)\footnote{Code, model, and subgraphs are available at \url{https://github.com/rajarshd/CBR-SUBG}}. Rajarshi Das, Ameya Godbole, Ankita Naik, Elliot Tower, Manzil Zaheer, Hannaneh Hajishirzi, Robin Jia, Andrew McCallum |
ICML | 6 |
| 2022 | Aligning to Social Norms and Values in Interactive NarrativesabstractPrithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Prithviraj Ammanabrolu, Maarten Sap, Hannaneh Hajishirzi, Yejin Choi 0001 |
NAACL-HLT | 4 |
| 2022 | Evidentiality-guided Generation for Knowledge-Intensive NLP TasksabstractRetrieval-augmented generation models have shown state-of-the-art performance across many knowledge-intensive NLP tasks such as open-domain question answering and fact verification.These models are trained to generate a final output given retrieved passages that can be irrelevant to an input query, leading to learning spurious cues or memorization.This work introduces a method to incorporate evidentiality of passages-whether a passage contains correct evidence to support the outputinto training the generator.We introduce a multi-task learning framework to jointly generate the final output and predict the evidentiality of each passage.Furthermore, we introduce a new task-agnostic method for obtaining high-quality silver evidentiality labels, addressing the issues of gold evidentiality labels being unavailable in most domains.Our experiments on five datasets across three knowledgeintensive tasks show that our new evidentialityguided generator significantly outperforms its direct counterpart on all of them, and advances the state of the art on three of them.Our analysis shows that the multi-task learning and silver evidentiality mining play key roles. Akari Asai, Matt Gardner 0001, Hannaneh Hajishirzi |
NAACL-HLT | 3 |
| 2022 | Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous PromptsabstractDaniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson 0001, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh 0001, Yejin Choi 0001 |
NAACL-HLT | 7 |
| 2022 | MetaICL: Learning to Learn In ContextabstractSewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Sewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi |
NAACL-HLT | 4 |
| 2022 | Patching open-vocabulary models by interpolating weightsabstractOpen-vocabulary models like CLIP achieve high accuracy across many image classification tasks. However, there are still settings where their zero-shot performance is far from optimal. We study model patching, where the goal is to improve accuracy on specific tasks without degrading accuracy on tasks where performance is already adequate. Towards this goal, we introduce PAINT, a patching method that uses interpolations between the weights of a model before fine-tuning and the weights after fine-tuning on a task to be patched. On nine tasks where zero-shot CLIP performs poorly, PAINT increases accuracy by 15 to 60 percentage points while preserving accuracy on ImageNet within one percentage point of the zero-shot model. PAINT also allows a single model to be patched on multiple tasks and improves with model scale. Furthermore, we identify cases of broad transfer, where patching on one task increases accuracy on other tasks even when the tasks have disjoint classes. Finally, we investigate applications beyond common benchmarks such as counting or reducing the impact of typographic attacks on CLIP. Our findings demonstrate that it is possible to expand the set of tasks on which open-vocabulary models achieve high accuracy without re-training them from scratch. Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, Ludwig Schmidt |
NeurIPS | 5 |
| 2022 | NaturalProver: Grounded Mathematical Proof Generation with Language ModelsabstractTheorem proving in natural mathematical language – the mixture of symbolic and natural language used by humans – plays a central role in mathematical advances and education, and tests aspects of reasoning that are core to intelligence. Yet it has remained underexplored with modern generative models. We study large-scale language models on two new generation tasks: suggesting the next step in a mathematical proof, and full proof generation. We develop NaturalProver, a language model that generates proofs by conditioning on background references (e.g. theorems and definitions that are either retrieved or human-provided), and optionally enforces their presence with constrained decoding. On theorems from the NaturalProofs benchmark, NaturalProver improves the quality of next-step suggestions and generated proofs over fine-tuned GPT-3, according to human evaluations from university-level mathematics students. NaturalProver is capable of proving some theorems that require short (2-6 step) proofs, and providing next-step suggestions that are rated as correct and useful over 40% of the time, which is to our knowledge the first demonstration of these capabilities using neural language models. Sean Welleck, Jiacheng Liu 0010, Ximing Lu, Hannaneh Hajishirzi, Yejin Choi 0001 |
NeurIPS | 4 |
| 2022 | End-to-End diagnosis of breast biopsy images with transformers
Sachin Mehta, Ximing Lu, Donald L. Weaver, Hannaneh Hajishirzi, Joann G. Elmore, Linda G. Shapiro |
Medical Image Anal. | 5 |
| 2022 | DiCENet: Dimension-Wise Convolutions for Efficient NetworksabstractWe introduce a novel and generic convolutional unit, DiCE unit, that is built using dimension-wise convolutions and dimension-wise fusion. The dimension-wise convolutions apply light-weight convolutional filtering across each dimension of the input tensor while dimension-wise fusion efficiently combines these dimension-wise representations; allowing the DiCE unit to efficiently encode spatial and channel-wise information contained in the input tensor. The DiCE unit is simple and can be seamlessly integrated with any architecture to improve its efficiency and performance. Compared to depth-wise separable convolutions, the DiCE unit shows significant improvements across different architectures. When DiCE units are stacked to build the DiCENet model, we observe significant improvements over state-of-the-art models across various computer vision tasks including image classification, object detection, and semantic segmentation. On the ImageNet dataset, the DiCENet delivers 2-4 percent higher accuracy than state-of-the-art manually designed models (e.g., MobileNetv2 and ShuffleNetv2). Also, DiCENet generalizes better to tasks (e.g., object detection) that are often used in resource-constrained devices in comparison to state-of-the-art separable convolution-based efficient networks, including neural search-based methods (e.g., MobileNetv3 and MixNet). Sachin Mehta, Hannaneh Hajishirzi, Mohammad Rastegari |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 9 |
| 2021 | Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and TextabstractChristopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Jonghyun Choi, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi |
EMNLP (1) | 13 |
| 2021 | Joint Passage Ranking for Diverse Multi-Answer RetrievalabstractWe study multi-answer retrieval, an underexplored problem that requires retrieving passages to cover multiple distinct answers for a given question.This task requires joint modeling of retrieved passages, as models should not repeatedly retrieve passages containing the same answer at the cost of missing a different valid answer.In this paper, we introduce JPR, the first joint passage retrieval model for multi-answer retrieval.JPR makes use of an autoregressive reranker that selects a sequence of passages, each conditioned on previously selected passages.JPR is trained to select passages that cover new answers at each timestep and uses a tree-decoding algorithm to enable flexibility in the degree of diversity.Compared to prior approaches, JPR achieves significantly better answer coverage on three multianswer datasets.When combined with downstream question answering, the improved retrieval enables larger answer generation models since they need to consider fewer passages, establishing a new state-of-the-art. Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, Hannaneh Hajishirzi |
EMNLP (1) | 5 |
| 2021 | DIALKI: Knowledge Identification in Conversational Systems through Dialogue-Document ContextualizationabstractIdentifying relevant knowledge to be used in conversational systems that are grounded in long documents is critical to effective response generation.We introduce a knowledge identification model that leverages the document structure to provide dialogue-contextualized passage encodings and better locate knowledge relevant to the conversation.An auxiliary loss captures the history of dialogue-document connections.We demonstrate the effectiveness of our model on two document-grounded conversational datasets and provide analyses showing generalization to unseen documents and long dialogue contexts. Zeqiu Wu, Bo-Ru Lu, Hannaneh Hajishirzi, Mari Ostendorf |
EMNLP (1) | 3 |
| 2021 | DeLighT: Deep and Light-weight Transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer 0001, Luke Zettlemoyer, Hannaneh Hajishirzi |
ICLR | 5 |
| 2021 | MultiModalQA: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, Jonathan Berant |
ICLR | 8 |
| 2021 | XOR QA: Cross-lingual Open-Retrieval Question AnsweringabstractAkari Asai, Jungo Kasai, Jonathan Clark, Kenton Lee, Eunsol Choi, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Akari Asai, Jungo Kasai, Jonathan H. Clark, Kenton Lee, Eunsol Choi, Hannaneh Hajishirzi |
NAACL-HLT | 6 |
| 2021 | Extracting a Knowledge Base of Mechanisms from COVID-19 PapersabstractTom Hope, Aida Amini, David Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel Weld, Roy Schwartz, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tom Hope, Aida Amini, Dave Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel S. Weld, Roy Schwartz 0001, Hannaneh Hajishirzi |
NAACL-HLT | 9 |
| 2021 | Probing Contextual Language Models for Common Ground with Visual RepresentationsabstractGabriel Ilharco, Rowan Zellers, Ali Farhadi, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Gabriel Ilharco, Rowan Zellers, Ali Farhadi, Hannaneh Hajishirzi |
NAACL-HLT | 4 |
| 2021 | One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalabstractWe present Cross-lingual Open-Retrieval Answer Generation (CORA), the first unified many-to-many question answering (QA) model that can answer questions across many languages, even for ones without language-specific annotated data or knowledge sources.We introduce a new dense passage retrieval algorithm that is trained to retrieve documents across languages for a question.Combined with a multilingual autoregressive generation model, CORA answers directly in the target language without any translation or in-language retrieval modules as used in prior work. We propose an iterative training method that automatically extends annotated data available only in high-resource languages to low-resource ones. Our results show that CORA substantially outperforms the previous state of the art on multilingual open QA benchmarks across 26 languages, 9 of which are unseen during training. Our analyses show the significance of cross-lingual retrieval and generation in many languages, particularly under low-resource settings. Akari Asai, Xinyan Yu 0001, Jungo Kasai, Hannaneh Hajishirzi |
NeurIPS | 4 |
| 2020 | Logic-Guided Data Augmentation and Regularization for Consistent Question AnsweringabstractMany natural language questions require qualitative, quantitative or logical comparisons between two entities or events.This paper addresses the problem of improving the accuracy and consistency of responses to comparison questions by integrating logic rules and neural models.Our method leverages logical and linguistic knowledge to augment labeled training data and then uses a consistency-based regularizer to train the model.Improving the global consistency of predictions, our approach achieves large improvements over previous methods in a variety of question answering (QA) tasks including multiple-choice qualitative reasoning, cause-effect reasoning, and extractive machine reading comprehension.In particular, our method significantly improves the performance of RoBERTa-based models by 1-5% across datasets.We advance state of the art by around 5-8% on WIQA and QuaRel and reduce consistency violations by 58% on HotpotQA.We further demonstrate that our approach can learn effectively from limited data. 1 Q: The ceramic vase was less flexible than the plastic ball so it was A: more breakable Q: The ceramic vase was more flexible than the plastic ball so it was A: less breakable Q: If it is silent, does the outer ear collect less sound waves?A: more [positive causal relationship] Q: If the outer ear collect less sound waves, is less sound being detected?A: more [positive causal relationship] Q: If it is silent, is less sound being detected?A: more [positive causal relationship] Akari Asai, Hannaneh Hajishirzi |
ACL | 2 |
| 2020 | SciREX: A Challenge Dataset for Document-Level Information ExtractionabstractExtracting information from full documents is an important problem in many domains, but most previous work focus on identifying relationships within a sentence or a paragraph.It is challenging to create a large-scale information extraction (IE) dataset at the document level since it requires an understanding of the whole document to annotate entities and their document-level relationships that usually span beyond sentences or even sections.In this paper, we introduce SCIREX, a document level IE dataset that encompasses multiple IE tasks, including salient entity identification and document level N -ary relation identification from scientific articles.We annotate our dataset by integrating automatic and human annotations, leveraging existing scientific knowledge resources.We develop a neural model as a strong baseline that extends previous state-of-the-art IE models to documentlevel IE.Analyzing the model performance shows a significant gap between human performance and current baselines, inviting the community to use our dataset as a challenge to develop document-level IE models. Madeleine van Zuylen, Hannaneh Hajishirzi, Iz Beltagy |
ACL | 3 |
| 2020 | Contextualized Sparse Representations for Real-Time Open-Domain Question AnsweringabstractOpen-domain question answering can be formulated as a phrase retrieval problem, in which we can expect huge scalability and speed benefit but often suffer from low accuracy due to the limitation of existing phrase representation models.In this paper, we aim to improve the quality of each phrase embedding by augmenting it with a contextualized sparse representation (SPARC).Unlike previous sparse vectors that are term-frequencybased (e.g., tf-idf) or directly learned (only few thousand dimensions), we leverage rectified self-attention to indirectly learn sparse vectors in n-gram vocabulary space.By augmenting the previous phrase retrieval model (Seo et al., 2019) with SPARC, we show 4%+ improvement in CuratedTREC and SQuAD-Open.Our CuratedTREC score is even better than the best known retrieve & read model with at least 45x faster inference speed. 1 Jinhyuk Lee, Minjoon Seo, Hannaneh Hajishirzi, Jaewoo Kang |
ACL | 3 |
| 2020 | ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured WebpagesabstractIn many documents, such as semi-structured webpages, textual semantics are augmented with additional information conveyed using visual elements including layout, font size, and color.Prior work on information extraction from semi-structured websites has required learning an extraction model specific to a given template via either manually labeled or distantly supervised data from that template.In this work, we propose a solution for "zero-shot" open-domain relation extraction from webpages with a previously unseen template, including from websites with little overlap with existing sources of knowledge for distant supervision and websites in entirely new subject verticals.Our model uses a graph neural network-based approach to build a rich representation of text fields on a webpage and the relationships between them, enabling generalization to new templates.Experiments show this approach provides a 31% F1 gain over a baseline for zero-shot extraction in a new subject vertical. Colin Lockard, Prashant Shiralkar, Xin Dong 0001, Hannaneh Hajishirzi |
ACL | 4 |
| 2020 | X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal TransformersabstractMirroring the success of masked language models, vision-and-language counterparts like VILBERT, LXMERT and UNITER have achieved state of the art performance on a variety of multimodal discriminative tasks like visual question answering and visual grounding.Recent work has also successfully adapted such models towards the generative task of image captioning.This begs the question: Can these models go the other way and generate images from pieces of text?Our analysis of a popular representative from this model family -LXMERT -finds that it is unable to generate rich and semantically meaningful imagery with its current training setup.We introduce X-LXMERT, an extension to LXMERT with training refinements including: discretizing visual representations, using uniform masking with a large range of masking ratios and aligning the right pre-training datasets to the right objectives which enables it to paint.X-LXMERT's image generation capabilities rival state of the art generative models while its question answering and captioning abilities remains comparable to LXMERT.Finally, we demonstrate the generality of these training refinements by adding image generation capabilities into UNITER to produce X-UNITER. Jaemin Cho 0001, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi, Aniruddha Kembhavi |
EMNLP (1) | 4 |
| 2020 | IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsabstractHumans often have to read multiple documents to address their information needs.However, most existing reading comprehension (RC) tasks only focus on questions for which the contexts provide all the information required to answer them, thus not evaluating a system's performance at identifying a potential lack of sufficient information and locating sources for that information.To fill this gap, we present a dataset, IIRC, with more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.The questions were written by crowd workers who did not have access to any of the linked documents, leading to questions that have little lexical overlap with the contexts where the answers appear.This process also gave many questions without answers, and those that require discrete reasoning, increasing the difficulty of the task.We follow recent modeling work on various reading comprehension datasets to construct a baseline model for this dataset, finding that it achieves 31.1% F1 on this task, while estimated human performance is 88.4%.The dataset, code for the baseline system, and a leaderboard can be found at https://allennlp.org/iirc. James Ferguson, Matt Gardner 0001, Hannaneh Hajishirzi, Tushar Khot, Pradeep Dasigi |
EMNLP (1) | 3 |
| 2020 | AmbigQA: Answering Ambiguous Open-domain QuestionsabstractAmbiguity is inherent to open-domain question answering; especially when exploring new topics, it can be difficult to ask questions that have a single, unambiguous answer.In this paper, we introduce AMBIGQA, a new open-domain question answering task which involves finding every plausible answer, and then rewriting the question for each one to resolve the ambiguity.To study this task, we construct AMBIGNQ, a dataset covering 14,042 questions from NQ-OPEN, an existing opendomain QA benchmark.We find that over half of the questions in NQ-OPEN are ambiguous, with diverse sources of ambiguity such as event and entity references.We also present strong baseline models for AMBIGQA which we show benefit from weakly supervised learning that incorporates NQ-OPEN, strongly suggesting our new task and data will support significant future research effort.Our data and baselines are available at https://nlp.cs. washington.edu/ambigqa.Type Example Event references (39%) What season does meredith and derek get married in grey's anatomy?Q: In what season do Meredith and Derek get informally married in Grey's Anatomy? / A: Season 5 Q: In what season do Meredith and Derek get legally married in Grey's Anatomy? / A: Season 7 Properties (27%) How many episode in seven deadly sins season 2? Q: How many episodes were there in seven deadly sins season 2, not including the OVA episode?/ A: 25 Q: How many episodes were there in seven deadly sins season 2, including the OVA episode?/ A: 26 Entity references (23%) How many sacks does clay matthews have in his career?Q: How many sacks does Clay Matthews Jr. have in his career?/ A: 69.5 Q: How many sacks does Clay Matthews III have in his career?/ A: 91.5 Answer types (16%) Who sings the song what a beautiful name it is?Q: Which group sings the song what a beautiful name it is?/ A: Hillsong Live Q: Who is the lead singer of the song what a beautiful name it is?/ A: Brooke Ligertwood Sewon Min, Julian Michael, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP (1) | 3 |
| 2020 | An Information Bottleneck Approach for Controlling Conciseness in Rationale ExtractionabstractDecisions of complex models for language understanding can be explained by limiting the inputs they are provided to a relevant subsequence of the original text -a rationale.Models that condition predictions on a concise rationale, while being more interpretable, tend to be less accurate than models that are able to use the entire context.In this paper, we show that it is possible to better manage the trade-off between concise explanations and high task accuracy by optimizing a bound on the Information Bottleneck (IB) objective.Our approach jointly learns an explainer that predicts sparse binary masks over input sentences without explicit supervision, and an end-task predictor that considers only the residual sentences.Using IB, we derive a learning objective that allows direct control of mask sparsity levels through a tunable sparse prior.Experiments on the ERASER benchmark demonstrate significant gains over previous work for both task performance and agreement with human rationales.Furthermore, we find that in the semi-supervised setting, a modest amount of gold rationales (25% of training examples with gold masks) can close the performance gap with a model that uses the full input.1 Bhargavi Paranjape, Mandar Joshi, John Thickstun, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP (1) | 4 |
| 2020 | Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsabstractSwabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Swabha Swayamdipta, Roy Schwartz 0001, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi 0001 |
EMNLP (1) | 5 |
| 2020 | Fact or Fiction: Verifying Scientific ClaimsabstractDavid Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, Hannaneh Hajishirzi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Dave Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, Hannaneh Hajishirzi |
EMNLP (1) | 7 |
| 2020 | Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, Caiming Xiong |
ICLR | 3 |
| 2020 | DeFINE: Deep Factorized Input Token Embeddings for Neural Sequence Modeling
Sachin Mehta, Rik Koncel-Kedziorski, Mohammad Rastegari, Hannaneh Hajishirzi |
ICLR | 4 |
| 2020 | Multi-modal Information Extraction from Text, Semi-structured, and Tabular Data on the WebabstractHow do we surface the large amount of information present in HTML documents on the Web, from news articles to Rotten Tomatoes pages to tables of sports scores? Such information can enable a variety of applications including knowledge base construction, question answering, recommendation, and more. In this tutorial, we present approaches for information extraction (IE) from Web data that can be differentiated along two key dimensions: 1) the diversity in data modality that is leveraged, e.g. text, visual, XML/HTML, and 2) the thrust to develop scalable approaches with zero to limited human supervision. Xin Dong 0001, Hannaneh Hajishirzi, Colin Lockard, Prashant Shiralkar |
KDD | 2 |
| 2020 | Web-scale Knowledge CollectionabstractHow do we surface the large amount of information present in HTML documents on the Web, from news articles to scientific papers to Rotten Tomatoes pages to tables of sports scores? Such information can enable a variety of applications including knowledge base construction, question answering, recommendation, and more. In this tutorial, we present approaches for Information Extraction (IE) from Web data that can be differentiated along two key dimensions: 1) the diversity in data modality that is leveraged, e.g. text, visual, XML/HTML, and 2) the thrust to develop scalable approaches with zero to limited human supervision. We cover the key ideas and intuition behind existing approaches to emphasize their applicability and potential in various settings. Colin Lockard, Prashant Shiralkar, Xin Dong 0001, Hannaneh Hajishirzi |
WSDM | 4 |
| 2019 | Compositional Questions Do Not Necessitate Multi-hop ReasoningabstractMulti-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs.We argue that it can be difficult to construct large multi-hop RC datasets.For example, even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant.Our analysis is centered on HOTPOTQA, where we show that single-hop reasoning can solve much more of the dataset than previously thought.We introduce a single-hop BERT-based RC model that achieves 67 F1-comparable to state-of-theart multi-hop models.We also design an evaluation setting where humans are not shown all of the necessary paragraphs for the intended multi-hop reasoning but can still answer over 80% of questions.Together with detailed error analysis, these results suggest there should be an increasing focus on the role of evidence in multi-hop reasoning and possibly even a shift towards information retrieval style evaluations with large and diverse evidence collections. Sewon Min, Eric Wallace, Sameer Singh 0001, Matt Gardner 0001, Hannaneh Hajishirzi, Luke Zettlemoyer |
ACL (1) | 5 |
| 2019 | Multi-hop Reading Comprehension through Question Decomposition and RescoringabstractMulti-hop Reading Comprehension (RC) requires reasoning and aggregation across several paragraphs.We propose a system for multi-hop RC that decomposes a compositional question into simpler sub-questions that can be answered by off-the-shelf single-hop RC models.Since annotations for such decomposition are expensive, we recast subquestion generation as a span prediction problem and show that our method, trained using only 400 labeled examples, generates sub-questions that are as effective as humanauthored sub-questions.We also introduce a new global rescoring approach that considers each decomposition (i.e. the sub-questions and their answers) to select the best final answer, greatly improving overall performance.Our experiments on HOTPOTQA show that this approach achieves the state-of-the-art results, while providing explainable evidence for its decision making in the form of sub-questions. Sewon Min, Victor Zhong, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 4 |
| 2019 | Real-Time Open-Domain Question Answering with Dense-Sparse Phrase IndexabstractExisting open-domain question answering (QA) models are not suitable for real-time usage because they need to process several long documents on-demand for every input query, which is computationally prohibitive. In this paper, we introduce query-agnostic indexable representations of document phrases that can drastically speed up open-domain QA. In particular, our dense-sparse phrase encoding effectively captures syntactic, semantic, and lexical information of the phrases and eliminates the pipeline filtering of context documents. Leveraging strategies for optimizing training and inference time, our model can be trained and deployed even in a single 4-GPU server. Moreover, by representing phrases as pointers to their start and end tokens, our model indexes phrases in the entire English Wikipedia (up to 60 billion phrases) using under 2TB. Our experiments on SQuAD-Open show that our model is on par with or more accurate than previous models with 6000x reduced computational cost, which translates into at least 68x faster end-to-end inference benchmark on CPUs. Code and demo are available at nlp.cs.washington.edu/denspi Minjoon Seo, Jinhyuk Lee, Tom Kwiatkowski, Ankur P. Parikh, Ali Farhadi, Hannaneh Hajishirzi |
ACL (1) | 6 |
| 2019 | ESPNetv2: A Light-Weight, Power Efficient, and General Purpose Convolutional Neural NetworkabstractWe introduce a light-weight, power efficient, and general purpose convolutional neural network, ESPNetv2, for modeling visual and sequential data. Our network uses group point-wise and depth-wise dilated separable convolutions to learn representations from a large effective receptive field with fewer FLOPs and parameters. The performance of our network is evaluated on four different tasks: (1) object classification, (2) semantic segmentation, (3) object detection, and (4) language modeling. Experiments on these tasks, including image classification on the ImageNet and language modeling on the PenTree bank dataset, demonstrate the superior performance of our method over the state-of-the-art methods. Our network outperforms ESPNet by 4-5% and has 2-4x fewer FLOPs on the PASCAL VOC and the Cityscapes dataset. Compared to YOLOv2 on the MS-COCO object detection, ESPNetv2 delivers 4.4% higher accuracy with 6x fewer FLOPs. Our experiments show that ESPNetv2 is much more power efficient than existing state-of-the-art efficient methods including ShuffleNets and MobileNets. Our code is open-source and available at https://github.com/sacmehta/ESPNetv2. Sachin Mehta, Mohammad Rastegari, Linda G. Shapiro, Hannaneh Hajishirzi |
CVPR | 4 |
| 2019 | Mixture Content Selection for Diverse Sequence GenerationabstractJaemin Cho, Minjoon Seo, Hannaneh Hajishirzi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jaemin Cho 0001, Minjoon Seo, Hannaneh Hajishirzi |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Discrete Hard EM Approach for Weakly Supervised Question AnsweringabstractSewon Min, Danqi Chen, Hannaneh Hajishirzi, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sewon Min, Danqi Chen 0001, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Entity, Relation, and Event Extraction with Contextualized Span RepresentationsabstractDavid Wadden, Ulme Wennberg, Yi Luan, Hannaneh Hajishirzi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dave Wadden, Ulme Wennberg, Yi Luan, Hannaneh Hajishirzi |
EMNLP/IJCNLP (1) | 4 |
| 2018 | ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation
Sachin Mehta, Mohammad Rastegari, Anat Caspi, Linda G. Shapiro, Hannaneh Hajishirzi |
ECCV (10) | 5 |
| 2018 | Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph ConstructionabstractWe introduce a multi-task setup of identifying and classifying entities, relations, and coreference clusters in scientific articles.We create SCIERC, a dataset that includes annotations for all three tasks and develop a unified framework called Scientific Information Extractor (SCIIE) for with shared span representations.The multi-task setup reduces cascading errors between tasks and leverages cross-sentence relations through coreference links.Experiments show that our multi-task model outperforms previous models in scientific information extraction without using any domain-specific features.We further show that the framework supports construction of a scientific knowledge graph, which we use to analyze information in scientific literature. 1 Extracting nodes (entities) The SCIIE model extracts entities, their relations, and coreference Yi Luan, Luheng He, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2018 | Pyramidal Recurrent Unit for Language ModelingabstractLSTMs are powerful tools for modeling contextual information, as evidenced by their success at the task of language modeling.However, modeling contexts in very high dimensional space can lead to poor generalizability.We introduce the Pyramidal Recurrent Unit (PRU), which enables learning representations in high dimensional space with more generalization power and fewer parameters.PRUs replace the linear transformation in LSTMs with more sophisticated interactions including pyramidal and grouped linear transformations.This architecture gives strong results on wordlevel language modeling while reducing the number of parameters significantly.In particular, PRU improves the perplexity of a recent state-of-the-art language model Merity et al. (2018) by up to 1.3 points while learning 15-20% fewer parameters.For similar number of model parameters, PRU outperforms all previous RNN models that exploit different gating mechanisms and transformations.We provide a detailed examination of the PRU and its behavior on the language modeling tasks. Sachin Mehta, Rik Koncel-Kedziorski, Mohammad Rastegari, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2018 | Phrase-Indexed Question Answering: A New Challenge for Scalable Document ComprehensionabstractWe formalize a new modular variant of current question answering tasks by enforcing complete independence of the document encoder from the question encoder.This formulation addresses a key challenge in machine comprehension by requiring a standalone representation of the document discourse.It additionally leads to a significant scalability advantage since the encoding of the answer candidate phrases in the document can be pre-computed and indexed offline for efficient retrieval.We experiment with baseline models for the new task, which achieve a reasonable accuracy but significantly underperform unconstrained QA models.We invite the QA research community to engage in Phrase-Indexed Question Answering (PIQA, pika) for closing the gap.The leaderboard is at: nlp. Minjoon Seo, Tom Kwiatkowski, Ankur P. Parikh, Ali Farhadi, Hannaneh Hajishirzi |
EMNLP | 5 |
| 2018 | Neural Speed Reading via Skim-RNN
Minjoon Seo, Sewon Min, Ali Farhadi, Hannaneh Hajishirzi |
ICLR (Poster) | 4 |
| 2017 | Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine ComprehensionabstractWe introduce the task of Multi-Modal Machine Comprehension (M3C), which aims at answering multimodal questions given a context of text, diagrams and images. We present the Textbook Question Answering (TQA) dataset that includes 1,076 lessons and 26,260 multi-modal questions, taken from middle school science curricula. Our analysis shows that a significant portion of questions require complex parsing of the text and the diagrams and reasoning, indicating that our dataset is more complex compared to previous machine comprehension and visual question answering datasets. We extend state-of-the-art methods for textual machine comprehension and visual question answering to the TQA dataset. Our experiments show that these models do not perform well on TQA. The presented dataset opens new challenges for research in question answering and reasoning across multiple modalities. Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Ali Farhadi, Hannaneh Hajishirzi |
CVPR | 6 |
| 2017 | Scientific Information Extraction with Semi-supervised Neural TaggingabstractThis paper addresses the problem of extracting keyphrases from scientific articles and categorizing them as corresponding to a task, process, or material.We cast the problem as sequence tagging and introduce semi-supervised methods to a neural tagging model, which builds on recent advances in named entity recognition.Since annotated training data is scarce in this domain, we introduce a graph-based semi-supervised algorithm together with a data selection scheme to leverage unannotated articles.Both inductive and transductive semi-supervised learning strategies outperform state-of-the-art information extraction performance on the 2017 SemEval Task 10 ScienceIE task. Yi Luan, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 3 |
| 2017 | Bidirectional Attention Flow for Machine Comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, Hannaneh Hajishirzi |
ICLR (Poster) | 4 |
| 2017 | Query-Reduction Networks for Question Answering
Minjoon Seo, Sewon Min, Ali Farhadi, Hannaneh Hajishirzi |
ICLR (Poster) | 4 |
| 2016 | Are Elephants Bigger than Butterflies? Reasoning about Sizes of ObjectsabstractHuman vision greatly benefits from the information about sizes of objects. The role of size in several visual reasoning tasks has been thoroughly explored in human perception and cognition. However, the impact of the information about sizes of objects is yet to be determined in AI. We postulate that this is mainly attributed to the lack of a comprehensive repository of size information. In this paper, we introduce a method to automatically infer object sizes, leveraging visual and textual information from web. By maximizing the joint likelihood of textual and visual observations, our method learns reliable relative size estimates, with no explicit human supervision. We introduce the relative size dataset and show that our method outperforms competitive textual and visual baselines in reasoning about size comparisons. Hessam Bagherinezhad, Hannaneh Hajishirzi, Yejin Choi 0001, Ali Farhadi |
AAAI | 2 |
| 2016 | Learning Prototypical Event Structure from Photo AlbumsabstractActivities and events in our lives are structural, be it a vacation, a camping trip, or a wedding.While individual details vary, there are characteristic patterns that are specific to each of these scenarios.For example, a wedding typically consists of a sequence of events such as walking down the aisle, exchanging vows, and dancing.In this paper, we present a data-driven approach to learning event knowledge from a large collection of photo albums.We formulate the task as constrained optimization to induce the prototypical temporal structure of an event, integrating both visual and textual cues.Comprehensive evaluation demonstrates that it is possible to learn multimodal knowledge of event structure from noisy web content. Antoine Bosselut, Jianfu Chen, David Scott Warren, Hannaneh Hajishirzi, Yejin Choi 0001 |
ACL (1) | 4 |
| 2016 | A Task-Oriented Approach for Cost-Sensitive RecognitionabstractWith the recent progress in visual recognition, we have already started to see a surge of vision related real-world applications. These applications, unlike general scene understanding, are task oriented and require specific information from visual data. Considering the current growth in new sensory devices, feature designs, feature learning methods, and algorithms, the search in the space of features and models becomes combinatorial. In this paper, we propose a novel cost-sensitive task-oriented recognition method that is based on a combination of linguistic semantics and visual cues. Our task-oriented framework is able to generalize to unseen tasks for which there is no training data and outperforms state-of-the-art cost-based recognition baselines on our new task-based dataset. Roozbeh Mottaghi, Hannaneh Hajishirzi, Ali Farhadi |
CVPR | 2 |
| 2016 | A Diagram is Worth a Dozen Images
Aniruddha Kembhavi, Michael Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi |
ECCV (4) | 5 |
| 2016 | A Theme-Rewriting Approach for Generating Algebra Word ProblemsabstractTexts present coherent stories that have a particular theme or overall setting, for example science fiction or western.In this paper, we present a text generation method called rewriting that edits existing human-authored narratives to change their theme without changing the underlying story.We apply the approach to math word problems, where it might help students stay more engaged by quickly transforming all of their homework assignments to the theme of their favorite movie without changing the math concepts that are being taught.Our rewriting method uses a twostage decoding process, which proposes new words from the target theme and scores the resulting stories according to a number of factors defining aspects of syntactic, semantic, and thematic coherence.Experiments demonstrate that the final stories typically represent the new theme well while still testing the original math concepts, outperforming a number of baselines.We also release a new dataset of human-authored rewrites of math word problems in several themes. Rik Koncel-Kedziorski, Ioannis Konstas, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2016 | Disfluency Detection Using a Bidirectional LSTMabstractWe introduce a new approach for disfluency detection using a Bidirectional Long-Short Term Memory neural network (BLSTM). In addition to the word sequence, the model takes as input pattern match features that were developed to reduce sensitivity to vocabulary size in training, which lead to improved performance over the word sequence alone. The BLSTM takes advantage of explicit repair states in addition to the standard reparandum states. The final output leverages integer linear programming to incorporate constraints of disfluency structure. In experiments on the Switchboard corpus, the model achieves state-of-the-art performance for both the standard disfluency detection task and the correction detection task. Analysis shows that the model has better detection of non-repetition disfluencies, which tend to be much harder to detect. Victoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi |
INTERSPEECH | 3 |
| 2016 | MAWPS: A Math Word Problem RepositoryabstractRik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, Hannaneh Hajishirzi. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, Hannaneh Hajishirzi |
HLT-NAACL | 5 |
| 2015 | Discriminative and consistent similarities in instance-level Multiple Instance LearningabstractIn this paper we present a bottom-up method to instance-level Multiple Instance Learning (MIL) that learns to discover positive instances with globally constrained reasoning about local pairwise similarities. We discover positive instances by optimizing for a ranking such that positive (top rank) instances are highly and consistently similar to each other and dissimilar to negative instances. Our approach takes advantage of a discriminative notion of pairwise similarity coupled with a structural cue in the form of a consistency metric that measures the quality of each similarity. We learn a similarity function for every pair of instances in positive bags by how similarly they differ from instances in negative bags, the only certain labels in MIL. Our experiments demonstrate that our method consistently outperforms state-of-the-art MIL methods both at bag-level and instance-level predictions in standard benchmarks, image category recognition, and text categorization datasets. Mohammad Rastegari, Hannaneh Hajishirzi, Ali Farhadi |
CVPR | 2 |
| 2015 | Talking to the crowd: What do people react to in online discussions?abstractThis paper addresses the question of how language use affects community reaction to comments in online discussion forums, and the relative importance of the message vs. the messenger.A new comment ranking task is proposed based on community annotated karma in Reddit discussions, which controls for topic and timing of comments.Experimental work with discussion threads from six subreddits shows that the importance of different types of language features varies with the community of interest. Aaron Jaech, Victoria Zayats, Hao Fang 0002, Mari Ostendorf, Hannaneh Hajishirzi |
EMNLP | 5 |
| 2015 | Solving Geometry Problems: Combining Text and Diagram InterpretationabstractThis paper introduces GEOS, the first automated system to solve unaltered SAT geometry questions by combining text understanding and diagram interpretation.We model the problem of understanding geometry questions as submodular optimization, and identify a formal problem description likely to be compatible with both the question text and diagram.GEOS then feeds the description to a geometric solver that attempts to determine the correct answer.In our experiments, GEOS achieves a 49% score on official SAT questions, and a score of 61% on practice questions. 1 Finally, we show that by integrating textual and visual information, GEOS boosts the accuracy of dependency and semantic parsing of the question text. Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, Clint Malcolm |
EMNLP | 2 |
| 2015 | Segment-Phrase Table for Semantic Segmentation, Visual Entailment and ParaphrasingabstractWe introduce Segment-Phrase Table (SPT), a large collection of bijective associations between textual phrases and their corresponding segmentations. Leveraging recent progress in object recognition and natural language semantics, we show how we can successfully build a high-quality segment-phrase table using minimal human supervision. More importantly, we demonstrate the unique value unleashed by this rich bimodal resource, for both vision as well as natural language understanding. First, we show that fine-grained textual labels facilitate contextual reasoning that helps in satisfying semantic constraints across image segments. This feature enables us to achieve state-of-the-art segmentation results on benchmark datasets. Next, we show that the association of high-quality segmentations to textual phrases aids in richer semantic understanding and reasoning of these textual phrases. Leveraging this feature, we motivate the problem of visual entailment and visual paraphrasing, and demonstrate its utility on a large dataset. Hamid Izadinia, Fereshteh Sadeghi, Santosh Kumar Divvala, Hannaneh Hajishirzi, Yejin Choi 0001, Ali Farhadi |
ICCV | 4 |
| 2015 | Learning Knowledge Graphs for Question Answering through Conversational DialogabstractWe describe how a question-answering system can learn about its domain from conversational dialogs.Our system learns to relate concepts in science questions to propositions in a fact corpus, stores new concepts and relations in a knowledge graph (KG), and uses the graph to solve questions.We are the first to acquire knowledge for question-answering from open, natural language dialogs without a fixed ontology or domain model that predetermines what users can say.Our relation-based strategies complete more successful dialogs than a query expansion baseline, our taskdriven relations are more effective for solving science questions than relations from general knowledge sources, and our method is practical enough to generalize to other domains. Ben Hixon, Peter Clark, Hannaneh Hajishirzi |
HLT-NAACL | 3 |
| 2015 | Aligning Sentences from Standard Wikipedia to Simple WikipediaabstractWilliam Hwang, Hannaneh Hajishirzi, Mari Ostendorf, Wei Wu. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. William Hwang, Hannaneh Hajishirzi, Mari Ostendorf |
HLT-NAACL | 2 |
| 2015 | Unediting: Detecting Disfluencies Without Careful TranscriptsabstractVictoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Victoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi |
HLT-NAACL | 3 |
| 2015 | Parsing Algebraic Word Problems into EquationsabstractThis paper formalizes the problem of solving multi-sentence algebraic word problems as that of generating and scoring equation trees. We use integer linear programming to generate equation trees and score their likelihood by learning local and global discriminative models. These models are trained on a small set of word problems and their answers, without any manual annotation, in order to choose the equation that best matches the problem text. We refer to the overall system as Alges. We compare Alges with previous work and show that it covers the full gamut of arithmetic operations whereas Hosseini et al. (2014) only handle addition and subtraction. In addition, Alges overcomes the brittleness of the Kushman et al. (2014) approach on single-equation problems, yielding a 15% to 50% reduction in error. Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, Siena Dumas Ang |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Diagram Understanding in Geometry QuestionsabstractAutomatically solving geometry questions is a long-standing AI problem. A geometry question typically includes a textual description accompanied by a diagram. The first step in solving geometry questions is diagram understanding, which consists of identifying visual elements in the diagram, their locations, their geometric properties, and aligning them to corresponding textual descriptions. In this paper, we present a method for diagram understanding that identifies visual elements in a diagram while maximizing agreement between textual and visual data. We show that the method's objective function is submodular; thus we are able to introduce an efficient method for diagram understanding that is close to optimal. To empirically evaluate our method, we compile a new dataset of geometry questions (textual descriptions and diagrams) and compare with baselines that utilize standard vision techniques. Our experimental evaluation shows an F1 boost of more than 17% in identifying visual elements and 25% in aligning visual elements with their textual descriptions. Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni |
AAAI | 2 |
| 2014 | Learning to Solve Arithmetic Word Problems with Verb CategorizationabstractThis paper presents a novel approach to learning to solve simple arithmetic word problems.Our system, ARIS, analyzes each of the sentences in the problem statement to identify the relevant variables and their values.ARIS then maps this information into an equation that represents the problem, and enables its (trivial) solution as shown in Figure 1.The paper analyzes the arithmetic-word problems "genre", identifying seven categories of verbs used in such problems.ARIS learns to categorize verbs with 81.2% accuracy, and is able to solve 77.7% of the problems in a corpus of standard primary school test questions.We report the first learning results on this task without reliance on predefined templates and make our data publicly available. Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, Nate Kushman |
EMNLP | 2 |
| 2014 | Multi-Resolution Language Grounding with Weak SupervisionabstractLanguage is given meaning through its correspondence with a world representation.This correspondence can be at multiple levels of granularity or resolutions.In this paper, we introduce an approach to multi-resolution language grounding in the extremely challenging domain of professional soccer commentaries.We define and optimize a factored objective function that allows us to leverage discourse structure and the compositional nature of both language and game events.We show that finer resolution grounding helps coarser resolution grounding, and vice versa.Our method results in an F1 improvement of more than 48% versus the previous state of the art for fine-resolution grounding 1 . Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ali Farhadi |
EMNLP | 2 |
| 2014 | Multi-domain disfluency and repair detectionabstractThis paper investigates automatic detection of different types of self-repairs in spontaneous speech under different social contexts, from casual conversations to government hearings. The work shows that a simple CRF-based model is effective for cross-domain training, which is important for contexts where annotated data is not available. The approach explicitly represents common types of disfluencies observed in multi-domain data both in the model state space and the features extracted. In addition, the model incorporates an expanded state space for recognizing the repair structure, unlike prior work that annotates only the reparandum. Victoria Zayats, Mari Ostendorf, Hannaneh Hajishirzi |
INTERSPEECH | 3 |
| 2013 | Joint Coreference Resolution and Named-Entity Linking with Multi-Pass SievesabstractMany errors in coreference resolution come from semantic mismatches due to inadequate world knowledge.Errors in named-entity linking (NEL), on the other hand, are often caused by superficial modeling of entity context.This paper demonstrates that these two tasks are complementary.We introduce NECO, a new model for named entity linking and coreference resolution, which solves both problems jointly, reducing the errors made on each.NECO extends the Stanford deterministic coreference system by automatically linking mentions to Wikipedia and introducing new NEL-informed mention-merging sieves.Linking improves mention-detection and enables new semantic attributes to be incorporated from Freebase, while coreference provides better context modeling by propagating named-entity links within mention clusters.Experiments show consistent improvements across a number of datasets and experimental conditions, including over 11% reduction in MUC coreference error and nearly 21% reduction in F1 NEL error on ACE 2004 newswire data. Hannaneh Hajishirzi, Leila Zilles, Daniel S. Weld, Luke Zettlemoyer |
EMNLP | 1 |
| 2013 | Managing chaos: models of turn-taking in character-multichild interactionsabstractTurn-taking decisions in multiparty settings are complex, especially when the participants are children. Our goal is to endow an interactive character with appropriate turn-taking behavior using visual, audio and contextual features. To that end, we investigate three distinct turn-taking models: a baseline model grounded in established turn-taking rules for adults and two machine learning models, one trained with data collected in situ and the other trained with data collected in more controlled conditions. The three models are shown to have different profiles of behavior during silences, overlapping speech, and at the end of participants' turns. An exploratory user evaluation focusing on the decision points where the models differ showed clear preference for the machine learning models over the baseline model. The results indicate that the rules for language interactions with small groups of children are not simply an extension of the rules for interacting with small groups of adults. Iolanda Leite, Hannaneh Hajishirzi, Sean Andrist, Jill Fain Lehman |
ICMI | 2 |
| 2012 | Using Group History to Identify Character-Directed Utterances in Multi-Child Interactions
Hannaneh Hajishirzi, Jill Fain Lehman, Jessica K. Hodgins |
SIGDIAL Conference | 1 |
| 2012 | Semantic Understanding of Professional Soccer Commentaries
Hannaneh Hajishirzi, Mohammad Rastegari, Ali Farhadi, Jessica K. Hodgins |
UAI | 1 |
| 2011 | Reasoning about RoboCup Soccer Narratives
Hannaneh Hajishirzi, Julia Hockenmaier, Erik T. Mueller, Eyal Amir |
UAI | 1 |
| 2010 | Reasoning about Deterministic Actions with Probabilistic Prior and Application to Stochastic Filtering
Hannaneh Hajishirzi, Eyal Amir |
KR | 1 |
| 2010 | Adaptive near-duplicate detection via similarity learningabstractIn this paper, we present a novel near-duplicate document detection method that can easily be tuned for a particular domain. Our method represents each document as a real-valued sparse k-gram vector, where the weights are learned to optimize for a specified similarity function, such as the cosine similarity or the Jaccard coefficient. Near-duplicate documents can be reliably detected through this improved similarity measure. In addition, these vectors can be mapped to a small number of hash-values as document signatures through the locality sensitive hashing scheme for efficient similarity computation. We demonstrate our approach in two target domains: Web news articles and email messages. Our method is not only more accurate than the commonly used methods such as Shingles and I-Match, but also shows consistent improvement across the domains, which is a desired property lacked by existing methods. Hannaneh Hajishirzi, Scott Yih, Alek Kolcz |
SIGIR | 1 |
| 2009 | Greedy Algorithms for Sequential Sensing Decisions
Hannaneh Hajishirzi, Afsaneh Shirazi, Jaesik Choi, Eyal Amir |
IJCAI | 1 |
| 2008 | Sampling First Order Logical Particles
Hannaneh Hajishirzi, Eyal Amir |
UAI | 1 |
| 2007 | Stochastic Filtering in a Probabilistic Action Model
Hannaneh Hajishirzi, Eyal Amir |
AAAI | 1 |