EDBT 2026 Demo / reviewers in the wild / expert
Oyvind Tafjord
dblp:178/8640
· DBLP profile ↗
27ranked-venue papers
3as first author
18since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OLMoE: Open Mixture-of-Experts Language ModelsabstractWe introduce OLMoE, a fully open, state-of-the-art language model leveraging sparse Mixture-of-Experts (MoE). OLMoE-1B-7B has 7 billion (B) parameters but uses only 1B per input token. We pretrain it on 5 trillion tokens and further adapt it to create OLMoE-1B-7B-Instruct. Our models outperform all available models with similar active parameters, even surpassing larger ones like Llama2-13B-Chat and DeepSeekMoE-16B. We present novel findings on MoE training, define and analyze new routing properties showing high specialization in our model, and open-source all our work: model weights, training data, code, and logs. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Pete Walsh 0001, Oyvind Tafjord, Nathan Lambert 0001, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, Dave Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi |
ICLR | 9 |
| 2025 | From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-AnsweringabstractRecent reasoning methods (e.g., chain-of-thought) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM’s overall understanding, or “theory,” about the question’s topic, making it still hard to trust the model. Our goal is to materialize such theories - here called microtheories (a linguistic analog of logical microtheories) - as a set of sentences encapsulating an LM’s core knowledge about a topic. These statements systematically work together to entail answers to a set of questions to both engender trust and improve performance. Our approach is to first populate a knowledge store with (model-generated) sentences that entail answers to training questions, and then distill those down to a core microtheory which is concise, general, and non-redundant. We show that, when added to a general corpus (e.g., Wikipedia), microtheories can supply critical information not necessarily present in the corpus, improving both a model’s ability to ground its answers to verifiable knowledge (i.e., show how answers are systematically entailed by documents in the corpus, grounding up to +8% more answers), and the accuracy of those grounded answers (up to +8% absolute). We also show that, in a human evaluation in the medical domain, our distilled microtheories contain a significantly higher concentration of topically critical facts than the non-distilled knowledge store. Finally, we show we can quantify the coverage of a microtheory for a topic (characterized by a dataset) using a notion of p-relevance. Together, these suggest that microtheories are an efficient distillation of an LM’s topic-relevant knowledge, that they can usefully augment existing corpora, and can provide both performance gains and an interpretable, verifiable window into the model’s knowledge of a topic. Nathaniel Weir, Bhavana Dalvi, Orion Weller, Oyvind Tafjord, Sam Hornstein, Alexander Sabol, Peter A. Jansen, Benjamin Van Durme, Peter Clark |
ICLR | 4 |
| 2025 | Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice QuestionsabstractMultiple-choice question answering (MCQA) is a key competence of performant transformer language models that is tested by mainstream benchmarks. However, recent evidence shows that models can have quite a range of performance, particularly when the task format is diversified slightly (such as by shuffling answer choice order). In this work we ask: how do successful models perform formatted MCQA? We employ vocabulary projection and activation patching methods to localize key hidden states that encode relevant information for predicting the correct answer. We find that prediction of a specific answer symbol is causally attributed to a few middle layers, and specifically their multi-head self-attention mechanisms. We show that subsequent layers increase the probability of the predicted answer symbol in vocabulary space, and that this probability increase is associated with a sparse set of attention heads with unique roles. We additionally uncover differences in how different models adjust to alternative symbols. Finally, we demonstrate that a synthetic task can disentangle sources of model error to pinpoint when a model has learned formatted MCQA, and show that logit differences between answer choice tokens continue to grow over the course of training. Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi, Ashish Sabharwal |
ICLR | 2 |
| 2025 | DataDecide: How to Predict Best Pretraining Data with Small ExperimentsabstractBecause large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and methods of making decisions from observed performance at small scale most accurately predict the datasets that yield the best large models? To empower open exploration of this question, we release models, data, and evaluations in DataDecide—the most extensive open suite of models over differences in data and scale. We conduct controlled pretraining experiments across 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, model sizes up to 1B parameters, and 3 random seeds. We find that the ranking of models at a single, small size (e.g., 150M parameters) is a strong baseline for predicting best models at our larger target scale (1B) ($\tilde$ 80% of comparisons correct). No scaling law methods among 8 baselines exceed the compute-decision frontier of single-scale predictions, but DataDecide can measure improvement in future scaling laws. We also identify that using continuous likelihood metrics as proxies in small experiments makes benchmarks including MMLU, ARC, HellaSwag, MBPP, and HumanEval $>$ 80% predictable at the target 1B scale with just 0.01% of the compute. Ian Magnusson, Nguyen Tai, Ben Bogin, David Heineman, Jena D. Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu 0010, Dirk Groeneveld, Oyvind Tafjord, Noah A. Smith, Pang Wei Koh, Jesse Dodge |
ICML | 10 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 6 |
| 2024 | Digital Socrates: Evaluating LLMs through Explanation CritiquesabstractWhile LLMs can provide reasoned explanations along with their answers, the nature and quality of those explanations are still poorly understood.In response, our goal is to define a detailed way of characterizing the explanation capabilities of modern models and to create a nuanced, interpretable explanation evaluation tool that can generate such characterizations automatically, without relying on expensive API calls or human annotations.Our approach is to (a) define the new task of explanation critiquing -identifying and categorizing any main flaw in an explanation and providing suggestions to address the flaw, (b) create a sizeable, humanverified dataset for this task, and (c) train an open-source, automatic critique model (called Digital Socrates) using this data.Through quantitative and qualitative analysis, we demonstrate how Digital Socrates is useful for revealing insights about student models by examining their reasoning chains, and how it can provide highquality, nuanced, automatic evaluation of those model explanations for the first time.Digital Socrates thus fills an important gap in evaluation tools for understanding and improving the explanation behavior of models. Yuling Gu, Oyvind Tafjord, Peter Clark |
ACL (1) | 2 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 28 |
| 2024 | Enhancing Systematic Decompositional Natural Language Inference Using Informal LogicabstractNathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, Benjamin Van Durme. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nathaniel Weir, Kate Sanders 0002, Orion Weller, Shreya Sharma 0010, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi, Oyvind Tafjord, Peter A. Jansen, Peter Clark, Benjamin Van Durme |
EMNLP | 8 |
| 2024 | DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery AgentsabstractAutomated scientific discovery promises to accelerate progress across scientific domains, but evaluating an agent's capacity for end-to-end scientific reasoning is challenging as running real-world experiments is often prohibitively expensive or infeasible. In this work we introduce DiscoveryWorld, a virtual environment that enables benchmarking an agent's ability to perform complete cycles of novel scientific discovery in an inexpensive, simulated, multi-modal, long-horizon, and fictional setting.DiscoveryWorld consists of 24 scientific tasks across three levels of difficulty, each with parametric variations that provide new discoveries for agents to make across runs. Tasks require an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. Task difficulties are normed to range from straightforward to challenging for human scientists with advanced degrees. DiscoveryWorld further provides three automatic metrics for evaluating performance, including: (1) binary task completion, (2) fine-grained report cards detailing procedural scoring of task-relevant actions, and (3) the accuracy of discovered explanatory knowledge.While simulated environments such as DiscoveryWorld are low-fidelity compared to the real world, we find that strong baseline agents struggle on most DiscoveryWorld tasks, highlighting the utility of using simulated environments as proxy tasks for near-term development of scientific discovery competency in agents. Peter A. Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi, Bodhisattwa Prasad Majumder, Oyvind Tafjord, Peter Clark |
NeurIPS | 7 |
| 2024 | Paloma: A Benchmark for Evaluating Language Model FitabstractEvaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Pete Walsh 0001, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson 0001, Jesse Dodge |
NeurIPS | 6 |
| 2023 | Language Models with RationalityabstractWhile large language models (LLMs) are proficient at question-answering (QA), it is not always clear how (or even if) an answer follows from their latent "beliefs".This lack of interpretability is a growing impediment to widespread use of LLMs.To address this, our goals are to make model beliefs and their inferential relationships explicit, and to resolve inconsistencies that may exist, so that answers are supported by interpretable chains of reasoning drawn from a consistent network of beliefs.Our approach, which we call REFLEX, is to add a rational, self-reflecting layer on top of the LLM.First, given a question, we construct a belief graph using a backward-chaining process to materialize relevant model beliefs (including beliefs about answer candidates) and their inferential relationships.Second, we identify and minimize contradictions in that graph using a formal constraint reasoner.We find that REFLEX significantly improves consistency (by 8%-11% absolute) without harming overall answer accuracy, resulting in answers supported by faithful chains of reasoning drawn from a more consistent belief system.This suggests a new style of system architecture in which an LLM extended with a rational layer can provide an interpretable window into system beliefs, add a systematic reasoning capability, and repair latent inconsistencies present in the LLM. Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson 0001, Hinrich Schütze, Peter Clark |
EMNLP | 2 |
| 2023 | Increasing Probability Mass on Answer Choices Does Not Always Improve AccuracyabstractWhen pretrained language models (LMs) are applied to discriminative tasks such as multiplechoice questions, they place probability mass on vocabulary tokens that aren't among the given answer choices.Spreading probability mass across multiple surface forms with identical meaning (such as "bath" and "bathtub") is thought to cause an underestimation of a model's true performance, referred to as the "surface form competition" (SFC) hypothesis.This has motivated the introduction of various probability normalization methods.However, many core questions remain unanswered.How do we measure SFC? Are there direct ways of reducing it, and does doing so improve task performance?We propose a mathematical formalism for SFC which allows us to quantify and bound its impact for the first time.We identify a simple method for reducing it-namely, increasing probability mass on the given answer choices by a) including them in the prompt and b) using in-context learning with even just one example.We show this method eliminates the impact of SFC in the majority of instances.Our experiments on three diverse datasets and six LMs reveal several additional surprising findings.For example, both normalization and prompting methods for reducing SFC can be ineffective or even detrimental to task performance for some LMs.We conclude with practical insights for effectively prompting LMs for multiple-choice tasks. 1 * Work done at AI2. 1 Code available at https://github.com/allenai/ revisiting_surface_form_competition. Sarah Wiegreffe, Matthew Finlayson, Oyvind Tafjord, Peter Clark, Ashish Sabharwal |
EMNLP | 3 |
| 2022 | LILA: A Unified Benchmark for Mathematical ReasoningabstractSwaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan |
EMNLP | 8 |
| 2022 | Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System ImprovementabstractOur goal is a teachable reasoning system for question-answering (QA), where a user can interact with faithful answer explanations, and correct its errors so that the system improves over time.Our approach is to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections to erroneous model beliefs that users identify during interaction.Retrievals from memory are used as additional context for QA, to help avoid previous mistakes in similar new situationsa novel application of memory-based continuous learning.With simulated feedback, we find that our system (called TeachMe 1 ) continually improves with time, and without model retraining, requiring feedback on only 25% of training examples to reach within 1% of the upper-bound (feedback on all examples).Similarly, in experiments with real users, we observe a similar trend, with performance improving by over 15% on a hidden test set after teaching.This suggests new opportunities for using frozen language models in an interactive setting where users can inspect, debug, and correct the model's beliefs, leading to improved system's performance over time. Bhavana Dalvi, Oyvind Tafjord, Peter Clark |
EMNLP | 2 |
| 2022 | Entailer: Answering Questions with Faithful and Truthful Chains of ReasoningabstractOur goal is a question-answering (QA) system that can show how its answers are implied by its own internal beliefs via a systematic chain of reasoning.Such a capability would allow better understanding of why a model produced the answer it did.Our approach is to recursively combine a trained backward-chaining model, capable of generating a set of premises entailing an answer hypothesis, with a verifier that checks that the model itself believes those premises (and the entailment itself) through self-querying.To our knowledge, this is the first system to generate multistep chains that are both faithful (the answer follows from the reasoning) and truthful (the chain reflects the system's own internal beliefs).In evaluation using two different datasets, users judge that a majority (70%+) of generated chains clearly show how an answer follows from a set of facts -substantially better than a high-performance baseline -while preserving answer accuracy.By materializing model beliefs that systematically support an answer, new opportunities arise for understanding the model's system of belief, and diagnosing and correcting its misunderstandings when an answer is wrong. Oyvind Tafjord, Bhavana Dalvi, Peter Clark |
EMNLP | 1 |
| 2022 | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringabstractWhen answering a question, humans utilize the information available across different modalities to synthesize a consistent and complete chain of thought (CoT). This process is normally a black box in the case of deep learning models like large-scale language models. Recently, science question benchmarks have been used to diagnose the multi-hop reasoning ability and interpretability of an AI system. However, existing datasets fail to provide annotations for the answers, or are restricted to the textual-only modality, small scales, and limited domain diversity. To this end, we present Science Question Answering (ScienceQA), a new benchmark that consists of ~21k multimodal multiple choice questions with a diverse set of science topics and annotations of their answers with corresponding lectures and explanations. We further design language models to learn to generate lectures and explanations as the chain of thought (CoT) to mimic the multi-hop reasoning process when answering ScienceQA questions. ScienceQA demonstrates the utility of CoT in language models, as CoT improves the question answering performance by 1.20% in few-shot GPT-3 and 3.99% in fine-tuned UnifiedQA. We also explore the upper bound for models to leverage explanations by feeding those in the input; we observe that it improves the few-shot performance of GPT-3 by 18.96%. Our analysis further shows that language models, similar to humans, benefit from explanations to learn from fewer data and achieve the same performance with just 40% of the data. The data and code are available at https://scienceqa.github.io. Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 0001, Kai-Wei Chang 0001, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, Ashwin Kalyan |
NeurIPS | 7 |
| 2021 | Explaining Answers with Entailment TreesabstractBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, Peter Clark. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Bhavana Dalvi, Peter A. Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, Peter Clark |
EMNLP (1) | 3 |
| 2021 | BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of BeliefabstractAlthough pretrained language models (PTLMs) contain significant amounts of world knowledge, they can still produce inconsistent answers to questions when probed, even after specialized training.As a result, it can be hard to identify what the model actually "believes" about the world, making it susceptible to inconsistent behavior and simple errors.Our goal is to reduce these problems.Our approach is to embed a PTLM in a broader system that also includes an evolving, symbolic memory of beliefs -a BeliefBank -that records but then may modify the raw PTLM answers.We describe two mechanisms to improve belief consistency in the overall system.First, a reasoning component -a weighted MaxSAT solver -revises beliefs that significantly clash with others.Second, a feedback component issues future queries to the PTLM using known beliefs as context.We show that, in a controlled experimental setting, these two mechanisms result in more consistent beliefs in the overall system, improving both the accuracy and consistency of its answers over time.This is significant as it is a first step towards PTLM-based architectures with a systematic notion of belief, enabling them to construct a more coherent picture of the world, and improve over time without model retraining. Nora Kassner, Oyvind Tafjord, Hinrich Schütze, Peter Clark |
EMNLP (1) | 2 |
| 2020 | "You are grounded!": Latent Name Artifacts in Pre-trained Language ModelsabstractPre-trained language models (LMs) may perpetuate biases originating in their training corpus to downstream models.We focus on artifacts associated with the representation of given names (e.g., Donald), which, depending on the corpus, may be associated with specific entities, as indicated by next token prediction (e.g., Trump).While helpful in some contexts, grounding happens also in underspecified or inappropriate contexts.For example, endings generated for 'Donald is a' substantially differ from those of other names, and often have more-than-average negative sentiment.We demonstrate the potential effect on downstream tasks with reading comprehension probes where name perturbation changes the model answers.As a silver lining, our experiments suggest that additional pre-training on different corpora may mitigate this bias. ModelMain Corpus Type Gen. Cls.Named Entities from News Named Entities from History Model Minimal News History Infrml Avg Minimal News History Infrml Avg GPT 0.0 7.0 12.7 1.4 5.3 0.0 21.9 39.1 7.8 17.2 GPT2-small 22.5 63.4 50.7 15.5 38.0 12.5 29.7 56.2 12.5 27.7 GPT2-medium 33.8 64.8 49.3 12.7 40.2 21.9 32.8 62.5 4.7 30.5 GPT2-large 43.7 66.2 47.9 16.9 43.7 29.7 29.7 56.2 12.5 32.0 GPT2-XL 50.7 62.0 45.1 21.1 44.7 28.1 31.2 60.9 14.1 33.6 TransformerXL 14.1 18.3 15.5 12.7 15.2 35.9 43.8 51.6 37.5 42.2 XLNet-base 4.2 33.8 12.7 4.2 13.7 0.0 34.4 23.4 3.1 15.2 XLNet-large 11.3 40.8 23.9 9.9 21.5 6.2 29.7 31.2 7.8 18.7 Average 22.5 44.5 32.2 11.8 27.7 16.8 31.7 47.6 12.5 27.1 Vered Shwartz, Rachel Rudinger, Oyvind Tafjord |
EMNLP (1) | 3 |
| 2020 | Transformers as Soft Reasoners over LanguageabstractBeginning with McCarthy's Advice Taker (1959), AI has pursued the goal of providing a system with explicit, general knowledge and having the system reason over that knowledge. However, expressing the knowledge in a formal (logical or probabilistic) representation has been a major obstacle to this research. This paper investigates a modern approach to this problem where the facts and rules are provided as natural language sentences, thus bypassing a formal representation. We train transformers to reason (or emulate reasoning) over these sentences using synthetically generated data. Our models, that we call RuleTakers, provide the first empirical demonstration that this kind of soft reasoning over language is learnable, can achieve high (99%) accuracy, and generalizes to test data requiring substantially deeper chaining than seen during training (95%+ scores). We also demonstrate that the models transfer well to two hand-authored rulebases, and to rulebases paraphrased into more natural language. These findings are significant as it suggests a new role for transformers, namely as limited "soft theorem provers" operating over explicit theories in language. This in turn suggests new possibilities for explainability, correctability, and counterfactual reasoning in question-answering. Peter Clark, Oyvind Tafjord, Kyle Richardson 0001 |
IJCAI | 2 |
| 2020 | Multi-class Hierarchical Question Classification for Multiple Choice Science ExamsabstractPrior work has demonstrated that question classification (QC), recognizing the problem domain of a question, can help answer it more accurately. However, developing strong QC algorithms has been hindered by the limited size and complexity of annotated data available. To address this, we present the largest challenge dataset for QC, containing 7,787 science exam questions paired with detailed classification labels from a fine-grained hierarchical taxonomy of 406 problem domains. We then show that a BERT-based model trained on this dataset achieves a large (+0.12 MAP) gain compared with previous methods, while also achieving state-of-the-art performance on benchmark open-domain and biomedical QC datasets. Finally, we show that using this model’s predictions of question topic significantly improves the accuracy of a question answering system by +1.7% P@1, with substantial future gains possible as QC performance improves. Dongfang Xu, Peter A. Jansen, Jaycie Martin, Zhengnan Xie, Vikas Yadav, Harish Tayyar Madabushi, Oyvind Tafjord, Peter Clark |
LREC | 7 |
| 2020 | Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit KnowledgeabstractTo what extent can a neural network systematically reason over symbolic facts? Evidence suggests that large pre-trained language models (LMs) acquire some reasoning capacity, but this ability is difficult to control. Recently, it has been shown that Transformer-based models succeed in consistent reasoning over explicit symbolic facts, under a "closed-world" assumption. However, in an open-domain setup, it is desirable to tap into the vast reservoir of implicit knowledge already encoded in the parameters of pre-trained LMs. In this work, we provide a first demonstration that LMs can be trained to reliably perform systematic reasoning combining both implicit, pre-trained knowledge and explicit natural language statements. To do this, we describe a procedure for automatically generating datasets that teach a model new reasoning skills, and demonstrate that models learn to effectively perform inference which involves implicit taxonomic and world knowledge, chaining and counting. Finally, we show that "teaching" models to reason generalizes beyond the training distribution: they successfully compose the usage of multiple reasoning skills in single examples. Our work paves a path towards open-domain systems that constantly improve by interacting with users who can instantly correct a model by adding simple natural language statements. Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, Jonathan Berant |
NeurIPS | 2 |
| 2019 | Declarative Question Answering over Knowledge Bases Containing Natural Language Text with Answer Set ProgrammingabstractWhile in recent years machine learning (ML) based approaches have been the popular approach in developing endto-end question answering systems, such systems often struggle when additional knowledge is needed to correctly answer the questions. Proposed alternatives involve translating the question and the natural language text to a logical representation and then use logical reasoning. However, this alternative falters when the size of the text gets bigger. To address this we propose an approach that does logical reasoning over premises written in natural language text. The proposed method uses recent features of Answer Set Programming (ASP) to call external NLP modules (which may be based on ML) which perform simple textual entailment. To test our approach we develop a corpus based on the life cycle questions and showed that Our system achieves up to 18% performance gain when compared to standard MCQ solvers. Arindam Mitra, Peter Clark, Oyvind Tafjord, Chitta Baral |
AAAI | 3 |
| 2019 | QUAREL: A Dataset and Models for Answering Questions about Qualitative RelationshipsabstractMany natural la guage questions require recognizing and reasoning with qualitative relationships (e.g., in science, economics, and medicine), but are challenging to answer with corpus-based methods. Qualitative modeling provides tools that support such reasoning, but the semantic parsing task of mapping questions into those models has formidable challenges. We present QUAREL, a dataset of diverse story questions involving qualitative relationships that characterize these challenges, and techniques that begin to address them. The dataset has 2771 questions relating 19 different types of quantities. For example, “Jenny observes that the robot vacuum cleaner moves slower on the living room carpet than on the bedroom carpet. Which carpet has more friction?” We contribute (1) a simple and flexible conceptual framework for representing these kinds of questions; (2) the QUAREL dataset, including logical forms, exemplifying the parsing challenges; and (3) two novel models for this task, built as extensions of type-constrained semantic parsing. The first of these models (called QUASP+) significantly outperforms off-the-shelf tools on QUAREL. The second (QUASP+ZERO) demonstrates zero-shot capability, i.e., the ability to handle new qualitative relationships without requiring additional training data, something not possible with previous models. This work thus makes inroads into answering complex, qualitative questions that require reasoning, and scaling to new relationships at low cost. The dataset and models are available at http://data.allenai.org/quarel. Oyvind Tafjord, Peter Clark, Matt Gardner 0001, Scott Yih, Ashish Sabharwal |
AAAI | 1 |
| 2019 | QuaRTz: An Open-Domain Dataset of Qualitative Relationship QuestionsabstractOyvind Tafjord, Matt Gardner, Kevin Lin, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Oyvind Tafjord, Matt Gardner 0001, Peter Clark |
EMNLP/IJCNLP (1) | 1 |
| 2016 | Combining Retrieval, Statistics, and Inference to Answer Elementary Science QuestionsabstractWhat capabilities are required for an AI system to pass standard 4th Grade Science Tests? Previous work has examined the use of Markov Logic Networks (MLNs) to represent the requisite background knowledge and interpret test questions, but did not improve upon an information retrieval (IR) baseline. In this paper, we describe an alternative approach that operates at three levels of representation and reasoning: information retrieval, corpus statistics, and simple inference over a semi-automatically constructed knowledge base, to achieve substantially improved results. We evaluate the methods on six years of unseen, unedited exam questions from the NY Regents Science Exam (using only non-diagram, multiple choice questions), and show that our overall system’s score is 71.3%, an improvement of 23.8% (absolute) over the MLN-based method described in previous work. We conclude with a detailed analysis, illustrating the complementary strengths of each method in the ensemble. Our datasets are being released to enable further research. Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter D. Turney, Daniel Khashabi |
AAAI | 5 |
| 2016 | Semantic Parsing to Probabilistic Programs for Situated Question AnsweringabstractSituated question answering is the problem of answering questions about an environment such as an image or diagram.This problem requires jointly interpreting a question and an environment using background knowledge to select the correct answer.We present Parsing to Probabilistic Programs (P 3 ), a novel situated question answering model that can use background knowledge and global features of the question/environment interpretation while retaining efficient approximate inference.Our key insight is to treat semantic parses as probabilistic programs that execute nondeterministically and whose possible executions represent environmental uncertainty.We evaluate our approach on a new, publicly-released data set of 5000 science diagram questions, outperforming several competitive classical and neural baselines. Jayant Krishnamurthy, Oyvind Tafjord, Aniruddha Kembhavi |
EMNLP | 2 |