EDBT 2026 Demo / reviewers in the wild / expert
Peter Clark
dblp:34/1184
· DBLP profile ↗
103ranked-venue papers
15as first author
43since 2021 · last 2026
0000-0003-1001-9226ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 97 · 12 first-author · 43 since 2021Databases, data management, data science and information retrieval · 12 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 1 since 2021Theory of computation · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Literature-Driven Scientific Theories at ScaleabstractContemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theory building remain underexplored.In this work, we formulate the problem of synthesizing theories consisting of qualitative and quantitative laws from large corpora of scientific literature.We study theory generation at scale, using 13.7k source papers to synthesize 2.9k theories, examining how generation using literaturegrounding versus parametric knowledge, and accuracy-focused versus novelty-focused generation objectives change theory properties.Our experiments show that, compared to using parametric LLM memory for generation, our literature-supported method creates theories that are significantly better at both matching existing evidence and at predicting future results from 4.6k subsequently-written papers. 1 Peter A. Jansen, Peter Clark, Doug Downey, Daniel S. Weld |
ACL (1) | 2 |
| 2025 | HypER: Literature-grounded Hypothesis Generation and Distillation with ProvenanceabstractLarge Language models have demonstrated promising performance in research ideation across scientific domains.Hypothesis development, the process of generating a highly specific declarative statement connecting a research idea with empirical validation, has received relatively less attention.Existing approaches trivially deploy retrieval augmentation and focus only on the quality of the final output ignoring the underlying reasoning process behind ideation.We present HypER (Hypothesis Generation with Explanation and Reasoning), a small language model (SLM) trained for literature-guided reasoning and evidence-based hypothesis generation.HypER is trained in a multi-task setting to discriminate between valid and invalid scientific reasoning chains in presence of controlled distractions.We find that HypER outperforms the base model, distinguishing valid from invalid reasoning chains (+22% average absolute F1), generates better evidence-grounded hypotheses (0.327 vs. 0.305 base model) with high feasibility and impact as judged by human experts (>3.5 on 5-point Likert scale).Resource at . Example of a valid reasoning chainTitle: Evidence suggesting that a chronic disease self-management program can improve health status while reducing hospitalization Abstract: This study evaluated the effectiveness (changes in health behaviors, health status, and health service utilization) of a self-management program for chronic disease ... Rosni Vasu, Chandrayee Basu, Bhavana Dalvi, Cristina Sarasua, Peter Clark, Abraham Bernstein |
EMNLP | 5 |
| 2025 | DiscoveryBench: Towards Data-Driven Discovery with Large Language ModelsabstractCan the rapid advances in code generation, function calling, and data analysis using large language models (LLMs) help automate the search and verification of hypotheses purely from a set of provided datasets? To evaluate this question, we present DiscoveryBench, the first comprehensive benchmark that formalizes the multi-step process of data-driven discovery. The benchmark is designed to systematically assess current model capabilities in discovery tasks and provide a useful resource for improving them. Our benchmark contains 264 tasks collected across 6 diverse domains, such as sociology and engineering, by manually deriving discovery workflows from published papers to approximate the real-world challenges faced by researchers, where each task is defined by a dataset, its metadata, and a discovery goal in natural language. We additionally provide 903 synthetic tasks to conduct controlled evaluations on data-driven workflows that are not covered in the manually collected split. Furthermore, our structured formalism of data-driven discovery enables a facet-based evaluation that provides useful insights into different failure modes. We evaluate several popular LLM-based reasoning frameworks using both open and closed LLMs as baselines on DiscoveryBench and find that even the best system scores only 25%. Our benchmark, thus, illustrates the challenges in autonomous data-driven discovery and serves as a valuable resource for the community to make progress. Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal 0003, Bhavana Dalvi, Abhijeetsingh Meena, Aryan Prakhar, Tirth Vora, Tushar Khot, Ashish Sabharwal, Peter Clark |
ICLR | 10 |
| 2025 | From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-AnsweringabstractRecent reasoning methods (e.g., chain-of-thought) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM’s overall understanding, or “theory,” about the question’s topic, making it still hard to trust the model. Our goal is to materialize such theories - here called microtheories (a linguistic analog of logical microtheories) - as a set of sentences encapsulating an LM’s core knowledge about a topic. These statements systematically work together to entail answers to a set of questions to both engender trust and improve performance. Our approach is to first populate a knowledge store with (model-generated) sentences that entail answers to training questions, and then distill those down to a core microtheory which is concise, general, and non-redundant. We show that, when added to a general corpus (e.g., Wikipedia), microtheories can supply critical information not necessarily present in the corpus, improving both a model’s ability to ground its answers to verifiable knowledge (i.e., show how answers are systematically entailed by documents in the corpus, grounding up to +8% more answers), and the accuracy of those grounded answers (up to +8% absolute). We also show that, in a human evaluation in the medical domain, our distilled microtheories contain a significantly higher concentration of topically critical facts than the non-distilled knowledge store. Finally, we show we can quantify the coverage of a microtheory for a topic (characterized by a dataset) using a notion of p-relevance. Together, these suggest that microtheories are an efficient distillation of an LM’s topic-relevant knowledge, that they can usefully augment existing corpora, and can provide both performance gains and an interpretable, verifiable window into the model’s knowledge of a topic. Nathaniel Weir, Bhavana Dalvi, Orion Weller, Oyvind Tafjord, Sam Hornstein, Alexander Sabol, Peter A. Jansen, Benjamin Van Durme, Peter Clark |
ICLR | 9 |
| 2025 | ZebraLogic: On the Scaling Limits of LLMs for Logical ReasoningabstractWe investigate the logical reasoning capabilities of Large Language Models (LLMs) and their scalability across complex deductive tasks. Using ZebraLogic, a newly developed benchmark dataset of logic grid puzzles derived from constraint satisfaction problems (CSPs), we systematically evaluate LLM performance. ZebraLogic spans a broad range of search space complexities and incorporates diverse logical constraints, providing a controlled environment to assess reasoning abilities. Our results reveal a significant decline in accuracy as problem complexity increases—a phenomenon we term the “curse of complexity.” Notably, this limitation persists even with scaling model size and inference-time computation, suggesting fundamental constraints in current LLM reasoning capabilities. Additionally, we explore strategies such as Best-of-N sampling, backtracking mechanisms, and self-verification prompts to enhance logical reasoning performance. Our findings provide critical insights into the scaling behavior of LLMs, highlight their limitations, and outline potential directions for advancing their reasoning capabilities. Bill Y. Lin, Ronan Le Bras 0001, Kyle Richardson 0001, Ashish Sabharwal, Radha Poovendran, Peter Clark, Yejin Choi 0001 |
ICML | 6 |
| 2025 | Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task SupervisionabstractZhouhang Xie, Tushar Khot, Bhavana Dalvi Mishra, Harshit Surana, Julian McAuley, Peter Clark, Bodhisattwa Prasad Majumder. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhouhang Xie, Tushar Khot, Bhavana Dalvi, Harshit Surana, Julian J. McAuley, Peter Clark, Bodhisattwa Prasad Majumder |
NAACL (Long Papers) | 6 |
| 2025 | AutoDiscovery: Open-ended Scientific Discovery via Bayesian SurpriseabstractThe promise of autonomous scientific discovery (ASD) hinges not only on answering questions, but also on knowing which questions to ask. Most recent works in ASD explore the use of large language models (LLMs) in goal-driven settings, relying on human-specified research questions to guide hypothesis generation. However, scientific discovery may be accelerated further by allowing the AI system to drive exploration by its own criteria. The few existing approaches in open-ended ASD select hypotheses based on diversity heuristics or subjective proxies for human interestingness, but the former struggles to meaningfully navigate the typically vast hypothesis space, and the latter suffers from imprecise definitions. This paper presents AutoDiscovery—a method for open-ended ASD that instead drives scientific exploration using Bayesian surprise. Here, we quantify the epistemic shift from the LLM’s prior beliefs about a hypothesis to its posterior beliefs after gathering experimental results. To efficiently explore the space of nested hypotheses, our method employs a Monte Carlo tree search (MCTS) strategy with progressive widening using surprisal as the reward function. We evaluate AutoDiscovery in the setting of data-driven discovery across 21 real-world datasets spanning domains such as biology, economics, finance, and behavioral science. Our results demonstrate that under a fixed budget, AutoDiscovery substantially outperforms competitors by producing 5-29% more discoveries deemed surprising by the LLM. Our human evaluation further reveals that two-thirds of discoveries made by our system are surprising to domain experts as well, suggesting this is an important step towards building open-ended ASD systems. Dhruv Agarwal 0003, Bodhisattwa Prasad Majumder, Reece Adamson, Megha Chakravorty, Satvika Reddy Gavireddy, Aditya Parashar, Harshit Surana, Bhavana Dalvi, Andrew McCallum, Ashish Sabharwal, Peter Clark |
NeurIPS | 11 |
| 2025 | Language Modeling by Language Modelsabstract*Can we leverage LLMs to model the process of discovering novel language model (LM) architectures?* Inspired by real research, we propose a multi-agent LLM approach that simulates the conventional stages of research, from ideation and literature search (proposal stage) to design implementation (code generation), generative pre-training, and downstream evaluation (verification). Using ideas from scaling laws, our system *Genesys* employs a *Ladder of Scales* approach; new designs are proposed, adversarially reviewed, implemented, and selectively verified at increasingly larger model scales (14M$\sim$350M parameters) with a narrowing budget (the number of models we can train at each scale). To help make discovery efficient and factorizable, Genesys uses a novel genetic programming backbone, which we show has empirical advantages over commonly used direct prompt generation workflows (e.g., $\sim$86\% percentage point improvement in successful design generation, a key bottleneck). We report experiments involving 1,162 newly discovered designs (1,062 fully verified) and find the best designs to be competitive with known architectures (e.g., outperform GPT2, Mamba2, etc., on 6/9 common benchmarks). We couple these results with comprehensive system-level ablations and formal results, which give broader insights into the design of effective autonomous discovery systems. Junyan Cheng, Peter Clark, Kyle Richardson 0001 |
NeurIPS | 2 |
| 2024 | Digital Socrates: Evaluating LLMs through Explanation CritiquesabstractWhile LLMs can provide reasoned explanations along with their answers, the nature and quality of those explanations are still poorly understood.In response, our goal is to define a detailed way of characterizing the explanation capabilities of modern models and to create a nuanced, interpretable explanation evaluation tool that can generate such characterizations automatically, without relying on expensive API calls or human annotations.Our approach is to (a) define the new task of explanation critiquing -identifying and categorizing any main flaw in an explanation and providing suggestions to address the flaw, (b) create a sizeable, humanverified dataset for this task, and (c) train an open-source, automatic critique model (called Digital Socrates) using this data.Through quantitative and qualitative analysis, we demonstrate how Digital Socrates is useful for revealing insights about student models by examining their reasoning chains, and how it can provide highquality, nuanced, automatic evaluation of those model explanations for the first time.Digital Socrates thus fills an important gap in evaluation tools for understanding and improving the explanation behavior of models. Yuling Gu, Oyvind Tafjord, Peter Clark |
ACL (1) | 3 |
| 2024 | The Unreasonable Effectiveness of Easy Training Data for Hard TasksabstractHow can we train models to perform well on hard test data when hard training data is by definition difficult to label correctly?This question has been termed the scalable oversight problem and has drawn increasing attention as language models have continually improved.In this paper, we present the surprising conclusion that current pretrained language models often generalize relatively well from easy to hard data, even performing as well as oracle models finetuned on hard data.We demonstrate this kind of easy-to-hard generalization using simple finetuning methods like in-context learning, linear classifier heads, and QLoRA for seven different measures of datapoint hardness, including six empirically diverse human hardness measures (like grade level) and one model-based measure (loss-based).Furthermore, we show that even if one cares most about model performance on hard data, it can be better to collect easy data rather than hard data for finetuning, since hard data is generally noisier and costlier to collect.Our experiments use open models up to 70b in size and four publicly available question-answering datasets with questions ranging in difficulty from 3rd grade science questions to college level STEM questions and general-knowledge trivia.We conclude that easy-to-hard generalization in LMs is surprisingly strong for the tasks studied.1Test Input LM Generated Answer Q: John hires a driving service to get him to work each day.His work is 30 miles away and he has to go there and back each day.He goes to work 5 days a week for 50 weeks a year.He gets charged $2 per mile driven and he also gives his driver a $150 bonus per month How much does he pay a year for driving?A: John goes to work 5 days a week for 50 weeks a year.John goes to work 5 x 50 = <<5*50=250>>250 times a year.John pays $2 per mile driven.John pays $2 x 30 x 2 = <<2*30*2=120>>120 for each trip.John pays $120 x 250 = <<120*250=30000>>30000 for each year.John pays $150 bonus per month.John pays $150 x 12 = <<150*12=1800>>1800 for each year.John pays $30000 + $1800 = <<30000+1800=31800>>31800 for each year. So the answer is 31800Correct on hard problem given only easy data in prompt High-school supervision improves college performance Correct! Peter Hase, Mohit Bansal, Peter Clark, Sarah Wiegreffe |
ACL (1) | 3 |
| 2024 | SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research RepositoriesabstractBen Bogin, Kejuan Yang, Shashank Gupta, Kyle Richardson, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ben Bogin, Kejuan Yang, Kyle Richardson 0001, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot |
EMNLP | 6 |
| 2024 | Enhancing Systematic Decompositional Natural Language Inference Using Informal LogicabstractNathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, Benjamin Van Durme. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nathaniel Weir, Kate Sanders 0002, Orion Weller, Shreya Sharma 0010, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi, Oyvind Tafjord, Peter A. Jansen, Peter Clark, Benjamin Van Durme |
EMNLP | 10 |
| 2024 | Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsabstractRecent works have showcased the ability of large-scale language models (LLMs) to embody diverse personas in their responses, exemplified by prompts like ‘_You are Yoda. Explain the Theory of Relativity._’ While this ability allows personalization of LLMs and enables human behavior simulation, its effect on LLMs’ capabilities remains unclear. To fill this gap, we present the first extensive study of the unintended side-effects of persona assignment on the ability of LLMs to perform _basic reasoning tasks_. Our study covers 24 reasoning datasets (spanning mathematics, law, medicine, morals, and more), 4 LLMs (2 versions of ChatGPT-3.5, GPT-4-Turbo, and Llama-2-70b-chat), and 19 diverse personas (e.g., ‘an Asian person’) spanning 5 socio-demographic groups: race, gender, religion, disability, and political affiliation. Our experiments unveil that LLMs harbor deep rooted bias against various socio-demographics underneath a veneer of fairness. While they overtly reject stereotypes when explicitly asked (‘_Are Black people less skilled at mathematics?_’), they manifest stereotypical and often erroneous presumptions when prompted to answer questions while adopting a persona. These can be observed as abstentions in the model’s response, e.g., ‘_As a Black person, I am unable to answer this question as it requires math knowledge_’, and generally result in a substantial drop in performance on reasoning tasks. Our experiments with ChatGPT-3.5 show that this bias is _ubiquitous_—80% of our personas demonstrate bias; it is _significant_—some datasets show performance drops of 70%+; and can be especially _harmful for certain groups_—some personas suffer statistically significant drops on 80%+ of the datasets. Overall, all four LLMs exhibit persona-induced bias to varying extents, with GPT-4-Turbo showing the least but still a problematic amount of bias (evident in 42% of the personas). Further analysis shows that these persona-induced errors can be hard-to-discern as they do not always manifest as explicit abstentions, and can also be hard-to-avoid—we find de-biasing prompts to have minimal to no effect. Our findings serve as a cautionary tale that the practice of assigning personas to LLMs—a trend on the rise—can surface their deep-rooted biases and have unforeseeable and detrimental side-effects. Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, Tushar Khot |
ICLR | 5 |
| 2024 | Position: Data-driven Discovery with Large Generative ModelsabstractWith the accumulation of data at an unprecedented rate, its potential to fuel scientific discovery is growing exponentially. This position paper urges the Machine Learning (ML) community to exploit the capabilities of large generative models (LGMs) to develop automated systems for end-to-end data-driven discovery—a paradigm encompassing the search and verification of hypotheses purely from a set of provided datasets, without the need for additional data collection or physical experiments. We first outline several desiderata for an ideal data-driven discovery system. Then, through DataVoyager, a proof-of-concept utilizing GPT-4, we demonstrate how LGMs fulfill several of these desiderata—a feat previously unattainable—while also highlighting important limitations in the current system that open up opportunities for novel ML research. We contend that achieving accurate, reliable, and robust end-to-end discovery systems solely through the current capabilities of LGMs is challenging. We instead advocate for fail-proof tool integration, along with active user moderation through feedback mechanisms, to foster data-driven scientific discoveries with efficiency and reproducibility. Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal 0003, Sanchaita Hazra, Ashish Sabharwal, Peter Clark |
ICML | 6 |
| 2024 | Skill Set Optimization: Reinforcing Language Model Behavior via Transferable SkillsabstractLarge language models (LLMs) have recently been used for sequential decision making in interactive environments. However, leveraging environment reward signals for continual LLM actor improvement is not straightforward. We propose Skill Set Optimization (SSO) for improving LLM actor performance through constructing and refining sets of transferable skills. SSO constructs skills by extracting common subtrajectories with high rewards and generating subgoals and instructions to represent each skill. These skills are provided to the LLM actor in-context to reinforce behaviors with high rewards. Then, SSO further refines the skill set by pruning skills that do not continue to result in high rewards. We evaluate our method in the classic videogame NetHack and the text environment ScienceWorld to demonstrate SSO's ability to optimize a set of skills and perform in-context policy improvement. SSO outperforms baselines by 40% in our custom NetHack task and outperforms the previous state-of-the-art in ScienceWorld by 35%. Kolby Nottingham, Bodhisattwa Prasad Majumder, Bhavana Dalvi, Sameer Singh 0001, Peter Clark, Roy Fox |
ICML | 5 |
| 2024 | NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
Nathaniel Weir, Peter Clark, Benjamin Van Durme |
IJCAI | 2 |
| 2024 | Leveraging Code to Improve In-Context Learning for Semantic ParsingabstractBen Bogin, Shivanshu Gupta, Peter Clark, Ashish Sabharwal. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ben Bogin, Shivanshu Gupta, Peter Clark, Ashish Sabharwal |
NAACL-HLT | 3 |
| 2024 | QualEval: Qualitative Evaluation for Model ImprovementabstractVishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Vishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan |
NAACL-HLT | 3 |
| 2024 | DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery AgentsabstractAutomated scientific discovery promises to accelerate progress across scientific domains, but evaluating an agent's capacity for end-to-end scientific reasoning is challenging as running real-world experiments is often prohibitively expensive or infeasible. In this work we introduce DiscoveryWorld, a virtual environment that enables benchmarking an agent's ability to perform complete cycles of novel scientific discovery in an inexpensive, simulated, multi-modal, long-horizon, and fictional setting.DiscoveryWorld consists of 24 scientific tasks across three levels of difficulty, each with parametric variations that provide new discoveries for agents to make across runs. Tasks require an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. Task difficulties are normed to range from straightforward to challenging for human scientists with advanced degrees. DiscoveryWorld further provides three automatic metrics for evaluating performance, including: (1) binary task completion, (2) fine-grained report cards detailing procedural scoring of task-relevant actions, and (3) the accuracy of discovered explanatory knowledge.While simulated environments such as DiscoveryWorld are low-fidelity compared to the real world, we find that strong baseline agents struggle on most DiscoveryWorld tasks, highlighting the utility of using simulated environments as proxy tasks for near-term development of scientific discovery competency in agents. Peter A. Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi, Bodhisattwa Prasad Majumder, Oyvind Tafjord, Peter Clark |
NeurIPS | 8 |
| 2024 | Learning to Reason via Program Generation, Emulation, and SearchabstractProgram synthesis with language models (LMs) has unlocked a large set of reasoning abilities; code-tuned LMs have proven adept at generating programs that solve a wide variety of algorithmic symbolic manipulation tasks (e.g. word concatenation). However, not all reasoning tasks are easily expressible as code, e.g. tasks involving commonsense reasoning, moral decision-making, and sarcasm understanding. Our goal is to extend a LM’s program synthesis skills to such tasks and evaluate the results via pseudo-programs, namely Python programs where some leaf function calls are left undefined. To that end, we propose, Code Generation and Emulated EXecution (COGEX). COGEX works by (1) training LMs to generate pseudo-programs and (2) teaching them to emulate their generated program’s execution, including those leaf functions, allowing the LM’s knowledge to fill in the execution gaps; and (3) using them to search over many programs to find an optimal one. To adapt the COGEX model to a new task, we introduce a method for performing program search to find a single program whose pseudo-execution yields optimal performance when applied to all the instances of a given dataset. We show that our approach yields large improvements compared to standard in-context learning approaches on a battery of tasks, both algorithmic and soft reasoning. This result thus demonstrates that code synthesis can be applied to a much broader class of problems than previously considered. Nathaniel Weir, Muhammad Khalifa, Linlu Qiu, Orion Weller, Peter Clark |
NeurIPS | 5 |
| 2023 | RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model OutputsabstractAfra Feyza Akyurek, Ekin Akyurek, Ashwin Kalyan, Peter Clark, Derry Tanti Wijaya, Niket Tandon. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Afra Feyza Akyürek, Ekin Akyürek, Ashwin Kalyan, Peter Clark, Derry Wijaya, Niket Tandon |
ACL (1) | 4 |
| 2023 | Do language models have coherent mental models of everyday things?abstractWhen people think of everyday things like an egg, they typically have a mental image associated with it.This allows them to correctly judge, for example, that "the yolk surrounds the shell" is a false statement.Do language models similarly have a coherent picture of such everyday things?To investigate this, we propose a benchmark dataset consisting of 100 everyday things, their parts, and the relationships between these parts, expressed as 11,720 "X relation Y?" true/false questions.Using these questions as probes, we observe that state-ofthe-art pre-trained language models (LMs) like GPT-3 and Macaw have fragments of knowledge about these everyday things, but do not have fully coherent "parts mental models" (54-59% accurate, 19-43% conditional constraint violation).We propose an extension where we add a constraint satisfaction layer on top of the LM's raw predictions to apply commonsense constraints.As well as removing inconsistencies, we find that this also significantly improves accuracy (by 16-20%), suggesting how the incoherence of the LM's pictures of everyday things can be significantly reduced. 1 Yuling Gu, Bhavana Dalvi, Peter Clark |
ACL (1) | 3 |
| 2023 | IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsabstractAlthough counterfactual reasoning is a fundamental aspect of intelligence, the lack of largescale counterfactual open-domain questionanswering (QA) benchmarks makes it difficult to evaluate and improve models on this ability.To address this void, we introduce the first such dataset, named IfQA, where each question is based on a counterfactual presupposition via an "if" clause.Such questions require models to go beyond retrieving direct factual knowledge from the Web: they must identify the right information to retrieve and reason about an imagined situation that may even go against the facts built into their parameters.The IfQA dataset contains 3,800 questions that were annotated by crowdworkers on relevant Wikipedia passages.Empirical analysis reveals that the IfQA dataset is highly challenging for existing open-domain QA methods, including supervised retrieve-then-read pipeline methods (F1 score 44.5), as well as recent few-shot approaches such as chain-of-thought prompting with ChatGPT (F1 score 57.2).We hope the unique challenges posed by IfQA will push open-domain QA research on both retrieval and reasoning fronts, while also helping endow counterfactual reasoning abilities to today's language understanding models.The IfQA dataset can be found and downloaded at https://allenai.org/data/ifqa. Wenhao Yu 0002, Meng Jiang 0001, Peter Clark, Ashish Sabharwal |
EMNLP | 3 |
| 2023 | Language Models with RationalityabstractWhile large language models (LLMs) are proficient at question-answering (QA), it is not always clear how (or even if) an answer follows from their latent "beliefs".This lack of interpretability is a growing impediment to widespread use of LLMs.To address this, our goals are to make model beliefs and their inferential relationships explicit, and to resolve inconsistencies that may exist, so that answers are supported by interpretable chains of reasoning drawn from a consistent network of beliefs.Our approach, which we call REFLEX, is to add a rational, self-reflecting layer on top of the LLM.First, given a question, we construct a belief graph using a backward-chaining process to materialize relevant model beliefs (including beliefs about answer candidates) and their inferential relationships.Second, we identify and minimize contradictions in that graph using a formal constraint reasoner.We find that REFLEX significantly improves consistency (by 8%-11% absolute) without harming overall answer accuracy, resulting in answers supported by faithful chains of reasoning drawn from a more consistent belief system.This suggests a new style of system architecture in which an LLM extended with a rational layer can provide an interpretable window into system beliefs, add a systematic reasoning capability, and repair latent inconsistencies present in the LLM. Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson 0001, Hinrich Schütze, Peter Clark |
EMNLP | 6 |
| 2023 | Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise GenerationabstractIn this paper, we present a novel approach for distilling math word problem solving capabilities from large language models (LLMs) into smaller, more efficient student models. Our approach is designed to consider the student model's weaknesses and foster a tailored learning experience by generating targeted exercises aligned with educational science principles, such as knowledge tracing and personalized learning. Concretely, we let GPT-3 be a math tutor and run two steps iteratively: 1) assessing the student model's current learning status on a GPT-generated exercise book, and 2) improving the student model by training it with tailored exercise samples generated by GPT-3. Experimental results reveal that our approach outperforms LLMs (e.g., GPT-3 and PaLM) in accuracy across three distinct benchmarks while employing significantly fewer parameters. Furthermore, we provide a comprehensive analysis of the various components within our methodology to substantiate their efficacy. Zhenwen Liang, Wenhao Yu 0002, Tanmay Rajpurohit, Peter Clark, Xiangliang Zhang 0001, Ashwin Kalyan |
EMNLP | 4 |
| 2023 | Increasing Probability Mass on Answer Choices Does Not Always Improve AccuracyabstractWhen pretrained language models (LMs) are applied to discriminative tasks such as multiplechoice questions, they place probability mass on vocabulary tokens that aren't among the given answer choices.Spreading probability mass across multiple surface forms with identical meaning (such as "bath" and "bathtub") is thought to cause an underestimation of a model's true performance, referred to as the "surface form competition" (SFC) hypothesis.This has motivated the introduction of various probability normalization methods.However, many core questions remain unanswered.How do we measure SFC? Are there direct ways of reducing it, and does doing so improve task performance?We propose a mathematical formalism for SFC which allows us to quantify and bound its impact for the first time.We identify a simple method for reducing it-namely, increasing probability mass on the given answer choices by a) including them in the prompt and b) using in-context learning with even just one example.We show this method eliminates the impact of SFC in the majority of instances.Our experiments on three diverse datasets and six LMs reveal several additional surprising findings.For example, both normalization and prompting methods for reducing SFC can be ineffective or even detrimental to task performance for some LMs.We conclude with practical insights for effectively prompting LMs for multiple-choice tasks. 1 * Work done at AI2. 1 Code available at https://github.com/allenai/ revisiting_surface_form_competition. Sarah Wiegreffe, Matthew Finlayson, Oyvind Tafjord, Peter Clark, Ashish Sabharwal |
EMNLP | 4 |
| 2023 | Complexity-Based Prompting for Multi-step Reasoning
Hao Peng 0018, Ashish Sabharwal, Peter Clark, Tushar Khot |
ICLR | 4 |
| 2023 | Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
ICLR | 6 |
| 2023 | Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Pan Lu, Liang Qiu 0001, Kai-Wei Chang 0001, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, Ashwin Kalyan |
ICLR | 7 |
| 2023 | Self-Refine: Iterative Refinement with Self-FeedbackabstractLike humans, large language models (LLMs) do not always generate the best output on their first try. Motivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement. The main idea is to generate an initial output using an LLMs; then, the same LLMs provides *feedback* for its output and uses it to *refine* itself, iteratively. Self-Refine does not require any supervised training data, additional training, or reinforcement learning, and instead uses a single LLM as the generator, refiner and the feedback provider. We evaluate Self-Refine across 7 diverse tasks, ranging from dialog response generation to mathematical reasoning, using state-of-the-art (GPT-3.5, ChatGPT, and GPT-4) LLMs. Across all evaluated tasks, outputs generated with Self-Refine are preferred by humans and automatic metrics over those generated with the same LLM using conventional one-step generation, improving by $\sim$20\% absolute on average in task performance. Our work demonstrates that even state-of-the-art LLMs like GPT-4 can be further improved at test-time using our simple, standalone approach. Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon 0002, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang 0002, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark |
NeurIPS | 16 |
| 2022 | NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning TasksabstractSwaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan |
ACL (1) | 5 |
| 2022 | What Makes Instruction Learning Hard? An Investigation and a New Challenge in a Synthetic EnvironmentabstractThe instruction learning paradigm-where a model learns to perform new tasks from task descriptions alone-has become popular in research on general-purpose models.The capabilities of large transformer models as instruction learners, however, remain poorly understood.We use a controlled synthetic environment to characterize such capabilities.Specifically, we use the task of deciding whether a given string matches a regular expression (viewed as an instruction) to identify properties of tasks, instructions, and instances that make instruction learning challenging.For instance, we find that our model, a fine-tuned T5-based text2text transformer, struggles with large regular languages, suggesting that less precise instructions are challenging for models.Instruction executions that require tracking longer contexts of prior steps are also difficult.We use our findings to systematically construct a challenging instruction learning dataset, which we call Hard RegSet.Fine-tuning on Hard RegSet, our large transformer learns to correctly interpret (with at least 90% accuracy) only 65.6% of test instructions, and 11%-24% of the instructions in out-of-distribution generalization settings.We thus propose Hard RegSet as a challenging instruction learning dataset, and a controlled environment for studying instruction learning.1 Matthew Finlayson, Kyle Richardson 0001, Ashish Sabharwal, Peter Clark |
EMNLP | 4 |
| 2022 | Memory-assisted prompt editing to improve GPT-3 after deploymentabstractLarge LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans.For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homophone, while the user intended a synonym.Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly.We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification.Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past.On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3.Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. 1 * Equal Contribution 1 Code, data, and instructions to implement MemPrompt for a new task at https://www.memprompt.com/ Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang 0002 |
EMNLP | 3 |
| 2022 | LILA: A Unified Benchmark for Mathematical ReasoningabstractSwaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan |
EMNLP | 10 |
| 2022 | Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System ImprovementabstractOur goal is a teachable reasoning system for question-answering (QA), where a user can interact with faithful answer explanations, and correct its errors so that the system improves over time.Our approach is to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections to erroneous model beliefs that users identify during interaction.Retrievals from memory are used as additional context for QA, to help avoid previous mistakes in similar new situationsa novel application of memory-based continuous learning.With simulated feedback, we find that our system (called TeachMe 1 ) continually improves with time, and without model retraining, requiring feedback on only 25% of training examples to reach within 1% of the upper-bound (feedback on all examples).Similarly, in experiments with real users, we observe a similar trend, with performance improving by over 15% on a hidden test set after teaching.This suggests new opportunities for using frozen language models in an interactive setting where users can inspect, debug, and correct the model's beliefs, leading to improved system's performance over time. Bhavana Dalvi, Oyvind Tafjord, Peter Clark |
EMNLP | 3 |
| 2022 | Entailer: Answering Questions with Faithful and Truthful Chains of ReasoningabstractOur goal is a question-answering (QA) system that can show how its answers are implied by its own internal beliefs via a systematic chain of reasoning.Such a capability would allow better understanding of why a model produced the answer it did.Our approach is to recursively combine a trained backward-chaining model, capable of generating a set of premises entailing an answer hypothesis, with a verifier that checks that the model itself believes those premises (and the entailment itself) through self-querying.To our knowledge, this is the first system to generate multistep chains that are both faithful (the answer follows from the reasoning) and truthful (the chain reflects the system's own internal beliefs).In evaluation using two different datasets, users judge that a majority (70%+) of generated chains clearly show how an answer follows from a set of facts -substantially better than a high-performance baseline -while preserving answer accuracy.By materializing model beliefs that systematically support an answer, new opportunities arise for understanding the model's system of belief, and diagnosing and correcting its misunderstandings when an answer is wrong. Oyvind Tafjord, Bhavana Dalvi, Peter Clark |
EMNLP | 3 |
| 2022 | DREAM: Improving Situational QA by First Elaborating the SituationabstractWhen people answer questions about a specific situation, e.g., "I cheated on my mid-term exam last week.Was that wrong?", cognitive science suggests that they form a mental picture of that situation before answering.While we do not know how language models (LMs) answer such questions, we conjecture that they may answer more accurately if they are also provided with additional details about the question situation, elaborating the "scene".To test this conjecture, we train a new model, DREAM, to answer questions that elaborate the scenes that situated questions are about, and then provide those elaborations as additional context to a question-answering (QA) model.We find that DREAM is able to create better scene elaborations (more accurate, useful, and consistent) than a representative state-of-the-art, zero-shot model (Macaw).We also find that using the scene elaborations as additional context improves the answer accuracy of a downstream QA system, including beyond that obtainable by simply further fine-tuning the QA system on DREAM's training data.These results suggest that adding focused elaborations about a situation can improve a system's reasoning about it, and may serve as an effective way of injecting new scenario-based knowledge into QA models.Finally, our approach is dataset-neutral; we observe improved QA performance across different models, with even bigger gains on models with fewer parameters. 1 Yuling Gu, Bhavana Dalvi, Peter Clark |
NAACL-HLT | 3 |
| 2022 | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringabstractWhen answering a question, humans utilize the information available across different modalities to synthesize a consistent and complete chain of thought (CoT). This process is normally a black box in the case of deep learning models like large-scale language models. Recently, science question benchmarks have been used to diagnose the multi-hop reasoning ability and interpretability of an AI system. However, existing datasets fail to provide annotations for the answers, or are restricted to the textual-only modality, small scales, and limited domain diversity. To this end, we present Science Question Answering (ScienceQA), a new benchmark that consists of ~21k multimodal multiple choice questions with a diverse set of science topics and annotations of their answers with corresponding lectures and explanations. We further design language models to learn to generate lectures and explanations as the chain of thought (CoT) to mimic the multi-hop reasoning process when answering ScienceQA questions. ScienceQA demonstrates the utility of CoT in language models, as CoT improves the question answering performance by 1.20% in few-shot GPT-3 and 3.99% in fine-tuned UnifiedQA. We also explore the upper bound for models to leverage explanations by feeding those in the input; we observe that it improves the few-shot performance of GPT-3 by 18.96%. Our analysis further shows that language models, similar to humans, benefit from explanations to learn from fewer data and achieve the same performance with just 40% of the data. The data and code are available at https://scienceqa.github.io. Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 0001, Kai-Wei Chang 0001, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, Ashwin Kalyan |
NeurIPS | 8 |
| 2021 | Explaining Answers with Entailment TreesabstractBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, Peter Clark. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Bhavana Dalvi, Peter A. Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, Peter Clark |
EMNLP (1) | 7 |
| 2021 | How much coffee was consumed during EMNLP 2019? Fermi Problems: A New Reasoning Challenge for AIabstractMany real-world problems require the combined application of multiple reasoning abilities-employing suitable abstractions, commonsense knowledge, and creative synthesis of problem-solving strategies.To help advance AI systems towards such capabilities, we propose a new reasoning challenge, namely Fermi Problems (FPs), which are questions whose answers can only be approximately estimated because their precise computation is either impractical or impossible.For example, "How much would the sea level rise if all ice in the world melted?"FPs are commonly used in quizzes and interviews to bring out and evaluate the creative reasoning abilities of humans.To do the same for AI systems, we present two datasets: 1) A collection of 1k real-world FPs sourced from quizzes and olympiads; and 2) a bank of 10k synthetic FPs of intermediate complexity to serve as a sandbox for the harder real-world challenge.In addition to question-answer pairs, the datasets contain detailed solutions in the form of an executable program and supporting facts, helping in supervision and evaluation of intermediate steps.We demonstrate that even extensively fine-tuned large-scale language models perform poorly on these datasets, on average making estimates that are off by two orders of magnitude.Our contribution is thus the crystallization of several unsolved AI problems into a single, new challenge that we hope will spur further advances in building systems that can reason. Solving a Fermi Problem How much would the sea level rise if all the ice melted?Ice on land causes rise in sea levels. Div(a, b)Vol. of ice in the world? million mi²Area of ice?Area of ocean?Thickness of ice? 3 mi Area of Antarctica?Vol. of ice on land?How many Antarcticas fit in the world map?Vol. of ice in Antarctica?Mul(a, b) Ice on land Ice in Antarctica ≈ Area of Antarctica Area of ice in Antarctica Ashwin Kalyan, Arjun Chandrasekaran, Ashish Sabharwal, Peter Clark |
EMNLP (1) | 5 |
| 2021 | BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of BeliefabstractAlthough pretrained language models (PTLMs) contain significant amounts of world knowledge, they can still produce inconsistent answers to questions when probed, even after specialized training.As a result, it can be hard to identify what the model actually "believes" about the world, making it susceptible to inconsistent behavior and simple errors.Our goal is to reduce these problems.Our approach is to embed a PTLM in a broader system that also includes an evolving, symbolic memory of beliefs -a BeliefBank -that records but then may modify the raw PTLM answers.We describe two mechanisms to improve belief consistency in the overall system.First, a reasoning component -a weighted MaxSAT solver -revises beliefs that significantly clash with others.Second, a feedback component issues future queries to the PTLM using known beliefs as context.We show that, in a controlled experimental setting, these two mechanisms result in more consistent beliefs in the overall system, improving both the accuracy and consistency of its answers over time.This is significant as it is a first step towards PTLM-based architectures with a systematic notion of belief, enabling them to construct a more coherent picture of the world, and improve over time without model retraining. Nora Kassner, Oyvind Tafjord, Hinrich Schütze, Peter Clark |
EMNLP (1) | 4 |
| 2021 | Think about it! Improving defeasible reasoning by first modeling the question scenarioabstractDefeasible reasoning is the mode of reasoning where conclusions can be overturned by taking into account new evidence.Existing cognitive science literature on defeasible reasoning suggests that a person forms a mental model of the problem scenario before answering questions.Our research goal asks whether neural models can similarly benefit from envisioning the question scenario before answering a defeasible query.Our approach is, given a question, to have a model first create a graph of relevant influences, and then leverage that graph as an additional input when answering the question.Our system, CURIOUS, achieves a new stateof-the-art on three different defeasible reasoning datasets.This result is significant as it illustrates that performance can be improved by guiding a system to "think about" a question and explicitly model the scenario, rather than answering reflexively. 1 Aman Madaan, Niket Tandon, Dheeraj Rajagopal, Peter Clark, Yiming Yang 0002, Eduard H. Hovy |
EMNLP (1) | 4 |
| 2021 | Text Modular Networks: Learning to Decompose Tasks in the Language of Existing ModelsabstractTushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, Ashish Sabharwal. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tushar Khot, Daniel Khashabi, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
NAACL-HLT | 4 |
| 2020 | QASC: A Dataset for Question Answering via Sentence CompositionabstractComposing knowledge from multiple pieces of texts is a key challenge in multi-hop question answering. We present a multi-hop reasoning dataset, Question Answering via Sentence Composition (QASC), that requires retrieving facts from a large corpus and composing them to answer a multiple-choice question. QASC is the first dataset to offer two desirable properties: (a) the facts to be composed are annotated in a large corpus, and (b) the decomposition into these facts is not evident from the question itself. The latter makes retrieval challenging as the system must introduce new concepts or relations in order to discover potential decompositions. Further, the reasoning model must then learn to identify valid compositions of these retrieved facts using common-sense reasoning. To help address these challenges, we provide annotation for supporting facts as well as their composition. Guided by these annotations, we present a two-step approach to mitigate the retrieval challenges. We use other multiple-choice datasets as additional training data to strengthen the reasoning model. Our proposed approach improves over current state-of-the-art language models by 11% (absolute). The reasoning and retrieval problems, however, remain unsolved as this model still lags by 20% behind human performance. Tushar Khot, Peter Clark, Michal Guerquin, Peter A. Jansen, Ashish Sabharwal |
AAAI | 2 |
| 2020 | Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-AnsweringabstractDespite the rapid progress in multihop question-answering (QA), models still have trouble explaining why an answer is correct, with limited explanation training data available to learn from.To address this, we introduce three explanation datasets in which explanations formed from corpus facts are annotated.Our first dataset, eQASC, contains over 98K explanation annotations for the multihop question answering dataset QASC, and is the first that annotates multiple candidate explanations for each answer.The second dataset eQASC-perturbed is constructed by crowd-sourcing perturbations (while preserving their validity) of a subset of explanations in QASC, to test consistency and generalization of explanation prediction models.The third dataset eOBQA is constructed by adding explanation annotations to the OBQA dataset to test generalization of models trained on eQASC.We show that this data can be used to significantly improve explanation quality (+14% absolute F1 over a strong retrieval baseline) using a BERT-based classifier, but still behind the upper bound, offering a new challenge for future research.We also explore a delexicalized chain representation in which repeated noun phrases are replaced by variables, thus turning them into generalized reasoning chains (for example: "X is a Y" AND "Y has Z" IMPLIES "X has Z").We find that generalized chains maintain performance while also being more robust to certain perturbations. 1 Harsh Jhamtani, Peter Clark |
EMNLP (1) | 2 |
| 2020 | A Dataset for Tracking Entities in Open Domain Procedural TextabstractNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, Eduard Hovy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson 0001, Eduard H. Hovy |
EMNLP (1) | 5 |
| 2020 | Transformers as Soft Reasoners over LanguageabstractBeginning with McCarthy's Advice Taker (1959), AI has pursued the goal of providing a system with explicit, general knowledge and having the system reason over that knowledge. However, expressing the knowledge in a formal (logical or probabilistic) representation has been a major obstacle to this research. This paper investigates a modern approach to this problem where the facts and rules are provided as natural language sentences, thus bypassing a formal representation. We train transformers to reason (or emulate reasoning) over these sentences using synthetically generated data. Our models, that we call RuleTakers, provide the first empirical demonstration that this kind of soft reasoning over language is learnable, can achieve high (99%) accuracy, and generalizes to test data requiring substantially deeper chaining than seen during training (95%+ scores). We also demonstrate that the models transfer well to two hand-authored rulebases, and to rulebases paraphrased into more natural language. These findings are significant as it suggests a new role for transformers, namely as limited "soft theorem provers" operating over explicit theories in language. This in turn suggests new possibilities for explainability, correctability, and counterfactual reasoning in question-answering. Peter Clark, Oyvind Tafjord, Kyle Richardson 0001 |
IJCAI | 1 |
| 2020 | Multi-class Hierarchical Question Classification for Multiple Choice Science ExamsabstractPrior work has demonstrated that question classification (QC), recognizing the problem domain of a question, can help answer it more accurately. However, developing strong QC algorithms has been hindered by the limited size and complexity of annotated data available. To address this, we present the largest challenge dataset for QC, containing 7,787 science exam questions paired with detailed classification labels from a fine-grained hierarchical taxonomy of 406 problem domains. We then show that a BERT-based model trained on this dataset achieves a large (+0.12 MAP) gain compared with previous methods, while also achieving state-of-the-art performance on benchmark open-domain and biomedical QC datasets. Finally, we show that using this model’s predictions of question topic significantly improves the accuracy of a question answering system by +1.7% P@1, with substantial future gains possible as QC performance improves. Dongfang Xu, Peter A. Jansen, Jaycie Martin, Zhengnan Xie, Vikas Yadav, Harish Tayyar Madabushi, Oyvind Tafjord, Peter Clark |
LREC | 8 |
| 2020 | Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit KnowledgeabstractTo what extent can a neural network systematically reason over symbolic facts? Evidence suggests that large pre-trained language models (LMs) acquire some reasoning capacity, but this ability is difficult to control. Recently, it has been shown that Transformer-based models succeed in consistent reasoning over explicit symbolic facts, under a "closed-world" assumption. However, in an open-domain setup, it is desirable to tap into the vast reservoir of implicit knowledge already encoded in the parameters of pre-trained LMs. In this work, we provide a first demonstration that LMs can be trained to reliably perform systematic reasoning combining both implicit, pre-trained knowledge and explicit natural language statements. To do this, we describe a procedure for automatically generating datasets that teach a model new reasoning skills, and demonstrate that models learn to effectively perform inference which involves implicit taxonomic and world knowledge, chaining and counting. Finally, we show that "teaching" models to reason generalizes beyond the training distribution: they successfully compose the usage of multiple reasoning skills in single examples. Our work paves a path towards open-domain systems that constantly improve by interacting with users who can instantly correct a model by adding simple natural language statements. Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, Jonathan Berant |
NeurIPS | 3 |
| 2019 | Declarative Question Answering over Knowledge Bases Containing Natural Language Text with Answer Set ProgrammingabstractWhile in recent years machine learning (ML) based approaches have been the popular approach in developing endto-end question answering systems, such systems often struggle when additional knowledge is needed to correctly answer the questions. Proposed alternatives involve translating the question and the natural language text to a logical representation and then use logical reasoning. However, this alternative falters when the size of the text gets bigger. To address this we propose an approach that does logical reasoning over premises written in natural language text. The proposed method uses recent features of Answer Set Programming (ASP) to call external NLP modules (which may be based on ML) which perform simple textual entailment. To test our approach we develop a corpus based on the life cycle questions and showed that Our system achieves up to 18% performance gain when compared to standard MCQ solvers. Arindam Mitra, Peter Clark, Oyvind Tafjord, Chitta Baral |
AAAI | 2 |
| 2019 | QUAREL: A Dataset and Models for Answering Questions about Qualitative RelationshipsabstractMany natural la guage questions require recognizing and reasoning with qualitative relationships (e.g., in science, economics, and medicine), but are challenging to answer with corpus-based methods. Qualitative modeling provides tools that support such reasoning, but the semantic parsing task of mapping questions into those models has formidable challenges. We present QUAREL, a dataset of diverse story questions involving qualitative relationships that characterize these challenges, and techniques that begin to address them. The dataset has 2771 questions relating 19 different types of quantities. For example, “Jenny observes that the robot vacuum cleaner moves slower on the living room carpet than on the bedroom carpet. Which carpet has more friction?” We contribute (1) a simple and flexible conceptual framework for representing these kinds of questions; (2) the QUAREL dataset, including logical forms, exemplifying the parsing challenges; and (3) two novel models for this task, built as extensions of type-constrained semantic parsing. The first of these models (called QUASP+) significantly outperforms off-the-shelf tools on QUAREL. The second (QUASP+ZERO) demonstrates zero-shot capability, i.e., the ability to handle new qualitative relationships without requiring additional training data, something not possible with previous models. This work thus makes inroads into answering complex, qualitative questions that require reasoning, and scaling to new relationships at low cost. The dataset and models are available at http://data.allenai.org/quarel. Oyvind Tafjord, Peter Clark, Matt Gardner 0001, Scott Yih, Ashish Sabharwal |
AAAI | 2 |
| 2019 | Exploiting Explicit Paths for Multi-hop Reading ComprehensionabstractWe propose a novel, path-based reasoning approach for the multi-hop reading comprehension task where a system needs to combine facts from multiple passages to answer a question.Although inspired by multi-hop reasoning over knowledge graphs, our proposed approach operates directly over unstructured text.It generates potential paths through passages and scores them without any direct path supervision.The proposed model, named PathNet, attempts to extract implicit relations from text through entity pair representations, and compose them to encode each path.To capture additional context, Path-Net also composes the passage representations along each path to compute a passage-based representation.Unlike previous approaches, our model is then able to explain its reasoning via these explicit paths through the passages.We show that our approach outperforms prior models on the multi-hop Wikihop dataset, and also can be generalized to apply to the OpenBookQA dataset, matching stateof-the-art performance. Souvik Kundu 0003, Tushar Khot, Ashish Sabharwal, Peter Clark |
ACL (1) | 4 |
| 2019 | Everything Happens for a Reason: Discovering the Purpose of Actions in Procedural TextabstractBhavana Dalvi, Niket Tandon, Antoine Bosselut, Wen-tau Yih, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bhavana Dalvi, Niket Tandon, Antoine Bosselut, Scott Yih, Peter Clark |
EMNLP/IJCNLP (1) | 5 |
| 2019 | What's Missing: A Knowledge Gap Guided Approach for Multi-hop Question AnsweringabstractTushar Khot, Ashish Sabharwal, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tushar Khot, Ashish Sabharwal, Peter Clark |
EMNLP/IJCNLP (1) | 3 |
| 2019 | QuaRTz: An Open-Domain Dataset of Qualitative Relationship QuestionsabstractOyvind Tafjord, Matt Gardner, Kevin Lin, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Oyvind Tafjord, Matt Gardner 0001, Peter Clark |
EMNLP/IJCNLP (1) | 4 |
| 2019 | WIQA: A dataset for "What if..." reasoning over procedural textabstractNiket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Niket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Project Aristo: Towards Machines that Capture and Reason with Science KnowledgeabstractAI2's Project Aristo seeks to build a system that has a deep understanding of science, using knowledge captured mainly from large-scale text. Recently, Aristo achieved surprising success on the Grade 8 New York Regents Science Exams, scoring over 90% on the exam's non-diagram, multiple choice (NDMC) questions, where even 3 years ago the best systems scored less than 60%. In this talk, I will describe the journey of Aristo through various knowledge capture technologies that have helped it, including acquiring if/then rules, tables, knowledge graphs, and latent neural representations. I will also discuss the growing tension between capturing structured knowledge vs. capturing knowledge latently using neural models, the latter proving highly effective but hard to interpret. Finally I will speculate on the larger quest towards knowledgable machines that can reason, explain, and discuss, and how structured and latent knowledge can interact to help reach this goal. Peter Clark |
K-CAP | 1 |
| 2018 | SciTaiL: A Textual Entailment Dataset from Science Question AnsweringabstractWe present a new dataset and model for textual entailment, derived from treating multiple-choice question-answering as an entailment problem. SciTail is the first entailment set that is created solely from natural sentences that already exist independently ``in the wild'' rather than sentences authored specifically for the entailment task. Different from existing entailment datasets, we create hypotheses from science questions and the corresponding answer candidates, and premises from relevant web sentences retrieved from a large corpus. These sentences are often linguistically challenging. This, combined with the high lexical similarity of premise and hypothesis for both entailed and non-entailed pairs, makes this new entailment task particularly difficult. The resulting challenge is evidenced by state-of-the-art textual entailment systems achieving mediocre performance on SciTail, especially in comparison to a simple majority class baseline. As a step forward, we demonstrate that one can improve accuracy on SciTail by 5% using a new neural model that exploits linguistic structure. Tushar Khot, Ashish Sabharwal, Peter Clark |
AAAI | 3 |
| 2018 | Bridging Knowledge Gaps in Neural Entailment via Symbolic ModelsabstractMost textual entailment models focus on lexical gaps between the premise text and the hypothesis, but rarely on knowledge gaps.We focus on filling these knowledge gaps in the Science Entailment task, by leveraging an external structured knowledge base (KB) of science facts.Our new architecture combines standard neural entailment models with a knowledge lookup module.To facilitate this lookup, we propose a fact-level decomposition of the hypothesis, and verifying the resulting sub-facts against both the textual premise and the structured KB.Our model, NSnet, learns to aggregate predictions from these heterogeneous data formats.On the SciTail dataset, NSnet outperforms a simpler combination of the two predictions by 3% and the base entailment model by 5%. Dongyeop Kang, Tushar Khot, Ashish Sabharwal, Peter Clark |
EMNLP | 4 |
| 2018 | Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question AnsweringabstractWe present a new kind of question answering dataset, OpenBookQA, modeled after open book exams for assessing human understanding of a subject.The open book that comes with our questions is a set of 1326 elementary level science facts.Roughly 6000 questions probe an understanding of these facts and their application to novel situations.This requires combining an open book fact (e.g., metals conduct electricity) with broad common knowledge (e.g., a suit of armor is made of metal) obtained from other sources.While existing QA datasets over documents or knowledge bases, being generally self-contained, focus on linguistic understanding, OpenBookQA probes a deeper understanding of both the topic-in the context of common knowledge-and the language it is expressed in.Human performance on OpenBookQA is close to 92%, but many state-of-the-art pre-trained QA methods perform surprisingly poorly, worse than several simple neural baselines we develop.Our oracle experiments designed to circumvent the knowledge retrieval bottleneck demonstrate the value of both the open book and additional facts.We leave it as a challenge to solve the retrieval problem in this multi-hop setting and to close the large gap to human performance.Question: Which of these would let the most heat travel through?A) a new pair of jeans.B) a steel spoon in a cafeteria.C) a cotton candy at a store.D) a calvin klein cotton hat. Todor Mihaylov, Peter Clark, Tushar Khot, Ashish Sabharwal |
EMNLP | 2 |
| 2018 | Reasoning about Actions and State Changes by Injecting Commonsense KnowledgeabstractComprehending procedural text, e.g., a paragraph describing photosynthesis, requires modeling actions and the state changes they produce, so that questions about entities at different timepoints can be answered.Although several recent systems have shown impressive progress in this task, their predictions can be globally inconsistent or highly improbable.In this paper, we show how the predicted effects of actions in the context of a paragraph can be improved in two ways: (1) by incorporating global, commonsense constraints (e.g., a non-existent entity cannot be destroyed), and (2) by biasing reading with preferences from large-scale corpora (e.g., trees rarely move).Unlike earlier methods, we treat the problem as a neural structured prediction task, allowing hard and soft constraints to steer the model away from unlikely predictions.We show that the new model significantly outperforms earlier systems on a benchmark dataset for procedural text comprehension (+8% relative gain), and that it also avoids some of the nonsensical predictions that earlier systems make. Niket Tandon, Bhavana Dalvi, Joel Grus, Scott Yih, Antoine Bosselut, Peter Clark |
EMNLP | 6 |
| 2018 | Knowledge Representation and Reasoning in Answering Science Questions: A Case Study for Food Web Questions
Arindam Mitra, Chitta Baral, Peter Clark |
KR | 3 |
| 2018 | Tracking State Changes in Procedural Text: a Challenge Dataset and Models for Process Paragraph ComprehensionabstractBhavana Dalvi, Lifu Huang, Niket Tandon, Wen-tau Yih, Peter Clark. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Bhavana Dalvi, Lifu Huang, Niket Tandon, Scott Yih, Peter Clark |
NAACL-HLT | 5 |
| 2017 | Tell Me Why: Using Question Answering as Distant Supervision for Answer JustificationabstractRebecca Sharp, Mihai Surdeanu, Peter Jansen, Marco A. Valenzuela-Escárcega, Peter Clark, Michael Hammond. Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017). 2017. Rebecca Sharp, Mihai Surdeanu, Peter A. Jansen, Marco Antonio Valenzuela-Escárcega, Peter Clark, Michael Hammond |
CoNLL | 5 |
| 2017 | Framing QA as Building and Ranking Intersentence Answer JustificationsabstractWe propose a question answering (QA) approach for standardized science exams that both identifies correct answers and produces compelling human-readable justifications for why those answers are correct. Our method first identifies the actual information needed in a question using psycholinguistic concreteness norms, then uses this information need to construct answer justifications by aggregating multiple sentences from different knowledge bases using syntactic and lexical information. We then jointly rank answers and their justifications using a reranking perceptron that treats justification quality as a latent variable. We evaluate our method on 1,000 multiple-choice questions from elementary school science exams, and empirically demonstrate that it performs better than several strong baselines, including neural network approaches. Our best configuration answers 44% of the questions correctly, where the top justifications for 57% of these correct answers contain a compelling human-readable justification that explains the inference required to arrive at the correct answer. We include a detailed characterization of the justification quality for both our method and a strong baseline, and show that information aggregation is key to addressing the information need in complex questions. Peter A. Jansen, Rebecca Sharp, Mihai Surdeanu, Peter Clark |
Comput. Linguistics | 4 |
| 2017 | Domain-Targeted, High Precision Knowledge ExtractionabstractOur goal is to construct a domain-targeted, high precision knowledge base (KB), containing general (subject,predicate,object) statements about the world, in support of a downstream question-answering (QA) application. Despite recent advances in information extraction (IE) techniques, no suitable resource for our task already exists; existing resources are either too noisy, too named-entity centric, or too incomplete, and typically have not been constructed with a clear scope or purpose. To address these, we have created a domain-targeted, high precision knowledge extraction pipeline, leveraging Open IE, crowdsourcing, and a novel canonical schema learning algorithm (called CASI), that produces high precision knowledge targeted to a particular domain - in our case, elementary science. To measure the KB’s coverage of the target domain’s knowledge (its “comprehensiveness” with respect to science) we measure recall with respect to an independent corpus of domain text, and show that our pipeline produces output with over 80% precision and 23% recall with respect to that target, a substantially higher coverage of tuple-expressible science knowledge than other comparable resources. We have made the KB publicly available. Bhavana Dalvi, Niket Tandon, Peter Clark |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | Combining Retrieval, Statistics, and Inference to Answer Elementary Science QuestionsabstractWhat capabilities are required for an AI system to pass standard 4th Grade Science Tests? Previous work has examined the use of Markov Logic Networks (MLNs) to represent the requisite background knowledge and interpret test questions, but did not improve upon an information retrieval (IR) baseline. In this paper, we describe an alternative approach that operates at three levels of representation and reasoning: information retrieval, corpus statistics, and simple inference over a semi-automatically constructed knowledge base, to achieve substantially improved results. We evaluate the methods on six years of unseen, unedited exam questions from the NY Regents Science Exam (using only non-diagram, multiple choice questions), and show that our overall system’s score is 71.3%, an improvement of 23.8% (absolute) over the MLN-based method described in previous work. We conclude with a detailed analysis, illustrating the complementary strengths of each method in the ensemble. Our datasets are being released to enable further research. Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter D. Turney, Daniel Khashabi |
AAAI | 1 |
| 2016 | What's in an Explanation? Characterizing Knowledge and Inference Requirements for Elementary Science ExamsabstractQA systems have been making steady advances in the challenging elementary science exam domain. In this work, we develop an explanation-based analysis of knowledge and inference requirements, which supports a fine-grained characterization of the challenges. In particular, we model the requirements based on appropriate sources of evidence to be used for the QA task. We create requirements by first identifying suitable sentences in a knowledge base that support the correct answer, then use these to build explanations, filling in any necessary missing information. These explanations are used to create a fine-grained categorization of the requirements. Using these requirements, we compare a retrieval and an inference solver on 212 questions. The analysis validates the gains of the inference solver, demonstrating that it answers more questions requiring complex inference, while also providing insights into the relative strengths of the solvers and knowledge sources. We release the annotated questions and explanations as a resource with broad utility for science exam QA, including determining knowledge base construction targets, as well as supporting information aggregation in automated inference. Peter A. Jansen, Niranjan Balasubramanian, Mihai Surdeanu, Peter Clark |
COLING | 4 |
| 2016 | Cross Sentence Inference for Process KnowledgeabstractFor AI systems to reason about real world situations, they need to recognize which processes are at play and which entities play key roles in them.Our goal is to extract this kind of rolebased knowledge about processes, from multiple sentence-level descriptions.This knowledge is hard to acquire; while semantic role labeling (SRL) systems can extract sentence level role information about individual mentions of a process, their results are often noisy and they do not attempt create a globally consistent characterization of a process.To overcome this, we extend standard within sentence joint inference to inference across multiple sentences.This cross sentence inference promotes role assignments that are compatible across different descriptions of the same process.When formulated as an Integer Linear Program, this leads to improvements over within-sentence inference by nearly 3% in F1.The resulting role-based knowledge is of high quality (with a F1 of nearly 82). Samuel Louvan, Chetan Naik, Sadhana Kumaravel, Heeyoung Kwon, Niranjan Balasubramanian, Peter Clark |
EMNLP | 6 |
| 2016 | Creating Causal Embeddings for Question Answering with Minimal SupervisionabstractA common model for question answering (QA) is that a good answer is one that is closely related to the question, where relatedness is often determined using generalpurpose lexical models such as word embeddings.We argue that a better approach is to look for answers that are related to the question in a relevant way, according to the information need of the question, which may be determined through task-specific embeddings.With causality as a use case, we implement this insight in three steps.First, we generate causal embeddings cost-effectively by bootstrapping cause-effect pairs extracted from free text using a small set of seed patterns.Second, we train dedicated embeddings over this data, by using task-specific contexts, i.e., the context of a cause is its effect.Finally, we extend a state-of-the-art reranking approach for QA to incorporate these causal embeddings.We evaluate the causal embedding models both directly with a casual implication task, and indirectly, in a downstream causal QA task using data from Yahoo! Answers.We show that explicitly modeling causality improves performance in both tasks.In the QA task our best model achieves 37.3% P@1, significantly outperforming a strong baseline by 7.7% (relative). Rebecca Sharp, Mihai Surdeanu, Peter A. Jansen, Peter Clark, Michael Hammond |
EMNLP | 4 |
| 2016 | Question Answering via Integer Programming over Semi-Structured Knowledge
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni, Dan Roth 0001 |
IJCAI | 4 |
| 2015 | Elementary School Science and Math Tests as a Driver for AI: Take the Aristo Challenge!abstractWhile there has been an explosion of impressive, data-driven AI applications in recent years, machines still largely lack a deeper understanding of the world to answer questions that go beyond information explicitly stated in text, and to explain and discuss those answers. To reach this next generation of AI applications, it is imperative to make faster progress in areas of knowledge, modeling, reasoning, and language. Standardized tests have often been proposed as a driver for such progress, with good reason: Many of the questions require sophisticated understanding of both language and the world, pushing the boundaries of AI, while other questions are easier, supporting incremental progress. In Project Aristo at the Allen Institute for AI, we are working on a specific version of this challenge, namely having the computer pass Elementary School Science and Math exams. Even at this level there is a rich variety of problems and question types, the most difficult requiring significant progress in AI. Here we propose this task as a challenge problem for the community, and are providing supporting datasets. Solutions to many of these problems would have a major impact on the field so we encourage you: Take the Aristo Challenge! Peter Clark |
AAAI | 1 |
| 2015 | Exploring Markov Logic Networks for Question AnsweringabstractElementary-level science exams pose sig-nificant knowledge acquisition and rea-soning challenges for automatic question answering. We develop a system that rea-sons with knowledge derived from text-books, represented in a subset of first-order logic. Automatic extraction, while scalable, often results in knowledge that is incomplete and noisy, motivating use of reasoning mechanisms that handle uncer-tainty. Markov Logic Networks (MLNs) seem a natural model for expressing such knowl-edge, but the exact way of leveraging MLNs is by no means obvious. We in-vestigate three ways of applying MLNs to our task. First, we simply use the extracted science rules directly as MLN clauses and exploit the structure present in hard con-straints to improve tractability. Second, we interpret science rules as describing prototypical entities, resulting in a drasti-cally simplified but brittle network. Our third approach, called Praline, uses MLNs to align lexical elements as well as define and control how inference should be per-formed in this task. Praline demonstrates a 15 % accuracy boost and a 10x reduction in runtime as compared to other MLN-based methods, and comparable accuracy to word-based baseline approaches. Tushar Khot, Niranjan Balasubramanian, Eric Gribkoff, Ashish Sabharwal, Peter Clark, Oren Etzioni |
EMNLP | 5 |
| 2015 | Answering Elementary Science Questions by Constructing Coherent Scenes using Background KnowledgeabstractMuch of what we understand from text is not explicitly stated.Rather, the reader uses his/her knowledge to fill in gaps and create a coherent, mental picture or "scene" depicting what text appears to convey.The scene constitutes an understanding of the text, and can be used to answer questions that go beyond the text.Our goal is to answer elementary science questions, where this requirement is pervasive; A question will often give a partial description of a scene and ask the student about implicit information.We show that by using a simple "knowledge graph" representation of the question, we can leverage several large-scale linguistic resources to provide missing background knowledge, somewhat alleviating the knowledge bottleneck in previous approaches.The coherence of the best resulting scene, built from a question/answer-candidate pair, reflects the confidence that the answer candidate is correct, and thus can be used to answer multiple choice questions.Our experiments show that this approach outperforms competitive algorithms on several datasets tested.The significance of this work is thus to show that a simple "knowledge graph" representation allows a version of "interpretation as scene construction" to be made viable. Yang Li 0150, Peter Clark |
EMNLP | 2 |
| 2015 | Learning Knowledge Graphs for Question Answering through Conversational DialogabstractWe describe how a question-answering system can learn about its domain from conversational dialogs.Our system learns to relate concepts in science questions to propositions in a fact corpus, stores new concepts and relations in a knowledge graph (KG), and uses the graph to solve questions.We are the first to acquire knowledge for question-answering from open, natural language dialogs without a fixed ontology or domain model that predetermines what users can say.Our relation-based strategies complete more successful dialogs than a query expansion baseline, our taskdriven relations are more effective for solving science questions than relations from general knowledge sources, and our method is practical enough to generalize to other domains. Ben Hixon, Peter Clark, Hannaneh Hajishirzi |
HLT-NAACL | 2 |
| 2015 | Spinning Straw into Gold: Using Free Text to Train Monolingual Alignment Models for Non-factoid Question AnsweringabstractRebecca Sharp, Peter Jansen, Mihai Surdeanu, Peter Clark. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Rebecca Sharp, Peter A. Jansen, Mihai Surdeanu, Peter Clark |
HLT-NAACL | 4 |
| 2015 | Higher-order Lexical Semantic Models for Non-factoid Answer RerankingabstractLexical semantic models provide robust performance for question answering, but, in general, can only capitalize on direct evidence seen during training. For example, monolingual alignment models acquire term alignment probabilities from semi-structured data such as question-answer pairs; neural network language models learn term embeddings from unstructured text. All this knowledge is then used to estimate the semantic similarity between question and answer candidates. We introduce a higher-order formalism that allows all these lexical semantic models to chain direct evidence to construct indirect associations between question and answer texts, by casting the task as the traversal of graphs that encode direct term associations. Using a corpus of 10,000 questions from Yahoo! Answers, we experimentally demonstrate that higher-order methods are broadly applicable to alignment and language models, across both word and syntactic representations. We show that an important criterion for success is controlling for the semantic drift that accumulates during graph traversal. All in all, the proposed higher-order approach improves five out of the six lexical semantic models investigated, with relative gains of up to +13% over their first-order variants. Daniel Fried, Peter A. Jansen, Gus Hahn-Powell, Mihai Surdeanu, Peter Clark |
Trans. Assoc. Comput. Linguistics | 5 |
| 2014 | Discourse Complements Lexical Semantics for Non-factoid Answer RerankingabstractWe propose a robust answer reranking model for non-factoid questions that integrates lexical semantics with discourse information, driven by two representations of discourse: a shallow representation centered around discourse markers, and a deep one based on Rhetorical Structure Theory.We evaluate the proposed model on two corpora from different genres and domains: one from Yahoo! Answers and one from the biology domain, and two types of non-factoid questions: manner and reason.We experimentally demonstrate that the discourse structure of nonfactoid answers provides information that is complementary to lexical semantic similarity between question and answer, improving performance up to 24% (relative) over a state-of-the-art model that exploits lexical semantic similarity alone.We further demonstrate excellent domain transfer of discourse information, suggesting these discourse features have general utility to non-factoid question answering. Peter A. Jansen, Mihai Surdeanu, Peter Clark |
ACL (1) | 3 |
| 2014 | Modeling Biological Processes for Reading ComprehensionabstractJonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014. Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning |
EMNLP | 7 |
| 2013 | Learning Biological Processes with Global ConstraintsabstractAju Thalappillil Scaria, Jonathan Berant, Mengqiu Wang, Peter Clark, Justin Lewis, Brittany Harding, Christopher D. Manning. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 2013. Aju Thalappillil Scaria, Jonathan Berant, Mengqiu Wang, Peter Clark, Justin Lewis, Brittany Harding, Christopher D. Manning |
EMNLP | 4 |
| 2013 | Semi-Markov Phrase-Based Monolingual AlignmentabstractWe introduce a novel discriminative model for phrase-based monolingual alignment using a semi-Markov CRF.Our model achieves stateof-the-art alignment accuracy on two phrasebased alignment datasets (RTE and paraphrase), while doing significantly better than other strong baselines in both non-identical alignment and phrase-only alignment.Additional experiments highlight the potential benefit of our alignment model to RTE, paraphrase identification and question answering, where even a naive application of our model's alignment score approaches the state of the art. Xuchen Yao, Benjamin Van Durme, Chris Callison-Burch, Peter Clark |
EMNLP | 4 |
| 2013 | Answer Extraction as Sequence Tagging with Tree Edit Distance
Xuchen Yao, Benjamin Van Durme, Chris Callison-Burch, Peter Clark |
HLT-NAACL | 4 |
| 2011 | Inquire for iPad: A Biology Textbook That Answers Questions
Aaron Spaulding, Adam Overholtzer, John Pacheco, Jing Tien, Vinay K. Chaudhri, David Gunning, Peter Clark |
AIED | 7 |
| 2011 | Preliminary steps towards a knowledge factory processabstractIn the fall 2010 issue of the AI Magazine, we reported the design, implementation and evaluation of a knowledge acquisition system called AURA. AURA enables domain experts in Physics, Chemistry and Biology to author their knowledge, and a different set of experts to pose questions against that knowledge. The evaluation results previously reported were from 50 pages each from science textbooks in Physics, Chemistry and Biology. The results were most promising for Biology. Based on those results we undertook a content building effort to capture knowledge from approximately 315 pages (or 20 chapters) of the same Biology textbook [2] and incorporated the resulting content in the electronic version of that book. In this demo/poster session, we will demonstrate the biology knowledge base (KB) created using AURA, the electronic textbook application Inquire, and discuss the knowledge engineering process we used to construct the KB. Vinay K. Chaudhri, Nikhil Dinesh, John Pacheco, Gary Ng, Peter Clark, Andrew Goldenkranz, A. Patrice Seyed, Naveen Sharma |
K-CAP | 5 |
| 2009 | Large-scale extraction and use of knowledge from textabstractMany AI tasks, in particular natural language processing, require a large amount of world knowledge to create expectations, assess plausibility, and guide disambiguation. However, acquiring this world knowledge remains a formidable challenge. Building on ideas by Schubert, we have developed a system called DART (Discovery and Aggregation of Relations in Text) that extracts simple, semi-formal statements of world knowledge (e.g., "airplanes can fly", "people can drive cars") from text by abstracting from a parser's output, and we have used it to create a database of 23 million propositions of this kind. An evaluation of the DART database on two language processing tasks (parsing and textual entailment) shows that it improves performance, and a human evaluation shows that over half the facts in it are considered true or partially true, rising to 70% for facts seen with high frequency. The significance of this work is two-fold: First it has created a new, publically available knowledge resource for language processing and other data interpretation tasks, and second it provides empirical evidence of the utility of this type of knowledge, going beyond Schubert et al's earlier evaluations which were based solely on human inspection of its contents. Peter Clark, Philip Harrison |
K-CAP | 1 |
| 2008 | Knowledge Patterns
Peter Clark |
EKAW | 1 |
| 2007 | AURA: Enabling Subject Matter Experts to Construct Declarative Knowledge Bases from Science Textbooks
Ken Barker 0002, Vinay K. Chaudhri, Shaw Yi Chaw, Peter Clark, Daniel Hansch, Bonnie E. John, Sunil Mishra, John Pacheco, Bruce W. Porter, Aaron Spaulding, Moritz Weiten |
AAAI | 4 |
| 2007 | Capturing and answering questions posed to a knowledge-based systemabstractAs part of the ongoing project, Project Halo, our goal is to build a system capable of answering questions posed by novice users to a formal knowledge base. In our current context, the knowledge base covers selected topics in physics, chemistry, and biology, and our question set consists of AP (advanced high-school) level examination questions. The task is challenging because the questions are linguistically complex and are often incomplete (assume unstated knowledge), and because the users do not have prior knowledge of the system's contents. Our solution involves two parts: a controlled language interface, in which users reformulate the original natural language questions in a simplified version of English, and a novel problem solver that can elaborate initially inadequate logical interpretations of a question by selecting relevant pieces of knowledge in the knowledge base. An evaluation of the work in 2006 showed that this approach is feasible and that complex, multisentence questions can be posed and answered, thus illustrating novel ways of dealing with the knowledge capture impedance between users and a formal knowledge base, while also revealing challenges that still remain. Peter Clark, Shaw Yi Chaw, Ken Barker 0002, Vinay K. Chaudhri, Philip Harrison, James Fan, Bonnie E. John, Bruce W. Porter, Aaron Spaulding, John A. Thompson, Peter Z. Yeh |
K-CAP | 1 |
| 2004 | Graph-Based Acquisition of Expressive Knowledge
Vinay K. Chaudhri, Kenneth S. Murray, John Pacheco, Peter Clark, Bruce W. Porter, Patrick J. Hayes |
EKAW | 4 |
| 2004 | A Question-Answering System for AP Chemistry: Assessing KR&R Technologies
Ken Barker 0002, Vinay K. Chaudhri, Shaw Yi Chaw, Peter Clark, James Fan, David J. Israel, Sunil Mishra, Bruce W. Porter, Pedro Romero, Dan Tecuci, Peter Z. Yeh |
KR | 4 |
| 2004 | Towards a Quantitative, Platform-Independent Analysis of Knowledge Systems
Noah S. Friedland, Paul G. Allen, Michael Witbrock, Gavin Matthews, Nancy Salay, Pierluigi Miraglia, Jürgen Angele, Steffen Staab, David J. Israel, Vinay K. Chaudhri, Bruce W. Porter, Ken Barker 0002, Peter Clark |
KR | 13 |
| 2003 | A Knowledge Acquisition Tool for Course of Action Analysis
Kim Barker, Jim Blythe, Gary C. Borchardt, Vinay K. Chaudhri, Peter Clark, Paul R. Cohen, Julie Fitzgerald, Kenneth D. Forbus, Yolanda Gil, Boris Katz, Jihie Kim, Gary W. King, Sunil Mishra, Clayton T. Morrison, Kenneth S. Murray, Charley Otstott, Bruce W. Porter, Robert Schrag, Tomás E. Uribe, Jeffrey M. Usher, Peter Z. Yeh |
IAAI | 5 |
| 2003 | Enabling domain experts to convey questions to a machine: a modified, template-based approachabstractIn order for a knowledge capture system to be effective, it needs to not only acquire general domain knowledge from experts, but also capture the specific problem-solving scenarios and questions which those experts are interested in solving using that knowledge. For some tasks, this latter aspect of knowledge capture is straightforward. In other cases, in particular for systems aimed at a wide variety of tasks, the question-posing aspect of knowledge capture can be a challenge in its own right. In this paper, we present the approach we have developed to address this challenge, based on the creation of a catalog of domain-independent question types and the extension of question template methods with graphical tools. Our goal was that domain experts could directly convey complex questions to a machine, in a form which it could then reason with. We evaluated the resulting system over several weeks, and in this paper we report some important lessons learned from this evaluation, revealing several interesting strengths and weaknesses of the approach. Peter Clark, Vinay K. Chaudhri, Sunil Mishra, Jérôme Thoméré, Ken Barker 0002, Bruce W. Porter |
K-CAP | 1 |
| 2003 | A Semantic Infosphere
Michael Uschold, Peter Clark, Fred Dickey, Casey K. Fung, Sonia Smith, Stephen A. Uczekaj, Michael Wilke, Sean Bechhofer, Ian Horrocks 0001 |
ISWC | 2 |
| 2001 | A library of generic concepts for composing knowledge basesabstractBuilding a knowledge base for a given domain traditionally involves a subject matter expert and a knowledge engineer. One of the goals of our research is to eliminate the knowledge engineer. There are at least two ways to achieve this goal: train domain experts to write axioms (i.e., turn them into knowledge engineers) or create tools that allow users to build knowledge bases without having to write axioms. Our strategy is to create tools that allow users to build knowledge bases through instantiation and assembly of generic knowledge components from a small library.In many ways, creating such a library is like designing an ontology: What are the most general kinds of events and entities? How are these things related hierarchically? What is their meaning and how is it represented? The pressures of making the library usable by domain experts, however, leads to departures from the traditional ontology design goals of coverage, consensus and elegance. In this paper we describe our component library, a hierarchy of reusable, composable, domain-independent knowledge units. The library emphasizes coverage (what is an appropriate set of components for our task), access (how can a domain expert find appropriate components) and semantics (what knowledge and what kind of representation permit useful composition). We have begun building a library on these principles, influenced heavily by linguistic resources. In early evaluations we have put the library into the hands of domain experts (in Biology) having no experience with knowledge bases or knowledge acquisition. Ken Barker 0002, Bruce W. Porter, Peter Clark |
K-CAP | 3 |
| 2001 | Knowledge entry as the graphical assembly of componentsabstractDespite some successes, the lack of tools to allow subject matter experts to directly enter, query, and debug formal domain knowledge in a knowledge-base still remains a major obstacle to their deployment. Our goal is to create such tools, so that a trained knowledge engineer is no longer required to mediate the interaction. This paper presents our work on the knowledge entry part of this overall knowledge capture task, which is based on several claims: that users can construct representations by connecting pre-fabricated, representational components, rather than writing low-level axioms; that these components can be presented to users as graphs; and the user can then perform composition through graph manipulation operations. To operationalize this, we have developed a novel technique of graphical dialog using examples of the component concepts, followed by an automated process for generalizing the user's graphically-entered assertions into axioms. We present these claims, our approach, the system (called SHAKEN) that we are developing, and an evaluation of our progress based on having users encode knowledge using the system. Keywords Graphical knowledge entry, knowledge acquisition, components, composition, knowledge-based systems. Peter Clark, John A. Thompson, Ken Barker 0002, Bruce W. Porter, Vinay K. Chaudhri, Andres C. Rodriguez, Jérôme Thoméré, Sunil Mishra, Yolanda Gil, Patrick J. Hayes, Thomas Reichherzer |
K-CAP | 1 |
| 2001 | Representing roles and purposeabstractOntology designers often distinguish Entities (things that are) from Events (things that happen). It is not obvious how this division admits Roles (things that are, but only in the context of things that happen). For example, Person might be considered an Entity, while Employee is a Role. A Person remains a Person independent of the Events in which he participates. Someone is an Employee only by virtue of participating in an Employment Event. The problem of how to represent Roles is not new, but there is little consensus on a solution. In this paper, we present an ontology that finds a place for Roles as well as a representation that allows Roles to be related to Entities and Events to express the teleological notion of purpose. James Fan, Ken Barker 0002, Bruce W. Porter, Peter Clark |
K-CAP | 4 |
| 2000 | Knowledge Patterns
Peter Clark, John A. Thompson, Bruce W. Porter |
KR | 1 |
| 1997 | Improving Image Classification by Combining Statistical, Case-Based and Model Based Prediction MethodsabstractEvidence for image classification can be considered to come from two sources: traditional statistical information derived algorithmically from image data, and model-based evidence arising from previous expertise and experience in a given application Peter Clark, Cao Feng, Stan Matwin, Ko Fung |
Fundam. Informaticae | 1 |
| 1993 | Learning Domain Theories using Abstract Beckground Knowledge
Peter Clark, Stan Matwin |
ECML | 1 |
| 1993 | Using Qualitative Models to Guide Inductive Learning
Peter Clark, Stan Matwin |
ICML | 1 |
| 1992 | Lazy Partial Evaluation: An Integration of Explanation-Based Generalization and Partial Evaluation
Peter Clark, Robert C. Holte |
ML | 1 |
| 1989 | The CN2 Induction Algorithm
Peter Clark, Tim Niblett |
Mach. Learn. | 1 |