VLDB 2026 Research / reviewers in the wild / expert
Peter A. Jansen
dblp:72/5962 · also Peter Jansen 0001
· DBLP profile ↗
25ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 9 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Literature-Driven Scientific Theories at ScaleabstractContemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theory building remain underexplored.In this work, we formulate the problem of synthesizing theories consisting of qualitative and quantitative laws from large corpora of scientific literature.We study theory generation at scale, using 13.7k source papers to synthesize 2.9k theories, examining how generation using literaturegrounding versus parametric knowledge, and accuracy-focused versus novelty-focused generation objectives change theory properties.Our experiments show that, compared to using parametric LLM memory for generation, our literature-supported method creates theories that are significantly better at both matching existing evidence and at predicting future results from 4.6k subsequently-written papers. 1 Peter A. Jansen, Peter Clark, Doug Downey, Daniel S. Weld |
ACL (1) | 1 |
| 2025 | Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials ScienceabstractContemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically generating code-based experiments to test those hypotheses.While hypotheses can be comparatively inexpensive to generate, automated experiments can be costly, particularly when run at scale (i.e.thousands of experiments).Developing the capacity to filter hypotheses based on their feasibility would allow discovery systems to run at scale, while increasing their likelihood of making significant discoveries.In this work we introduce MATTER-OF-FACT, a challenge dataset for determining the feasibility of hypotheses framed as claims, while operationalizing feasibility assessment as a temporally-filtered claim verification task using backtesting.MATTER-OF-FACT includes 8.4K claims extracted from scientific articles spanning four high-impact contemporary materials science topics, including superconductors, semiconductors, batteries, and aerospace materials, while including qualitative and quantitative claims from theoretical, experimental, and code/simulation results.We show that strong baselines that include retrieval augmented generation over scientific literature and code generation fail to exceed 72% performance on this task (chance performance is 50%), while domain-expert verification suggests nearly all are solvable -highlighting both the difficulty of this task for current models, and the potential to accelerate scientific discovery by making near-term progress.1 Batteries Battery Claim #320Claim: In Li10Ge(PS6)2, Li+ ion motion becomes less correlated at lower temps., with Haven ratios closer to 1 below the sublattice phase trans.temp.Gold Label: True (feasible) Explanation: Demonstrated through Haven ratio calculations, Li+ ion motion becomes less correlated at lower temperatures.The Haven ratio approaches 1 below the sublattice phase transition temperature (~400K). Peter A. Jansen, Samiah Hassan, Ruoyao Wang |
EMNLP | 1 |
| 2025 | From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-AnsweringabstractRecent reasoning methods (e.g., chain-of-thought) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM’s overall understanding, or “theory,” about the question’s topic, making it still hard to trust the model. Our goal is to materialize such theories - here called microtheories (a linguistic analog of logical microtheories) - as a set of sentences encapsulating an LM’s core knowledge about a topic. These statements systematically work together to entail answers to a set of questions to both engender trust and improve performance. Our approach is to first populate a knowledge store with (model-generated) sentences that entail answers to training questions, and then distill those down to a core microtheory which is concise, general, and non-redundant. We show that, when added to a general corpus (e.g., Wikipedia), microtheories can supply critical information not necessarily present in the corpus, improving both a model’s ability to ground its answers to verifiable knowledge (i.e., show how answers are systematically entailed by documents in the corpus, grounding up to +8% more answers), and the accuracy of those grounded answers (up to +8% absolute). We also show that, in a human evaluation in the medical domain, our distilled microtheories contain a significantly higher concentration of topically critical facts than the non-distilled knowledge store. Finally, we show we can quantify the coverage of a microtheory for a topic (characterized by a dataset) using a notion of p-relevance. Together, these suggest that microtheories are an efficient distillation of an LM’s topic-relevant knowledge, that they can usefully augment existing corpora, and can provide both performance gains and an interpretable, verifiable window into the model’s knowledge of a topic. Nathaniel Weir, Bhavana Dalvi, Orion Weller, Oyvind Tafjord, Sam Hornstein, Alexander Sabol, Peter A. Jansen, Benjamin Van Durme, Peter Clark |
ICLR | 7 |
| 2024 | Enhancing Systematic Decompositional Natural Language Inference Using Informal LogicabstractNathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, Benjamin Van Durme. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nathaniel Weir, Kate Sanders 0002, Orion Weller, Shreya Sharma 0010, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi, Oyvind Tafjord, Peter A. Jansen, Peter Clark, Benjamin Van Durme |
EMNLP | 9 |
| 2024 | DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery AgentsabstractAutomated scientific discovery promises to accelerate progress across scientific domains, but evaluating an agent's capacity for end-to-end scientific reasoning is challenging as running real-world experiments is often prohibitively expensive or infeasible. In this work we introduce DiscoveryWorld, a virtual environment that enables benchmarking an agent's ability to perform complete cycles of novel scientific discovery in an inexpensive, simulated, multi-modal, long-horizon, and fictional setting.DiscoveryWorld consists of 24 scientific tasks across three levels of difficulty, each with parametric variations that provide new discoveries for agents to make across runs. Tasks require an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. Task difficulties are normed to range from straightforward to challenging for human scientists with advanced degrees. DiscoveryWorld further provides three automatic metrics for evaluating performance, including: (1) binary task completion, (2) fine-grained report cards detailing procedural scoring of task-relevant actions, and (3) the accuracy of discovered explanatory knowledge.While simulated environments such as DiscoveryWorld are low-fidelity compared to the real world, we find that strong baseline agents struggle on most DiscoveryWorld tasks, highlighting the utility of using simulated environments as proxy tasks for near-term development of scientific discovery competency in agents. Peter A. Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi, Bodhisattwa Prasad Majumder, Oyvind Tafjord, Peter Clark |
NeurIPS | 1 |
| 2023 | Behavior Cloned Transformers are Neurosymbolic ReasonersabstractIn this work, we explore techniques for augmenting interactive AI agents with information from symbolic modules, much like humans use tools like calculators and GPS systems to assist with arithmetic and navigation.We test our agent's abilities in text games-challenging benchmarks for evaluating the multi-step reasoning abilities of game agents in grounded, language-based environments.Our experimental study indicates that injecting the actions from these symbolic modules into the action space of a behavior cloned transformer agent increases performance on four text game benchmarks that test arithmetic, navigation, sorting, and common sense reasoning by an average of 22%, allowing an agent to reach the highest possible performance on unseen games.This action injection technique is easily extended to new agents, environments, and symbolic modules. 1 Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj Ammanabrolu |
EACL | 2 |
| 2023 | ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text GamesabstractIn this work we investigate the capacity of language models to generate explicit, interpretable, and interactive world models of scientific and common-sense reasoning tasks.We operationalize this as a task of generating text games, expressed as hundreds of lines of PYTHON code.To facilitate this task, we introduce BYTESIZED32 1 , a corpus of 32 reasoning-focused text games totalling 20k lines of PYTHON code.We empirically demonstrate that GPT-4 can use these games as templates for single-shot in-context learning, successfully producing runnable games on unseen topics in 28% of cases.When allowed to selfreflect on program errors, game runnability substantially increases to 57%.While evaluating simulation fidelity is labor intensive, we introduce a suite of automated metrics to assess game fidelity, technical validity, adherence to task specifications, and winnability, showing a high-degree of agreement with expert human ratings.We pose this as a challenge task to spur further development at the juncture of world modeling and code generation. Ruoyao Wang, Graham Todd, Xingdi Yuan, Ziang Xiao, Marc-Alexandre Côté, Peter A. Jansen |
EMNLP | 6 |
| 2022 | ScienceWorld: Is your Agent Smarter than a 5th Grader?abstractWe present SCIENCEWORLD, a benchmark to test agents' scientific reasoning abilities in a new interactive text environment at the level of a standard elementary school science curriculum.Despite the transformer-based progress seen in question-answering and scientific text processing, we find that current models cannot reason about or explain learned science concepts in novel contexts.For instance, models can easily answer what the conductivity of a known material is but struggle when asked how they would conduct an experiment in a grounded environment to find the conductivity of an unknown material.This begs the question of whether current models are simply retrieving answers by way of seeing a large number of similar examples or if they have learned to reason about concepts in a reusable manner.We hypothesize that agents need to be grounded in interactive environments to achieve such reasoning capabilities.Our experiments provide empirical evidence supporting this hypothesisshowing that a 1.5 million parameter agent trained interactively for 100k steps outperforms a 11 billion parameter model statically trained for scientific question-answering and reasoning from millions of expert demonstrations.12 Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj Ammanabrolu |
EMNLP | 2 |
| 2022 | Extracting Space Situational Awareness Events from News TextabstractSpace situational awareness typically makes use of physical measurements from radar, telescopes, and other assets to monitor satellites and other spacecraft for operational, navigational, and defense purposes. In this work we explore using textual input for the space situational awareness task. We construct a corpus of 48.5k news articles spanning all known active satellites between 2009 and 2020. Using a dependency-rule-based extraction system designed to target three high-impact events – spacecraft launches, failures, and decommissionings, we identify 1,787 space-event sentences that are then annotated by humans with 15.9k labels for event slots. We empirically demonstrate a state-of-the-art neural extraction system achieves an overall F1 between 53 and 91 per slot for event extraction in this low-resource, high-impact domain. Zhengnan Xie, Alice Saebom Kwak, Enfa George, Laura W. Dozal, Hoang Van, Moriba K. Jah, Roberto Furfaro, Peter A. Jansen |
LREC | 8 |
| 2021 | Explaining Answers with Entailment TreesabstractBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, Peter Clark. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Bhavana Dalvi, Peter A. Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, Peter Clark |
EMNLP (1) | 2 |
| 2021 | On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert RatingsabstractBuilding compositional explanations requires models to combine two or more facts that, together, describe why the answer to a question is correct.Typically, these "multi-hop" explanations are evaluated relative to one (or a small number of) gold explanations.In this work, we show these evaluations substantially underestimate model performance, both in terms of the relevance of included facts, as well as the completeness of model-generated explanations, because models regularly discover and produce valid explanations that are different than gold explanations.To address this, we construct a large corpus of 126k domain-expert (science teacher) relevance ratings that augment a corpus of explanations to standardized science exam questions, discovering 80k additional relevant facts not rated as gold.We build three strong models based on different methodologies (generation, ranking, and schemas), and empirically show that while expert-augmented ratings provide better estimates of explanation quality, both original (gold) and expertaugmented automatic evaluations still substantially underestimate performance by up to 36% when compared with full manual expert judgements, with different models being disproportionately affected.This poses a significant methodological challenge to accurately evaluating explanations produced by compositional reasoning models. Peter A. Jansen, Kelly J. Smith, Dan Moreno, Huitzilin Ortiz |
EMNLP (1) | 1 |
| 2020 | QASC: A Dataset for Question Answering via Sentence CompositionabstractComposing knowledge from multiple pieces of texts is a key challenge in multi-hop question answering. We present a multi-hop reasoning dataset, Question Answering via Sentence Composition (QASC), that requires retrieving facts from a large corpus and composing them to answer a multiple-choice question. QASC is the first dataset to offer two desirable properties: (a) the facts to be composed are annotated in a large corpus, and (b) the decomposition into these facts is not evident from the question itself. The latter makes retrieval challenging as the system must introduce new concepts or relations in order to discover potential decompositions. Further, the reasoning model must then learn to identify valid compositions of these retrieved facts using common-sense reasoning. To help address these challenges, we provide annotation for supporting facts as well as their composition. Guided by these annotations, we present a two-step approach to mitigate the retrieval challenges. We use other multiple-choice datasets as additional training data to strengthen the reasoning model. Our proposed approach improves over current state-of-the-art language models by 11% (absolute). The reasoning and retrieval problems, however, remain unsolved as this model still lags by 20% behind human performance. Tushar Khot, Peter Clark, Michal Guerquin, Peter A. Jansen, Ashish Sabharwal |
AAAI | 4 |
| 2020 | ScienceExamCER: A High-Density Fine-Grained Science-Domain Corpus for Common Entity RecognitionabstractNamed entity recognition identifies common classes of entities in text, but these entity labels are generally sparse, limiting utility to downstream tasks. In this work we present ScienceExamCER, a densely-labeled semantic classification corpus of 133k mentions in the science exam domain where nearly all (96%) of content words have been annotated with one or more fine-grained semantic class labels including taxonomic groups, meronym groups, verb/action groups, properties and values, and synonyms. Semantic class labels are drawn from a manually-constructed fine-grained typology of 601 classes generated through a data-driven analysis of 4,239 science exam questions. We show an off-the-shelf BERT-based named entity recognition model modified for multi-label classification achieves an accuracy of 0.85 F1 on this task, suggesting strong utility for downstream tasks in science domain question answering requiring densely-labeled semantic classification. Hannah Smith, Zeyu Zhang 0002, John Culnan, Peter A. Jansen |
LREC | 4 |
| 2020 | WorldTree V2: A Corpus of Science-Domain Structured Explanations and Inference Patterns supporting Multi-Hop InferenceabstractExplainable question answering for complex questions often requires combining large numbers of facts to answer a question while providing a human-readable explanation for the answer, a process known as multi-hop inference. Standardized science questions require combining an average of 6 facts, and as many as 16 facts, in order to answer and explain, but most existing datasets for multi-hop reasoning focus on combining only two facts, significantly limiting the ability of multi-hop inference algorithms to learn to generate large inferences. In this work we present the second iteration of the WorldTree project, a corpus of 5,114 standardized science exam questions paired with large detailed multi-fact explanations that combine core scientific knowledge and world knowledge. Each explanation is represented as a lexically-connected “explanation graph” that combines an average of 6 facts drawn from a semi-structured knowledge base of 9,216 facts across 66 tables. We use this explanation corpus to author a set of 344 high-level science domain inference patterns similar to semantic frames supporting multi-hop inference. Together, these resources provide training data and instrumentation for developing many-fact multi-hop inference models for question answering. Zhengnan Xie, Sebastian Thiem, Jaycie Martin, Elizabeth Wainwright, Steven Marmorstein, Peter A. Jansen |
LREC | 6 |
| 2020 | Multi-class Hierarchical Question Classification for Multiple Choice Science ExamsabstractPrior work has demonstrated that question classification (QC), recognizing the problem domain of a question, can help answer it more accurately. However, developing strong QC algorithms has been hindered by the limited size and complexity of annotated data available. To address this, we present the largest challenge dataset for QC, containing 7,787 science exam questions paired with detailed classification labels from a fine-grained hierarchical taxonomy of 406 problem domains. We then show that a BERT-based model trained on this dataset achieves a large (+0.12 MAP) gain compared with previous methods, while also achieving state-of-the-art performance on benchmark open-domain and biomedical QC datasets. Finally, we show that using this model’s predictions of question topic significantly improves the accuracy of a question answering system by +1.7% P@1, with substantial future gains possible as QC performance improves. Dongfang Xu, Peter A. Jansen, Jaycie Martin, Zhengnan Xie, Vikas Yadav, Harish Tayyar Madabushi, Oyvind Tafjord, Peter Clark |
LREC | 2 |
| 2018 | Controlling Information Aggregation for Complex Question Answering
Heeyoung Kwon, Harsh Trivedi, Peter A. Jansen, Mihai Surdeanu, Niranjan Balasubramanian |
ECIR | 3 |
| 2018 | WorldTree: A Corpus of Explanation Graphs for Elementary Science Questions supporting Multi-hop Inference
Peter A. Jansen, Elizabeth Wainwright, Steven Marmorstein, Clayton T. Morrison |
LREC | 1 |
| 2017 | Tell Me Why: Using Question Answering as Distant Supervision for Answer JustificationabstractRebecca Sharp, Mihai Surdeanu, Peter Jansen, Marco A. Valenzuela-Escárcega, Peter Clark, Michael Hammond. Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017). 2017. Rebecca Sharp, Mihai Surdeanu, Peter A. Jansen, Marco Antonio Valenzuela-Escárcega, Peter Clark, Michael Hammond |
CoNLL | 3 |
| 2017 | Framing QA as Building and Ranking Intersentence Answer JustificationsabstractWe propose a question answering (QA) approach for standardized science exams that both identifies correct answers and produces compelling human-readable justifications for why those answers are correct. Our method first identifies the actual information needed in a question using psycholinguistic concreteness norms, then uses this information need to construct answer justifications by aggregating multiple sentences from different knowledge bases using syntactic and lexical information. We then jointly rank answers and their justifications using a reranking perceptron that treats justification quality as a latent variable. We evaluate our method on 1,000 multiple-choice questions from elementary school science exams, and empirically demonstrate that it performs better than several strong baselines, including neural network approaches. Our best configuration answers 44% of the questions correctly, where the top justifications for 57% of these correct answers contain a compelling human-readable justification that explains the inference required to arrive at the correct answer. We include a detailed characterization of the justification quality for both our method and a strong baseline, and show that information aggregation is key to addressing the information need in complex questions. Peter A. Jansen, Rebecca Sharp, Mihai Surdeanu, Peter Clark |
Comput. Linguistics | 1 |
| 2016 | What's in an Explanation? Characterizing Knowledge and Inference Requirements for Elementary Science ExamsabstractQA systems have been making steady advances in the challenging elementary science exam domain. In this work, we develop an explanation-based analysis of knowledge and inference requirements, which supports a fine-grained characterization of the challenges. In particular, we model the requirements based on appropriate sources of evidence to be used for the QA task. We create requirements by first identifying suitable sentences in a knowledge base that support the correct answer, then use these to build explanations, filling in any necessary missing information. These explanations are used to create a fine-grained categorization of the requirements. Using these requirements, we compare a retrieval and an inference solver on 212 questions. The analysis validates the gains of the inference solver, demonstrating that it answers more questions requiring complex inference, while also providing insights into the relative strengths of the solvers and knowledge sources. We release the annotated questions and explanations as a resource with broad utility for science exam QA, including determining knowledge base construction targets, as well as supporting information aggregation in automated inference. Peter A. Jansen, Niranjan Balasubramanian, Mihai Surdeanu, Peter Clark |
COLING | 1 |
| 2016 | Creating Causal Embeddings for Question Answering with Minimal SupervisionabstractA common model for question answering (QA) is that a good answer is one that is closely related to the question, where relatedness is often determined using generalpurpose lexical models such as word embeddings.We argue that a better approach is to look for answers that are related to the question in a relevant way, according to the information need of the question, which may be determined through task-specific embeddings.With causality as a use case, we implement this insight in three steps.First, we generate causal embeddings cost-effectively by bootstrapping cause-effect pairs extracted from free text using a small set of seed patterns.Second, we train dedicated embeddings over this data, by using task-specific contexts, i.e., the context of a cause is its effect.Finally, we extend a state-of-the-art reranking approach for QA to incorporate these causal embeddings.We evaluate the causal embedding models both directly with a casual implication task, and indirectly, in a downstream causal QA task using data from Yahoo! Answers.We show that explicitly modeling causality improves performance in both tasks.In the QA task our best model achieves 37.3% P@1, significantly outperforming a strong baseline by 7.7% (relative). Rebecca Sharp, Mihai Surdeanu, Peter A. Jansen, Peter Clark, Michael Hammond |
EMNLP | 3 |
| 2015 | Spinning Straw into Gold: Using Free Text to Train Monolingual Alignment Models for Non-factoid Question AnsweringabstractRebecca Sharp, Peter Jansen, Mihai Surdeanu, Peter Clark. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Rebecca Sharp, Peter A. Jansen, Mihai Surdeanu, Peter Clark |
HLT-NAACL | 2 |
| 2015 | Higher-order Lexical Semantic Models for Non-factoid Answer RerankingabstractLexical semantic models provide robust performance for question answering, but, in general, can only capitalize on direct evidence seen during training. For example, monolingual alignment models acquire term alignment probabilities from semi-structured data such as question-answer pairs; neural network language models learn term embeddings from unstructured text. All this knowledge is then used to estimate the semantic similarity between question and answer candidates. We introduce a higher-order formalism that allows all these lexical semantic models to chain direct evidence to construct indirect associations between question and answer texts, by casting the task as the traversal of graphs that encode direct term associations. Using a corpus of 10,000 questions from Yahoo! Answers, we experimentally demonstrate that higher-order methods are broadly applicable to alignment and language models, across both word and syntactic representations. We show that an important criterion for success is controlling for the semantic drift that accumulates during graph traversal. All in all, the proposed higher-order approach improves five out of the six lexical semantic models investigated, with relative gains of up to +13% over their first-order variants. Daniel Fried, Peter A. Jansen, Gus Hahn-Powell, Mihai Surdeanu, Peter Clark |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Discourse Complements Lexical Semantics for Non-factoid Answer RerankingabstractWe propose a robust answer reranking model for non-factoid questions that integrates lexical semantics with discourse information, driven by two representations of discourse: a shallow representation centered around discourse markers, and a deep one based on Rhetorical Structure Theory.We evaluate the proposed model on two corpora from different genres and domains: one from Yahoo! Answers and one from the biology domain, and two types of non-factoid questions: manner and reason.We experimentally demonstrate that the discourse structure of nonfactoid answers provides information that is complementary to lexical semantic similarity between question and answer, improving performance up to 24% (relative) over a state-of-the-art model that exploits lexical semantic similarity alone.We further demonstrate excellent domain transfer of discourse information, suggesting these discourse features have general utility to non-factoid question answering. Peter A. Jansen, Mihai Surdeanu, Peter Clark |
ACL (1) | 1 |
| 2012 | Strong systematicity through sensorimotor conceptual grounding: an unsupervised, developmental approach to connectionist sentence processingabstractConnectionist language modelling typically has difficulty with syntactic systematicity, or the ability to generalise language learning to untrained sentences. This work develops an unsupervised connectionist model of infant grammar learning. Following the semantic boostrapping hypothesis, the network distils word category using a developmentally plausible infant-scale database of grounded sensorimotor conceptual representations, as well as a biologically plausible semantic co-occurrence activation function. The network then uses this knowledge to acquire an early benchmark clausal grammar using correlational learning, and further acquires separate conceptual and grammatical category representations. The network displays strongly systematic behaviour indicative of the general acquisition of the combinatorial systematicity present in the grounded infant-scale language stream, outperforms previous contemporary models that contain primarily noun and verb word categories, and successfully generalises broadly to novel untrained sensorimotor grounded sentences composed of unfamiliar nouns and verbs. Limitations as well as implications to later grammar learning are discussed. Peter A. Jansen, Scott Watter |
Connect. Sci. | 1 |