VLDB 2026 Research / reviewers in the wild / expert
Tushar Khot
dblp:83/8117
· DBLP profile ↗
56ranked-venue papers
9as first author
21since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-authorDatabases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Theory of computation · 6Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of DocumentsabstractAbstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco. Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | DiscoveryBench: Towards Data-Driven Discovery with Large Language ModelsabstractCan the rapid advances in code generation, function calling, and data analysis using large language models (LLMs) help automate the search and verification of hypotheses purely from a set of provided datasets? To evaluate this question, we present DiscoveryBench, the first comprehensive benchmark that formalizes the multi-step process of data-driven discovery. The benchmark is designed to systematically assess current model capabilities in discovery tasks and provide a useful resource for improving them. Our benchmark contains 264 tasks collected across 6 diverse domains, such as sociology and engineering, by manually deriving discovery workflows from published papers to approximate the real-world challenges faced by researchers, where each task is defined by a dataset, its metadata, and a discovery goal in natural language. We additionally provide 903 synthetic tasks to conduct controlled evaluations on data-driven workflows that are not covered in the manually collected split. Furthermore, our structured formalism of data-driven discovery enables a facet-based evaluation that provides useful insights into different failure modes. We evaluate several popular LLM-based reasoning frameworks using both open and closed LLMs as baselines on DiscoveryBench and find that even the best system scores only 25%. Our benchmark, thus, illustrates the challenges in autonomous data-driven discovery and serves as a valuable resource for the community to make progress. Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal 0003, Bhavana Dalvi, Abhijeetsingh Meena, Aryan Prakhar, Tirth Vora, Tushar Khot, Ashish Sabharwal, Peter Clark |
ICLR | 8 |
| 2025 | Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task SupervisionabstractZhouhang Xie, Tushar Khot, Bhavana Dalvi Mishra, Harshit Surana, Julian McAuley, Peter Clark, Bodhisattwa Prasad Majumder. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhouhang Xie, Tushar Khot, Bhavana Dalvi, Harshit Surana, Julian J. McAuley, Peter Clark, Bodhisattwa Prasad Majumder |
NAACL (Long Papers) | 2 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 20 |
| 2024 | AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsabstractHarsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, Niranjan Balasubramanian. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Ashish Sabharwal, Niranjan Balasubramanian |
ACL (1) | 2 |
| 2024 | SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research RepositoriesabstractBen Bogin, Kejuan Yang, Shashank Gupta, Kyle Richardson, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ben Bogin, Kejuan Yang, Kyle Richardson 0001, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot |
EMNLP | 8 |
| 2024 | Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsabstractRecent works have showcased the ability of large-scale language models (LLMs) to embody diverse personas in their responses, exemplified by prompts like ‘_You are Yoda. Explain the Theory of Relativity._’ While this ability allows personalization of LLMs and enables human behavior simulation, its effect on LLMs’ capabilities remains unclear. To fill this gap, we present the first extensive study of the unintended side-effects of persona assignment on the ability of LLMs to perform _basic reasoning tasks_. Our study covers 24 reasoning datasets (spanning mathematics, law, medicine, morals, and more), 4 LLMs (2 versions of ChatGPT-3.5, GPT-4-Turbo, and Llama-2-70b-chat), and 19 diverse personas (e.g., ‘an Asian person’) spanning 5 socio-demographic groups: race, gender, religion, disability, and political affiliation. Our experiments unveil that LLMs harbor deep rooted bias against various socio-demographics underneath a veneer of fairness. While they overtly reject stereotypes when explicitly asked (‘_Are Black people less skilled at mathematics?_’), they manifest stereotypical and often erroneous presumptions when prompted to answer questions while adopting a persona. These can be observed as abstentions in the model’s response, e.g., ‘_As a Black person, I am unable to answer this question as it requires math knowledge_’, and generally result in a substantial drop in performance on reasoning tasks. Our experiments with ChatGPT-3.5 show that this bias is _ubiquitous_—80% of our personas demonstrate bias; it is _significant_—some datasets show performance drops of 70%+; and can be especially _harmful for certain groups_—some personas suffer statistically significant drops on 80%+ of the datasets. Overall, all four LLMs exhibit persona-induced bias to varying extents, with GPT-4-Turbo showing the least but still a problematic amount of bias (evident in 42% of the personas). Further analysis shows that these persona-induced errors can be hard-to-discern as they do not always manifest as explicit abstentions, and can also be hard-to-avoid—we find de-biasing prompts to have minimal to no effect. Our findings serve as a cautionary tale that the practice of assigning personas to LLMs—a trend on the rise—can surface their deep-rooted biases and have unforeseeable and detrimental side-effects. Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, Tushar Khot |
ICLR | 7 |
| 2024 | DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery AgentsabstractAutomated scientific discovery promises to accelerate progress across scientific domains, but evaluating an agent's capacity for end-to-end scientific reasoning is challenging as running real-world experiments is often prohibitively expensive or infeasible. In this work we introduce DiscoveryWorld, a virtual environment that enables benchmarking an agent's ability to perform complete cycles of novel scientific discovery in an inexpensive, simulated, multi-modal, long-horizon, and fictional setting.DiscoveryWorld consists of 24 scientific tasks across three levels of difficulty, each with parametric variations that provide new discoveries for agents to make across runs. Tasks require an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. Task difficulties are normed to range from straightforward to challenging for human scientists with advanced degrees. DiscoveryWorld further provides three automatic metrics for evaluating performance, including: (1) binary task completion, (2) fine-grained report cards detailing procedural scoring of task-relevant actions, and (3) the accuracy of discovered explanatory knowledge.While simulated environments such as DiscoveryWorld are low-fidelity compared to the real world, we find that strong baseline agents struggle on most DiscoveryWorld tasks, highlighting the utility of using simulated environments as proxy tasks for near-term development of scientific discovery competency in agents. Peter A. Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi, Bodhisattwa Prasad Majumder, Oyvind Tafjord, Peter Clark |
NeurIPS | 3 |
| 2023 | Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsabstractPrompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA).They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters.While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA.Here, what to retrieve depends on what has already been derived, which in turn may depend on what was previously retrieved.To address this, we propose IRCoT, a new approach for multi-step QA that interleaves retrieval with steps (sentences) in a CoT, guiding the retrieval with CoT and in turn using retrieved results to improve CoT.Using IRCoT with GPT3 substantially improves retrieval (up to 21 points) as well as downstream QA (up to 15 points) on four datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC.We observe similar substantial gains in out-ofdistribution (OOD) settings as well as with much smaller models such as Flan-T5-large without additional training.IRCoT reduces model hallucination, resulting in factually more accurate CoT reasoning.1 .erdotii Nostri Primordia died? Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
ACL (1) | 3 |
| 2023 | Complexity-Based Prompting for Multi-step Reasoning
Hao Peng 0018, Ashish Sabharwal, Peter Clark, Tushar Khot |
ICLR | 5 |
| 2023 | Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
ICLR | 1 |
| 2023 | Specializing Smaller Language Models towards Multi-Step ReasoningabstractThe surprising ability of Large Language Models (LLMs) to perform well on complex reasoning with only few-shot chain-of-thought prompts is believed to emerge only in very large-scale models. We show that such abilities can, in fact, be distilled down from GPT-3.5 (≥ 175B) to T5 variants (≤ 11B). We propose model specialization, to specialize the model’s ability towards a target task. The hypothesis is that large models (commonly viewed as larger than 100B) have strong modeling power such that they can perform a large spectrum of tasks. Small models (commonly viewed as smaller than 10B) have limited model capacity, but if we specialize their capacity towards a target task, the model can achieve decent performance improvements. We use multi-step math reasoning as our testbed because it is a very typical emergent ability. We show two important aspects of model abilities: (1) balancing language model’s performance on multiple tasks is a delicate matter, as improvements on one task may compromise other tasks; (2) yet by intentionally paying the price of decreased generic ability, we can clearly improve across different model scales smaller than 10B towards a specialized multi-step math reasoning ability. We further give comprehensive discussions about important design choices for better generalization, including the data format mixture and the start model checkpoint. We hope our practice and discoveries can serve as an important attempt towards specialized smaller models in the new research paradigm set by LLMs. Hao Peng 0018, Litu Ou, Ashish Sabharwal, Tushar Khot |
ICML | 5 |
| 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open ResourcesabstractIn this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu, our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.Our experiments show that different instruction-tuning datasets can uncover or enhance specific skills, while no single dataset (or combination) provides the best performance across all evaluations. Interestingly, we find that model and human preference-based evaluations fail to reflect differences in model capabilities exposed by benchmark-based evaluations, suggesting the need for the type of systemic evaluation performed in this work. Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and instruction-tuning data is required to close the gap. We release our instruction-tuned models, including a fully finetuned 65B Tülu, along with our code, data, and evaluation framework to facilitate future research. Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, Dave Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, Hannaneh Hajishirzi |
NeurIPS | 5 |
| 2022 | Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard ContextsabstractQuestion-answering datasets require a broad set of reasoning skills.We show how to use question decompositions to teach language models these broad reasoning skills in a robust fashion.Specifically, we use widely available QDMR representations to programmatically create hard-to-cheat synthetic contexts for real questions in six multi-step reasoning datasets.These contexts are carefully designed to avoid common reasoning shortcuts prevalent in real contexts that prevent models from learning the right skills.This results in a pretraining dataset, named TeaBReaC, containing 525K multi-step questions (with associated formal programs) covering about 900 reasoning patterns.We show that pretraining standard language models (LMs) on TeaBReaC before fine-tuning them on target datasets improves their performance by up to 13 F1 points across 4 multi-step QA datasets, with up to 21 point gain on more complex questions.The resulting models also demonstrate higher robustness, with a 5-8 F1 point improvement on two contrast sets.Furthermore, TeaBReaC pretraining substantially improves model performance and robustness even when starting with numerate LMs pretrained using recent methods (e.g., PReasM, POET).Our work thus shows how to effectively use decomposition-guided contexts to robustly teach multi-step reasoning.1 Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
EMNLP | 3 |
| 2022 | Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous PromptsabstractDaniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson 0001, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh 0001, Yejin Choi 0001 |
NAACL-HLT | 8 |
| 2022 | ♫ MuSiQue: Multihop Questions via Single-hop Question CompositionabstractAbstract Multihop reasoning remains an elusive goal as existing multihop benchmarks are known to be largely solvable via shortcuts. Can we create a question answering (QA) dataset that, by construction, requires proper multihop reasoning? To this end, we introduce a bottom–up approach that systematically selects composable pairs of single-hop questions that are connected, that is, where one reasoning step critically relies on information from another. This bottom–up methodology lets us explore a vast space of questions and add stringent filters as well as other mechanisms targeting connected reasoning. It provides fine-grained control over the construction process and the properties of the resulting k-hop questions. We use this methodology to create MuSiQue-Ans, a new multihop QA dataset with 25K 2–4 hop questions. Relative to existing datasets, MuSiQue-Ans is more difficult overall (3× increase in human–machine gap), and harder to cheat via disconnected reasoning (e.g., a single-hop model has a 30-point drop in F1). We further add unanswerable contrast questions to produce a more stringent dataset, MuSiQue-Full. We hope our datasets will help the NLP community develop models that perform genuine multihop reasoning.1 Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | ReadOnce Transformers: Reusable Representations of Text for TransformersabstractShih-Ting Lin, Ashish Sabharwal, Tushar Khot. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shih-Ting Lin, Ashish Sabharwal, Tushar Khot |
ACL/IJCNLP (1) | 3 |
| 2021 | Text Modular Networks: Learning to Decompose Tasks in the Language of Existing ModelsabstractTushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, Ashish Sabharwal. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tushar Khot, Daniel Khashabi, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
NAACL-HLT | 1 |
| 2021 | Temporal Reasoning on Implicit Events from Distant SupervisionabstractBen Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ben Zhou, Kyle Richardson 0001, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
NAACL-HLT | 4 |
| 2021 | Structure learning for relational logistic regression: an ensemble approach
Nandini Ramanan, Gautam Kunapuli, Tushar Khot, Bahare Fatemi, Mehran Kazemi, David Poole 0001, Kristian Kersting, Sriraam Natarajan |
Data Min. Knowl. Discov. | 3 |
| 2021 | Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning StrategiesabstractAbstract A key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce StrategyQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies. We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, StrategyQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in StrategyQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of ∼ 66%. Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | QASC: A Dataset for Question Answering via Sentence CompositionabstractComposing knowledge from multiple pieces of texts is a key challenge in multi-hop question answering. We present a multi-hop reasoning dataset, Question Answering via Sentence Composition (QASC), that requires retrieving facts from a large corpus and composing them to answer a multiple-choice question. QASC is the first dataset to offer two desirable properties: (a) the facts to be composed are annotated in a large corpus, and (b) the decomposition into these facts is not evident from the question itself. The latter makes retrieval challenging as the system must introduce new concepts or relations in order to discover potential decompositions. Further, the reasoning model must then learn to identify valid compositions of these retrieved facts using common-sense reasoning. To help address these challenges, we provide annotation for supporting facts as well as their composition. Guided by these annotations, we present a two-step approach to mitigate the retrieval challenges. We use other multiple-choice datasets as additional training data to strengthen the reasoning model. Our proposed approach improves over current state-of-the-art language models by 11% (absolute). The reasoning and retrieval problems, however, remain unsolved as this model still lags by 20% behind human performance. Tushar Khot, Peter Clark, Michal Guerquin, Peter A. Jansen, Ashish Sabharwal |
AAAI | 1 |
| 2020 | IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsabstractHumans often have to read multiple documents to address their information needs.However, most existing reading comprehension (RC) tasks only focus on questions for which the contexts provide all the information required to answer them, thus not evaluating a system's performance at identifying a potential lack of sufficient information and locating sources for that information.To fill this gap, we present a dataset, IIRC, with more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.The questions were written by crowd workers who did not have access to any of the linked documents, leading to questions that have little lexical overlap with the contexts where the answers appear.This process also gave many questions without answers, and those that require discrete reasoning, increasing the difficulty of the task.We follow recent modeling work on various reading comprehension datasets to construct a baseline model for this dataset, finding that it achieves 31.1% F1 on this task, while estimated human performance is 88.4%.The dataset, code for the baseline system, and a leaderboard can be found at https://allennlp.org/iirc. James Ferguson, Matt Gardner 0001, Hannaneh Hajishirzi, Tushar Khot, Pradeep Dasigi |
EMNLP (1) | 4 |
| 2020 | A Simple Yet Strong Pipeline for HotpotQAabstractState-of-the-art models for multi-hop question answering typically augment large-scale language models like BERT with additional, intuitively useful capabilities such as named entity recognition, graph-based reasoning, and question decomposition.However, does their strong performance on popular multihop datasets really justify this added design complexity?Our results suggest that the answer may be no, because even our simple pipeline based on BERT, named QUARK, performs surprisingly well.Specifically, on Hot-potQA, QUARK outperforms these models on both question answering and support identification (and achieves performance very close to a RoBERTa model).Our pipeline has three steps: 1) use BERT to identify potentially relevant sentences independently of each other; 2) feed the set of selected sentences as context into a standard BERT span prediction model to choose an answer; and 3) use the sentence selection model, now with the chosen answer, to produce supporting sentences.The strong performance of QUARK resurfaces the importance of carefully exploring simple model designs before using popular benchmarks to justify the value of complex techniques. Dirk Groeneveld, Tushar Khot, Mausam, Ashish Sabharwal |
EMNLP (1) | 2 |
| 2020 | More Bang for Your Buck: Natural Perturbation for Robust Question AnsweringabstractDeep learning models for linguistic tasks require large training datasets, which are expensive to create.As an alternative to the traditional approach of creating new instances by repeating the process of creating one instance, we propose doing so by first collecting a set of seed examples and then applying humandriven natural perturbations (as opposed to rule-based machine perturbations), which often change the gold label as well.Such perturbations have the advantage of being relatively easier (and hence cheaper) to create than writing out completely new examples.Further, they help address the issue that even models achieving human-level scores on NLP datasets are known to be considerably sensitive to small changes in input.To evaluate the idea, we consider a recent question-answering dataset (BOOLQ) and study our approach as a function of the perturbation cost ratio, the relative cost of perturbing an existing question vs. creating a new one from scratch.We find that when natural perturbations are moderately cheaper to create (cost ratio under 60%), it is more effective to use them for training BOOLQ models: such models exhibit 9% higher robustness and 4.5% stronger generalization, while retaining performance on the original BOOLQ dataset. Daniel Khashabi, Tushar Khot, Ashish Sabharwal |
EMNLP (1) | 2 |
| 2020 | Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected ReasoningabstractHas there been real progress in multi-hop question-answering?Models often exploit dataset artifacts to produce correct answers, without connecting information across multiple supporting facts.This limits our ability to measure true progress and defeats the purpose of building multi-hop QA datasets.We make three contributions towards addressing this.First, we formalize such undesirable behavior as disconnected reasoning across subsets of supporting facts.This allows developing a model-agnostic probe for measuring how much any model can cheat via disconnected reasoning.Second, using a notion of contrastive support sufficiency, we introduce an automatic transformation of existing datasets that reduces the amount of disconnected reasoning.Third, our experiments 1 suggest that there hasn't been much progress in multifact QA in the reading comprehension setting.For a recent large-scale model (XLNet), we show that only 18 points out of its answer F1 score of 72 on HotpotQA are obtained through multifact reasoning, roughly the same as that of a simpler RNN baseline.Our transformation substantially reduces disconnected reasoning (19 points in answer F1).It is complementary to adversarial approaches, yielding further reductions in conjunction.Original Dataset D ⇒ Question q = (Q, C; A) in D is assumed to be annotated with supporting facts {f 1 , f 2 }.Probing Dataset P ans+supp (D) for Answer Prediction and Support Identification tests: ⇒ Probing question collection P ans+supp (q) has only one group, corresponding to the unique bi-partition {{f 1 }, {f 2 }}, containing:Transformed Dataset T(D) for evaluating Constrastive Support Sufficiency: ⇒ Transformed question group T(q) in T(D) is defined using a single replacement fact f r ∈ C \ {f 1 , f 2 }:Probing Dataset P ans+supp+suff (T(D)) for all three tests: ⇒ Probing question collection P ans+supp+suff (T(q)) for the transformed question T(q) has only one group, corresponding to the unique bi-partition {{f 1 }, {f 2 }}, and is defined as: Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
EMNLP (1) | 3 |
| 2019 | Exploiting Explicit Paths for Multi-hop Reading ComprehensionabstractWe propose a novel, path-based reasoning approach for the multi-hop reading comprehension task where a system needs to combine facts from multiple passages to answer a question.Although inspired by multi-hop reasoning over knowledge graphs, our proposed approach operates directly over unstructured text.It generates potential paths through passages and scores them without any direct path supervision.The proposed model, named PathNet, attempts to extract implicit relations from text through entity pair representations, and compose them to encode each path.To capture additional context, Path-Net also composes the passage representations along each path to compute a passage-based representation.Unlike previous approaches, our model is then able to explain its reasoning via these explicit paths through the passages.We show that our approach outperforms prior models on the multi-hop Wikihop dataset, and also can be generalized to apply to the OpenBookQA dataset, matching stateof-the-art performance. Souvik Kundu 0003, Tushar Khot, Ashish Sabharwal, Peter Clark |
ACL (1) | 2 |
| 2019 | What's Missing: A Knowledge Gap Guided Approach for Multi-hop Question AnsweringabstractTushar Khot, Ashish Sabharwal, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tushar Khot, Ashish Sabharwal, Peter Clark |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Question Answering as Global Reasoning Over Semantic AbstractionsabstractWe propose a novel method for exploiting the semantic structure of text to answer multiple-choice questions. The approach is especially suitable for domains that require reasoning over a diverse set of linguistic constructs but have limited training data. To address these challenges, we present the first system, to the best of our knowledge, that reasons over a wide range of semantic abstractions of the text, which are derived using off-the-shelf, general-purpose, pre-trained natural language modules such as semantic role labelers, coreference resolvers, and dependency parsers. Representing multiple abstractions as a family of graphs, we translate question answering (QA) into a search for an optimal subgraph that satisfies certain global and local properties. This formulation generalizes several prior structured QA systems. Our system, SEMANTICILP, demonstrates strong performance on two domains simultaneously. In particular, on a collection of challenging science QA datasets, it outperforms various state-of-the-art approaches, including neural models, broad coverage information retrieval, and specialized techniques using structured knowledge bases, by 2%-6%. Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
AAAI | 2 |
| 2018 | SciTaiL: A Textual Entailment Dataset from Science Question AnsweringabstractWe present a new dataset and model for textual entailment, derived from treating multiple-choice question-answering as an entailment problem. SciTail is the first entailment set that is created solely from natural sentences that already exist independently ``in the wild'' rather than sentences authored specifically for the entailment task. Different from existing entailment datasets, we create hypotheses from science questions and the corresponding answer candidates, and premises from relevant web sentences retrieved from a large corpus. These sentences are often linguistically challenging. This, combined with the high lexical similarity of premise and hypothesis for both entailed and non-entailed pairs, makes this new entailment task particularly difficult. The resulting challenge is evidenced by state-of-the-art textual entailment systems achieving mediocre performance on SciTail, especially in comparison to a simple majority class baseline. As a step forward, we demonstrate that one can improve accuracy on SciTail by 5% using a new neural model that exploits linguistic structure. Tushar Khot, Ashish Sabharwal, Peter Clark |
AAAI | 1 |
| 2018 | AdvEntuRe: Adversarial Training for Textual Entailment with Knowledge-Guided ExamplesabstractWe consider the problem of learning textual entailment models with limited supervision (5K-10K training examples), and present two complementary approaches for it.First, we propose knowledge-guided adversarial example generators for incorporating large lexical resources in entailment models via only a handful of rule templates.Second, to make the entailment model-a discriminator-more robust, we propose the first GAN-style approach for training it using a natural language example generator that iteratively adjusts based on the discriminator's performance.We demonstrate effectiveness using two entailment datasets, where the proposed methods increase accuracy by 4.7% on SciTail and by 2.8% on a 1% training sub-sample of SNLI.Notably, even a single hand-written rule, negate, improves the accuracy on the negation examples in SNLI by 6.1%.P: The dog did not eat all of the chickens.H: The dog ate all of the chickens.S: entails (score 56:5%) P: The red box is in the blue box.H: The blue box is in the red box. Dongyeop Kang, Tushar Khot, Ashish Sabharwal, Eduard H. Hovy |
ACL (1) | 2 |
| 2018 | Bridging Knowledge Gaps in Neural Entailment via Symbolic ModelsabstractMost textual entailment models focus on lexical gaps between the premise text and the hypothesis, but rarely on knowledge gaps.We focus on filling these knowledge gaps in the Science Entailment task, by leveraging an external structured knowledge base (KB) of science facts.Our new architecture combines standard neural entailment models with a knowledge lookup module.To facilitate this lookup, we propose a fact-level decomposition of the hypothesis, and verifying the resulting sub-facts against both the textual premise and the structured KB.Our model, NSnet, learns to aggregate predictions from these heterogeneous data formats.On the SciTail dataset, NSnet outperforms a simpler combination of the two predictions by 3% and the base entailment model by 5%. Dongyeop Kang, Tushar Khot, Ashish Sabharwal, Peter Clark |
EMNLP | 2 |
| 2018 | Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question AnsweringabstractWe present a new kind of question answering dataset, OpenBookQA, modeled after open book exams for assessing human understanding of a subject.The open book that comes with our questions is a set of 1326 elementary level science facts.Roughly 6000 questions probe an understanding of these facts and their application to novel situations.This requires combining an open book fact (e.g., metals conduct electricity) with broad common knowledge (e.g., a suit of armor is made of metal) obtained from other sources.While existing QA datasets over documents or knowledge bases, being generally self-contained, focus on linguistic understanding, OpenBookQA probes a deeper understanding of both the topic-in the context of common knowledge-and the language it is expressed in.Human performance on OpenBookQA is close to 92%, but many state-of-the-art pre-trained QA methods perform surprisingly poorly, worse than several simple neural baselines we develop.Our oracle experiments designed to circumvent the knowledge retrieval bottleneck demonstrate the value of both the open book and additional facts.We leave it as a challenge to solve the retrieval problem in this multi-hop setting and to close the large gap to human performance.Question: Which of these would let the most heat travel through?A) a new pair of jeans.B) a steel spoon in a cafeteria.C) a cotton candy at a store.D) a calvin klein cotton hat. Todor Mihaylov, Peter Clark, Tushar Khot, Ashish Sabharwal |
EMNLP | 3 |
| 2018 | Structure Learning for Relational Logistic Regression: An Ensemble Approach
Nandini Ramanan, Gautam Kunapuli, Tushar Khot, Bahare Fatemi, Mehran Kazemi, David Poole 0001, Kristian Kersting, Sriraam Natarajan |
KR | 3 |
| 2017 | Learning What is Essential in QuestionsabstractQuestion answering (QA) systems are easily distracted by irrelevant or redundant words in questions, especially when faced with long or multi-sentence questions in difficult domains. This paper introduces and studies the notion of essential question terms with the goal of improving such QA solvers. We illustrate the importance of essential question terms by showing that humans' ability to answer questions drops significantly when essential terms are eliminated from questions.We then develop a classifier that reliably (90% mean average precision) identifies and ranks essential terms in questions. Finally, we use the classifier to demonstrate that the notion of question term essentiality allows state-of-the-art QA solver for elementary-level science questions to make better and more informed decisions,improving performance by up to 5%.We also introduce a new dataset of over 2,200 crowd-sourced essential terms annotated science questions. Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
CoNLL | 2 |
| 2017 | Relational Restricted Boltzmann Machines: A Probabilistic Logic Learning Approach
Gautam Kunapuli, Tushar Khot, Kristian Kersting, William Cohen, Sriraam Natarajan |
ILP | 3 |
| 2017 | Markov logic networks for adverse drug event extraction from text
Sriraam Natarajan, Vishal Bangera, Tushar Khot, Jose Picado, Anurag Wazalwar, Vítor Santos Costa, David Page, Michael Caldwell |
Knowl. Inf. Syst. | 3 |
| 2016 | Combining Retrieval, Statistics, and Inference to Answer Elementary Science QuestionsabstractWhat capabilities are required for an AI system to pass standard 4th Grade Science Tests? Previous work has examined the use of Markov Logic Networks (MLNs) to represent the requisite background knowledge and interpret test questions, but did not improve upon an information retrieval (IR) baseline. In this paper, we describe an alternative approach that operates at three levels of representation and reasoning: information retrieval, corpus statistics, and simple inference over a semi-automatically constructed knowledge base, to achieve substantially improved results. We evaluate the methods on six years of unseen, unedited exam questions from the NY Regents Science Exam (using only non-diagram, multiple choice questions), and show that our overall system’s score is 71.3%, an improvement of 23.8% (absolute) over the MLN-based method described in previous work. We conclude with a detailed analysis, illustrating the complementary strengths of each method in the ensemble. Our datasets are being released to enable further research. Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter D. Turney, Daniel Khashabi |
AAAI | 3 |
| 2016 | Learning Continuous-Time Bayesian Networks in Relational Domains: A Non-Parametric ApproachabstractMany real world applications in medicine, biology, communication networks, web mining, and economics, among others, involve modeling and learning structured stochastic processes that evolve over continuous time. Existing approaches, however, have focused on propositional domains only. Without extensive feature engineering, it is difficult-if not impossible-to apply them within relational domains where we may have varying number of objects and relations among them. We therefore develop the first relational representation called Relational Continuous-Time Bayesian Networks (RCTBNs) that can address this challenge. It features a nonparametric learning method that allows for efficiently learning the complex dependencies and their strengths simultaneously from sequence data. Our experimental results demonstrate that RCTBNs can learn as effectively as state-of-the-art approaches for propositional tasks while modeling relational tasks faithfully. Shuo Yang 0004, Tushar Khot, Kristian Kersting, Sriraam Natarajan |
AAAI | 2 |
| 2016 | Question Answering via Integer Programming over Semi-Structured Knowledge
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni, Dan Roth 0001 |
IJCAI | 2 |
| 2016 | Inductive Logic Programming Meets Relational Databases: Efficient Learning of Markov Logic Networks
Marcin Malec, Tushar Khot, James G. Nagy, Erik Blask, Sriraam Natarajan |
ILP | 2 |
| 2016 | Scaling Lifted Probabilistic Inference and Learning Via Graph DatabasesabstractOver the past decade, exploiting relations and symmetries within probabilistic models has been proven to be surprisingly effective at solving large scale data mining problems. One of the key operations inside these lifted approaches is counting - be it for parameter/structure learning or for efficient inference. Typically, however, they just count exploiting the logical structure using adhoc operators. This paper investigates whether ‘Compilation to Graph Databases’ could be a practical technique for scaling lifted probabilistic inference and learning methods. We demonstrate that the proposed approach achieves reasonable speed-ups for both inference and learning, without sacrificing performance. Mayukh Das, Yuqing Wu, Tushar Khot, Kristian Kersting, Sriraam Natarajan |
SDM | 3 |
| 2015 | Knowledge-Based Probabilistic Logic LearningabstractAdvice giving has been long explored in artificial intelligence to build robust learning algorithms. We consider advice giving in relational domains where the noise is systematic. The advice is provided as logical statements that are then explicitly considered by the learning algorithm at every update. Our empirical evidence proves that human advice can effectively accelerate learning in noisy structured domains where so far humans have been merely used as labelers or as designers of initial structure of the model. Phillip Odom, Tushar Khot, Reid B. Porter, Sriraam Natarajan |
AAAI | 2 |
| 2015 | Extracting Adverse Drug Events from Text Using Human Advice
Phillip Odom, Vishal Bangera, Tushar Khot, David Page, Sriraam Natarajan |
AIME | 3 |
| 2015 | Exploring Markov Logic Networks for Question AnsweringabstractElementary-level science exams pose sig-nificant knowledge acquisition and rea-soning challenges for automatic question answering. We develop a system that rea-sons with knowledge derived from text-books, represented in a subset of first-order logic. Automatic extraction, while scalable, often results in knowledge that is incomplete and noisy, motivating use of reasoning mechanisms that handle uncer-tainty. Markov Logic Networks (MLNs) seem a natural model for expressing such knowl-edge, but the exact way of leveraging MLNs is by no means obvious. We in-vestigate three ways of applying MLNs to our task. First, we simply use the extracted science rules directly as MLN clauses and exploit the structure present in hard con-straints to improve tractability. Second, we interpret science rules as describing prototypical entities, resulting in a drasti-cally simplified but brittle network. Our third approach, called Praline, uses MLNs to align lexical elements as well as define and control how inference should be per-formed in this task. Praline demonstrates a 15 % accuracy boost and a 10x reduction in runtime as compared to other MLN-based methods, and comparable accuracy to word-based baseline approaches. Tushar Khot, Niranjan Balasubramanian, Eric Gribkoff, Ashish Sabharwal, Peter Clark, Oren Etzioni |
EMNLP | 1 |
| 2015 | Gradient-based boosting for statistical relational learning: the Markov logic network and missing data cases
Tushar Khot, Sriraam Natarajan, Kristian Kersting, Jude W. Shavlik |
Mach. Learn. | 1 |
| 2014 | Relational One-Class Classification: A Non-Parametric ApproachabstractOne-class classification approaches have been proposed in the literature to learn classifiers from examples of only one class. But these approaches are not directly applicable to relational domains due to their reliance on a feature vector or a distance measure. We propose a non-parametric relational one-class classification approach based on first-order trees. We learn a tree-based distance measure that iteratively introduces new relational features to differentiate relational examples. We update the distance measure so as to maximize the one-class classification performance of our model. We also relate our model definition to existing work on probabilistic combination functions and density estimation. We experimentally show that our approach can discover relevant features and outperform three baseline approaches. Tushar Khot, Sriraam Natarajan, Jude W. Shavlik |
AAAI | 1 |
| 2014 | Learning from Imbalanced Data in Relational Domains: A Soft Margin ApproachabstractWe consider the problem of learning probabilistic models from relational data. One of the key issues with relational data is class imbalance where the number of negative examples far outnumbers the number of positive examples. The common approach for dealing with this problem is the use of sub-sampling of negative examples. We, on the other hand, consider a soft margin approach that explicitly trades off between the false positives and false negatives. We apply this approach to the recently successful formalism of relational functional gradient boosting. Specifically, we modify the objective function of the learning problem to explicitly include the trade-off between false positives and negatives. We show empirically that this approach is more successful in handling the class imbalance problem than the original framework that weighed all the examples equally. Shuo Yang 0004, Tushar Khot, Kristian Kersting, Gautam Kunapuli, Kris Hauser, Sriraam Natarajan |
ICDM | 2 |
| 2014 | Effectively Creating Weakly Labeled Training Examples via Approximate Domain Knowledge
Sriraam Natarajan, Jose Picado, Tushar Khot, Kristian Kersting, Christopher Ré, Jude W. Shavlik |
ILP | 3 |
| 2014 | Statistical Relational Learning for Handwriting Recognition
Arti Shivram, Tushar Khot, Sriraam Natarajan, Venu Govindaraju |
ILP | 2 |
| 2013 | Accelerating Imitation Learning in Relational Domains via Transfer by Initialization
Sriraam Natarajan, Phillip Odom, Saket Joshi, Tushar Khot, Kristian Kersting, Prasad Tadepalli |
ILP | 4 |
| 2012 | A Machine Learning Pipeline for Three-Way Classification of Alzheimer Patients from Structural Magnetic Resonance Images of the BrainabstractMagnetic resonance imaging (MRI) has emerged as an important tool to identify intermediate biomarkers of Alzheimer's disease (AD) due to its ability to measure regional changes in the brain that are thought to reflect disease severity and progression. In this paper, we set out a novel pipeline that uses volumetric MRI data collected from different subjects as input and classifies them into one of three classes: AD, mild cognitive impairment (MCI) and cognitively normal (CN). Our pipeline consists of three stages -- (1) a segmentation layer where brain MRI data is divided into clinically relevant regions, (2) a classification layer that uses relational learning algorithms to make pair wise predictions between the three classes, and (3)a combination layer that combines the results of the different classes to obtain the final classification. One of the key features of our proposed approach is that it allows for domain expert's knowledge to guide the learning in all the layers. We evaluate our pipeline on 397 patients acquired from the Alzheimer's Disease Neuroimaging Initiative and demonstrate that it obtains state-of the-art performance with minimal feature engineering. Sriraam Natarajan, Saket Joshi, Baidya Nath Saha, Adam Edwards, Tushar Khot, Elizabeth M. Davenport, Kristian Kersting, Christopher T. Whitlow, Joseph A. Maldjian |
ICMLA (1) | 5 |
| 2012 | Gradient-based boosting for statistical relational learning: The relational dependency network case
Sriraam Natarajan, Tushar Khot, Kristian Kersting, Bernd Gutmann, Jude W. Shavlik |
Mach. Learn. | 2 |
| 2011 | Learning Markov Logic Networks via Functional Gradient BoostingabstractRecent years have seen a surge of interest in Statistical Relational Learning (SRL) models that combine logic with probabilities. One prominent example is Markov Logic Networks (MLNs). While MLNs are indeed highly expressive, this expressiveness comes at a cost. Learning MLNs is a hard problem and therefore has attracted much interest in the SRL community. Current methods for learning MLNs follow a two-step approach: first, perform a search through the space of possible clauses and then learn appropriate weights for these clauses. We propose to take a different approach, namely to learn both the weights and the structure of the MLN simultaneously. Our approach is based on functional gradient boosting where the problem of learning MLNs is turned into a series of relational functional approximation problems. We use two kinds of representations for the gradients: clause-based and tree-based. Our experimental evaluation on several benchmark data sets demonstrates that our new approach can learn MLNs as good or better than those found with state-of-the-art methods, but often in a fraction of the time. Tushar Khot, Sriraam Natarajan, Kristian Kersting, Jude W. Shavlik |
ICDM | 1 |
| 2010 | Exploiting Causal Independence in Markov Logic Networks: Combining Undirected and Directed Models
Sriraam Natarajan, Tushar Khot, Daniel Lowd, Prasad Tadepalli, Kristian Kersting, Jude W. Shavlik |
ECML/PKDD (2) | 2 |
| 2009 | Some new directions in graph-based semi-supervised learningabstractIn this position paper, we first review the state-of-the-art in graph-based semi-supervised learning, and point out three limitations that are particularly relevant to multimedia analysis: (1) rich data is restricted to live on a single manifold; (2) learning must happen in batch mode; and (3) the target label is assumed smooth on the manifold. We then discuss new directions in semi-supervised learning research that can potentially overcome these limitations: (i) modeling data as a mixture of multiple manifolds that may intersect or overlap; (ii) online semi-supervised learning that learns incrementally with low computation and memory needs; and (iii) learning spectrally sparse but non-smooth labels with compressive sensing. We give concrete examples in each new direction. We hope this article will inspire new research that makes semi-supervised learning an even more valuable tool for multimedia analysis. Xiaojin Zhu 0001, Andrew B. Goldberg, Tushar Khot |
ICME | 3 |