Harsh Trivedi

dblp:95/9586 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-1603-8591ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
abstract
Abstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco.
Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics2
2024 AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
abstract
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, Niranjan Balasubramanian. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Ashish Sabharwal, Niranjan Balasubramanian
ACL (1)1
2023 Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
abstract
Prompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA).They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters.While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA.Here, what to retrieve depends on what has already been derived, which in turn may depend on what was previously retrieved.To address this, we propose IRCoT, a new approach for multi-step QA that interleaves retrieval with steps (sentences) in a CoT, guiding the retrieval with CoT and in turn using retrieved results to improve CoT.Using IRCoT with GPT3 substantially improves retrieval (up to 21 points) as well as downstream QA (up to 15 points) on four datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC.We observe similar substantial gains in out-ofdistribution (OOD) settings as well as with much smaller models such as Flan-T5-large without additional training.IRCoT reduces model hallucination, resulting in factually more accurate CoT reasoning.1 .erdotii Nostri Primordia died?
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal
ACL (1)1
2023 Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal
ICLR2
2022 Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard Contexts
abstract
Question-answering datasets require a broad set of reasoning skills.We show how to use question decompositions to teach language models these broad reasoning skills in a robust fashion.Specifically, we use widely available QDMR representations to programmatically create hard-to-cheat synthetic contexts for real questions in six multi-step reasoning datasets.These contexts are carefully designed to avoid common reasoning shortcuts prevalent in real contexts that prevent models from learning the right skills.This results in a pretraining dataset, named TeaBReaC, containing 525K multi-step questions (with associated formal programs) covering about 900 reasoning patterns.We show that pretraining standard language models (LMs) on TeaBReaC before fine-tuning them on target datasets improves their performance by up to 13 F1 points across 4 multi-step QA datasets, with up to 21 point gain on more complex questions.The resulting models also demonstrate higher robustness, with a 5-8 F1 point improvement on two contrast sets.Furthermore, TeaBReaC pretraining substantially improves model performance and robustness even when starting with numerate LMs pretrained using recent methods (e.g., PReasM, POET).Our work thus shows how to effectively use decomposition-guided contexts to robustly teach multi-step reasoning.1
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal
EMNLP1
2022 ♫ MuSiQue: Multihop Questions via Single-hop Question Composition
abstract
Abstract Multihop reasoning remains an elusive goal as existing multihop benchmarks are known to be largely solvable via shortcuts. Can we create a question answering (QA) dataset that, by construction, requires proper multihop reasoning? To this end, we introduce a bottom–up approach that systematically selects composable pairs of single-hop questions that are connected, that is, where one reasoning step critically relies on information from another. This bottom–up methodology lets us explore a vast space of questions and add stringent filters as well as other mechanisms targeting connected reasoning. It provides fine-grained control over the construction process and the properties of the resulting k-hop questions. We use this methodology to create MuSiQue-Ans, a new multihop QA dataset with 25K 2–4 hop questions. Relative to existing datasets, MuSiQue-Ans is more difficult overall (3× increase in human–machine gap), and harder to cheat via disconnected reasoning (e.g., a single-hop model has a 30-point drop in F1). We further add unanswerable contrast questions to produce a more stringent dataset, MuSiQue-Full. We hope our datasets will help the NLP community develop models that perform genuine multihop reasoning.1
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal
Trans. Assoc. Comput. Linguistics1
2021 IrEne: Interpretable Energy Prediction for Transformers
abstract
Qingqing Cao, Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian
ACL/IJCNLP (1)3
2021 What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
abstract
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman
ACL/IJCNLP (1)3
2021 Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading Comprehension
abstract
How can we generate concise explanations for multi-hop Reading Comprehension (RC)?The current strategies of identifying supporting sentences can be seen as an extractive questionfocused summarization of the input text.However, these extractive explanations are not necessarily concise i.e. not minimally sufficient for answering a question.Instead, we advocate for an abstractive approach, where we propose to generate a question-focused, abstractive summary of input paragraphs and then feed it to an RC system.Given a limited amount of human-annotated abstractive explanations, we train the abstractive explainer in a semi-supervised manner, where we start from the supervised model and then train it further through trial and error maximizing a conciseness-promoted reward function.Our experiments demonstrate that the proposed abstractive explainer can generate more compact explanations than an extractive explainer with limited supervision (only 2k instances) while maintaining sufficiency.1 Our implementation is publicly available at https:// github.com/StonyBrookNLP/suqa.Charlie Rowe plays Billy Costa in a film based on what novel?[P1] [1] The Golden Compass is a 2007 British-American fantasy adventure film based on "Northern Lights", the first novel in Philip Pullman's trilogy "His Dark Materials".
Naoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian, Kentaro Inui
EMNLP (1)2
2020 DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering
abstract
Transformer-based QA models use input-wide self-attention -i.e.across both the question and the input passage -at all layers, causing them to be slow and memory-intensive.It turns out that we can get by without inputwide self-attention at all layers, especially in the lower layers.We introduce DeFormer, a decomposed transformer, which substitutes the full self-attention with question-wide and passage-wide self-attentions in the lower layers.This allows for question-independent processing of the input text representations, which in turn enables pre-computing passage representations reducing runtime compute drastically.Furthermore, because DeFormer is largely similar to the original model, we can initialize DeFormer with the pre-training weights of a standard transformer, and directly fine-tune on the target QA dataset.We show DeFormer versions of BERT and XLNet can be used to speed up QA by over 4.3x and with simple distillation-based losses they incur only a 1% drop in accuracy.We open source the code at https://github.com/ StonyBrookNLP/deformer.
Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian
ACL2
2020 Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected Reasoning
abstract
Has there been real progress in multi-hop question-answering?Models often exploit dataset artifacts to produce correct answers, without connecting information across multiple supporting facts.This limits our ability to measure true progress and defeats the purpose of building multi-hop QA datasets.We make three contributions towards addressing this.First, we formalize such undesirable behavior as disconnected reasoning across subsets of supporting facts.This allows developing a model-agnostic probe for measuring how much any model can cheat via disconnected reasoning.Second, using a notion of contrastive support sufficiency, we introduce an automatic transformation of existing datasets that reduces the amount of disconnected reasoning.Third, our experiments 1 suggest that there hasn't been much progress in multifact QA in the reading comprehension setting.For a recent large-scale model (XLNet), we show that only 18 points out of its answer F1 score of 72 on HotpotQA are obtained through multifact reasoning, roughly the same as that of a simpler RNN baseline.Our transformation substantially reduces disconnected reasoning (19 points in answer F1).It is complementary to adversarial approaches, yielding further reductions in conjunction.Original Dataset D ⇒ Question q = (Q, C; A) in D is assumed to be annotated with supporting facts {f 1 , f 2 }.Probing Dataset P ans+supp (D) for Answer Prediction and Support Identification tests: ⇒ Probing question collection P ans+supp (q) has only one group, corresponding to the unique bi-partition {{f 1 }, {f 2 }}, containing:Transformed Dataset T(D) for evaluating Constrastive Support Sufficiency: ⇒ Transformed question group T(q) in T(D) is defined using a single replacement fact f r ∈ C \ {f 1 , f 2 }:Probing Dataset P ans+supp+suff (T(D)) for all three tests: ⇒ Probing question collection P ans+supp+suff (T(q)) for the transformed question T(q) has only one group, corresponding to the unique bi-partition {{f 1 }, {f 2 }}, and is defined as:
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal
EMNLP (1)1
2018 Controlling Information Aggregation for Complex Question Answering
Heeyoung Kwon, Harsh Trivedi, Peter A. Jansen, Mihai Surdeanu, Niranjan Balasubramanian
ECIR2