VLDB 2026 Research / reviewers in the wild / expert
Harsh Trivedi
dblp:95/9586
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-1603-8591ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of DocumentsabstractAbstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco. Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsabstractHarsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, Niranjan Balasubramanian. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Ashish Sabharwal, Niranjan Balasubramanian |
ACL (1) | 1 |
| 2023 | Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsabstractPrompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA).They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters.While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA.Here, what to retrieve depends on what has already been derived, which in turn may depend on what was previously retrieved.To address this, we propose IRCoT, a new approach for multi-step QA that interleaves retrieval with steps (sentences) in a CoT, guiding the retrieval with CoT and in turn using retrieved results to improve CoT.Using IRCoT with GPT3 substantially improves retrieval (up to 21 points) as well as downstream QA (up to 15 points) on four datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC.We observe similar substantial gains in out-ofdistribution (OOD) settings as well as with much smaller models such as Flan-T5-large without additional training.IRCoT reduces model hallucination, resulting in factually more accurate CoT reasoning.1 .erdotii Nostri Primordia died? Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
ACL (1) | 1 |
| 2023 | Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
ICLR | 2 |
| 2022 | Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard ContextsabstractQuestion-answering datasets require a broad set of reasoning skills.We show how to use question decompositions to teach language models these broad reasoning skills in a robust fashion.Specifically, we use widely available QDMR representations to programmatically create hard-to-cheat synthetic contexts for real questions in six multi-step reasoning datasets.These contexts are carefully designed to avoid common reasoning shortcuts prevalent in real contexts that prevent models from learning the right skills.This results in a pretraining dataset, named TeaBReaC, containing 525K multi-step questions (with associated formal programs) covering about 900 reasoning patterns.We show that pretraining standard language models (LMs) on TeaBReaC before fine-tuning them on target datasets improves their performance by up to 13 F1 points across 4 multi-step QA datasets, with up to 21 point gain on more complex questions.The resulting models also demonstrate higher robustness, with a 5-8 F1 point improvement on two contrast sets.Furthermore, TeaBReaC pretraining substantially improves model performance and robustness even when starting with numerate LMs pretrained using recent methods (e.g., PReasM, POET).Our work thus shows how to effectively use decomposition-guided contexts to robustly teach multi-step reasoning.1 Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
EMNLP | 1 |
| 2022 | ♫ MuSiQue: Multihop Questions via Single-hop Question CompositionabstractAbstract Multihop reasoning remains an elusive goal as existing multihop benchmarks are known to be largely solvable via shortcuts. Can we create a question answering (QA) dataset that, by construction, requires proper multihop reasoning? To this end, we introduce a bottom–up approach that systematically selects composable pairs of single-hop questions that are connected, that is, where one reasoning step critically relies on information from another. This bottom–up methodology lets us explore a vast space of questions and add stringent filters as well as other mechanisms targeting connected reasoning. It provides fine-grained control over the construction process and the properties of the resulting k-hop questions. We use this methodology to create MuSiQue-Ans, a new multihop QA dataset with 25K 2–4 hop questions. Relative to existing datasets, MuSiQue-Ans is more difficult overall (3× increase in human–machine gap), and harder to cheat via disconnected reasoning (e.g., a single-hop model has a 30-point drop in F1). We further add unanswerable contrast questions to produce a more stringent dataset, MuSiQue-Full. We hope our datasets will help the NLP community develop models that perform genuine multihop reasoning.1 Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | IrEne: Interpretable Energy Prediction for TransformersabstractQingqing Cao, Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian |
ACL/IJCNLP (1) | 3 |
| 2021 | What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?abstractNikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman |
ACL/IJCNLP (1) | 3 |
| 2021 | Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading ComprehensionabstractHow can we generate concise explanations for multi-hop Reading Comprehension (RC)?The current strategies of identifying supporting sentences can be seen as an extractive questionfocused summarization of the input text.However, these extractive explanations are not necessarily concise i.e. not minimally sufficient for answering a question.Instead, we advocate for an abstractive approach, where we propose to generate a question-focused, abstractive summary of input paragraphs and then feed it to an RC system.Given a limited amount of human-annotated abstractive explanations, we train the abstractive explainer in a semi-supervised manner, where we start from the supervised model and then train it further through trial and error maximizing a conciseness-promoted reward function.Our experiments demonstrate that the proposed abstractive explainer can generate more compact explanations than an extractive explainer with limited supervision (only 2k instances) while maintaining sufficiency.1 Our implementation is publicly available at https:// github.com/StonyBrookNLP/suqa.Charlie Rowe plays Billy Costa in a film based on what novel?[P1] [1] The Golden Compass is a 2007 British-American fantasy adventure film based on "Northern Lights", the first novel in Philip Pullman's trilogy "His Dark Materials". Naoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian, Kentaro Inui |
EMNLP (1) | 2 |
| 2020 | DeFormer: Decomposing Pre-trained Transformers for Faster Question AnsweringabstractTransformer-based QA models use input-wide self-attention -i.e.across both the question and the input passage -at all layers, causing them to be slow and memory-intensive.It turns out that we can get by without inputwide self-attention at all layers, especially in the lower layers.We introduce DeFormer, a decomposed transformer, which substitutes the full self-attention with question-wide and passage-wide self-attentions in the lower layers.This allows for question-independent processing of the input text representations, which in turn enables pre-computing passage representations reducing runtime compute drastically.Furthermore, because DeFormer is largely similar to the original model, we can initialize DeFormer with the pre-training weights of a standard transformer, and directly fine-tune on the target QA dataset.We show DeFormer versions of BERT and XLNet can be used to speed up QA by over 4.3x and with simple distillation-based losses they incur only a 1% drop in accuracy.We open source the code at https://github.com/ StonyBrookNLP/deformer. Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian |
ACL | 2 |
| 2020 | Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected ReasoningabstractHas there been real progress in multi-hop question-answering?Models often exploit dataset artifacts to produce correct answers, without connecting information across multiple supporting facts.This limits our ability to measure true progress and defeats the purpose of building multi-hop QA datasets.We make three contributions towards addressing this.First, we formalize such undesirable behavior as disconnected reasoning across subsets of supporting facts.This allows developing a model-agnostic probe for measuring how much any model can cheat via disconnected reasoning.Second, using a notion of contrastive support sufficiency, we introduce an automatic transformation of existing datasets that reduces the amount of disconnected reasoning.Third, our experiments 1 suggest that there hasn't been much progress in multifact QA in the reading comprehension setting.For a recent large-scale model (XLNet), we show that only 18 points out of its answer F1 score of 72 on HotpotQA are obtained through multifact reasoning, roughly the same as that of a simpler RNN baseline.Our transformation substantially reduces disconnected reasoning (19 points in answer F1).It is complementary to adversarial approaches, yielding further reductions in conjunction.Original Dataset D ⇒ Question q = (Q, C; A) in D is assumed to be annotated with supporting facts {f 1 , f 2 }.Probing Dataset P ans+supp (D) for Answer Prediction and Support Identification tests: ⇒ Probing question collection P ans+supp (q) has only one group, corresponding to the unique bi-partition {{f 1 }, {f 2 }}, containing:Transformed Dataset T(D) for evaluating Constrastive Support Sufficiency: ⇒ Transformed question group T(q) in T(D) is defined using a single replacement fact f r ∈ C \ {f 1 , f 2 }:Probing Dataset P ans+supp+suff (T(D)) for all three tests: ⇒ Probing question collection P ans+supp+suff (T(q)) for the transformed question T(q) has only one group, corresponding to the unique bi-partition {{f 1 }, {f 2 }}, and is defined as: Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal |
EMNLP (1) | 1 |
| 2018 | Controlling Information Aggregation for Complex Question Answering
Heeyoung Kwon, Harsh Trivedi, Peter A. Jansen, Mihai Surdeanu, Niranjan Balasubramanian |
ECIR | 2 |