VLDB 2026 Research / reviewers in the wild / expert
Chaitanya Malaviya
dblp:194/3061
· DBLP profile ↗
15ranked-venue papers
9as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | R ESEARCH QA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and RubricsabstractAbstract Evaluating long-form responses to research queries is increasingly important for LLM agents, particularly emerging deep research systems. Such evaluation heavily relies on expert annotators, restricting attention to areas like AI where researchers can conveniently enlist colleagues. Yet, research expertise is abundant: survey articles consolidate knowledge spread across the literature. We introduce RESEARCHQA, a resource for evaluating LLM systems by distilling survey articles from 75 research felds into 21K queries and 160K rubric items. Queries and rubrics are jointly derived from survey sections, where rubric items list query-specific answer evaluation criteria, i.e., citing papers, making explanations, and describing limitations. 31 Ph.D. annotators in 8 fields judge that 90% of queries reflect Ph.D. information needs and 87% of rubric items warrant emphasis of a sentence or longer. We leverage RESEARCHQA to evaluate 18 systems in 7.6K head-to-heads. No parametric or retrieval-augmented system we evaluate exceeds 70% on covering rubric items, and the highest-ranking system shows 75% coverage. Error analysis reveals that the highest-ranking system fully addresses less than 11% of citation rubric items, 48% of limitation items, and 49% of comparison items. We release our data to facilitate more comprehensive multi-field evaluations. Li S. Yifei, Allen Chang, Chaitanya Malaviya, Mark Yatskar |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 3 |
| 2025 | LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense RetrievalabstractWhile significant progress has been made with dual-and bi-encoder dense retrievers, they often struggle on queries with logical connectives, a use case often overlooked yet important in downstream applications.In this paper, we introduce LOGICOL, a logically informed contrastive learning objective for dense retrievers.LOGICOL builds upon in-batch supervised contrastive learning and learns dense retrievers to respect the subset and mutually exclusive set relation between query results.We evaluated the effectiveness of LOGICOL in the entity retrieval task, where the model is expected to retrieve a set of Wikipedia entities that satisfy the implicit logical constraints of the query.We show that models trained with LOGICOL show improvements both in terms of retrieval performance and logical consistency in the results.We provide detailed analysis and insights to uncover why queries with logical connectives are challenging for dense retrievers and why LOGI-COL is effective.Our codes and data are available at https://github.com/yanzhen4/LogiCoL. A not B: Species of orchids inMalaysia but not Thailand. Yanzhen Shen, Xueqiang Xu, Yunyi Zhang 0001, Chaitanya Malaviya, Dan Roth 0001 |
EMNLP | 5 |
| 2025 | Dolomites: Domain-Specific Long-Form Methodical TasksabstractAbstract Experts in various fields routinely perform methodical writing tasks to plan, organize, and report their work. From a clinician writing a differential diagnosis for a patient, to a teacher writing a lesson plan for students, these tasks are pervasive, requiring to methodically generate structured long-form output for a given input. We develop a typology of methodical tasks structured in the form of a task objective, procedure, input, and output, and introduce DoLoMiTes, a novel benchmark with specifications for 519 such tasks elicited from hundreds of experts from across 25 fields. Our benchmark further contains specific instantiations of methodical tasks with concrete input and output examples (1,857 in total) which we obtain by collecting expert revisions of up to 10 model-generated examples of each task. We use these examples to evaluate contemporary language models, highlighting that automating methodical tasks is a challenging long-form generation problem, as it requires performing complex inferences, while drawing upon the given context as well as domain knowledge. Our dataset is available at https://dolomites-benchmark.github.io/. Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das 0001, Mirella Lapata, Christopher Alberti |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Contextualized Evaluations: Judging Language Model Responses to Underspecified QueriesabstractAbstract Language model users often issue queries that lack specification, where the context under which a query was issued—such as the user’s identity, the query’s intent, and the criteria for a response to be useful—is not explicit. For instance, a good response to a subjective query like “What book should I read next?” would depend on the user’s preferences, and a good response to an open-ended query like “How do antibiotics work against bacteria?” would depend on the user’s expertise. This makes evaluation of responses to such queries an ill-posed task, as evaluators may make arbitrary judgments about the response quality. To remedy this, we present contextualized evaluations, a protocol that synthetically constructs context surrounding an underspecified query and provides it during evaluation. We find that the presence of context can 1) alter conclusions drawn from evaluation, even flipping benchmark rankings between model pairs, 2) nudge evaluators to make fewer judgments based on surface-level criteria, like style, and 3) provide new insights about model behavior across diverse contexts. Specifically, our procedure suggests a potential bias towards WEIRD (Western, Educated, Industrialized, Rich and Democratic) contexts in models’ “default” responses and we find that models are not equally sensitive to following different contexts, even when they are provided in prompts.1 Chaitanya Malaviya, Joseph Chee Chang, Dan Roth 0001, Mohit Iyyer, Mark Yatskar, Kyle Lo |
Trans. Assoc. Comput. Linguistics | 1 |
| 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?abstractLanguage agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web.In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets or locating relevant nearby businesses.We introduce ASSISTANTBENCH, a challenging new benchmark consisting of 214 realistic tasks that can be automatically evaluated, covering different scenarios and domains.We find that AS-SISTANTBENCH exposes the limitations of current systems, including language models and retrieval-augmented language models, as no model reaches an accuracy of more than 25 points.While closed-book LMs perform well in terms of accuracy, they exhibit low precision and tend to hallucinate facts.State-of-the-art web agents reach a score of near zero.Additionally, we introduce SEEPLANACT (SPA), a new web agent that significantly outperforms previous agents, and an ensemble of SPA and closed-book models reaches the best overall performance.Moreover, we analyze failures of current systems and highlight that open web navigation remains a major challenge. Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, Jonathan Berant |
EMNLP | 3 |
| 2024 | ExpertQA: Expert-Curated Questions and Attributed AnswersabstractChaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chaitanya Malaviya, Elizabeth Sieber, Mark Yatskar, Dan Roth 0001 |
NAACL-HLT | 1 |
| 2024 | What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User PerceptionabstractChaitanya Malaviya, Subin Lee, Dan Roth, Mark Yatskar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chaitanya Malaviya, Dan Roth 0001, Mark Yatskar |
NAACL-HLT | 1 |
| 2023 | QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set OperationsabstractFormulating selective information needs results in queries that implicitly specify set operations, such as intersection, union, and difference.For instance, one might search for "shorebirds that are not sandpipers" or "science-fiction films shot in England".To study the ability of retrieval systems to meet such information needs, we construct QUEST, a dataset of 3357 natural language queries with implicit set operations, that map to a set of entities corresponding to Wikipedia documents.The dataset challenges models to match multiple constraints mentioned in queries with corresponding evidence in documents and correctly perform various set operations.The dataset is constructed semi-automatically using Wikipedia category names.Queries are automatically composed from individual categories, then paraphrased and further validated for naturalness and fluency by crowdworkers.Crowdworkers also assess the relevance of entities based on their documents and highlight attribution of query constraints to spans of document text.We analyze several modern retrieval systems, finding that they often struggle on such queries.Queries involving negation and conjunction are particularly challenging and systems are further challenged with combinations of these operations. 1 * Work done during an internship at Google.retrieving an exhaustive document set, instead lim-65 iting annotation to the top few results of a baseline 66 information retrieval system.67 To analyze how well retrieval systems handle 68 such queries, we present QUEST, a dataset with 69 natural language queries from four domains, that 70 are mapped to relatively comprehensive sets of en-71 tities corresponding to Wikipedia pages.We use 72 Wikipedia categories and their mapping to entities 73 in Wikipedia as a building block for our dataset 74 construction approach, but do not allow access to 75 this semi-structured data source at inference time, 76 to simulate text-based retrieval.Wikipedia cate-77 gories represent a broad set of natural language 78 descriptions of entity properties and often corre-79 spond to selective information need queries that 80 could be plausibly issued by a search engine user 81 ([At least 90% of the time based on our filtering?]).82 The correspondence between property names and 83 document text is also often subtle and requires so-84 phisticated reasoning to determine relevance, rep-85 resenting the natural language inference challenge 86 inherent in the task, while the knowledge of cate-87 gory membership allows us to construct relatively 88 comprehensive sets of candidate entities for atomic 89 categories and their combinations.90 Our dataset construction process is outlined in 91 Figure 1.The base queries in our dataset are 92 semi-automatically generated using Wikipedia cat-93 egory names.To construct queries, we sample 94 category names and compose them into complex 95 queries by using pre-defined templates (for exam-96 ple, A \ B \ C).Next, we ask crowdworkers to 97 paraphrase these automatically generated queries, 98 while ensuring that the paraphrased queries are 99 fluent and clearly describe what a user could be 00 looking for.These are then validated for natural-01 ness and fluency by a different set of crowdworkers, 02 and filtered according to those criteria.Finally, for 03 a large subset of our dataset, we collect scalar rel-04 evance labels based on the entity documents, and 05 textual attributions mapping query constraints to 06 spans of document text, to aid the development of 07 systems that can make precise inferences based on 08 trusted sources.09 Performing well on this dataset requires sys-10 tems that can match query constraints with cor-11 Chaitanya Malaviya, Peter Shaw 0004, Ming-Wei Chang, Kenton Lee, Kristina Toutanova |
ACL (1) | 1 |
| 2022 | Cascading Biases: Investigating the Effect of Heuristic Annotation Strategies on Data and ModelsabstractCognitive psychologists have documented that humans use cognitive heuristics, or mental shortcuts, to make quick decisions while expending less effort.While performing annotation work on crowdsourcing platforms, we hypothesize that such heuristic use among annotators cascades on to data quality and model robustness.In this work, we study cognitive heuristic use in the context of annotating multiple-choice reading comprehension datasets.We propose tracking annotator heuristic traces, where we tangibly measure low-effort annotation strategies that could indicate usage of various cognitive heuristics.We find evidence that annotators might be using multiple such heuristics, based on correlations with a battery of psychological tests.Importantly, heuristic use among annotators determines data quality along several dimensions: (1) known biased models, such as partial input models, more easily solve examples authored by annotators that rate highly on heuristic use, (2) models trained on annotators scoring highly on heuristic use don't generalize as well, and (3) heuristic-seeking annotators tend to create qualitatively less challenging examples.Our findings suggest that tracking heuristic usage among annotators can potentially help with collecting challenging datasets and diagnosing model biases. Chaitanya Malaviya, Sudeep Bhatia, Mark Yatskar |
EMNLP | 1 |
| 2020 | Commonsense Knowledge Base Completion with Structural and Semantic ContextabstractAutomatic KB completion for commonsense knowledge graphs (e.g., ATOMIC and ConceptNet) poses unique challenges compared to the much studied conventional knowledge bases (e.g., Freebase). Commonsense knowledge graphs use free-form text to represent nodes, resulting in orders of magnitude more nodes compared to conventional KBs ( ∼18x more nodes in ATOMIC compared to Freebase (FB15K-237)). Importantly, this implies significantly sparser graph structures — a major challenge for existing KB completion methods that assume densely connected graphs over a relatively smaller set of nodes.In this paper, we present novel KB completion models that can address these challenges by exploiting the structural and semantic context of nodes. Specifically, we investigate two key ideas: (1) learning from local graph structure, using graph convolutional networks and automatic graph densification and (2) transfer learning from pre-trained language models to knowledge graphs for enhanced contextual representation of knowledge. We describe our method to incorporate information from both these sources in a joint model and provide the first empirical results for KB completion on ATOMIC and evaluation with ranking metrics on ConceptNet. Our results demonstrate the effectiveness of language model representations in boosting link prediction performance and the advantages of learning from local graph structure (+1.5 points in MRR for ConceptNet) when training on subgraphs for computational efficiency. Further analysis on model predictions shines light on the types of commonsense knowledge that language models capture well. Chaitanya Malaviya, Chandra Bhagavatula, Antoine Bosselut, Yejin Choi 0001 |
AAAI | 1 |
| 2020 | Abductive Commonsense Reasoning
Chandra Bhagavatula, Ronan Le Bras 0001, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, Yejin Choi 0001 |
ICLR | 3 |
| 2019 | COMET: Commonsense Transformers for Automatic Knowledge Graph ConstructionabstractWe present the first comprehensive study on automatic knowledge base construction for two prevalent commonsense knowledge graphs: ATOMIC (Sap et al., 2019) and Con-ceptNet (Speer et al., 2017).Contrary to many conventional KBs that store knowledge with canonical templates, commonsense KBs only store loosely structured open-text descriptions of knowledge.We posit that an important step toward automatic commonsense completion is the development of generative models of commonsense knowledge, and propose COMmonsEnse Transformers (COMET ) that learn to generate rich and diverse commonsense descriptions in natural language.Despite the challenges of commonsense modeling, our investigation reveals promising results when implicit knowledge from deep pre-trained language models is transferred to generate explicit knowledge in commonsense knowledge graphs.Empirical results demonstrate that COMET is able to generate novel knowledge that humans rate as high quality, with up to 77.5% (ATOMIC) and 91.7% (ConceptNet) precision at top 1, which approaches human performance for these resources.Our findings suggest that using generative commonsense models for automatic commonsense KB completion could soon be a plausible alternative to extractive methods. Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, Yejin Choi 0001 |
ACL (1) | 4 |
| 2018 | Neural Factor Graph Models for Cross-lingual Morphological TaggingabstractMorphological analysis involves predicting the syntactic traits of a word (e.g.{POS: Noun, Case: Acc, Gender: Fem}).Previous work in morphological tagging improves performance for low-resource languages (LRLs) through cross-lingual training with a high-resource language (HRL) from the same family, but is limited by the strict-often false-assumption that tag sets exactly overlap between the HRL and LRL.In this paper we propose a method for cross-lingual morphological tagging that aims to improve information sharing between languages by relaxing this assumption.The proposed model uses factorial conditional random fields with neural network potentials, making it possible to (1) utilize the expressive power of neural network representations to smooth over superficial differences in the surface forms, (2) model pairwise and transitive relationships between tags, and (3) accurately generate tag sets that are unseen or rare in the training data.Experiments on four languages from the Universal Dependencies Treebank (Nivre et al., 2017) demonstrate superior tagging accuracies over existing cross-lingual approaches. 1 Chaitanya Malaviya, Matthew R. Gormley, Graham Neubig |
ACL (1) | 1 |
| 2017 | Learning Language Representations for Typology PredictionabstractOne central mystery of neural NLP is what neural models "know" about their subject matter.When a neural machine translation system learns to translate from one language to another, does it learn the syntax or semantics of the languages?Can this knowledge be extracted from the system to fill holes in human scientific knowledge?Existing typological databases contain relatively full feature specifications for only a few hundred languages.Exploiting the existence of parallel texts in more than a thousand languages, we build a massive many-to-one neural machine translation (NMT) system from 1017 languages into English, and use this to predict information missing from typological databases.Experiments show that the proposed method is able to infer not only syntactic, but also phonological and phonetic inventory features, and improves over a baseline that has access to information about the languages' geographic and phylogenetic neighbors.1 Chaitanya Malaviya, Graham Neubig, Patrick Littell |
EMNLP | 1 |