VLDB 2026 Research / reviewers in the wild / expert
Madeleine Grunde-McLaughlin
dblp:271/8198
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-7290-068XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RDoFlow: Automatically assessing under-specified statistical analyses in HCIabstractWhen designing and analyzing a study, researchers must navigate a large space of methodological decisions, or “researcher degrees of freedom.” If these choices are not preregistered or transparently reported, they can increase the risk of inflated false-positive rates and exaggerated effect sizes, undermining scientific credibility. Drawing on psychology research that characterizes these degrees of freedom, we create a protocol for scoring how hypotheses are reported in the HCI literature (ReportDoF). We manually apply ReportDoF to 100 hypotheses from HCI texts authored between 2015-2025, including both preregistrations and papers. Based on this experience, we contribute an LLM workflow and proof-of-concept interactive interface (RDoFlow) that applies ReportDoF to new texts, enabling large-scale analysis of the composition and quality of reported analysis specifications. For example, RDoFlow reveals that HCI research more frequently tests multiple dependent variables for a single hypothesis than psychology research does—a practice that increases the risk of false positives. Madeleine Grunde-McLaughlin, Weixuan Liu, Ria Patil, Nino Migineishvili, Emily Reif, Ranjay Krishna, Daniel S. Weld, Jeffrey Heer |
IUI | 1 |
| 2026 | Wildfire and Forest Management: Opportunities for HCI ResearchabstractWildfire and forest management increasingly rely on geospatial technologies, i.e., data and tools contributing to the geographic mapping and analysis of the Earth, to inform measures for the control of wildfires. Nevertheless, challenges arising from domain experts adopting these complex, non-intuitive technologies are not well understood. We interviewed 12 participants in wildfire and forest management to explore the technical and socio-technical nature of these challenges, revealing that (1) knowledge and data are fragmented across stakeholders, ranging from governmental agencies to small landowners. This fragmentation causes participants to (2) struggle in sharing knowledge and expertise. Participants (3) voice concerns about model bias since decisions informed by geospatial technologies can have far-reaching impacts. Yet, they (4) face barriers engaging people most impacted by these decisions. We detail an HCI research agenda that includes: exploring opportunities to connect stakeholders and sharing knowledge, standardizing decision-making, and engaging local communities. Nino Migineishvili, Madeleine Grunde-McLaughlin, Emmanuel Azuh, Spencer Wood, René Just, Katharina Reinecke |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2025 | Designing LLM Chains by Adapting Techniques from Crowdsourcing WorkflowsabstractLLM chains enable complex tasks by decomposing work into a sequence of subtasks. Similarly, the more established techniques of crowdsourcing workflows decompose complex tasks into smaller tasks for human crowdworkers. Chains address LLM errors analogously to the way crowdsourcing workflows address human error. To characterize opportunities for LLM chaining, we survey 107 papers across the crowdsourcing and chaining literature to construct a design space for chain development. The design space covers a designer’s objectives and the tactics used to build workflows. We then surface strategies that mediate how workflows use tactics to achieve objectives. To explore how techniques from crowdsourcing may apply to chaining, we adapt crowdsourcing workflows to implement LLM chains across three case studies: creating a taxonomy, shortening text, and writing a short story. From the design space and our case studies, we identify takeaways for effective chain design and raise implications for future research and development. Madeleine Grunde-McLaughlin, Michelle S. Lam, Ranjay Krishna, Daniel S. Weld, Jeffrey Heer |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2024 | How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz StudyabstractData analysis is challenging as analysts must navigate nuanced decisions that may yield divergent conclusions. AI assistants have the potential to support analysts in planning their analyses, enabling more robust decision making. Though AI-based assistants that target code execution (e.g., Github Copilot) have received significant attention, limited research addresses assistance for both analysis execution and planning. In this work, we characterize helpful planning suggestions and their impacts on analysts’ workflows. We first review the analysis planning literature and crowd-sourced analysis studies to categorize suggestion content. We then conduct a Wizard-of-Oz study (n=13) to observe analysts’ preferences and reactions to planning assistance in a realistic scenario. Our findings highlight subtleties in contextual factors that impact suggestion helpfulness, emphasizing design implications for supporting different abstractions of assistance, forms of initiative, increased engagement, and alignment of goals between analysts and assistants. Ken Gu, Madeleine Grunde-McLaughlin, Andrew M. McNutt, Jeffrey Heer, Tim Althoff |
CHI | 2 |
| 2023 | Explanations Can Reduce Overreliance on AI Systems During Decision-MakingabstractPrior work has identified a resilient phenomenon that threatens the performance of human-AI decision-making teams: overreliance, when people agree with an AI, even when it is incorrect. Surprisingly, overreliance does not reduce when the AI produces explanations for its predictions, compared to only providing predictions. Some have argued that overreliance results from cognitive biases or uncalibrated trust, attributing overreliance to an inevitability of human cognition. By contrast, our paper argues that people strategically choose whether or not to engage with an AI explanation, demonstrating empirically that there are scenarios where AI explanations reduce overreliance. To achieve this, we formalize this strategic choice in a cost-benefit framework, where the costs and benefits of engaging with the task are weighed against the costs and benefits of relying on the AI. We manipulate the costs and benefits in a maze task, where participants collaborate with a simulated AI to find the exit of a maze. Through 5 studies (N = 731), we find that costs such as task difficulty (Study 1), explanation difficulty (Study 2, 3), and benefits such as monetary compensation (Study 4) affect overreliance. Finally, Study 5 adapts the Cognitive Effort Discounting paradigm to quantify the utility of different explanations, providing further support for our framework. Our results suggest that some of the null effects found in literature could be due in part to the explanation not sufficiently reducing the costs of verifying the AI's prediction. Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, Ranjay Krishna |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Measuring Compositional Consistency for Video Question AnsweringabstractRecent video question answering benchmarks indicate that state-of-the-art models struggle to answer compositional questions. However, it remains unclear which types of compositional reasoning cause models to mispredict. Furthermore, it is difficult to discern whether models arrive at answers using compositional reasoning or by leveraging data biases. In this paper, we develop a question decomposition engine that programmatically deconstructs a compositional question into a directed acyclic graph of sub-questions. The graph is designed such that each parent question is a composition of its children. We present AGQA-Decomp, a benchmark containing 2.3M question graphs, with an average of 11.49 sub-questions per graph, and 4.55M total new sub-questions. Using question graphs, we evaluate three state-of-the-art models with a suite of novel compositional consistency metrics. We find that models either cannot reason correctly through most compositions or are reliant on incorrect reasoning to reach answers, frequently contradicting themselves or achieving high accuracies when failing at intermediate reasoning steps. Mona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala |
CVPR | 4 |
| 2021 | AGQA: A Benchmark for Compositional Spatio-Temporal ReasoningabstractVisual events are a composition of temporal actions involving actors spatially interacting with objects. When developing computer vision models that can reason about compositional spatio-temporal events, we need benchmarks that can analyze progress and uncover shortcomings. Existing video question answering benchmarks are useful, but they often conflate multiple sources of error into one accuracy metric and have strong biases that models can exploit, making it difficult to pinpoint model weaknesses. We present Action Genome Question Answering (AGQA), a new benchmark for compositional spatio-temporal reasoning. AGQA contains 192M unbalanced question answer pairs for 9.6K videos. We also provide a balanced subset of 3.9M question answer pairs, 3 orders of magnitude larger than existing benchmarks, that minimizes bias by balancing the answer distributions and types of question structures. Although human evaluators marked 86.02% of our question-answer pairs as correct, the best model achieves only 47.74% accuracy. In addition, AGQA introduces multiple training/test splits to test for various reasoning abilities, including generalization to novel compositions, to indirect references, and to more compositional steps. Using AGQA, we evaluate modern visual reasoning systems, demonstrating that the best models barely perform better than non-visual baselines exploiting linguistic biases and that none of the existing models generalize to novel compositions unseen during training. Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala |
CVPR | 1 |
| 2021 | Bayesian-Assisted Inference from Visualized DataabstractA Bayesian view of data interpretation suggests that a visualization user should update their existing beliefs about a parameter's value in accordance with the amount of information about the parameter value captured by the new observations. Extending recent work applying Bayesian models to understand and evaluate belief updating from visualizations, we show how the predictions of Bayesian inference can be used to guide more rational belief updating. We design a Bayesian inference-assisted uncertainty analogy that numerically relates uncertainty in observed data to the user's subjective uncertainty, and a posterior visualization that prescribes how a user should update their beliefs given their prior beliefs and the observed data. In a pre-registered experiment on 4,800 people, we find that when a newly observed data sample is relatively small (N=158), both techniques reliably improve people's Bayesian updating on average compared to the current best practice of visualizing uncertainty in the observed data. For large data samples (N=5208), where people's updated beliefs tend to deviate more strongly from the prescriptions of a Bayesian model, we find evidence that the effectiveness of the two forms of Bayesian assistance may depend on people's proclivity toward trusting the source of the data. We discuss how our results provide insight into individual processes of belief updating and subjective uncertainty, and how understanding these aspects of interpretation paves the way for more sophisticated interactive visualizations for analysis and communication. Yea-Seul Kim, Paula Kayongo, Madeleine Grunde-McLaughlin, Jessica Hullman |
IEEE Trans. Vis. Comput. Graph. | 3 |