VLDB 2026 Research / reviewers in the wild / expert
Bhavya Chopra
dblp:292/9532
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-8638-3863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Why Do Multi-Agent LLM Systems Fail?abstractDespite enthusiasm for Multi-Agent LLM Systems (MAS), their performance gains on popular benchmarks are often minimal. This gap highlights a critical need for a principled understanding of why MAS fail. Addressing this question requires systematic identification and analysis of failure patterns. We introduce MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks. MAST-Data is the first multi-agent system dataset to outline the failure dynamics in MAS for guiding the development of better future systems. To enable systematic classification of failures for MAST-Data, we build the first Multi-Agent System Failure Taxonomy (MAST). We develop MAST through rigorous analysis of 150 traces, guided closely by expert human annotators andvalidated by high inter-annotator agreement (κ = 0.88). This process identifies 14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification. To enable scalable annotation, we develop an LLM-as-a-Judge pipeline with high agreement with human annotations. We leverage MAST and MAST-Data to analyze failure patterns across models (GPT4, Claude 3, Qwen2.5, CodeLlama) and tasks (coding, math, general agent), demonstrating improvement headrooms from better MAS design. Our analysis provides insights revealing that identified failures require more sophisticated solutions, highlighting a clear roadmap for future research. We publicly release our comprehensive dataset (MAST-Data), the MAST, and our LLM annotator to facilitate widespread research and development in MAS. Mert Cemri, Melissa Z. Pan, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya G. Parameswaran, Daniel Klein 0001, Kannan Ramchandran, Matei Zaharia, Joseph Gonzalez 0001, Ion Stoica |
NeurIPS | 5 |
| 2025 | Rethinking Dataset Discovery with DataScout
Rachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar, Madelon Hulsebos, Aditya G. Parameswaran |
UIST | 2 |
| 2025 | Steering Semantic Data Processing With DocWranglerabstractintent is hard to communicat Prompts need to be detailed and dataspecific LLM behavior varies across document Requires fine-grained decomposition A B C Docs & Outputs condition discomfort_level symptoms Shreya Shankar, Bhavya Chopra, Mawil Hasan, Stephen Lee, Björn Hartmann, Joseph M. Hellerstein, Aditya G. Parameswaran, Eugene Wu 0002 |
UIST | 2 |
| 2024 | Let's Fix this Together: Conversational Debugging with GitHub CopilotabstractDespite advancements in IDE tooling, code understanding, generation, and automated repair, debugging continues to present significant challenges. Existing debugging strategies available to developers in literature are often too mechanical and rigid for day-to-day issues. Recent advances in Large Language Models (LLMs) promise practical solutions that allow for more free-form debugging strategies. While LLMs offer satisfactory assistance in some cases, they often leap to action without sufficient context, making implicit assumptions and providing inaccurate responses. Moreover, the dialogue between developers and LLMs predominantly takes the form of question-answer pairs, placing the burden of formulating the correct questions and sustaining multi-turn conversations on the developer. We introduce Robin, a novel multi-agent conversational AI-assistant within GitHub Copilot Chat, specifically designed for debugging. Robin moves beyond the question-answer pairs by introducing the investigate & respond pattern, that focuses on using information gathered automatically from the IDE or gathered interactively from the developer before responding. Robin incorporates a general debugging strategy to systematically analyze bugs to sustain collaborative interactions while ensuring that the conversation does not deviate from the debugging task at hand. Through a within-subjects user study with 16 industry professionals, we find that equipping Robin to-(1) leverage the insert expansion interaction pattern, (2) facilitate turn-taking, and (3) utilize debugging strategies-leads to lowered conversation barriers, a 2.5 x improvement in bug localization and a substantial 3.5x improvement in bug resolution compared to AI-assisted debugging in Visual Studio prior to Robin. Yasharth Bajpai, Bhavya Chopra, Param Biyani, Cagri Aslan, Dustin Coleman, Sumit Gulwani, Chris Parnin, Arjun Radhakrishna, Gustavo Soares |
VL/HCC | 2 |
| 2023 | Detangler: Helping Data Scientists Explore, Understand, and Debug Data Wrangling PipelinesabstractData scientists spend significant time on data wrangling-a process involving data cleaning, shaping, and pre-processing. Data wrangling requires meticulous exploration and backtracking to assess data quality by applying and validating numerous data transformation chains, making it a tedious and error-prone process. In this paper, we present Detangler, an interactive tool within the RStudio IDE that helps data scientists identify and debug data quality issues and wrangling code. The design of Detangler is informed via formative interviews, and it presents data scientists with (i) insights into potential data quality issues, and (ii) always-on visual summaries of the effects of individual data transformations, enabling interactive exploration of data and wrangling code. Through a laboratory study with 18 data scientists, triangulated with telemetry data, we find that Detangler improves exploration and debugging of data smells and data wrangling code. We discuss design implications for future tools for data science programming. Nischal Shrestha, Bhavya Chopra, Austin Z. Henley, Chris Parnin |
VL/HCC | 2 |