VLDB 2026 Research / reviewers in the wild / expert
Victor S. Bursztyn
dblp:154/7800 · also Victor Soares Bursztyn
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-6187-6415ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TeamFusion: Supporting Open-ended Teamwork with Multi-Agent SystemsabstractJiale Liu, Victor Bursztyn, Lin Ai, Haoliang Wang, Sunav Choudhary, Saayan Mitra, Qingyun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Victor S. Bursztyn, Lin Ai, Sunav Choudhary, Saayan Mitra, Qingyun Wu |
ACL (1) | 2 |
| 2025 | A Flash in the Pan: Better Prompting Strategies to Deploy Out-of-the-Box LLMs as Conversational Recommendation SystemsabstractConversational Recommendation Systems (CRSs) are a particularly interesting application for out-of-the-box LLMs due to their potential for eliciting user preferences and making recommendations in natural language across a wide set of domains. Somewhat surprisingly, we find however that in such a conversational application, the more questions a user answers about their preferences, the worse the model’s recommendations become. We demonstrate this phenomenon on a previously published dataset as well as two novel datasets which we contribute. We also explain why earlier benchmarks failed to detect this round-over-round performance loss, highlighting the importance of the evaluation strategy we use and expanding upon Li et al. (2023a). We also present preference elicitation and recommendation strategies that mitigate this degradation in performance, beating state-of-the-art results, and show how three underlying models, GPT-3.5, GPT-4, and Claude 3.5 Sonnet, differently impact these strategies. Our datasets and code are available at https://github.com/CtrlVGustavo/A-Flash- in-the-Pan-CRS. Gustavo Adolpho Lucas de Carvalho, Simon Ben Igeri, Jennifer A. Healey, Victor S. Bursztyn, David Demeter, Lawrence Birnbaum |
COLING | 4 |
| 2025 | Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A SurveyabstractMehrab Tanjim, Yeonjun In, Xiang Chen, Victor Bursztyn, Ryan A. Rossi, Sungchul Kim, Guang-Jie Ren, Vaishnavi Muppala, Shun Jiang, Yongsung Kim, Chanyoung Park. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Md. Mehrab Tanjim, Yeonjun In, Xiang Chen 0010, Victor S. Bursztyn, Ryan Rossi, Sungchul Kim, Vaishnavi Muppala, Shun Jiang, Yongsung Kim, Chanyoung Park 0001 |
EMNLP | 4 |
| 2025 | ConvoMap: Interactive Visualizations for Exploring Complex Conversations in Multi-Agent Systemsabstract—Following the rapid emergence of large language models, Multi-Agent Systems (MASs) became a promising approach for accomplishing complex tasks. In MASs, multiple autonomous agents with predetermined roles collaborate by dividing responsibilities. However, MAS developers often struggle to understand and diagnose agents’ behavior from thousands of inter-agent messages across multiple complex conversations. To identify key requirements and challenges related to evaluating, debugging, and managing MASs, we conducted a formative study with six MAS developers. We then introduce ConvoMap, a prototype that addresses a key challenge of MAS development-understanding agents’ behaviors across multiple conversations. ConvoMap integrates automated qualitative coding to enable multi-level inspection of agents’ behavior. ConvoMap can then visualize hundreds of MAS conversations by representing messages as points on a 2D map that encode their semantic meanings and interactions between agents. To better support navigation and deeper analysis, ConvoMap provides topic overviews and highlights relevant text segments. A comparison study showed that ConvoMap helped to understand agents’ behavior more accurately than the baseline. Ashley Ge Zhang, Victor S. Bursztyn, Gromit Yeuk-Yin Chan, Shunan Guo, Eunyee Koh, Steve Oney, Jane Hoffswell |
VL/HCC | 2 |
| 2025 | How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying LayoutsabstractLarge Language Models (LLMs) have been adopted for a variety of visualizations tasks, but how far are we from perceptually aware LLMs that can predict human takeaways? Graphical perception literature has shown that human chart takeaways are sensitive to visualization design choices, such as spatial layouts. In this work, we examine the extent to which LLMs exhibit such sensitivity when generating takeaways, using bar charts with varying spatial layouts as a case study. We conducted three experiments and tested four common bar chart layouts: vertically juxtaposed, horizontally juxtaposed, overlaid, and stacked. In Experiment 1, we identified the optimal configurations to generate meaningful chart takeaways by testing four LLMs, two temperature settings, nine chart specifications, and two prompting strategies. We found that even state-of-the-art LLMs struggled to generate semantically diverse and factually accurate takeaways. In Experiment 2, we used the optimal configurations to generate 30 chart takeaways each for eight visualizations across four layouts and two datasets in both zero-shot and one-shot settings. Compared to human takeaways, we found that the takeaways LLMs generated often did not match the types of comparisons made by humans. In Experiment 3, we examined the effect of chart context and data on LLM takeaways. We found that LLMs, unlike humans, exhibited variation in takeaway comparison types for different bar charts using the same bar layout. Overall, our case study evaluates the ability of LLMs to emulate human interpretations of data and points to challenges and opportunities in using LLMs to predict human chart takeaways. Huichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn, Cindy Xiong Bearfield |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | ToolChain*: Efficient Action Space Navigation in Large Language Models with A* SearchabstractLarge language models (LLMs) have demonstrated powerful decision-making and planning capabilities in solving complicated real-world problems. LLM-based autonomous agents can interact with diverse tools (e.g., functional APIs) and generate solution plans that execute a series of API function calls in a step-by-step manner. The multitude of candidate API function calls significantly expands the action space, amplifying the critical need for efficient action space navigation. However, existing methods either struggle with unidirectional exploration in expansive action spaces, trapped into a locally optimal solution, or suffer from exhaustively traversing all potential actions, causing inefficient navigation. To address these issues, we propose ToolChain*, an efficient tree search-based planning algorithm for LLM-based agents. It formulates the entire action space as a decision tree, where each node represents a possible API function call involved in a solution plan. By incorporating the A$^*$ search algorithm with task-specific cost function design, it efficiently prunes high-cost branches that may involve incorrect actions, identifying the most low-cost valid path as the solution. Extensive experiments on multiple tool-use and reasoning tasks demonstrate that ToolChain* efficiently balances exploration and exploitation within an expansive action space. It outperforms state-of-the-art baselines on planning and reasoning tasks by 3.1% and 3.5% on average while requiring 7.35x and 2.31x less time, respectively. Yuchen Zhuang, Xiang Chen 0010, Tong Yu 0001, Saayan Mitra, Victor S. Bursztyn, Ryan Rossi, Somdeb Sarkhel, Chao Zhang 0014 |
ICLR | 5 |
| 2024 | Flexible And Faithful Data Insights GenerationabstractWith the ever-growing volume of data, corporate users want to understand their data more efficiently. In addition to data visualization, they like text-based insights (e.g., a summary of trends, seasonality, and anomalies in the data) so they can digest them easily and/or share them with other stakeholders. In this paper, we introduce an efficient framework that generates faithful and flexible insights that resemble human writing. It is a hybrid approach that takes advantage of both template-based summaries and the latest Generative AI technology. With a fine-tuned compact Large Language Model (LLM) and a Gatekeeper, we can achieve a lower error rate than state-of-the-art LLMs, while using only a fraction of resources. We have released a service based on the proposed approach and received very positive feedback. Victor S. Bursztyn |
ISM | 2 |
| 2024 | "The Data Says Otherwise" - Towards Automated Fact-checking and Communication of Data ClaimsabstractFact-checking data claims requires data evidence retrieval and analysis, which can become tedious and intractable when done manually. This work presents Aletheia, an automated fact-checking prototype designed to facilitate data claims verification and enhance data evidence communication. For verification, we utilize a pre-trained LLM to parse the semantics for evidence retrieval. To effectively communicate the data evidence, we design representations in two forms: data tables and visualizations, tailored to various data fact types. Additionally, we design interactions that showcase a real-world application of these techniques. We evaluate the performance of two core NLP tasks with a curated dataset comprising 400 data claims and compare the two representation forms regarding viewers’ assessment time, confidence, and preference via a user study with 20 participants. The evaluation offers insights into the feasibility and bottlenecks of using LLMs for data fact-checking tasks, potential advantages and disadvantages of using visualizations over data tables, and design recommendations for presenting data evidence. Yu Fu 0010, Shunan Guo, Jane Hoffswell, Victor S. Bursztyn, Ryan Rossi, John T. Stasko |
UIST | 4 |
| 2024 | Representing Charts as Text for Language Models: An In-Depth Study of Question Answering for Bar ChartsabstractMachine Learning models for chart-grounded Q&A (CQA) often treat charts as images, but performing CQA on pixel values has proven challenging. We thus investigate a resource overlooked by current ML-based approaches: the declarative documents describing how charts should visually encode data (i.e., chart specifications). In this work, we use chart specifications to enhance language models (LMs) for chart-reading tasks, such that the resulting system can robustly understand language for CQA. Through a case study with 359 bar charts, we test novel fine tuning schemes on both GPT-3 and T5 using a new dataset curated for two CQA tasks: question-answering and visual explanation generation. Our text-only approaches strongly outperform vision-based GPT-4 on explanation generation (99% vs. 63% accuracy), and show promising results for question-answering (57–67% accuracy). Through in-depth experiments, we also show that our text-only approaches are mostly robust to natural language variation. Victor S. Bursztyn, Jane Hoffswell, Eunyee Koh, Shunan Guo |
IEEE VIS | 1 |
| 2021 | "It doesn't look good for a date": Transforming Critiques into Preferences for Conversational Recommendation SystemsabstractConversations aimed at determining good recommendations are iterative in nature.People often express their preferences in terms of a critique of the current recommendation (e.g., "It doesn't look good for a date"), requiring some degree of common sense for a preference to be inferred.In this work, we present a method for transforming a user critique into a positive preference (e.g., "I prefer more romantic") in order to retrieve reviews pertaining to potentially better recommendations (e.g., "Perfect for a romantic dinner").We leverage a large neural language model (LM) in a fewshot setting to perform critique-to-preference transformation, and we test two methods for retrieving recommendations: one that matches embeddings, and another that fine-tunes an LM for the task.We instantiate this approach in the restaurant domain and evaluate it using a new dataset of restaurant critiques.In an ablation study, we show that utilizing critiqueto-preference transformation improves recommendations, and that there are at least three general cases that explain this improved performance. Victor S. Bursztyn, Jennifer A. Healey, Nedim Lipka, Eunyee Koh, Doug Downey, Lawrence Birnbaum |
EMNLP (1) | 1 |
| 2019 | Thousands of small, constant rallies: a large-scale analysis of partisan WhatsApp groupsabstractThere is growing concern about the use of social platforms to push political narratives during elections. One very recent case is Brazil's, where WhatsApp is now widely perceived as a key enabler of the far-right's rise to power. In this paper, we perform a large-scale analysis of partisan WhatsApp groups to shed light on how both right-wingers and left-wingers used the platform in the 2018 Brazilian presidential election. Across its two rounds, we collected +2.8M messages from +45k users in 232 public groups (175 right-wing vs. 57 left-wing). After describing how we obtained a sample that is many times larger than previous works, we contrast right-wingers and left-wingers on their social network metrics, regional distribution of users, content-sharing habits, and most characteristic news sources. Victor S. Bursztyn, Lawrence Birnbaum |
ASONAM | 1 |