VLDB 2026 Research / reviewers in the wild / expert
Mary Beth Kery
dblp:149/9443
· DBLP profile ↗
17ranked-venue papers
9as first author
4since 2021 · last 2025
0000-0002-1771-0565ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 16 · 8 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Policy Maps: Tools for Guiding the Unbounded Space of LLM BehaviorsabstractFigure 1: Policy maps chart LLM policy coverage over an unbounded space of model behaviors.Here, an AI practitioner is designing a policy for how an LLM should summarize violent text.Policy map abstractions (right) allow the policy designer to interactively author and test policies that govern a model's behavior using if-then rules over concepts.The designer can create any desired concept by providing a simple text definition to capture cases of model behavior.Our Policy Projector tool (center) renders cases, concepts, and policies as visual map layers to aid iterative policy design. Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery |
UIST | 6 |
| 2024 | Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning ExperiencesabstractOn-device machine learning (ML) promises to improve the privacy, responsiveness, and proliferation of new, intelligent user experiences by moving ML computation onto everyday personal devices. However, today’s large ML models must be drastically compressed to run efficiently on-device, a hurtle that requires deep, yet currently niche expertise. To engage the broader human-centered ML community in on-device ML experiences, we present the results from an interview study with 30 experts at Apple that specialize in producing efficient models. We compile tacit knowledge that experts have developed through practical experience with model compression across different hardware platforms. Our findings offer pragmatic considerations missing from prior work, covering the design process, trade-offs, and technical strategies that go into creating efficient models. Finally, we distill design recommendations for tooling to help ease the difficulty of this work and bring on-device ML into to more widespread practice. Fred Hohman, Mary Beth Kery, Donghao Ren, Dominik Moritz |
CHI | 2 |
| 2023 | Collaborative Machine Learning Model Building with Families Using Co-MLabstractExisting novice-friendly machine learning (ML) modeling tools center around a solo user experience, where a single user collects only their own data to build a model. However, solo modeling experiences limit valuable opportunities for encountering alternative ideas and approaches that can arise when learners work together; consequently, it often precludes encountering critical issues in ML around data representation and diversity that can surface when different perspectives are manifested in a group-constructed data set. To address this issue, we created Co-ML – a tablet-based app for learners to collaboratively build ML image classifiers through an end-to-end, iterative model-building process. In this paper, we illustrate the feasibility and potential richness of collaborative modeling by presenting an in-depth case study of a family (two children 11 and 14-years-old working with their parents) using Co-ML in a facilitated introductory ML activity at home. We share the Co-ML system design and contribute a discussion of how using Co-ML in a collaborative activity enabled beginners to collectively engage with dataset design considerations underrepresented in prior work such as data diversity, class imbalance, and data quality. We discuss how a distributed collaborative process, in which individuals can take on different model-building responsibilities, provides a rich context for children and adults to learn ML dataset design. Tiffany Tseng, Jennifer King Chen, Mona Abdelrahman, Mary Beth Kery, Fred Hohman, Adriana Hilliard, R. Benjamin Shapiro |
IDC | 4 |
| 2023 | Angler: Helping Machine Translation Practitioners Prioritize Model ImprovementsabstractMachine learning (ML) models can fail in unexpected ways in the real world, but not all model failures are equal. With finite time and resources, ML practitioners are forced to prioritize their model debugging and improvement efforts. Through interviews with 13 ML practitioners at Apple, we found that practitioners construct small targeted test sets to estimate an error’s nature, scope, and impact on users. We built on this insight in a case study with machine translation models, and developed Angler, an interactive visual analytics tool to help practitioners prioritize model improvements. In a user study with 7 machine translation experts, we used Angler to understand prioritization practices when the input space is infinite, and obtaining reliable signals of model quality is expensive. Our study revealed that participants could form more interesting and user-focused hypotheses for prioritization by analyzing quantitative summary statistics and qualitatively assessing data by reading sentences. Samantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery, Fred Hohman |
CHI | 4 |
| 2020 | Understanding and Visualizing Data Iteration in Machine LearningabstractSuccessful machine learning (ML) applications require iterations on both modeling and the underlying data. While prior visualization tools for ML primarily focus on modeling, our interviews with 23 ML practitioners reveal that they improve model performance frequently by iterating on their data (e.g., collecting new data, adding labels) rather than their models. We also identify common types of data iterations and associated analysis tasks and challenges. To help attribute data iterations to model performance, we design a collection of interactive visualizations and integrate them into a prototype, Chameleon, that lets users compare data features, training/testing splits, and performance across data versions. We present two case studies where developers apply \system to their own evolving datasets on production ML projects. Our interface helps them verify data collection efforts, find failure cases stretching across data versions, capture data processing changes that impacted performance, and identify opportunities for future data iterations. Fred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur Patel |
CHI | 3 |
| 2020 | mage: Fluid Moves Between Code and Graphical Work in Computational NotebooksabstractWe aim to increase the flexibility at which a data worker can choose the right tool for the job, regardless of whether the tool is a code library or an interactive graphical user interface (GUI). To achieve this flexibility, we extend computational notebooks with a new API mage, which supports tools that can represent themselves as both code and GUI as needed. We discuss the design of mage as well as design opportunities in the space of flexible code/GUI tools for data work. To understand tooling needs, we conduct a study with nine professional practitioners and elicit their feedback on mage and potential areas for flexible code/GUI tooling. We then implement six client tools for mage that illustrate the main themes of our study findings. Finally, we discuss open challenges in providing flexible code/GUI interactions for data workers. Mary Beth Kery, Donghao Ren, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat, Kayur Patel |
UIST | 1 |
| 2019 | Towards Effective Foraging by Data Scientists to Find Past Analysis ChoicesabstractData scientists are responsible for the analysis decisions they make, but it is hard for them to track the process by which they achieved a result. Even when data scientists keep logs, it is onerous to make sense of the resulting large number of history records full of overlapping variants of code, output, plots, etc. We developed algorithmic and visualization techniques for notebook code environments to help data scientists forage for information in their history. To test these interventions, we conducted a think-aloud evaluation with 15 data scientists, where participants were asked to find specific information from the history of another person's data science project. The participants succeed on a median of 80% of the tasks they performed. The quantitative results suggest promising aspects of our design, while qualitative results motivated a number of design improvements. The resulting system, called Verdant, is released as an open-source extension for JupyterLab. Mary Beth Kery, Bonnie E. John, Patrick O'Flaherty, Amber Horvath, Brad A. Myers |
CHI | 1 |
| 2019 | The Long Tail: Understanding the Discoverability of API FunctionalityabstractAlmost all software development revolves around the discovery and use of application programming interfaces (APIs). Once a suitable API is selected, programmers must begin the process of determining what functionality in the API is relevant to a programmer's task and how to use it. Our work aims to understand how API functionality is discovered by programmers and where tooling may be appropriate. We employed a mixed-methods approach to investigate Apache Beam, a distributed data processing API, by mining Beam client code and running a lab study to see how people discover Beam's available functionality. We found that programmers' prior experience with similar APIs significantly impacted their ability to find relevant features in an API and attempting to form a top-down mental model of an API resulted in less discovery of features. Amber Horvath, Sachin Grover, Sihan Dong, Emily Zhou, Finn Voichick, Mary Beth Kery, Shwetha Shinju, Daye Nam, Mariann Nagy, Brad A. Myers |
VL/HCC | 6 |
| 2018 | The Story in the Notebook: Exploratory Data Science using a Literate Programming ToolabstractLiterate programming tools are used by millions of programmers today, and are intended to facilitate presenting data analyses in the form of a narrative. We interviewed 21 data scientists to study coding behaviors in a literate programming environment and how data scientists kept track of variants they explored. For participants who tried to keep a detailed history of their experimentation, both informal and formal versioning attempts led to problems, such as reduced notebook readability. During iteration, participants actively curated their notebooks into narratives, although primarily through cell structure rather than markdown explanations. Next, we surveyed 45 data scientists and asked them to envision how they might use their past history in an future version control system. Based on these results, we give design guidance for future literate programming tools, such as providing history search based on how programmers recall their explorations, through contextual details including images and parameters. Mary Beth Kery, Marissa Radensky, Mahima Arya, Bonnie E. John, Brad A. Myers |
CHI | 1 |
| 2018 | Towards Scaffolding Complex Exploratory Data Science Programming PracticesabstractAlthough a wide range of professional and end-user programmers want to engage today with data science programming, this form of programming presents unique challenges. For instance, data science tasks typically require exploratory iterations: coding and running many different approaches to reach a desired result [1]-[3]. In a body of research building towards my thesis, I have interleaved behavioral studies of data scientists with systems building research towards scaffolding new forms of support for keeping track of iterations during this experiment-driven form of work. Mary Beth Kery |
VL/HCC | 1 |
| 2018 | Interactions for Untangling Messy History in a Computational NotebookabstractExperimentation through code is central to data scientists' work. Prior work has identified the need for interaction techniques for quickly exploring multiple versions of the code and the associated outputs. Yet previous approaches that provide history information have been challenging to scale: real use produces a high number of versions of different code and non-code artifacts with dependency relationships and a convoluted mix of different analysis intents. Prior work has found that navigating these records to pick out the relevant information for a given task is difficult and time consuming. We introduce Verdant, a new system with a novel versioning model to support fast retrieval and sensemaking of messy version data. Verdant provides light-weight interactions for comparing, replaying, and tracing relationships among many versions of different code and non-code artifacts in the editor. We implemented Verdant into Jupyter Notebooks, and validated the usability of Verdant's interactions through a usability study. Mary Beth Kery, Brad A. Myers |
VL/HCC | 1 |
| 2018 | API Designers in the Field: Design Practices and Challenges for Creating Usable APIsabstractApplication Programming Interfaces (APIs) are a rapidly growing industry and the usability of the APIs is crucial to programmer productivity. Although prior research has shown that APIs commonly suffer from significant usability problems, little attention has been given to studying how APIs are designed and created in the first place. We interviewed 24 professionals involved with API design from 7 major companies to identify their training and design processes. Interviewees had insights into many different aspects of designing for API usability and areas of significant struggle. For example, they learned to do API design on the job, and had little training for it in school. During the design phase they found it challenging to discern which potential use cases of the API users will value most. After an API is released, designers lack tools to gather aggregate feedback from this data even as developers openly discuss the API online. Lauren Murphy, Mary Beth Kery, Oluwatosin Alliyu, Andrew Macvean, Brad A. Myers |
VL/HCC | 2 |
| 2017 | Variolite: Supporting Exploratory Programming by Data ScientistsabstractHow do people ideate through code? Using semi-structured interviews and a survey, we studied data scientists who program, often with small scripts, to experiment with data. These studies show that data scientists frequently code new analysis ideas by building off of their code from a previous idea. They often rely on informal versioning interactions like copying code, keeping unused code, and commenting out code to repurpose older analysis code while attempting to keep those older analyses intact. Unlike conventional version control, these informal practices allow for fast versioning of any size code snippet, and quick comparisons by interchanging which versions are run. However, data scientists must maintain a strong mental map of their code in order to distinguish versions, leading to errors and confusion. We explore the needs for improving version control tools for exploratory tasks, and demonstrate a tool for lightweight local versioning, called Variolite, which programmers found usable and desirable in a preliminary usability study. Mary Beth Kery, Amber Horvath, Brad A. Myers |
CHI | 1 |
| 2017 | Tools to support exploratory programming with dataabstractA diverse range of people, from students to engineers to designers, are interested in using programming to analyze, visualize, and build new intelligent systems from data. However, when working with data, a programmer must typically experiment heavily: writing out and running many different approaches in code to reach a desired result [1][2]. This form of exploratory programming presents extra challenges and pitfalls for programmers. For example, as a person iterates on a problem over a long period of time, it can become difficult to answer questions like: “Through what steps did I achieve this result?” [3] or “What different analyses have I tried or not tried so far?”. Currently there is a sparsity of tools for programmers to keep track of all their code experimentation, and intermediary data analysis steps are easily lost. Programmers also struggle with making many small highly exploratory code changes, such as trying 10 different values of a parameter [2]. Current code version tools are not designed for rapid iteration on small sections of code, thus these programmers often resort to informal ad-hoc methods of versioning such as leaving 10 different values of a parameter in comments in their code, or copying the 10 different values to a separate text file. As programmers work to make sense of challenging data questions, managing experiments in code adds an extra degree of overhead and confusion that can make this work all the more difficult. Mary Beth Kery |
VL/HCC | 1 |
| 2017 | Exploring exploratory programmingabstractIn open-ended tasks where a program's behavior cannot be specified in advance, exploratory programming is a key practice in which programmers actively experiment with different possibilities using code. Exploratory programming is highly relevant today to a variety of professional and end-user programmer domains, including prototyping, learning through play, digital art, and data science. However, prior research has largely lacked clarity on what exploratory programming is, and what behaviors are characteristic of this practice. Drawing on this data and prior literature, we provide an organized description of what exploratory programming has meant historically and a framework of four dimensions for studying exploratory programming tasks: (1) applications, (2) required code quality, (3) ease or difficulty of exploration, and (4) the exploratory process. This provides a basis for better analyzing tool support for exploratory programming. Mary Beth Kery, Brad A. Myers |
VL/HCC | 1 |
| 2017 | Moonstone: Support for understanding and writing exception handling codeabstractMoonstone is a new plugin for Eclipse that supports developers in understanding exception flow and in writing exception handlers in Java. Understanding exception control flow is paramount for writing robust exception handlers, a task many developers struggle with. To help with this understanding, we present two new kinds of information: ghost comments, which are transient overlays that reveal potential sources of exceptions directly in code, and annotated highlights of skipped code and associated handlers. To help developers write better handlers, Moonstone additionally provides project-specific recommendations, detects common bad practices, such as empty or inadequate handlers, and provides automatic resolutions, introducing programmers to advanced Java exception handling features, such as try-with-resources. We present findings from two formative studies that informed the design of Moonstone. We then show with a user study that Moonstone improves users' understanding in certain areas and enables developers to amend exception handling code more quickly and correctly. Florian Kistner, Mary Beth Kery, Michael Puskas, Steven Moore, Brad A. Myers |
VL/HCC | 2 |
| 2016 | Examining programmer practices for locally handling exceptionsabstractMany have argued that the current try/catch mechanism for handling exceptions in Java is flawed. A major complaint is that programmers often write minimal and low quality handlers. We used the Boa tool to examine a large number of Java projects on GitHub to provide empirical evidence about how programmers currently deal with exceptions. We found that programmers handle exceptions locally in catch blocks much of the time, rather than propagating by throwing an Exception. Programmers make heavy use of actions like Log, Print, Return, or Throw in catch blocks, and also frequently copy code between handlers. We found bad practices like empty catch blocks or catching Exception are indeed widespread. We discuss evidence that programmers may misjudge risk when catching Exception, and face a tension between handlers that directly address local program statement failure and handlers that consider the program-wide implications of an exception. Some of these issues might be addressed by future tools which autocomplete more complete handlers. Mary Beth Kery, Claire Le Goues, Brad A. Myers |
MSR | 1 |