VLDB 2026 Research / reviewers in the wild / expert
Daniel G. Goldstein
dblp:46/9704 · also Daniel Gray Goldstein
· DBLP profile ↗
19ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-0970-5598ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 12 · 8 since 2021Artificial intelligence and machine learning · 6 · 5 first-authorTheory of computation · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Math Education With Large Language Models: Peril or Promise?
David M. Rothschild, Daniel G. Goldstein, Jake M. Hofman |
AIED (4) | 3 |
| 2025 | Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
Sofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. Hofman |
CHI | 3 |
| 2024 | Using Open Data to Automatically Generate Localized AnalogiesabstractNumerical analogies (or “perspectives”) that translate unfamiliar measurements into comparisons with familiar reference objects (e.g., “275,000 square miles is roughly as large as Texas”) have been shown to aid readers’ recall, estimation, and error detection for numbers. However, because familiar reference objects are culture-specific, analogies do not always generalize across audiences. Crowdsourcing perspectives has proven effective but is limited by scalability issues and a lack of crowdworking markets in many regions. In this research, we develop an automated technique for generating localized perspectives. We utilize several open data sources for relevance signals and develop a surprisingly simple model capable of localizing analogies to new audiences without any retraining from human judges. We validate the model by testing it in both a new domain and with a different linguistic audience residing in another country. We release the compiled dataset of 400,000 reference objects to the research community. Sofia Eleni Spatharioti, Daniel G. Goldstein, Jake M. Hofman |
CHI | 2 |
| 2022 | Putting scientific results in perspective: Improving the communication of standardized effect sizesabstractHow do people form impressions of effect size when reading scientific results? We present a series of studies on how people perceive treatment effectiveness when scientific results are summarized in various ways. We first show that a prevalent form of summarizing results—presenting mean differences between conditions—can lead to significant overestimation of treatment effectiveness, and that including confidence intervals can exacerbate the problem. We attempt to remedy potential misperceptions by displaying information about variability in individual outcomes in different formats: statements about variance, a quantitative measure of standardized effect size, and analogies that compare the treatment with more familiar effects (e.g., height differences by age). We find that all of these formats substantially reduce potential misperceptions and that analogies can be as helpful as more precise quantitative statements of standardized effect size. These findings can be applied by scientists in HCI and beyond to improve the communication of results to laypeople. Yea-Seul Kim, Jake M. Hofman, Daniel G. Goldstein |
CHI | 3 |
| 2022 | Round Numbers Can Sharpen CognitionabstractScientists and journalists strive to report numbers with high precision to keep readers well-informed. Our work investigates whether this practice can backfire due to the cognitive costs of processing multi-digit precise numbers. In a pre-registered randomized experiment, we presented readers with several news stories containing numbers in either precise or round versions. We then measured their ability to approximately recall these numbers and make estimates based on what they read. Our results revealed a counter-intuitive effect where reading round numbers helped people better approximate the precise values, while seeing precise numbers made them worse. We also conducted two surveys to elicit individual preferences for the ideal degree of rounding for numbers spanning seven orders of magnitude in various contexts. From the surveys, we found that people tended to prefer more precision when the rounding options contained only digits (e.g., ”2,500,000”) than when they contained modifier terms (e.g., ”2.5 million”). We conclude with a discussion of how these findings can be leveraged to enhance numeracy in digital content consumption. Huy Anh Nguyen, Jake M. Hofman, Daniel G. Goldstein |
CHI | 3 |
| 2022 | Investigating Perceptual Biases in Icon ArraysabstractIcon arrays are graphical displays in which a subset of identical shapes are filled to convey probabilities. They are widely used for communicating probabilities to the general public. A primary design decision concerning icon arrays is how to fill and arrange these shapes. For example, a designer could fill the shapes from top to bottom or in a random fashion. We investigated the effect of different arrangements in icon arrays on probability perception. We showed participants icon arrays depicting probabilities between 0% and 100% in six different arrangements. Participants were more accurate in estimating probabilities when viewing the top, row, and diagonal arrangements, but they overestimated the proportions with the central arrangement and underestimated the proportions with the edge arrangement. They were biased to either overestimate or underestimate when viewing the random arrangement depending on the objective proportions, following a cyclical pattern consistent with existing findings in the psychophysics literature. Cindy Xiong Bearfield, Ali Sarvghad, Daniel G. Goldstein, Jake M. Hofman, Çagatay Demiralp |
CHI | 3 |
| 2021 | Manipulating and Measuring Model InterpretabilityabstractWith machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Although many supposedly interpretable models have been proposed, there have been relatively few experimental studies investigating whether these models achieve their intended effects, such as making people more closely follow a model’s predictions when it is beneficial for them to do so or enabling them to detect when a model has made a mistake. We present a sequence of pre-registered experiments (N = 3, 800) in which we showed participants functionally identical models that varied only in two factors commonly thought to make machine learning models more or less interpretable: the number of features and the transparency of the model (i.e., whether the model internals are clear or black box). Predictably, participants who saw a clear model with few features could better simulate the model’s predictions. However, we did not find that participants more closely followed its predictions. Furthermore, showing participants a clear model meant that they were less able to detect and correct for the model’s sizable mistakes, seemingly due to information overload. These counterintuitive findings emphasize the importance of testing over intuition when developing interpretable models. Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, Hanna M. Wallach |
CHI | 2 |
| 2021 | Datamations: Animated Explanations of Data Analysis PipelinesabstractPlots and tables are commonplace in today’s data-driven world, and much research has been done on how to make these figures easy to read and understand. Often times, however, the information they contain conveys only the end result of a complex and subtle data analysis pipeline. This can leave the reader struggling to understand what steps were taken to arrive at a figure, and what implications this has for the underlying results. In this paper, we introduce datamations, which are animations designed to explain the steps that led to a given plot or table. We present the motivation and concept behind datamations, discuss how to programmatically generate them, and provide the results of two large-scale randomized experiments investigating how datamations affect people’s abilities to understand potentially puzzling results compared to seeing only final plots and tables containing those results. Xiaoying Pu, Sean Kross, Jake M. Hofman, Daniel G. Goldstein |
CHI | 4 |
| 2020 | How Visualizing Inferential Uncertainty Can Mislead Readers About Treatment Effects in Scientific ResultsabstractWhen presenting visualizations of experimental results, scientists often choose to display either inferential uncertainty (e.g., uncertainty in the estimate of a population mean) or outcome uncertainty (e.g., variation of outcomes around that mean) about their estimates. How does this choice impact readers' beliefs about the size of treatment effects? We investigate this question in two experiments comparing 95% confidence intervals (means and standard errors) to 95% prediction intervals (means and standard deviations). The first experiment finds that participants are willing to pay more for and overestimate the effect of a treatment when shown confidence intervals relative to prediction intervals. The second experiment evaluates how alternative visualizations compare to standard visualizations for different effect sizes. We find that axis rescaling reduces error, but not as well as prediction intervals or animated hypothetical outcome plots (HOPs), and that depicting inferential uncertainty causes participants to underestimate variability in individual outcomes. Jake M. Hofman, Daniel G. Goldstein, Jessica Hullman |
CHI | 2 |
| 2018 | To Put That in Perspective: Generating Analogies that Make Numbers Easier to UnderstandabstractLaypeople are frequently exposed to unfamiliar numbers published by journalists, social media users, and algorithms. These figures can be difficult for readers to comprehend, especially when they are extreme in magnitude or contain unfamiliar units. Prior work has shown that adding "perspective sentences" that employ ratios, ranks, and unit changes to such measurements can improve people's ability to understand unfamiliar numbers (e.g., "695,000 square kilometers is about the size of Texas"). However, there are many ways to provide context for a measurement. In this paper we systematically test what factors influence the quality of perspective sentences through randomized experiments involving over 1,000 participants. We develop a statistical model for generating perspectives and test it against several alternatives, finding beneficial effects of perspectives on comprehension that persist for six weeks. We conclude by discussing future work in deploying and testing perspectives at scale. Christopher Riederer, Jake M. Hofman, Daniel G. Goldstein |
CHI | 3 |
| 2017 | VoxPL: Programming with the Wisdom of the CrowdabstractHaving a crowd estimate a numeric value is the original inspiration for the notion of "the wisdom of the crowd." Quality control for such estimated values is challenging because prior, consensus-based approaches for quality control in labeling tasks are not applicable in estimation tasks. We present VoxPL, a high-level programming framework that automatically obtains high-quality crowdsourced estimates of values. The VoxPL domain-specific language lets programmers concisely specify complex estimation tasks with a desired level of confidence and budget. VoxPL's runtime system implements a novel quality control algorithm that automatically computes sample sizes and obtains high quality estimates from the crowd at low cost. To evaluate VoxPL, we implement four estimation applications, ranging from facial feature recognition to calorie counting. The resulting programs are concise---under 200 lines of code---and obtain high quality estimates from the crowd quickly and inexpensively. Daniel W. Barowy, Emery D. Berger, Daniel G. Goldstein, Siddharth Suri |
CHI | 3 |
| 2017 | Learning in the Repeated Secretary ProblemabstractIn the classical secretary problem, one attempts to find the maximum of an unknown and unlearnable distribution through sequential search. In many real-world searches, however, distributions are not entirely unknown and can be learned through experience. To investigate learning in such a repeated secretary problem we conduct a large-scale behavioral experiment in which people search repeatedly from fixed distributions. In contrast to prior investigations that find no evidence for learning in the classical scenario, in the repeated setting we observe substantial learning resulting in near-optimal stopping behavior. We conduct a Bayesian comparison of multiple behavioral models which shows that participants' behavior is best described by a class of threshold-based models that contains the theoretically optimal strategy. In fact, fitting such a threshold-based model to data reveals players' estimated thresholds to be surprisingly close to the optimal thresholds after only a small number of games. Daniel G. Goldstein, R. Preston McAfee, Siddharth Suri, James R. Wright |
EC | 1 |
| 2016 | Improving Comprehension of Numbers in the NewsabstractHow many guns are there in the USA? What is the incidence of breast cancer? Is a billion dollar budget cut large or small? Advocates of scientific and civic literacy are concerned with improving how people estimate and comprehend risks, measurements, and frequencies, but relatively little progress has been made in this direction. In this article we describe and test a framework to help people comprehend numerical measurements in everyday settings through simple sentences, termed perspectives, that employ ratios, ranks, and unit changes to make them easier to understand. We use a crowdsourced system to generate perspectives for a wide range of numbers taken from online news articles. We then test the effectiveness of these perspectives in three randomized, online experiments involving over 3,200 participants. We find that perspective clauses substantially improve people's ability to recall measurements they have read, estimate ones they have not, and detect errors in manipulated measurements. We see this as the first of many steps in leveraging digital platforms to improve numeracy among online readers. Pablo Barrio 0002, Daniel G. Goldstein, Jake M. Hofman |
CHI | 2 |
| 2014 | The wisdom of smaller, smarter crowdsabstractThe "wisdom of crowds" refers to the phenomenon that aggregated predictions from a large group of people can rival or even beat the accuracy of experts. In domains with substantial stochastic elements, such as stock picking, crowd strategies (e.g. indexing) are difficult to beat. However, in domains in which some crowd members have demonstrably more skill than others, smart sub-crowds could possibly outperform the whole. The central question this work addresses is whether such smart subsets of a crowd can be identified a priori in a large-scale prediction contest that has substantial skill and luck components. We study this question with data obtained from fantasy soccer, a game in which millions of people choose professional players from the English Premier League to be on their fantasy soccer teams. The better the professional players do in real life games, the more points fantasy teams earn. Fantasy soccer is ideally suited to this investigation because it comprises millions of individual-level, within-subject predictions, past performance indicators, and the ability to test the effectiveness of arbitrary player-selection strategies. We find that smaller, smarter crowds can be identified in advance and that they beat the wisdom of the larger crowd. We also show that many players would do better by simply imitating the strategy of a player who has done well in the past. Finally, we provide a theoretical model that explains the results we see from our empirical analyses. Daniel G. Goldstein, R. Preston McAfee, Siddharth Suri |
EC | 1 |
| 2013 | Improving the Effectiveness of Time-Based Display Advertising
Daniel G. Goldstein, R. Preston McAfee, Siddharth Suri |
IJCAI | 1 |
| 2013 | The cost of annoying adsabstractDisplay advertisements vary in the extent to which they annoy users. While publishers know the payment they receive to run annoying ads, little is known about the cost such ads incur due to user abandonment. We conducted a two-experiment investigation to analyze ad features that relate to annoyingness and to put a monetary value on the cost of annoying ads. The first experiment asked users to rate and comment on a large number of ads taken from the Web. This allowed us to establish sets of annoying and innocuous ads for use in the second experiment, in which users were given the opportunity to categorize emails for a per-message wage and quit at any time. Participants were randomly assigned to one of three different pay rates and also randomly assigned to categorize the emails in the presence of no ads, annoying ads, or innocuous ads. Since each email categorization constituted an impression, this design, inspired by Toomim et al., allowed us to determine how much more one must pay a person to generate the same number of impressions in the presence of annoying ads compared to no ads or innocuous ads. We conclude by proposing a theoretical model which relates ad quality to publisher market share, illustrating how our empirical findings could affect the economics of Internet advertising. Daniel G. Goldstein, R. Preston McAfee, Siddharth Suri |
WWW | 1 |
| 2012 | The structure of online diffusion networksabstractModels of networked diffusion that are motivated by analogy with the spread of infectious disease have been applied to a wide range of social and economic adoption processes, including those related to new products, ideas, norms and behaviors. However, it is unknown how accurately these models account for the empirical structure of diffusion over networks. Here we describe the diffusion patterns arising from seven online domains, ranging from communications platforms to networked games to microblogging services, each involving distinct types of content and modes of sharing. We find strikingly similar patterns across all domains. Sharad Goel, Duncan J. Watts, Daniel G. Goldstein |
EC | 3 |
| 2012 | Improving the effectiveness of time-based display advertisingabstractDisplay advertisements are typically sold by the impression, where one impression is simply one download of an ad. Previous work has shown that the longer an ad is in view, the more likely a user is to remember it and that there are diminishing returns to increased exposure time [Goldstein et al. 2011]. Since a pricing scheme that is at least partially based on time is more exact than one based solely on impressions, time- based advertising may become an industry standard. We answer an open question concerning time-based pricing schemes: how should time slots for advertisements be divided? We provide evidence that ads can be scheduled in a way that leads to greater total recollection, which advertisers value, and increased revenue, which publishers value. We document two main findings. First, we show that displaying two shorter ads results in more total recollection than displaying one longer ad of twice the duration. Second, we show that this effect disappears as the duration of these ads increases. We conclude with a theoretical prediction regarding the circumstances under which the display advertising industry would benefit if it moved to a partially or fully time-based standard. Daniel G. Goldstein, R. Preston McAfee, Siddharth Suri |
EC | 1 |
| 2011 | The effects of exposure time on memory of display advertisementsabstractDisplay advertising is a multi-billion dollar industry that has traditionally used a pricing scheme based on the number of impressions delivered. The number of impressions of an ad is simply the number of downloads of that ad. One impression, however, does not differentiate between an ad that is in view for five seconds or five minutes. Since advertisers seek brand recognition and recall, we ask whether a time-based accounting of advertising can better align with advertisers' goals. This work aims to model the basic relationship between ad exposure time and the probability that a viewer will remember an advertisement. We investigate this question via two behavioral experiments, conducted using Amazon Mechanical Turk, in which people viewed Web pages accompanied by ads. The amount of time the ads were in view was either determined endogenously (as a function of reading speed) or exogenously (as a function of a timer and random assignment). Our results suggest that for exposure times of up to one minute, there is a strong, causal influence of exposure time on ad recognition and recall, with the marginal effects diminishing at durations beyond this level. Simple models describing memory response as a function of the logarithm of exposure time provide a good fit. In addition, we find that advertisements that are displayed when the Web page loads attain greater marginal increases in recognition per unit time than do ads that come into view second in a sequence. Nonetheless, for both types of ads, exposure time has a substantial effect. A psychologically-informed accounting system based on ad exposure duration, sequence and onset time may more closely align with advertiser goals than the industry standard of impression-based accounting. Daniel G. Goldstein, R. Preston McAfee, Siddharth Suri |
EC | 1 |