Jake M. Hofman

dblp:63/5865 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0002-9364-9604ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7Artificial intelligence and machine learning · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 ChatBench: From Static Benchmarks to Human-AI Evaluation
abstract
With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together.However, standard benchmarks, such as MMLU, measure LLM capabilities in isolation (i.e., "AIalone").Here, we design and conduct a user study to convert MMLU questions into user-AI conversations, by seeding the user with the question and having them carry out a conversation with the LLM to answer their question.We release ChatBench, a new dataset with AIalone, user-alone, and user-AI data for 396 questions and two LLMs, including 144K answers and 7,336 user-AI conversations.We find that AI-alone accuracy fails to predict user-AI accuracy, with significant differences across multiple subjects (math, physics, and moral reasoning), and we analyze the user-AI conversations to provide insight into how they diverge from AI-alone benchmarks.Finally, we show that fine-tuning a user simulator on a subset of ChatBench improves its ability to estimate user-AI accuracies, increasing correlation on heldout questions by more than 20 points, creating possibilities for scaling interactive evaluation. 1 A number is rounded to the nearest hundred.The result is 500.Which number could not be the number before it was rounded to the nearest hundred? A. 458 B. 463 C. 547 D. 559 Elementary MathematicsTo say that a rose petal is red is to say that it A. absorbs red B. reflects red C. emits red D. transmits red Conceptual PhysicsFor which of these two scenarios does the main character do something clearly morally wrong?Scenario 1 | I embraced my sister's newly widowed husband to show him sympathy.Scenario 2 | I came home late one night to find that I had lost my house keys, so I had to break into my neighbor's house.
Serina Chang, Ashton Anderson, Jake M. Hofman
ACL (1)3
2025 Math Education With Large Language Models: Peril or Promise?
David M. Rothschild, Daniel G. Goldstein, Jake M. Hofman
AIED (4)4
2025 Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
Sofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. Hofman
CHI4
2024 Using Open Data to Automatically Generate Localized Analogies
abstract
Numerical analogies (or “perspectives”) that translate unfamiliar measurements into comparisons with familiar reference objects (e.g., “275,000 square miles is roughly as large as Texas”) have been shown to aid readers’ recall, estimation, and error detection for numbers. However, because familiar reference objects are culture-specific, analogies do not always generalize across audiences. Crowdsourcing perspectives has proven effective but is limited by scalability issues and a lack of crowdworking markets in many regions. In this research, we develop an automated technique for generating localized perspectives. We utilize several open data sources for relevance signals and develop a surprisingly simple model capable of localizing analogies to new audiences without any retraining from human judges. We validate the model by testing it in both a new domain and with a different linguistic audience residing in another country. We release the compiled dataset of 400,000 reference objects to the research community.
Sofia Eleni Spatharioti, Daniel G. Goldstein, Jake M. Hofman
CHI3
2022 Putting scientific results in perspective: Improving the communication of standardized effect sizes
abstract
How do people form impressions of effect size when reading scientific results? We present a series of studies on how people perceive treatment effectiveness when scientific results are summarized in various ways. We first show that a prevalent form of summarizing results—presenting mean differences between conditions—can lead to significant overestimation of treatment effectiveness, and that including confidence intervals can exacerbate the problem. We attempt to remedy potential misperceptions by displaying information about variability in individual outcomes in different formats: statements about variance, a quantitative measure of standardized effect size, and analogies that compare the treatment with more familiar effects (e.g., height differences by age). We find that all of these formats substantially reduce potential misperceptions and that analogies can be as helpful as more precise quantitative statements of standardized effect size. These findings can be applied by scientists in HCI and beyond to improve the communication of results to laypeople.
Yea-Seul Kim, Jake M. Hofman, Daniel G. Goldstein
CHI2
2022 Round Numbers Can Sharpen Cognition
abstract
Scientists and journalists strive to report numbers with high precision to keep readers well-informed. Our work investigates whether this practice can backfire due to the cognitive costs of processing multi-digit precise numbers. In a pre-registered randomized experiment, we presented readers with several news stories containing numbers in either precise or round versions. We then measured their ability to approximately recall these numbers and make estimates based on what they read. Our results revealed a counter-intuitive effect where reading round numbers helped people better approximate the precise values, while seeing precise numbers made them worse. We also conducted two surveys to elicit individual preferences for the ideal degree of rounding for numbers spanning seven orders of magnitude in various contexts. From the surveys, we found that people tended to prefer more precision when the rounding options contained only digits (e.g., ”2,500,000”) than when they contained modifier terms (e.g., ”2.5 million”). We conclude with a discussion of how these findings can be leveraged to enhance numeracy in digital content consumption.
Huy Anh Nguyen, Jake M. Hofman, Daniel G. Goldstein
CHI2
2022 Investigating Perceptual Biases in Icon Arrays
abstract
Icon arrays are graphical displays in which a subset of identical shapes are filled to convey probabilities. They are widely used for communicating probabilities to the general public. A primary design decision concerning icon arrays is how to fill and arrange these shapes. For example, a designer could fill the shapes from top to bottom or in a random fashion. We investigated the effect of different arrangements in icon arrays on probability perception. We showed participants icon arrays depicting probabilities between 0% and 100% in six different arrangements. Participants were more accurate in estimating probabilities when viewing the top, row, and diagonal arrangements, but they overestimated the proportions with the central arrangement and underestimated the proportions with the edge arrangement. They were biased to either overestimate or underestimate when viewing the random arrangement depending on the objective proportions, following a cyclical pattern consistent with existing findings in the psychophysics literature.
Cindy Xiong Bearfield, Ali Sarvghad, Daniel G. Goldstein, Jake M. Hofman, Çagatay Demiralp
CHI4
2021 Manipulating and Measuring Model Interpretability
abstract
With machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Although many supposedly interpretable models have been proposed, there have been relatively few experimental studies investigating whether these models achieve their intended effects, such as making people more closely follow a model’s predictions when it is beneficial for them to do so or enabling them to detect when a model has made a mistake. We present a sequence of pre-registered experiments (N = 3, 800) in which we showed participants functionally identical models that varied only in two factors commonly thought to make machine learning models more or less interpretable: the number of features and the transparency of the model (i.e., whether the model internals are clear or black box). Predictably, participants who saw a clear model with few features could better simulate the model’s predictions. However, we did not find that participants more closely followed its predictions. Furthermore, showing participants a clear model meant that they were less able to detect and correct for the model’s sizable mistakes, seemingly due to information overload. These counterintuitive findings emphasize the importance of testing over intuition when developing interpretable models.
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, Hanna M. Wallach
CHI3
2021 Datamations: Animated Explanations of Data Analysis Pipelines
abstract
Plots and tables are commonplace in today’s data-driven world, and much research has been done on how to make these figures easy to read and understand. Often times, however, the information they contain conveys only the end result of a complex and subtle data analysis pipeline. This can leave the reader struggling to understand what steps were taken to arrive at a figure, and what implications this has for the underlying results. In this paper, we introduce datamations, which are animations designed to explain the steps that led to a given plot or table. We present the motivation and concept behind datamations, discuss how to programmatically generate them, and provide the results of two large-scale randomized experiments investigating how datamations affect people’s abilities to understand potentially puzzling results compared to seeing only final plots and tables containing those results.
Xiaoying Pu, Sean Kross, Jake M. Hofman, Daniel G. Goldstein
CHI3
2020 How Visualizing Inferential Uncertainty Can Mislead Readers About Treatment Effects in Scientific Results
abstract
When presenting visualizations of experimental results, scientists often choose to display either inferential uncertainty (e.g., uncertainty in the estimate of a population mean) or outcome uncertainty (e.g., variation of outcomes around that mean) about their estimates. How does this choice impact readers' beliefs about the size of treatment effects? We investigate this question in two experiments comparing 95% confidence intervals (means and standard errors) to 95% prediction intervals (means and standard deviations). The first experiment finds that participants are willing to pay more for and overestimate the effect of a treatment when shown confidence intervals relative to prediction intervals. The second experiment evaluates how alternative visualizations compare to standard visualizations for different effect sizes. We find that axis rescaling reduces error, but not as well as prediction intervals or animated hypothetical outcome plots (HOPs), and that depicting inferential uncertainty causes participants to underestimate variability in individual outcomes.
Jake M. Hofman, Daniel G. Goldstein, Jessica Hullman
CHI1
2018 To Put That in Perspective: Generating Analogies that Make Numbers Easier to Understand
abstract
Laypeople are frequently exposed to unfamiliar numbers published by journalists, social media users, and algorithms. These figures can be difficult for readers to comprehend, especially when they are extreme in magnitude or contain unfamiliar units. Prior work has shown that adding "perspective sentences" that employ ratios, ranks, and unit changes to such measurements can improve people's ability to understand unfamiliar numbers (e.g., "695,000 square kilometers is about the size of Texas"). However, there are many ways to provide context for a measurement. In this paper we systematically test what factors influence the quality of perspective sentences through randomized experiments involving over 1,000 participants. We develop a statistical model for generating perspectives and test it against several alternatives, finding beneficial effects of perspectives on comprehension that persist for six weeks. We conclude by discussing future work in deploying and testing perspectives at scale.
Christopher Riederer, Jake M. Hofman, Daniel G. Goldstein
CHI2
2016 Improving Comprehension of Numbers in the News
abstract
How many guns are there in the USA? What is the incidence of breast cancer? Is a billion dollar budget cut large or small? Advocates of scientific and civic literacy are concerned with improving how people estimate and comprehend risks, measurements, and frequencies, but relatively little progress has been made in this direction. In this article we describe and test a framework to help people comprehend numerical measurements in everyday settings through simple sentences, termed perspectives, that employ ratios, ranks, and unit changes to make them easier to understand. We use a crowdsourced system to generate perspectives for a wide range of numbers taken from online news articles. We then test the effectiveness of these perspectives in three randomized, online experiments involving over 3,200 participants. We find that perspective clauses substantially improve people's ability to recall measurements they have read, estimate ones they have not, and detect errors in manipulated measurements. We see this as the first of many steps in leveraging digital platforms to improve numeracy among online readers.
Pablo Barrio 0002, Daniel G. Goldstein, Jake M. Hofman
CHI3
2016 Exploring Limits to Prediction in Complex Social Systems
abstract
How predictable is success in complex social systems? In spite of a recent profusion of prediction studies that exploit online social and information network data, this question remains unanswered, in part because it has not been adequately specified. In this paper we attempt to clarify the question by presenting a simple stylized model of success that attributes prediction error to one of two generic sources: insufficiency of available data and/or models on the one hand; and inherent unpredictability of complex social systems on the other. We then use this model to motivate an illustrative empirical study of information cascade size prediction on Twitter. Despite an unprecedented volume of information about users, content, and past performance, our best performing models can explain less than half of the variance in cascade sizes. In turn, this result suggests that even with unlimited data predictive performance would be bounded well below deterministic accuracy. Finally, we explore this potential bound theoretically using simulations of a diffusion process on a random scale free network similar to Twitter. We show that although higher predictive power is possible in theory, such performance requires a homogeneous system and perfect ex-ante knowledge of it: even a small degree of uncertainty in estimating product quality or slight variation in quality across products leads to substantially more restrictive bounds on predictability. We conclude that realistic bounds on predictive accuracy are not dissimilar from those we have obtained empirically, and that such bounds for other complex social systems for which data is more difficult to obtain are likely even lower.
Travis Martin, Jake M. Hofman, Amit Sharma 0007, Ashton Anderson, Duncan J. Watts
WWW2
2015 Estimating the Causal Impact of Recommendation Systems from Observational Data
abstract
Recommendation systems are an increasingly prominent part of the web, accounting for up to a third of all traffic on several of the world's most popular sites. Nevertheless, little is known about how much activity such systems actually cause over and above activity that would have occurred via other means (e.g., search) if recommendations were absent. Although the ideal way to estimate the causal impact of recommendations is via randomized experiments, such experiments are costly and may inconvenience users. In this paper, therefore, we present a method for estimating causal effects from purely observational data. Specifically, we show that causal identification through an instrumental variable is possible when a product experiences an instantaneous shock in direct traffic and the products recommended next to it do not. We then apply our method to browsing logs containing anonymized activity for 2.1 million users on Amazon.com over a 9 month period and analyze over 4,000 unique products that experience such shocks. We find that although recommendation click-throughs do account for a large fraction of traffic among these products, at least 75% of this activity would likely occur in the absence of recommendations. We conclude with a discussion about the assumptions under which the method is appropriate and caveats around extrapolating results to other products, sites, or settings.
Amit Sharma 0007, Jake M. Hofman, Duncan J. Watts
EC2
2015 Scalable Recommendation with Hierarchical Poisson Factorization
Prem Gopalan, Jake M. Hofman, David M. Blei
UAI2
2013 Sharding social networks
abstract
Online social networking platforms regularly support hundreds of millions of users, who in aggregate generate substantially more data than can be stored on any single physical server. As such, user data are distributed, or sharded, across many machines. A key requirement in this setting is rapid retrieval not only of a given user's information, but also of all data associated with his or her social contacts, suggesting that one should consider the topology of the social network in selecting a sharding policy. In this paper we formalize the problem of efficiently sharding large social network databases, and evaluate several sharding strategies, both analytically and empirically. We find that random sharding---the de facto standard---results in provably poor performance even when frequently accessed nodes are replicated to many shards. By contrast, we demonstrate that one can substantially reduce querying costs by identifying and assigning tightly knit communities to shards. In particular, our theoretical analysis motivates a novel, scalable sharding algorithm that outperforms both random and location-based sharding schemes.
Quang Duong 0004, Sharad Goel, Jake M. Hofman, Sergei Vassilvitskii
WSDM3
2012 Who Does What on the Web: A Large-Scale Study of Browsing Behavior
Sharad Goel, Jake M. Hofman, M. Irmak Sirer
ICWSM2
2011 Everyone's an influencer: quantifying influence on twitter
abstract
In this paper we investigate the attributes and relative influence of 1.6M Twitter users by tracking 74 million diffusion events that took place on the Twitter follower graph over a two month interval in 2009. Unsurprisingly, we find that the largest cascades tend to be generated by users who have been influential in the past and who have a large number of followers. We also find that URLs that were rated more interesting and/or elicited more positive feelings by workers on Mechanical Turk were more likely to spread. In spite of these intuitive results, however, we find that predictions of which particular user or URL will generate large cascades are relatively unreliable. We conclude, therefore, that word-of-mouth diffusion can only be harnessed reliably by targeting large numbers of potential influencers, thereby capturing average effects. Finally, we consider a family of hypothetical marketing strategies, defined by the relative cost of identifying versus compensating potential "influencers." We find that although under some circumstances, the most influential users are also the most cost-effective, under a wide range of plausible assumptions the most cost-effective performance can be realized using "ordinary influencers"---individuals who exert average or even less-than-average influence.
Eytan Bakshy, Jake M. Hofman, Winter A. Mason, Duncan J. Watts
WSDM2
2011 Who says what to whom on twitter
abstract
We study several longstanding questions in media communications research, in the context of the microblogging service Twitter, regarding the production, flow, and consumption of information. To do so, we exploit a recently introduced feature of Twitter known as "lists" to distinguish between elite users - by which we mean celebrities, bloggers, and representatives of media outlets and other formal organizations - and ordinary users. Based on this classification, we find a striking concentration of attention on Twitter, in that roughly 50% of URLs consumed are generated by just 20K elite users, where the media produces the most information, but celebrities are the most followed. We also find significant homophily within categories: celebrities listen to celebrities, while bloggers listen to bloggers etc; however, bloggers in general rebroadcast more information than the other categories. Next we re-examine the classical "two-step flow" theory of communications, finding considerable support for it on Twitter. Third, we find that URLs broadcast by different categories of users or containing different types of content exhibit systematically different lifespans. And finally, we examine the attention paid by the different user categories to different news topics.
Shaomei Wu, Jake M. Hofman, Winter A. Mason, Duncan J. Watts
WWW2
2010 Inferring relevant social networks from interpersonal communication
abstract
Researchers increasingly use electronic communication data to construct and study large social networks, effectively inferring unobserved ties (e.g. i is connected to j) from observed communication events (e.g. i emails j). Often overlooked, however, is the impact of tie definition on the corresponding network, and in turn the relevance of the inferred network to the research question of interest. Here we study the problem of network inference and relevance for two email data sets of different size and origin. In each case, we generate a family of networks parameterized by a threshold condition on the frequency of emails exchanged between pairs of individuals. After demonstrating that different choices of the threshold correspond to dramatically different network structures, we then formulate the relevance of these networks in terms of a series of prediction tasks that depend on various network features. In general, we find: a) that prediction accuracy is maximized over a non-trivial range of thresholds corresponding to 5-10 reciprocated emails per year; b) that for any prediction task, choosing the optimal value of the threshold yields a sizable (~30%) boost in accuracy over naive choices; and c) that the optimal threshold value appears to be (somewhat surprisingly) consistent across data sets and prediction tasks. We emphasize the practical utility in defining ties via their relevance to the prediction task(s) at hand and discuss implications of our empirical results.
Munmun De Choudhury, Winter A. Mason, Jake M. Hofman, Duncan J. Watts
WWW3
2010 Graphical models for inferring single molecule dynamics
abstract
BACKGROUND: The recent explosion of experimental techniques in single molecule biophysics has generated a variety of novel time series data requiring equally novel computational tools for analysis and inference. This article describes in general terms how graphical modeling may be used to learn from biophysical time series data using the variational Bayesian expectation maximization algorithm (VBEM). The discussion is illustrated by the example of single-molecule fluorescence resonance energy transfer (smFRET) versus time data, where the smFRET time series is modeled as a hidden Markov model (HMM) with Gaussian observables. A detailed description of smFRET is provided as well. RESULTS: The VBEM algorithm returns the model's evidence and an approximating posterior parameter distribution given the data. The former provides a metric for model selection via maximum evidence (ME), and the latter a description of the model's parameters learned from the data. ME/VBEM provide several advantages over the more commonly used approach of maximum likelihood (ML) optimized by the expectation maximization (EM) algorithm, the most important being a natural form of model selection and a well-posed (non-divergent) optimization problem. CONCLUSIONS: The results demonstrate the utility of graphical modeling for inference of dynamic processes in single molecule biophysics.
Jonathan E. Bronson, Jake M. Hofman, Jingyi Fei, Ruben L. Gonzalez, Chris Wiggins 0001
BMC Bioinform.2
2009 Characterizing individual communication patterns
abstract
The increasing availability of electronic communication data, such as that arising from e-mail exchange, presents social and information scientists with new possibilities for characterizing individual behavior and, by extension, identifying latent structure in human populations. Here, we propose a model of individual e-mail communication that is sufficiently rich to capture meaningful variability across individuals, while remaining simple enough to be interpretable. We show that the model, a cascading non-homogeneous Poisson process, can be formulated as a double-chain hidden Markov model, allowing us to use an efficient inference algorithm to estimate the model parameters from observed data. We then apply this model to two e-mail data sets consisting of 404 and 6,164 users, respectively, that were collected from two universities in different countries and years. We find that the resulting best-estimate parameter distributions for both data sets are surprisingly similar, indicating that at least some features of communication dynamics generalize beyond specific contexts. We also find that variability of individual behavior over time is significantly less than variability across the population, suggesting that individuals can be classified into persistent "types". We conclude that communication patterns may prove useful as an additional class of attribute data, complementing demographic and network data, for user classification and outlier detection-a point that we illustrate with an interpretable clustering of users based on their inferred model parameters.
R. Dean Malmgren, Jake M. Hofman, Luis A. Nunes Amaral, Duncan J. Watts
KDD2