EDBT 2026 Demo / reviewers in the wild / expert
Matthew Kay 0001
dblp:85/2585
· DBLP profile ↗
58ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0001-9446-0419ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 40 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 14 since 2021Security and privacy · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Through a Live Elections Dashboard, Darkly: Managing Expectations and Trust in Progressive Vote Counting During the 2024 U.S. Election
Mandi Cai, Chloe Mortenson, Fumeng Yang, Erik C. Nisbet, Matthew Kay 0001 |
CHI | 6 |
| 2026 | Codesigning Ripplet: an LLM-Assisted Assessment Authoring System Grounded in a Conceptual Model of Teachers' WorkflowsabstractAssessments are critical in education, but creating them can be difficult. To address this challenge in a grounded way, we partnered with 13 teachers in a seven-month codesign process. We developed a conceptual model that characterizes the iterative dual process where teachers develop assessments while simultaneously refining requirements. To enact this model in practice, we built Ripplet,1 a web-based tool with multilevel reusable interactions to support assessment authoring. The extended codesign revealed that Ripplet enabled teachers to create formative assessments they would not have otherwise made, shifted their practices from generation to curation, and helped them reflect more on assessment quality. In a user study with 15 additional teachers, compared to their current practices, teachers felt the results were more worth their effort and that assessment quality improved. Annabel Goldman, Jovy Zhou, Clarissa M. Shieh, Joshua Yao, Mia Lillian Cheng, Matthew Kay 0001, Fumeng Yang |
CHI | 8 |
| 2026 | An Autoethnography on Visualization Literacy: A Wicked Measurement ProblemabstractWe contribute an autoethnographic reflection on the complexity of defining and measuring visualization literacy (i.e., the ability to interpret and construct visualizations) to expose our tacit thoughts that often exist in-between polished works and remain unreported in individual research papers. Our work is inspired by the growing number of empirical studies in visualization research that rely on visualization literacy as a basis for developing effective data representations or educational interventions. Researchers have already made various efforts to assess this construct, yet it is often hard to pinpoint either what we want to measure or what we are effectively measuring. In this autoethnography, we gather insights from 14 internal interviews with researchers who are users or designers of visualization literacy tests. We aim to identify what makes visualization literacy assessment a "wicked" problem. We further reflect on the fluidity of visualization literacy and discuss how this property may lead to misalignment between what the construct is and how measurements of it are used or designed. We also examine potential threats to measurement validity from conceptual, operational, and methodological perspectives. Based on our experiences and reflections, we propose several calls to action aimed at tackling the wicked problem of visualization literacy measurement, such as by broadening test scopes and modalities, improving test ecological validity, making it easier to use tests, seeking interdisciplinary collaboration, and drawing from continued dialogue on visualization literacy to expect and be more comfortable with its fluidity. Lily W. Ge, Anne-Flore Cabouat, Karen Bonilla, Yiren Ding, Noëlle Rakotondravony, Mackenzie Michael Creamer, Jasmine Otto, Maryam Hedayati, Bum Chul Kwon, Angela Locoro, Lane Harrison, Petra Isenberg, Michael Correll, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 15 |
| 2025 | AVEC: An Assessment of Visual Encoding Ability in Visualization Construction
Lily W. Ge, Matthew Kay 0001 |
CHI | 3 |
| 2025 | Seeing Eye to AI? Applying Deep-Feature-Based Similarity Metrics to Information VisualizationabstractJudging the similarity of visualizations is crucial to various applications, such as visualization-based search and visualization recommendation systems.Recent studies show deep-feature-based similarity metrics correlate well with perceptual judgments of image similarity and serve as effective loss functions for tasks like image super-resolution and style transfer.We explore the application of such metrics to judgments of visualization similarity.We extend a similarity metric using five ML architectures and three pre-trained weight sets.We replicate results from previous crowdsourced studies on scatterplot and visual channel similarity perception.Notably, our metric using pre-trained ImageNet weights outperformed gradient-descent tuned MS-SSIM, a multi-scale similarity metric based on luminance, contrast, and structure.Our work contributes to understanding how deep-feature-based metrics can enhance similarity assessments in visualization, potentially improving visual analysis tools and techniques.Supplementary materials are available at https://osf.io/dj2ms/. Sheng Long 0001, Angelos Chatzimparmpas, Emma Alexander, Matthew Kay 0001, Jessica Hullman |
CHI | 4 |
| 2025 | More Forecasts, More (Decision) Problems: How Uncertainty Representations for Multiple Forecasts Impact Decision MakingabstractUsers often have access to multiple forecasts regarding an event.Different forecasts incorporate different assumptions and epistemic information.A growing body of work argues against decisionmaking solely based on expected utility maximisation strategies in multiple forecasts scenarios, in favour of other strategies such as the maximin expected utility.In this work, we compare two different approaches for depicting epistemic uncertainty-ensembles (a direct representation of multiple forecasts) and p-boxes (a representation which only communicates the bounds of epistemic uncertainty)-in plots where individual distributions are represented as cumulative distribution plots (CDFs).We conduct three experiments to investigate the impact of the visual representation on the decision-making strategies that people adopt.Our results suggest that participants adopt conservative decision-making strategies (i.e.place greater weight on the worst-case forecast than the best-case forecast) for both p-boxes and ensembles if the set of forecasts are uniformly distributed.However, if a majority of the forecasts are clustered near one of the bounds, participants may discount the forecast which appears as a visual outlier. Abhraneel Sarma, Maryam Hedayati, Matthew Kay 0001 |
CHI | 3 |
| 2025 | Promises and Pitfalls: Using Large Language Models to Generate Visualization ItemsabstractVisualization items-factual questions about visualizations that ask viewers to accomplish visualization tasks-are regularly used in the field of information visualization as educational and evaluative materials. For example, researchers of visualization literacy require large, diverse banks of items to conduct studies where the same skill is measured repeatedly on the same participants. Yet, generating a large number of high-quality, diverse items requires significant time and expertise. To address the critical need for a large number of diverse visualization items in education and research, this paper investigates the potential for large language models (LLMS) to automate the generation of multiple-choice visualization items. Through an iterative design process, we develop the VILA (Visualization Items Generated by Large LAnguage Models) pipeline, for efficiently generating visualization items that measure people's ability to accomplish visualization tasks. We use the VILA pipeline to generate 1,404 candidate items across 12 chart types and 13 visualization tasks. In collaboration with 11 visualization experts, we develop an evaluation rulebook which we then use to rate the quality of all candidate items. The result is the VILA bank of ~1, 100 items. From this evaluation, we also identify and classify current limitations of the VILA pipeline, and discuss the role of human oversight in ensuring quality. In addition, we demonstrate an application of our work by creating a visualization literacy test, VILA-VLAT, which measures people's ability to complete a diverse set of tasks on various types of visualizations; comparing it to the existing VLAT, VILA-VLAT shows moderate to high convergent validity (R = 0.70). Lastly, we discuss the application areas of the VILA pipeline and the VILA bank and provide practical recommendations for their use. All supplemental materials are available at https://osf.io/ysrhq/. Lily W. Ge, Yiren Ding, Lane Harrison, Fumeng Yang, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | VMC: A Grammar for Visualizing Statistical Model ChecksabstractVisualizations play a critical role in validating and improving statistical models. However, the design space of model check visualizations is not well understood, making it difficult for authors to explore and specify effective graphical model checks. VMC defines a model check visualization using four components: (1) samples of distributions of checkable quantities generated from the model, including predictive distributions for new data and distributions of model parameters; (2) transformations on observed data to facilitate comparison; (3) visual representations of distributions; and (4) layouts to facilitate comparing model samples and observed data. We contribute an implementation of VMC as an R package. We validate VMC by reproducing a set of canonical model check examples, and show how using VMC to generate model checks reduces the edit distance between visualizations relative to existing visualization toolkits. The findings of an interview study with three expert modelers who used VMC highlight challenges and opportunities for encouraging exploration of correct, effective model check visualizations. Alex Kale, Matthew Kay 0001, Jessica Hullman |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | What University Students Learn In Visualization ClassesabstractAs a step towards improving visualization literacy, this work investigates how students approach reading visualizations differently after taking a university-level visualization course. We asked students to verbally walk through their process of making sense of unfamiliar visualizations, and conducted a qualitative analysis of these walkthroughs. Our qualitative analysis found that after taking a visualization course, students engaged with visualizations in more sophisticated ways: they were more likely to exhibit design empathy by thinking critically about the tradeoffs behind why a chart was designed in a particular way, and were better able to deconstruct a chart to make sense of it. We also gave students a quantitative assessment of visualization literacy and found no evidence of scores improving after the class, likely because the test we used focused on a different set of skills than those emphasized in visualization classes. While current measurement instruments for visualization literacy are useful, we propose developing standardized assessments for additional aspects of visualization literacy, such as deconstruction and design empathy. We also suggest that these additional aspects could be incorporated more explicitly in visualization courses. All supplemental materials are available at https://osf.io/w5pum/. Maryam Hedayati, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | PrefaceabstractThis January 2025 issue of the IEEE Transactions on Visualization and Computer Graphics (TVCG) contains the proceedings of IEEE VIS 2024, held on October 1318 October, 2024 in St. Pete Beach, Florida, USA, with the three General Chairs Paul Rosen (University of Utah), Kristi Potter (U.S. National Renewable Energy Laboratory), and Remco Chang (Tufts University). With IEEE VIS 2024, the conference series is in its 35th year. Tamara Munzner, Niklas Elmqvist, Holger Theisel, Matthew Kay 0001, Adam Perer, Tatiana von Landesberger, Jiawan Zhang, Christoph Garth, Chaoli Wang 0001, Pierre Dragicevic, Daniel F. Keefe, Filip Sadlo, Ivan Viola, Wenwen Dou, Steffen Koch 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | The Backstory to "Swaying the Public": A Design Chronicle of Election Forecast VisualizationsabstractA year ago, we submitted an IEEE VIS paper entitled "Swaying the Public? Impacts of Election Forecast Visualizations on Emotion, Trust, and Intention in the 2022 U.S. Midterms" [50], which was later bestowed with the honor of a best paper award. Yet, studying such a complex phenomenon required us to explore many more design paths than we could count, and certainly more than we could document in a single paper. This paper, then, is the unwritten prequel-the backstory. It chronicles our journey from a simple idea-to study visualizations for election forecasts-through obstacles such as developing meaningfully different, easy-to-understand forecast visualizations, crafting professional-looking forecasts, and grappling with how to study perceptions of the forecasts before, during, and after the 2022 U.S. midterm elections. This journey yielded a rich set of original knowledge. We formalized a design space for two-party election forecasts, navigating through dimensions like data transformations, visual channels, and types of animated narratives. Through qualitative evaluation of ten representative prototypes with 13 participants, we then identified six core insights into the interpretation of uncertainty visualizations in a U.S. election context. These insights informed our revisions to remove ambiguity in our visual encodings and to prepare a professional-looking forecasting website. As part of this story, we also distilled challenges faced and design lessons learned to inform both designers and practitioners. Ultimately, we hope our methodical approach could inspire others in the community to tackle the hard problems inherent to designing and evaluating visualizations for the general public. Fumeng Yang, Mandi Cai, Chloe Mortenson, Hoda Fakhari, Ayse D. Lokmanoglu, Nicholas Diakopoulos, Erik C. Nisbet, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | Watching the Election Sausage Get Made: How Data Journalists Visualize the Vote Counting Process in U.S. ElectionsabstractElection results in the United States are visualized online in real time by news outlets as vote counting persists over days or weeks. They are a massive public-facing exercise in managing audience understanding of uncertainty in partial data, breaking news web traffic records as the public seeks information about winners. We categorize designs of real-time election results from 19 U.S. news outlets and election results providers for the 2020 and 2022 general elections to create a visual vocabulary of live results. We then use this vocabulary to guide interviews with data journalists who worked on these designs to understand their design goals and challenges. Tying these conversations back to our visual vocabulary, we map out how communication goals like balancing certainty and uncertainty in the journey towards finding out winners, alongside challenges like determining thresholds at which information is shown, manifest in the designs displayed. Mandi Cai, Matthew Kay 0001 |
CHI | 2 |
| 2024 | V-FRAMER: Visualization Framework for Mitigating Reasoning Errors in Public PolicyabstractExisting data visualization design guidelines focus primarily on constructing grammatically-correct visualizations that faithfully convey the values and relationships in the underlying data. However, a designer may create a grammatically-correct visualization that still leaves audiences susceptible to reasoning misleaders, e.g. by failing to normalize data or using unrepresentative samples. Reasoning misleaders are especially pernicious when presenting public policy data, where data-driven decisions can affect public health, safety, and economic development. Through textual analysis, a formative evaluation, and iterative design with 19 policy communicators, we construct an actionable visualization design framework, V-FRAMER, that effectively synthesizes ways of mitigating reasoning misleaders. We discuss important design considerations for frameworks like V-FRAMER, including using concrete examples to help designers understand reasoning misleaders, and using a hierarchical structure to support example-based accessing. We further describe V-FRAMER’s congruence with current practice and how practitioners might integrate the framework into their existing workflows. Related materials available at: https://osf.io/q3uta/. Lily W. Ge, Matthew W. Easterday, Matthew Kay 0001, Evanthia Dimara, Peter C.-H. Cheng, Steven Franconeri |
CHI | 3 |
| 2024 | Authors' Values and Attitudes Towards AI-bridged Scalable Personalization of Creative Language ArtsabstractGenerative AI has the potential to create a new form of interactive media: AI-bridged creative language arts (CLA), which bridge the author and audience by personalizing the author’s vision to the audience’s context and taste at scale. However, it is unclear what the authors’ values and attitudes would be regarding AI-bridged CLA. To identify these values and attitudes, we conducted an interview study with 18 authors across eight genres (e.g., poetry, comics) by presenting speculative but realistic AI-bridged CLA scenarios. We identified three benefits derived from the dynamics between author, artifact, and audience: those that 1) authors get from the process, 2) audiences get from the artifact, and 3) authors get from the audience. We found how AI-bridged CLA would either promote or reduce these benefits, along with authors’ concerns. We hope our investigation hints at how AI can provide intriguing experiences to CLA audiences while promoting authors’ values. Taewook Kim 0001, Hyomin Han, Eytan Adar, Matthew Kay 0001, John Joon Young Chung |
CHI | 4 |
| 2024 | To Cut or Not To Cut? A Systematic Exploration of Y-Axis TruncationabstractY-axis truncation is a well-known, much-debated visualization practice. Our work complements existing empirical work by providing a systematic analysis of y-axis truncation on grouped bar charts. Drawing upon theoretical frameworks such as Algebraic Visualization Design, we examine how structure-preserving modifications to visualization affect user performance by systematically dividing the space of possible truncations according to their monotonicity and the type of relations in the underlying data. Our results demonstrate that for comparing and estimating the difference between the lengths of two bars, truncating the y-axis does not affect task performance. For comparing or estimating the relative growth between two bars, truncating monotonically has similar performance to no truncation, while truncating non-monotonically is very likely to impair performance. We discuss possible extensions of our work and recommendations for y-axis truncation. All supplementary materials are available at https://osf.io/k4hjd/?view_only=008b087fc3d94be7ba0ce7aea95012a7. Sheng Long 0001, Matthew Kay 0001 |
CHI | 2 |
| 2024 | Milliways: Taming Multiverses through Principled Evaluation of Data Analysis PathsabstractMultiverse analyses involve conducting all combinations of reasonable choices in a data analysis process. A reader of a study containing a multiverse analysis might question—are all the choices included in the multiverse reasonable and equally justifiable? How much do results vary if we make different choices in the analysis process? In this work, we identify principles for validating the composition of, and interpreting the uncertainty in, the results of a multiverse analysis. We present Milliways, a novel interactive visualisation system to support principled evaluation of multiverse analyses. Milliways provides interlinked panels presenting result distributions, individual analysis composition, multiverse code specification, and data summaries. Milliways supports interactions to sort, filter and aggregate results based on the analysis specification to identify decisions in the analysis process to which the results are sensitive. To represent the two qualitatively different types of uncertainty that arise in multiverse analyses—probabilistic uncertainty from estimating unknown quantities of interest such as regression coefficients, and possibilistic uncertainty from choices in the data analysis—Milliways uses consonance curves and probability boxes. Through an evaluative study with five users familiar with multiverse analysis, we demonstrate how Milliways can support multiverse analysis tasks, including a principled assessment of the results of a multiverse analysis. Abhraneel Sarma, Kyle Hwang, Jessica Hullman, Matthew Kay 0001 |
CHI | 4 |
| 2024 | Odds and Insights: Decision Quality in Exploratory Data Analysis Under UncertaintyabstractRecent studies have shown that users of visual analytics tools can have difficulty distinguishing robust findings in the data from statistical noise, but the true extent of this problem is likely dependent on both the incentive structure motivating their decisions, and the ways that uncertainty and variability are (or are not) represented in visualisations. In this work, we perform a crowd-sourced study measuring decision-making quality in visual analytics, testing both an explicit structure of incentives designed to reward cautious decision-making as well as a variety of designs for communicating uncertainty. We find that, while participants are unable to perfectly control for false discoveries as well as idealised statistical models such as the Benjamini-Hochberg, certain forms of uncertainty visualisations can improve the quality of participants’ decisions and lead to fewer false discoveries than not correcting for multiple comparisons. We conclude with a call for researchers to further explore visual analytics decision quality under different decision-making contexts, and for designers to directly present uncertainty and reliability information to users of visual analytics tools. The supplementary materials are available at: https://osf.io/xtsfz/. Abhraneel Sarma, Xiaoying Pu, Michael Correll, Eli T. Brown, Matthew Kay 0001 |
CHI | 6 |
| 2024 | In Dice We Trust: Uncertainty Displays for Maintaining Trust in Election Forecasts Over TimeabstractTrust in high-profile election forecasts influences the public’s confidence in democratic processes and electoral integrity. Yet, maintaining trust after unexpected outcomes like the 2016 U.S. presidential election is a significant challenge. Our work confronts this challenge through three experiments that gauge trust in election forecasts. We generate simulated U.S. presidential election forecasts, vary win probabilities and outcomes, and present them to participants in a professional-looking website interface. In this website interface, we explore (1) four different uncertainty displays, (2) a technique for subjective probability correction, and (3) visual calibration that depicts an outcome with its forecast distribution. Our quantitative results suggest that text summaries and quantile dotplots engender the highest trust over time, with observable partisan differences. The probability correction and calibration show small-to-null effects on average. Complemented by our qualitative results, we provide design recommendations for conveying U.S. presidential election forecasts and discuss long-term trust in uncertainty communication. We provide preregistration, code, data, model files, and videos at https://doi.org/10.17605/OSF.IO/923E7. Fumeng Yang, Chloe Mortenson, Erik C. Nisbet, Nicholas Diakopoulos, Matthew Kay 0001 |
CHI | 5 |
| 2024 | Opportunities in Mental Health Support for Informal Dementia Caregivers Suffering from Verbal AgitationabstractPeople with dementia (PwD) often present verbal agitation such as cursing, screaming, and persistently complaining. Verbal agitation can impose mental distress on informal caregivers (e.g., family, friends), which may cause severe mental illnesses, such as depression and anxiety disorders. To improve informal caregivers' mental health, we explore design opportunities by interviewing 11 informal caregivers suffering from verbal agitation of PwD. In particular, we first characterize how the predictability of verbal agitation impacts informal caregivers' mental health and how caregivers' coping strategies vary before, during, and after verbal agitation. Based on our findings, we propose design opportunities to improve the mental health of informal caregivers suffering from verbal agitation: distracting PwD (in-situ support; before), prompting just-in-time maneuvers (information support; during), and comfort and education (social & information support; after). We discuss our reflections on cultural disparities between participants. Our work envisions a broader design space for supporting informal caregivers' well-being and describes when and how that support could be provided. Taewook Kim 0001, Hyeok Kim, Angela Roberts 0001, Maia L. Jacobs, Matthew Kay 0001 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Adaptive Assessment of Visualization LiteracyabstractVisualization literacy is an essential skill for accurately interpreting data to inform critical decisions. Consequently, it is vital to understand the evolution of this ability and devise targeted interventions to enhance it, requiring concise and repeatable assessments of visualization literacy for individuals. However, current assessments, such as the Visualization Literacy Assessment Test (VLAT), are time-consuming due to their fixed, lengthy format. To address this limitation, we develop two streamlined computerized adaptive tests (CATs) for visualization literacy, A-VLAT and A-CALVI, which measure the same set of skills as their original versions in half the number of questions. Specifically, we (1) employ item response theory (IRT) and non-psychometric constraints to construct adaptive versions of the assessments, (2) finalize the configurations of adaptation through simulation, (3) refine the composition of test items of A-CALVI via a qualitative study, and (4) demonstrate the test-retest reliability (ICC: 0.98 and 0.98) and convergent validity (correlation: 0.81 and 0.66) of both CATs via four online studies. We discuss practical recommendations for using our CATs and opportunities for further customization to leverage the full potential of adaptive assessments. All supplemental materials are available at https://osf.io/a6258/. Lily W. Ge, Yiren Ding, Fumeng Yang, Lane Harrison, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | The Risks of Ranking: Revisiting Graphical Perception to Model Individual Differences in Visualization PerformanceabstractGraphical perception studies typically measure visualization encoding effectiveness using the error of an "average observer", leading to canonical rankings of encodings for numerical attributes: e.g., position area angle volume. Yet different people may vary in their ability to read different visualization types, leading to variance in this ranking across individuals not captured by population-level metrics using "average observer" models. One way we can bridge this gap is by recasting classic visual perception tasks as tools for assessing individual performance, in addition to overall visualization performance. In this article we replicate and extend Cleveland and McGill's graphical comparison experiment using Bayesian multilevel regression, using these models to explore individual differences in visualization skill from multiple perspectives. The results from experiments and modeling indicate that some people show patterns of accuracy that credibly deviate from the canonical rankings of visualization effectiveness. We discuss implications of these findings, such as a need for new ways to communicate visualization effectiveness to designers, how patterns in individuals' responses may show systematic biases and strategies in visualization judgment, and how recasting classic visual perception tasks as tools for assessing individual performance may offer new ways to quantify aspects of visualization literacy. Experiment data, source code, and analysis scripts are available at the following repository: https://osf.io/8ub7t/?view_only=9be4798797404a4397be3c6fc2a68cc0. Russell Davis, Xiaoying Pu, Yiren Ding, Brian D. Hall, Karen Bonilla, Mi Feng, Matthew Kay 0001, Lane Harrison |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | ggdist: Visualizations of Distributions and Uncertainty in the Grammar of GraphicsabstractThe grammar of graphics is ubiquitous, providing the foundation for a variety of popular visualization tools and toolkits. Yet support for uncertainty visualization in the grammar graphics-beyond simple variations of error bars, uncertainty bands, and density plots-remains rudimentary. Research in uncertainty visualization has developed a rich variety of improved uncertainty visualizations, most of which are difficult to create in existing grammar of graphics implementations. ggdist, an extension to the popular ggplot2 grammar of graphics toolkit, is an attempt to rectify this situation. ggdist unifies a variety of uncertainty visualization types through the lens of distributional visualization, allowing functions of distributions to be mapped to directly to visual channels (aesthetics), making it straightforward to express a variety of (sometimes weird!) uncertainty visualization types. This distributional lens also offers a way to unify Bayesian and frequentist uncertainty visualization by formalizing the latter with the help of confidence distributions. In this paper, I offer a description of this uncertainty visualization paradigm and lessons learned from its development and adoption: ggdist has existed in some form for about six years (originally as part of the tidybayes R package for post-processing Bayesian models), and it has evolved substantially over that time, with several rewrites and API re-organizations as it changed in response to user feedback and expanded to cover increasing varieties of uncertainty visualization types. Ultimately, given the huge expressive power of the grammar of graphics and the popularity of tools built on it, I hope a catalog of my experience with ggdist will provide a catalyst for further improvements to formalizations and implementations of uncertainty visualization in grammar of graphics ecosystems. A free copy of this paper is available at https://osf.io/2gsz6. All supplemental materials are available at https://github.com/mjskay/ggdist-paper and are archived on Zenodo at doi:10.5281/zenodo.7770984. Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Swaying the Public? Impacts of Election Forecast Visualizations on Emotion, Trust, and Intention in the 2022 U.S. MidtermsabstractWe conducted a longitudinal study during the 2022 U.S. midterm elections, investigating the real-world impacts of uncertainty visualizations. Using our forecast model of the governor elections in 33 states, we created a website and deployed four uncertainty visualizations for the election forecasts: single quantile dotplot (1-Dotplot), dual quantile dotplots (2-Dotplot), dual histogram intervals (2-Interval), and Plinko quantile dotplot (Plinko), an animated design with a physical and probabilistic analogy. Our online experiment ran from Oct. 18, 2022, to Nov. 23, 2022, involving 1,327 participants from 15 states. We use Bayesian multilevel modeling and post-stratification to produce demographically-representative estimates of people's emotions, trust in forecasts, and political participation intention. We find that election forecast visualizations can heighten emotions, increase trust, and slightly affect people's intentions to participate in elections. 2-Interval shows the strongest effects across all measures; 1-Dotplot increases trust the most after elections. Both visualizations create emotional and trust gaps between different partisan identities, especially when a Republican candidate is predicted to win. Our qualitative analysis uncovers the complex political and social contexts of election forecast visualizations, showcasing that visualizations may provoke polarization. This intriguing interplay between visualization types, partisanship, and trust exemplifies the fundamental challenge of disentangling visualization from its context, underscoring a need for deeper investigation into the real-world impacts of visualizations. Our preprint and supplements are available at https://doi.org/osf.io/ajq8f. Fumeng Yang, Mandi Cai, Chloe Mortenson, Hoda Fakhari, Ayse D. Lokmanoglu, Jessica Hullman, Steven Franconeri, Nicholas Diakopoulos, Erik C. Nisbet, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2023 | CALVI: Critical Thinking Assessment for Literacy in VisualizationsabstractVisualization misinformation is a prevalent problem, and combating it requires understanding people’s ability to read, interpret, and reason about erroneous or potentially misleading visualizations, which lacks a reliable measurement: existing visualization literacy tests focus on well-formed visualizations. We systematically develop an assessment for this ability by: (1) developing a precise definition of misleaders (decisions made in the construction of visualizations that can lead to conclusions not supported by the data), (2) constructing initial test items using a design space of misleaders and chart types, (3) trying out the provisional test on 497 participants, and (4) analyzing the test tryout results and refining the items using Item Response Theory, qualitative analysis, a wrong-due-to-misleader score, and the content validity index. Our final bank of 45 items shows high reliability, and we provide item bank usage recommendations for future tests and different use cases. Related materials are available at: https://osf.io/pv67z/. Lily W. Ge, Matthew Kay 0001 |
CHI | 3 |
| 2023 | How Data Analysts Use a Visualization Grammar in PracticeabstractVisualization grammars, often based on the Grammar of Graphics (GoG), have much potential for augmenting data analysis in a programming environment. However, we do not know how analysts conceptualize grammar abstractions, or how a visualization grammar works with data analysis in practice. Therefore, we qualitatively analyzed how experienced analysts (N = 6) from TidyTuesday, a social data project, wrangled and visualized data using GoG-based ggplot2 without given tasks in R Markdown. Though participants’ analysis and customization needs could mismatch with GoG component design, their analysis processes aligned with the goal of GoG to expedite visualization iteration. We also found a feedback loop and tight coupling between visualization and data transformation code, explaining both participants’ productivity and their errors. From these results, we discuss how future visualization grammars can become more practical for analysts and how visualization grammar and analysis tools can better integrate within a programming (i.e., computational notebook) environment. Xiaoying Pu, Matthew Kay 0001 |
CHI | 2 |
| 2023 | "It can bring you in the right direction": Episode-Driven Data Narratives to Help Patients Navigate Multidimensional Diabetes Data to Make Care DecisionsabstractEngaging with multiple streams of personal health data to inform self-care of chronic health conditions remains a challenge. Existing informatics tools provide limited support for patients to make data actionable. To design better tools, we conducted two studies with Type 1 diabetes patients and their clinicians. In the first study, we observed data review sessions between patients and clinicians to articulate the tasks involved in assessing different types of data from diabetes devices to make care decisions. Drawing upon these tasks, we designed novel data interfaces called episode-driven data narratives and performed a task-driven evaluation. We found that as compared to the commercially available diabetes data reports, episode-driven data narratives improved engagement and decision-making with data. We discuss implications for designing data interfaces to support interaction with multidimensional health data to inform self-care. Shriti Raj, Toshi Gupta, Joyce M. Lee, Matthew Kay 0001, Mark W. Newman |
CHI | 4 |
| 2023 | multiverse: Multiplexing Alternative Data Analyses in R NotebooksabstractThere are myriad ways to analyse a dataset. But which one to trust? In the face of such uncertainty, analysts may adopt multiverse analysis: running all reasonable analyses on the dataset. Yet this is cognitively and technically difficult with existing tools—how does one specify and execute all combinations of reasonable analyses of a dataset?—and often requires discarding existing workflows. We present multiverse, a tool for implementing multiverse analyses in R with expressive syntax supporting existing computational notebook workflows. multiverse supports building up a multiverse through local changes to a single analysis and optimises execution by pruning redundant computations. We evaluate how multiverse supports programming multiverse analyses using (a) principles of cognitive ergonomics to compare with two existing multiverse tools; and (b) case studies based on semi-structured interviews with researchers who have successfully implemented an end-to-end analysis using multiverse. We identify design tradeoffs (e.g. increased flexibility versus learnability), and suggest future directions for multiverse tool design. Abhraneel Sarma, Alex Kale, Michael Jongho Moon, Nathan Taback, Fanny Chevalier, Jessica Hullman, Matthew Kay 0001 |
CHI | 7 |
| 2023 | Subjective Probability Correction for Uncertainty RepresentationsabstractWe propose a new approach to uncertainty communication: we keep the uncertainty representation fixed, but adjust the distribution displayed to compensate for biases in people’s subjective probability in decision-making. To do so, we adopt a linear-in-probit model of subjective probability and derive two corrections to a Normal distribution based on the model’s intercept and slope: one correcting all right-tailed probabilities, and the other preserving the mode and one focal probability. We then conduct two experiments on U.S. demographically-representative samples. We show participants hypothetical U.S. Senate election forecasts as text or a histogram and elicit their subjective probabilities using a betting task. The first experiment estimates the linear-in-probit intercepts and slopes, and confirms the biases in participants’ subjective probabilities. The second, preregistered follow-up shows participants the bias-corrected forecast distributions. We find the corrections substantially improve participants’ decision quality by reducing the integrated absolute error of their subjective probabilities compared to the true probabilities. These corrections can be generalized to any univariate probability or confidence distribution, giving them broad applicability. Our preprint, code, data, and preregistration are available at https://doi.org/10.17605/osf.io/kcwxm Fumeng Yang, Maryam Hedayati, Matthew Kay 0001 |
CHI | 3 |
| 2023 | Evaluating the Use of Uncertainty Visualisations for Imputations of Data Missing At Random in ScatterplotsabstractMost real-world datasets contain missing values yet most exploratory data analysis (EDA) systems only support visualising data points with complete cases. This omission may potentially lead the user to biased analyses and insights. Imputation techniques can help estimate the value of a missing data point, but introduces additional uncertainty. In this work, we investigate the effects of visualising imputed values in charts using different ways of representing data imputations and imputation uncertainty-no imputation, mean, 95% confidence intervals, probability density plots, gradient intervals, and hypothetical outcome plots. We focus on scatterplots, which is a commonly used chart type, and conduct a crowdsourced study with 202 participants. We measure users' bias and precision in performing two tasks-estimating average and detecting trend-and their self-reported confidence in performing these tasks. Our results suggest that, when estimating averages, uncertainty representations may reduce bias but at the cost of decreasing precision. When estimating trend, only hypothetical outcome plots may lead to a small probability of reducing bias while increasing precision. Participants in every uncertainty representation were less certain about their response when compared to the baseline. The findings point towards potential trade-offs in using uncertainty encodings for datasets with a large number of missing values. This paper and the associated analysis materials are available at: https://osf.io/q4y5r/. Abhraneel Sarma, Shunan Guo, Jane Hoffswell, Ryan Rossi, Fan Du, Eunyee Koh, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | A Survey of Tasks and Visualizations in Multiverse Analysis ReportsabstractAbstract Analysing data from experiments is a complex, multi‐step process, often with multiple defensible choices available at each step. While analysts often report a single analysis without documenting how it was chosen, this can cause serious transparency and methodological issues. To make the sensitivity of analysis results to analytical choices transparent, some statisticians and methodologists advocate the use of ‘multiverse analysis’: reporting the full range of outcomes that result from all combinations of defensible analytic choices. Summarizing this combinatorial explosion of statistical results presents unique challenges; several approaches to visualizing the output of multiverse analyses have been proposed across a variety of fields (e.g. psychology, statistics, economics, neuroscience). In this article, we (1) introduce a consistent conceptual framework and terminology for multiverse analyses that can be applied across fields; (2) identify the tasks researchers try to accomplish when visualizing multiverse analyses and (3) classify multiverse visualizations into ‘archetypes’, assessing how well each archetype supports each task. Our work sets a foundation for subsequent research on developing visualization tools and techniques to support multiverse analysis and its reporting. Brian D. Hall, Yvonne Jansen, Pierre Dragicevic, Fanny Chevalier, Matthew Kay 0001 |
Comput. Graph. Forum | 6 |
| 2021 | An Aligned Rank Transform Procedure for Multifactor Contrast TestsabstractData from multifactor HCI experiments often violates the assumptions of parametric tests (i.e., nonconforming data). The Aligned Rank Transform (ART) has become a popular nonparametric analysis in HCI that can find main and interaction effects in nonconforming data, but leads to incorrect results when used to conduct post hoc contrast tests. We created a new algorithm called ART-C for conducting contrast tests within the ART paradigm and validated it on 72,000 synthetic data sets. Our results indicate that ART-C does not inflate Type I error rates, unlike contrasts based on ART, and that ART-C has more statistical power than a t-test, Mann-Whitney U test, Wilcoxon signed-rank test, and ART. We also extended an open-source tool called ARTool with our ART-C algorithm for both Windows and R. Our validation had some limitations (e.g., only six distribution types, no mixed factorial designs, no random slopes), and data drawn from Cauchy distributions should not be analyzed with ART-C. Lisa A. Elkin, Matthew Kay 0001, James J. Higgins, Jacob O. Wobbrock |
UIST | 2 |
| 2021 | Visual Reasoning Strategies for Effect Size Judgments and DecisionsabstractUncertainty visualizations often emphasize point estimates to support magnitude estimates or decisions through visual comparison. However, when design choices emphasize means, users may overlook uncertainty information and misinterpret visual distance as a proxy for effect size. We present findings from a mixed design experiment on Mechanical Turk which tests eight uncertainty visualization designs: 95% containment intervals, hypothetical outcome plots, densities, and quantile dotplots, each with and without means added. We find that adding means to uncertainty visualizations has small biasing effects on both magnitude estimation and decision-making, consistent with discounting uncertainty. We also see that visualization designs that support the least biased effect size estimation do not support the best decision-making, suggesting that a chart user's sense of effect size may not necessarily be identical when they use the same information for different tasks. In a qualitative analysis of users' strategy descriptions, we find that many users switch strategies and do not employ an optimal strategy when one exists. Uncertainty visualizations which are optimally designed in theory may not be the most effective in practice because of the ways that users satisfice with heuristics, suggesting opportunities to better understand visualization effectiveness by modeling sets of potential strategies. Alex Kale, Matthew Kay 0001, Jessica Hullman |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Revealing Perceptual Proxies with Adversarial ExamplesabstractData visualizations convert numbers into visual marks so that our visual system can extract data from an image instead of raw numbers. Clearly, the visual system does not compute these values as a computer would, as an arithmetic mean or a correlation. Instead, it extracts these patterns using perceptual proxies; heuristic shortcuts of the visual marks, such as a center of mass or a shape envelope. Understanding which proxies people use would lead to more effective visualizations. We present the results of a series of crowdsourced experiments that measure how powerfully a set of candidate proxies can explain human performance when comparing the mean and range of pairs of data series presented as bar charts. We generated datasets where the correct answer-the series with the larger arithmetic mean or range-was pitted against an "adversarial" series that should be seen as larger if the viewer uses a particular candidate proxy. We used both Bayesian logistic regression models and a robust Bayesian mixed-effects linear model to measure how strongly each adversarial proxy could drive viewers to answer incorrectly and whether different individuals may use different proxies. Finally, we attempt to construct adversarial datasets from scratch, using an iterative crowdsourcing procedure to perform black-box optimization. Brian D. Ondov, Fumeng Yang, Matthew Kay 0001, Niklas Elmqvist, Steven Franconeri |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | A Probabilistic Grammar of GraphicsabstractVisualizations depicting probabilities and uncertainty are used everywhere from medical risk communication to machine learning, yet these probabilistic visualizations are difficult to specify, prone to error, and their designs are cumbersome to explore. We propose a Probabilistic Grammar of Graphics (PGoG), an extension to Wilkinson's original framework. Inspired by the success of probabilistic programming languages, PGoG makes probability expressions, such as P(A|B), a first-class citizen in the language. PGoG abstractions also reflect the distinction between probability and frequency framing, a concept from the uncertainty communication literature. It is expressive, encompassing product plots, density plots, icon arrays, and dotplots, among other visualizations. Its coherent syntax ensures correctness (that the proportions of visual elements and their spatial placement reflect the underlying probability distribution) and reduces edit distance between probabilistic visualization specifications, potentially supporting more design exploration. We provide a proof-of-concept implementation of PGoG in R. Xiaoying Pu, Matthew Kay 0001 |
CHI | 2 |
| 2020 | Prior Setting in Practice: Strategies and Rationales Used in Choosing Prior Distributions for Bayesian AnalysisabstractBayesian statistical analysis is steadily growing in popularity and use. Choosing priors is an integral part of Bayesian inference. While there exist extensive normative recommendations for prior setting, little is known about how priors are chosen in practice. We conducted a survey (N = 50) and interviews (N = 9) where we used interactive visualizations to elicit prior distributions from researchers experienced withBayesian statistics and asked them for rationales for those priors. We found that participants' experience and philosophy influence how much and what information they are willing to incorporate into their priors, manifesting as different levels of informativeness and skepticism. We also identified three broad strategies participants use to set their priors: centrality matching, interval matching, and visual mass allocation. We discovered that participants' understanding of the notion of 'weakly informative priors"-a commonly-recommended normative approach to prior setting-manifests very differently across participants. Our results have implications both for how to develop prior setting recommendations and how to design tools to elicit priors in Bayesian analysis. Abhraneel Sarma, Matthew Kay 0001 |
CHI | 2 |
| 2020 | How patterns of students dashboard use are related to their achievement and self-regulatory engagementabstractThe aim of student-facing dashboards is to support learning by providing students with actionable information and promoting self-regulated learning. We created a new dashboard design aligned with SRL theory, called MyLA, to better understand how students use a learning analytics tool. We conducted sequence analysis on students' interactions with three different visualizations in the dashboard, implemented in a LMS, for a large number of students (860) in ten courses representing different disciplines. To evaluate different students' experiences with the dashboard, we computed chi-squared tests of independence on dashboard users (52%) to find frequent patterns that discriminate students by their differences in academic achievement and self-regulated learning behaviors. The results revealed discriminating patterns in dashboard use among different levels of academic achievement and self-regulated learning, particularly for low achieving students and high self-regulated learners. Our findings highlight the importance of differences in students' experience with a student-facing dashboard, and emphasize that one size does not fit all in the design of learning analytics tools. Fatemeh Salehian Kia, Stephanie D. Teasley, Marek Hatala, Stuart A. Karabenick, Matthew Kay 0001 |
LAK | 5 |
| 2019 | Increasing the Transparency of Research Papers with Explorable Multiverse AnalysesabstractWe present explorable multiverse analysis reports, a new approach to statistical reporting where readers of research papers can explore alternative analysis options by interacting with the paper itself. This approach draws from two recent ideas: i) multiverse analysis, a philosophy of statistical reporting where paper authors report the outcomes of many different statistical analyses in order to show how fragile or robust their findings are; and ii) explorable explanations, narratives that can be read as normal explanations but where the reader can also become active by dynamically changing some elements of the explanation. Based on five examples and a design space analysis, we show how combining those two ideas can complement existing reporting approaches and constitute a step towards more transparent research papers. Pierre Dragicevic, Yvonne Jansen, Abhraneel Sarma, Matthew Kay 0001, Fanny Chevalier |
CHI | 4 |
| 2019 | Decision-Making Under Uncertainty in Research Synthesis: Designing for the Garden of Forking PathsabstractTo make evidence-based recommendations to decision-makers, researchers conducting systematic reviews and meta-analyses must navigate a garden of forking paths: a series of analytical decision-points, each of which has the potential to influence findings. To identify challenges and opportunities related to designing systems to help researchers manage uncertainty around which of multiple analyses is best, we interviewed 11 professional researchers who conduct research synthesis to inform decision-making within three organizations. We conducted a qualitative analysis identifying 480 analytical decisions made by researchers throughout the scientific process. We present descriptions of current practices in applied research synthesis and corresponding design challenges: making it more feasible for researchers to try and compare analyses, shifting researchers' attention from rationales for decisions to impacts on results, and supporting communication techniques that acknowledge decision-makers' aversions to uncertainty. We identify opportunities to design systems which help researchers explore, reason about, and communicate uncertainty in decision-making about possible analyses in research synthesis. Alex Kale, Matthew Kay 0001, Jessica Hullman |
CHI | 2 |
| 2019 | Some Prior(s) Experience Necessary: Templates for Getting Started With Bayesian AnalysisabstractBayesian statistical analysis has gained attention in recent years, including in HCI. The Bayesian approach has several advantages over traditional statistics, including producing results with more intuitive interpretations. Despite growing interest, few papers in CHI use Bayesian analysis. Existing tools to learn Bayesian statistics require significant time investment, making it difficult to casually explore Bayesian methods. Here, we present a tool that lowers the barrier to exploration: a set of R code templates that guide Bayesian novices through their first analysis. The templates are tailored to CHI, supporting analyses found to be most common in recent CHI papers. In a user study, we found that the templates were easy to understand and use. However, we found that participants without a statistical background were not confident in their use. Together our contributions provide a concise analysis tool and empirical results for understanding and addressing barriers to using Bayesian analysis in HCI. Chanda Phelan, Jessica Hullman, Matthew Kay 0001, Paul Resnick |
CHI | 3 |
| 2019 | In Pursuit of Error: A Survey of Uncertainty Visualization EvaluationabstractUnderstanding and accounting for uncertainty is critical to effectively reasoning about visualized data. However, evaluating the impact of an uncertainty visualization is complex due to the difficulties that people have interpreting uncertainty and the challenge of defining correct behavior with uncertainty information. Currently, evaluators of uncertainty visualization must rely on general purpose visualization evaluation frameworks which can be ill-equipped to provide guidance with the unique difficulties of assessing judgments under uncertainty. To help evaluators navigate these complexities, we present a taxonomy for characterizing decisions made in designing an evaluation of an uncertainty visualization. Our taxonomy differentiates six levels of decisions that comprise an uncertainty visualization evaluation: the behavioral targets of the study, expected effects from an uncertainty visualization, evaluation goals, measures, elicitation techniques, and analysis approaches. Applying our taxonomy to 86 user studies of uncertainty visualizations, we find that existing evaluation practice, particularly in visualization research, focuses on Performance and Satisfaction-based measures that assume more predictable and statistically-driven judgment behavior than is suggested by research on human judgment and decision making. We reflect on common themes in evaluation practice concerning the interpretation and semantics of uncertainty, the use of confidence reporting, and a bias toward evaluating performance as accuracy rather than decision quality. We conclude with a concrete set of recommendations for evaluators designed to reduce the mismatch between the conceptualization of uncertainty in visualization versus other fields. Jessica Hullman, Xiaoli Qiao, Michael Correll, Alex Kale, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2019 | Hypothetical Outcome Plots Help Untrained Observers Judge Trends in Ambiguous DataabstractAnimated representations of outcomes drawn from distributions (hypothetical outcome plots, or HOPs) are used in the media and other public venues to communicate uncertainty. HOPs greatly improve multivariate probability estimation over conventional static uncertainty visualizations and leverage the ability of the visual system to quickly, accurately, and automatically process the summary statistical properties of ensembles. However, it is unclear how well HOPs support applied tasks resembling real world judgments posed in uncertainty communication. We identify and motivate an appropriate task to investigate realistic judgments of uncertainty in the public domain through a qualitative analysis of uncertainty visualizations in the news. We contribute two crowdsourced experiments comparing the effectiveness of HOPs, error bars, and line ensembles for supporting perceptual decision-making from visualized uncertainty. Participants infer which of two possible underlying trends is more likely to have produced a sample of time series data by referencing uncertainty visualizations which depict the two trends with variability due to sampling error. By modeling each participant's accuracy as a function of the level of evidence presented over many repeated judgments, we find that observers are able to correctly infer the underlying trend in samples conveying a lower level of evidence when using HOPs rather than static aggregate uncertainty visualizations as a decision aid. Modeling approaches like ours contribute theoretically grounded and richly descriptive accounts of user perceptions to visualization evaluation. Alex Kale, Francis Nguyen, Matthew Kay 0001, Jessica Hullman |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Uncertainty Displays Using Quantile Dotplots or CDFs Improve Transit Decision-MakingabstractEveryday predictive systems typically present point predictions, making it hard for people to account for uncertainty when making decisions. Evaluations of uncertainty displays for transit prediction have assessed people's ability to extract probabilities, but not the quality of their decisions. In a controlled, incentivized experiment, we had subjects decide when to catch a bus using displays with textual uncertainty, uncertainty visualizations, or no-uncertainty (control). Frequency-based visualizations previously shown to allow people to better extract probabilities (quantile dotplots) yielded better decisions. Decisions with quantile dotplots with 50 outcomes were(1) better on average, having expected payoffs 97% of optimal(95% CI: [95%,98%]), 5 percentage points more than control (95% CI: [2,8]); and (2) more consistent, having within-subject standard deviation of 3 percentage points (95% CI:[2,4]), 4 percentage points less than control (95% CI: [2,6]).Cumulative distribution function plots performed nearly as well, and both outperformed textual uncertainty, which was sensitive to the probability interval communicated. We discuss implications for real time transit predictions and possible generalization to other domains. Michael Fernandes, Logan Walls, Sean A. Munson, Jessica Hullman, Matthew Kay 0001 |
CHI | 5 |
| 2018 | Imagining Replications: Graphical Prediction & Discrete Visualizations Improve Recall & Estimation of Effect UncertaintyabstractPeople often have erroneous intuitions about the results of uncertain processes, such as scientific experiments. Many uncertainty visualizations assume considerable statistical knowledge, but have been shown to prompt erroneous conclusions even when users possess this knowledge. Active learning approaches been shown to improve statistical reasoning, but are rarely applied in visualizing uncertainty in scientific reports. We present a controlled study to evaluate the impact of an interactive, graphical uncertainty prediction technique for communicating uncertainty in experiment results. Using our technique, users sketch their prediction of the uncertainty in experimental effects prior to viewing the true sampling distribution from an experiment. We find that having a user graphically predict the possible effects from experiment replications is an effective way to improve one's ability to make predictions about replications of new experiments. Additionally, visualizing uncertainty as a set of discrete outcomes, as opposed to a continuous probability distribution, can improve recall of a sampling distribution from a single experiment. Our work has implications for various applications where it is important to elicit peoples' estimates of probability distributions and to communicate uncertainty effectively. Jessica Hullman, Matthew Kay 0001, Yea-Seul Kim, Samana Shrestha |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Self-Experimentation for Behavior Change: Design and Formative Evaluation of Two ApproachesabstractDesirable outcomes such as health are tightly linked to behaviors, thus inspiring research on technologies that support people in changing those behaviors. Many behavior-change technologies are designed by HCI experts but this approach can make it difficult to personalize support to each user's unique goals and needs. This paper reports on the iterative design of two complementary support strategies for helping users create their own personalized behavior-change plans via self-experimentation: One emphasized the use of interactive instructional materials, and the other additionally introduced context-aware computing to enable user creation of "just in time" home-based interventions. In a formative trial with 27 users, we compared these two approaches to an unstructured sleep education control. Results suggest great promise in both strategies and provide insights on how to develop personalized behavior-change technologies. Jisoo Lee, Erin Walker, Winslow Burleson, Matthew Kay 0001, Matthew P. Buman, Eric B. Hekler |
CHI | 4 |
| 2016 | When (ish) is My Bus?: User-centered Visualizations of Uncertainty in Everyday, Mobile Predictive SystemsabstractUsers often rely on realtime predictions in everyday contexts like riding the bus, but may not grasp that such predictions are subject to uncertainty. Existing uncertainty visualizations may not align with user needs or how they naturally reason about probability. We present a novel mobile interface design and visualization of uncertainty for transit predictions on mobile phones based on discrete outcomes. To develop it, we identified domain specific design requirements for visualizing uncertainty in transit prediction through: 1) a literature review, 2) a large survey of users of a popular realtime transit application, and 3) an iterative design process. We present several candidate visualizations of uncertainty for realtime transit predictions in a mobile context, and we propose a novel discrete representation of continuous outcomes designed for small screens, quantile dotplots. In a controlled experiment we find that quantile dotplots reduce the variance of probabilistic estimates by ~1.15 times compared to density plots and facilitate more confident estimation by end-users in the context of realtime transit prediction scenarios. Matthew Kay 0001, Tara Kola, Jessica Hullman, Sean A. Munson |
CHI | 1 |
| 2016 | Researcher-Centered Design of Statistics: Why Bayesian Statistics Better Fit the Culture and Incentives of HCIabstractA core tradition of HCI lies in the experimental evaluation of the effects of techniques and interfaces to determine if they are useful for achieving their purpose. However, our individual analyses tend to stand alone, and study results rarely accrue in more precise estimates via meta-analysis: in a literature search, we found only 56 meta-analyses in HCI in the ACM Digital Library, 3 of which were published at CHI (often called the top HCI venue). Yet meta-analysis is the gold standard for demonstrating robust quantitative knowledge. We treat this as a user-centered design problem: the failure to accrue quantitative knowledge is not the users' (i.e. researchers') failure, but a failure to consider those users' needs when designing statistical practice. Using simulation, we compare hypothetical publication worlds following existing frequentist against Bayesian practice. We show that Bayesian analysis yields more precise effects with each new study, facilitating knowledge accrual without traditional meta-analyses. Bayesian practices also allow more principled conclusions from small-n studies of novel techniques. These advantages make Bayesian practices a likely better fit for the culture and incentives of the field. Instead of admonishing ourselves to spend resources on larger studies, we propose using tools that more appropriately analyze small studies and encourage knowledge accrual from one study to the next. We also believe Bayesian methods can be adopted from the bottom up without the need for new incentives for replication or meta-analysis. These techniques offer the potential for a more user- (i.e. researcher-) centered approach to statistical analysis in HCI. Matthew Kay 0001, Gregory L. Nelson, Eric B. Hekler |
CHI | 1 |
| 2016 | Cognitive rhythms: unobtrusive and continuous sensing of alertness using a mobile phoneabstractThroughout the day, our alertness levels change and our cognitive performance fluctuates. The creation of technology that can adapt to such variations requires reliable measurement with ecological validity. Our study is the first to collect alertness data in the wild using the clinically validated Psychomotor Vigilance Test. With 20 participants over 40 days, we find that alertness can oscillate approximately 30% depending on time and body clock type and that Daylight Savings Time, hours slept, and stimulant intake can influence alertness as well. Based on these findings, we develop novel methods for unobtrusively and continuously assessing alertness. In estimating response time, our model achieves a root-mean-square error of 80.64 milliseconds, which is significantly lower than the 500ms threshold used as a standard indicator of impaired cognitive ability. Finally, we discuss how such real-time detection of alertness is a key first step towards developing systems that are sensitive to our biological variations. Saeed Abdullah, Elizabeth L. Murnane, Mark Matthews, Matthew Kay 0001, Julie A. Kientz, Geri Gay, Tanzeem Choudhury |
UbiComp | 4 |
| 2016 | Mobile manifestations of alertness: connecting biological rhythms with patterns of smartphone app useabstractOur body clock causes considerable variations in our behavioral, mental, and physical processes, including alertness, throughout the day. While much research has studied technology usage patterns, the potential impact of underlying biological processes on these patterns is under-explored. Using data from 20 participants over 40 days, this paper presents the first study to connect patterns of mobile application usage with these contributing biological factors. Among other results, we find that usage patterns vary for individuals with different body clock types, that usage correlates with rhythms of alertness, that app use features such as duration and switching can distinguish periods of low and high alertness, and that app use reflects sleep interruptions as well as sleep duration. We conclude by discussing how our findings inform the design of biologically-friendly technology that can better support personal rhythms of performance. Elizabeth L. Murnane, Saeed Abdullah, Mark Matthews, Matthew Kay 0001, Julie A. Kientz, Tanzeem Choudhury, Geri Gay, Dan Cosley |
MobileHCI | 4 |
| 2016 | Beyond Weber's Law: A Second Look at Ranking Visualizations of CorrelationabstractModels of human perception - including perceptual "laws" - can be valuable tools for deriving visualization design recommendations. However, it is important to assess the explanatory power of such models when using them to inform design. We present a secondary analysis of data previously used to rank the effectiveness of bivariate visualizations for assessing correlation (measured with Pearson's r) according to the well-known Weber-Fechner Law. Beginning with the model of Harrison et al. [1], we present a sequence of refinements including incorporation of individual differences, log transformation, censored regression, and adoption of Bayesian statistics. Our model incorporates all observations dropped from the original analysis, including data near ceilings caused by the data collection process and entire visualizations dropped due to large numbers of observations worse than chance. This model deviates from Weber's Law, but provides improved predictive accuracy and generalization. Using Bayesian credibility intervals, we derive a partial ranking that groups visualizations with similar performance, and we give precise estimates of the difference in performance between these groups. We find that compared to other visualizations, scatterplots are unique in combining low variance between individuals and high precision on both positively- and negatively-correlated data. We conclude with a discussion of the value of data sharing and replication, and share implications for modeling similar experimental data. Matthew Kay 0001, Jeffrey Heer |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | Unequal Representation and Gender Stereotypes in Image Search Results for OccupationsabstractInformation environments have the power to affect people's perceptions and behaviors. In this paper, we present the results of studies in which we characterize the gender bias present in image search results for a variety of occupations. We experimentally evaluate the effects of bias in image search results on the images people choose to represent those careers and on people's perceptions of the prevalence of men and women in each occupation. We find evidence for both stereotype exaggeration and systematic underrepresentation of women in search results. We also find that people rate search results higher when they are consistent with stereotypes for a career, and shifting the representation of gender in image search results can shift people's perceptions about real-world distributions. We also discuss tensions between desires for high-quality results and broader societal goals for equality of representation in this space. Matthew Kay 0001, Cynthia Matuszek, Sean A. Munson |
CHI | 1 |
| 2015 | How Good is 85%?: A Survey Tool to Connect Classifier Evaluation to Acceptability of AccuracyabstractMany HCI and ubiquitous computing systems are characterized by two important properties: their output is uncertain-it has an associated accuracy that researchers attempt to optimize-and this uncertainty is user-facing-it directly affects the quality of the user experience. Novel classifiers are typically evaluated using measures like the F1 score-but given an F-score of (e.g.) 0.85, how do we know whether this performance is good enough? Is this level of uncertainty actually tolerable to users of the intended application-and do people weight precision and recall equally? We set out to develop a survey instrument that can systematically answer such questions. We introduce a new measure, acceptability of accuracy, and show how to predict it based on measures of classifier accuracy. Out tool allows us to systematically select an objective function to optimize during classifier evaluation, but can also offer new insights into how to design feedback for user-facing classification systems (e.g., by combining a seemingly-low-performing classifier with appropriate feedback to make a highly usable system). It also reveals potential issues with the ubiquitous F1-measure as applied to user-facing systems. Matthew Kay 0001, Shwetak N. Patel, Julie A. Kientz |
CHI | 1 |
| 2015 | SleepTight: low-burden, self-monitoring technology for capturing and reflecting on sleep behaviorsabstractManual tracking of health behaviors affords many benefits, including increased awareness and engagement. However, the capture burden makes long-term manual tracking challenging. In this study on sleep tracking, we examine ways to reduce the capture burden of manual tracking while leveraging its benefits. We report on the design and evaluation of SleepTight, a low-burden, self-monitoring tool that leverages the Android's widgets both to reduce the capture burden and to improve access to information. Through a four-week deployment study (N = 22), we found that participants who used SleepTight with the widgets enabled had a higher sleep diary compliance rate (92%) than participants who used SleepTight without the widgets (73%). In addition, the widgets improved information access and encouraged self-reflection. We discuss how to leverage widgets to help people collect more data and improve access to information, and more broadly, how to design successful manual self-monitoring tools that support self-reflection. Eun Kyoung Choe, Bongshin Lee, Matthew Kay 0001, Wanda Pratt, Julie A. Kientz |
UbiComp | 3 |
| 2013 | There's no such thing as gaining a pound: reconsidering the bathroom scale user interfaceabstractThe weight scale is perhaps the most ubiquitous health sensor of all and is important to many health and lifestyle decisions, but its fundamental interface--a single numerical estimate of a person's current weight--has remained largely unchanged for 100 years. An opportunity exists to impact public health by re-considering this pervasive interface. Toward that end, we investigated the correspondence between consumers' perceptions of weight data and the realities of weight fluctuation. Through an analysis of online product reviews, a journaling study on weight fluctuations, expert interviews, and a large-scale survey of scale users, we found that consumers' perception of weight scale behavior is often disconnected from scales' capabilities and from clinical relevance, and that accurate understanding of weight fluctuation is associated with greater trust in the scale itself. We propose significant changes to how weight data should be presented and discuss broader implications for the design of other ubiquitous health sensing devices. Matthew Kay 0001, Dan Morris 0001, m. c. schraefel, Julie A. Kientz |
UbiComp | 1 |
| 2012 | Lullaby: a capture & access system for understanding the sleep environmentabstractThe bedroom environment can have a significant impact on the quality of a person's sleep. Experts recommend sleeping in a room that is cool, dark, quiet, and free from disruptors to ensure the best quality sleep. However, it is sometimes difficult for a person to assess which factors in the environment may be causing disrupted sleep. In this paper, we present the design, implementation, and initial evaluation of a capture and access system, called Lullaby. Lullaby combines temperature, light, and motion sensors, audio and photos, and an off-the-shelf sleep sensor to provide a comprehensive recording of a person's sleep. Lullaby allows users to review graphs and access recordings of factors relating to their sleep quality and environmental conditions to look for trends and potential causes of sleep disruptions. In this paper, we report results of a feasibility study where participants (N=4) used Lullaby in their homes for two weeks. Based on our experiences, we discuss design insights for sleep technologies, capture and access applications, and personal informatics tools. Matthew Kay 0001, Eun Kyoung Choe, Jesse Shepherd, Ben Greenstein, Nathaniel F. Watson, Sunny Consolvo, Julie A. Kientz |
UbiComp | 1 |
| 2010 | Perceptions and practices of usability in the free/open source software (FoSS) communityabstractThis paper presents results from a study examining perceptions and practices of usability in the free/open source software (FOSS) community. 27 individuals associated with 11 different FOSS projects were interviewed to understand how they think about, act on, and are motivated to address usability issues. Our results indicate that FOSS project members possess rather sophisticated notions of software usability, which collectively mirror definitions commonly found in HCI textbooks. Our study also uncovered a wide range of practices that ultimately work to improve software usability. Importantly, these activities are typically based on close, direct interpersonal relationships between developers and their core users, a group of users who closely follow the project and provide high quality, respected feedback. These relationships, along with positive feedback from other users, generate social rewards that serve as the primary motivations for attending to usability issues on a day-to-day basis. These findings suggest a need to reconceptualize HCI methods to better fit this culture of practice and its corresponding value system. Michael A. Terry, Matthew Kay 0001, Benjamin J. Lafreniere |
CHI | 2 |
| 2010 | Textured agreements: re-envisioning electronic consentabstractResearch indicates that less than 2% of the population reads license agreements during software installation [12]. To address this problem, we developed textured agreements, visually redesigned agreements that employ factoids, vignettes, and iconic symbols to accentuate information and highlight its personal relevance. Notably, textured agreements accomplish these goals without requiring modification of the underlying text. A between-subjects experimental study with 84 subjects indicates these agreements can significantly increase reading times. In our study, subjects spent approximately 37 seconds on agreement screens with textured agreements, compared to 7 seconds in the plain text control condition. A follow-up study examined retention of agreement content, finding that median scores on a comprehension quiz increased by 4 out of 16 points for textured agreements. These results provide convincing evidence of the potential for textured agreements to positively impact software agreement processes. Matthew Kay 0001, Michael A. Terry |
SOUPS | 1 |
| 2009 | Textured agreements: re-envisioning electronic consentabstractNo abstract available. Matthew Kay 0001, Michael A. Terry |
SOUPS | 1 |
| 2008 | Ingimp: introducing instrumentation to an end-user open source applicationabstractOpen source projects are gradually incorporating usability methods into their development practices, but there are still many unmet needs. One particular need for nearly any open source project is data that describes its user base, including information indicating how the software is actually used in practice. This paper presents the concept of open instrumentation, or the augmentation of an open source application to openly collect and publicly disseminate rich application usage data. We demonstrate the concept of open instrumentation in ingimp, a version of the open source GNU Image Manipulation Program that has been modified to collect end-user usage data. ingimp automatically collects five types of data: The commands used, high-level user interface events, overall features of the user's documents, summaries of the user's general computing environment, and users' own descriptions of their planned tasks. In the spirit of open source software, all collected data are made available for anyone to download and analyze. This paper's primary contributions lie in presenting the overall design of ingimp, with a particular focus on how the design addresses two prominent issues in open instrumentation: privacy and motivating use. Michael A. Terry, Matthew Kay 0001, Brad Van Vugt, Brandon Slack, Terry Park |
CHI | 2 |