Katharina Reinecke

dblp:33/318 · DBLP profile ↗
← Back
68ranked-venue papers
8as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 57 · 8 first-author · 25 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Hopeful Failure: How Collaborative Design Fiction Reimagines AI
abstract
While the number of people using AI is growing, the number of people making core AI decisions remains limited. Advocates call for opening up the development of algorithmic systems to a wider range of perspectives, interests, and methods, with particular attention to racial exclusions and harms. This paper responds to this suggestion with two design fiction workshops where 10 Black American participants imagine futures with and against AI. We introduce Exquisite Tellings, selectively reading in-progress stories while co-developing design fiction plots. Across both workshops, participants repeatedly imagined moments of technological failure, including algorithmic breakdowns and mechanical malfunctions. Rather than signaling collapse, these failures surfaced forms of resourcefulness, enabling characters to reconnect with personal and collective capacities obscured by automation. We argue that analyzing specific instances of ‘hopeful failure’—where challenges in AI development reveal broader social possibilities—can help scholars and critics better understand the emerging effects of AI on society.
Jeffrey Basoah, Katharina Reinecke, Daniela Rosner, Ihudiya Finda Ogbonnaya-Ogburu
DIS2
2026 Decoupling of Usefulness and Novelty: Evaluating the Impact of Generative AI on Design Outputs and Novice Designers' Creative Thinking
Tony Zhou, Bin Han 0011, Marx Wang, Zelia Gomes Da Costa Lai, Rock Yuren Pang, Katharina Reinecke, Jacob O. Wobbrock, Alexis Hiniker
CHI8
2026 Regulating AI: Where U.S. State Policy and HCI (Mis)align
abstract
Artificial intelligence (AI) technologies are increasingly adopted into everyday life, with most investment and development concentrated in the U.S. In response to rapid AI integration and scant federal guidelines, U.S. states have formed AI committees charged with studying AI-related societal trade-offs. We analyzed the 18 existing state-level AI committee reports to understand how policymakers discuss AI-related benefits and risks. We then compared the risks surfaced by policymakers to an established taxonomy of AI risks aggregated from literature and examined how policymakers’ concerns align—or misalign—from those of HCI scholars. These insights provide important mileposts for shaping currently ongoing policy initiatives and future research. Our findings reveal important gaps: while committees invoke responsible AI, their framings often omit broader socio-technical concerns emphasized in HCI. We discuss opportunities for HCI to support socio-technical perspectives, employ participatory design, and close the gap between research and policy.
Nino Migineishvili, Alice Gao, Adinawa Adjagbodjou, Dhanaraj Thakur, René Just, Katharina Reinecke
CHI6
2026 Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
abstract
The output quality of large language models (LLMs) can be improved via “reasoning”: generating segments of chain-of-thought (CoT) content to further condition the model prior to producing user-facing output. While these chains contain valuable information, they are verbose and lack explicit organization, making them tedious to review. Moreover, they lack opportunities for user feedback, such as removing unwanted considerations, adding desired ones, or clarifying unclear assumptions. We introduce Interactive Reasoning, an interaction design that visualizes chain-of-thought outputs as a hierarchy of topics and enables user review and modification. We implement interactive reasoning in Hippo, a prototype for AI-assisted decision making in the face of uncertain trade-offs. In a user study with 16 participants, we find that interactive reasoning in Hippo allows users to quickly identify and interrupt erroneous generations, efficiently steer the model towards customized responses, and better understand both model reasoning and model outputs. Our work contributes to a new paradigm that incorporates user oversight into LLM reasoning processes.
Rock Yuren Pang, K. J. Kevin Feng, Shangbin Feng, Chu Li 0001, Yulia Tsvetkov, Jeffrey Heer, Katharina Reinecke
IUI8
2026 Passing the Buck to AI: How Individuals' Decision-Making Patterns Affect Reliance on AI
abstract
Psychological research has identified different patterns individuals have while making decisions, such as vigilance (making decisions after thorough information gathering), hypervigilance (rushed and anxious decision-making), and buckpassing (deferring decisions to others). We examine whether these decision-making patterns affect peoples’ engagement with AI-generated information in decision-making. In an online experiment with 810 participants tasked with distinguishing food facts from myths, we found that a higher buckpassing tendency was positively correlated with the likelihood of seeking AI information and reported reliance on AI, while being negatively correlated with the time spent reading AI explanations. In contrast, the higher a participant tended towards vigilance, the more carefully they scrutinized the AI’s information, as indicated by an increased time spent looking through the AI’s explanations. These findings suggest that a person’s decision-making pattern plays a significant role in their interactions with AI suggestions, which provides a new understanding of individual differences in AI-assisted decision-making.
Katelyn Mei, Rock Yuren Pang, Alex Lyford, Lucy Lu Wang, Katharina Reinecke
ACM Trans. Comput. Hum. Interact.5
2026 Wildfire and Forest Management: Opportunities for HCI Research
abstract
Wildfire and forest management increasingly rely on geospatial technologies, i.e., data and tools contributing to the geographic mapping and analysis of the Earth, to inform measures for the control of wildfires. Nevertheless, challenges arising from domain experts adopting these complex, non-intuitive technologies are not well understood. We interviewed 12 participants in wildfire and forest management to explore the technical and socio-technical nature of these challenges, revealing that (1) knowledge and data are fragmented across stakeholders, ranging from governmental agencies to small landowners. This fragmentation causes participants to (2) struggle in sharing knowledge and expertise. Participants (3) voice concerns about model bias since decisions informed by geospatial technologies can have far-reaching impacts. Yet, they (4) face barriers engaging people most impacted by these decisions. We detail an HCI research agenda that includes: exploring opportunities to connect stakeholders and sharing knowledge, standardizing decision-making, and engaging local communities.
Nino Migineishvili, Madeleine Grunde-McLaughlin, Emmanuel Azuh, Spencer Wood, René Just, Katharina Reinecke
ACM Trans. Comput. Hum. Interact.6
2025 Biased LLMs can Influence Political Decision-Making
abstract
Jillian Fisher, Shangbin Feng, Robert Aron, Thomas Richardson, Yejin Choi, Daniel W Fisher, Jennifer Pan, Yulia Tsvetkov, Katharina Reinecke. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jillian Fisher, Shangbin Feng, Robert Aron, Yejin Choi 0001, Daniel W. Fisher, Jennifer Pan, Yulia Tsvetkov, Katharina Reinecke
ACL (1)9
2025 NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
abstract
Abhinav Sukumar Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, Maarten Sap. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, Maarten Sap
NAACL (Long Papers)4
2025 Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black Users
abstract
AI-supported writing technologies (AISWT) that provide grammatical suggestions, autocomplete sentences, or generate and rewrite text are now a regular feature integrated into many people's workflows. However, little is known about how people perceive the suggestions these tools provide. In this paper, we investigate how Black American users perceive AISWT, motivated by prior findings in natural language processing that highlight how the underlying large language models can contain racial biases. Using interviews and observational user studies with 13 Black American users of AISWT, we found a strong tradeoff between the perceived benefits of using AISWT to enhance their writing style and feeling like ''it wasn't built for us'''. Specifically, participants reported AISWT's failure to recognize commonly used names and expressions in African American Vernacular English, experiencing its corrections as hurtful and alienating and fearing it might further minoritize their culture. We end with a reflection on the tension between AISWT that fail to include Black American culture and language, and AISWT that attempt to mimic it, with attention to accuracy, authenticity, and the production of social difference.
Jeffrey Basoah, Jay L. Cunningham, Erica Adams, Alisha Bose, Kaustubh Yadav, Zhengyang Yang, Katharina Reinecke, Daniela Karin Rosner
Proc. ACM Hum. Comput. Interact.8
2024 Challenges and Considerations for Accessibility Research Across Cultures and Regions
abstract
Postcolonial and decolonial computing examines how technology design and adoption can perpetuate subtle dimensions of coloniality, under-represent certain regions (e.g., the Global South, non-Western regions, Indigenous societies), and marginalize them. There has been a growing interest in interdisciplinary research focusing on marginalized communities, including accessibility and participatory research. Despite the rapid expansion of accessibility research in the last decades, little focus is placed on accessibility issues within marginalized societies, hindering them from effectively benefiting from accessibility research discussions and outcomes. The accessibility and HCI communities still lack comprehensive knowledge on conducting interdisciplinary research that includes diverse cultures and experiences from some of the systematically marginalized regions. This workshop will explore the intersection of accessibility, HCI, and cross-regional studies, bringing together researchers and practitioners to foster collaborations, identify under-explored research areas, and develop guidelines to support inclusive research practices.
Laleh Nourian, Yulia Goldenberg, Muhammad Sadi Adamu, Vikram Kamath Cannanure, Catherine Holloway, Neha Kumar 0001, Katharina Reinecke, Garreth W. Tigwell
ASSETS7
2024 Know Your Audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audience
abstract
Language models (LMs) show promise as tools for communicating science to the general public by simplifying and summarizing complex language. Because models can be prompted to generate text for a specific audience (e.g., college-educated adults), LMs might be used to create multiple versions of plain language summaries for people with different familiarities of scientific topics. However, it is not clear what the benefits and pitfalls of adaptive plain language are. When is simplifying necessary, what are the costs in doing so, and do these costs differ for readers with different background knowledge? Through three within-subjects studies in which we surface summaries for different envisioned audiences to participants of different backgrounds, we found that while simpler text led to the best reading experience for readers with little to no familiarity in a topic, high familiarity readers tended to ignore certain details in overly plain summaries (e.g., study limitations). Our work provides methods and guidance on ways of adapting plain language summaries beyond the single “general” audience.
Tal August, Kyle Lo, Noah A. Smith, Katharina Reinecke
CHI4
2024 BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
abstract
Digital technologies have positively transformed society, but they have also led to undesirable consequences not anticipated at the time of design or development. We posit that insights into past undesirable consequences can help researchers and practitioners gain awareness and anticipate potential adverse effects. To test this assumption, we introduce Blip, a system that extracts real-world undesirable consequences of technology from online articles, summarizes and categorizes them, and presents them in an interactive, web-based interface. In two user studies with 15 researchers in various computer science disciplines, we found that Blip substantially increased the number and diversity of undesirable consequences they could list in comparison to relying on prior knowledge or searching online. Moreover, Blip helped them identify undesirable consequences relevant to their ongoing projects, made them aware of undesirable consequences they “had never considered,” and inspired them to reflect on their own experiences with technology.
Rock Yuren Pang, Sebastin Santy, René Just, Katharina Reinecke
CHI4
2023 NLPositionality: Characterizing Design Biases of Datasets and Models
abstract
Design biases in NLP systems, such as performance differences for different populations, often stem from their creator's positionality, i.e., views and lived experiences shaped by identity and background.Despite the prevalence and risks of design biases, they are hard to quantify because researcher, system, and dataset positionality is often unobserved.We introduce NLPositionality, a framework for characterizing design biases and quantifying the positionality of NLP datasets and models.Our framework continuously collects annotations from a diverse pool of volunteer participants on LabintheWild, and statistically quantifies alignment with dataset labels and model predictions.We apply NLPositionality to existing datasets and models for two tasks-social acceptability and hate speech detection.To date, we have collected 16, 299 annotations in over a year for 600 instances from 1, 096 annotators across 87 countries.We find that datasets and models align predominantly with Western, White, college-educated, and younger populations.Additionally, certain groups, such as nonbinary people and non-native English speakers, are further marginalized by datasets and models as they rank least in alignment across all tasks.Finally, we draw from prior literature to discuss how researchers can examine their own positionality and that of their datasets and models, opening the door for more inclusive NLP systems.
Sebastin Santy, Jenny T. Liang, Ronan Le Bras 0001, Katharina Reinecke, Maarten Sap
ACL (1)4
2023 Why, when, and from whom: considerations for collecting and reporting race and ethnicity data in HCI
abstract
Engaging diverse participants in HCI research is critical for creating safe, inclusive, and equitable technology. However, there is a lack of guidelines on when, why, and how HCI researchers collect study participants’ race and ethnicity. Our paper aims to take the first step toward such guidelines by providing a systematic review and discussion of the status quo of race and ethnicity data collection in HCI. Through an analysis of 2016–2021 CHI proceedings and a survey with 15 authors who published in these proceedings, we found that reporting race and ethnicity of participants is very rare (<3%) and that researchers are far from consensus. Drawing from multidisciplinary literature and our findings, we devise considerations for HCI researchers to decide why, when, and from whom to collect race and ethnicity data. For truly inclusive, equitable technologies, we encourage deliberate decisions rather than default omissions.
Yiqun Chen 0001, Angela D. R. Smith, Katharina Reinecke, Alexandra To
CHI3
2023 "That's important, but...": How Computer Science Researchers Anticipate Unintended Consequences of Their Research Innovations
abstract
Computer science research has led to many breakthrough innovations but has also been scrutinized for enabling technology that has negative, unintended consequences for society. Given the increasing discussions of ethics in the news and among researchers, we interviewed 20 researchers in various CS sub-disciplines to identify whether and how they consider potential unintended consequences of their research innovations. We show that considering unintended consequences is generally seen as important but rarely practiced. Principal barriers are a lack of formal process and strategy as well as the academic practice that prioritizes fast progress and publications. Drawing on these findings, we discuss approaches to support researchers in routinely considering unintended consequences, from bringing diverse perspectives through community participation to increasing incentives to investigate potential consequences. We intend for our work to pave the way for routine explorations of the societal implications of technological innovations before, during, and after the research process.
Kimberly Do, Rock Yuren Pang, Jiachen Jiang, Katharina Reinecke
CHI4
2023 How Language Formality in Security and Privacy Interfaces Impacts Intended Compliance
abstract
Strong end-user security practices benefit both the user and hosting platform, but it is not well understood how companies communicate with their users to encourage these practices. This paper explores whether web companies and their platforms use different levels of language formality in these communications and tests the hypothesis that higher language formality leads to users’ increased intention to comply. We contribute a dataset and systematic analysis of 1,817 English language strings in web security and privacy interfaces across 13 web platforms, showing strong variations in language. An online study with 512 participants further demonstrated that people perceive differences in the language formality across platforms and that a higher language formality is associated with higher self-reported intention to comply. Our findings suggest that formality can be an important factor in designing effective security and privacy prompts. We discuss implications of these results, including how to balance formality with platform language style. In addition to being the first piece of work to analyze language formality in user security, these findings provide valuable insights into how platforms can best communicate with users about account security.
Jackson Stokes, Tal August, Robert A Marver, Alexei Czeskis, Franziska Roesner, Tadayoshi Kohno, Katharina Reinecke
CHI7
2023 PACMHCI V7, CSCW1, April 2023 Editorial
Munmun De Choudhury, Xianghua Ding, Shion Guha, Aparecido Fabiano Pinatti de Carvalho, Hideaki Kuzuoka, Katharina Reinecke, Hao-Chuan Wang, Naomi Yamashita
Proc. ACM Hum. Comput. Interact.6
2022 Generating Scientific Definitions with Controllable Complexity
abstract
Unfamiliar terminology and complex language can present barriers to understanding science. Natural language processing stands to help address these issues by automatically defining unfamiliar terms. We introduce a new task and dataset for defining scientific terms and controlling the complexity of generated definitions as a way of adapting to a specific reader's background knowledge. We test four definition generation methods for this new task, finding that a sequence-to-sequence approach is most successful. We then explore the version of the task in which definitions are generated at a target complexity level. We introduce a novel reranking approach and find in human evaluations that it offers superior fluency while also controlling complexity, compared to several controllable generation baselines.
Tal August, Katharina Reinecke, Noah A. Smith
ACL (1)2
2022 Understanding and Improving Information Extraction From Online Geospatial Data Visualizations for Screen-Reader Users
abstract
Prior work has studied the interaction experiences of screen-reader users with simple online data visualizations (e.g., bar charts, line graphs, scatter plots), highlighting the disenfranchisement of screen-reader users in accessing information from these visualizations. However, the interactions of screen-reader users with online geospatial data visualizations, commonly used by visualization creators to represent geospatial data (e.g., COVID-19 cases per US state), remain unexplored. In this work, we study the interactions of and information extraction by screen-reader users from online geospatial data visualizations. Specifically, we conducted a user study with 12 screen-reader users to understand the information they seek from online geospatial data visualizations and the questions they ask to extract that information. We utilized our findings to generate a taxonomy of information sought from our participants’ interactions. Additionally, we extended the functionalities of VoxLens—an open-source multi-modal solution that improves data visualization accessibility—to enable screen-reader users to extract information from online geospatial data visualizations.
Ather Sharif, Andrew Mingwei Zhang, Anna Shih, Jacob O. Wobbrock, Katharina Reinecke
ASSETS5
2022 Apéritif: Scaffolding Preregistrations to Automatically Generate Analysis Code and Methods Descriptions
abstract
The HCI community has been advocating preregistration as a practice to improve the credibility of scientific research. However, it remains unclear how HCI researchers preregister studies and what preregistration users perceive as benefits and challenges. By systematically reviewing the past four CHI proceedings and surveying 11 researchers, we found that only 1.11% of papers presented preregistered studies, though both authors and reviewers of preregistered studies perceive it as beneficial. Our formative studies revealed key challenges ranging from a lack of detail about the study design, hindering comprehensibility, to inconsistencies between preregistrations and published papers. To explore ways for addressing these issues, we developed Apéritif, a research prototype that scaffolds the preregistration process and automatically generates analysis code and a methods description. In an evaluation with 17 HCI researchers, we found that Apéritif reduces the effort of preregistering a study, facilitates researchers’ workflows, and promotes consistency between research artifacts.
Rock Yuren Pang, Katharina Reinecke, René Just
CHI2
2022 VoxLens: Making Online Data Visualizations Accessible with an Interactive JavaScript Plug-In
abstract
JavaScript visualization libraries are widely used to create online data visualizations but provide limited access to their information for screen-reader users. Building on prior findings about the experiences of screen-reader users with online data visualizations, we present VoxLens, an open-source JavaScript plug-in that—with a single line of code—improves the accessibility of online data visualizations for screen-reader users using a multi-modal approach. Specifically, VoxLens enables screen-reader users to obtain a holistic summary of presented information, play sonified versions of the data, and interact with visualizations in a “drill-down” manner using voice-activated commands. Through task-based experiments with 21 screen-reader users, we show that VoxLens improves the accuracy of information extraction and interaction time by 122% and 36%, respectively, over existing conventional interaction with online data visualizations. Our interviews with screen-reader users suggest that VoxLens is a “game-changer” in making online data visualizations accessible to screen-reader users, saving them time and effort.
Ather Sharif, Olivia H. Wang, Alida T. Muongchan, Katharina Reinecke, Jacob O. Wobbrock
CHI4
2022 Gendered Mental Health Stigma in Masked Language Models
abstract
Mental health stigma prevents many individuals from receiving the appropriate care, and social psychology studies have shown that mental health tends to be overlooked in men.In this work, we investigate gendered mental health stigma in masked language models.In doing so, we operationalize mental health stigma by developing a framework grounded in psychology research: we use clinical psychology literature to curate prompts, then evaluate the models' propensity to generate gendered words.We find that masked language models capture societal stigma about gender in mental health: models are consistently more likely to predict female subjects than male in sentences about having a mental health condition (32% vs. 19%), and this disparity is exacerbated for sentences that indicate treatment-seeking behavior.Furthermore, we find that different models capture dimensions of stigma differently for men and women, associating stereotypes like anger, blame, and pity more with women with mental health conditions than with men.In showing the complex nuances of models' gendered mental health stigma, we demonstrate that context and overlapping dimensions of identity are important considerations when assessing computational models' social biases.
Inna Wanyin Lin, Lucille Njoo, Anjalie Field, Ashish Sharma 0004, Katharina Reinecke, Tim Althoff, Yulia Tsvetkov
EMNLP5
2022 PACMHCI V6, CSCW1, April 2022 Editorial
abstract
No abstract available.
Shaowen Bardzell, Siân E. Lindley, Aleksandra Sarcevic, Hideaki Kuzuoka, Katharina Reinecke, Hao-Chuan Wang, Naomi Yamashita
Proc. ACM Hum. Comput. Interact.5
2022 PACMHCI V6, CSCW2, November 2022 Editorial
abstract
We are delighted to present this issue of the Proceedings of the ACM on Human-Computer Interaction, which contains scholarship from the Computer-Supported Cooperative Work and Social Computing (CSCW) community. This issue has 293 papers, 94 that were accepted from the April 2021 cycle, 61 that were accepted from the July 2021 cycle, and 138 that were accepted from the January 2022 cycle. It reflects great efforts and contributions from external reviewers, Associate Chairs and Editors, who together have conducted a rigorous review process. As Papers Chairs, we are grateful for the community's collective efforts to continue shaping and sharing CSCW's tradition of high-quality scholarship during an ongoing global pandemic.
Hideaki Kuzuoka, Katharina Reinecke, Hao-Chuan Wang, Naomi Yamashita, Shaowen Bardzell, Siân E. Lindley, Aleksandra Sarcevic
Proc. ACM Hum. Comput. Interact.2
2022 An HCI Research Agenda for Online Science Communication
abstract
Social media, blogs, podcasts, and other computer-mediated communication technology have become an integral way for the public to access and engage with research. However, despite the evolving challenges researchers face navigating these platforms, and the high stakes of online science communication, relatively little HCI research has focused on understanding and supporting online science communication through these participatory platforms. Through a review of the literature and a set of interviews with HCI researchers (n = 24), we identify challenges currently facing researchers who try to engage with the public about their work, and establish a research agenda for HCI to study, design, and evaluate technology to support science communication. Specifically, we advocate for the design of tools to support audience analytics, automated summary and outreach workflows, and providing quantitative and qualitative feedback about online outreach efforts, as well as additional research to elucidate the impacts of self-directed science communication efforts and the evolving roles of scientists on the participatory web. With shifting online platforms placing researchers in the role of advocates and participants in science communication, understanding and supporting these interactions is now more important than ever.
Spencer Williams, Ridley Jones, Katharina Reinecke, Gary Hsieh
Proc. ACM Hum. Comput. Interact.3
2021 Respectful Language as Perceived by People with Disabilities
abstract
Respectfully and adequately referring to people with various disabilities is difficult due to societal norms and constantly evolving languages. In this work, we address the question of how expert researchers in the field of accessibility are referring to people with disabilities and whether this terminology corresponds to how people with disabilities prefer to be addressed. By conducting a systematic literature review of the past three ASSETS proceeding, we summarize how accessibility researchers are currently referring to people with disabilities in English. A survey of 63 people with disabilities further revealed that while researchers from ASSETS are using terms that are mostly aligned with participants’ expectations, the same terminologies can be perceived both respectful and disrespectful by varying participants. Through this preliminary work, we pave the path for researchers to further explore respectful terminology and encourage researchers to improve the inclusivity and diversity of language use in our community.
Lior Levy, Qisheng Li, Ather Sharif, Katharina Reinecke
ASSETS4
2021 How Online Tests Contribute to the Support System for People With Cognitive and Mental Disabilities
abstract
Roughly 1 in 3 people around the world are affected by cognitive or mental disabilities at some point in their lives, yet people often face a variety of barriers when seeking support and receiving diagnosis from healthcare professionals. While prior work found that people with such disabilities assess themselves using online tests and assessments, it remains unknown whether and how effectively these tests fill gaps in healthcare and general support systems. To find out, we interviewed 17 adults with cognitive or mental disabilities about their motivation for and experience using online tests. We learned that online tests act as an important resource that address the shortcomings in support systems for people with professionally diagnosed or suspected cognitive or mental disabilities. In particular, online tests can lower barriers to a professional diagnosis, provide valuable information about the nuances of a disability, and support people in forming a disability identity – an invaluable step towards a positive acceptance of oneself. Our results also uncovered challenges and risks that prevent people with known or suspected health conditions from fully taking advantage of online tests. Based on these findings, we discuss how online tests can be better leveraged to support people with cognitive or mental disabilities before and after professional diagnosis.
Qisheng Li, Josephine Lee, Christina Zhang, Katharina Reinecke
ASSETS4
2021 Understanding Screen-Reader Users' Experiences with Online Data Visualizations
abstract
Online data visualizations are widely used to communicate information from simple statistics to complex phenomena, supporting people in gaining important insights from data. However, due to the defining visual nature of data visualizations, extracting information from visualizations can be difficult or impossible for screen-reader users. To assess screen-reader users’ challenges with online data visualizations, we conducted two empirical studies: (1) A qualitative study with nine screen-reader users, and (2) a quantitative study with 36 screen-reader and 36 non-screen-reader users. Our results show that due to the inaccessibility of online data visualizations, screen-reader users extract information 61.48% less accurately and spend 210.96% more time interacting with online data visualizations compared to non-screen-reader users. Additionally, our findings show that online data visualizations are commonly indiscoverable to screen readers. In visualizations that are discoverable and comprehensible, screen-reader users suggested tabular and textual representation of data as techniques to improve the accessibility of online visualizations. Taken together, our results provide empirical evidence of the inequalities screen-readers users face in their interaction with online data visualizations.
Ather Sharif, Sanjana Shivani Chintalapati, Jacob O. Wobbrock, Katharina Reinecke
ASSETS4
2021 Do Cross-Cultural Differences in Visual Attention Patterns Affect Search Efficiency on Websites?
abstract
Prior work in cross-cultural psychology and neuroscience has shown robust variations in visual attention patterns. People from East Asian societies, in which a holistic thinking style predominates, have been found to attend to contextual information in scenes more than Westerners, whose tendency to think analytically expresses itself in greater attention to foreground objects. This paper applies these findings to website design, using an online study to evaluate whether Japanese (N=65) remember more and are faster at finding contextual website information than US Americans (N=84). Our results do not support this hypothesis. Instead, Japanese overall took significantly longer to find information than US participants—a difference that was exacerbated by an increase in website complexity—suggesting that Japanese may holistically take in a website before engaging with detailed information. We discuss implications of these findings for website design and cross-cultural research.
Amanda Baughan, Nigini Oliveira, Tal August, Naomi Yamashita, Katharina Reinecke
CHI5
2021 How WEIRD is CHI?
abstract
Computer technology is often designed in technology hubs in Western countries, invariably making it “WEIRD”, because it is based on the intuition, knowledge, and values of people who are Western, Educated, Industrialized, Rich, and Democratic. Developing technology that is universally useful and engaging requires knowledge about members of WEIRD and non-WEIRD societies alike. In other words, it requires us, the CHI community, to generate this knowledge by studying representative participant samples. To find out to what extent CHI participant samples are from Western societies, we analyzed papers published in the CHI proceedings between 2016-2020. Our findings show that 73% of CHI study findings are based on Western participant samples, representing less than 12% of the world’s population. Furthermore, we show that most participant samples at CHI tend to come from industrialized, rich, and democratic countries with generally highly educated populations. Encouragingly, recent years have seen a slight increase in non-Western samples and those that include several countries. We discuss suggestions for further broadening the international representation of CHI participant samples.
Sebastian Linxen, Christian Sturm 0001, Florian Brühlmann, Vincent Cassau, Klaus Opwis, Katharina Reinecke
CHI6
2020 The Reliability of Fitts's Law as a Movement Model for People with and without Limited Fine Motor Function
abstract
For over six decades, Fitts’s law (1954) has been utilized by researchers to quantify human pointing performance in terms of “throughput,” a combined speed-accuracy measure of aimed movement efficiency. Throughput measurements are commonly used to evaluate pointing techniques and devices, helping to inform software and hardware developments. Although Fitts’s law has been used extensively in HCI and beyond, its test-retest reliability, both in terms of throughput and model fit, from one session to the next, is still unexplored. Additionally, despite the fact that prior work has shown that Fitts’s law provides good model fits, with Pearson correlation coefficients commonly at r=.90 or above, the model fitness of Fitts’s law has not been thoroughly investigated for people who exhibit limited fine motor function in their dominant hand. To fill these gaps, we conducted a study with 21 participants with limited fine motor function and 34 participants without such limitations. Each participant performed a classic reciprocal pointing task comprising vertical ribbons in a 1-D layout in two sessions, which were at least four hours and at most 48 hours apart. Our findings indicate that the throughput values between the two sessions were statistically significantly different, both for people with and without limited fine motor function, suggesting that Fitts’s law provides low test-retest reliability. Importantly, the test-retest reliability of Fitts’s throughput metric was 4.7% lower for people with limited fine motor function. Additionally, we found that the model fitness of Fitts’s law as measured by Pearson correlation coefficient, r, was .89 (SD=0.08) for people without limited fine motor function, and .81 (SD=0.09) for people with limited fine motor function. Taken together, these results indicate that Fitts’s law should be used with caution and, if possible, over multiple sessions, especially when used in assistive technology evaluations.
Ather Sharif, Victoria Pao, Katharina Reinecke, Jacob O. Wobbrock
ASSETS3
2020 Explain like I am a Scientist: The Linguistic Barriers of Entry to r/science
abstract
As an online community for discussing research findings, r/science has the potential to contribute to science outreach and communication with a broad audience. Yet previous work suggests that most of the active contributors on r/science are science-educated people rather than a lay general public. One potential reason is that r/science contributors might use a different, more specialized language than used in other subreddits. To investigate this possibility, we analyzed the language used in more than 68 million posts and comments from 12 subreddits from 2018. We show that r/science uses a specialized language that is distinct from other subreddits. Transient (newer) authors of posts and comments on r/science use less specialized language than more frequent authors, and those that leave the community use less specialized language than those that stay, even when comparing their first comments. These findings suggest that the specialized language used in r/science has a gatekeeping effect, preventing participation by people whose language does not align with that used in r/science. By characterizing r/science's specialized language, we contribute guidelines and tools for increasing the number of contributors in r/science.
Tal August, Dallas Card, Gary Hsieh, Noah A. Smith, Katharina Reinecke
CHI5
2020 Keep it Simple: How Visual Complexity and Preferences Impact Search Efficiency on Websites
abstract
We conducted an online study with 165 participants in which we tested their search efficiency and information recall. We confirm that the visual complexity of a website has a significant negative effect on search efficiency and information recall. However, the search efficiency of those who preferred simple websites was more negatively affected by highly complex websites than those who preferred high visual complexity. Our results suggest that diverse visual preferences need to be accounted for when assessing search response time and information recall in HCI experiments, testing software, or A/B tests.
Amanda Baughan, Tal August, Naomi Yamashita, Katharina Reinecke
CHI4
2020 Writing Strategies for Science Communication: Data and Computational Analysis
abstract
Communicating complex scientific ideas without misleading or overwhelming the public is challenging.While science communication guides exist, they rarely offer empirical evidence for how their strategies are used in practice.Writing strategies that can be automatically recognized could greatly support science communication efforts by enabling tools to detect and suggest strategies for writers.We compile a set of writing strategies drawn from a wide range of prescriptive sources and develop an annotation scheme allowing humans to recognize them.We collect a corpus of 128K science writing documents in English and annotate a subset of this corpus.1 We use the annotations to train transformer-based classifiers and measure the strategies' use in the larger corpus.We find that the use of strategies, such as storytelling and emphasizing the most important findings, varies significantly across publications with different reader audiences.
Tal August, Lauren Kim, Katharina Reinecke, Noah A. Smith
EMNLP (1)3
2019 Pay Attention, Please: Formal Language Improves Attention in Volunteer and Paid Online Experiments
abstract
Participant engagement in online studies is key to collecting reliable data, yet achieving it remains an often discussed challenge in the research community. One factor that might impact engagement is the formality of language used to communicate with participants throughout the study. Prior work has found that language formality can convey social cues and power hierarchies, affecting people's responses and actions. We explore how formality influences engagement, measured by attention, dropout, time spent on the study and participant performance, in an online study with 369 participants on Mechanical Turk (paid) and LabintheWild (volunteer). Formal language improves participant attention compared to using casual language in both paid and volunteer conditions, but does not affect dropout, time spent, or participant performance. We suggest using more formal language in studies containing complex tasks where fully reading instructions is especially important. We also highlight trade-offs that different recruitment incentives provide in online experimentation.
Tal August, Katharina Reinecke
CHI2
2019 r/science: Challenges and Opportunities in Online Science Communication
abstract
Online discussion websites, such as Reddit's r/science forum, have the potential to foster science communication between researchers and the general public. However, little is known about who participates, what is discussed, and whether such websites are successful in achieving meaningful science discussions. To find out, we conducted a mixed-methods study analyzing 11,859 r/science posts and conducting interviews with 18 community members. Our results show that r/science facilitates rich information exchange and that the comments section provides a unique science communication document that guides engagement with scientific research. However, this community-sourced science communication comes largely from a knowledgeable public. We conclude with design suggestions for a number of critical problems that we uncovered: addressing the problem of topic newsworthiness and balancing broader participation and rigor.
Ridley Jones, Lucas Colusso, Katharina Reinecke, Gary Hsieh
CHI3
2019 The Impact of Web Browser Reader Views on Reading Speed and User Experience
abstract
As reading increasingly shifts from paper to online media, many web browsers now provide a "Reader View,'' which modifies web page layout and design for better readability. However, research has yet to establish whether Reader Views are effective in improving readability and how they might change the user experience. We characterize how Mozilla Firefox's Reader View significantly reduces the visual complexity of websites by excluding menus, images, and content. We then conducted an online study with 391 participants (including 42 who self-reported having been diagnosed with dyslexia), showing that compared to standard websites the Reader View increased reading speed by 5% for readers on average, and significantly improved perceived readability and visual appeal. We suggest guidelines for the design of websites and browsers that better support people with varying reading skills.
Qisheng Li, Meredith Ringel Morris, Adam Fourney, Kevin Larson, Katharina Reinecke
CHI5
2019 Tea: A High-level Language and Runtime System for Automating Statistical Analysis
abstract
Though statistical analyses are centered on research questions and hypotheses, current statistical analysis tools are not. Users must first translate their hypotheses into specific statistical tests and then perform API calls with functions and parameters. To do so accurately requires that users have statistical expertise. To lower this barrier to valid, replicable statistical analysis, we introduce Tea, a high-level declarative language and runtime system. In Tea, users express their study design, any parametric assumptions, and their hypotheses. Tea compiles these high-level specifications into a constraint satisfaction problem that determines the set of valid statistical tests and then executes them to test the hypothesis. We evaluate Tea using a suite of statistical analyses drawn from popular tutorials. We show that Tea generally matches the choices of experts while automatically switching to non-parametric tests when parametric assumptions are not met. We simulate the effect of mistakes made by non-expert users and show that Tea automatically avoids both false negatives and false positives that could be produced by the application of incorrect statistical tests.
Eunice Jun, Maureen Daum, Jared Roesch, Sarah E. Chasins, Emery D. Berger, René Just, Katharina Reinecke
UIST7
2018 Volunteer-Based Online Studies With Older Adults and People with Disabilities
abstract
There are few large-scale empirical studies with people with disabilities or older adults, mainly because recruiting partici­pants with specific characteristics is even harder than recruit­ing young and/or non-disabled populations. Analyzing four online experiments on LabintheWild with a total of 355,656 participants, we show that volunteer-based online experiments that provide personalized feedback attract large numbers of participants with diverse disabilities and ages and allow ro­bust studies with these populations that replicate and extend the findings of prior laboratory studies. To find out what mo­tivates people with disabilities to take part, we additionally analyzed participants' feedback and forum entries that discuss LabintheWild experiments. The results show that participants use the studies to diagnose themselves, compare their abilities to others, quantify potential impairments, self-experiment, and share their own stories -- findings that we use to inform design guidelines for online experiment platforms that adequately support and engage people with disabilities.
Qisheng Li, Krzysztof Z. Gajos, Katharina Reinecke
ASSETS3
2018 A Large Inclusive Study of Human Listening Rates
abstract
As conversational agents and digital assistants become increasingly pervasive, understanding their synthetic speech becomes increasingly important. Simultaneously, speech synthesis is becoming more sophisticated and manipulable, providing the opportunity to optimize speech rate to save users time. However, little is known about people's abilities to understand fast speech. In this work, we provide the first large-scale study on human listening rates. Run on LabintheWild, it used volunteer participants, was screen reader accessible, and measured listening rate by accuracy at answering questions spoken by a screen reader at various rates. Our results show that blind and low-vision people, who often rely on audio cues and access text aurally, generally have higher listening rates than sighted people. The findings also suggest a need to expand the range of rates available on personal devices. These results demonstrate the potential for users to learn to listen to faster rates, expanding the possibilities for human-conversational agent interaction.
Danielle Bragg, Cynthia L. Bennett, Katharina Reinecke, Richard E. Ladner
CHI3
2018 A Case for Design Localization: Diversity of Website Aesthetics in 44 Countries
abstract
Adapting the visual designs of websites to a local target audience can be beneficial, because such design localization increases users' appeal, trust, and work efficiency. Yet designers often find it difficult to decide when to adapt and how to adapt the designs, mainly because there are currently no guidelines that describe common website designs in various countries. We contribute the first large-scale analysis of 80,901 website designs across 44 countries, made available via an interactive web-based design catalog. Using computational image metrics to compare the ~2,000 most visited websites per country, we found significant differences between several design aspects, such as a website's colorfulness, visual complexity, the number of text areas and the average saturation of colors. Our results contribute a snapshot of web designs that users in 44 countries frequently see, showing that the design of websites with a global reach are more homogenized compared to local websites between countries.
Manuel Nordhoff, Tal August, Nigini Oliveira, Katharina Reinecke
CHI4
2018 Massive Online Experiment in Cognitive Science
Joshua K. Hartshorne, Josh de Leeuw, Laura T. Germine, Katharina Reinecke, Mariela Jennings
CogSci4
2018 The potential for scientific outreach and learning in mechanical turk experiments
abstract
The global reach of online experiments and their wide adoption in fields ranging from political science to computer science poses an underexplored opportunity for learning at scale: the possibility of participants learning about the research to which they contribute data. We conducted three experiments on Amazon's Mechanical Turk to evaluate whether participants of paid online experiments are interested in learning about research, what information they find most interesting, and whether providing them with such information actually leads to learning gains. Our findings show that 40% of our participants on Mechanical Turk actively sought out post-experiment learning opportunities despite having already received their financial compensation. Participants expressed high interest in a range of research topics, including previous research and experimental design. Finally, we find that participants comprehend and accurately recall facts from post-experiment learning opportunities. Our findings suggest that Mechanical Turk can be a valuable platform for learning at scale and scientific outreach.
Eunice Jun, Morelle Arian, Katharina Reinecke
L@S3
2018 The Impact of Culture on Learner Behavior in Visual Debuggers
abstract
People around the world are learning to code using online resources. However, research has found that these learners might not gain equal benefit from such resources, in particular because culture may affect how people learn from and use online resources. We therefore expect to see cultural differences in how people use and benefit from visual debuggers. We investigated the use of one popular online debugger which allows users to execute Python code and navigate bidirectionally through the execution using forward-steps and back-steps. We examined behavioral logs of 78,369 users from 69 countries and conducted an experiment with 522 participants from 82 countries. We found that people from countries that tend to prefer self-directed learning (such as those from countries with a low Power Distance, which tend to be less hierarchical than others) used about twice as many back-steps. We also found that for individuals whose values aligned with instructor-directed learning (those who scored high on a “Conservation” scale), back-steps were associated with less debugging success.
Kyle Thayer, Philip J. Guo, Katharina Reinecke
VL/HCC3
2018 Framing Effects: Choice of Slogans Used to Advertise Online Experiments Can Boost Recruitment and Lead to Sample Biases
abstract
Online experimentation with volunteers relies on participants' non-financial motivations to complete a study, such as to altruistically support science or to compare oneself to others. Researchers rely on these motivations to attract study participants and often use incentives, like performance comparisons, to encourage participation. Often, these study incentives are advertised using a slogan (e.g., "What is your thinking style?''). Research on framing effects suggests that advertisement slogans attract people with varying demographics and motivations. Could the slogan advertisements for studies risk attracting only specific users? To investigate the existence of potential sample biases, we measured how different slogan frames affected which participants self-selected into studies. We found that slogan frames impact recruitment significantly; changing the slogan frame from a 'supporting science' frame to a 'comparing oneself to others' frame lead to a 9% increase in recruitment for some studies. Additionally, slogans framed as learning more about oneself attract participants significantly more motivated by boredom compared to other slogan frames. We discuss design implications for using frames to improve recruitment and mitigate sources of sample bias in online research with volunteers.
Tal August, Nigini Oliveira, Chenhao Tan, Noah A. Smith, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.5
2018 Digestif: Promoting Science Communication in Online Experiments
abstract
Online experiments allow researchers to collect data from large, demographically diverse global populations. Unlike in-lab studies, however, online experiments often fail to inform participants about the research to which they contribute. This paper is the first to investigate barriers that prevent researchers from providing such science communication in online experiments. We found that the main obstacles preventing researchers from including such information are assumptions about participant disinterest, limited time, concerns about losing anonymity, and concerns about experimental bias. Researchers also noted the dearth of tools to help them close the information loop with their study participants. Based on these findings, we formulated design requirements and implemented Digestif, a new web-based tool that supports researchers in providing their participants with science communication pages. Our evaluation shows that Digestif's scaffolding, examples, and nudges to focus on participants make researchers more aware of their participants' curiosity about research and more likely to disclose pertinent research information.
Eunice Jun, Blue A. Jo, Nigini Oliveira, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.4
2018 The Exchange in StackExchange: Divergences between Stack Overflow and its Culturally Diverse Participants
abstract
StackExchange is a network of Question & Answer (Q&A) sites that support collaborative knowledge exchange on a variety of topics. Prior research found a significant imbalance between those who contribute content to Q&A sites (predominantly people from Western countries) and those who passively use the site (the so-called "lurkers"). One possible explanation for such participation differences between countries could be a mismatch between culturally related preferences of some users and the values ingrained in the design of the site. To examine this hypothesis, we conducted a value-sensitive analysis of the design of the StackExchange site Stack Overflow and contrasted our findings with those of participants from societies with varying cultural backgrounds using a series of focus groups and interviews. Our results reveal tensions between collectivist values, such as the openness for social interactions, and the performance-oriented, individualist values embedded in Stack Overflow's design and community guidelines. This finding confirms that socio-technical sites like Stack Overflow reflect the inherent values of their designers, knowledge that can be leveraged to foster participation equity.
Nigini Oliveira, Michael J. Muller, Nazareno Andrade, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.4
2018 Data Through Others' Eyes: The Impact of Visualizing Others' Expectations on Visualization Interpretation
abstract
In addition to visualizing input data, interactive visualizations have the potential to be social artifacts that reveal other people's perspectives on the data. However, how such social information embedded in a visualization impacts a viewer's interpretation of the data remains unknown. Inspired by recent interactive visualizations that display people's expectations of data against the data, we conducted a controlled experiment to evaluate the effect of showing social information in the form of other people's expectations on people's ability to recall the data, the degree to which they adjust their expectations to align with the data, and their trust in the accuracy of the data. We found that social information that exhibits a high degree of consensus lead participants to recall the data more accurately relative to participants who were exposed to the data alone. Additionally, participants trusted the accuracy of the data less and were more likely to maintain their initial expectations when other people's expectations aligned with their own initial expectations but not with the data. We conclude by characterizing the design space for visualizing others' expectations alongside data.
Yea-Seul Kim, Katharina Reinecke, Jessica Hullman
IEEE Trans. Vis. Comput. Graph.2
2017 The Effect of Performance Feedback on Social Media Sharing at Volunteer-Based Online Experiment Platforms
abstract
As an alternative to online labor markets, several platforms recruit unpaid online volunteers to participate in behavioral experiments that provide personalized feedback. These platforms rely on word-of-mouth sharing by previous participants for recruitment of new participants. We analyzed the impact of performance feedback provided at the end of an experiment on 81,131 participants' sharing behavior. We show that higher performing participants share significantly more. We also show that self-verification has a moderating effect: people who expected to do poorly are not affected by a high score, but people who expected to do as well as others or better, are. In a second experiment, we evaluate three distinct social comparison designs for the presentation of the results. As expected, the design that most emphasized participants' relative success led to most sharing. Contrary to our expectations, people who expected to do poorly benefited from the most optimistic social comparison more than participants who expected to do better than others.
Bernd Huber, Katharina Reinecke, Krzysztof Z. Gajos
CHI2
2017 Explaining the Gap: Visualizing One's Predictions Improves Recall and Comprehension of Data
abstract
Information visualizations use interactivity to enable user-driven querying of visualized data. However, users' interactions with their internal representations, including their expectations about data, are also critical for a visualization to support learning. We present multiple graphically-based techniques for eliciting and incorporating a user's prior knowledge about data into visualization interaction. We use controlled experiments to evaluate how graphically eliciting forms of prior knowledge and presenting feedback on the gap between prior knowledge and the observed data impacts a user's ability to recall and understand the data. We find that participants who are prompted to reflect on their prior knowledge by predicting and self-explaining data outperform a control group in recall and comprehension. These effects persist when participants have moderate or little prior knowledge on the datasets. We discuss how the effects differ based on text versus visual presentations of data. We characterize the design space of graphical prediction and feedback techniques and describe design recommendations.
Yea-Seul Kim, Katharina Reinecke, Jessica Hullman
CHI2
2017 No Such Thing as Too Much Chocolate: Evidence Against Choice Overload in E-Commerce
abstract
E-commerce designers must decide how many products to display at one time. Choice overload research has demonstrated the surprising finding that more choice is not necessarily better?selecting from larger choice sets can be more cognitively demanding and can result in lower levels of choice satisfaction. This research tests the choice overload effect in an e-commerce context and explores how the choice overload effect is influenced by an individual's tendency to maximize or satisfice decisions. We conducted an online experiment with 611 participants randomly assigned to select a gourmet chocolate bar from either 12, 24, 40, 50, 60, or 72 different options. Consistent with prior work, we find that maximizers are less satisfied with their product choice than satisficers. However, using Bayesian analysis, we find that it's unlikely that choice set size affects choice satisfaction by much, if at all. We discuss why the decision-making process may be different in e-commerce contexts than the physical settings used in previous choice overload experiments.
Carol Moser, Chanda Phelan, Paul Resnick, Sarita Yardi Schoenebeck, Katharina Reinecke
CHI5
2017 Citizen Science Opportunities in Volunteer-Based Online Experiments
abstract
Online experimentation with volunteers could be described as a form of citizen science in which participants take part in behavioral studies without financial compensation. However, while citizen science projects aim to improve scientific understanding, volunteer-based online experiment platforms currently provide minimal possibilities for research involvement and learning. The goal of this paper is to uncover opportunities for expanding participant involvement and learning in the research process. Analyzing comments from 8,288 volunteers who took part in four online experiments on LabintheWild, we identified six themes that reveal needs and opportunities for closer interaction between researchers and participants. Our findings demonstrate opportunities for research involvement, such as engaging participants in refining experiment implementations, and learning opportunities, such as providing participants with possibilities to learn about research aims. We translate these findings into ideas for the design of future volunteer-based online experiment platforms that are more mutually beneficial to citizen scientists and researchers.
Nigini Oliveira, Eunice Jun, Katharina Reinecke
CHI3
2017 The Influence of Early Respondents: Information Cascade Effects in Online Event Scheduling
abstract
Sequential group decision-making processes, such as online event scheduling, can be subject to social influence if the decisions involve individuals? subjective preferences and values. Indeed, prior work has shown that scheduling polls that allow respondents to see others' answers are more likely to succeed than polls that hide other responses, suggesting the impact of social influence and coordination. In this paper, we investigate whether this difference is due to information cascade effects in which later respondents adopt the decisions of earlier respondents. Analyzing more than 1.3 million Doodle polls, we found evidence that cascading effects take place during event scheduling, and in particular, that early respondents have a larger influence on the outcome of a poll than people who come late. Drawing on simulations of an event scheduling model, we compare possible interventions to mitigate this bias and show that we can optimize the success of polls by hiding the responses of a small percentage of low availability respondents.
Daniel M. Romero, Katharina Reinecke, Lionel P. Robert Jr.
WSDM2
2017 Types of Motivation Affect Study Selection, Attention, and Dropouts in Online Experiments
abstract
Understanding whether and how motivation affects participation in online experiments is critical because who contributes and how they contribute can affect the validity of findings. Analyzing data from 7,674 participants across three different studies on the volunteer-based online experiment platform LabintheWild, we identified five motivation types for participating: boredom, comparison, fun, science, and self-learning. We found that these motivation types affect study selection, attention, and dropouts. Participants who were highly motivated by boredom paid less attention and were more likely to dropout than those who were motivated by the possibility of contributing to science. We additionally show that motivation can impact study results and suggest how researchers can take participants' motivation into account when designing and analyzing data from volunteer-based online experiments.
Eunice Jun, Gary Hsieh, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.3
2016 Technology at the Table: Attitudes about Mobile Phone Use at Mealtimes
abstract
Mealtimes are a cherished part of everyday life around the world. Often centered on family, friends, or special occasions, sharing meals is a practice embedded with traditions and values. However, as mobile phone adoption becomes increasingly pervasive, tensions emerge about how appropriate it is to use personal devices while sharing a meal with others. Furthermore, while personal devices have been designed to support awareness for the individual user (e.g., notifications), little is known about how to support shared awareness in acceptability in social settings such as meals. In order to understand attitudes about mobile phone use during shared mealtimes, we conducted an online survey with 1,163 English-speaking participants. We find that attitudes about mobile phone use at meals differ depending on the particular phone activity and on who at the meal is engaged in that activity, children versus adults. We also show that three major factors impact participants' attitudes: 1) their own mobile phone use; 2) their age; and 3) whether a child is present at the meal. We discuss the potential for incorporating social awareness features into mobile phone systems to ease tensions around conflicting mealtime behaviors and attitudes.
Carol Moser, Sarita Yardi Schoenebeck, Katharina Reinecke
CHI3
2016 Enabling Designers to Foresee Which Colors Users Cannot See
abstract
Users frequently experience situations in which their ability to differentiate screen colors is affected by a diversity of situations, such as when bright sunlight causes glare, or when monitors are dimly lit. However, designers currently have no way of choosing colors that will be differentiable by users of various demographic backgrounds and abilities and in the wide range of situations where their designs may be viewed. Our goal is to provide designers with insight into the effect of real-world situational lighting conditions on people's ability to differentiate colors in applications and imagery. We therefore developed an online color differentiation test that includes a survey of situational lighting conditions, verified our test in a lab study, and deployed it in an online environment where we collected data from around 30,000 participants. We then created ColorCheck, an image-processing tool that shows designers the proportion of the population they include (or exclude) by their color choices.
Katharina Reinecke, David R. Flatla, Christopher Brooks 0001
CHI1
2015 Infographic Aesthetics: Designing for the First Impression
abstract
Information graphics, or infographics, combine elements of data visualization with design and have become an increasingly popular means for disseminating data. While several studies have suggested that aesthetics in visualization and infographics relate to desirable outcomes like engagement and memorability, it remains unknown how quickly aesthetic impressions are formed, and what it is that makes an infographic appealing. We address these questions by analyzing 1,278 participants' ratings on appeal after seeing infographics for 500ms. Our results establish that: 1) people form a reliable first impression of the appeal of an infographic based on a mere exposure effect, 2) this first impression is largely based on colorfulness and visual complexity, and 3) age, gender, and education level influence the preferred level of colorfulness and complexity. More generally, these findings suggest that outcomes such as engagement and memorability might be determined much earlier than previously thought.
Lane Harrison, Katharina Reinecke, Remco Chang
CHI2
2015 LabintheWild: Conducting Large-Scale Online Experiments With Uncompensated Samples
abstract
Web-based experimentation with uncompensated and unsupervised samples has the potential to support the replication, verification, extension and generation of new results with larger and more diverse sample populations than previously seen. We introduce the experimental online platform LabintheWild, which provides participants with personalized feedback in exchange for participation in behavioral studies. In comparison to conventional in-lab studies, LabintheWild enables the recruitment of participants at larger scale and from more diverse demographic and geographic backgrounds. We analyze Google Analytics data, participants' comments, and tweets to discuss how participants hear about the platform, and why they might choose to participate. Analyzing three example experiments, we additionally show that these experiments replicate previous in-lab study results with comparable data quality.
Katharina Reinecke, Krzysztof Z. Gajos
CSCW1
2014 Quantifying visual preferences around the world
abstract
Website aesthetics have been recognized as an influential moderator of people's behavior and perception. However, what users perceive as "good design" is subject to individual preferences, questioning the feasibility of universal design guidelines. To better understand how people's visual preferences differ, we collected 2.4 million ratings of the visual appeal of websites from nearly 40 thousand participants of diverse backgrounds. We address several gaps in the knowledge about design preferences of previously understudied groups. Among other findings, our results show that the level of colorfulness and visual complexity at which visual appeal is highest strongly varies: Females, for example, liked colorful websites more than males. A high education level generally lowers this preference for colorfulness. Russians preferred a lower visual complexity, and Macedonians liked highly colorful designs more than any other country in our dataset. We contribute a computational model and estimates of peak appeal that can be used to support rapid evaluations of website design prototypes for specific target groups.
Katharina Reinecke, Krzysztof Z. Gajos
CHI1
2014 Demographic differences in how students navigate through MOOCs
abstract
The current generation of Massive Open Online Courses (MOOCs) attract a diverse student audience from all age groups and over 196 countries around the world. Researchers, educators, and the general public have recently become interested in how the learning experience in MOOCs differs from that in traditional courses. A major component of the learning experience is how students navigate through course content.
Philip J. Guo, Katharina Reinecke
L@S2
2013 SPRWeb: preserving subjective responses to website colour schemes through automatic recolouring
abstract
Colours are an important part of user experiences on the Web. Colour schemes influence the aesthetics, first impressions and long-term engagement with websites. However, five percent of people perceive a subset of all colours because they have colour vision deficiency (CVD), resulting in an unequal and less-rich user experience on the Web. Traditionally, people with CVD have been supported by recolouring tools that improve colour differentiability, but do not consider the subjective properties of colour schemes while recolouring. To address this, we developed SPRWeb, a tool that recolours websites to preserve subjective responses and improve colour differentiability - thus enabling users with CVD to have similar online experiences. To develop SPRWeb, we extended existing models of non-CVD subjective responses to CVD, then used this extended model to steer the recolouring process. In a lab study, we found that SPRWeb did significantly better than a standard recolouring tool at preserving the temperature and naturalness of websites, while achieving similar weight and differentiability preservation. We also found that recolouring did not preserve activity, and hypothesize that visual complexity influences activity more than colour. SPRWeb is the first tool to automatically preserve the subjective and perceptual properties of website colour schemes thereby equalizing the colour-based web experience for people with CVD.
David R. Flatla, Katharina Reinecke, Carl Gutwin, Krzysztof Z. Gajos
CHI2
2013 Crowdsourcing performance evaluations of user interfaces
abstract
Online labor markets, such as Amazon's Mechanical Turk (MTurk), provide an attractive platform for conducting human subjects experiments because the relative ease of recruitment, low cost, and a diverse pool of potential participants enable larger-scale experimentation and faster experimental revision cycle compared to lab-based settings. However, because the experimenter gives up the direct control over the participants' environments and behavior, concerns about the quality of the data collected in online settings are pervasive. In this paper, we investigate the feasibility of conducting online performance evaluations of user interfaces with anonymous, unsupervised, paid participants recruited via MTurk. We implemented three performance experiments to re-evaluate three previously well-studied user interface designs. We conducted each experiment both in lab and online with participants recruited via MTurk. The analysis of our results did not yield any evidence of significant or substantial differences in the data collected in the two settings: All statistically significant differences detected in lab were also present on MTurk and the effect sizes were similar. In addition, there were no significant differences between the two settings in the raw task completion times, error rates, consistency, or the rates of utilization of the novel interaction mechanisms introduced in the experiments. These results suggest that MTurk may be a productive setting for conducting performance evaluations of user interfaces providing a complementary approach to existing methodologies.
Steven Komarov, Katharina Reinecke, Krzysztof Z. Gajos
CHI2
2013 Predicting users' first impressions of website aesthetics with a quantification of perceived visual complexity and colorfulness
abstract
Users make lasting judgments about a website's appeal within a split second of seeing it for the first time. This first impression is influential enough to later affect their opinions of a site's usability and trustworthiness. In this paper, we demonstrate a means to predict the initial impression of aesthetics based on perceptual models of a website's colorfulness and visual complexity. In an online study, we collected ratings of colorfulness, visual complexity, and visual appeal of a set of 450 websites from 548 volunteers. Based on these data, we developed computational models that accurately measure the perceived visual complexity and colorfulness of website screenshots. In combination with demographic variables such as a user's education level and age, these models explain approximately half of the variance in the ratings of aesthetic appeal given after viewing a website for 500ms only.
Katharina Reinecke, Tom Yeh, Luke Miratrix, Rahmatri Mardiko, Yuechen Zhao, Jenny Liu, Krzysztof Z. Gajos
CHI1
2013 Doodle around the world: online scheduling behavior reflects cultural differences in time perception and group decision-making
abstract
Event scheduling is a group decision-making process in which social dynamics influence people's choices and the overall outcome. As a result, scheduling is not simply a matter of finding a mutually agreeable time, but a process that is shaped by social norms and values, which can highly vary between countries. To investigate the influence of national culture on people's scheduling behavior we analyzed more than 1.5 million Doodle date/time polls from 211 countries. We found strong correlations between characteristics of national culture and several behavioral phenomena, such as that poll participants from collectivist countries respond earlier, agree to fewer options but find more consensus than predominantly individualist societies. Our study provides empirical evidence of behavioral differences in group decision-making and time perception with implications for cross-cultural collaborative work.
Katharina Reinecke, Minh Khoa Nguyen, Abraham Bernstein, Michael Näf, Krzysztof Z. Gajos
CSCW1
2012 Accurate measurements of pointing performance from in situ observations
abstract
We present a method for obtaining lab-quality measurements of pointing performance from unobtrusive observations of natural in situ interactions. Specifically, we have developed a set of user-independent classifiers for discriminating between deliberate, targeted mouse pointer movements and those movements that were affected by any extraneous factors. To develop and validate these classifiers, we developed logging software to unobtrusively record pointer trajectories as participants naturally interacted with their computers over the course of several weeks. Each participant also performed a set of pointing tasks in a formal study set-up. For each movement, we computed a set of measures capturing nuances of the trajectory and the speed, acceleration, and jerk profiles. Treating the observations from the formal study as positive examples of deliberate, targeted movements and the in situ observations as unlabeled data with an unknown mix of deliberate and distracted interactions, we used a recent advance in machine learning to develop the classifiers. Our results show that, on four distinct metrics, the data collected in-situ and filtered with our classifiers closely matches the results obtained from the formal experiment.
Krzysztof Z. Gajos, Katharina Reinecke, Charles Herrmann
CHI2
2011 MOCCA - a system that learns and recommends visual preferences based on cultural similarity
abstract
We demonstrate our culturally adaptive system MOCCA, which is able to automatically adapt its visual appearance to the user's national culture. Rather than only adapting to one nationality, MOCCA takes into account a person's current and previous countries of residences, and uses this information to calculate user-specific preferences. In addition, the system is able to learn new, and refine existing adaptation rules from users' manual modifications of the user interface based on a collaborative filtering mechanism, and from observing the user's interaction with the interface.
Katharina Reinecke, Patrick Minder, Abraham Bernstein
IUI1
2011 Improving performance, perceived usability, and aesthetics with culturally adaptive user interfaces
abstract
When we investigate the usability and aesthetics of user interfaces, we rarely take into account that what users perceive as beautiful and usable strongly depends on their cultural background. In this paper, we argue that it is not feasible to design one interface that appeals to all users of an increasingly global audience. Instead, we propose to design culturally adaptive systems, which automatically generate personalized interfaces that correspond to cultural preferences. In an evaluation of one such system, we demonstrate that a majority of international participants preferred their personalized versions over a nonadapted interface of the same Website. Results show that users were 22% faster using the culturally adapted interface, needed fewer clicks, and made fewer errors, in line with subjective results demonstrating that they found the adapted version significantly easier to use. Our findings show that interfaces that adapt to cultural preferences can immensely increase the user experience.
Katharina Reinecke, Abraham Bernstein
ACM Trans. Comput. Hum. Interact.1
2009 Tell Me Where You've Lived, and I'll Tell You What You Like: Adapting Interfaces to Cultural Preferences
Katharina Reinecke, Abraham Bernstein
UMAP1