Hong Shen 0004

dblp:74/3247-4 · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
26since 2021 · last 2026
0000-0002-5364-3718ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 24 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PeerCoPilot: A Language Model-Powered Assistant for Behavioral Health Organizations
abstract
Behavioral health conditions, which include mental health and substance use disorders, are the leading disease burden in the United States. Peer-run behavioral health organizations (PROs) critically assist individuals facing these conditions by combining mental health services with assistance for needs such as income, employment, and housing. However, limited funds and staffing make it difficult for PROs to address all service user needs. To assist peer providers at PROs with their day-to-day tasks, we introduce PeerCoPilot, a large language model (LLM)-powered assistant that helps peer providers create wellness plans, construct step-by-step goals, and locate organizational resources to support these goals. PeerCoPilot ensures information reliability through a retrieval-augmented generation pipeline backed by a large database of over 1,300 vetted resources. We conducted human evaluations with 15 peer providers and 6 service users and found that over 90% of users supported using PeerCoPilot. Moreover, we demonstrate that PeerCoPilot provides more reliable and specific information than a baseline LLM. PeerCoPilot is now used by a group of 5-10 peer providers at CSPNJ, a large behavioral health organization serving over 10,000 service users, and we are actively expanding PeerCoPilot's use.
Gao Mo, Naveen Raman 0001, Megan Chai, Cindy Peng, Shannon Pagdon, Nev Jones, Hong Shen 0004, Margaret Swarbrick, Fei Fang 0001
AAAI7
2026 When Your Boss Is an AI Bot: Exploring Opportunities and Risks of Manager Clone Agents in the Future Workplace
abstract
As Generative AI (GenAI) becomes increasingly embedded in the workplace, managers are beginning to create Manager Clone Agents—AI-powered digital surrogates trained on their work communications and decision patterns to perform managerial tasks on their behalf. To investigate this emerging phenomenon, we conducted six design fiction workshops (n = 23) with managers and workers, in which participants co-created speculative scenarios and discussed how Manager Clone Agents might transform collaborative work. We identified four potential roles that participants envisioned for Manager Clone Agents: proxy presence, informational conveyor, productivity engine, and leadership amplifier, while highlighting concerns spanning individual, interpersonal, and organizational levels. We provide design recommendations envisioned by both parties for integrating Manager Clone Agents responsibly into the future workplace, emphasizing the need to prioritize workers’ perspectives and nurture interpersonal bonds while also anticipating alternative futures that may disrupt managerial hierarchies.
Qing Xiao 0002, Hancheng Cao, Hong Shen 0004
CHI4
2026 Large Language Models in Peer-Run Community Behavioral Health Services: Understanding Peer Specialists and Service Users' Perspectives on Opportunities, Risks, and Mitigation Strategies
abstract
Peer-run organizations (PROs) provide critical, recovery-based behavioral health support rooted in lived experience. As large language models (LLMs) enter this domain, their scale, conversationality, and opacity introduce new challenges for situatedness, trust, and autonomy. Partnering with Collaborative Support Programs of New Jersey (CSPNJ), a statewide PRO in the Northeastern United States, we used comicboarding, a co-design method, to conduct workshops with 16 peer specialists and 10 service users exploring perceptions of integrating an LLM-based recommendation system into peer support. Findings show that depending on how LLMs are introduced, constrained, and co-used, they can reconfigure in-room dynamics by sustaining, undermining, or amplifying the relational authority that grounds peer support. We identify opportunities, risks, and mitigation strategies across three tensions: bridging scale and locality, protecting trust and relational dynamics, and preserving peer autonomy amid efficiency gains. We contribute design implications that center lived-experience-in-the-loop, reframe trust as co-constructed, and position LLMs not as clinical tools but as relational collaborators in high-stakes, community-led care.
Cindy Peng, Megan Chai, Gao Mo, Naveen Raman 0001, Ningjing Tang, Shannon Pagdon, Margaret Swarbrick, Nev Jones, Fei Fang 0001, Hong Shen 0004
CHI10
2026 Worker Discretion Advised: Co-designing Risk Disclosure in Crowdsourced Responsible AI (RAI) Content Work
abstract
Responsible AI (RAI) content work, such as annotation, moderation, or red teaming for AI safety, often exposes crowd workers to potentially harmful content. While prior work has underscored the importance of communicating well-being risk to employed content moderators, designing effective disclosure mechanisms for crowd workers while balancing worker protection with the needs of task designers and platforms remains largely unexamined. To address this gap, we conducted individual co-design sessions with 15 task designers, 11 crowdworkers, and 3 platform representatives. We investigated task designer preferences for support in disclosing tasks, worker preferences for receiving risk disclosure warnings, and how platform representatives envision their role in shaping risk disclosure practices. We identify design tensions and map the sociotechnical tradeoffs that shape disclosure practices. We contribute design recommendations and feature concepts for risk disclosure mechanisms in the context of RAI content work.
Alice Qian Zhang, Ryland Shaw, Jina Suh, Laura A. Dabbish, Hong Shen 0004
CHI6
2026 The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
abstract
Large language models can influence users through conversation, creating new forms of dark patterns that differ from traditional UX dark patterns. We define LLM dark patterns as manipulative or deceptive behaviors enacted in dialogue. Drawing on prior work and AI incident reports, we outline a diverse set of categories with real-world examples. Using them, we conducted a scenario-based study where participants (N=34) compared manipulative and neutral LLM responses. Our results reveal that recognition of LLM dark patterns often hinged on conversational cues such as exaggerated agreement, biased framing, or privacy intrusions, but these behaviors were also sometimes normalized as ordinary assistance. Users’ perceptions of these dark patterns shaped how they respond to them. Responsibilities for these behaviors were also attributed in different ways, with participants assigning it to companies and developers, the model itself, or to users. We conclude with implications for design, advocacy, and governance to safeguard user autonomy.
Yike Shi, Qing Xiao 0002, Hong Shen 0004, Hua Shen 0005
CHI4
2026 Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source Platforms
abstract
Model documentation plays a crucial role in promoting responsible AI (RAI) development. The emergence of Generative AI (GenAI) models has reshaped the conditions under which documentation is produced, particularly on open-source platforms where models are hosted and shared. To examine how these changes have manifested in developers’ documentation practices, we interviewed 17 GenAI developers who document models on open-source platforms. Our findings illustrate that uncertainties have become a defining feature of developers’ GenAI documentation practices and that these uncertainties unfold in three interrelated forms: (1) normative and epistemic uncertainties in determining documentation content; (2) methodological uncertainties in evaluating and communicating model properties; and (3) ecosystemic uncertainties about who should document. We argue that these uncertainties in GenAI documentation require coordinated interventions, including infrastructural support to address epistemic and methodological uncertainties, community-based mechanisms to cultivate RAI documentation norms, and collaboration across supply-chain actors to address ecosystemic uncertainties.
Ningjing Tang, Megan Li, Amy A. Winecoff, Michael A. Madaio, Hoda Heidari, Hong Shen 0004
CHI6
2026 Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
abstract
Theory of Mind (ToM) -- the ability to infer what others are thinking (e.g., intentions) from observable cues -- is traditionally considered fundamental to human social interactions. This has sparked growing efforts in building and benchmarking AI's ToM capability, yet little is known about how such capability could translate into the design and experience of everyday user-facing AI products and services. We conducted 13 co-design sessions with 26 U.S.-based AI practitioners to envision, reflect, and distill design recommendations for ToM-enabled everyday AI products and services that are both future-looking and grounded in the realities of AI design and development practices. Analysis revealed three interrelated design recommendations: ToM-enabled AI should 1) be situated in the social context that shape users' mental states, 2) be responsive to the dynamic nature of mental states, and 3) be attuned to subjective individual differences. We surface design tensions within each recommendation that reveal a broader gap between practitioners' envisioned futures of ToM-enabled AI and the realities of current AI design and development practices. These findings point toward the need to move beyond static, inference-driven approach to ToM and toward designing ToM as a pervasive capability that supports continuous human-AI interaction loops.
Qiaosi Wang, Jini Kim, Avanita Sharma, Alicia (Hyun Jin) Lee, Jodi Forlizzi, Hong Shen 0004
CHI6
2026 Can GenAI Move from Individual Use to Collaborative Work? Experiences, Challenges, and Opportunities of Coordinating GenAI into Collaborative Newswork
abstract
Generative AI (GenAI) is reshaping work, but adoption remains largely individual and experimental rather than coordinated into collaborative work. Whether GenAI can move from individual use to collaborative work is a critical question for future organizations. Journalism offers a compelling site to examine this shift: individual journalists have already been disrupted by GenAI tools; yet newswork is inherently collaborative relying on shared norms and coordinated workflows. We conducted 27 interviews with newsroom managers, editors and front-line journalists in China. We found that journalists frequently used GenAI to support daily tasks, but value alignment was safeguarded mainly through individual discretion. At the organizational level, GenAI use remained disconnected from team workflows, hindered by structural barriers and cultural reluctance to share practices. These findings underscore the gap between individual and collaborative work, pointing to the need to account for organizational structures, cultural norms, and workflow when coordinating GenAI for collaborative work.
Qing Xiao 0002, Jingjia Xiao, Hancheng Cao, Hong Shen 0004
CHI5
2026 Do Teachers Dream of GenAI Widening Educational (In)equality? Envisioning the Future of K-12 GenAI Education from Global Teachers' Perspectives
abstract
Generative artificial intelligence (GenAI) is rapidly entering K-12 classrooms worldwide, initiating urgent debates about its potential to either reduce or exacerbate educational inequalities. Drawing on interviews with 30 K-12 teachers across the United States, South Africa, and Taiwan, this study examines how teachers navigate this GenAI tension around educational equalities. We found teachers actively framed GenAI education as an equality-oriented practice: they used it to alleviate pre-existing inequalities while simultaneously working to prevent new inequalities from emerging. Despite these efforts, teachers confronted persistent systemic barriers, i.e., unequal infrastructure, insufficient professional training, and restrictive social norms, that individual initiative alone could not overcome. Teachers thus articulated normative visions for more inclusive GenAI education. By centering teachers’ practices, constraints, and future envisions, this study contributes a global account of how GenAI education is being integrated into K-12 contexts and highlights what is required to make its adoption genuinely equal.
Ruiwei Xiao, Qing Xiao 0002, Xinying Hou, Phenyo Phemelo Moletsane, Hanqi Jane Li, Hong Shen 0004, John C. Stamper
CHI6
2026 Content Creation with Generative AI: How Do Content Creators Responsibly Use Generative AI Tools? CSCW009
abstract
The rise of Generative AI (GenAI) has demonstrated significant potential to improve productivity and foster creativity among content creators, social media influencers with large audiences on platforms such as Instagram, TikTok, and YouTube. However, as GenAI tools became increasingly integrated into creative workflows, significant concerns have emerged about potential risks and harms, including misinformation, social biases, and threats to authenticity. While prior research in HCI and CSCW has documented the pressures content creators face within algorithmic ecosystems, relatively little is known about how creators practically manage responsibility work when using GenAI tools. To address this gap, we conducted semi-structured interviews (N = 16) with content creators active on popular social media platforms such as YouTube, Instagram, and TikTok, examining their motivations, practices, and specific challenges related to responsible GenAI use. Our findings reveal that creators’ motivations for practicing responsible AI use span personal reputation management, audience trust-building, and broader social responsibility. However, they face persistent tensions, as integrating GenAI significantly intensifies conflicts between responsible AI practices and the pressures of visibility, engagement, and monetization imposed by platform algorithms. Content creators are required to perform extensive and often invisible responsibility work, which directly conflicts with the rapid production cycles and engagement demands of algorithm-driven platforms. Based on these insights, we propose concrete socio-technical design implications at the individual, community, and institutional levels, advocating solutions that shift responsibility beyond individual creators alone.
Jini Kim, Manqing Yu, Jiayin Zhi, Stephanie Milani, Jingwen Cheng, Xianzhe Fan, Hong Shen 0004, Jodi Forlizzi
Proc. ACM Hum. Comput. Interact.7
2026 Locating Risk: Task Designers and the Challenge of Risk Disclosure in Crowdsourced RAI Content Work CSCW029
abstract
As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highlighted the risks to worker well-being associated with RAI content work, far less attention has been paid to how these risks are communicated to workers by task designers or individuals who design and post RAI tasks. Existing transparency frameworks and guidelines, such as model cards, datasheets, and crowdworksheets, focus on documenting model information and dataset collection processes, but they overlook an important aspect of disclosing well-being risks to workers. In the absence of standard workflows or clear guidance, the consistent application of content warnings, consent flows, or other forms of well-being risk disclosure remains unclear. This study investigates how task designers approach risk disclosure in crowdsourced RAI tasks. Drawing on interviews with 23 task designers across academic and industry sectors, we examine how well-being risk is recognized, interpreted, and communicated in practice. Our findings highlight the need to support task designers in identifying and communicating risks not only to support crowdworker well-being but also to strengthen the ethical integrity and technical efficacy of AI development pipelines.
Alice Qian Zhang, Ryland Shaw, Laura A. Dabbish, Jina Suh, Hong Shen 0004
Proc. ACM Hum. Comput. Interact.5
2025 User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
abstract
Peer Reviewed
Xianzhe Fan, Qing Xiao 0002, Jiaxin Pei, Maarten Sap, Zhicong Lu, Hong Shen 0004
CHI7
2025 POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
Evans Xu Han, Alice Qian Zhang, Haiyi Zhu, Hong Shen 0004, Paul Pu Liang, Jane Hsieh
UIST4
2025 Navigating Security and Privacy Threats in Homeless Service Provision
Ruoxi Zhang, Shiyue Liu, Mufei He, Aidan Hong, Jeremy J. Northup, Calla Kainaroi, Fei Fang 0001, Hong Shen 0004
USENIX Security Symposium9
2025 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research Practices
abstract
Large language models are increasingly applied in real-world scenarios, including research and education. These models, however, come with well-known ethical issues, which may manifest in unexpected ways in human-computer interaction research due to the extensive engagement with human subjects. This paper reports on research practices related to LLM use, drawing on 16 semi-structured interviews and a survey with 50 HCI researchers. We discuss the ways in which LLMs are already being utilized throughout the entire HCI research pipeline, from ideation to system development and paper writing. While researchers described nuanced understandings of ethical issues, they were rarely or only partially able to identify and address those ethical concerns in their own projects. This lack of action and reliance on workarounds was explained through the perceived lack of control and distributed responsibility in the LLM supply chain, the conditional nature of engaging with ethics, and competing priorities. Finally, we reflect on the implications of our findings and present opportunities to shape emerging norms of engaging with large language models in HCI research.
Shivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li 0001, Hong Shen 0004
Proc. ACM Hum. Comput. Interact.5
2025 "You're in a Ferrari. I'm Waiting for the Bus": Confronting Tensions in Community-University Partnerships
abstract
There have been increasing calls within HCI to build sustained partnerships with communities that go beyond surface-level engagement. However, little is known about how communities view such partnerships and their outcomes. In collaboration with a community-based organization, we co-analyzed a series of interviews to understand the impacts of university-led research initiatives and publicly deployed technologies on local communities, and to explore strategies for more equitable community-university partnerships. Our findings reveal that local communities often perceive technology companies and academic institutions as potential threats due to their shared role in a series of projects, including predictive policing, surveillance, and broader concerns on technological bias and exclusion against minoritized groups. While interviewees named material benefits, sustained relationships, and meaningful accountability as desirable from universities, they pointed to academia's institutional priorities that pose barriers to forming effective partnerships. Drawing from la paperson's concept of a Third University, we argue that researchers and academic institutions must contend with these complexities, while taking a decolonizing approach to community-university partnerships through the lens of revestment.
Cella Monet Sum, Jiayin Zhi, Amil N. T. Cook, Patrick James Cooper, Arturo Lozano, Tj Johnson, Jason Perez, Rayid Ghani, Michael Skirpan, Motahhare Eslami, Hong Shen 0004, Sarah E. Fox
Proc. ACM Hum. Comput. Interact.11
2025 AURA: Amplifying Understanding, Resilience, and Awareness for Responsible AI Content Work
abstract
Behind the scenes of maintaining the safety of technology products from harmful and illegal digital content lies unrecognized human labor. The recent rise in the use of generative AI technologies and the accelerating demands to meet responsible AI (RAI) aims necessitates an increased focus on the labor behind such efforts in the age of AI. This study investigates the nature and challenges of content work that supports RAI efforts, or "RAI content work," that spans content moderation, data labeling, and red teaming -- through the lived experiences of content workers. We conduct a formative survey and semi-structured interview studies to develop a conceptualization of RAI content work and a subsequent framework of recommendations for providing holistic support for content workers. We validate our recommendations through a series of workshops with content workers and derive considerations for and examples of implementing such recommendations. We discuss how our framework may guide future innovation to support the well-being and professional development of the RAI content workforce.
Alice Qian Zhang, Judith Amores, Hong Shen 0004, Mary Czerwinski, Mary L. Gray, Jina Suh
Proc. ACM Hum. Comput. Interact.3
2024 PATIENT-ψ: Using Large Language Models to Simulate Patients for Training Mental Health Professionals
abstract
Ruiyi Wang, Stephanie Milani, Jamie C. Chiu, Jiayin Zhi, Shaun M. Eack, Travis Labrum, Samuel M Murphy, Nev Jones, Kate V Hardy, Hong Shen, Fei Fang, Zhiyu Chen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ruiyi Wang, Stephanie Milani, Jamie C. Chiu, Jiayin Zhi, Shaun M. Eack, Travis Labrum, Samuel M. Murphy, Nev Jones, Kate Hardy, Hong Shen 0004, Fei Fang 0001, Zhiyu Chen 0002
EMNLP10
2024 Predicting and Presenting Task Difficulty for Crowdsourcing Food Rescue Platforms
abstract
Food waste and food insecurity are two problems that co-exist worldwide. A major force to combat food waste and insecurity, food rescue platforms (FRP) match food donations to low-resource communities. Since they rely on external volunteers to deliver the food, communicating rescue task difficulty to volunteers is very important for volunteer engagement and retention. We develop a hybrid model with tabular and natural language data to predict the difficulty of a given rescue trip, which significantly outperforms baselines in identifying easy and hard rescues. Furthermore, using storyboards, we conducted interviews with different stakeholders to understand their perspectives on how to integrate such predictions into volunteers' workflow. Motivated by our findings, we developed three explanation methods to generate interpretable insights for volunteers to better understand the predictions. The results from this study are in the process of being adopted at Food Rescue Hero, a large FRP serving over 25 cities across the United States.
Zheyuan Shi, Jiayin Zhi, Siqi Zeng 0001, Zhicheng Zhang 0003, Ameesh Kapoor, Sean Hudson, Hong Shen 0004, Fei Fang 0001
WWW7
2023 Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry Practice
abstract
Recent years have seen growing interest among both researchers and practitioners in user-engaged approaches to algorithm auditing, which directly engage users in detecting problematic behaviors in algorithmic systems. However, we know little about industry practitioners’ current practices and challenges around user-engaged auditing, nor what opportunities exist for them to better leverage such approaches in practice. To investigate, we conducted a series of interviews and iterative co-design activities with practitioners who employ user-engaged auditing approaches in their work. Our findings reveal several challenges practitioners face in appropriately recruiting and incentivizing user auditors, scaffolding user audits, and deriving actionable insights from user-engaged audit reports. Furthermore, practitioners shared organizational obstacles to user-engaged auditing, surfacing a complex relationship between practitioners and user auditors. Based on these findings, we discuss opportunities for future HCI research to help realize the potential (and mitigate risks) of user-engaged auditing in industry practice.
Wesley Deng, Bill Boyuan Guo, Alicia DeVrio, Hong Shen 0004, Motahhare Eslami, Kenneth Holstein
CHI4
2023 Understanding Frontline Workers' and Unhoused Individuals' Perspectives on AI Used in Homeless Services
abstract
Recent years have seen growing adoption of AI-based decision-support systems (ADS) in homeless services, yet we know little about stakeholder desires and concerns surrounding their use. In this work, we aim to understand impacted stakeholders’ perspectives on a deployed ADS that prioritizes scarce housing resources. We employed AI lifecycle comicboarding, an adapted version of the comicboarding method, to elicit stakeholder feedback and design ideas across various components of an AI system’s design. We elicited feedback from county workers who operate the ADS daily, service providers whose work is directly impacted by the ADS, and unhoused individuals in the region. Our participants shared concerns and design suggestions around the AI system’s overall objective, specific model design choices, dataset selection, and use in deployment. Our findings demonstrate that stakeholders, even without AI knowledge, can provide specific and critical feedback on an AI system’s design and deployment, if empowered to do so.
Tzu-Sheng Kuo, Hong Shen 0004, Jisoo Geum, Nev Jones, Jason I. Hong, Haiyi Zhu, Kenneth Holstein
CHI2
2023 Participation and Division of Labor in User-Driven Algorithm Audits: How Do Everyday Users Work together to Surface Algorithmic Harms?
abstract
Recent years have witnessed an interesting phenomenon in which users come together to interrogate potentially harmful algorithmic behaviors they encounter in their everyday lives. Researchers have started to develop theoretical and empirical understandings of these user-driven audits, with a hope to harness the power of users in detecting harmful machine behaviors. However, little is known about users’ participation and their division of labor in these audits, which are essential to support these collective efforts in the future. Through collecting and analyzing 17,984 tweets from four recent cases of user-driven audits, we shed light on patterns of users’ participation and engagement, especially with the top contributors in each case. We also identified the various roles users’ generated content played in these audits, including hypothesizing, data collection, amplification, contextualization, and escalation. We discuss implications for designing tools to support user-driven audits and users who labor to raise awareness of algorithm bias.
Rena Li, Sara Kingsley, Chelsea Fan, Proteeti Sinha, Nora Wai, Jaimie Lee, Hong Shen 0004, Motahhare Eslami, Jason I. Hong
CHI7
2022 Toward User-Driven Algorithm Auditing: Investigating users' strategies for uncovering harmful algorithmic behavior
abstract
Recent work in HCI suggests that users can be powerful in surfacing harmful algorithmic behaviors that formal auditing approaches fail to detect. However, it is not well understood how users are often able to be so effective, nor how we might support more effective user-driven auditing. To investigate, we conducted a series of think-aloud interviews, diary studies, and workshops, exploring how users find and make sense of harmful behaviors in algorithmic systems, both individually and collectively. Based on our findings, we present a process model capturing the dynamics of and influences on users’ search and sensemaking behaviors. We find that 1) users’ search strategies and interpretations are heavily guided by their personal experiences with and exposures to societal bias; and 2) collective sensemaking amongst multiple users is invaluable in user-driven algorithm audits. We offer directions for the design of future methods and tools that can better support user-driven auditing.
Alicia DeVos, Aditi Dhabalia, Hong Shen 0004, Kenneth Holstein, Motahhare Eslami
CHI3
2021 More Kawaii than a Real-Person Live Streamer: Understanding How the Otaku Community Engages with and Perceives Virtual YouTubers
abstract
Live streaming has become increasingly popular, with most streamers presenting their real-life appearance. However, Virtual YouTubers (VTubers), virtual 2D or 3D avatars that are voiced by humans, are emerging as live streamers and attracting a growing viewership in East Asia. Although prior research has found that many viewers seek real-life interpersonal interactions with real-person streamers, it is currently unknown what makes VTuber live streams engaging or how they are perceived differently than real-person streamers. We conducted an interview study to understand how viewers engage with VTubers and perceive the identities of the voice actors behind the avatars (i.e., Nakanohito). The data revealed that Virtual avatars bring unique performative opportunities which result in different viewer expectations and interpretations of VTuber behavior. Viewers intentionally upheld the disembodiment of VTuber avatars from their voice actors. We uncover the nuances in viewer perceptions and attitudes and further discuss the implications of VTuber practices to the understanding of live streaming in general.
Zhicong Lu, Chenxinran Shen, Jiannan Li, Hong Shen 0004, Daniel J. Wigdor
CHI4
2021 Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic Behaviors
abstract
A growing body of literature has proposed formal approaches to audit algorithmic systems for biased and harmful behaviors. While formal auditing approaches have been greatly impactful, they often suffer major blindspots, with critical issues surfacing only in the context of everyday use once systems are deployed. Recent years have seen many cases in which everyday users of algorithmic systems detect and raise awareness about harmful behaviors that they encounter in the course of their everyday interactions with these systems. However, to date little academic attention has been granted to these bottom-up, user-driven auditing processes. In this paper, we propose and explore the concept of everyday algorithm auditing, a process in which users detect, understand, and interrogate problematic machine behaviors via their day-to-day interactions with algorithmic systems. We argue that everyday users are powerful in surfacing problematic machine behaviors that may elude detection via more centrally-organized forms of auditing, regardless of users' knowledge about the underlying algorithms. We analyze several real-world cases of everyday algorithm auditing, drawing lessons from these cases for the design of future platforms and tools that facilitate such auditing behaviors. Finally, we discuss work that lies ahead, toward bridging the gaps between formal auditing approaches and the organic auditing behaviors that emerge in everyday use of algorithmic systems.
Hong Shen 0004, Alicia DeVos, Motahhare Eslami, Kenneth Holstein
Proc. ACM Hum. Comput. Interact.1
2021 Lean Privacy Review: Collecting Users' Privacy Concerns of Data Practices at a Low Cost
abstract
Today, industry practitioners (e.g., data scientists, developers, product managers) rely on formal privacy reviews (a combination of user interviews, privacy risk assessments, etc.) in identifying potential customer acceptance issues with their organization’s data practices. However, this process is slow and expensive, and practitioners often have to make ad-hoc privacy-related decisions with little actual feedback from users. We introduce Lean Privacy Review (LPR), a fast, cheap, and easy-to-access method to help practitioners collect direct feedback from users through the proxy of crowd workers in the early stages of design. LPR takes a proposed data practice, quickly breaks it down into smaller parts, generates a set of questionnaire surveys, solicits users’ opinions, and summarizes those opinions in a compact form for practitioners to use. By doing so, LPR can help uncover the range and magnitude of different privacy concerns actual people have at a small fraction of the cost and wait-time for a formal review. We evaluated LPR using 12 real-world data practices with 240 crowd users and 24 data practitioners. Our results show that (1) the discovery of privacy concerns saturates as the number of evaluators exceeds 14 participants, which takes around 5.5 hours to complete (i.e., latency) and costs 3.7 hours of total crowd work ( $80 in our experiments); and (2) LPR finds 89% of privacy concerns identified by data practitioners as well as 139% additional privacy concerns that practitioners are not aware of, at a 6% estimated false alarm rate.
Haojian Jin, Hong Shen 0004, Swarun Kumar, Jason I. Hong
ACM Trans. Comput. Hum. Interact.2
2020 'I Can't Even Buy Apples If I Don't Use Mobile Pay?': When Mobile Payments Become Infrastructural in China
abstract
Despite slow adoption in the US, mobile payments are thede facto solution for hundreds of millions of users in China for everything from paying bills to riding buses, from sending virtual "Red Packets'' to buying money-market funds. In this paper, we use the theoretical lens of infrastructure to study users' interactions with ubiquitous mobile payment systems in China, focusing on Alipay and WeChat Pay, the two dominant apps on the market. Based on data from a survey (n=466) and follow-up interviews (n=12) with users in China, we describe the diverse usage patterns across physical, social, and digital ubiquity, and a series of challenges people face. Reflecting on the lessons we learned from the Chinese case -- in particular, problems and pitfalls -- we discuss some implications both for design and for policy. Our findings have important implications for other countries that have been moving towards greater adoption of mobile payments.
Hong Shen 0004, Cori Faklaris, Haojian Jin, Laura A. Dabbish, Jason I. Hong
Proc. ACM Hum. Comput. Interact.1
2020 Designing Alternative Representations of Confusion Matrices to Support Non-Expert Public Understanding of Algorithm Performance
abstract
Ensuring effective public understanding of algorithmic decisions that are powered by machine learning techniques has become an urgent task with the increasing deployment of AI systems into our society. In this work, we present a concrete step toward this goal by redesigning confusion matrices for binary classification to support non-experts in understanding the performance of machine learning models. Through interviews (n=7) and a survey (n=102), we mapped out two major sets of challenges lay people have in understanding standard confusion matrices: the general terminologies and the matrix design. We further identified three sub-challenges regarding the matrix design, namely, confusion about the direction of reading the data, layered relations and quantities involved. We then conducted an online experiment with 483 participants to evaluate how effective a series of alternative representations target each of those challenges in the context of an algorithm for making recidivism predictions. We developed three levels of questions to evaluate users' objective understanding. We assessed the effectiveness of our alternatives for accuracy in answering those questions, completion time, and subjective understanding. Our results suggest that (1) only by contextualizing terminologies can we significantly improve users' understanding and (2) flow charts, which help point out the direction of reading the data, were most useful in improving objective understanding. Our findings set the stage for developing more intuitive and generally understandable representations of the performance of machine learning models.
Hong Shen 0004, Haojian Jin, Ángel Alexander Cabrera, Adam Perer, Haiyi Zhu, Jason I. Hong
Proc. ACM Hum. Comput. Interact.1