Jina Suh

dblp:151/6630 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0002-7646-5563ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 28 · 1 first-author · 20 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Worker Discretion Advised: Co-designing Risk Disclosure in Crowdsourced Responsible AI (RAI) Content Work
abstract
Responsible AI (RAI) content work, such as annotation, moderation, or red teaming for AI safety, often exposes crowd workers to potentially harmful content. While prior work has underscored the importance of communicating well-being risk to employed content moderators, designing effective disclosure mechanisms for crowd workers while balancing worker protection with the needs of task designers and platforms remains largely unexamined. To address this gap, we conducted individual co-design sessions with 15 task designers, 11 crowdworkers, and 3 platform representatives. We investigated task designer preferences for support in disclosing tasks, worker preferences for receiving risk disclosure warnings, and how platform representatives envision their role in shaping risk disclosure practices. We identify design tensions and map the sociotechnical tradeoffs that shape disclosure practices. We contribute design recommendations and feature concepts for risk disclosure mechanisms in the context of RAI content work.
Alice Qian Zhang, Ryland Shaw, Jina Suh, Laura A. Dabbish, Hong Shen 0004
CHI4
2026 Locating Risk: Task Designers and the Challenge of Risk Disclosure in Crowdsourced RAI Content Work CSCW029
abstract
As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highlighted the risks to worker well-being associated with RAI content work, far less attention has been paid to how these risks are communicated to workers by task designers or individuals who design and post RAI tasks. Existing transparency frameworks and guidelines, such as model cards, datasheets, and crowdworksheets, focus on documenting model information and dataset collection processes, but they overlook an important aspect of disclosing well-being risks to workers. In the absence of standard workflows or clear guidance, the consistent application of content warnings, consent flows, or other forms of well-being risk disclosure remains unclear. This study investigates how task designers approach risk disclosure in crowdsourced RAI tasks. Drawing on interviews with 23 task designers across academic and industry sectors, we examine how well-being risk is recognized, interpreted, and communicated in practice. Our findings highlight the need to support task designers in identifying and communicating risks not only to support crowdworker well-being but also to strengthen the ethical integrity and technical efficacy of AI development pipelines.
Alice Qian Zhang, Ryland Shaw, Laura A. Dabbish, Jina Suh, Hong Shen 0004
Proc. ACM Hum. Comput. Interact.4
2026 SENSE-7: Taxonomy and Dataset for Measuring User Perceptions of Empathy in Sustained Human-AI Conversations
abstract
Empathy is increasingly recognized as a key factor in human–AI communication, yet conventional approaches to “digital empathy” often focus on simulating internal, human like emotional states while overlooking the inherently subjective, contextual, and relational facets of empathy as perceived by users. In this work, we propose a human-centered taxonomy that emphasizes observable empathic behaviors and introduce a new dataset, SENSE-7, of real-world conversations between information workers and Large Language Models (LLMs), which includes per-turn empathy annotations directly from the users, along with user characteristics, and contextual details, offering a more user-grounded representation of empathy. Analysis of 695 conversations from 109 participants reveals that empathy judgments are highly individualized, context-sensitive, and vulnerable to disruption when conversational continuity fails or user expectations go unmet. To promote further research, we provide a subset of 672 anonymized conversation and provide exploratory classification analysis, showing that an LLM-based classifier can recognize 5 levels of empathy with an encouraging average Spearman ρ = 0.369 and Accuracy = 0.487 over this set. Overall, our findings underscore the need for AI designs that dynamically tailor empathic behaviors to user contexts and goals, offering a roadmap for future research and practical development of socially attuned, human-centered artificial agents.
Jina Suh, Lindy Le, Erfan Shayegani, Gonzalo A. Ramos, Judith Amores, Desmond C. Ong, Mary Czerwinski, Javier Hernandez
IEEE Trans. Affect. Comput.1
2025 AI on My Shoulder: Supporting Emotional Labor in Front-Office Roles with an LLM-based Empathetic Coworker
abstract
Client-Service Representatives (CSRs) are vital to organizations.Frequent interactions with disgruntled clients, however, disrupt their mental well-being.To help CSRs regulate their emotions while interacting with uncivil clients, we designed Care-Pilot, an LLM-powered assistant, and evaluated its efficacy, perception, and use.Our comparative analyses between 665 human and Care-Pilotgenerated support messages highlight Care-Pilot's ability to adapt to and demonstrate empathy in various incivility incidents.Additionally, 143 CSRs assessed Care-Pilot's empathy as more sincere and actionable than human messages.Finally, we interviewed 20 CSRs who interacted with Care-Pilot in a simulation exercise.They reported that Care-Pilot helped them avoid negative thinking, recenter thoughts, and humanize clients; showing potential for bridging gaps in coworker support.Yet, they also noted deployment challenges and emphasized the indispensability of shared experiences.We discuss future designs and societal implications of AI-mediated emotional labor, underscoring empathy as a critical function for AI assistants for worker mental health.
Vedant Das Swain, Qiuyue Joy Zhong, Jash Rajesh Parekh, Yechan Jeon, Roy Zimmermann, Mary Czerwinski, Jina Suh, Varun Mishra 0001, Koustuv Saha, Javier Hernandez
CHI7
2025 SCOPE: Examining Technology-Enhanced Collaborative Care Management of Depression in the Cancer Setting
abstract
Collaborative care management is an evidence-based approach to integrated psychosocial care for patients with comorbid cancer and depression. Prior work highlights challenges in patient-provider collaboration in navigating parallel cancer care and psychosocial care journeys of these patients. We design and deploy SCOPE , a platform for technology-enhanced collaborative care combining a patient-facing mobile app with a provider-facing registry. We examine SCOPE through a total of 45 interviews with patients and providers conducted in SCOPE 's 15 months of design and development and 24 Months of SCOPE 's deployment for actual care in 6 cancer clinics. We find that: (1) SCOPE supported patient engagement in its underlying collaborative care and behavioral activation interventions, (2) patient-generated data in SCOPE improved patient-provider collaboration between and within in-person sessions, (3) SCOPE supported providers in delivering care and improved care team collaboration, (4) experience with SCOPE created evolving expectations for collaboration around data, and (5) SCOPE 's deployment in actual care surfaced important implementation barriers. We discuss the implications of our findings in terms of designing for engagement with behavioral health interventions, negotiating patient data sharing and provider responsiveness, supporting personalized self-tracking goals in evidence-based interventions, exploring the role of digital health navigators in technology-enhanced care, and the need for flexibility in aligning technology-supported interventions to patient needs.
Anant Mittal, Tae Jones, Ravi Karkar, Jina Suh, Spencer Williams, Yihao Zheng 0004, Lydia M. Andris, Nicole Bates, Amy M. Bauer, Ty W. Lostuter, Jesse R. Fann, James Fogarty, Gary Hsieh
Proc. ACM Hum. Comput. Interact.4
2025 From User Surveys to Telemetry-Driven AI Agents: Exploring the Potential of Personalized Productivity Solutions
abstract
Information workers increasingly struggle with productivity challenges in modern workplaces, facing difficulties in managing time and effectively utilizing workplace analytics data for behavioral improvement. Despite the availability of productivity metrics through enterprise tools, workers often fail to translate this data into actionable insights. We present a comprehensive, user-centric approach to address these challenges through AI-based productivity agents tailored to users' needs. Utilizing a two-phase method, we first conducted a survey with 363 participants, exploring various aspects of productivity, communication style, agent approach, personality traits, personalization, and privacy. Drawing on the survey insights, we developed a GPT-4 powered personalized productivity agent that utilizes telemetry data gathered via Viva Insights from information workers to provide tailored assistance. We compared its performance with alternative productivity-assistive tools, such as dashboard and narrative, in a study involving 40 participants. Our findings highlight the importance of user-centric design, adaptability, and the balance between personalization and privacy in AI-assisted productivity tools. By building on these insights, our work provides important guidance for developing more effective productivity solutions, ultimately leading to optimized efficiency and user experiences for information workers.
Subigya Nepal, Javier Hernandez, Talie Massachi, Kael Rowan, Judith Amores, Jina Suh, Gonzalo A. Ramos, Brian Houck, Shamsi T. Iqbal, Mary Czerwinski
Proc. ACM Hum. Comput. Interact.6
2025 AURA: Amplifying Understanding, Resilience, and Awareness for Responsible AI Content Work
abstract
Behind the scenes of maintaining the safety of technology products from harmful and illegal digital content lies unrecognized human labor. The recent rise in the use of generative AI technologies and the accelerating demands to meet responsible AI (RAI) aims necessitates an increased focus on the labor behind such efforts in the age of AI. This study investigates the nature and challenges of content work that supports RAI efforts, or "RAI content work," that spans content moderation, data labeling, and red teaming -- through the lived experiences of content workers. We conduct a formative survey and semi-structured interview studies to develop a conceptualization of RAI content work and a subsequent framework of recommendations for providing holistic support for content workers. We validate our recommendations through a series of workshops with content workers and derive considerations for and examples of implementing such recommendations. We discuss how our framework may guide future innovation to support the well-being and professional development of the RAI content workforce.
Alice Qian Zhang, Judith Amores, Hong Shen 0004, Mary Czerwinski, Mary L. Gray, Jina Suh
Proc. ACM Hum. Comput. Interact.6
2025 Triple Peak Day: Work Rhythms of Software Developers in Hybrid Work
abstract
The future of work is rapidly changing, with remote and hybrid settings blurring the boundaries between professional and personal life. To understand how work rhythms vary across different work settings, we conducted a month-long study of 65 software developers, collecting anonymized computer activity data as well as daily ratings for perceived stress, productivity, and work setting. In addition to confirming the double-peak pattern of activity at 10:00 am and 2:00 pm observed in prior research, we observed a significant third peak around 9:00 pm. This third peak was associated with higher perceived productivity during remote days but increased stress during onsite and hybrid days, highlighting a nuanced interplay between work demands and work settings. Additionally, we found strong correlations between computer activity, productivity, and stress, including an inverted U-shaped relationship where productivity peaked at around six hours of computer activity before declining on more active days. These findings provide new insights into evolving work rhythms and highlight the impact of different work settings on productivity and stress.
Javier Hernandez, Vedant Das Swain, Jina Suh, Daniel McDuff, Judith Amores, Gonzalo A. Ramos, Kael Rowan, Brian Houck, Shamsi T. Iqbal, Mary Czerwinski
IEEE Trans. Software Eng.3
2024 Large Language Models Produce Responses Perceived to be Empathic
abstract
Large Language Models (LLMs) have demonstrated surprising performance on many tasks, including writing supportive messages that display empathy. Here, we had these models generate empathic messages in response to posts describing common life experiences, such as workplace situations, parenting, relationships, and other anxiety- and anger-eliciting situations. Across two studies (N=192, 202), we showed human raters a variety of responses written by several models (GPT4 Turbo, Llama2, and Mistral), and had people rate these responses on how empathic they seemed to be. We found that LLM-generated responses were consistently rated as more empathic than human-written responses. Linguistic analyses also show that these models write in distinct, predictable “styles”, in terms of their use of punctuation, emojis, and certain words. These results highlight the potential of using LLMs to enhance human peer support in contexts where empathy is important.
Yoon Kyung Lee, Jina Suh, Hongli Zhan, Junyi Jessy Li, Desmond C. Ong
ACII2
2024 IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction
abstract
Navigating certain communication situations can be challenging due to individuals' lack of skills and the interference of strong emotions.However, effective learning opportunities are rarely accessible.In this work, we conduct a human-centered study that uses language models to simulate bespoke communication training and provide just-in-time feedback to support the practice and learning of interpersonal effectiveness skills.We apply the interpersonal effectiveness framework from Dialectical Behavioral Therapy (DBT), DEAR MAN, which focuses on both conversational and emotional skills.We present IMBUE, an interactive training system that provides feedback 25% more similar to experts' feedback, compared to that generated by GPT-4.IMBUE is the first to focus on communication skills and emotion management simultaneously, incorporate experts' domain knowledge in providing feedback, and be grounded in psychology theory.Through a randomized trial of 86 participants, we find that IMBUE's simulation-only variant significantly improves participants' self-efficacy (up to 17%) and reduces negative emotions (up to 25%).With IMBUE's additional just-in-time feedback, participants demonstrate 17% improvement in skill mastery, along with greater enhancements in self-efficacy (27% more) and reduction of negative emotions (16% more) compared to simulation-only.The improvement in skill mastery is the only measure that is transferred to new and more difficult situations; situation-specific training is necessary for improving self-efficacy and emotion reduction.
Inna Wanyin Lin, Ashish Sharma 0004, Christopher Michael Rytting, Adam S. Miner, Jina Suh, Tim Althoff
ACL (1)5
2024 DISCERN: Designing Decision Support Interfaces to Investigate the Complexities of Workplace Social Decision-Making With Line Managers
abstract
Line managers form the first level of management in organizations, and must make complex decisions, while maintaining relationships with those impacted by their decisions. Amidst growing interest in technology-supported decision-making at work, their needs remain understudied. Further, most existing design knowledge for supporting social decision-making comes from domains where decision-makers are more socially detached from those they decide for. We conducted iterative design research with line managers within a technology organization, investigating decision-making practices, and opportunities for technological support. Through formative research, development of a decision-representation tool—DISCERN—and user enactments, we identify their communication and analysis needs that lack adequate support. We found they preferred tools for externalizing reasoning rather than tools that replace interpersonal interactions, and they wanted tools to support a range of intuitive and calculative decision-making. We discuss how design of social decision-making supports, especially in the workplace, can more explicitly support highly interactional social decision-making.
Pranav Khadpe, Lindy Le, Kate Nowak, Shamsi T. Iqbal, Jina Suh
CHI5
2024 Towards Inclusive Futures for Worker Wellbeing
abstract
The global COVID-19 pandemic has spurred on new collaborations across borders, and emphasized the importance of supporting wellbeing in the workplace, whether that workplace is hybrid, remote, or in-person. Work in CSCW, HCI, and organizational psychology has explored how people come to understand their wellbeing at work, and the role of identity, culture, and organizational factors in that process. In this study, we build on this past research and explore the importance of these factors when designing tools that support worker wellbeing for location-independent teams. We ask the question: how did organizational, cultural, and individual factors influence how workers understood their workplace wellbeing needs during the move to remote work? To investigate this question, we conduct a large scale linguistic analysis of 13,265 diary entries collected between 2020 - 2022, and complement it with in-depth interviews with 26 global employees, exploring intersections between technology, context, and wellbeing needs. We utilize this data to analyze the broader human infrastructure supporting hybrid and remote work, demonstrating how ideas around wellbeing are influenced by the (often technology-mediated) environment around both information and essential workers, and power differentials within it. Building on our findings, we provide recommendations for how technology design can better support more diverse and inclusive forms of worker wellbeing.
Sachin R. Pendse, Talie Massachi, Jalehsadat Mahdavimoghaddam, Jenna L. Butler, Jina Suh, Mary Czerwinski
Proc. ACM Hum. Comput. Interact.5
2024 Improving Work-Nonwork Balance with Data-Driven Implementation Intention and Mental Contrasting
abstract
Work-nonwork balance is an important aspect of workplace well-being with associations to improved physical and mental health, job performance, and quality of life. However, realizing work-nonwork balance goals is challenging due to competing demands and limited resources within organizational and interpersonal contexts. These challenges are compounded by technologies that blur the boundaries of work and nonwork in the always-on work cultures. At an individual level, such challenges can be subsided through the effective application of self-regulation techniques, such as implementation intentions and mental contrasting (IIMC). Further supporting these techniques through reflection on personal data, we implement the idea of data-driven IIMC into a self-tracking and behavior planning system and evaluate it in a three-week between-participant study with 43 information workers who used our system for improving work-nonwork balance. We find evidence that reflection on personal data improves awareness of behavior plan compliance and rescheduling, which are important in realizing work-nonwork balance goals. We also observe the value of micro-reflection, reflection on limited data of the very recent past, for IIMC. Our findings highlight opportunities for automation in data collection and sense-making and for further exploring the role of data-driven IIMC as boundary negotiating artifacts in support of work-nonwork balance goals.
Yasaman S. Sefidgar, Matthew Jörke, Jina Suh, Koustuv Saha, Shamsi T. Iqbal, Gonzalo A. Ramos, Mary Czerwinski
Proc. ACM Hum. Comput. Interact.3
2024 "I Want It That Way": Enabling Interactive Decision Support Using Large Language Models and Constraint Programming
abstract
A critical factor in the success of many decision support systems is the accurate modeling of user preferences. Psychology research has demonstrated that users often develop their preferences during the elicitation process, highlighting the pivotal role of system-user interaction in developing personalized systems. This paper introduces a novel approach, combining Large Language Models (LLMs) with Constraint Programming to facilitate interactive decision support. We study this hybrid framework through the lens of meeting scheduling, a time-consuming daily activity faced by a multitude of information workers. We conduct three studies to evaluate the novel framework, including a diary study to characterize contextual scheduling preferences, a quantitative evaluation of the system’s performance, and a user study to elicit insights with a technology probe that encapsulates our framework. Our work highlights the potential for a hybrid LLM and optimization approach for iterative preference elicitation, and suggests design considerations for building systems that support human-system collaborative decision-making processes.
Connor Lawless, Jakob Schöffer, Lindy Le, Kael Rowan, Shilad Sen, Cristina St. Hill, Jina Suh, Bahareh Sarrafzadeh
ACM Trans. Interact. Intell. Syst.7
2023 Do You Even Need Sensors?: Synthetic Biomusic as an Empathic Technology
abstract
Previous research suggests that biomusic, a type of biosignal sharing, is effective at promoting empathy and closeness among individuals. However, it is unclear whether these effects are due to the information it encodes or other emotional aspects of its resulting music. To explore this question, we developed a Generative Adversarial Network (GAN) to create synthetic biomusic that approximates real biomusic, and employed deception to evaluate its effects on 24 pairs of participants engaged in real-time emotional disclosure. Users reported that both real and synthetic biomusic provided the same amount of information about their conversational partner as observing body language, facial expressions, or vocal tone. Further, both conditions increased users’ ratings of closeness and empathy with each other compared to listening to no music. However, we found no statistically significant differences between the two biomusic conditions across any of our metrics. We discuss the implications of these results for the design of future biomusic systems.
Daway Chou-Ren, Mike Winters, Javier Hernandez, Daniel McDuff, Jina Suh, Vanessa Rodriguez, Gonzalo A. Ramos, Mary Czerwinski
ACII5
2023 Focused Time Saves Nine: Evaluating Computer-Assisted Protected Time for Hybrid Information Work
abstract
Information workers often struggle to balance their time for a variety of activities like focused work, communication, and caring. This study analyzes the impact of a commercially available computer-assisted time protection intervention that automatically and preemptively schedules calendar time for self-determined activities. We analyzed the behaviors and self-reports of workers in two naturalistic studies. First, we studied 27 workers who were already using Computer-Assisted Protected Time (CAP time) and found that they mainly used it for focused work. Second, we analyzed the effect of CAP time as a randomized intervention on 89 workers who never had CAP time and found that those with it self-reported an increase in performance, job resources, and immersion. In both studies, workers with CAP time exhibited a rearrangement of activities leading to an overall reduction in work activity. This study highlights new opportunities for intelligent time-management interventions and the importance of protected time at work.
Vedant Das Swain, Javier Hernandez, Brian Houck, Koustuv Saha, Jina Suh, Ahad Chaudhry, Tenny Cho, Wendy Guo, Shamsi T. Iqbal, Mary Czerwinski
CHI5
2023 Pearl: A Technology Probe for Machine-Assisted Reflection on Personal Data
abstract
Reflection on one’s personal data can be an effective tool for supporting wellbeing. However, current wellbeing reflection support tools tend to offer a one-size-fits-all approach, ignoring the diversity of people’s wellbeing goals and their agency in the self-reflection process. In this work, we identify an opportunity to help people work toward their wellbeing goals by empowering them to reflect on their data on their own terms. Through a formative study, we inform the design and implementation of Pearl, a workplace wellbeing reflection support tool that allows users to explore their personal data in relation to their wellbeing goal. Pearl is a calendar-based interactive machine teaching system that allows users to visualize data sources and tag regions of interest on their calendar. In return, the system provides insights about these tags that can be saved to a reflection journal. We used Pearl as a technology probe with 12 participants without data science expertise and found that all participants successfully gained insights into their workplace wellbeing. In our analysis, we discuss how Pearl’s capabilities facilitate insights, the role of machine assistance in the self-reflection process, and the data sources that participants found most insightful. We conclude with design dimensions for intelligent reflection support systems as inspiration for future work.
Matthew Jörke, Yasaman S. Sefidgar, Talie Massachi, Jina Suh, Gonzalo A. Ramos
IUI4
2023 Sensing Wellbeing in the Workplace, Why and For Whom? Envisioning Impacts with Organizational Stakeholders
abstract
With the heightened digitization of the workplace, alongside the rise of remote and hybrid work prompted by the pandemic, there is growing corporate interest in using passive sensing technologies for workplace wellbeing. Existing research on these technologies often focus on understanding or improving interactions between an individual user and the technology. Workplace settings can, however, introduce a range of complexities that challenge the potential impact and in-practice desirability of wellbeing sensing technologies. Today, there is an inadequate empirical understanding of how everyday workers---including those who are impacted by, and impact the deployment of workplace technologies--envision its broader socio-ecological impacts. In this study, we conduct storyboard-driven interviews with 33 participants across three stakeholder groups: organizational governors, AI builders, and worker data subjects. Overall, our findings surface how workers envisioned wellbeing sensing technologies may lead to cascading impacts on their broader organizational culture, interpersonal relationships with colleagues, and individual day-to-day lives. Participants anticipated harms arising from ambiguity and misalignment around scaled notions of "worker wellbeing,'' underlying technical limitations to workplace-situated sensing, and assumptions regarding how social structures and relationships may shape the impacts and use of these technologies. Based on our findings, we discuss implications for designing worker-centered data-driven wellbeing technologies.
Anna Kawakami, Shreya Chowdhary, Shamsi T. Iqbal, Qingzi Vera Liao, Alexandra Olteanu, Jina Suh, Koustuv Saha
Proc. ACM Hum. Comput. Interact.6
2022 Advancing the Understanding and Measurement of Workplace Stress in Remote Information Workers from Passive Sensors and Behavioral Data
abstract
Workplace stress has been increasing in recent decades and has worsened by the unique demands imposed by COVID-19 and the new remote/hybrid work settings. High-stress working conditions can be detrimental to the health and wellness of workers and can lead to significant business costs in terms of productivity loss and medical expenses. An essential step toward managing stress involves finding comfortable ways to sense workers and recognizing stress as soon as it happens. This work explores the potential value of using pervasive sensors such as keyboards, webcams, and behavioral data such as calendar and e-mail activity to passively assess individual stress levels of work in real-life. In particular, we collected a large corpus of such data from 46 remote information workers over one month and asked them to self-report their stress levels and other relevant factors several times a day. Analysis of the data demonstrates that passive sensors can effectively detect both triggers and manifestations of workplace stress and that having access to prior data of the worker is critical for developing well-performing stress recognition models. Furthermore, we provide qualitative feedback capturing workers' preferences in workplace stress monitoring.
Mehrab Bin Morshed, Javier Hernandez, Daniel McDuff, Jina Suh, Esther Howe, Kael Rowan, Marah Ihab Abdin, Gonzalo A. Ramos, Tracy Tran, Mary Czerwinski
ACII4
2022 Design of Digital Workplace Stress-Reduction Intervention Systems: Effects of Intervention Type and Timing
abstract
Workplace stress-reduction interventions have produced mixed results due to engagement and adherence barriers. Leveraging technology to integrate such interventions into the workday may address these barriers and help mitigate the mental, physical, and monetary effects of workplace stress. To inform the design of a workplace stress-reduction intervention system, we conducted a four-week longitudinal study with 86 participants, examining the effects of intervention type and timing on usage, stress reduction impact, and user preferences. We compared three intervention types and two delivery timing conditions: Pre-scheduled (PS) by users and Just-in-time (JIT) prompted by the system-identified user stress-levels. We found JIT participants completed significantly more interventions than PS participants, but post-intervention and study-long stress reduction was not significantly different between conditions. Participants rated low-effort interventions highest, but high-effort interventions reduced the most stress. Participants felt JIT provided accountability but desired partial agency over timing. We present type and timing implications.
Esther Howe, Jina Suh, Mehrab Bin Morshed, Daniel McDuff, Kael Rowan, Javier Hernandez, Marah Ihab Abdin, Gonzalo A. Ramos, Tracy Tran, Mary Czerwinski
CHI2
2022 ForSense: Accelerating Online Research Through Sensemaking Integration and Machine Research Support
abstract
Online research is a frequent and important activity people perform on the Internet, yet current support for this task is basic, fragmented and not well integrated into web browser experiences. Guided by sensemaking theory, we present ForSense, a browser extension for accelerating people’s online research experience. The two primary sources of novelty of ForSense are the integration of multiple stages of online research and providing machine assistance to the user by leveraging recent advances in neural-driven machine reading. We use ForSense as a design probe to explore (1) the benefits of integrating multiple stages of online research, (2) the opportunities to accelerate online research using current advances in machine reading, (3) the opportunities to support online research tasks in the presence of imprecise machine suggestions, and (4) insights about the behaviors people exhibit when performing online research, the pages they visit, and the artifacts they create. Through our design probe, we observe people performing online research tasks, and see that they benefit from ForSense’s integration and machine support for online research. From the information and insights we collected, we derive and share key recommendations for designing and supporting imprecise machine assistance for research tasks.
Gonzalo A. Ramos, Napol Rachatasumrit, Jina Suh, Rachel Ng, Christopher Meek
ACM Trans. Interact. Intell. Syst.3
2021 Guidelines for Assessing and Minimizing Risks of Emotion Recognition Applications
abstract
Society has witnessed a rapid increase in the adoption of commercial uses of emotion recognition. Tools that were traditionally used by domain experts are now being used by individuals who are often unaware of the technology’s limitations and may use them in potentially harmful settings. The change in scale and agency, paired with gaps in regulation, urge the research community to rethink how we design, position, implement and ultimately deploy emotion recognition to anticipate and minimize potential risks. To help understand the current ecosystem of applied emotion recognition, this work provides an overview of some of the most frequent commercial applications and identifies some of the potential sources of harm. Informed by these, we then propose 12 guidelines for systematically assessing and reducing the risks presented by emotion recognition applications. These guidelines can help identify potential misuses and inform future deployments of emotion recognition.
Javier Hernandez, Josh Lovejoy, Daniel McDuff, Jina Suh, Tim O'Brien, Arathi Sethumadhavan, Gretchen Greene, Rosalind W. Picard, Mary Czerwinski
ACII4
2021 AffectiveSpotlight: Facilitating the Communication of Affective Responses from Audience Members during Online Presentations
abstract
The ability to monitor audience reactions is critical when delivering presentations. However, current videoconferencing platforms offer limited solutions to support this. This work leverages recent advances in affect sensing to capture and facilitate communication of relevant audience signals. Using an exploratory survey (N=175), we assessed the most relevant audience responses such as confusion, engagement, and head-nods. We then implemented AffectiveSpotlight, a Microsoft Teams bot that analyzes facial responses and head gestures of audience members and dynamically spotlights the most expressive ones. In a within-subjects study with 14 groups (N=117), we observed that the system made presenters significantly more aware of their audience, speak for a longer period of time, and self-assess the quality of their talk more similarly to the audience members, compared to two control conditions (randomly-selected spotlight and default platform UI). We provide design recommendations for future affective interfaces for online presentations based on feedback from the study.
Prasanth Murali, Javier Hernandez, Daniel McDuff, Kael Rowan, Jina Suh, Mary Czerwinski
CHI5
2021 MeetingCoach: An Intelligent Dashboard for Supporting Effective & Inclusive Meetings
abstract
Video-conferencing is essential for many companies, but its limitations in conveying social cues can lead to ineffective meetings. We present MeetingCoach, an intelligent post-meeting feedback dashboard that summarizes contextual and behavioral meeting information. Through an exploratory survey (N=120), we identified important signals (e.g., turn taking, sentiment) and used these insights to create a wireframe dashboard. The design was evaluated with in situ participants (N=16) who helped identify the components they would prefer in a post-meeting dashboard. After recording video-conferencing meetings of eight teams over four weeks, we developed an AI system to quantify the meeting features and created personalized dashboards for each participant. Through interviews and surveys (N=23), we found that reviewing the dashboard helped improve attendees’ awareness of meeting dynamics, with implications for improved effectiveness and inclusivity. Based on our findings, we provide suggestions for future feedback system designs of video-conferencing meetings.
Samiha Samrose, Daniel McDuff, Robert Sim, Jina Suh, Kael Rowan, Javier Hernandez, Sean Rintel, Kevin Moynihan, Mary Czerwinski
CHI4
2021 ForSense: Accelerating Online Research Through Sensemaking Integration and Machine Research Support
abstract
Online research is a frequent and important activity people perform on the Internet, yet current support for this task is basic, fragmented and not well integrated into web browser experiences. Guided by sensemaking theory, we present ForSense, a browser extension for accelerating people’s online research experience. The two primary sources of novelty of ForSense are the integration of multiple stages of online research and providing machine assistance to the user by leveraging recent advances in neural-driven machine reading. We use ForSense as a design probe to explore (1) the benefits of integrating multiple stages of online research, (2) the opportunities to accelerate online research using current advances in machine reading, and (3) the opportunities to support online research tasks under the presence of imprecise machine suggestions. In our study, we observe people performing online research tasks, and see that they benefit from ForSense’s integration and machine support for online research. From our study, we derive and share key recommendations for designing and supporting imprecise machine assistance for research tasks.
Napol Rachatasumrit, Gonzalo A. Ramos, Jina Suh, Rachel Ng, Christopher Meek
IUI3
2021 Population-Scale Study of Human Needs During the COVID-19 Pandemic: Analysis and Implications
abstract
Most work to date on mitigating the COVID-19 pandemic is focused urgently on biomedicine and epidemiology. Yet, pandemic-related policy decisions cannot be made on health information alone. Decisions need to consider the broader impacts on people and their needs. Quantifying human needs across the population is challenging as it requires high geo-temporal granularity, high coverage across the population, and appropriate adjustment for seasonal and other external effects. Here, we propose a computational methodology, building on Maslow's hierarchy of needs, that can capture a holistic view of relative changes in needs following the pandemic through a difference-in-differences approach that corrects for seasonality and volume variations. We apply this approach to characterize changes in human needs across physiological, socioeconomic, and psychological realms in the US, based on more than 35 billion search interactions spanning over 36,000 ZIP codes over a period of 14 months. The analyses reveal that the expression of basic human needs has increased exponentially while higher-level aspirations declined during the pandemic in comparison to the pre-pandemic period. In exploring the timing and variations in statewide policies, we find that the durations of shelter-in-place mandates have influenced social and emotional needs significantly. We demonstrate that potential barriers to addressing critical needs, such as support for unemployment and domestic violence, can be identified through web search interactions. Our approach and results suggest that population-scale monitoring of shifts in human needs can inform policies and recovery efforts for current and anticipated needs.
Jina Suh, Eric Horvitz, Ryen W. White, Tim Althoff
WSDM1
2020 Understanding and Supporting Knowledge Decomposition for Machine Teaching
abstract
Machine teaching (MT) is an emerging field that studies non-machine learning (ML) experts incrementally building semantic ML models in efficient ways. While MT focuses on the types of knowledge a human teacher provides a machine learner, not much is known about how people perform or can be supported in this essential task of identifying and expressing useful knowledge. We refer to this process as knowledge decomposition. To address the challenges of this type of Human-AI collaboration, we seek to build foundational frameworks for understanding and supporting knowledge decomposition. We present results of a study investigating what types of knowledge people teach, what cognitive processes they use, and what challenges they encounter when teaching a learner to classify text documents. From our observations, we introduce design opportunities for new tools to support knowledge decomposition. Our findings carry implications for applying the benefits of knowledge decomposition to MT and ML.
Felicia Ng, Jina Suh, Gonzalo A. Ramos
Conference on Designing Interactive Systems2
2020 Design and evaluation of intelligent agent prototypes for assistance with focus and productivity at work
abstract
Current research on building intelligent agents for aiding with productivity and focus in the workplace is quite limited, despite the ubiquity of information workers across the globe. In our work, we present a productivity agent which helps users schedule and block out time on their calendar to focus on important tasks, monitor and intervene with distractions, and reflect on their daily mood and goals in a single, standalone application. We created two different prototype versions of our agent: a text-based (TB) agent with a similar UI to a standard chatbot, and a more emotionally expressive virtual agent (VA) that employs a video avatar and the ability to detect and respond appropriately to users' emotions. We evaluated these two agent prototypes against an existing product (control) condition through a three-week, within subjects study design with 40 participants, across different work roles in a large organization. We found that participants scheduled 134% more time with the TB prototype, and 110% more time with the VA prototype for focused tasks compared to the control condition. Users reported that they felt more satisfied and productive with the VA agent. However, The perception of anthropomorphism in the VA was polarized, with several participants suggesting that the human appearance was unnecessary. We discuss important insights from our work for the future design of conversational agents for productivity, wellbeing, and focus in the workplace.
Ted Grover, Kael Rowan, Jina Suh, Daniel McDuff, Mary Czerwinski
IUI3
2020 Interactive machine teaching: a human-centered approach to building machine-learned models
abstract
Modern systems can augment people’s capabilities by using machine-learned models to surface intelligent behaviors. Unfortunately, building these models remains challenging and beyond the reach of non-machine learning experts. We describe interactive machine teaching (IMT) and its potential to simplify the creation of machine-learned models. One of the key characteristics of IMT is its iterative process in which the human-in-the-loop takes the role of a teacher teaching a machine how to perform a task. We explore alternative learning theories as potential theoretical foundations for IMT, the intrinsic human capabilities related to teaching, and how IMT systems might leverage them. We argue that IMT processes that enable people to leverage these capabilities have a variety of benefits, including making machine learning methods accessible to subject-matter experts and the creation of semantic and debuggable machine learning (ML) models. We present an integrated teaching environment (ITE) that embodies principles from IMT, and use it as a design probe to observe how non-ML experts do IMT and as the basis of a system that helps us study how to guide teachers. We explore and highlight the benefits and challenges of IMT systems. We conclude by outlining six research challenges to advance the field of IMT.
Gonzalo A. Ramos, Christopher Meek, Patrice Y. Simard, Jina Suh, Soroush Ghorashi
Hum. Comput. Interact.4
2020 Parallel Journeys of Patients with Cancer and Depression: Challenges and Opportunities for Technology-Enabled Collaborative Care
abstract
Depression is common but under-treated in patients with cancer, despite being a major modifiable contributor to morbidity and early mortality. Integrating psychosocial care into cancer services through the team-based Collaborative Care Management (CoCM) model has been proven to be effective in improving patient outcomes in cancer centers. However, there is currently a gap in understanding the challenges that patients and their care team encounter in managing co-morbid cancer and depression in integrated psycho-oncology care settings. Our formative study examines the challenges and needs of CoCM in cancer settings with perspectives from patients, care managers, oncologists, psychiatrists, and administrators, with a focus on technology opportunities to support CoCM. We find that: (1) patients with co-morbid cancer and depression struggle to navigate between their cancer and psychosocial care journeys, and (2) conceptualizing co-morbidities as separate and independent care journeys is insufficient for characterizing this complex care context. We then propose the parallel journeys framework as a conceptual design framework for characterizing challenges that patients and their care team encounter when cancer and psychosocial care journeys interact. We use the challenges discovered through the lens of this framework to highlight and prioritize technology design opportunities for supporting whole-person care for patients with co-morbid cancer and depression.
Jina Suh, Spencer Williams, Jesse R. Fann, James Fogarty, Amy M. Bauer, Gary Hsieh
Proc. ACM Hum. Comput. Interact.1
2020 AnchorViz: Facilitating Semantic Data Exploration and Concept Discovery for Interactive Machine Learning
abstract
When building a classifier in interactive machine learning (iML), human knowledge about the target class can be a powerful reference to make the classifier robust to unseen items. The main challenge lies in finding unlabeled items that can either help discover or refine concepts for which the current classifier has no corresponding features (i.e., it has feature blindness ). Yet it is unrealistic to ask humans to come up with an exhaustive list of items, especially for rare concepts that are hard to recall. This article presents AnchorViz , an interactive visualization that facilitates the discovery of prediction errors and previously unseen concepts through human-driven semantic data exploration. By creating example-based or dictionary-based anchors representing concepts, users create a topology that (a) spreads data based on their similarity to the concepts and (b) surfaces the prediction and label inconsistencies between data points that are semantically related. Once such inconsistencies and errors are discovered, users can encode the new information as labels or features and interact with the retrained classifier to validate their actions in an iterative loop. We evaluated AnchorViz through two user studies. Our results show that AnchorViz helps users discover more prediction errors than stratified random and uncertainty sampling methods. Furthermore, during the beginning stages of a training task, an iML tool with AnchorViz can help users build classifiers comparable to the ones built with the same tool with uncertainty sampling and keyword search, but with fewer labels and more generalizable features. We discuss exploration strategies observed during the two studies and how AnchorViz supports discovering, labeling, and refining of concepts through a sensemaking loop.
Jina Suh, Soroush Ghorashi, Gonzalo A. Ramos, Nan-Chen Chen, Steven Mark Drucker, Johan Verwey, Patrice Y. Simard
ACM Trans. Interact. Intell. Syst.1
2019 Guidelines for Human-AI Interaction
abstract
Advances in artificial intelligence (AI) frame opportunities and challenges for user interface design. Principles for human-AI interaction have been discussed in the human-computer interaction community for over two decades, but more study and innovation are needed in light of advances in AI and the growing uses of AI technologies in human-facing applications. We propose 18 generally applicable design guidelines for human-AI interaction. These guidelines are validated through multiple rounds of evaluation including a user study with 49 design practitioners who tested the guidelines against 20 popular AI-infused products. The results verify the relevance of the guidelines over a spectrum of interaction scenarios and reveal gaps in our knowledge, highlighting opportunities for further research. Based on the evaluations, we believe the set of design guidelines can serve as a resource to practitioners working on the design of applications and features that harness AI technologies, and to researchers interested in the further development of human-AI interaction design principles.
Saleema Amershi, Daniel S. Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi T. Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, Eric Horvitz
CHI7
2018 Grounding Interactive Machine Learning Tool Design in How Non-Experts Actually Build Models
abstract
Machine learning (ML) promises data-driven insights and solutions for people from all walks of life, but the skill of crafting these solutions is possessed by only a few. Emerging research addresses this issue by creating ML tools that are easy and accessible to people who are not formally trained in ML (non-experts). This work investigated how non-experts build ML solutions for themselves in real life. Our interviews and surveys revealed unique potentials of non-expert ML, as well several pitfalls that non-experts are susceptible to. For example, many perceived percentage accuracy as a sole measure of performance, thus problematic models proceeded to deployment. These observations suggested that, while challenging, making ML easy and robust should both be important goals of designing novice-facing ML tools. To advance on this insight, we discuss design implications and created a sensitizing concept to demonstrate how designers might guide non-experts to easily build robust solutions.
Qian Yang 0004, Jina Suh, Nan-Chen Chen, Gonzalo A. Ramos
Conference on Designing Interactive Systems2
2018 AnchorViz: Facilitating Classifier Error Discovery through Interactive Semantic Data Exploration
abstract
When building a classifier in interactive machine learning, human knowledge about the target class can be a powerful reference to make the classifier robust to unseen items. The main challenge lies in finding unlabeled items that can either help discover or refine concepts for which the current classifier has no corresponding features (i.e., it has feature blindness). Yet it is unrealistic to ask humans to come up with an exhaustive list of items, especially for rare concepts that are hard to recall. This paper presents AnchorViz, an interactive visualization that facilitates error discovery through semantic data exploration. By creating example-based anchors, users create a topology to spread data based on their similarity to the anchors and examine the inconsistencies between data points that are semantically related. The results from our user study show that AnchorViz helps users discover more prediction errors than stratified random and uncertainty sampling methods.
Nan-Chen Chen, Jina Suh, Johan Verwey, Gonzalo A. Ramos, Steven Mark Drucker, Patrice Y. Simard
IUI2
2018 Using Machine Learning to Support Qualitative Coding in Social Science: Shifting the Focus to Ambiguity
abstract
Machine learning (ML) has become increasingly influential to human society, yet the primary advancements and applications of ML are driven by research in only a few computational disciplines. Even applications that affect or analyze human behaviors and social structures are often developed with limited input from experts outside of computational fields. Social scientists—experts trained to examine and explain the complexity of human behavior and interactions in the world—have considerable expertise to contribute to the development of ML applications for human-generated data, and their analytic practices could benefit from more human-centered ML methods. Although a few researchers have highlighted some gaps between ML and social sciences [51, 57, 70], most discussions only focus on quantitative methods. Yet many social science disciplines rely heavily on qualitative methods to distill patterns that are challenging to discover through quantitative data. One common analysis method for qualitative data is qualitative coding . In this article, we highlight three challenges of applying ML to qualitative coding. Additionally, we utilize our experience of designing a visual analytics tool for collaborative qualitative coding to demonstrate the potential in using ML to support qualitative coding by shifting the focus to identifying ambiguity. We illustrate dimensions of ambiguity and discuss the relationship between disagreement and ambiguity. Finally, we propose three research directions to ground ML applications for social science as part of the progression toward human-centered machine learning.
Nan-Chen Chen, Margaret Drouhard, Rafal Kocielnik, Jina Suh, Cecilia R. Aragon
ACM Trans. Interact. Intell. Syst.4
2017 Aeonium: Visual analytics to support collaborative qualitative coding
abstract
Qualitative coding offers the potential to obtain deep insights into social media, but the technique can be inconsistent and hard to scale. Researchers using qualitative coding impose structure on unstructured data through “codes” that represent categories for analysis. Our visual analytics interface, Aeonium, supports human insight in collaborative coding through visual overviews of codes assigned by multiple researchers and distributions of important keywords and codes. The underlying machine learning model highlights ambiguity and inconsistency. Our goal was not to reduce qualitative coding to a machine-solvable problem, but rather to bolster human understanding gained from coding and reinterpreting the data collaboratively. We conducted an experimental study with 39 participants who coded tweets using our interface. In addition to increased understanding of the topic, participants reported that Aeonium's collaborative coding functionality helped them reflect on their own interpretations. Feedback from participants demonstrates that visual analytics can help facilitate rich qualitative analysis and suggests design implications for future exploration.
Margaret Drouhard, Nan-Chen Chen, Jina Suh, Rafal Kocielnik, Vanessa Peña Araya, Keting Cen, Xiangyi Zheng, Cecilia R. Aragon
PacificVis3
2017 Squares: Supporting Interactive Performance Analysis for Multiclass Classifiers
abstract
Performance analysis is critical in applied machine learning because it influences the models practitioners produce. Current performance analysis tools suffer from issues including obscuring important characteristics of model behavior and dissociating performance from data. In this work, we present Squares, a performance visualization for multiclass classification problems. Squares supports estimating common performance metrics while displaying instance-level distribution information necessary for helping practitioners prioritize efforts and access data. Our controlled study shows that practitioners can assess performance significantly faster and more accurately with Squares than a confusion matrix, a common performance analysis tool in machine learning.
Donghao Ren, Saleema Amershi, Bongshin Lee, Jina Suh, Jason D. Williams
IEEE Trans. Vis. Comput. Graph.4
2016 The Label Complexity of Mixed-Initiative Classifier Training
abstract
Mixed-initiative classifier training, where the human teacher can choose which items to label or to label items chosen by the computer, has enjoyed empirical success but without a rigorous statistical learning theoretical justification. We analyze the label complexity of a simple mixed-initiative training mechanism using teach- ing dimension and active learning. We show that mixed-initiative training is advantageous com- pared to either computer-initiated (represented by active learning) or human-initiated classifier training. The advantage exists across all human teaching abilities, from optimal to completely unhelpful teachers. We further improve classifier training by educating the human teachers. This is done by showing, or explaining, optimal teaching sets to the human teachers. We conduct Mechanical Turk human experiments on two stylistic classifier training tasks to illustrate our approach.
Jina Suh, Xiaojin Zhu 0001, Saleema Amershi
ICML1
2015 ModelTracker: Redesigning Performance Analysis Tools for Machine Learning
abstract
Model building in machine learning is an iterative process. The performance analysis and debugging step typically involves a disruptive cognitive switch from model building to error analysis, discouraging an informed approach to model building. We present ModelTracker, an interactive visualization that subsumes information contained in numerous traditional summary statistics and graphs while displaying example-level performance and enabling direct error examination and debugging. Usage analysis from machine learning practitioners building real models with ModelTracker over six months shows ModelTracker is used often and throughout model building. A controlled experiment focusing on ModelTracker's debugging capabilities shows participants prefer ModelTracker over traditional tools without a loss in model performance.
Saleema Amershi, David Maxwell Chickering, Steven Mark Drucker, Bongshin Lee, Patrice Y. Simard, Jina Suh
CHI6