EDBT 2026 Demo / reviewers in the wild / expert
Mingrui Ray Zhang
dblp:239/7686
· DBLP profile ↗
17ranked-venue papers
9as first author
9since 2021 · last 2023
0000-0003-2557-5903ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 17 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Self-Talk with Superhero Zip: Supporting Children's Socioemotional Learning with Conversational AgentsabstractSocioemotional competencies are fundamental for children’s growth and success, and prior work shows that in some instances, technology can support children in acquiring these skills. Here, we examine whether children can learn to use a socioemotional strategy known as “self-talk” from a conversational agent (CA). To investigate this question, we designed and built “Self-Talk with Superhero Zip,” an interactive CA experience, and deployed it for one week in ten family homes to pairs of siblings between the ages of five and ten (N = 20). We found that children could recall and accurately describe the lessons taught by the intervention, and we saw indications of children applying self-talk in daily life. Targeting sibling pairs rather than individual users proved to be a design challenge in its own right, and families suggested design ideas for supporting this context, such as UI to manage conversational flow and reduce competition, and visuals and embodied activities to encourage focus. The dual-user context coupled with the audio modality prompted “preinput huddles” in which children conversed in whispers before responding to the system. We contribute evidence that CAs can support children in learning to use self-talk as well as design guidance for creating multi-user conversational interfaces. Mingrui Ray Zhang, Lynn K. Nguyen, Rebecca Michelson, Tala June Tayebi, Alexis Hiniker |
IDC | 2 |
| 2022 | "I Don't Even Remember What I Read": How Design Influences Dissociation on Social MediaabstractMany people have experienced mindlessly scrolling on social media. We investigated these experiences through the lens of normative dissociation: total cognitive absorption, characterized by diminished self-awareness and reduced sense of agency. To explore user experiences of normative dissociation and how design affects the likelihood of normative dissociation, we deployed Chirp, a custom Twitter client, to 43 U.S. participants. Experience sampling and interviews revealed that sometimes, becoming absorbed in normative dissociation on social media felt like a beneficial break. However, people also reported passively slipping into normative dissociation, such that they failed to absorb any content and were left feeling like they had wasted their time. We found that designed interventions–including custom lists, reading history labels, time limit dialogs, and usage statistics–reduced normative dissociation. Our findings demonstrate that interaction designs intended to capture attention likely do so by harnessing people’s natural inclination to seek normative dissociation experiences. This suggests that normative dissociation may be a more productive framing than addiction for discussing social media overuse. Amanda Baughan, Mingrui Ray Zhang, Raveena Rao, Kai Lukoff, Anastasia Schaadhardt, Lisa D. Butler, Alexis Hiniker |
CHI | 2 |
| 2022 | Ga11y: An Automated GIF Annotation System for Visually Impaired UsersabstractAnimated GIF images have become prevalent in internet culture, often used to express richer and more nuanced meanings than static images. But animated GIFs often lack adequate alternative text descriptions, and it is challenging to generate such descriptions automatically, resulting in inaccessible GIFs for blind or low-vision (BLV) users. To improve the accessibility of animated GIFs for BLV users, we provide a system called Ga11y (pronounced “galley”), for creating GIF annotations. Ga11y combines the power of machine intelligence and crowdsourcing and has three components: an Android client for submitting annotation requests, a backend server and database, and a web interface where volunteers can respond to annotation requests. We evaluated three human annotation interfaces and employ the one that yielded the best annotation quality. We also conducted a multi-stage evaluation with 12 BLV participants from the United States and China, receiving positive feedback. Mingrui Ray Zhang, Mingyuan Zhong 0001, Jacob O. Wobbrock |
CHI | 1 |
| 2022 | Monitoring Screen Time or Redesigning It?: Two Approaches to Supporting Intentional Social Media UseabstractExisting designs helping people manage their social media use include: 1) external supports that monitor and limit use; 2) internal supports that change the interface itself. Here, we design and deploy Chirp, a mobile Twitter client, to independently examine how users experience external and internal supports. To develop Chirp, we identified 16 features that influence users’ sense of agency on Twitter through a survey of 129 participants and a design workshop. We then conducted a four-week within-subjects deployment with 31 participants. Our internal supports (including features to filter tweets and inform users when they have exhausted new content) significantly increased users’ sense of agency, while our external supports (a usage dashboard and nudges to close the app) did not. Participants valued our internal supports and said that our external supports were for “other people.” Our findings suggest that design patterns promoting agency may serve users better than screen time tools. Mingrui Ray Zhang, Kai Lukoff, Raveena Rao, Amanda Baughan, Alexis Hiniker |
CHI | 1 |
| 2022 | TypeAnywhere: A QWERTY-Based Text Entry Solution for Ubiquitous ComputingabstractWe present a QWERTY-based text entry system, TypeAnywhere, for use in off-desktop computing environments. Using a wearable device that can detect finger taps, users can leverage their touch-typing skills from physical keyboards to perform text entry on any surface. TypeAnywhere decodes typing sequences based only on finger-tap sequences without relying on tap locations. To achieve optimal decoding performance, we trained a neural language model and achieved a 1.6% character error rate (CER) in an offline evaluation, compared to a 5.3% CER from a traditional n-gram language model. Our user study showed that participants achieved an average performance of 70.6 WPM, or 80.4% of their physical keyboard speed, and 1.50% CER after 2.5 hours of practice over five days on a table surface. They also achieved 43.9 WPM and 1.37% CER when typing on their laps. Our results demonstrate the strong potential of QWERTY typing as a ubiquitous text entry solution. Mingrui Ray Zhang, Shumin Zhai, Jacob O. Wobbrock |
CHI | 1 |
| 2021 | Can Conversational Agents Change the Way Children Talk to People?abstractMillions of children now use conversational agents (CAs), leading researchers and the public alike to ask how interactions with these devices might shape children’s communication with people. We conducted a single-session observational lab study with 22 five-to-ten-year-old children as a step toward understanding whether and how children might transfer a linguistic routine they learned from a CA to a conversation with another person. We found that 68% of children spontaneously used this routine in a conversation with their parent in the lab, and 55% continued to use it at home. When addressing parents, children infused the routine with warmth and playfulness that they did not use when addressing the CA, adapting it to suit their relationship with their parent. However, only 18% of children used it in conversation with an unfamiliar researcher, where they instead were more likely to follow conventional conversational norms. These findings suggest children are quick to learn linguistic routines from CAs but use social differentiation when they apply them. Children’s willingness to expand on and share the routine with their parent is consistent with the principles of the Joint Media Engagement (JME) framework and suggests CAs may be a productive medium for creating JME experiences. Alexis Hiniker, Amelia Wang, Jonathan A. Tran, Mingrui Ray Zhang, Jenny S. Radesky, Kiley Sobel, Sungsoo Ray Hong |
IDC | 4 |
| 2021 | Revamp: Enhancing Accessible Information Seeking Experience of Online Shopping for Blind or Low Vision UsersabstractOnline shopping has become a valuable modern convenience, but blind or low vision (BLV) users still face significant challenges using it, because of: 1) inadequate image descriptions and 2) the inability to filter large amounts of information using screen readers. To address those challenges, we propose Revamp, a system that leverages customer reviews for interactive information retrieval. Revamp is a browser integration that supports review-based question-answering interactions on a reconstructed product page. From our interview, we identified four main aspects (color, logo, shape, and size) that are vital for BLV users to understand the visual appearance of a product. Based on the findings, we formulated syntactic rules to extract review snippets, which were used to generate image descriptions and responses to users’ queries. Evaluations with eight BLV users showed that Revamp 1) provided useful descriptive information for understanding product appearance and 2) helped the participants locate key information efficiently. Ruolin Wang, Mingrui Ray Zhang, Zhaoheng Li, Zhixiu Liu, Zihan Dang, Chun Yu, Xiang 'Anthony' Chen |
CHI | 3 |
| 2021 | Voicemoji: Emoji Entry Using Voice for Visually Impaired PeopleabstractKeyboard-based emoji entry can be challenging for people with visual impairments: users have to sequentially navigate emoji lists using screen readers to find their desired emojis, which is a slow and tedious process. In this work, we explore the design and benefits of emoji entry with speech input, a popular text entry method among people with visual impairments. After conducting interviews to understand blind or low vision (BLV) users’ current emoji input experiences, we developed Voicemoji, which (1) outputs relevant emojis in response to voice commands, and (2) provides context-sensitive emoji suggestions through speech output. We also conducted a multi-stage evaluation study with six BLV participants from the United States and six BLV participants from China, finding that Voicemoji significantly reduced entry time by 91.2% and was preferred by all participants over the Apple iOS keyboard. Based on our findings, we present Voicemoji as a feasible solution for voice-based emoji entry. Mingrui Ray Zhang, Ruolin Wang, Xuhai Xu, Qisheng Li, Ather Sharif, Jacob O. Wobbrock |
CHI | 1 |
| 2021 | PhraseFlow: Designs and Empirical Studies of Phrase-Level InputabstractDecoding on phrase-level may afford more correction accuracy than on word-level according to previous research. However, how phrase-level input affects the user typing behavior, and how to design the interaction to make it practical remain under explored. We present PhraseFlow, a phrase-level input keyboard that is able to correct previous text based on the subsequently input sequences. Computational studies show that phrase-level input reduces the error rate of autocorrection by over 16%. We found that phrase-level input introduced extra cognitive load to the user that hindered their performance. Through an iterative design-implement-research process, we optimized the design of PhraseFlow that alleviated the cognitive load. An in-lab study shows that users could adopt PhraseFlow quickly, resulting in 19% fewer error without losing speed. In real-life settings, we conducted a six-day deployment study with 42 participants, showing that 78.6% of the users would like to have the phrase-level input feature in future keyboards. Mingrui Ray Zhang, Shumin Zhai |
CHI | 1 |
| 2020 | Gedit: Keyboard Gestures for Mobile Text EditingabstractText editing on mobile devices can be a tedious process. To perform various editing operations, a user must repeatedly move his or her fingers between the text input area and the keyboard, making multiple round trips and breaking the flow of typing. In this work, we present Gedit, a system of on-keyboard gestures for convenient mobile text editing. Our design includes a ring gesture and flicks for cursor control, bezel gestures for mode switching, and four gesture shortcuts for copy, paste, cut, and undo. Variations of our gestures exist for one and two hands. We conducted an experiment to compare Gedit with the de facto touch+widget based editing interactions. Our results showed that Gedit's gestures were easy to learn, 24% and 17% faster than the de facto interactions for oneand two-handed use, respectively, and preferred by participants. Mingrui Ray Zhang, Jacob O. Wobbrock |
Graphics Interface | 1 |
| 2020 | JustCorrect: Intelligent Post Hoc Text Correction Techniques on SmartphonesabstractCorrecting errors in entered text is a common task but usually diffcult to perform on mobile devices due to tedious cursor navigation steps. In this paper, we present JustCorrect, an intelligent post hoc text correction technique for smartphones. To make a correction, the user simply types the correct text at the end of their current input, and JustCorrect will automatically detect the error and apply the correction in the form of an insertion or a substitution. In this way, manual navigation steps are bypassed, and the correction can be committed with a single tap. We solved two critical problems to support JustCorrect: (1) Correction Algorithm: we propose an algorithm that infers the user's correction intention from the last typed word. (2) Input Modalities: our study revealed that both tap and gesture were suitable input modalities for performing JustCorrect. Based on our fndings, we integrated JustCorrect into a soft keyboard. Our user studies show that using JustCorrect reduces the text correction time by 12.8% over the stock Android keyboard and by 9.7% over the "Type, then Correct" text correction technique by Zhang et al. (2019). Overall, JustCorrect complements existing post hoc text correction techniques, making error correction more automatic and intelligent. Wenzhe Cui, Suwen Zhu, Mingrui Ray Zhang, H. Andrew Schwartz, Jacob O. Wobbrock, Xiaojun Bi 0001 |
UIST | 3 |
| 2019 | Communication Breakdowns Between Families and AlexaabstractWe investigate how families repair communication breakdowns with digital home assistants. We recruited 10 diverse families to use an Amazon Echo Dot in their homes for four weeks. All families had at least one child between four and 17 years old. Each family participated in pre- and post- deployment interviews. Their interactions with the Echo Dot (Alexa) were audio recorded throughout the study. We analyzed 59 communication breakdown interactions between family members and Alexa, framing our analysis with concepts from HCI and speech-language pathology. Our findings indicate that family members collaborate using discourse scaffolding (supportive communication guidance) and a variety of speech and language modifications in their attempts to repair communication breakdowns with Alexa. Alexa's responses also influence the repair strategies that families use. Designers can relieve the communication repair burden that primarily rests with families by increasing digital home assistants' abilities to collaborate together with users to repair communication breakdowns. Erin Beneteau, Olivia K. Richards, Mingrui Ray Zhang, Julie A. Kientz, Jason C. Yip 0001, Alexis Hiniker |
CHI | 3 |
| 2019 | Anchored Audio Sampling: A Seamless Method for Exploring Children's Thoughts During Deployment StudiesabstractMany traditional HCI methods, such as surveys and interviews, are of limited value when working with preschoolers. In this paper, we present anchored audio sampling (AAS), a remote data collection technique for extracting qualitative audio samples during field deployments with young children. AAS offers a developmentally sensitive way of understanding how children make sense of technology and situates their use in the larger context of daily life. AAS is defined by an anchor event, around which audio is collected. A sliding window surrounding this anchor captures both antecedent and ensuing recording, providing the researcher insight into the activities that led up to the event of interest as well as those that followed. We present themes from three deployments that leverage this technique. Based on our experiences using AAS, we have also developed a reusable open-source library for embedding AAS into any Android application. Alexis Hiniker, Jon Froehlich, Mingrui Ray Zhang, Erin Beneteau |
CHI | 3 |
| 2019 | Text Entry Throughput: Towards Unifying Speed and Accuracy in a Single Performance MetricabstractHuman-computer input performance inherently involves speed-accuracy tradeoffs---the faster users act, the more inaccurate those actions are. Therefore, comparing speeds and accuracies separately can result in ambiguous outcomes: Does a fast but inaccurate technique perform better or worse overall than a slow but accurate one? For pointing, speed and accuracy has been unified for over 60 years as throughput (bits/s) (Crossman 1957, Welford 1968), but to date, no similar metric has been established for text entry. In this paper, we introduce a text entry method-independent throughput metric based on Shannon information theory (1948). To explore the practical usability of the metric, we conducted an experiment in which 16 participants typed with a laptop keyboard using different cognitive sets, i.e., speed-accuracy biases. Our results show that as a performance metric, text entry throughput remains relatively stable under different speed-accuracy conditions. We also evaluated a smartphone keyboard with 12 participants, finding that throughput varied least compared to other text entry metrics. This work allows researchers to characterize text entry performance with a single unified measure of input efficiency. Mingrui Ray Zhang, Shumin Zhai, Jacob O. Wobbrock |
CHI | 1 |
| 2019 | Beyond the Input Stream: Making Text Entry Evaluations More Flexible with Transcription SequencesabstractMethod-independent text entry evaluation tools are often used to conduct text entry experiments and compute performance metrics, like words per minute and error rates. The input stream paradigm of Soukoreff & MacKenzie (2001, 2003) still remains prevalent, which presents a string for transcription and uses a strictly serial character representation for encoding the text entry process. Although an advance over prior paradigms, the input stream paradigm is unable to support many modern text entry features. To address these limitations, we present transcription sequences: for each new input, a snapshot of the entire transcribed string unto that point is captured. By comparing adjacent strings within a transcription sequence, we can compute all prior metrics, reduce artificial constraints on text entry evaluations, and introduce new metrics. We conducted a study with 18 participants who typed 1620 phrases using a laptop keyboard, on-screen keyboard, and smartphone keyboard using features such as auto-correction, word prediction, and copy/paste. We also evaluated non-keyboard methods Dasher, gesture typing, and T9. Our results show that modern text entry methods and features can be accommodated, prior metrics can be correctly computed, and new metrics can reveal insights. We validated our algorithms using ground truth based on cursor positioning, confirming 100% accuracy. We also provide a new tool, TextTest++, to facilitate web-based evaluations. Mingrui Ray Zhang, Jacob O. Wobbrock |
UIST | 1 |
| 2019 | Type, Then Correct: Intelligent Text Correction Techniques for Mobile Text Entry Using Neural NetworksabstractCurrent text correction processes on mobile touch devices are laborious: users either extensively use backspace, or navigate the cursor to the error position, make a correction, and navigate back, usually by employing multiple taps or drags over small targets. In this paper, we present three novel text correction techniques to improve the correction process: Drag-n-Drop, Drag-n-Throw, and Magic Key. All of the techniques skip error-deletion and cursor-positioning procedures, and instead allow the user to type the correction first, and then apply that correction to a previously committed error. Specifically, Drag-n-Drop allows a user to drag a correction and drop it on the error position. Drag-n-Throw lets a user drag a correction from the keyboard suggestion list and "throw" it to the approximate area of the error text, with a neural network determining the most likely error in that area. Magic Key allows a user to type a correction and tap a designated key to highlight possible error candidates, which are also determined by a neural network. The user can navigate among these candidates by directionally dragging from atop the key, and can apply the correction by simply tapping the key. We evaluated these techniques in both text correction and text composition tasks. Our results show that correction with the new techniques was faster than de facto cursor and backspace-based correction. Our techniques apply to any touch-based text entry method. Mingrui Ray Zhang, Jacob O. Wobbrock |
UIST | 1 |
| 2015 | ATK: Enabling Ten-Finger Freehand Typing in Air Based on 3D Hand Tracking DataabstractTen-finger freehand mid-air typing is a potential solution for post-desktop interaction. However, the absence of tactile feedback as well as the inability to accurately distinguish tapping finger or target keys exists as the major challenge for mid-air typing. In this paper, we present ATK, a novel interaction technique that enables freehand ten-finger typing in the air based on 3D hand tracking data. Our hypothesis is that expert typists are able to transfer their typing ability from physical keyboards to mid-air typing. We followed an iterative approach in designing ATK. We first empirically investigated users' mid-air typing behavior, and examined fingertip kinematics during tapping, correlated movement among fingers and 3D distribution of tapping endpoints. Based on the findings, we proposed a probabilistic tap detection algorithm, and augmented Goodman's input correction model to account for the ambiguity in distinguishing tapping finger. We finally evaluated the performance of ATK with a 4-block study. Participants typed 23.0 WPM with an uncorrected word-level error rate of 0.3% in the first block, and later achieved 29.2 WPM in the last block without sacrificing accuracy. Xin Yi 0001, Chun Yu, Mingrui Ray Zhang, Sida Gao, Ke Sun 0003, Yuanchun Shi |
UIST | 3 |