VLDB 2026 Research / reviewers in the wild / expert
Yukang Yan
dblp:199/3032
· DBLP profile ↗
39ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0001-7515-3755ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 34 · 4 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tuning the Face: Modulating Facial Expressions for Realistic Self-Avatars in Virtual RealityabstractPhotorealistic self-avatars with facial tracking bring VR communication closer to face-to-face interaction, but also expose a fundamental limitation: people often lack accurate awareness and control of their own facial expressions. As a result, directly mirroring facial movements may not reliably convey users’ intended expressions in social VR. We propose modulating rendered expressions on avatars to improve social interaction in VR (e.g., exaggerating positive ones). To design effective modulation strategies and understand how modulation impacts user behavior and perception, we invited 18 participants to perform emotional expressions under varying modulation levels. We collected their self-reported ratings on accuracy and social appropriateness, as well as naturalness ratings and emotion recognition results from 18 additional observers who viewed the expressions. Results show that suppressing rendered expressions amplified participants’ actual facial movements, with effects differing by expression type. Based on our findings, we calculated the modulation ranges for different expressions and developed a real-time modulation system that exaggerates users’ facial expressions with a mild amplitude. We evaluated the system through a VR speech task where 12 participants delivered speeches with and without modulation. Results showed that the applied modulation improved emotional expressiveness and performance. Finally, we developed three application prototypes to illustrate the broader practical implications of our findings. Jiamu Tang, Jiankun Yang, Shijian Luo, Chenliang Xu, Yukang Yan |
DIS | 6 |
| 2026 | Capability at a Glance: Design Guidelines for Intuitive Avatars Communicating Augmented Actions in Virtual RealityabstractVirtual Reality (VR) enables users to engage with capabilities beyond human limitations, but it is not always obvious how to trigger these capabilities. Taking the lens of Affordance [35], we believe avatar design is the key to solving this issue, which ideally should communicate its capabilities and how to activate them. To understand the current practice, we selected eight capabilities across four categories and invited twelve professional designers to design avatars that communicate the capabilities and their corresponding interactions. From the resulting designs, we formed 16 guidelines to provide general and category-specific recommendations. Then, we validated these guidelines by letting two groups of twelve participants design avatars with and without guidelines. Participants rated the guidelines’ clarity and usefulness highly. External judges confirmed that avatars designed with the guidelines were more intuitive in conveying the capabilities and interaction methods. Finally, we demonstrated the applicability of the guidelines in avatar design for four VR applications. Jiamu Tang, Yanna Lin, Jiankun Yang, Longyu Zhang, Shijian Luo, Yukang Yan |
CHI | 8 |
| 2025 | Modeling the Impact of Visual Stimuli on Redirection Noticeability with Gaze Behavior in Virtual Reality
Zhipeng Li 0001, Yishu Ji, Ruijia Chen, Yuntao Wang 0001, Yuanchun Shi, Yukang Yan |
CHI | 7 |
| 2025 | DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio DiscussionsabstractBy overlaying time-synced user comments on videos, Danmu creates a co-watching experience for online viewers. However, its visual-centric design poses significant challenges for blind and low vision (BLV) viewers. Our formative study identified three primary challenges that hinder BLV viewers' engagement with Danmu: the lack of visual context, the speech interference between comments and videos, and the disorganization of comments. To address these challenges, we present DanmuA11y, a system that makes Danmu accessible by transforming it into multi-viewer audio discussions. DanmuA11y incorporates three core features: (1) Augmenting Danmu with visual context, (2) Seamlessly integrating Danmu into videos, and (3) Presenting Danmu via multi-viewer discussions. Evaluation with twelve BLV viewers demonstrated that DanmuA11y significantly improved Danmu comprehension, provided smooth viewing experiences, and fostered social connections among viewers. We further highlight implications for enhancing commentary accessibility in video-based social media and live-streaming platforms. Shuchang Xu, Xiaofu Jin, Huamin Qu, Yukang Yan |
CHI | 4 |
| 2025 | From Following to Understanding: Investigating the Role of Reflective Prompts in AR-Guided Tasks to Promote User UnderstandingabstractAugmented Reality (AR) is a promising medium for guiding users through tasks, yet its impact on fostering deeper task understanding remains underexplored. This paper investigates the impact of reflective prompts -- strategic questions that encourage users to challenge assumptions, connect actions to outcomes, and consider hypothetical scenarios -- on task comprehension and performance. We conducted a two-phase study: a formative survey and co-design sessions (N=9) to develop reflective prompts, followed by a within-subject evaluation (N=16) comparing AR instructions with and without these prompts in coffee-making and circuit assembly tasks. Our results show that reflective prompts significantly improved objective task understanding and resulted in more proactive information acquisition behaviors during task completion. These findings highlight the potential of incorporating reflective elements into AR instructions to foster deeper engagement and learning. Based on data from both studies, we synthesized design guidelines for integrating reflective elements into AR systems to enhance user understanding without compromising task performance. Nandi Zhang, Yukang Yan, Ryo Suzuki 0001 |
CHI | 2 |
| 2025 | Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision UsersabstractFigure 1: Branch Explorer transforms 360° videos into branching narratives-stories that dynamically unfold based on viewer choices-to create an engaging experience for blind and low vision (BLV) users.It employs a multi-modal machine learning pipeline to generate diverse narrative paths, enabling BLV users to make choices at key branching points and explore each storyline through immersive audio guidance.The figure shows the 360° video HELP (available at: https://youtu.be/G-XZhKqQAHU). Shuchang Xu, Xiaofu Jin, Huamin Qu, Yukang Yan |
UIST | 5 |
| 2025 | A Comparison Study Understanding the Impact of Mixed Reality Collaboration on Sense of Co-PresenceabstractSense of co-presence refers to the perceived closeness and interaction between participants in a collaborative context, which critically impacts the collaboration experience and task performance. With the emergence of Mixed Reality (MR) technologies, we would like to investigate the effect of MR immersive collaboration environment on promoting co-presence in a remote setting by comparing it with non-MR methods, such as video conferencing. We conduct a comparison study, where we invited 14 dyads of participants to collaborate on block assembly tasks with video conferencing, MR system, and in a physically co-located scenario. Each participant of a dyad was assigned either a local worker to assemble the blocks or a remote helper to give the instructions. Results show that MR system can create comparable sense of co-presence with co-located situation, and allow users to interact more naturally with both the environment and each other. The adoption of mixed reality enhances collaboration and task performance by reducing reliance on verbal communication and favoring action-based interactions through gestures and direct manipulation of virtual objects. Weicheng Zheng, Yuntao Wang 0001, Xin Tong 0004, Yukang Yan |
VR | 5 |
| 2025 | Systematic Literature Review of Using Virtual Reality as a Social Platform in HCI CommunityabstractVirtual reality (VR) is increasingly used as a social platform for users to interact and build connections with one another in an immersive virtual environment. Reflecting on the empirical progress in this area of study, a comprehensive review of how VR could be used to support social interaction is required to consolidate existing practices and identify research gaps to inspire future studies. In this work, we conducted a systematic review of 94 publications in the HCI field to examine how VR is designed and evaluated for social purposes. We found that VR influences social interaction through self-representation, interpersonal interactions, and interaction environments. We summarized four positive effects of using VR for socializing, which are relaxation, engagement, intimacy, and accessibility, and showed that it could also negatively affect user social experiences by intensifying harassment experiences and amplifying privacy concerns. We introduce an evaluation framework that outlines the key aspects of social experience: intrapersonal, interpersonal, and interaction experiences. According to the results, we uncover several research gaps and propose future directions for designing and developing VR to enhance social experience. Xiaoying Wei, Xiaofu Jin, Ge Lin 0001, Yukang Yan, Mingming Fan 0001 |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | AmplitudeArrow: On-the-Go AR Menu Selection Using Consecutive Simple Head Gestures and Amplitude VisualizationabstractHeads-up computing aims to provide synergistic digital assistance that minimally interferes with users' on-the-go daily activities. Currently, the input modalities of heads-up computing are mainly voice and finger gestures. In this work, we propose and evaluate the AmplitudeArrow (AA) technique designed for on-the-go AR menu selection to demonstrate that consecutive simple head gestures can also be an effective input modality for heads-up computing. Specifically, AA arranges menu icons into one/two row(s). To select a target icon, the user first makes their head yaw to pre-select the target icon or the column containing it and then makes their head pitch to make the arrow in the target icon expand until the arrow covers the target icon completely, i.e., the pitch amplitude surpasses the selection confirmation threshold. User studies indicated that AA demonstrated robust resistance to walking-caused head perturbation and external factors such as other people/obstacles, delivering high accuracy (error rate $< $< 5$\%$%) and fast speed ($< $< 1.5s per selection) when there were no more than six icon columns (twelve icons) distributed horizontally and evenly in a menu area with a horizontal visual angle of $43^{\circ }$43∘. Yang Tian 0008, Yukang Yan, Shengdong Zhao 0001, Xiaojuan Ma, Yuanchun Shi |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | MineXR: Mining Personalized Extended Reality InterfacesabstractExtended Reality (XR) interfaces offer engaging user experiences, but their effective design requires a nuanced understanding of user behavior and preferences. This knowledge is challenging to obtain without the widespread adoption of XR devices. We introduce MineXR, a design mining workflow and data analysis platform for collecting and analyzing personalized XR user interaction and experience data. MineXR enables elicitation of personalized interfaces from participants of a data collection: for any particular context, participants create interface elements using application screenshots from their own smartphone, place them in the environment, and simultaneously preview the resulting XR layout on a headset. Using MineXR, we contribute a dataset of personalized XR interfaces collected from 31 participants, consisting of 695 XR widgets created from 178 unique applications. We provide insights for XR widget functionalities, categories, clusters, UI element types, and placement. Our open-source tools and data support researchers and designers in developing future XR interfaces. Hyunsung Cho, Yukang Yan, Kashyap Todi, Mark Parent, Missie Smith, Tanya R. Jonker, Hrvoje Benko, David Lindlbauer |
CHI | 2 |
| 2024 | Predicting the Noticeability of Dynamic Virtual Elements in Virtual RealityabstractWhile Virtual Reality (VR) systems can present virtual elements such as notifications anywhere, designing them so they are not missed by or distracting to users is highly challenging for content creators. To address this challenge, we introduce a novel approach to predict the noticeability of virtual elements. It computes the visual saliency distribution of what users see, and analyzes the temporal changes of the distribution with respect to the dynamic virtual elements that are animated. The computed features serve as input for a long short-term memory (LSTM) model that predicts whether a virtual element will be noticed. Our approach is based on data collected from 24 users in different VR environments performing tasks such as watching a video or typing. We evaluate our approach (n = 12), and show that it can predict the timing of when users notice a change to a virtual element within 2.56 sec compared to a ground truth, and demonstrate the versatility of our approach with a set of applications. We believe that our predictive approach opens the path for computational design tools that assist VR content creators in creating interfaces that automatically adapt virtual elements based on noticeability. Zhipeng Li 0001, Yi Fei Cheng 0001, Yukang Yan, David Lindlbauer |
CHI | 3 |
| 2024 | Time2Stop: Adaptive and Explainable Human-AI Loop for Smartphone Overuse InterventionabstractDespite a rich history of investigating smartphone overuse intervention techniques, AI-based just-in-time adaptive intervention (JITAI) methods for overuse reduction are lacking. We develop Time2Stop, an intelligent, adaptive, and explainable JITAI system that leverages machine learning to identify optimal intervention timings, introduces interventions with transparent AI explanations, and collects user feedback to establish a human-AI loop and adapt the intervention model over time. We conducted an 8-week field experiment (N=71) to evaluate the effectiveness of both the adaptation and explanation aspects of Time2Stop. Our results indicate that our adaptive models significantly outperform the baseline methods on intervention accuracy (>32.8% relatively) and receptivity (>8.0%). In addition, incorporating explanations further enhances the effectiveness by 53.8% and 11.4% on accuracy and receptivity, respectively. Moreover, Time2Stop significantly reduces overuse, decreasing app visit frequency by 7.0 ∼ 8.9%. Our subjective data also echoed these quantitative measures. Participants preferred the adaptive interventions and rated the system highly on intervention time accuracy, effectiveness, and level of trust. We envision our work can inspire future research on JITAI systems with a human-AI loop to evolve with users. Adiba Orzikulova, Zhipeng Li 0001, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Marzyeh Ghassemi, Sung-Ju Lee 0001, Anind K. Dey, Xuhai Xu |
CHI | 4 |
| 2024 | AdaptiveVoice: Cognitively Adaptive Voice Interface for Driving AssistanceabstractCurrent voice assistants present messages in a predefined format without considering users’ mental states. This paper presents an optimization-based approach to alleviate this issue which adjusts the level of details and speech speed of the voice messages according to the estimated cognitive load of the user. In the first user study (N = 12), we investigated the impact of cognitive load on user performance. The findings reveal significant differences in preferred message formats across five cognitive load levels, substantiating the need for voice message adaptation. We then implemented AdaptiveVoice, an algorithm based on combinatorial optimization to generate adaptive voice messages in real time. In the second user study (N = 30) conducted in a VR-simulated driving environment, we compare AdaptiveVoice with a fixed format baseline, with and without visual guidance on the Heads-up display (HUD). Results indicate that users benefit from AdaptiveVoice with reduced response time and improved driving performance, particularly when it is augmented with HUD. Shaoyue Wen, Songming Ping, Jialin Wang 0002, Hai-Ning Liang, Xuhai Xu, Yukang Yan |
CHI | 6 |
| 2024 | Computational Trichromacy Reconstruction: Empowering the Color-Vision Deficient to Recognize Colors Using Augmented RealityabstractWe propose an assistive technology that helps individuals with Color Vision Deficiencies (CVD) to recognize/name colors. A dichromat’s color perception is a reduced two-dimensional (2D) subset of a normal trichromat’s three dimensional color (3D) perception, leading to confusion when visual stimuli that appear identical to the dichromat are referred to by different color names. Using our proposed system, CVD individuals can interactively induce distinct perceptual changes to originally confusing colors via a computational color space transformation. By combining their original 2D precepts for colors with the discriminative changes, a three dimensional color space is reconstructed, where the dichromat can learn to resolve color name confusions and accurately recognize colors. Our system is implemented as an Augmented Reality (AR) interface on smartphones, where users interactively control the rotation through swipe gestures and observe the induced color shifts in the camera view or in a displayed image. Through psychophysical experiments and a longitudinal user study, we demonstrate that such rotational color shifts have discriminative power (initially confusing colors become distinct under rotation) and exhibit structured perceptual shifts dichromats can learn with modest training. The AR App is also evaluated in two real-world scenarios (building with lego blocks and interpreting artistic works); users all report positive experience in using the App to recognize object colors that they otherwise could not. Yuhao Zhu 0001, Ethan Chen, Colin Hascup, Yukang Yan, Gaurav Sharma 0001 |
UIST | 4 |
| 2024 | Auptimize: Optimal Placement of Spatial Audio Cues for Extended RealityabstractSpatial audio in Extended Reality (XR) provides users with better awareness of where virtual elements are placed, and efficiently guides them to events such as notifications, system alerts from different windows, or approaching avatars. Humans, however, are inaccurate in localizing sound cues, especially with multiple sources due to limitations in human auditory perception such as angular discrimination error and front-back confusion. This decreases the efficiency of XR interfaces because users misidentify from which XR element a sound is coming from. To address this, we propose Auptimize, a novel computational approach for placing XR sound sources, which mitigates such localization errors by utilizing the ventriloquist effect. Auptimize disentangles the sound source locations from the visual elements and relocates the sound sources to optimal positions for unambiguous identification of sound cues, avoiding errors due to inter-source proximity and front-back confusion. Our evaluation shows that Auptimize decreases spatial audio-based source identification errors compared to playing sound cues at the paired visual-sound locations. We demonstrate the applicability of Auptimize for diverse spatial audio-based interactive XR scenarios. Hyunsung Cho, Alexander Wang, Divya Kartik, Emily Liying Xie, Yukang Yan, David Lindlbauer |
UIST | 5 |
| 2024 | Memory Reviver: Supporting Photo-Collection Reminiscence for People with Visual Impairment via a Proactive ChatbotabstractReminiscing with photo collections offers significant psychological benefits but poses challenges for people with visual impairment (PVI). Their current reliance on sighted help restricts the flexibility of this activity. In response, we explored using a chatbot in a preliminary study. We identified two primary challenges that hinder effective reminiscence with a chatbot: the scattering of information and a lack of proactive guidance. To address these limitations, we present Memory Reviver, a proactive chatbot that helps PVI reminisce with a photo collection through natural language communication. Memory Reviver incorporates two novel features: (1) a Memory Tree, which uses a hierarchical structure to organize the information in a photo collection; and (2) a Proactive Strategy, which actively delivers information to users at proper conversation rounds. Evaluation with twelve PVI demonstrated that Memory Reviver effectively facilitated engaging reminiscence, enhanced understanding of photo collections, and delivered natural conversational experiences. Based on our findings, we distill implications for supporting photo reminiscence and designing chatbots for PVI. Shuchang Xu, Chang Chen 0005, Xiaofu Jin, Linping Yuan, Yukang Yan, Huamin Qu |
UIST | 6 |
| 2024 | HCI Research and Innovation in China: A 10-Year PerspectiveabstractIn the past years, human computer interaction (HCI) research and innovation have developed substantially, leading to a number of fruitful research topics. In this paper, we surveyed the HCI research and innovation in China from a 10-year perspective. We analyzed the popular research methodology and topics among Chinese researchers, including human modeling, user interface techniques, context awareness, user acceptance and performance, user experience design, human-AI interaction, HCI applications and social influences. We also conducted a bibliography analysis on the published papers in top-tier conferences and journals, which revealed a significant rising trend, and a generally broad distribution of research types. Moreover, we described typical applications and the industry influence of the research outcomes. We concluded with implications and reflections for HCI researchers across the world and shared the future research trends envisioned by Chinese researchers. Yuanchun Shi, Xin Yi 0001, Yuntao Wang 0001, Yukang Yan, Zhimin Cheng, Pengye Zhu, Yongjuan Li, Yanci Liu, Weixuan Zhou, Diya Zhao |
Int. J. Hum. Comput. Interact. | 6 |
| 2024 | InputJump: Augmented reality-facilitated cross-device input fusion based on spatial and semantic informationabstractThe proliferation of computing devices requires seamless cross-device interactions. Augmented reality (AR) headsets can facilitate interactions with existing computers owing to their user-centered views and natural inputs. In this study, we propose InputJump, a user-centered cross-device input fusion method that maps multi-modal cross-device inputs to interactive elements on graphical interfaces. The input jump calculates the spatial coordinates of the input target positions and the interactive elements within the coordinate system of the AR headset. It also extracts semantic descriptions of inputs and elements using large language models (LLMs). Two types of information from different inputs (e.g., gaze, gesture, mouse, and keyboard) were fused to map onto an interactive element. The proposed method is explained in detail and implemented on both an AR headset and a desktop PC. We then conducted a user study and extensive simulations to validate our proposed method. The results showed that InputJump can accurately associate a fused input with the target interactive element, enabling a more natural and flexible interaction experience. Tengxiang Zhang, Yukang Yan, Yiqiang Chen 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2023 | Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution ImagesabstractA computer vision system using low-resolution image sensors can provide intelligent services (e.g., activity recognition) but preserve unnecessary visual privacy information from the hardware level. However, preserving visual privacy and enabling accurate machine recognition have adversarial needs on image resolution. Modeling the trade-off of privacy preservation and machine recognition performance can guide future privacy-preserving computer vision systems using low-resolution image sensors. In this paper, using the at-home activity of daily livings (ADLs) as the scenario, we first obtained the most important visual privacy features through a user survey. Then we quantified and analyzed the effects of image resolution on human and machine recognition performance in activity recognition and privacy awareness tasks. We also investigated how modern image super-resolution techniques influence these effects. Based on the results, we proposed a method for modeling the trade-off of privacy preservation and activity recognition on low-resolution images. Yuntao Wang 0001, Zirui Cheng, Xin Yi 0001, Yan Kong, Xuhai Xu, Yukang Yan, Chun Yu, Shwetak N. Patel, Yuanchun Shi |
CHI | 7 |
| 2023 | HandAvatar: Embodying Non-Humanoid Virtual Avatars through HandsabstractWe propose HandAvatar to enable users to embody non-humanoid avatars using their hands. HandAvatar leverages the high dexterity and coordination of users’ hands to control virtual avatars, enabled through our novel approach for automatically-generated joint-to-joint mappings. We contribute an observation study to understand users’ preferences on hand-to-avatar mappings on eight avatars. Leveraging insights from the study, we present an automated approach that generates mappings between users’ hands and arbitrary virtual avatars by jointly optimizing control precision, structural similarity, and comfort. We evaluated HandAvatar on static posing, dynamic animation, and creative exploration tasks. Results indicate that HandAvatar enables more precise control, requires less physical effort, and brings comparable embodiment compared to a state-of-the-art body-to-avatar control method. We demonstrate HandAvatar’s potential with applications including non-humanoid avatar based social interaction in VR, 3D animation composition, and VR scene design with physical proxies. We believe that HandAvatar unlocks new interaction opportunities, especially for usage in Virtual Reality, by letting users become the avatar in applications including virtual social interaction, animation, gaming, or education. Yu Jiang 0010, Zhipeng Li 0001, Mufei He, David Lindlbauer, Yukang Yan |
CHI | 5 |
| 2023 | Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device SensingabstractGestures performed accompanying the voice are essential for voice interaction to convey complementary semantics for interaction purposes such as wake-up state and input modality. In this paper, we investigated voice-accompanying hand-to-face (VAHF) gestures for voice interaction. We targeted on hand-to-face gestures because such gestures relate closely to speech and yield significant acoustic features (e.g., impeding voice propagation). We conducted a user study to explore the design space of VAHF gestures, where we first gathered candidate gestures and then applied a structural analysis to them in different dimensions (e.g., contact position and type), outputting a total of 8 VAHF gestures with good usability and least confusion. To facilitate VAHF gesture recognition, we proposed a novel cross-device sensing method that leverages heterogeneous channels (vocal, ultrasound, and IMU) of data from commodity devices (earbuds, watches, and rings). Our recognition model achieved an accuracy of 97.3% for recognizing 3 gestures and 91.5% for recognizing 8 gestures (excluding the "empty" gesture), proving the high applicability. Quantitative analysis also shed light on the recognition capability of each sensor channel and their different combinations. In the end, we illustrated the feasible use cases and their design principles to demonstrate the applicability of our system in various scenarios. Zisu Li, Yuntao Wang 0001, Chun Yu, Yukang Yan, Mingming Fan 0001, Yuanchun Shi |
CHI | 6 |
| 2023 | Selecting Real-World Objects via User-Perspective Phone OcclusionabstractPerceiving the region of interest (ROI) and target object by smartphones from the user’s first-person perspective can enable diverse spatial interactions. In this paper, we propose a novel ROI input method and a target selecting method for smartphones by utilizing the user-perspective phone occlusion. This concept of turning the phone into real-world physical cursor benefits from the proprioception, gets rid of the constraint of camera preview, and allows users to rapidly and accurately select the target object. Meanwhile, our method can provide a resizable and rotatable rectangular ROI to disambiguate dense targets. We implemented the prototype system by positioning the user’s iris with the front camera and estimating the rectangular area blocked by the phone with the rear camera simultaneously, followed by a target prediction algorithm with the distance-weighted Jaccard index. We analyzed the behavioral models of using our method and evaluated our prototype system’s pointing accuracy and usability. Results showed that our method is well-accepted by the users for its convenience, accuracy, and efficiency. Chun Yu, Jiachen Yao, Yueting Weng, Yukang Yan, Yuanchun Shi |
CHI | 7 |
| 2023 | A Survey on Remote Assistance and Training in Mixed Reality EnvironmentsabstractThe recent pandemic, war, and oil crises have caused many to reconsider their need to travel for education, training, and meetings. Providing assistance and training remotely has thus gained importance for many applications, from industrial maintenance to surgical telemonitoring. Current solutions such as video conferencing platforms lack essential communication cues such as spatial referencing, which negatively impacts both time completion and task performance. Mixed Reality (MR) offers opportunities to improve remote assistance and training, as it opens the way to increased spatial clarity and large interaction space. We contribute a survey of remote assistance and training in MR environments through a systematic literature review to provide a deeper understanding of current approaches, benefits and challenges. We analyze 62 articles and contextualize our findings along a taxonomy based on degree of collaboration, perspective sharing, MR space symmetry, time, input and output modality, visual display, and application domain. We identify the main gaps and opportunities in this research area, such as exploring collaboration scenarios beyond one-expert-to-one-trainee, enabling users to move across the reality-virtuality spectrum during a task, or exploring advanced interaction techniques that resort to hand or eye tracking. Our survey informs and helps researchers in different domains, including maintenance, medicine, engineering, or education, build and evaluate novel MR approaches to remote training and assistance. All supplemental materials are available at https://augmented-perception.org/publications/2023-training-survey.html. Catarina G. Fidalgo, Yukang Yan, Hyunsung Cho, Maurício Sousa, David Lindlbauer, Joaquim Jorge 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | ConeSpeech: Exploring Directional Speech Interaction for Multi-Person Remote Communication in Virtual RealityabstractRemote communication is essential for efficient collaboration among people at different locations. We present ConeSpeech, a virtual reality (VR) based multi-user remote communication technique, which enables users to selectively speak to target listeners without distracting bystanders. With ConeSpeech, the user looks at the target listener and only in a cone-shaped area in the direction can the listeners hear the speech. This manner alleviates the disturbance to and avoids overhearing from surrounding irrelevant people. Three featured functions are supported, directional speech delivery, size-adjustable delivery range, and multiple delivery areas, to facilitate speaking to more than one listener and to listeners spatially mixed up with bystanders. We conducted a user study to determine the modality to control the cone-shaped delivery area. Then we implemented the technique and evaluated its performance in three typical multi-user communication tasks by comparing it to two baseline methods. Results show that ConeSpeech balanced the convenience and flexibility of voice communication. Yukang Yan, Haohua Liu, Yingtian Shi, Ruici Guo, Zisu Li, Xuhai Xu, Chun Yu, Yuntao Wang 0001, Yuanchun Shi |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | FaceOri: Tracking Head Position and Orientation Using Ultrasonic Ranging on EarphonesabstractFace orientation can often indicate users’ intended interaction target. In this paper, we propose FaceOri, a novel face tracking technique based on acoustic ranging using earphones. FaceOri can leverage the speaker on a commodity device to emit an ultrasonic chirp, which is picked up by the set of microphones on the user’s earphone, and then processed to calculate the distance from each microphone to the device. These measurements are used to derive the user’s face orientation and distance with respect to the device. We conduct a ground truth comparison and user study to evaluate FaceOri’s performance. The results show that the system can determine whether the user orients to the device at a 93.5% accuracy within a 1.5 meters range. Furthermore, FaceOri can continuously track user’s head orientation with a median absolute error of 10.9 mm in the distance, 3.7° in yaw, and 5.8° in pitch. FaceOri can allow for convenient hands-free control of devices and produce more intelligent context-aware interactions. Yuntao Wang 0001, Jiexin Ding, Ishan Chatterjee, Farshid Salemi Parizi, Yuzhou Zhuang, Yukang Yan, Shwetak N. Patel, Yuanchun Shi |
CHI | 6 |
| 2022 | Towards Understanding Diminished RealityabstractDiminished reality (DR) refers to the concept of removing content from a user’s visual environment. While its implementation is becoming feasible, it is still unclear how users perceive and interact in DR-enabled environments and what applications it benefits. To address this challenge, we first conduct a formative study to compare user perceptions of DR and mediated reality effects (e. g., changing the color or size of target elements) in four example scenarios. Participants preferred removing objects through opacity reduction (i. e., the standard DR implementation) and appreciated mechanisms for maintaining a contextual understanding of diminished items (e. g., outlining). In a second study, we explore the user experience of performing tasks within DR-enabled environments. Participants selected which objects to diminish and the magnitude of the effects when performing two separate tasks (video viewing, assembly). Participants were comfortable with decreased contextual understanding, particularly for less mobile tasks. Based on the results, we define guidelines for creating general DR-enabled environments. Yi Fei Cheng 0001, Yukang Yan, Jan Gugenheimer, David Lindlbauer |
CHI | 3 |
| 2022 | Color-to-Depth Mappings as Depth Cues in Virtual RealityabstractDespite significant improvements to Virtual Reality (VR) technologies, most VR displays are fixed focus and depth perception is still a key issue that limits the user experience and the interaction performance. To supplement humans’ inherent depth cues (e.g., retinal blur, motion parallax), we investigate users’ perceptual mappings of distance to virtual objects’ appearance to generate visual cues aimed to enhance depth perception. As a first step, we explore color-to-depth mappings for virtual objects so that their appearance differs in saturation and value to reflect their distance. Through a series of controlled experiments, we elicit and analyze users’ strategies of mapping a virtual object’s hue, saturation, value and a combination of saturation and value to its depth. Based on the collected data, we implement a computational model that generates color-to-depth mappings fulfilling adjustable requirements on confusion probability, number of depth levels, and consistent saturation/value changing tendency. We demonstrate the effectiveness of color-to-depth mappings in a 3D sketching task, showing that compared to single-colored targets and strokes, with our mappings, the users were more confident in the accuracy without extra cognitive load and reduced the perceived depth error by 60.8%. We also implement four VR applications and demonstrate how our color cues can benefit the user experience and interaction performance in VR. Zhipeng Li 0001, Yikai Cui, Tianze Zhou, Yu Jiang 0010, Yuntao Wang 0001, Yukang Yan, Michael Nebeling, Yuanchun Shi |
UIST | 6 |
| 2021 | ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single MicrophoneabstractWake-up-free techniques (e.g., Raise-to-Speak) are important for improving the voice input experience. We present ProxiMic, a close-to-mic (within 5 cm) speech sensing technique using only one microphone. With ProxiMic, a user keeps a microphone-embedded device close to the mouth and speaks directly to the device without wake-up phrases or button presses. To detect close-to-mic speech, we use the feature from pop noise observed when a user speaks and blows air onto the microphone. Sound input is first passed through a low-pass adaptive threshold filter, then analyzed by a CNN which detects subtle close-to-mic features (mainly pop noise). Our two-stage algorithm can achieve 94.1% activation recall, 12.3 False Accepts per Week per User (FAWU) with 68 KB memory size, which can run at 352 fps on the smartphone. The user study shows that ProxiMic is efficient, user-friendly, and practical. Chun Yu, Zhaoheng Li, Mingyuan Zhong 0001, Yukang Yan, Yuanchun Shi |
CHI | 5 |
| 2021 | FaceSight: Enabling Hand-to-Face Gesture Interaction on AR Glasses with a Downward-Facing Camera VisionabstractWe present FaceSight, a computer vision-based hand-to-face gesture sensing technique for AR glasses. FaceSight fixes an infrared camera onto the bridge of AR glasses to provide extra sensing capability of the lower face and hand behaviors. We obtained 21 hand-to-face gestures and demonstrated the potential interaction benefits through five AR applications. We designed and implemented an algorithm pipeline that segments facial regions, detects hand-face contact (f1 score: 98.36%), and trains convolutional neural network (CNN) models to classify the hand-to-face gestures. The input features include gesture recognition, nose deformation estimation, and continuous fingertip movement. Our algorithm achieves classification accuracy of all gestures at 83.06%, proved by the data of 10 users. Due to the compact form factor and rich gestures, we recognize FaceSight as a practical solution to augment input capability of AR glasses in the future. Yueting Weng, Chun Yu, Yingtian Shi, Yuhang Zhao 0001, Yukang Yan, Yuanchun Shi |
CHI | 5 |
| 2021 | HulaMove: Using Commodity IMU for Waist InteractionabstractWe present HulaMove, a novel interaction technique that leverages the movement of the waist as a new eyes-free and hands-free input method for both the physical world and the virtual world. We first conducted a user study (N=12) to understand users’ ability to control their waist. We found that users could easily discriminate eight shifting directions and two rotating orientations, and quickly confirm actions by returning to the original position (quick return). We developed a design space with eight gestures for waist interaction based on the results and implemented an IMU-based real-time system. Using a hierarchical machine learning model, our system could recognize waist gestures at an accuracy of 97.5%. Finally, we conducted a second user study (N=12) for usability testing in both real-world scenarios and virtual reality settings. Our usability study indicated that HulaMove significantly reduced interaction time by 41.8% compared to a touch screen method, and greatly improved users’ sense of presence in the virtual world. This novel technique provides an additional input method when users’ eyes or hands are busy, accelerates users’ daily operations, and augments their immersive experience in the virtual world. Xuhai Xu, Tianyi Yuan, Liang He 0005, Xin Liu 0034, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey |
CHI | 6 |
| 2021 | SemanticAdapt: Optimization-based Adaptation of Mixed Reality Layouts Leveraging Virtual-Physical Semantic ConnectionsabstractWe present an optimization-based approach that automatically adapts Mixed Reality (MR) interfaces to different physical environments. Current MR layouts, including the position and scale of virtual interface elements, need to be manually adapted by users whenever they move between environments, and whenever they switch tasks. This process is tedious and time consuming, and arguably needs to be automated for MR systems to be beneficial for end users. We contribute an approach that formulates this challenge as a combinatorial optimization problem and automatically decides the placement of virtual interface elements in new environments. To achieve this, we exploit the semantic association between the virtual interface elements and physical objects in an environment. Our optimization furthermore considers the utility of elements for users’ current task, layout factors, and spatio-temporal consistency to previous layouts. All those factors are combined in a single linear program, which is used to adapt the layout of MR interfaces in real time. We demonstrate a set of application scenarios, showcasing the versatility and applicability of our approach. Finally, we show that compared to a naive adaptive baseline approach that does not take semantic associations into account, our approach decreased the number of manual interface adaptations by 33%. Yi Fei Cheng 0001, Yukang Yan, Xin Yi 0001, Yuanchun Shi, David Lindlbauer |
UIST | 2 |
| 2021 | ReflecTrack: Enabling 3D Acoustic Position Tracking Using Commodity Dual-Microphone Smartphonesabstract3D position tracking on smartphones has the potential to unlock a variety of novel applications, but has not been made widely available due to limitations in smartphone sensors. In this paper, we propose ReflecTrack, a novel 3D acoustic position tracking method for commodity dual-microphone smartphones. A ubiquitous speaker (e.g., smartwatch or earbud) generates inaudible Frequency Modulated Continuous Wave (FMCW) acoustic signals that are picked up by both smartphone microphones. To enable 3D tracking with two microphones, we introduce a reflective surface that can be easily found in everyday objects near the smartphone. Thus, the microphones can receive sound from the speaker and echoes from the surface for FMCW-based acoustic ranging. To simultaneously estimate the distances from the direct and reflective paths, we propose the echo-aware FMCW technique with a new signal pattern and target detection process. Our user study shows that ReflecTrack achieves a median error of 28.4 mm in the 60cm × 60cm × 60cm space and 22.1 mm in the 30cm × 30cm × 30cm space for 3D positioning. We demonstrate the easy accessibility of ReflecTrack using everyday surfaces and objects with several typical applications of 3D position tracking, including 3D input for smartphones, fine-grained gesture recognition, and motion tracking in smartphone-based VR systems. Yuzhou Zhuang, Yuntao Wang 0001, Yukang Yan, Xuhai Xu, Yuanchun Shi |
UIST | 3 |
| 2020 | EarBuddy: Enabling On-Face Interaction via Wireless EarbudsabstractPast research regarding on-body interaction typically requires custom sensors, limiting their scalability and generalizability. We propose EarBuddy, a real-time system that leverages the microphone in commercial wireless earbuds to detect tapping and sliding gestures near the face and ears. We develop a design space to generate 27 valid gestures and conducted a user study (N=16) to select the eight gestures that were optimal for both human preference and microphone detectability. We collected a dataset on those eight gestures (N=20) and trained deep learning models for gesture detection and classification. Our optimized classifier achieved an accuracy of 95.3%. Finally, we conducted a user study (N=12) to evaluate EarBuddy's usability. Our results show that EarBuddy can facilitate novel interaction and that users feel very positively about the system. EarBuddy provides a new eyes-free, socially acceptable input method that is compatible with commercial wireless earbuds and has the potential for scalability and generalizability Xuhai Xu, Haitian Shi, Xin Yi 0001, Wenjia Liu, Yukang Yan, Yuanchun Shi, Alexander Mariakakis, Jennifer Mankoff, Anind K. Dey |
CHI | 5 |
| 2020 | FrownOnError: Interrupting Responses from Smart Speakers by Facial ExpressionsabstractIn the conversations with smart speakers, misunderstandings of users' requests lead to erroneous responses. We propose FrownOnError, a novel interaction technique that enables users to interrupt the responses by intentional but natural facial expressions. This method leverages the human nature that the facial expression changes when we receive unexpected responses. We conducted a first user study (N=12) to understand users' intuitive reactions to the correct and incorrect responses. Our results reveal the significant difference in the frequency of occurrence and intensity of users' facial expressions between two conditions, and frowning and raising eyebrows are intuitive to perform and easy to control. Our second user study (N=16) evaluated the user experience and interruption efficiency of FrownOnError and the third user study (N=12) explored suitable conversation recovery strategies after the interruptions. Our results show that FrownOnError can be accurately detected (precision: 97.4%, recall: 97.6%), provides the most timely interruption compared to the baseline methods of wake-up word and button press, and is rated as most intuitive and easiest to be performed by users. Yukang Yan, Chun Yu, Wengrui Zheng, Ruining Tang, Xuhai Xu, Yuanchun Shi |
CHI | 1 |
| 2019 | PrivateTalk: Activating Voice Input with Hand-On-Mouth Gesture Detected by Bluetooth EarphonesabstractWe introduce PrivateTalk, an on-body interaction technique that allows users to activate voice input by performing the Hand-On-Mouth gesture during speaking. The gesture is performed as a hand partially covering the mouth from one side. PrivateTalk provides two benefits simultaneously. First, it enhances privacy by reducing the spread of voice while also concealing the lip movements from the view of other people in the environment. Second, the simple gesture removes the need for speaking wake-up words and is more accessible than a physical/software button especially when the device is not in the user's hands. To recognize the Hand-On-Mouth gesture, we propose a novel sensing technique that leverages the difference of signals received by two Bluetooth earphones worn on the left and right ear. Our evaluation shows that the gesture can be accurately detected and users consistently like PrivateTalk and consider it intuitive and effective. Yukang Yan, Chun Yu, Yingtian Shi, Minxing Xie |
UIST | 1 |
| 2019 | Gesture-based target acquisition in virtual and augmented realityabstractBackground Gesture is a basic interaction channel that is frequently used by humans to communicate in daily life. In this paper, we explore to use gesture-based approaches for target acquisition in virtual and augmented reality. A typical process of gesture-based target acquisition is: when a user intends to acquire a target, she performs a gesture with her hands, head or other parts of the body, the computer senses and recognizes the gesture and infers the most possible target. Methods We build mental model and behavior model of the user to study two key parts of the interaction process. Mental model describes how user thinks up a gesture for acquiring a target, and can be the intuitive mapping between gestures and targets. Behavior model describes how user moves the body parts to perform the gestures, and the relationship between the gesture that user intends to perform and signals that computer senses. Results In this paper, we present and discuss three pieces of research that focus on the mental model and behavior model of gesture-based target acquisition in VR and AR. Conclusions We show that leveraging these two models, interaction experience and performance can be improved in VR and AR environments. Yukang Yan, Xin Yi 0001, Chun Yu, Yuanchun Shi |
Virtual Real. Intell. Hardw. | 1 |
| 2018 | Eyes-Free Target Acquisition in Interaction Space around the Body for Virtual RealityabstractEyes-free target acquisition is a basic and important human ability to interact with the surrounding physical world, relying on the sense of space and proprioception. In this research, we leverage this ability to improve interaction in virtual reality (VR), by allowing users to acquire a virtual object without looking at it. We expect this eyes-free approach can effectively reduce head movements and focus changes, so as to speed up the interaction and alleviate fatigue and VR sickness. We conduct three lab studies to progressively investigate the feasibility and usability of eyes-free target acquisition in VR. Results show that, compared with the eyes-engaged manner, the eyes-free approach is significantly faster, provides satisfying accuracy, and introduces less fatigue and sickness; Most participants (13/16) prefer this approach. We also measure the accuracy of motion control and evaluate subjective experience of users when acquiring targets at different locations around the body. Based on the results, we make suggestions on designing appropriate target layout and discuss several design issues for eyes-free target acquisition in VR. Yukang Yan, Chun Yu, Xiaojuan Ma, Yuanchun Shi |
CHI | 1 |
| 2018 | VirtualGrasp: Leveraging Experience of Interacting with Physical Objects to Facilitate Digital Object RetrievalabstractWe propose VirtualGrasp, a novel gestural approach to retrieve virtual objects in virtual reality. Using VirtualGrasp, a user retrieves an object by performing a barehanded gesture as if grasping its physical counterpart. The object-gesture mapping under this metaphor is of high intuitiveness, which enables users to easily discover, remember the gestures to retrieve the objects. We conducted three user studies to demonstrate the feasibility and effectiveness of the approach. Progressively, we investigated the consensus of the object-gesture mapping across users, the expressivity of grasping gestures, and the learnability and performance of the approach. Results showed that users achieved high agreement on the mapping, with an average agreement score [35] of 0.68 (SD=0.27). Without exposure to the gestures, users successfully retrieved 76% objects with VirtualGrasp. A week after learning the mapping, they could recall the gestures for 93% objects. Yukang Yan, Chun Yu, Xiaojuan Ma, Xin Yi 0001, Ke Sun 0003, Yuanchun Shi |
CHI | 1 |
| 2017 | Float: One-Handed and Touch-Free Target Selection on SmartwatchesabstractTouch interaction on smartwatches suffers from the awkwardness of having to use two hands and the "fat finger" problem. We present Float, a wrist-to-finger input approach that enables one-handed and touch-free target selection on smartwatches with high efficiency and precision using only commercially-available built-in sensors. With Float, a user tilts the wrist to point and performs an in-air finger tap to click. To realize Float, we first explore the appropriate motion space for wrist tilt and determine the clicking action (finger tap) through a user-elicitation study. We combine the photoplethysmogram (PPG) signal with accelerometer and gyroscope to detect finger taps with a recall of 97.9% and a false discovery rate of 0.4%. Experiments show that using just one hand, Float allows users to acquire targets with size ranging from 2mm to 10mm in less than 2s to 1s, meanwhile achieve much higher accuracy than direct touch in both stationary (>98.9%) and walking (>71.5%) contexts. Ke Sun 0003, Yuntao Wang 0001, Chun Yu, Yukang Yan, Hongyi Wen, Yuanchun Shi |
CHI | 4 |