VLDB 2026 Research / reviewers in the wild / expert
Yuntao Wang 0001
dblp:52/4107-1
· DBLP profile ↗
44ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0002-4249-8893ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 37 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Does Personalized Nudging Wear Off? A Longitudinal Study of AI Self-Modeling for Behavioral EngagementabstractSustaining the effectiveness of behavior change technologies remains a key challenge. AI self-modeling, which generates personalized portrayals of one’s ideal self, has shown promise for motivating behavior change, yet prior work largely examines short-term effects. We present one of the first longitudinal evaluations of AI self-modeling in fitness engagement through a two-stage empirical study. A 1-week, three-arm experiment (visual self-modeling (VSM), auditory self-modeling (ASM), Control; N=28) revealed that VSM drove initial performance gains, while ASM showed no significant effects. A subsequent 4-week study (VSM vs. Control; N=31) demonstrated that VSM sustained higher performance levels but exhibited diminishing improvement rates after two weeks. Interviews uncovered a catalyst effect that fostered early motivation through clear, attainable goals, followed by habituation and internalization which stabilized performance. These findings highlight the temporal dynamics of personalized nudging and inform the design of behavior change technologies for long-term engagement. Yuzhou Du, Jiahuan Ding, Yuanchun Shi, Yuntao Wang 0001 |
CHI | 6 |
| 2026 | Routine Computing: A Systematic Review of Sensing Daily Life Dimensions Towards Human-Centered GoalsabstractHuman routines structure daily life, yet remain challenging for computational systems to understand. This paper presents the first systematic review of routine computing, a previously implicit but increasingly recognized field that focuses on computationally sensing and modeling human behaviors. It synthesizes 203 studies published up to August 2025. The paper presents a new taxonomy of the literature, focusing on temporal structures, behavioral interactions, cognitive aspects, and how variability and deviations are addressed. The common goals of routine computing extend across four major application domains, including accessibility care, the promotion of healthy habits, adaptive and context-aware support, and large-scale population insights. Persistent challenges that limit the design of truly human-centered systems are identified, including the gap between low-level activity recognition and high-level intent, the tension between personalization and generalization, unresolved privacy concerns, and data-related limitations. By consolidating these findings, this paper provides a foundational framework for HCI researchers, outlining principles for designing ethical, adaptive, and human-centered routine-aware systems. Borislav Pavlov, Yuntao Wang 0001, Yuanchun Shi |
CHI | 4 |
| 2026 | ActivitySeeker: Towards Collaborative Personalized Human Activity Discovery and Recognition on SmartphonesabstractSmartphones provide an attractive yet challenging platform for human activity recognition (HAR). They are ubiquitous, but also limit the input of HAR systems to a single IMU. These systems are also challenged by the inherent diversity of human activities and varying phone placement on the user’s body. This results in traditional smartphone HAR systems having limited personalization potential or imposing a high user burden. We propose ActivitySeeker, a personalized smartphone HAR system that combines self-supervised activity discovery and low-burden user interaction to collaboratively label IMU data and adapt HAR models to individual users on-device through transfer learning. We evaluated ActivitySeeker through simulated online learning and in-the-wild user experiments, where it discovered 95.5% of personal activity types and achieved high recognition accuracy (93.3%) while maintaining a positive user experience. Leveraging the synergy between user and smartphone, ActivitySeeker opens up new possibilities for HAR-based applications like fitness, health and personalized recommendation. Zhoutong Ye, Yanwen Huang, Chun Yu, Yuntao Wang 0001, Yuqi Luo, Yuanchun Shi |
CHI | 4 |
| 2026 | Enabling Adaptive Cardio-Respiratory Biofeedback Training on Ubiquitous Hand-Worn DevicesabstractWe introduce an adaptive cardio-respiratory biofeedback system implemented on ubiquitous hand-worn devices such as smart watches and rings, enabling accessible and real-time physiological training outside clinical settings. Users place a hand on their abdomen to promote embodied awareness of breathing rhythms, while PPG and IMU sensors continuously capture cardio-respiratory signals. Unlike conventional open-loop biofeedback that delivers fixed breathing guidance irrespective of user response, our system employs a closed-loop adaptation: real-time physiological signals adjust breathing cues to optimize cardio-respiratory coupling, ensuring personalized training trajectories. This shift from static to adaptive guidance markedly improves user engagement and training efficacy. A user performance evaluation study further showed that adaptive biofeedback significantly boosts HRV, prolongs high-HRV states, and enhances user experience, demonstrating clear advantages over non-adaptive methods. Together, these findings position adaptive, hand-worn biofeedback as a promising approach for ubiquitous, user-centered mental health interventions. Ruotong Yu, Xintong Wu, Lily Sheng, Yuntao Wang 0001, Yuanchun Shi |
CHI | 5 |
| 2026 | LubDubDecoder: Bringing Micro-Mechanical Cardiac Monitoring to HearablesabstractWe present LubDubDecoder, a system that enables fine-grained monitoring of micro-cardiac vibrations associated with the opening and closing of heart valves across a range of hearables. Our system transforms the built-in speaker, the only transducer common to all hearables, into an acoustic sensor that captures the coarse “lub-dub” heart sounds, leverages their shared temporal and spectral structure to reconstruct the subtle seismocardiography (SCG) and gyrocardiography (GCG) waveforms, and extract the timing of key micro-cardiac events. In an IRB-approved feasibility study with 25 users, our system achieves correlations of 0.88–0.95 compared to chest-mounted reference measurements in within-user and cross-user evaluations, and generalizes to unseen hearables using a zero-effort adaptation scheme with a correlation of 0.91. Our system is robust across remounting sessions and music playback. Xiyuxing Zhang, Duc Nguyen Tien Vu, Tao Qiang, Clara Palacios, Jiangyifei Zhu, Yuntao Wang 0001, Mayank Goel, Justin Chan |
CHI | 7 |
| 2025 | Non-Contact Health Monitoring During Daily Personal Care RoutinesabstractRemote photoplethysmography (rPPG) enables noncontact, continuous monitoring of physiological signals and offers a practical alternative to traditional health sensing methods. Although rPPG is promising for daily health monitoring, its application in long-term personal care scenarios-such as mirrorfacing routines in high-altitude environments-remains challenging due to ambient lighting variations, frequent occlusions from hand movements, and dynamic facial postures. To address these challenges, we present the Long-term Altitude Daily Health (LADH) dataset, the first long-term rPPG dataset containing 240 synchronized RGB and infrared (IR) facial videos from 21 participants across five common personal care scenarios, along with ground-truth PPG, respiration, and blood oxygen signals. Our experiments demonstrate that combining RGB and IR video inputs improves the accuracy and robustness of non-contact physiological monitoring, achieving a mean absolute error (MAE) of 4.99 BPM in heart rate estimation. Furthermore, we find that multi-task learning enhances performance across multiple physiological indicators simultaneously. Dataset and code are open at https://github.com/McJackTang/FusionVitals. Xulin Ma, Jiankai Tang, Zhang Jiang, Songqin Cheng, Yuanchun Shi, Xin Liu 0034, Daniel McDuff, Yuntao Wang 0001 |
BSN | 10 |
| 2025 | Modeling the Impact of Visual Stimuli on Redirection Noticeability with Gaze Behavior in Virtual Reality
Zhipeng Li 0001, Yishu Ji, Ruijia Chen, Yuntao Wang 0001, Yuanchun Shi, Yukang Yan |
CHI | 5 |
| 2025 | The Odyssey Journey: Top-Tier Medical Resource Seeking for Specialized Disorder in ChinaabstractIt is pivotal for patients to receive accurate health information, diagnoses, and timely treatments. However, in China, the significant imbalanced doctor-to-patient ratio intensifies the information and power asymmetries in doctor-patient relationships. Health information-seeking, which enables patients to collect information from sources beyond doctors, is a potential approach to mitigate these asymmetries. While HCI research predominantly focuses on common chronic conditions, our study focuses on specialized disorders, which are often familiar to specialists but not to general practitioners and the public. With Hemifacial Spasm (HFS) as an example, we aim to understand patients' health information and top-tier1 medical resource seeking journeys in China. Through interviews with three neurosurgeons and 12 HFS patients from rural and urban areas, and applying Actor-Network Theory, we provide empirical insights into the roles, interactions, and workflows of various actors in the health information-seeking network. We also identified five strategies patients adopted to mitigate asymmetries and access top-tier medical resources, illustrating these strategies as subnetworks within the broader health information-seeking network and outlining their advantages and challenges. © 2025 Copyright held by the owner/author(s). Ka I Chan, Siying Hu, Yuntao Wang 0001, Xuhai Xu, Zhicong Lu, Yuanchun Shi |
CHI | 3 |
| 2025 | Unknown Word Detection for English as a Second Language (ESL) Learners using Gaze and Pre-trained Language Models
Jiexin Ding, Bowen Zhao 0004, Yuntao Wang 0001, Xinyun Liu, Ishan Chatterjee, Yuanchun Shi |
CHI | 3 |
| 2025 | VAction: A Lightweight and Integrated VR Training System for Authentic Film-Shooting Experience
Che Qu, Minjing Yu, Chao Zhou 0012, Yuntao Wang 0001, Yu-Hui Wen, Yuanchun Shi, Yong-Jin Liu 0001 |
CHI | 5 |
| 2025 | BIT: Battery-free, IC-less and Wireless Smart Textile Interface and Sensing SystemabstractThe development of smart textile interfaces is hindered by the inclusion of rigid hardware components and batteries within the fabric, which pose challenges in terms of manufacturability, usability, and environmental concerns related to electronic waste. To mitigate these issues, we propose a smart textile interface and its wireless sensing system to eliminate the need for ICs, batteries, and connectors embedded into textiles. Our technique is established on the integration of multi-resonant circuits in smart textile interfaces, and utilizing near-field electromagnetic coupling between two coils to facilitate wireless power transfer and data acquisition from smart textile interface. A key aspect of our system is the development of a mathematical model that accurately represents the equivalent circuit of the sensing system. Using this model, we developed a novel algorithm to accurately estimate sensor signals based on changes in system impedance. Through simulation-based experiments and a user study, we demonstrate that our technique effectively supports multiple textile sensors of various types. Weiye Xu 0002, Tony Li, Yuntao Wang 0001, Xing-Dong Yang, Te-Yen Wu |
CHI | 3 |
| 2025 | Actual Achieved Gain and Optimal Perceived Gain: Modeling Human Take-over Decisions Towards Automated Vehicles' SuggestionsabstractDriver decision quality in take-overs is critical for effective human-Autonomous Driving System (ADS) collaboration.However, current research lacks detailed analysis of its variations.This paper Xin Yi 0001, Chuye Hong, Gujun Chen, Yongquan Hu, Yuntao Wang 0001, Hewu Li |
CHI | 9 |
| 2025 | Understanding Users' Perceptions and Expectations toward a Social Balloon Robot via an Exploratory Study
Tianyi Xia, Manqiu Liao, Yuan Gao 0024, Chun Yu, Yuntao Wang 0001, Yuanchun Shi |
UIST | 11 |
| 2025 | Predicting Ray Pointer Landing Poses in VR Using Multimodal LSTM-Based Neural NetworksabstractTarget selection is one of the most fundamental tasks in VR interaction systems. Prediction heuristics can provide users with a smoother interaction experience in this process. Our work aims to predict the ray landing pose for hand-based raycasting selection in Virtual Reality (VR) using a Long Short-Term Memory (LSTM)-based neural network with time-series data input of speed and distance over time from three different pose channels: hand, Head-Mounted Display (HMD), and eye. We first conducted a study to collect motion data from these three input channels and analyzed these movement behaviors. Additionally, we evaluated which combination of input modalities yields the optimal result. A second study validates raycasting across a continuous range of distances, angles, and target sizes. On average, our technique’s predictions were within 4.6° of the true landing Pose when 50% of the way through the movement. We compared our LSTM neural network model to a kinematic information model and further validated its generalizability in two ways: by training the model on one user’s data and testing on other users (cross-user) and by training on a group of users and testing on entirely new users (unseen users). Compared to the baseline and a previous kinematic method, our model increased prediction accuracy by a factor of 3.5 and 1.9, re spectively, when 40% of the way through the movement. Wenxuan Xu 0001, Yushi Wei, Xuning Hu, Wolfgang Stuerzlinger, Yuntao Wang 0001, Hai-Ning Liang |
VR | 5 |
| 2025 | A Comparison Study Understanding the Impact of Mixed Reality Collaboration on Sense of Co-PresenceabstractSense of co-presence refers to the perceived closeness and interaction between participants in a collaborative context, which critically impacts the collaboration experience and task performance. With the emergence of Mixed Reality (MR) technologies, we would like to investigate the effect of MR immersive collaboration environment on promoting co-presence in a remote setting by comparing it with non-MR methods, such as video conferencing. We conduct a comparison study, where we invited 14 dyads of participants to collaborate on block assembly tasks with video conferencing, MR system, and in a physically co-located scenario. Each participant of a dyad was assigned either a local worker to assemble the blocks or a remote helper to give the instructions. Results show that MR system can create comparable sense of co-presence with co-located situation, and allow users to interact more naturally with both the environment and each other. The adoption of mixed reality enhances collaboration and task performance by reducing reliance on verbal communication and favoring action-based interactions through gestures and direct manipulation of virtual objects. Weicheng Zheng, Yuntao Wang 0001, Xin Tong 0004, Yukang Yan |
VR | 3 |
| 2025 | Spiking-PhysFormer: Camera-based remote photoplethysmography with parallel spike-driven transformer
Mingxuan Liu 0001, Jiankai Tang, Yongli Chen, Jiahao Qi, Kegang Wang, Yuntao Wang 0001, Hong Chen 0002 |
Neural Networks | 9 |
| 2025 | FlowRing: Integrated Microgesture and Surface Interaction Ring for Versatile XR Input MHCI010abstractAs Extended Reality (XR) advances, a device has the potential to be used across contexts from immersive productivity at a desk to on-the-go, public scenarios. Existing input solutions lack the versatility to provide both high-throughput, mouse-grade input and subtle, ergonomic interaction. We introduce FlowRing, a novel ring-form device that combines microgestures with precise 2D mouse-like input on surfaces. FlowRing supports five microgestures for discreet interaction and 2D input for richer tasks, using an optical flow sensor, skin-contact microphone, and IMU at the base of the finger. In a study with 11 participants, FlowRing achieved 93.6% microgesture recognition accuracy across sessions and 85.2% across unseen users, rising to 90.1% with just four gesture set examples from a new user. A separate 2D Fitts’ law study demonstrated its effectiveness for continuous input on various surfaces. FlowRing emerges as a versatile, user-friendly solution for the future of interactive technology. Ishan Chatterjee, Jiexin Ding, Anandghan Waghmare, Joseph Breda, Yuquan Deng, Bo Liu 0091, Yuntao Wang 0001, Shwetak N. Patel |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2024 | ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled ExpressionsabstractSilent speech interaction (SSI) allows users to discreetly input text without using their hands. Existing wearable SSI systems typically require custom devices and are limited to a small lexicon, limiting their utility to a small set of command words. This work proposes ReHEarSSE, an earbud-based ultrasonic SSI system capable of generalizing to words that do not appear in its training dataset, providing support for nearly an entire dictionary’s worth of words. As a user silently spells words, ReHEarSSE uses autoregressive features to identify subtle changes in ear canal shape. ReHEarSSE infers words using a deep learning model trained to optimize connectionist temporal classification (CTC) loss with an intermediate embedding that accounts for different letters and transitions between them. We find that ReHEarSSE recognizes 100 unseen words with an accuracy of 89.3%. Xuefu Dong, Yifei Chen 0008, Yuuki Nishiyama, Kaoru Sezaki, Yuntao Wang 0001, Kenneth Christofferson, Alexander Mariakakis |
CHI | 5 |
| 2024 | Time2Stop: Adaptive and Explainable Human-AI Loop for Smartphone Overuse InterventionabstractDespite a rich history of investigating smartphone overuse intervention techniques, AI-based just-in-time adaptive intervention (JITAI) methods for overuse reduction are lacking. We develop Time2Stop, an intelligent, adaptive, and explainable JITAI system that leverages machine learning to identify optimal intervention timings, introduces interventions with transparent AI explanations, and collects user feedback to establish a human-AI loop and adapt the intervention model over time. We conducted an 8-week field experiment (N=71) to evaluate the effectiveness of both the adaptation and explanation aspects of Time2Stop. Our results indicate that our adaptive models significantly outperform the baseline methods on intervention accuracy (>32.8% relatively) and receptivity (>8.0%). In addition, incorporating explanations further enhances the effectiveness by 53.8% and 11.4% on accuracy and receptivity, respectively. Moreover, Time2Stop significantly reduces overuse, decreasing app visit frequency by 7.0 ∼ 8.9%. Our subjective data also echoed these quantitative measures. Participants preferred the adaptive interventions and rated the system highly on intervention time accuracy, effectiveness, and level of trust. We envision our work can inspire future research on JITAI systems with a human-AI loop to evolve with users. Adiba Orzikulova, Zhipeng Li 0001, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Marzyeh Ghassemi, Sung-Ju Lee 0001, Anind K. Dey, Xuhai Xu |
CHI | 5 |
| 2024 | PepperPose: Full-Body Pose Estimation with a Companion RobotabstractAccurate full-body pose estimation across diverse actions in a user-friendly and location-agnostic manner paves the way for interactive applications in realms like sports, fitness, and healthcare. This task becomes challenging in real-world scenarios due to factors like the user’s dynamic positioning, the diversity of actions, and the varying acceptability of the pose-capturing system. In this context, we present PepperPose, a novel companion robot system tailored for optimized pose estimation. Unlike traditional methods, PepperPose actively tracks the user and refines its viewpoint, facilitating enhanced pose accuracy across different locations and actions. This allows users to enjoy a seamless action-sensing experience. Our evaluation, involving 30 participants undertaking daily functioning and exercise actions in a home-like space, underscores the robot’s promising capabilities. Moreover, we demonstrate the opportunities that PepperPose presents for human-robot interaction, its current limitations, and future developments. Lingxiao Zhong, Chun Yu, Yuntao Wang 0001, Yuan Gao 0024, Tin Lun Lam, Yuanchun Shi |
CHI | 6 |
| 2024 | DreamCatcher: A Wearer-aware Multi-modal Sleep Event Dataset Based on Earables in Non-restrictive EnvironmentsabstractPoor quality sleep can be characterized by the occurrence of events ranging from body movement to breathing impairment. Widely available earbuds equipped with sensors (also known as earables) can be combined with a sleep event detection algorithm to offer a convenient alternative to laborious clinical tests for individuals suffering from sleep disorders. Although various solutions utilizing such devices have been proposed to detect sleep events, they ignore the fact that individuals often share sleeping spaces with roommates or couples. To address this issue, we introduce DreamCatcher, the first publicly available dataset for wearer-aware sleep event algorithm development on earables. DreamCatcher encompasses eight distinct sleep events, including synchronous dual-channel audio and motion data collected from 12 pairs (24 participants) totaling 210 hours (420 hour.person) with fine-grained label. We tested multiple benchmark models on three tasks related to sleep event detection, demonstrating the usability and unique challenge of DreamCatcher. We hope that the proposed DreamCatcher can inspire other researchers to further explore efficient wearer-aware human vocal activity sensing on earables. DreamCatcher is publicly available at https://github.com/thuhci/DreamCatcher. Xiyuxing Zhang, Ruotong Yu, Yuntao Wang 0001, Kenneth Christofferson, Jingru Zhang 0005, Alexander Mariakakis, Yuanchun Shi |
NeurIPS | 4 |
| 2024 | Voila-A: Aligning Vision-Language Models with User's Gaze AttentionabstractIn recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challenges in handling real-world applications with complex scenes and multiple objects, as well as aligning their focus with the diverse attention patterns of human users. In this paper, we introduce gaze information, feasibly collected by ubiquitous wearable devices such as MR glasses, as a proxy for human attention to guide VLMs. We propose a novel approach, Voila-A, for gaze alignment to enhance the effectiveness of these models in real-world applications. First, we collect hundreds of minutes of gaze data to demonstrate that we can mimic human gaze modalities using localized narratives. We then design an automatic data annotation pipeline utilizing GPT-4 to generate the VOILA-COCO dataset. Additionally, we introduce a new model VOILA-A that integrate gaze information into VLMs while maintain pretrained knowledge from webscale dataset. We evaluate Voila-A using a hold-out validation set and a newly collected VOILA-GAZE testset, which features real-life scenarios captured with a gaze-tracking device. Our experimental results demonstrate that Voila-A significantly outperforms several baseline models. By aligning model attention with human gaze patterns, Voila-A paves the way for more intuitive, user-centric VLMs and fosters engaging human-AI interaction across a wide range of applications. Kun Yan 0004, Lei Ji 0001, Yuntao Wang 0001, Nan Duan 0001, Shuai Ma 0001 |
NeurIPS | 4 |
| 2024 | HCI Research and Innovation in China: A 10-Year PerspectiveabstractIn the past years, human computer interaction (HCI) research and innovation have developed substantially, leading to a number of fruitful research topics. In this paper, we surveyed the HCI research and innovation in China from a 10-year perspective. We analyzed the popular research methodology and topics among Chinese researchers, including human modeling, user interface techniques, context awareness, user acceptance and performance, user experience design, human-AI interaction, HCI applications and social influences. We also conducted a bibliography analysis on the published papers in top-tier conferences and journals, which revealed a significant rising trend, and a generally broad distribution of research types. Moreover, we described typical applications and the industry influence of the research outcomes. We concluded with implications and reflections for HCI researchers across the world and shared the future research trends envisioned by Chinese researchers. Yuanchun Shi, Xin Yi 0001, Yuntao Wang 0001, Yukang Yan, Zhimin Cheng, Pengye Zhu, Yongjuan Li, Yanci Liu, Weixuan Zhou, Diya Zhao |
Int. J. Hum. Comput. Interact. | 5 |
| 2023 | Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution ImagesabstractA computer vision system using low-resolution image sensors can provide intelligent services (e.g., activity recognition) but preserve unnecessary visual privacy information from the hardware level. However, preserving visual privacy and enabling accurate machine recognition have adversarial needs on image resolution. Modeling the trade-off of privacy preservation and machine recognition performance can guide future privacy-preserving computer vision systems using low-resolution image sensors. In this paper, using the at-home activity of daily livings (ADLs) as the scenario, we first obtained the most important visual privacy features through a user survey. Then we quantified and analyzed the effects of image resolution on human and machine recognition performance in activity recognition and privacy awareness tasks. We also investigated how modern image super-resolution techniques influence these effects. Based on the results, we proposed a method for modeling the trade-off of privacy preservation and activity recognition on low-resolution images. Yuntao Wang 0001, Zirui Cheng, Xin Yi 0001, Yan Kong, Xuhai Xu, Yukang Yan, Chun Yu, Shwetak N. Patel, Yuanchun Shi |
CHI | 1 |
| 2023 | Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device SensingabstractGestures performed accompanying the voice are essential for voice interaction to convey complementary semantics for interaction purposes such as wake-up state and input modality. In this paper, we investigated voice-accompanying hand-to-face (VAHF) gestures for voice interaction. We targeted on hand-to-face gestures because such gestures relate closely to speech and yield significant acoustic features (e.g., impeding voice propagation). We conducted a user study to explore the design space of VAHF gestures, where we first gathered candidate gestures and then applied a structural analysis to them in different dimensions (e.g., contact position and type), outputting a total of 8 VAHF gestures with good usability and least confusion. To facilitate VAHF gesture recognition, we proposed a novel cross-device sensing method that leverages heterogeneous channels (vocal, ultrasound, and IMU) of data from commodity devices (earbuds, watches, and rings). Our recognition model achieved an accuracy of 97.3% for recognizing 3 gestures and 91.5% for recognizing 8 gestures (excluding the "empty" gesture), proving the high applicability. Quantitative analysis also shed light on the recognition capability of each sensor channel and their different combinations. In the end, we illustrated the feasible use cases and their design principles to demonstrate the applicability of our system in various scenarios. Zisu Li, Yuntao Wang 0001, Chun Yu, Yukang Yan, Mingming Fan 0001, Yuanchun Shi |
CHI | 3 |
| 2023 | rPPG-Toolbox: Deep Remote PPG ToolboxabstractCamera-based physiological measurement is a fast growing field of computer vision. Remote photoplethysmography (rPPG) utilizes imaging devices (e.g., cameras) to measure the peripheral blood volume pulse (BVP) via photoplethysmography, and enables cardiac measurement via webcams and smartphones. However, the task is non-trivial with important pre-processing, modeling and post-processing steps required to obtain state-of-the-art results. Replication of results and benchmarking of new models is critical for scientific progress; however, as with many other applications of deep learning, reliable codebases are not easy to find or use. We present a comprehensive toolbox, rPPG-Toolbox, unsupervised and supervised rPPG models with support for public benchmark datasets, data augmentation and systematic evaluation: https://github.com/ubicomplab/rPPG-Toolbox. Xin Liu 0034, Girish Narayanswamy, Akshay Paruchuri, Jiankai Tang, Roni Sengupta, Shwetak N. Patel, Yuntao Wang 0001, Daniel McDuff |
NeurIPS | 9 |
| 2023 | ConeSpeech: Exploring Directional Speech Interaction for Multi-Person Remote Communication in Virtual RealityabstractRemote communication is essential for efficient collaboration among people at different locations. We present ConeSpeech, a virtual reality (VR) based multi-user remote communication technique, which enables users to selectively speak to target listeners without distracting bystanders. With ConeSpeech, the user looks at the target listener and only in a cone-shaped area in the direction can the listeners hear the speech. This manner alleviates the disturbance to and avoids overhearing from surrounding irrelevant people. Three featured functions are supported, directional speech delivery, size-adjustable delivery range, and multiple delivery areas, to facilitate speaking to more than one listener and to listeners spatially mixed up with bystanders. We conducted a user study to determine the modality to control the cone-shaped delivery area. Then we implemented the technique and evaluated its performance in three typical multi-user communication tasks by comparing it to two baseline methods. Results show that ConeSpeech balanced the convenience and flexibility of voice communication. Yukang Yan, Haohua Liu, Yingtian Shi, Ruici Guo, Zisu Li, Xuhai Xu, Chun Yu, Yuntao Wang 0001, Yuanchun Shi |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2022 | FaceOri: Tracking Head Position and Orientation Using Ultrasonic Ranging on EarphonesabstractFace orientation can often indicate users’ intended interaction target. In this paper, we propose FaceOri, a novel face tracking technique based on acoustic ranging using earphones. FaceOri can leverage the speaker on a commodity device to emit an ultrasonic chirp, which is picked up by the set of microphones on the user’s earphone, and then processed to calculate the distance from each microphone to the device. These measurements are used to derive the user’s face orientation and distance with respect to the device. We conduct a ground truth comparison and user study to evaluate FaceOri’s performance. The results show that the system can determine whether the user orients to the device at a 93.5% accuracy within a 1.5 meters range. Furthermore, FaceOri can continuously track user’s head orientation with a median absolute error of 10.9 mm in the distance, 3.7° in yaw, and 5.8° in pitch. FaceOri can allow for convenient hands-free control of devices and produce more intelligent context-aware interactions. Yuntao Wang 0001, Jiexin Ding, Ishan Chatterjee, Farshid Salemi Parizi, Yuzhou Zhuang, Yukang Yan, Shwetak N. Patel, Yuanchun Shi |
CHI | 1 |
| 2022 | TypeOut: Leveraging Just-in-Time Self-Affirmation for Smartphone Overuse ReductionabstractSmartphone overuse is related to a variety of issues such as lack of sleep and anxiety. We explore the application of Self-Affirmation Theory on smartphone overuse intervention in a just-in-time manner. We present TypeOut, a just-in-time intervention technique that integrates two components: an in-situ typing-based unlock process to improve user engagement, and self-affirmation-based typing content to enhance effectiveness. We hypothesize that the integration of typing and self-affirmation content can better reduce smartphone overuse. We conducted a 10-week within-subject field experiment (N=54) and compared TypeOut against two baselines: one only showing the self-affirmation content (a common notification-based intervention), and one only requiring typing non-semantic content (a state-of-the-art method). TypeOut reduces app usage by over 50%, and both app opening frequency and usage duration by over 25%, all significantly outperforming baselines. TypeOut can potentially be used in other domains where an intervention may benefit from integrating self-affirmation exercises with an engaging just-in-time mechanism. Xuhai Xu, Tianyuan Zou, Yanzhang Li, Ruolin Wang, Tianyi Yuan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey |
CHI | 7 |
| 2022 | Color-to-Depth Mappings as Depth Cues in Virtual RealityabstractDespite significant improvements to Virtual Reality (VR) technologies, most VR displays are fixed focus and depth perception is still a key issue that limits the user experience and the interaction performance. To supplement humans’ inherent depth cues (e.g., retinal blur, motion parallax), we investigate users’ perceptual mappings of distance to virtual objects’ appearance to generate visual cues aimed to enhance depth perception. As a first step, we explore color-to-depth mappings for virtual objects so that their appearance differs in saturation and value to reflect their distance. Through a series of controlled experiments, we elicit and analyze users’ strategies of mapping a virtual object’s hue, saturation, value and a combination of saturation and value to its depth. Based on the collected data, we implement a computational model that generates color-to-depth mappings fulfilling adjustable requirements on confusion probability, number of depth levels, and consistent saturation/value changing tendency. We demonstrate the effectiveness of color-to-depth mappings in a 3D sketching task, showing that compared to single-colored targets and strokes, with our mappings, the users were more confident in the accuracy without extra cognitive load and reduced the perceived depth error by 60.8%. We also implement four VR applications and demonstrate how our color cues can benefit the user experience and interaction performance in VR. Zhipeng Li 0001, Yikai Cui, Tianze Zhou, Yu Jiang 0010, Yuntao Wang 0001, Yukang Yan, Michael Nebeling, Yuanchun Shi |
UIST | 5 |
| 2022 | GazeDock: Gaze-Only Menu Selection in Virtual Reality using Auto-Triggering Peripheral MenuabstractGaze-only input techniques in VR face the challenge of avoiding false triggering due to continuous eye tracking while maintaining interaction performance. In this paper, we proposed GazeDock, a technique for enabling fast and robust gaze-based menu selection in VR. GazeDock features a view-fixed peripheral menu layout that automatically triggers appearing and selection when the user’s gaze approaches and leaves the menu zone, thus facilitating interaction speed and minimizing the false triggering rate. We built a dataset of 12 participants’ natural gaze movements in typical VR applications. By analyzing their gaze movement patterns, we designed the menu UI personalization and optimized selection detection algorithm of GazeDock. We also examined users’ gaze selection precision for targets on the peripheral menu and found that 4–8 menu items yield the highest throughput when considering both speed and accuracy. Finally, we validated the usability of GazeDock in a VR navigation game that contains both scene exploration and menu selection. Results showed that GazeDock achieved an average selection time of 471ms and a false triggering rate of 3.6%. And it received higher user preference ratings compared with dwell-based and pursuit-based techniques. Xin Yi 0001, Yiqin Lu, Ziyin Cai, Zihan Wu 0002, Yuntao Wang 0001, Yuanchun Shi |
VR | 5 |
| 2022 | Easily-add battery-free wireless sensors to everyday objects: system implementation and usability study
Tengxiang Zhang, Zi Qian, Hsuan-Wei Fan, Jie Ren 0017, Yuntao Wang 0001, Yuanchun Shi |
CCF Trans. Pervasive Comput. Interact. | 5 |
| 2021 | Understanding the Design Space of Mouth MicrogesturesabstractAs wearable devices move toward the face (i.e. smart earbuds, glasses), there is an increasing need to facilitate intuitive interactions with these devices. Current sensing techniques can already detect many mouth-based gestures; however, users’ preferences of these gestures are not fully understood. In this paper, we investigate the design space and usability of mouth-based microgestures. We first conducted brainstorming sessions (N=16) and compiled an extensive set of 86 user-defined gestures. Then, with an online survey (N=50), we assessed the physical and mental demand of our gesture set and identified a subset of 14 gestures that can be performed easily and naturally. Finally, we conducted a remote Wizard-of-Oz usability study (N=11) mapping gestures to various daily smartphone operations under a sitting and walking context. From these studies, we develop a taxonomy for mouth gestures, finalize a practical gesture set for common applications, and provide design guidelines for future mouth-based gesture interactions. Xuhai Xu, Richard Li 0002, Yuanchun Shi, Shwetak N. Patel, Yuntao Wang 0001 |
Conference on Designing Interactive Systems | 6 |
| 2021 | Facilitating Text Entry on Smartphones with QWERTY Keyboard for Users with Parkinson's DiseaseabstractQWERTY is the primary smartphone text input keyboard configuration. However, insertion and substitution errors caused by hand tremors, often experienced by users with Parkinson’s disease, can severely affect typing efficiency and user experience. In this paper, we investigated Parkinson’s users’ typing behavior on smartphones. In particular, we identified and compared the typing characteristics generated by users with and without Parkinson’s symptoms. We then proposed an elastic probabilistic model for input prediction. By incorporating both spatial and temporal features, this model generalized the classical statistical decoding algorithm to correct insertion, substitution and omission errors, while maintaining direct physical interpretation. User study results confirmed that the proposed algorithm outperformed baseline techniques: users reached 22.8 WPM typing speed with a significantly lower error rate and higher user-perceived performance and preference. We concluded that our method could effectively improve the text entry experience on smartphones for users with Parkinson’s disease. Yuntao Wang 0001, Ao Yu, Xin Yi 0001, Yuanwei Zhang, Ishan Chatterjee, Shwetak N. Patel, Yuanchun Shi |
CHI | 1 |
| 2021 | Auth+Track: Enabling Authentication Free Interaction on Smartphone by Continuous User TrackingabstractWe propose Auth+Track, a novel authentication model that aims to reduce redundant authentication in everyday smartphone usage. By sparse authentication and continuous tracking of the user’s status, Auth+Track eliminates the “gap” authentication between fragmented sessions and enables “Authentication Free when User is Around”. To instantiate the Auth+Track model, we present PanoTrack, a prototype that integrates body and near field hand information for user tracking. We install a fisheye camera on the top of the phone to achieve a panoramic vision that can capture both user’s body and on-screen hands. Based on the captured video stream, we develop an algorithm to extract 1) features for user tracking, including body keypoints and their temporal and spatial association, near field hand status, and 2) features for user identity assignment. The results of our user studies validate the feasibility of PanoTrack and demonstrate that Auth+Track not only improves the authentication efficiency but also enhances user experiences with better usability. Chun Yu, Xiaoying Wei, Xuhai Xu, Yongquan Hu, Yuntao Wang 0001, Yuanchun Shi |
CHI | 6 |
| 2021 | HulaMove: Using Commodity IMU for Waist InteractionabstractWe present HulaMove, a novel interaction technique that leverages the movement of the waist as a new eyes-free and hands-free input method for both the physical world and the virtual world. We first conducted a user study (N=12) to understand users’ ability to control their waist. We found that users could easily discriminate eight shifting directions and two rotating orientations, and quickly confirm actions by returning to the original position (quick return). We developed a design space with eight gestures for waist interaction based on the results and implemented an IMU-based real-time system. Using a hierarchical machine learning model, our system could recognize waist gestures at an accuracy of 97.5%. Finally, we conducted a second user study (N=12) for usability testing in both real-world scenarios and virtual reality settings. Our usability study indicated that HulaMove significantly reduced interaction time by 41.8% compared to a touch screen method, and greatly improved users’ sense of presence in the virtual world. This novel technique provides an additional input method when users’ eyes or hands are busy, accelerates users’ daily operations, and augments their immersive experience in the virtual world. Xuhai Xu, Tianyi Yuan, Liang He 0005, Xin Liu 0034, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey |
CHI | 7 |
| 2021 | ReflecTrack: Enabling 3D Acoustic Position Tracking Using Commodity Dual-Microphone Smartphonesabstract3D position tracking on smartphones has the potential to unlock a variety of novel applications, but has not been made widely available due to limitations in smartphone sensors. In this paper, we propose ReflecTrack, a novel 3D acoustic position tracking method for commodity dual-microphone smartphones. A ubiquitous speaker (e.g., smartwatch or earbud) generates inaudible Frequency Modulated Continuous Wave (FMCW) acoustic signals that are picked up by both smartphone microphones. To enable 3D tracking with two microphones, we introduce a reflective surface that can be easily found in everyday objects near the smartphone. Thus, the microphones can receive sound from the speaker and echoes from the surface for FMCW-based acoustic ranging. To simultaneously estimate the distances from the direct and reflective paths, we propose the echo-aware FMCW technique with a new signal pattern and target detection process. Our user study shows that ReflecTrack achieves a median error of 28.4 mm in the 60cm × 60cm × 60cm space and 22.1 mm in the 30cm × 30cm × 30cm space for 3D positioning. We demonstrate the easy accessibility of ReflecTrack using everyday surfaces and objects with several typical applications of 3D position tracking, including 3D input for smartphones, fine-grained gesture recognition, and motion tracking in smartphone-based VR systems. Yuzhou Zhuang, Yuntao Wang 0001, Yukang Yan, Xuhai Xu, Yuanchun Shi |
UIST | 2 |
| 2020 | MoveVR: Enabling Multiform Force Feedback in Virtual Reality using Household Cleaning RobotabstractHaptic feedback can significantly enhance the realism and immersiveness of virtual reality (VR) systems. In this paper, we propose MoveVR, a technique that enables realistic, multiform force feedback in VR leveraging commonplace cleaning robots. MoveVR can generate tension, resistance, impact and material rigidity force feedback with multiple levels of force intensity and directions. This is achieved by changing the robot's moving speed, rotation, position as well as the carried proxies. We demonstrated the feasibility and effectiveness of MoveVR through interactive VR gaming. In our quantitative and qualitative evaluation studies, participants found that MoveVR provides more realistic and enjoyable user experience when compared to commercially available haptic solutions such as vibrotactile haptic systems. Yuntao Wang 0001, Zichao (Tyson) Chen, Hanchuan Li, Zhengyi Cao, Huiyi Luo, Tengxiang Zhang, Ke Ou, John Raiti, Chun Yu, Shwetak N. Patel, Yuanchun Shi |
CHI | 1 |
| 2020 | ThermalRing: Gesture and Tag Inputs Enabled by a Thermal Imaging Smart RingabstractThe heterogeneous and ubiquitous input demands in smart spaces call for an input device that can enable rich and spontaneous interactions. We propose ThermalRing, a thermal imaging smart ring using low-resolution thermal camera for identity-anonymous, illumination-invariant, and power-efficient sensing of both dynamic and static gestures. We also design ThermalTag, thin and passive thermal imageable tags that reflect the heat from the human hand. ThermalTag can be easily made and applied onto everyday objects by users. We develop sensing techniques for three typical input demands: drawing gestures for device pairing, click and slide gestures for device control, and tag scan gestures for quick access. The study results show that ThermalRing can recognize nine drawing gestures with an overall accuracy of 90.9%, detect click gestures with an accuracy of 94.9%, and identify among six ThermalTags with an overall accuracy of 95.0%. Finally, we show the versatility and potential of ThermalRing through various applications. Tengxiang Zhang, Yinshuai Zhang, Ke Sun 0003, Yuntao Wang 0001, Yiqiang Chen 0001 |
CHI | 5 |
| 2017 | Float: One-Handed and Touch-Free Target Selection on SmartwatchesabstractTouch interaction on smartwatches suffers from the awkwardness of having to use two hands and the "fat finger" problem. We present Float, a wrist-to-finger input approach that enables one-handed and touch-free target selection on smartwatches with high efficiency and precision using only commercially-available built-in sensors. With Float, a user tilts the wrist to point and performs an in-air finger tap to click. To realize Float, we first explore the appropriate motion space for wrist tilt and determine the clicking action (finger tap) through a user-elicitation study. We combine the photoplethysmogram (PPG) signal with accelerometer and gyroscope to detect finger taps with a recall of 97.9% and a false discovery rate of 0.4%. Experiments show that using just one hand, Float allows users to acquire targets with size ranging from 2mm to 10mm in less than 2s to 1s, meanwhile achieve much higher accuracy than direct touch in both stationary (>98.9%) and walking (>71.5%) contexts. Ke Sun 0003, Yuntao Wang 0001, Chun Yu, Yukang Yan, Hongyi Wen, Yuanchun Shi |
CHI | 2 |
| 2017 | ViVo: Video-Augmented Dictionary for Vocabulary LearningabstractResearch on Computer-Assisted Language Learning (CALL) has shown that the use of multimedia materials such as images and videos can facilitate interpretation and memorization of new words and phrases by providing richer cues than text alone. We present ViVo, a novel video-augmented dictionary that provides an inexpensive, convenient, and scalable way to exploit huge online video resources for vocabulary learning. ViVo automatically generates short video clips from existing movies with the target word highlighted in the subtitles. In particular, we apply a word sense disambiguation algorithm to identify the appropriate movie scenes with adequate contextual information for learning. We analyze the challenges and feasibility of this approach and describe our interaction design. A user study showed that learners were able to retain nearly 30% more new words with ViVo than with a standard bilingual dictionary days after learning. They preferred our video-augmented dictionary for its benefits in memorization and enjoyable learning experience. Yeshuang Zhu, Yuntao Wang 0001, Chun Yu, Shaoyun Shi, Yankai Zhang, Shuang He, Peijun Zhao, Xiaojuan Ma, Yuanchun Shi |
CHI | 2 |
| 2017 | BitID: Easily Add Battery-Free Wireless Sensors to Everyday ObjectsabstractRadio-Frequency Identification (RFID) systems are becoming increasingly used within smart environments. In this paper, we propose BitID, a passive Ultra-High Frequency (UHF) RFID based sensing technique that can easily be made using off-the-shelf tags. BitID can be added to everyday objects to enable sensing and control capabilities. With a simple shorting mechanism, BitID is able to differentiate between two states of the object to which it is attached (for example, whether a door is open or closed). We explain the working principle of BitID, and demonstrate how to build and apply it to target objects. We also show that by using a three- layered system architecture, BitID can be used for various applications, including event detection, energy monitoring, fitness tracking, human behavior tracking and control. Tengxiang Zhang, Nicholas Becker, Yuntao Wang 0001, Yuanchun Shi |
SMARTCOMP | 3 |
| 2014 | FOCUS: enhancing children's engagement in reading by using contextual BCI training sessionsabstractReading is an important aspect of a child's development. Reading outcome is heavily dependent on the level of engagement while reading. In this paper, we present FOCUS, an EEG-augmented reading system which monitors a child's engagement level in real time, and provides contextual BCI training sessions to improve a child's reading engagement. A laboratory experiment was conducted to assess the validity of the system. Results showed that FOCUS could significantly improve engagement in terms of both EEG-based measurement and teachers' subjective measure on the reading outcome. Chun Yu, Yuntao Wang 0001, Yuhang Zhao 0001, Chou Mo, Jie Liu 0027, Lie Zhang, Yuanchun Shi |
CHI | 3 |
| 2013 | Understanding performance of eyes-free, absolute position control on touchable mobile phonesabstractMany eyes-free interaction techniques have been proposed for touchscreens, but few researches have studied human's eyes-free pointing ability with mobile phones. In this paper, we investigate the single-handed thumb performance of eyes-free, absolute position control on mobile touch screens. Both 1D and 2D experiments were conducted. We explored the effects of target size and location on eyes-free touch patterns and accuracy. Our findings show that variance of touch points per target will converge as target size decreases. The centroid of touch points per target tends to be offset to the left of target center along horizontal direction, and shift toward screen center along vertical direction. Average accuracy drops from 99.6% of 2×2 layout to 85.0% of 4×4 layout, and average per target varies depending on the location of target. Our findings and design implications provide a foundation for future researches based on eyes-free, absolute position control using thumb on mobile devices. Yuntao Wang 0001, Chun Yu, Jie Liu 0027, Yuanchun Shi |
Mobile HCI | 1 |