Xuhai Xu

dblp:198/0980 · also Xuhai Orson Xu · DBLP profile ↗
← Back
56ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0001-5930-3899ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 42 · 8 first-author · 34 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toward Scalable ASL Education: Egocentric Stereo Sensing with LLM Feedback for Error-Aware Learning
abstract
American Sign Language (ASL) is the primary language of many Deaf and Hard of Hearing (DHH) individuals. However, existing learning resources often lack timely, individualized feedback, leaving learners uncertain about signing accuracy. We introduce a novel egocentric ASL learning system that integrates stereo vision, error detection across four manual ASL parameters (handshape, orientation, location, movement), and large language model (LLM)–driven natural language feedback. To our knowledge, this is the first system to deliver error-aware, pedagogically grounded feedback for ASL learners. A formative study with 15 ASL teachers and 30 learners (both Deaf and hearing backgrounds) supports the motivation and design goals, while a system evaluation with 13 Deaf ASL participants (novice to advanced) practicing 230 signs provides initial evidence of system feasibility and short-term, pedagogically promising behavior within the primary user community. Across two complementary studies, we identify key design principles: prioritizing reliability over sensitivity, stratifying feedback by error severity, and leveraging egocentric alignment for natural practice. Collectively, these contributions establish a foundation for scalable ASL education and provide generalizable insights for designing AI-mediated feedback in Human-Computer Interaction (HCI).
Yongxiang Cai, Taiting Lu, Yanjun Zhu, Yi-Shan Wu 0004, Qingsen Zhang, Xuhai Xu, Zhanpeng Jin, Mahanth Gowda, Yincheng Jin
CHI7
2026 More than Decision Support: Exploring Patients' Longitudinal Usage of Large Language Models in Real-World Healthcare-Seeking Journeys
abstract
Large language models (LLMs) have been increasingly adopted to support patients' healthcare-seeking in recent years. While prior patient-centered studies have examined the capabilities and experience of LLM-based tools in specific health-related tasks such as information-seeking, diagnosis, or decision-supporting, the inherently longitudinal nature of healthcare in real-world practice has been underexplored. This paper presents a four-week diary study with 25 patients to examine LLMs' roles across healthcare-seeking trajectories. Our analysis reveals that patients integrate LLMs not just as simple decision-support tools, but as dynamic companions that scaffold their journey across behavioral, informational, emotional, and cognitive levels. Meanwhile, patients actively assign diverse socio-technical meanings to LLMs, altering the traditional dynamics of agency, trust, and power in patient-provider relationships. Drawing from these findings, we conceptualize future LLMs as a longitudinal boundary companion that continuously mediates between patients and clinicians throughout longitudinal healthcare-seeking trajectories.
Yancheng Cao, Yishu Ji, Chris Yue Fu, Sahiti Dharmavaram, Meghan Turchioe, Natalie C. Benda, Lena Mamykina, Yuling Sun, Xuhai Xu
CHI9
2026 Recovery is Relational: Digital Support Needs for Patients and Supporters in Eating Disorder Recovery
abstract
Eating disorder (ED) recovery extends beyond therapy sessions, unfolding in vulnerable moments embedded in everyday life and relationships. Yet empirical understanding of how these moments arise, how supporters contribute, and how technologies might offer timely, contextual assistance remains limited. To address this gap, we conducted a design session and two-week diary study with 27 individuals with ED and 12 social supporters. Our analysis identified diverse contexts in which patients and supporters perceived support to be needed, and the forms of support they envisioned digital tools could offer. While many needs were mutually recognized, the actual practice of support often involved mismatches, suggesting opportunities for technologies to help mediate supportive engagement. Our study contributes empirical insight into everyday support moments in ED recovery and highlights opportunities to design digital interventions that provide context-sensitive assistance, empower supporters, and extend care beyond clinical settings.
Ryuhaerang Choi, Seohyeon Yoo, Xuhai Xu, Sung-Ju Lee 0001
CHI3
2026 MindfulAgents: Personalizing Mindfulness Meditation via an Expert-Aligned Multi-Agent System
abstract
Mindfulness meditation is a widely accessible and evidence-based method for supporting mental health. Despite the proliferation of mindfulness meditation apps, sustaining user engagement remains a persistent challenge. Personalizing the meditation experience is a promising strategy to improve engagement, but it often requires costly and unscalable manual effort. We present MindfulAgents, a multi-agent system powered by large language models that: (1) generates guided meditation scripts based on an expert-established mindfulness framework, (2) encourages users’ reflection on emotional states and mindfulness skills, and (3) enables real-time personalization of the mindfulness meditation experience for each user. In a formative lab study (N=13), MindfulAgents significantly improved in-session engagement (p = 0.011) and self-awareness (p = 0.014), as well as reduced momentary stress (p = 0.020). Furthermore, a four-week deployment study (N=62) demonstrated a notable increase (p = 0.002) in long-term engagement and level of mindfulness (p = 0.023). Participants reported that MindfulAgents offered more relevant meditation sessions personalized to individual needs in various contexts, supporting sustained practice. Our findings highlight the potential of LLM-driven personalization for enhancing user engagement in digital mindfulness meditation interventions.
Mengyuan Millie Wu, Zhihan Jiang 0001, Yuang Fan, Richard Feng, Sahiti Dharmavaram, Mathew Polowitz, Shawn Fallon, Bashima Islam, Lizbeth Benson, Irene Tung, J. David Creswell, Xuhai Xu
CHI12
2026 MIND: Empowering Mental Health Clinicians with Multimodal Data Insights through a Narrative Dashboard
abstract
Advances in data collection enable the capture of rich patient-generated data: from passive sensing (e.g., wearables and smartphones) to active self-reports (e.g., cross-sectional surveys and ecological momentary assessments). Although prior research has demonstrated the utility of patient-generated data in mental healthcare, significant challenges remain in effectively presenting these data streams along with clinical data (e.g., clinical notes) for clinical decision-making. Through co-design sessions with five clinicians, we propose MIND, a large language model-powered dashboard designed to present clinically relevant multimodal data insights for mental healthcare. MIND presents multimodal insights through narrative text, complemented by charts communicating underlying data. Our user study (N=16) demonstrates that clinicians perceive MIND as a significant improvement over baseline methods, reporting improved performance to reveal hidden and clinically relevant data insights (p<.001) and support their decision-making (p=.004). Grounded in the study results, we discuss future research opportunities to integrate data narratives in broader clinical practices.
Ruishi Zou, Margaret E. Morris, Jihan Ryu, Timothy D. Becker, Nicholas Allen, Anne Marie Albano, Randy Auerbach, Daniel A. Adler, Varun Mishra 0001, Lace M. K. Padilla, Dakuo Wang, Ryan Sultan, Xuhai Xu
CHI14
2026 DietGlance: Dietary Monitoring and Personalized Analysis at a Glance with Knowledge-Empowered AI Assistant
abstract
Growing awareness of wellness has prompted people to consider whether their dietary patterns align with their health and fitness goals. In response, researchers have introduced various wearable dietary monitoring systems and dietary assessment approaches. However, these solutions are either limited to identifying foods with simple ingredients or insufficient in providing an analysis of individual dietary behaviors with domain-specific knowledge. In this article, we present DietGlance , a system that automatically monitors dietary behaviors in daily routines and delivers personalized analysis from knowledge sources. DietGlance first detects ingestive episodes from multimodal inputs using eyeglasses, capturing privacy-preserving meal images of various dishes being consumed. Based on the inferred food items and consumed quantities from these images, DietGlance further provides nutritional analysis and personalized dietary suggestions, empowered by the retrieval-augmented generation module on a reliable nutrition library. A short-term user study (N = 33) and a 4-week longitudinal study (N = 16) demonstrate the usability and effectiveness of DietGlance , offering insights and implications for future AI-assisted dietary monitoring and personalized healthcare intervention systems using eyewear.
Zhihan Jiang 0001, Running Zhao, Lin Lin 0012, Handi Chen, Xuhai Xu, Yifang Wang 0001, Xiaojuan Ma, Edith C. H. Ngai
ACM Trans. Comput. Heal.7
2026 Human-Inspired Perspectives: A Survey on AI Long-Term Memory
abstract
With the rapid advancement of AI systems, their abilities to store, retrieve, and utilize information over the long term - referred to as long-term memory - have become increasingly significant. These capabilities are crucial for enhancing the performance of AI systems across a wide range of tasks. However, there is currently no comprehensive survey that systematically investigates AI's long-term memory capabilities, formulates a theoretical framework, and inspires the development of next-generation AI long-term memory systems. This paper begins by introducing the mechanisms of human long-term memory, then explores AI long-term memory mechanisms, establishing a mapping between the two. Based on the mapping relationships identified, we extend the current cognitive architectures and propose the Cognitive Architecture of Self-Adaptive Long-term Memory (SALM). SALM provides a theoretical framework for the practice of AI long-term memory and holds potential for guiding the creation of next-generation long-term memory driven AI systems. Finally, we delve into the future directions and application prospects of AI long-term memory.
Zihong He, Weizhe Lin, Fan Zhang 0017, Matt W. Jones, Laurence Aitchison, Xuhai Xu, Miao Liu 0007, Hai-Ning Liang, Per Ola Kristensson, Junxiao Shen
Proc. IEEE7
2025 Substance over Style: Evaluating Proactive Conversational Coaching Agents
abstract
While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals that evolve through multi-turn interactions, subjective evaluation criteria, mixed-initiative dialogue. In this work, we describe and implement five multi-turn coaching agents that exhibit distinct conversational styles, and evaluate them through a user study, collecting first-person feedback on 155 conversations. We find that users highly value core functionality, and that stylistic components in absence of core components are viewed negatively. By comparing user feedback with third-person evaluations from health experts and an LM, we reveal significant misalignment across evaluation approaches. Our findings provide insights into design and evaluation of conversational coaching agents and contribute toward improving human-centered NLP applications.
Vidya Srinivas, Xuhai Xu, Xin Liu 0034, Kumar Ayush, Isaac R. Galatzer-Levy, Shwetak N. Patel, Daniel McDuff, Tim Althoff
ACL (1)2
2025 MedAI-SciTS: Enhancing Interdisciplinary Collaboration between AI Researchers and Medical Experts
Chen Cao 0005, Zoe Xiao Fang, Zhenwen Liang, Lena Mamykina, Laura Sbaffi, Xuhai Xu
CHI7
2025 The Odyssey Journey: Top-Tier Medical Resource Seeking for Specialized Disorder in China
abstract
It is pivotal for patients to receive accurate health information, diagnoses, and timely treatments. However, in China, the significant imbalanced doctor-to-patient ratio intensifies the information and power asymmetries in doctor-patient relationships. Health information-seeking, which enables patients to collect information from sources beyond doctors, is a potential approach to mitigate these asymmetries. While HCI research predominantly focuses on common chronic conditions, our study focuses on specialized disorders, which are often familiar to specialists but not to general practitioners and the public. With Hemifacial Spasm (HFS) as an example, we aim to understand patients' health information and top-tier1 medical resource seeking journeys in China. Through interviews with three neurosurgeons and 12 HFS patients from rural and urban areas, and applying Actor-Network Theory, we provide empirical insights into the roles, interactions, and workflows of various actors in the health information-seeking network. We also identified five strategies patients adopted to mitigate asymmetries and access top-tier medical resources, illustrating these strategies as subnetworks within the broader health information-seeking network and outlining their advantages and challenges. © 2025 Copyright held by the owner/author(s).
Ka I Chan, Siying Hu, Yuntao Wang 0001, Xuhai Xu, Zhicong Lu, Yuanchun Shi
CHI4
2025 Promoting Prosociality via Micro-acts of Joy: A Large-Scale Well-Being Intervention Study
abstract
Prosociality has been well-documented to positively impact mental, social, and physical well-being.However, existing studies of interventions for promoting prosociality have limitations such as
Hitesh Goel, Yoobin Park, Jin Liou, Darwin A. Guevarra, Peggy Callahan, Jolene Smith, Bingsheng Yao, Dakuo Wang, Xin Liu 0034, Daniel McDuff, Noémie Elhadad, Emiliana Simon-Thomas, Elissa Epel, Xuhai Xu
CHI14
2025 What Social Media Use Do People Regret? An Analysis of 34K Smartphone Screenshots with Multimodal LLM
abstract
Smartphone users often regret aspects of their phone use, especially social media use.However, pinpointing specific ways in which the design of an interface contributes to regrettable use can be challenging due to the complexity of social media app features and user intentions.We conducted a one-week study with 17 Android users, using a novel method where we passively collected screenshots every five seconds, which we analyzed via a multimodal large language model to understand participants' usage activity at a finegrained level.Triangulating this data with data from experience sampling, surveys, and interviews, we found that regret varies based on user intention, with non-intentional and social media use being especially regrettable.Regret also varies by social media activity; participants were most likely to regret viewing algorithmically recommended content and comments.Additionally, participants frequently deviated to browsing social media when their intention was direct communication, which slightly increased their regret.Our findings provide guidance to designers and policy-makers seeking to improve users' experience and autonomy.
Longjie Guo, Xiran Lin, Xuhai Xu, Yung-Ju Chang, Alexis Hiniker
CHI4
2025 Scaling Wearable Foundation Models
abstract
Wearable sensors have become ubiquitous thanks to a variety of health tracking features. The resulting continuous and longitudinal measurements from everyday life generate large volumes of data. However, making sense of these observations for scientific and actionable insights is non-trivial. Inspired by the empirical success of generative modeling, where large neural networks learn powerful representations from vast amounts of text, image, video, or audio data, we investigate the scaling properties of wearable sensor foundation models across compute, data, and model size. Using a dataset of up to 40 million hours of in-situ heart rate, heart rate variability, accelerometer, electrodermal activity, skin temperature, and altimeter per-minute data from over 165,000 people, we create LSM, a multimodal foundation model built on the largest wearable-signals dataset with the most extensive range of sensor modalities to date. Our results establish the scaling laws of LSM for tasks such as imputation, interpolation and extrapolation across both time and sensor modalities. Moreover, we highlight how LSM enables sample-efficient downstream learning for tasks including exercise and activity recognition.
Girish Narayanswamy, Xin Liu 0034, Kumar Ayush, Yuzhe Yang 0003, Xuhai Xu, Shun Liao, Jake Garrison, Shyam A. Tailor, Jacob E. Sunshine, Yun Liu 0013, Tim Althoff, Shri Narayanan, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Samy Abdel-Ghaffar, Daniel McDuff
ICLR5
2025 Poster: Split-and-Combine Rectification of Ultra-Wide Fisheye Images into Cubemaps
abstract
Ultra-wide fisheye cameras (FoV > 180°) offer unmatched scene coverage but introduce severe geometric distortions that degrade vision system performance. Existing rectification methods struggle with such extreme FoVs due to the inherent limitations of single-perspective projections and the scarcity of ground truth data. We propose a novel split-and-combine framework that rectifies ultra-wide fisheye images into 5-face cubemaps. Our approach begins with a lightweight CNN estimating geometric parameters to guide a structured decomposition of the fisheye image into directional regions. Each region is independently corrected using a two-stage pipeline: a flow-based pre-corrector for geometric warping and a diffusion-based enhancer for detail restoration. We train on a synthetic dataset derived from real-world perspective images, enabling partial supervision. Preliminary results demonstrate strong quantitative and qualitative performance, validating our modular architecture as an effective solution for high-fidelity rectification in ultra-wide fisheye imagery.
Yuang Fan, Xuhai Xu, Xiaofan Jiang 0001
MobiCom2
2025 RADAR: Benchmarking Language Models on Imperfect Tabular Data
abstract
Language models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness—the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies—remains underexplored. These artifacts are especially common in real-world tabular data and, if mishandled, can significantly compromise the validity of analytical conclusions. To address this gap, we present RADAR, a benchmark for systematically evaluating data-aware reasoning on tabular data. We develop a framework to simulate data artifacts via programmatic perturbations to enable targeted evaluation of model behavior. RADAR comprises 2,980 table-query pairs, grounded in real-world data spanning 9 domains and 5 data artifact types. In addition to evaluating artifact handling, RADAR systematically varies table size to study how reasoning performance holds when increasing table size. Our evaluation reveals that, despite decent performance on tables without data artifacts, frontier models degrade significantly when data artifacts are introduced, exposing critical gaps in their capacity for robust, data-aware analysis. Designed to be flexible and extensible, RADAR supports diverse perturbation types and controllable table sizes, offering a valuable resource for advancing tabular reasoning.
Ken Gu, Zhihan Zhang 0002, Kate Lin, Yuwei Zhang 0001, Akshay Paruchuri, Hong Yu 0001, Mehran Kazemi, Kumar Ayush, A. Ali Heydari, Maxwell A. Xu, Yun Liu 0013, Ming-Zher Poh, Yuzhe Yang 0003, Mark Malhotra, Shwetak N. Patel, Hamid Palangi, Xuhai Xu, Daniel McDuff, Tim Althoff, Xin Liu 0034
NeurIPS17
2025 SensorLM: Learning the Language of Wearable Sensors
abstract
We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks. Code is available at https://github.com/Google-Health/consumer-health-research/tree/main/sensorlm.
Yuwei Zhang 0001, Kumar Ayush, Siyuan Qiao, A. Ali Heydari, Girish Narayanswamy, Maxwell A. Xu, Ahmed Metwally 0002, Jinhua Xu, Jake Garrison, Xuhai Xu, Tim Althoff, Yun Liu 0013, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Cecilia Mascolo, Xin Liu 0034, Daniel McDuff, Yuzhe Yang 0003
NeurIPS10
2025 SignGlass: First-Person View Comprehensive and Generalizable ASL Translation Using Wearable Glass
Yongxiang Cai, Taiting Lu, Hao Zhou 0001, Kenneth DeHaan, Xuhai Xu, Mahanth Gowda, Yincheng Jin
UIST6
2025 TangibleTale: Designing Tangible Child-Parent Interactive Storytelling for Promoting Eating Behaviors
abstract
Tangible narrative allow children to interact with physical objects and provides an immersive storytelling approach to influence their cognition and behavior. However, research on using tangible narratives to encourage behaviors in children remains limited. To improve children’s eating behaviors, we combined narrative transportation theory with behavior change principles to conduct a design study, called TangibleTale. Specifically, we conducted a formative user study (N = 12 pairs) to identify the characteristics of children’s engagement with tangible narrative and their interactions with parents, which were then incorporated into a design workshop (N = 12) to develop an interactive product comprising tangible elements and an accompanying app. With the produced outcomes, we conducted a comparative experiment (N = 24 pairs) in a home setting to verify and explore the role of tangible narrative in child–parent mealtime interaction. Finally, we formulated design guidelines for a tangible narrative that can serve as a reference to assist in creating more impactful products that foster positive behavioral growth in children.
Mingxuan Liu 0008, Xuhai Xu, Danli Luo, Gigi Nathalie, Jiaji Li, Cheng Yang 0014, Ye Tao 0001, Guanyun Wang
Int. J. Hum. Comput. Interact.4
2024 From Text to Self: Users' Perception of AIMC Tools on Interpersonal Communication and Self
abstract
In the rapidly evolving landscape of AI-mediated communication (AIMC), tools powered by Large Language Models (LLMs) are becoming integral to interpersonal communication. Employing a mixed-methods approach, we conducted a one-week diary and interview study to explore users’ perceptions of these tools’ ability to: 1) support interpersonal communication in the short-term, and 2) lead to potential long-term effects. Our findings indicate that participants view AIMC support favorably, citing benefits such as increased communication confidence, finding precise language to express their thoughts, and navigating linguistic and cultural barriers. However, our findings also show current limitations of AIMC tools, including verbosity, unnatural responses, and excessive emotional intensity. These shortcomings are further exacerbated by user concerns about inauthenticity and potential overreliance on the technology. We identify four key communication spaces delineated by communication stakes (high or low) and relationship dynamics (formal or informal) that differentially predict users’ attitudes toward AIMC tools. Specifically, participants report that these tools are more suitable for communicating in formal relationships than informal ones and more beneficial in high-stakes than low-stakes communication.
Sami Foell, Xuhai Xu, Alexis Hiniker
CHI3
2024 InteractOut: Leveraging Interaction Proxies as Input Manipulation Strategies for Reducing Smartphone Overuse
abstract
Smartphone overuse poses risks to people’s physical and mental health. However, current intervention techniques mainly focus on explicitly changing screen content (i.e., output) and often fail to persistently reduce smartphone overuse due to being over-restrictive or over-flexible. We present the design and implementation of InteractOut, a suite of implicit input manipulation techniques that leverage interaction proxies to weakly inhibit the natural execution of common user gestures on mobile devices. We present a design space for input manipulations and demonstrate 8 Android implementations of input interventions. We first conducted a pilot lab study (N=30) to evaluate the usability of these interventions. Based on the results, we then performed a 5-week within-subject field experiment (N=42) to evaluate InteractOut in real-world scenarios. Compared to the traditional and common timed lockout technique, InteractOut significantly reduced the usage time by an additional 15.6% and opening frequency by 16.5% on participant-selected target apps. InteractOut also achieved a 25.3% higher user acceptance rate, and resulted in less frustration and better user experience according to participants’ subjective feedback. InteractOut demonstrates a new direction for smartphone overuse intervention and serves as a strong complementary set of techniques with existing methods.
Tao Lu 0013, Hongxiao Zheng, Tianying Zhang, Xuhai Xu, Anhong Guo
CHI4
2024 Time2Stop: Adaptive and Explainable Human-AI Loop for Smartphone Overuse Intervention
abstract
Despite a rich history of investigating smartphone overuse intervention techniques, AI-based just-in-time adaptive intervention (JITAI) methods for overuse reduction are lacking. We develop Time2Stop, an intelligent, adaptive, and explainable JITAI system that leverages machine learning to identify optimal intervention timings, introduces interventions with transparent AI explanations, and collects user feedback to establish a human-AI loop and adapt the intervention model over time. We conducted an 8-week field experiment (N=71) to evaluate the effectiveness of both the adaptation and explanation aspects of Time2Stop. Our results indicate that our adaptive models significantly outperform the baseline methods on intervention accuracy (>32.8% relatively) and receptivity (>8.0%). In addition, incorporating explanations further enhances the effectiveness by 53.8% and 11.4% on accuracy and receptivity, respectively. Moreover, Time2Stop significantly reduces overuse, decreasing app visit frequency by 7.0 ∼ 8.9%. Our subjective data also echoed these quantitative measures. Participants preferred the adaptive interventions and rated the system highly on intervention time accuracy, effectiveness, and level of trust. We envision our work can inspire future research on JITAI systems with a human-AI loop to evolve with users.
Adiba Orzikulova, Zhipeng Li 0001, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Marzyeh Ghassemi, Sung-Ju Lee 0001, Anind K. Dey, Xuhai Xu
CHI10
2024 Fast-Forward Reality: Authoring Error-Free Context-Aware Policies with Real-Time Unit Tests in Extended Reality
abstract
Advances in ubiquitous computing have enabled end-user authoring of context-aware policies (CAPs) that control smart devices based on specific contexts of the user and environment. However, authoring CAPs accurately and avoiding run-time errors is challenging for end-users as it is difficult to foresee CAP behaviors under complex real-world conditions. We propose Fast-Forward Reality, an Extended Reality (XR) based authoring workflow that enables end-users to iteratively author and refine CAPs by validating their behaviors via simulated unit test cases. We develop a computational approach to automatically generate test cases based on the authored CAP and the user’s context history. Our system delivers each test case with immersive visualizations in XR, facilitating users to verify the CAP behavior and identify necessary refinements. We evaluated Fast-Forward Reality in a user study (N=12). Our authoring and validation process improved the accuracy of CAPs and the users provided positive feedback on the system usability.
Xun Qian, Tianyi Wang 0004, Xuhai Xu, Tanya R. Jonker, Kashyap Todi
CHI3
2024 AdaptiveVoice: Cognitively Adaptive Voice Interface for Driving Assistance
abstract
Current voice assistants present messages in a predefined format without considering users’ mental states. This paper presents an optimization-based approach to alleviate this issue which adjusts the level of details and speech speed of the voice messages according to the estimated cognitive load of the user. In the first user study (N = 12), we investigated the impact of cognitive load on user performance. The findings reveal significant differences in preferred message formats across five cognitive load levels, substantiating the need for voice message adaptation. We then implemented AdaptiveVoice, an algorithm based on combinatorial optimization to generate adaptive voice messages in real time. In the second user study (N = 30) conducted in a VR-simulated driving environment, we compare AdaptiveVoice with a fixed format baseline, with and without visual guidance on the Heads-up display (HUD). Results indicate that users benefit from AdaptiveVoice with reduced response time and improved driving performance, particularly when it is augmented with HUD.
Shaoyue Wen, Songming Ping, Jialin Wang 0002, Hai-Ning Liang, Xuhai Xu, Yukang Yan
CHI5
2024 MindShift: Leveraging Large Language Models for Mental-States-Based Problematic Smartphone Use Intervention
abstract
Problematic smartphone use negatively affects physical and mental health. Despite the wide range of prior research, existing persuasive techniques are not flexible enough to provide dynamic persuasion content based on users’ physical contexts and mental states. We first conducted a Wizard-of-Oz study (N=12) and an interview study (N=10) to summarize the mental states behind problematic smartphone use: boredom, stress, and inertia. This informs our design of four persuasion strategies: understanding, comforting, evoking, and scaffolding habits. We leveraged large language models (LLMs) to enable the automatic and dynamic generation of effective persuasion content. We developed MindShift, a novel LLM-powered problematic smartphone use intervention technique. MindShift takes users’ in-the-moment app usage behaviors, physical contexts, mental states, goals & habits as input, and generates personalized and dynamic persuasive content with appropriate persuasion strategies. We conducted a 5-week field experiment (N=25) to compare MindShift with its simplified version (remove mental states) and baseline techniques (fixed reminder). The results show that MindShift improves intervention acceptance rates by 4.7-22.5% and reduces smartphone usage duration by 7.4-9.8%. Moreover, users have a significant drop in smartphone addiction scale scores and a rise in self-efficacy scale scores. Our study sheds light on the potential of leveraging LLMs for context-aware persuasion in other behavior change domains.
Ruolan Wu, Chun Yu, Xiaole Pan, Yujia Liu 0004, Ningning Zhang, Yuhan Wang 0015, Qiaolei Jiang, Xuhai Xu, Yuanchun Shi
CHI11
2024 Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis Diagnosis
abstract
Today's AI systems for medical decision support often succeed on benchmark datasets in research papers but fail in real-world deployment. This work focuses on the decision making of sepsis, an acute life-threatening systematic infection that requires an early diagnosis with high uncertainty from the clinician. Our aim is to explore the design requirements for AI systems that can support clinical experts in making better decisions for the early diagnosis of sepsis. The study begins with a formative study investigating why clinical experts abandon an existing AI-powered Sepsis predictive module in their electrical health record (EHR) system. We argue that a human-centered AI system needs to support human experts in the intermediate stages of a medical decision-making process (e.g., generating hypotheses or gathering data), instead of focusing only on the final decision. Therefore, we build SepsisLab based on a state-of-the-art AI algorithm and extend it to predict the future projection of sepsis development, visualize the prediction uncertainty, and propose actionable suggestions (i.e., which additional laboratory tests can be collected) to reduce such uncertainty. Through heuristic evaluation with six clinicians using our prototype system, we demonstrate that SepsisLab enables a promising human-AI collaboration paradigm for the future of AI-assisted sepsis diagnosis and other high-stakes medical decision making.
Shao Zhang, Xuhai Xu, Changchang Yin, Yuxuan Lu 0003, Bingsheng Yao, Melanie Tory, Lace M. K. Padilla, Jeffrey M. Caterino, Ping Zhang 0016, Dakuo Wang
CHI3
2024 Boosting Gesture Recognition with an Automatic Gesture Annotation Framework
abstract
Training a real-time gesture recognition model heavily relies on annotated data. However, manual data annotation is costly and demands substantial human effort. In order to address this challenge, we propose a framework that can automatically annotate gesture classes and identify their temporal ranges. Our framework consists of two key components: (1) a novel annotation model that leverages the Connectionist Temporal Classification (CTC) loss, and (2) a semi-supervised learning pipeline that enables the model to improve its performance by training on its own predictions, known as pseudo labels. These high-quality pseudo labels can also be used to enhance the accuracy of other downstream gesture recognition models. To evaluate our framework, we conducted experiments using two publicly available gesture datasets. Our ablation study demonstrates that our annotation model design surpasses the baseline in terms of both gesture classification accuracy (3–4 % improvement) and localization accuracy (71-75% improvement). Additionally, we illustrate that the pseudo-labeled dataset produced from the proposed framework significantly boosts the accuracy of a pre-trained downstream gesture recognition model by 11-18%. We believe that this annotation framework has immense potential to improve the training of downstream gesture recognition models using unlabeled datasets.
Junxiao Shen, Xuhai Xu, Ran Tan, Amy Karlson, Evan Strasnick
FG2
2024 Towards Open-World Gesture Recognition
abstract
Providing users with accurate gestural interfaces, such as gesture recognition based on wrist-worn devices, is a key challenge in mixed reality. However, static machine learning processes in gesture recognition assume that training and test data come from the same underlying distribution. Unfortunately, in real-world applications involving gesture recognition, such as gesture recognition based on wrist-worn devices, the data distribution may change over time. We formulate this problem of adapting recognition models to new tasks, where new data patterns emerge, as open-world gesture recognition (OWGR). We propose the use of continual learning to enable machine learning models to be adaptive to new tasks without degrading performance on previously learned tasks. However, the process of exploring parameters for questions around when, and how, to train and deploy recognition models requires resource-intensive user studies may be impractical. To address this challenge, we propose a design engineering approach that enables offline analysis on a collected large-scale dataset by systematically examining various parameters and comparing different continual learning methods. Finally, we provide design guidelines to enhance the development of an open-world wrist-worn gesture recognition process.
Junxiao Shen, Matthias De Lange, Xuhai Xu, Enmin Zhou, Ran Tan, Naveen Suda, Maciej Lazarewicz, Per Ola Kristensson, Amy Karlson, Evan Strasnick
ISMAR3
2024 MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
abstract
Foundation models are becoming valuable tools in medicine. Yet despite their promise, the best way to leverage Large Language Models (LLMs) in complex medical tasks remains an open question. We introduce a novel multi-agent framework, named **M**edical **D**ecision-making **Agents** (**MDAgents**) that helps to address this gap by automatically assigning a collaboration structure to a team of LLMs. The assigned solo or group collaboration structure is tailored to the medical task at hand, a simple emulation inspired by the way real-world medical decision-making processes are adapted to tasks of different complexities. We evaluate our framework and baseline methods using state-of-the-art LLMs across a suite of real-world medical knowledge and clinical diagnosis benchmarks, including a comparison of LLMs’ medical complexity classification against human physicians. MDAgents achieved the **best performance in seven out of ten** benchmarks on tasks requiring an understanding of medical knowledge and multi-modal reasoning, showing a significant **improvement of up to 4.2\%** ($p$ < 0.05) compared to previous methods' best performances. Ablation studies reveal that MDAgents effectively determines medical complexity to optimize for efficiency and accuracy across diverse medical tasks. Notably, the combination of moderator review and external medical knowledge in group collaboration resulted in an average accuracy **improvement of 11.8\%**. Our code can be found at https://github.com/mitmedialab/MDAgents.
Yubin Kim 0002, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, Hae Won Park 0001
NeurIPS5
2023 An Autoethnographic Case Study of Generative Artificial Intelligence's Utility for Accessibility
abstract
With the recent rapid rise in Generative Artificial Intelligence (GAI) tools, it is imperative that we understand their impact on people with disabilities, both positive and negative. However, although we know that AI in general poses both risks and opportunities for people with disabilities, little is known specifically about GAI in particular. To address this, we conducted a three-month autoethnography of our use of GAI to meet personal and professional needs as a team of researchers with and without disabilities. Our findings demonstrate a wide variety of potential accessibility-related uses for GAI while also highlighting concerns around verifiability, training data, ableism, and false promises.
Kate S. Glazko, Momona Yamagami, Aashaka Desai, Kelly Mack, Venkatesh Potluri, Xuhai Xu, Jennifer Mankoff
ASSETS6
2023 Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution Images
abstract
A computer vision system using low-resolution image sensors can provide intelligent services (e.g., activity recognition) but preserve unnecessary visual privacy information from the hardware level. However, preserving visual privacy and enabling accurate machine recognition have adversarial needs on image resolution. Modeling the trade-off of privacy preservation and machine recognition performance can guide future privacy-preserving computer vision systems using low-resolution image sensors. In this paper, using the at-home activity of daily livings (ADLs) as the scenario, we first obtained the most important visual privacy features through a user survey. Then we quantified and analyzed the effects of image resolution on human and machine recognition performance in activity recognition and privacy awareness tasks. We also investigated how modern image super-resolution techniques influence these effects. Based on the results, we proposed a method for modeling the trade-off of privacy preservation and activity recognition on low-resolution images.
Yuntao Wang 0001, Zirui Cheng, Xin Yi 0001, Yan Kong, Xuhai Xu, Yukang Yan, Chun Yu, Shwetak N. Patel, Yuanchun Shi
CHI6
2023 XAIR: A Framework of Explainable AI in Augmented Reality
abstract
Explainable AI (XAI) has established itself as an important component of AI-driven interactive systems. With Augmented Reality (AR) becoming more integrated in daily lives, the role of XAI also becomes essential in AR because end-users will frequently interact with intelligent services. However, it is unclear how to design effective XAI experiences for AR. We propose XAIR, a design framework that addresses when, what, and how to provide explanations of AI output in AR. The framework was based on a multi-disciplinary literature review of XAI and HCI research, a large-scale survey probing 500+ end-users’ preferences for AR-based explanations, and three workshops with 12 experts collecting their insights about XAI design in AR. XAIR’s utility and effectiveness was verified via a study with 10 designers and another study with 12 end-users. XAIR can provide guidelines for designers, inspiring them to identify new design opportunities and achieve effective XAI designs in AR.
Xuhai Xu, Anna Yu, Tanya R. Jonker, Kashyap Todi, Feiyu Lu 0001, Xun Qian, João Marcelo Evangelista Belo, Tianyi Wang 0004, Michelle Li, Aran Mun, Te-Yen Wu, Junxiao Shen, Ting Zhang 0013, Narine Kokhlikyan, Fulton Wang, Paul Sorenson, Sophie Kahyun Kim, Hrvoje Benko
CHI1
2023 Reviewing and Reflecting on Smart Home Research from the Human-Centered Perspective
abstract
While there has been rapid growth in smart home research from a technical perspective– focusing on home automation, devices, software, and protocols– few review papers examine the human-centered perspective. A human-centered focus is crucial for achieving the goals of providing natural, convenient, comfortable, friendly, and safe user experiences in the smart home. To understand key innovations in human-centered smart home research, we analyzed keyword changes over time via 19,091 papers from 2000 to 2022, then selected 55 papers from high-impact venues in the last five years, and summarized them through a combination of qualitative and quantitative methods. Our analysis revealed five research trends with unique characteristics and interdependence. Drawing on this review, we elaborate on the future of smart home design research with respect to multidisciplinary development, stakeholder involvement, and the shift of design implications.
Zhijun Ma, Xuhai Xu, Haipeng Mi
CHI5
2023 Exploring the Impact of User and System Factors on Human-AI Interactions in Head-Worn Displays
abstract
Empowered by the rich sensory capabilities and the advancements in artificial intelligence (AI), head-worn displays (HWD) could understand the user’s contexts and provide just-in-time assistance to users’ tasks to augment their everyday lives. However, there has been limited understanding of how users perceive interacting with AI services, and how different factors impact the user experience in HWD applications. In this research, we investigated broadly what user and system factors play important roles in human-AI experiences during an AI-assisted spatial task. We conducted a user study to simulate an everyday scenario where augmented reality (AR) glasses could provide suggestions/assistance. We researched three AI system factors (performance, initiation, transparency) with multiple user factors (personality traits, trust propensity, and prior trust with AI). We not only identified the impact of user traits such as the levels of conscientiousness and prior trust with the AI, but also found interesting interactions between them and system factors such as AI’s performance and initiation strategy. Based on the findings, we suggest that future AI assistance on HWD needs to take users’ individual characteristics into account and customize the system design accordingly.
Feiyu Lu 0001, Xuhai Xu, Brennan Jones, Laird Malamed
ISMAR3
2023 ConeSpeech: Exploring Directional Speech Interaction for Multi-Person Remote Communication in Virtual Reality
abstract
Remote communication is essential for efficient collaboration among people at different locations. We present ConeSpeech, a virtual reality (VR) based multi-user remote communication technique, which enables users to selectively speak to target listeners without distracting bystanders. With ConeSpeech, the user looks at the target listener and only in a cone-shaped area in the direction can the listeners hear the speech. This manner alleviates the disturbance to and avoids overhearing from surrounding irrelevant people. Three featured functions are supported, directional speech delivery, size-adjustable delivery range, and multiple delivery areas, to facilitate speaking to more than one listener and to listeners spatially mixed up with bystanders. We conducted a user study to determine the modality to control the cone-shaped delivery area. Then we implemented the technique and evaluated its performance in three typical multi-user communication tasks by comparing it to two baseline methods. Results show that ConeSpeech balanced the convenience and flexibility of voice communication.
Yukang Yan, Haohua Liu, Yingtian Shi, Ruici Guo, Zisu Li, Xuhai Xu, Chun Yu, Yuntao Wang 0001, Yuanchun Shi
IEEE Trans. Vis. Comput. Graph.7
2022 UnlockedMaps: Visualizing Real-Time Accessibility of Urban Rail Transit Using a Web-Based Map
abstract
Current web-based maps do not provide visibility into real-time elevator outages at urban rail transit stations, disenfranchising commuters (e.g., wheelchair users) who rely on functioning elevators at transit stations. In this paper, we demonstrate UnlockedMaps, an open-source and open-data web-based map that visualizes the real-time accessibility of urban rail transit stations in six North American cities, assisting users in making informed decisions regarding their commute. Specifically, UnlockedMaps uses a map to display transit stations, prominently highlighting their real-time accessibility status (accessible with functioning elevators, accessible but experiencing at least one elevator outage, or not-accessible) and surrounding accessible restaurants and restrooms. UnlockedMaps is the first system to collect elevator outage data from 2,336 transit stations over 23 months and make it publicly available via an API. We report on results from our pilot user studies with five stakeholder groups: (1) people with mobility disabilities; (2) pregnant people; (3) cyclists/stroller users/commuters with heavy equipment; (4) members of disability advocacy groups; and (5) civic hackers.
Ather Sharif, Aneesha Ramesh, Trung-Anh Nguyen, Luna Chen, Kent Richard Zeng, Lanqing Hou, Xuhai Xu
ASSETS7
2022 Enabling Hand Gesture Customization on Wrist-Worn Devices
abstract
We present a framework for gesture customization requiring minimal examples from users, all without degrading the performance of existing gesture sets. To achieve this, we first deployed a large-scale study (N=500+) to collect data and train an accelerometer-gyroscope recognition model with a cross-user accuracy of 95.7% and a false-positive rate of 0.6 per hour when tested on everyday non-gesture data. Next, we design a few-shot learning framework which derives a lightweight model from our pre-trained model, enabling knowledge transfer without performance degradation. We validate our approach through a user study (N=20) examining on-device customization from 12 new gestures, resulting in an average accuracy of 55.3%, 83.1%, and 87.2% on using one, three, or five shots when adding a new gesture, while maintaining the same recognition accuracy and false-positive rate from the pre-existing gesture set. We further evaluate the usability of our real-time implementation with a user experience study (N=20). Our results highlight the effectiveness, learnability, and usability of our customization framework. Our approach paves the way for a future where users are no longer bound to pre-existing gestures, freeing them to creatively introduce new gestures tailored to their preferences and abilities.
Xuhai Xu, Jun Gong 0002, Carolina Brum, Lilian Liang, Bongsoo Suh, Shivam Kumar Gupta, Yash Agarwal, Laurence Lindsey, Runchang Kang, Behrooz Shahsavari, Heriberto Nieto, Scott E. Hudson, Charlie Maalouf, Seyed Mousavi, Gierad Laput
CHI1
2022 TypeOut: Leveraging Just-in-Time Self-Affirmation for Smartphone Overuse Reduction
abstract
Smartphone overuse is related to a variety of issues such as lack of sleep and anxiety. We explore the application of Self-Affirmation Theory on smartphone overuse intervention in a just-in-time manner. We present TypeOut, a just-in-time intervention technique that integrates two components: an in-situ typing-based unlock process to improve user engagement, and self-affirmation-based typing content to enhance effectiveness. We hypothesize that the integration of typing and self-affirmation content can better reduce smartphone overuse. We conducted a 10-week within-subject field experiment (N=54) and compared TypeOut against two baselines: one only showing the self-affirmation content (a common notification-based intervention), and one only requiring typing non-semantic content (a state-of-the-art method). TypeOut reduces app usage by over 50%, and both app opening frequency and usage duration by over 25%, all significantly outperforming baselines. TypeOut can potentially be used in other domains where an intervention may benefit from integrating self-affirmation exercises with an engaging just-in-time mechanism.
Xuhai Xu, Tianyuan Zou, Yanzhang Li, Ruolin Wang, Tianyi Yuan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey
CHI1
2022 Ubilung: Multi-Modal Passive-Based Lung Health Assessment
abstract
Lung health assessment is traditionally done mainly through X-ray images and spirometry tests which are time-consuming, cumbersome, and costly. In this paper, we investigate the potential of passively recordable contents such as speech, cough and heart signal for such an assessment. Our regression model is the first in the literature to achieve mean absolute error (MAE) of 7.47% for estimation of forced expiratory volume in 1 sec. (FEV1) over forced vital capacity (FVC) ratio using these contents. This is comparable to the state of the art active phone-based spirometry methods. Additionally our classification models achieve a F1-score of 0.982 for healthy v.s. diseased, 0.881 for obstructive v.s. non-obstructive, 0.854 for chronic obstructive pulmonary disease (COPD) v.s. asthma, and 0.892 for severe v.s. non-severe obstruction classification.
Ebrahim Nemati, Xuhai Xu, Viswam Nathan, Korosh Vatanparvar, Tousif Ahmed, Daniel McCaffrey 0001, Jilong Kuang, Jun Alex Gao
ICASSP2
2022 GLOBEM Dataset: Multi-Year Datasets for Longitudinal Human Behavior Modeling Generalization
abstract
Recent research has demonstrated the capability of behavior signals captured by smartphones and wearables for longitudinal behavior modeling. However, there is a lack of a comprehensive public dataset that serves as an open testbed for fair comparison among algorithms. Moreover, prior studies mainly evaluate algorithms using data from a single population within a short period, without measuring the cross-dataset generalizability of these algorithms. We present the first multi-year passive sensing datasets, containing over 700 user-years and 497 unique users’ data collected from mobile and wearable sensors, together with a wide range of well-being metrics. Our datasets can support multiple cross-dataset evaluations of behavior modeling algorithms’ generalizability across different users and years. As a starting point, we provide the benchmark results of 18 algorithms on the task of depression detection. Our results indicate that both prior depression detection algorithms and domain generalization techniques show potential but need further research to achieve adequate cross-dataset generalizability. We envision our multi-year datasets can support the ML community in developing generalizable longitudinal behavior modeling algorithms.
Xuhai Xu, Han Zhang 0004, Yasaman S. Sefidgar, Yiyi Ren, Xin Liu 0034, Woosuk Seo, Kevin S. Kuehn, Mike A. Merrill, Paula S. Nurius, Shwetak N. Patel, Tim Althoff, Margaret E. Morris, Eve A. Riskin, Jennifer Mankoff, Anind K. Dey
NeurIPS1
2021 Understanding the Design Space of Mouth Microgestures
abstract
As wearable devices move toward the face (i.e. smart earbuds, glasses), there is an increasing need to facilitate intuitive interactions with these devices. Current sensing techniques can already detect many mouth-based gestures; however, users’ preferences of these gestures are not fully understood. In this paper, we investigate the design space and usability of mouth-based microgestures. We first conducted brainstorming sessions (N=16) and compiled an extensive set of 86 user-defined gestures. Then, with an online survey (N=50), we assessed the physical and mental demand of our gesture set and identified a subset of 14 gestures that can be performed easily and naturally. Finally, we conducted a remote Wizard-of-Oz usability study (N=11) mapping gestures to various daily smartphone operations under a sitting and walking context. From these studies, we develop a taxonomy for mouth gestures, finalize a practical gesture set for common applications, and provide design guidelines for future mouth-based gesture interactions.
Xuhai Xu, Richard Li 0002, Yuanchun Shi, Shwetak N. Patel, Yuntao Wang 0001
Conference on Designing Interactive Systems2
2021 LightWrite: Teach Handwriting to The Visually Impaired with A Smartphone
abstract
Learning to write is challenging for blind and low vision (BLV) people because of the lack of visual feedback. Regardless of the drastic advancement of digital technology, handwriting is still an essential part of daily life. Although tools designed for teaching BLV to write exist, many are expensive and require the help of sighted teachers. We propose LightWrite, a low-cost, easy-to-access smartphone application that uses voice-based descriptive instruction and feedback to teach BLV users to write English lowercase letters and Arabian digits in a specifically designed font. A two-stage study with 15 BLV users with little prior writing knowledge shows that LightWrite can successfully teach users to learn handwriting characters in an average of 1.09 minutes for each letter. After initial training and 20-minute daily practice for 5 days, participants were able to write an average of 19.9 out of 26 letters that are recognizable by sighted raters.
Zihan Wu 0002, Chun Yu, Xuhai Xu, Tianyuan Zou, Ruolin Wang, Yuanchun Shi
CHI3
2021 Auth+Track: Enabling Authentication Free Interaction on Smartphone by Continuous User Tracking
abstract
We propose Auth+Track, a novel authentication model that aims to reduce redundant authentication in everyday smartphone usage. By sparse authentication and continuous tracking of the user’s status, Auth+Track eliminates the “gap” authentication between fragmented sessions and enables “Authentication Free when User is Around”. To instantiate the Auth+Track model, we present PanoTrack, a prototype that integrates body and near field hand information for user tracking. We install a fisheye camera on the top of the phone to achieve a panoramic vision that can capture both user’s body and on-screen hands. Based on the captured video stream, we develop an algorithm to extract 1) features for user tracking, including body keypoints and their temporal and spatial association, near field hand status, and 2) features for user identity assignment. The results of our user studies validate the feasibility of PanoTrack and demonstrate that Auth+Track not only improves the authentication efficiency but also enhances user experiences with better usability.
Chun Yu, Xiaoying Wei, Xuhai Xu, Yongquan Hu, Yuntao Wang 0001, Yuanchun Shi
CHI4
2021 HulaMove: Using Commodity IMU for Waist Interaction
abstract
We present HulaMove, a novel interaction technique that leverages the movement of the waist as a new eyes-free and hands-free input method for both the physical world and the virtual world. We first conducted a user study (N=12) to understand users’ ability to control their waist. We found that users could easily discriminate eight shifting directions and two rotating orientations, and quickly confirm actions by returning to the original position (quick return). We developed a design space with eight gestures for waist interaction based on the results and implemented an IMU-based real-time system. Using a hierarchical machine learning model, our system could recognize waist gestures at an accuracy of 97.5%. Finally, we conducted a second user study (N=12) for usability testing in both real-world scenarios and virtual reality settings. Our usability study indicated that HulaMove significantly reduced interaction time by 41.8% compared to a touch screen method, and greatly improved users’ sense of presence in the virtual world. This novel technique provides an additional input method when users’ eyes or hands are busy, accelerates users’ daily operations, and augments their immersive experience in the virtual world.
Xuhai Xu, Tianyi Yuan, Liang He 0005, Xin Liu 0034, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey
CHI1
2021 Voicemoji: Emoji Entry Using Voice for Visually Impaired People
abstract
Keyboard-based emoji entry can be challenging for people with visual impairments: users have to sequentially navigate emoji lists using screen readers to find their desired emojis, which is a slow and tedious process. In this work, we explore the design and benefits of emoji entry with speech input, a popular text entry method among people with visual impairments. After conducting interviews to understand blind or low vision (BLV) users’ current emoji input experiences, we developed Voicemoji, which (1) outputs relevant emojis in response to voice commands, and (2) provides context-sensitive emoji suggestions through speech output. We also conducted a multi-stage evaluation study with six BLV participants from the United States and six BLV participants from China, finding that Voicemoji significantly reduced entry time by 91.2% and was preferred by all participants over the Apple iOS keyboard. Based on our findings, we present Voicemoji as a feasible solution for voice-based emoji entry.
Mingrui Ray Zhang, Ruolin Wang, Xuhai Xu, Qisheng Li, Ather Sharif, Jacob O. Wobbrock
CHI3
2021 ReflecTrack: Enabling 3D Acoustic Position Tracking Using Commodity Dual-Microphone Smartphones
abstract
3D position tracking on smartphones has the potential to unlock a variety of novel applications, but has not been made widely available due to limitations in smartphone sensors. In this paper, we propose ReflecTrack, a novel 3D acoustic position tracking method for commodity dual-microphone smartphones. A ubiquitous speaker (e.g., smartwatch or earbud) generates inaudible Frequency Modulated Continuous Wave (FMCW) acoustic signals that are picked up by both smartphone microphones. To enable 3D tracking with two microphones, we introduce a reflective surface that can be easily found in everyday objects near the smartphone. Thus, the microphones can receive sound from the speaker and echoes from the surface for FMCW-based acoustic ranging. To simultaneously estimate the distances from the direct and reflective paths, we propose the echo-aware FMCW technique with a new signal pattern and target detection process. Our user study shows that ReflecTrack achieves a median error of 28.4 mm in the 60cm × 60cm × 60cm space and 22.1 mm in the 30cm × 30cm × 30cm space for 3D positioning. We demonstrate the easy accessibility of ReflecTrack using everyday surfaces and objects with several typical applications of 3D position tracking, including 3D input for smartphones, fine-grained gesture recognition, and motion tracking in smartphone-based VR systems.
Yuzhou Zhuang, Yuntao Wang 0001, Yukang Yan, Xuhai Xu, Yuanchun Shi
UIST4
2021 Understanding practices and needs of researchers in human state modeling by passive mobile sensing
Xuhai Xu, Jennifer Mankoff, Anind K. Dey
CCF Trans. Pervasive Comput. Interact.1
2020 EarBuddy: Enabling On-Face Interaction via Wireless Earbuds
abstract
Past research regarding on-body interaction typically requires custom sensors, limiting their scalability and generalizability. We propose EarBuddy, a real-time system that leverages the microphone in commercial wireless earbuds to detect tapping and sliding gestures near the face and ears. We develop a design space to generate 27 valid gestures and conducted a user study (N=16) to select the eight gestures that were optimal for both human preference and microphone detectability. We collected a dataset on those eight gestures (N=20) and trained deep learning models for gesture detection and classification. Our optimized classifier achieved an accuracy of 95.3%. Finally, we conducted a user study (N=12) to evaluate EarBuddy's usability. Our results show that EarBuddy can facilitate novel interaction and that users feel very positively about the system. EarBuddy provides a new eyes-free, socially acceptable input method that is compatible with commercial wireless earbuds and has the potential for scalability and generalizability
Xuhai Xu, Haitian Shi, Xin Yi 0001, Wenjia Liu, Yukang Yan, Yuanchun Shi, Alexander Mariakakis, Jennifer Mankoff, Anind K. Dey
CHI1
2020 FrownOnError: Interrupting Responses from Smart Speakers by Facial Expressions
abstract
In the conversations with smart speakers, misunderstandings of users' requests lead to erroneous responses. We propose FrownOnError, a novel interaction technique that enables users to interrupt the responses by intentional but natural facial expressions. This method leverages the human nature that the facial expression changes when we receive unexpected responses. We conducted a first user study (N=12) to understand users' intuitive reactions to the correct and incorrect responses. Our results reveal the significant difference in the frequency of occurrence and intensity of users' facial expressions between two conditions, and frowning and raising eyebrows are intuitive to perform and easy to control. Our second user study (N=16) evaluated the user experience and interruption efficiency of FrownOnError and the third user study (N=12) explored suitable conversation recovery strategies after the interruptions. Our results show that FrownOnError can be accurately detected (precision: 97.4%, recall: 97.6%), provides the most timely interruption compared to the baseline methods of wake-up word and button press, and is rated as most intuitive and easiest to be performed by users.
Yukang Yan, Chun Yu, Wengrui Zheng, Ruining Tang, Xuhai Xu, Yuanchun Shi
CHI5
2020 Effects of Past Interactions on User Experience with Recommended Documents
abstract
Recommender systems are commonly used in entertainment, news, e-commerce, and social media. Document recommendation is a new and under-explored application area, in which both re-finding and discovery of documents need to be supported. In this paper we provide an initial exploration of users' experience with recommended documents, with a focus on how prior interactions influence recognition and interest. Through a field study of more than 100 users, we investigate the effects of past interactions with recommended documents on users' recognition of, prior intent to open, and interest in the documents. We examined different presentations of interaction history, and the recency and richness of prior interaction. We found that presentation only influenced recognition time. Our findings also indicate that people are more likely to recognize documents they had accessed recently and to do so more quickly. Similarly, documents that people had interacted with more deeply were also more frequently and quickly recognized. However, people were more interested in older documents or those with which they had less involved interactions. This finding suggests that in addition to helping users quickly access documents they intend to re-find, document recommendation can add value in helping users discover other documents. Our results offer implications for designing document recommendation systems that help users fulfil different needs.
Farnaz Jahanbakhsh, Ahmed Awadallah 0001, Susan T. Dumais, Xuhai Xu
CHIIR4
2020 Understanding User Behavior For Document Recommendation
abstract
Personalized document recommendation systems aim to provide users with a quick shortcut to the documents they may want to access next, usually with an explanation about why the document is recommended. Previous work explored various methods for better recommendations and better explanations in different domains. However, there are few efforts that closely study how users react to the recommended items in a document recommendation scenario. We conducted a large-scale log study of users’ interaction behavior with the explainable recommendation on one of the largest cloud document platforms office.com. Our analysis reveals a number of factors, including display position, file type, authorship, recency of last access, and most importantly, the recommendation explanations, that are associated with whether users will recognize or open the recommended documents. Moreover, we specifically focus on explanations and conduct an online experiment to investigate the influence of different explanations on user behavior. Our analysis indicates that the recommendations help users access their documents significantly faster, but sometimes users miss a recommendation and resort to other more complicated methods to open the documents. Our results suggest opportunities to improve explanations and more generally the design of systems that provide and explain recommendations for documents.
Xuhai Xu, Ahmed Awadallah 0001, Susan T. Dumais, Farheen Omar, Bogdan Popp, Robert Rounthwaite, Farnaz Jahanbakhsh
WWW1
2019 Clench Interface: Novel Biting Input Techniques
abstract
People eat every day and biting is one of the most fundamental and natural actions that they perform on a daily basis. Existing work has explored tooth click location and jaw movement as input techniques, however clenching has the potential to add control to this input channel. We propose clench interaction that leverages clenching as an actively controlled physiological signal that can facilitate interactions. We conducted a user study to investigate users' ability to control their clench force. We found that users can easily discriminate three force levels, and that they can quickly confirm actions by unclenching (quick release). We developed a design space for clench interaction based on the results and investigated the usability of the clench interface. Participants preferred the clench over baselines and indicated a willingness to use clench-based interactions. This novel technique can provide an additional input method in cases where users' eyes or hands are busy, augment immersive experiences such as virtual/augmented reality, and assist individuals with disabilities.
Xuhai Xu, Chun Yu, Anind K. Dey, Jennifer Mankoff
CHI1
2018 VMotion: Designing a Seamless Walking Experience in VR
abstract
Physically walking in virtual reality can provide a satisfying sense of presence. However, natural locomotion in virtual worlds larger than the tracked space remains a practical challenge. Numerous redirected walking techniques have been proposed to overcome space limitations but they often require rapid head rotation, sometimes induced by distractors, to keep the scene rotation imperceptible. We propose a design methodology of seamlessly integrating redirection into the virtual experience that takes advantage of the perceptual phenomenon of inattentional blindness. Additionally, we present four novel visibility control techniques that work with our design methodology to minimize disruption to the user experience commonly found in existing redirection techniques. A user study (N = 16) shows that our techniques are imperceptible and users report significantly less dizziness when using our methods. The illusion of unconstrained walking in a large area (16 x 8m) is maintained even though users are limited to a smaller (3.5 x 3.5m) physical space.
Misha Sra, Xuhai Xu, Aske Mottelson, Pattie Maes
Conference on Designing Interactive Systems2
2018 BreathVR: Leveraging Breathing as a Directly Controlled Interface for Virtual Reality Games
abstract
With virtual reality head-mounted displays rapidly becoming accessible to mass audiences, there is growing interest in new forms of natural input techniques to enhance immersion and engagement for players. Research has explored physiological input for enhancing immersion in single player games through indirectly controlled signals like heart rate or galvanic skin response. In this paper, we propose breathing as a directly controlled physiological signal that can facilitate unique and engaging play experiences through natural interaction in single and multiplayer virtual reality games. Our study (N = 16) shows that participants report a higher sense of presence and find the gameplay more fun and challenging when using our breathing actions. From study observations and analysis we present five design strategies that can aid virtual reality game designers interested in using directly controlled forms of physiological input.
Misha Sra, Xuhai Xu, Pattie Maes
CHI2
2018 ForceBoard: Subtle Text Entry Leveraging Pressure
abstract
We present ForceBoard, a pressure-based input technique that enables text entry by subtle finger motion. To enter text, users apply pressure to control a multi-letter-wide sliding cursor on a one-dimensional keyboard with alphabetical ordering, and confirm the selection with a quick release. We examined the error model of pressure control for successive and error-tolerant input, which was incorporated into a Bayesian algorithm to infer user input. A user study showed that, after a 10-minute training, the average text entry rate reached 4.2 wpm (Words Per Minute) for character-level input, and 11.0 wpm for word-level input. Users reported that ForceBoard was easy to learn and interesting to use. These results demonstrated the feasibility of applying pressure as the main channel for text entry. We conclude by discussing the limitation, as well as the potential of ForceBoard to support interaction with constraints from form factor, social concern and physical environments.
Mingyuan Zhong 0001, Chun Yu, Xuhai Xu, Yuanchun Shi
CHI4
2018 Hand range interface: information always at hand with a body-centric mid-air input surface
abstract
Most interfaces of our interactive devices such as phones and laptops are flat and are built as external devices in our environment, disconnected from our bodies. Therefore, we need to carry them with us in our pocket or in a bag and accommodate our bodies to their design by sitting at a desk or holding the device in our hand. We propose Hand Range Interface, an input surface that is always at our fingertips. This body-centric interface is a semi-sphere attached to a user's wrist, with a radius the same as the distance from the wrist to the index finger. We prototyped the concept in virtual reality and conducted a user study with a pointing task. The input surface can be designed as rotating with the wrist or fixed relative to the wrist. We evaluated and compared participants' subjective physical comfort level, pointing speed and pointing accuracy on the interface that was divided into 64 regions. We found that the interface whose orientation was fixed had a much better performance, with 41.2% higher average comfort score, 40.6% shorter average pointing time and 34.5% lower average error. Our results revealed interesting insights on user performance and preference of different regions on the interface. We concluded with a set of guidelines for future designers and developers on how to develop this type of new body-centric input surface.
Xuhai Xu, Alexandru Dancu, Pattie Maes, Suranga Nanayakkara
MobileHCI1
2017 GalVR: a novel collaboration interface using GVS
abstract
GalVR is a navigation interface that uses galvanic vestibular stimulation (GVS) during walking to cause users to turn from their planned trajectory. We explore GalVR for collaborative navigation in a two-player virtual reality (VR) game. The interface affords a novel game design that exploits the differences in first and third person perspectives, allowing VR and non-VR users to share a play experience. By introducing interdependence arising from dissimilar points of view, players can uniquely contribute to the shared experience based on their roles. We detail the design of our asymmetrical game, Dark Room and present some insights from a pilot study. Trust emerged as the defining factor for successful play.
Misha Sra, Xuhai Xu, Pattie Maes
VRST2