Baosheng James Hou

dblp:266/2756 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-7413-8921ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Aligner, Nodder, and Winker: Creating Complete, Hands-Free Interaction Techniques that Unify Precise, UI-Independent Selection and Dragging ETRA016
abstract
Hands-free interaction is essential when users’ hands are occupied with primary tasks. A key challenge is unifying precise, UI-independent selection and continuous dragging without explicit mode switching. This paper introduces two novel techniques: Aligner (spatial eye-head alignment) and Nodder (gestural decomposition of a nod); and evaluates them against an established Winker technique, a gestural clutch using single-eye closure. A user study with controlled Fitts’ law and dragging tasks, alongside a practical application, evaluated the techniques. Results show that Aligner offers superior selection precision but creates a visuomotor conflict during dragging; Nodder optimizes movement efficiency for large-amplitude targets despite higher physical effort; and Winker provides the fastest performance but is susceptible to accidental activation. These findings inform the design of future hands-free systems by highlighting the context-dependent nature of this interaction challenge.
Baosheng James Hou, Pavel Manakhov, Jinghui Hu, Hans-Werner Gellersen
Proc. ACM Hum. Comput. Interact.1
2026 Directed or Guided? Classification of Gaze Attention Shifts based on Eye and Head Movement ETRA002
abstract
Gaze can be directed at will, or guided by objects that draw our attention. We propose classifying gaze shifts as either directed or guided, since these two forms of attention have different implications for HCI. Directed attention may serve as stronger indicator of users’ planning and intent, whereas guided attention reflects interface efficacy in guiding information acquisition. We introduce a method based on eye and head movement features during gaze shifts, using data collected in virtual reality to train and evaluate a machine learning model, which we then validate in application to visual search. Our results show that this classification is both feasible and practical, extending established uses of eye tracking in HCI. This is significant because it enables a new level of analysis of visual attention.
Anam Ahmad Khan, Baosheng James Hou, Hans-Werner Gellersen
Proc. ACM Hum. Comput. Interact.2
2025 Online-EYE: Multimodal Implicit Eye Tracking Calibration for XR
abstract
Unlike other inputs for extended reality (XR) that work out of the box, eye tracking typically requires custom calibration per user or session. We present a multimodal inputs approach for implicit calibration of eye tracker in VR, leveraging UI interaction for continuous, background calibration. Our method analyzes gaze data alongside controller interaction with UI elements, and employing ML techniques it continuously refines the calibration matrix without interrupting users from their current tasks. Potentially eliminating the need for explicit calibration. We demonstrate the accuracy and effectiveness of this implicit approach across various tasks and real time applications achieving comparable eye tracking accuracy to native, explicit calibration. While our evaluation focuses on VR and controller-based interactions, we anticipate the broader applicability of this approach to various XR devices and input modalities.
Baosheng James Hou, Lucy Abramyan, Prasanthi Gurumurthy, Haley Adams, Ivana Tosic Rodgers, Eric J. Gonzalez, Khushman Patel, Andrea Colaco, Ken Pfeuffer, Hans-Werner Gellersen, Karan Ahuja, Mar González-Franco
CHI1
2024 Understanding the Impact of the Reality-Virtuality Continuum on Visual Search Using Fixation-Related Potentials and Eye Tracking Features
abstract
While Mixed Reality allows the seamless blending of digital content in users' surroundings, it is unclear if its fusion with physical information impacts users' perceptual and cognitive resources differently. While the fusion of digital and physical objects provides numerous opportunities to present additional information, it also introduces undesirable side effects, such as split attention and increased visual complexity. We conducted a visual search study in three manifestations of mixed reality (Augmented Reality, Augmented Virtuality, Virtual Reality) to understand the effects of the environment on visual search behavior. We conducted a multimodal evaluation measuring Fixation-Related Potentials (FRPs), alongside eye tracking to assess search efficiency, attention allocation, and behavioral measures. Our findings indicate distinct patterns in FRPs and eye-tracking data that reflect varying cognitive demands across environments. Specifically, AR environments were associated with increased workload, as indicated by decreased FRP - P3 amplitudes and more scattered eye movement patterns, impairing users' ability to identify target information efficiently. Participants reported AR as the most demanding and distracting environment. These insights inform design implications for MR adaptive systems, emphasizing the need for interfaces that dynamically respond to user cognitive load based on physiological inputs.
Francesco Chiossi, Uwe Gruenefeld, Baosheng James Hou, Joshua Newn, Changkun Ou, Rulu Liao, Robin Welsch, Sven Mayer
Proc. ACM Hum. Comput. Interact.3
2024 GazeSwitch: Automatic Eye-Head Mode Switching for Optimised Hands-Free Pointing
abstract
This paper contributes GazeSwitch, an ML-based technique that optimises the real-time switching between eye and head modes for fast and precise hands-free pointing. GazeSwitch reduces false positives from natural head movements and efficiently detects head gestures for input, resulting in an effective hands-free and adaptive technique for interaction. We conducted two user studies to evaluate its performance and user experience. Comparative analyses with baseline switching techniques, Eye+Head Pinpointing (manual) and BimodalGaze (threshold-based) revealed several trade-offs. We found that GazeSwitch provides a natural and effortless experience but trades off control and stability compared to manual mode switching, and requires less head movement compared to BimodalGaze. This work demonstrates the effectiveness of machine learning approach to learn and adapt to patterns in head movement, allowing us to better leverage the synergistic relation between eye and head input modalities for interaction in mixed and extended reality.
Baosheng James Hou, Joshua Newn, Ludwig Sidenmark, Anam Ahmad Khan, Hans-Werner Gellersen
Proc. ACM Hum. Comput. Interact.1
2023 Classifying Head Movements to Separate Head-Gaze and Head Gestures as Distinct Modes of Input
abstract
Head movement is widely used as a uniform type of input for human-computer interaction. However, there are fundamental differences between head movements coupled with gaze in support of our visual system, and head movements performed as gestural expression. Both Head-Gaze and Head Gestures are of utility for interaction but differ in their affordances. To facilitate the treatment of Head-Gaze and Head Gestures as separate types of input, we developed HeadBoost as a novel classifier, achieving high accuracy in classifying gaze-driven versus gestural head movement (F1-Score: 0.89). We demonstrate the utility of the classifier with three applications: gestural input while avoiding unintentional input by Head-Gaze; target selection with Head-Gaze while avoiding Midas Touch by head gestures; and switching of cursor control between Head-Gaze for fast positioning and Head Gesture for refinement. The classification of Head-Gaze and Head Gesture allows for seamless head-based interaction while avoiding false activation.
Baosheng James Hou, Joshua Newn, Ludwig Sidenmark, Anam Ahmad Khan, Per Baekgaard, Hans-Werner Gellersen
CHI1
2023 Exploring Eye Expressions for Enhancing EOG-Based Interaction
Joshua Newn, Sophia Quesada, Baosheng James Hou, Anam Ahmad Khan, Florian Weidner, Hans-Werner Gellersen
INTERACT (4)3
2022 Feasibility of a Device for Gaze Interaction by Visually-Evoked Brain Signals
abstract
A dry-electrode head-mounted sensor for visually-evoked electroencephalogram (EEG) signals has been introduced to the gamer market, and provides wireless, low-cost tracking of a user’s gaze fixation on target areas in real-time. Unlike traditional EEG sensors, this new device is easy to set up for non-professionals. We conducted a Fitts’ law study (N = 6) and found the mean throughput (TP) to be 0.82 bits/s. The sensor yielded robust performance with error rates below 1%. The overall median activation time (AT) was 2.35 s with a minuscule difference between one or nine concurrent targets. We discuss whether the method might supplement camera-based gaze interaction, for example, in gaze typing or wheelchair control, and note some limitations, such as a slow AT, the difficulty of calibration with thick hair, and the limit of 10 concurrent targets.
Baosheng James Hou, John Paulin Hansen, Cihan Uyanik, Per Baekgaard, Sadasivan Puthusserypady, Jacopo M. Araujo, I. Scott MacKenzie
ETRA1