EDBT 2026 Demo / reviewers in the wild / expert
Joshua Newn
dblp:173/9876
· DBLP profile ↗
23ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-5769-6297ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 19 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewabstractMultimodal interaction has long promised to make interfaces more intuitive and effective by combining complementary inputs. Among these, gaze and speech form a compelling pairing: gaze provides rapid spatial grounding, while speech conveys rich semantic information. Together, they offer rich cues for understanding user behaviour and intent. Yet despite decades of exploration, the research remains fragmented, making this synthesis timely as these inputs mature and are integrated into consumer-ready devices. This scoping review examined 103 studies published between 1991 and 2025, organised into explicit, where users intentionally provide gaze and speech, and implicit, where systems leverage users' natural behaviours to support interaction. Across both, we identified recurring ways for combining gaze and speech to resolve ambiguity, ground references, and support adaptivity. We contribute a synthesis of research on their combined use while highlighting challenges of temporal alignment, fusion and privacy, offering guidance for future research toward richer multimodal human-computer interaction. Anam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou, Andrea Bianchi, Eduardo Velloso, Hans-Werner Gellersen, Joshua Newn |
CHI | 8 |
| 2025 | HeadShift: Head Pointing with Dynamic Control-Display GainabstractHead pointing is widely used for hands-free input in head-mounted displays (HMDs). The primary role of head movement in an HMD is to control the viewport based on absolute mapping of head rotation to the 3D environment. Head pointing is conventionally supported by the same 1:1 mapping of input with a cursor fixed in the centre of the view, but this requires exaggerated head movement and limits input granularity. In this work, we propose to adopt dynamic gain to improve ergonomics and precision, and introduce the HeadShift technique. The design of HeadShift is grounded in natural eye-head coordination to manage control of the viewport and the cursor at different speeds. We evaluated HeadShift in a Fitts’ Law experiment and on three different applications in VR, finding the technique to reduce error rate and effort. The findings are significant as they show that gain can be adopted effectively for head pointing while ensuring that the cursor is maintained within a comfortable eye-in-head viewing range. Ludwig Sidenmark, Florian Weidner, Joshua Newn, Hans-Werner Gellersen |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2024 | Understanding the Impact of the Reality-Virtuality Continuum on Visual Search Using Fixation-Related Potentials and Eye Tracking FeaturesabstractWhile Mixed Reality allows the seamless blending of digital content in users' surroundings, it is unclear if its fusion with physical information impacts users' perceptual and cognitive resources differently. While the fusion of digital and physical objects provides numerous opportunities to present additional information, it also introduces undesirable side effects, such as split attention and increased visual complexity. We conducted a visual search study in three manifestations of mixed reality (Augmented Reality, Augmented Virtuality, Virtual Reality) to understand the effects of the environment on visual search behavior. We conducted a multimodal evaluation measuring Fixation-Related Potentials (FRPs), alongside eye tracking to assess search efficiency, attention allocation, and behavioral measures. Our findings indicate distinct patterns in FRPs and eye-tracking data that reflect varying cognitive demands across environments. Specifically, AR environments were associated with increased workload, as indicated by decreased FRP - P3 amplitudes and more scattered eye movement patterns, impairing users' ability to identify target information efficiently. Participants reported AR as the most demanding and distracting environment. These insights inform design implications for MR adaptive systems, emphasizing the need for interfaces that dynamically respond to user cognitive load based on physiological inputs. Francesco Chiossi, Uwe Gruenefeld, Baosheng James Hou, Joshua Newn, Changkun Ou, Rulu Liao, Robin Welsch, Sven Mayer |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2024 | GazeSwitch: Automatic Eye-Head Mode Switching for Optimised Hands-Free PointingabstractThis paper contributes GazeSwitch, an ML-based technique that optimises the real-time switching between eye and head modes for fast and precise hands-free pointing. GazeSwitch reduces false positives from natural head movements and efficiently detects head gestures for input, resulting in an effective hands-free and adaptive technique for interaction. We conducted two user studies to evaluate its performance and user experience. Comparative analyses with baseline switching techniques, Eye+Head Pinpointing (manual) and BimodalGaze (threshold-based) revealed several trade-offs. We found that GazeSwitch provides a natural and effortless experience but trades off control and stability compared to manual mode switching, and requires less head movement compared to BimodalGaze. This work demonstrates the effectiveness of machine learning approach to learn and adapt to patterns in head movement, allowing us to better leverage the synergistic relation between eye and head input modalities for interaction in mixed and extended reality. Baosheng James Hou, Joshua Newn, Ludwig Sidenmark, Anam Ahmad Khan, Hans-Werner Gellersen |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | Speech-Augmented Cone-of-Vision for Exploratory Data AnalysisabstractMutual awareness of visual attention is crucial for successful collaboration. Previous research has explored various ways to represent visual attention, such as field-of-view visualizations and cursor visualizations based on eye-tracking, but these methods have limitations. Verbal communication is often utilized as a complementary strategy to overcome such disadvantages. This paper proposes a novel method that combines verbal communication with the Cone of Vision to improve gaze inference and mutual awareness in VR. We conducted a within-group study with pairs of participants who performed a collaborative analysis of data visualizations in VR. We found that our proposed method provides a better approximation of eye gaze than the approximation provided by head direction. Furthermore, we release the first collaborative head, eyes, and verbal behaviour dataset. The results of this study provide a foundation for investigating the potential of verbal communication as a tool for enhancing visual cues for joint attention. Riccardo Bovo, Daniele Giunchi, Ludwig Sidenmark, Joshua Newn, Hans-Werner Gellersen, Enrico Costanza, Thomas Heinis |
CHI | 4 |
| 2023 | Classifying Head Movements to Separate Head-Gaze and Head Gestures as Distinct Modes of InputabstractHead movement is widely used as a uniform type of input for human-computer interaction. However, there are fundamental differences between head movements coupled with gaze in support of our visual system, and head movements performed as gestural expression. Both Head-Gaze and Head Gestures are of utility for interaction but differ in their affordances. To facilitate the treatment of Head-Gaze and Head Gestures as separate types of input, we developed HeadBoost as a novel classifier, achieving high accuracy in classifying gaze-driven versus gestural head movement (F1-Score: 0.89). We demonstrate the utility of the classifier with three applications: gestural input while avoiding unintentional input by Head-Gaze; target selection with Head-Gaze while avoiding Midas Touch by head gestures; and switching of cursor control between Head-Gaze for fast positioning and Head Gesture for refinement. The classification of Head-Gaze and Head Gesture allows for seamless head-based interaction while avoiding false activation. Baosheng James Hou, Joshua Newn, Ludwig Sidenmark, Anam Ahmad Khan, Per Baekgaard, Hans-Werner Gellersen |
CHI | 2 |
| 2023 | Vergence Matching: Inferring Attention to Objects in 3D Environments for Gaze-Assisted SelectionabstractGaze pointing is the de facto standard to infer attention and interact in 3D environments but is limited by motor and sensor limitations. To circumvent these limitations, we propose a vergence-based motion correlation method to detect visual attention toward very small targets. Smooth depth movements relative to the user are induced on 3D objects, which cause slow vergence eye movements when looked upon. Using the principle of motion correlation, the depth movements of the object and vergence eye movements are matched to determine which object the user is focussing on. In two user studies, we demonstrate how the technique can reliably infer gaze attention on very small targets, systematically explore how different stimulus motions affect attention detection, and show how the technique can be extended to multi-target selection. Finally, we provide example applications using the concept and design guidelines for small target and accuracy-independent attention detection in 3D environments. Ludwig Sidenmark, Christopher Clarke, Joshua Newn, Mathias N. Lystbæk, Ken Pfeuffer, Hans-Werner Gellersen |
CHI | 3 |
| 2023 | Exploring Eye Expressions for Enhancing EOG-Based Interaction
Joshua Newn, Sophia Quesada, Baosheng James Hou, Anam Ahmad Khan, Florian Weidner, Hans-Werner Gellersen |
INTERACT (4) | 1 |
| 2023 | Comparing Gaze, Head and Controller Selection of Dynamically Revealed Targets in Head-Mounted DisplaysabstractThis paper presents a head-mounted virtual reality study that compared gaze, head, and controller pointing for selection of dynamically revealed targets. Existing studies on head-mounted 3D interaction have focused on pointing and selection tasks where all targets are visible to the user. Our study compared the effects of screen width (field of view), target amplitude and width, and prior knowledge of target location on modality performance. Results show that gaze and controller pointing are significantly faster than head pointing and that increased screen width only positively impacts performance up to a certain point. We further investigated the applicability of existing pointing models. Our analysis confirmed the suitability of previously proposed two-component models for all modalities while uncovering differences for gaze at known and unknown target positions. Our findings provide new empirical evidence for understanding input with gaze, head, and controller and are significant for applications that extend around the user. Ludwig Sidenmark, Franziska Prummer, Joshua Newn, Hans-Werner Gellersen |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Integrating Gaze and Speech for Enabling Implicit InteractionsabstractGaze and speech are rich contextual sources of information that, when combined, can result in effective and rich multimodal interactions. This paper proposes a machine learning-based pipeline that leverages and combines users’ natural gaze activity, the semantic knowledge from their vocal utterances and the synchronicity between gaze and speech data to facilitate users’ interaction. We evaluated our proposed approach on an existing dataset, which involved 32 participants recording voice notes while reading an academic paper. Using a Logistic Regression classifier, we demonstrate that our proposed multimodal approach maps voice notes with accurate text passages with an average F1-Score of 0.90. Our proposed pipeline motivates the design of multimodal interfaces that combines natural gaze and speech patterns to enable robust interactions. Anam Ahmad Khan, Joshua Newn, James Bailey 0001, Eduardo Velloso |
CHI | 2 |
| 2022 | To type or to speak? The effect of input modality on text understanding during note-takingabstractThough recent technological advances have enabled note-taking through different modalities (e.g., keyboard, digital ink, voice), there is still a lack of understanding of the effect of the modality choice on learning. In this paper, we compared two note-taking input modalities—keyboard and voice—to study their effects on participants’ understanding of learning content. We conducted a study with 60 participants in which they were asked to take notes using voice or keyboard on two independent digital text passages while also making a judgment about their performance on an upcoming test. We built mixed-effects models to examine the effect of the note-taking modality on learners’ text comprehension, the content of notes and their meta-comprehension judgement. Our findings suggest that taking notes using voice leads to a higher conceptual understanding of the text when compared to typing the notes. We also found that using voice triggers generative processes that result in learners taking more elaborate and comprehensive notes. The findings of the study imply that note-taking tools designed for digital learning environments could incorporate voice as an input modality to promote effective note-taking and higher conceptual understanding of the text. Anam Ahmad Khan, Sadia Nawaz, Joshua Newn, Ryan Kelly 0001, Jason M. Lodge, James Bailey 0001, Eduardo Velloso |
CHI | 3 |
| 2021 | Are you with me? Measurement of Learners' Video-Watching Attention with Eye TrackingabstractVideo has become an essential medium for learning. However, there are challenges when using traditional methods to measure how learners attend to lecture videos in video learning analytics, such as difficulty in capturing learners’ attention at a fine-grained level. Therefore, in this paper, we propose a gaze-based metric—“with-me-ness direction” that can measure how learners’ gaze-direction changes when they listen to the instructor’s dialogues in a video-lecture. We analyze the gaze data of 45 participants as they watched a video lecture and measured both the sequences of with-me-ness direction and proportion of time a participant spent looking in each direction throughout the lecture at different levels. We found that although the majority of the time participants followed the instructor’s dialogues, their behaviour of looking-ahead, looking-behind or looking-outside differed by their prior knowledge. These findings open the possibility of using eye-tracking to measure learners’ video-watching attention patterns and examine factors that can influence their attention, thereby helping instructors to design effective learning materials. Namrata Srivastava, Sadia Nawaz, Joshua Newn, Jason M. Lodge, Eduardo Velloso, Sarah M. Erfani, Dragan Gasevic, James Bailey 0001 |
LAK | 3 |
| 2021 | GAVIN: Gaze-Assisted Voice-Based Implicit Note-takingabstractAnnotation is an effective reading strategy people often undertake while interacting with digital text. It involves highlighting pieces of text and making notes about them. Annotating while reading in a desktop environment is considered trivial but, in a mobile setting where people read while hand-holding devices, the task of highlighting and typing notes on a mobile display is challenging. In this article, we introduce GAVIN, a gaze-assisted voice note-taking application, which enables readers to seamlessly take voice notes on digital documents by implicitly anchoring them to text passages. We first conducted a contextual enquiry focusing on participants’ note-taking practices on digital documents. Using these findings, we propose a method which leverages eye-tracking and machine learning techniques to annotate voice notes with reference text passages. To evaluate our approach, we recruited 32 participants performing voice note-taking. Following, we trained a classifier on the data collected to predict text passage where participants made voice notes. Lastly, we employed the classifier to built GAVIN and conducted a user study to demonstrate the feasibility of the system. This research demonstrates the feasibility of using gaze as a resource for implicit anchoring of voice notes, enabling the design of systems that allow users to record voice notes with minimal effort and high accuracy. Anam Ahmad Khan, Joshua Newn, Ryan Kelly 0001, Namrata Srivastava, James Bailey 0001, Eduardo Velloso |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2020 | Faces of Focus: A Study on the Facial Cues of Attentional StatesabstractAutomatically detecting attentional states is a prerequisite for designing interventions to manage attention - knowledge workers' most critical resource. As a first step towards this goal, it is necessary to understand how different attentional states are made discernible through visible cues in knowledge workers. In this paper, we demonstrate the important facial cues to detect attentional states by evaluating a data set of 15 participants that we tracked over a whole workday, which included their challenge and engagement levels. Our evaluation shows that gaze, pitch, and lips part action units are indicators of engaged work; while pitch, gaze movements, gaze angle, and upper-lid raiser action units are indicators of challenging work. These findings reveal a significant relationship between facial cues and both engagement and challenge levels experienced by our tracked participants. Our work contributes to the design of future studies to detect attentional states based on facial cues. Ebrahim Babaei, Namrata Srivastava, Joshua Newn, Qiushi Zhou, Tilman Dingler, Eduardo Velloso |
CHI | 3 |
| 2020 | Combining gaze and AI planning for online human intention recognition
Ronal Singh, Tim Miller 0001, Joshua Newn, Eduardo Velloso, Frank Vetere, Liz Sonenberg |
Artif. Intell. | 3 |
| 2020 | Fully-Occluded Target Selection in Virtual RealityabstractThe presence of fully-occluded targets is common within virtual environments, ranging from a virtual object located behind a wall to a datapoint of interest hidden in a complex visualization. However, efficient input techniques for locating and selecting these targets are mostly underexplored in virtual reality (VR) systems. In this paper, we developed an initial set of seven techniques techniques for fully-occluded target selection in VR. We then evaluated their performance in a user study and derived a set of design implications for simple and more complex tasks from our results. Based on these insights, we refined the most promising techniques and conducted a second, more comprehensive user study. Our results show how factors, such as occlusion layers, target depths, object densities, and the estimation of target locations, can affect technique performance. Our findings from both studies and distilled recommendations can inform the design of future VR systems that offer selections for fully-occluded targets. Difeng Yu, Qiushi Zhou, Joshua Newn, Tilman Dingler, Eduardo Velloso, Jorge Gonçalves 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Eyes-free Target Acquisition During Walking in Immersive Mixed RealityabstractReaching towards out-of-sight objects during walking is a common task in daily life, however the same task can be challenging when wearing immersive Head-Mounted Displays (HMD). In this paper, we investigate the effects of spatial reference frame, walking path curvature, and target placement relative to the body on user performance of manually acquiring out-of-sight targets located around their bodies, as they walk in a spatial-mapping Mixed Reality (MR) environment wearing an immersive HMD. We found that walking and increased path curvature negatively affected the overall spatial accuracy of the performance, and that the performance benefited more from using the torso as the reference frame than the head. We also found that targets placed at maximum reaching distance yielded less error in angular rotation and depth of the reaching arm. We discuss our findings with regard to human walking kinesthetics and the sensory integration in the peripersonal space during locomotion in immersive MR. We provide design guidelines for future immersive MR experience featuring spatial mapping and full-body motion tracking to provide better embodied experience. Qiushi Zhou, Difeng Yu, Martin Reinoso, Joshua Newn, Jorge Gonçalves 0001, Eduardo Velloso |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Biometric Mirror: Exploring Ethical Opinions towards Facial Analysis and Automated Decision-MakingabstractFacial analysis applications are increasingly being applied to inform decision-making processes. However, as global reports of unfairness emerge, governments, academia and industry have recognized the ethical limitations and societal implications of this technology. Alongside initiatives that aim to formulate ethical frameworks, we believe that the public should be invited to participate in the debate. In this paper, we discuss Biometric Mirror, a case study that explored opinions about the ethics of an emerging technology. The interactive application distinguished demographic and psychometric information from people's facial photos and presented speculative scenarios with potential consequences based on their results. We analyzed the interactions with Biometric Mirror and media reports covering the study. Our findings demonstrate the nature of public opinion about the technology's possibilities, reliability, and privacy implications. Our study indicates an opportunity for case study-based digital ethics research, and we provide practical guidelines for designing future studies. Niels Wouters, Ryan Kelly 0001, Eduardo Velloso, Katrin Wolf 0001, Hasan Shahid Ferdous, Joshua Newn, Zaher Joukhadar, Frank Vetere |
Conference on Designing Interactive Systems | 6 |
| 2019 | Frame Analysis of Voice Interaction GameplayabstractVoice control is an increasingly common feature of digital games, but the experience of playing with voice control is often hampered by feelings of embarrassment and dissonance. Past research has recognised these tensions, but has not offered a general model of how they arise and how players respond to them. In this study, we use Erving Goffman's frame analysis, as adapted to the study of games by Conway and Trevillian, to understand the social experience of playing games by voice. Based on 24 interviews with participants who played voice-controlled games in a social setting, we put forward a frame analytic model of gameplay as a social event, along with seven themes that describe how voice interaction enhances or disrupts the player experience. Our results demonstrate the utility of frame analysis for understanding social dissonance in voice interaction gameplay, and point to practical considerations for designers to improve engagement with voice-controlled games. Fraser Allison, Joshua Newn, Wally Smith, Marcus Carter, Martin R. Gibbs |
CHI | 2 |
| 2019 | Designing Interactions with Intention-Aware Gaze-Enabled Artificial Agents
Joshua Newn, Ronal Singh, Fraser Allison, Prashan Madumal, Eduardo Velloso, Frank Vetere |
INTERACT (2) | 1 |
| 2018 | Looks Can Be Deceiving: Using Gaze Visualisation to Predict and Mislead Opponents in Strategic GameplayabstractIn competitive co-located gameplay, players use their opponents' gaze to make predictions about their plans while simultaneously managing their own gaze to avoid giving away their plans. This socially competitive dimension is lacking in most online games, where players are out of sight of each other. We conducted a lab study using a strategic online game; finding that (1) players are better at discerning their opponent's plans when shown a live visualisation of the opponent's gaze, and (2) players who are aware that their gaze is tracked will manipulate their gaze to keep their intentions hidden. We describe the strategies that players employed, to various degrees of success, to deceive their opponent through their gaze behaviour. This gaze-based deception adds an effortful and challenging aspect to the competition. Lastly, we discuss the various implications of our findings and its applicability for future game design. Joshua Newn, Fraser Allison, Eduardo Velloso, Frank Vetere |
CHI | 1 |
| 2017 | Motion Correlation: Selecting Objects by Matching Their MovementabstractSelection is a canonical task in user interfaces, commonly supported by presenting objects for acquisition by pointing. In this article, we consider motion correlation as an alternative for selection. The principle is to represent available objects by motion in the interface, have users identify a target by mimicking its specific motion, and use the correlation between the system’s output with the user’s input to determine the selection. The resulting interaction has compelling properties, as users are guided by motion feedback, and only need to copy a presented motion. Motion correlation has been explored in earlier work but only recently begun to feature in holistic interface designs. We provide a first comprehensive review of the principle, and present an analysis of five previously published works, in which motion correlation underpinned the design of novel gaze and gesture interfaces for diverse application contexts. We derive guidelines for motion correlation algorithms, motion feedback, choice of modalities, overall design of motion correlation interfaces, and identify opportunities and challenges identified for future research and design. Eduardo Velloso, Marcus Carter, Joshua Newn, Augusto Esteves, Christopher Clarke, Hans-Werner Gellersen |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2016 | Multimodal Segmentation on a Large Interactive Tabletop: Extending Interaction on Horizontal Surfaces with GazeabstractEye tracking is a promising input modality for interactive tabletops. However, issues such as eyelid occlusion and the viewing angle at distant positions present significant challenges for remote gaze tracking in this setting. We present the results of two studies that explore the way gaze interaction can be enabled. Our first study contributes the results from an empirical investigation of gaze accuracy on a large horizontal surface, finding gaze to be unusable close to the user (due to eyelid occlusion), accurate at arm's length, and only precise horizontally at large distances. In consideration of these results, we propose two solutions for the design of interactive systems that utilise remote gaze-tracking on the tabletop; multimodal segmentation and the use of X-Gaze-our novel technique-to interact with out-of-reach objects. Our second study evaluates and validates both these solutions in a Video-on-Demand application, presenting immediate opportunities for remote-gaze interaction on horizontal surfaces. Joshua Newn, Eduardo Velloso, Marcus Carter, Frank Vetere |
ISS | 1 |