Anam Ahmad Khan

dblp:211/4074 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-1620-1902ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 5 first-author · 10 since 2021
YearPublicationVenuePosition
2026 AtoMix: Fostering Structured Visual Ideation for Remote Groups through Atomic Composition and Cross‑Pollination
abstract
6-3-5 brainwriting is a structured, silent ideation method where 6 participants each generate 3 ideas in 5 minutes across 6 rounds, yielding up to 108 ideas and enabling equal participation, rapid ideation and cumulative development. Many ideas are fundamentally visual, making generative image models especially useful for accelerating visual exploration and making abstract ideas quickly tangible, but existing tools cannot fully support visual composition, where ideas are built from atomic elements and iteratively remixed. A formative study (N=18) showed that even with GenAI embedded in a 6-3-5 structure, participants invested effort in crafting monolithic prompts yet struggled to obtain intended results, build on others’ ideas, or work modularly, leading to redundancy and limited divergence. To address these issues, we present AtoMix, a collaborative visual interface for remote 6-3-5 brainwriting that supports granularity and cross-pollination through atomic prompting, canvas composition, and reuse of image fragments. In a comparative study (N=24), AtoMix afforded more fine-grained interaction than the baseline and better supported collaborative visual ideation.
Heeji Jolie Kim, Anam Ahmad Khan, SeungHui Huh, Andrea Bianchi
Creativity & Cognition2
2026 Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping Review
abstract
Multimodal interaction has long promised to make interfaces more intuitive and effective by combining complementary inputs. Among these, gaze and speech form a compelling pairing: gaze provides rapid spatial grounding, while speech conveys rich semantic information. Together, they offer rich cues for understanding user behaviour and intent. Yet despite decades of exploration, the research remains fragmented, making this synthesis timely as these inputs mature and are integrated into consumer-ready devices. This scoping review examined 103 studies published between 1991 and 2025, organised into explicit, where users intentionally provide gaze and speech, and implicit, where systems leverage users' natural behaviours to support interaction. Across both, we identified recurring ways for combining gaze and speech to resolve ambiguity, ground references, and support adaptivity. We contribute a synthesis of research on their combined use while highlighting challenges of temporal alignment, fusion and privacy, offering guidance for future research toward richer multimodal human-computer interaction.
Anam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou, Andrea Bianchi, Eduardo Velloso, Hans-Werner Gellersen, Joshua Newn
CHI1
2026 Directed or Guided? Classification of Gaze Attention Shifts based on Eye and Head Movement ETRA002
abstract
Gaze can be directed at will, or guided by objects that draw our attention. We propose classifying gaze shifts as either directed or guided, since these two forms of attention have different implications for HCI. Directed attention may serve as stronger indicator of users’ planning and intent, whereas guided attention reflects interface efficacy in guiding information acquisition. We introduce a method based on eye and head movement features during gaze shifts, using data collected in virtual reality to train and evaluate a machine learning model, which we then validate in application to visual search. Our results show that this classification is both feasible and practical, extending established uses of eye tracking in HCI. This is significant because it enables a new level of analysis of visual attention.
Anam Ahmad Khan, Baosheng James Hou, Hans-Werner Gellersen
Proc. ACM Hum. Comput. Interact.1
2024 GazeSwitch: Automatic Eye-Head Mode Switching for Optimised Hands-Free Pointing
abstract
This paper contributes GazeSwitch, an ML-based technique that optimises the real-time switching between eye and head modes for fast and precise hands-free pointing. GazeSwitch reduces false positives from natural head movements and efficiently detects head gestures for input, resulting in an effective hands-free and adaptive technique for interaction. We conducted two user studies to evaluate its performance and user experience. Comparative analyses with baseline switching techniques, Eye+Head Pinpointing (manual) and BimodalGaze (threshold-based) revealed several trade-offs. We found that GazeSwitch provides a natural and effortless experience but trades off control and stability compared to manual mode switching, and requires less head movement compared to BimodalGaze. This work demonstrates the effectiveness of machine learning approach to learn and adapt to patterns in head movement, allowing us to better leverage the synergistic relation between eye and head input modalities for interaction in mixed and extended reality.
Baosheng James Hou, Joshua Newn, Ludwig Sidenmark, Anam Ahmad Khan, Hans-Werner Gellersen
Proc. ACM Hum. Comput. Interact.4
2023 Classifying Head Movements to Separate Head-Gaze and Head Gestures as Distinct Modes of Input
abstract
Head movement is widely used as a uniform type of input for human-computer interaction. However, there are fundamental differences between head movements coupled with gaze in support of our visual system, and head movements performed as gestural expression. Both Head-Gaze and Head Gestures are of utility for interaction but differ in their affordances. To facilitate the treatment of Head-Gaze and Head Gestures as separate types of input, we developed HeadBoost as a novel classifier, achieving high accuracy in classifying gaze-driven versus gestural head movement (F1-Score: 0.89). We demonstrate the utility of the classifier with three applications: gestural input while avoiding unintentional input by Head-Gaze; target selection with Head-Gaze while avoiding Midas Touch by head gestures; and switching of cursor control between Head-Gaze for fast positioning and Head Gesture for refinement. The classification of Head-Gaze and Head Gesture allows for seamless head-based interaction while avoiding false activation.
Baosheng James Hou, Joshua Newn, Ludwig Sidenmark, Anam Ahmad Khan, Per Baekgaard, Hans-Werner Gellersen
CHI4
2023 HotFoot: Foot-Based User Identification Using Thermal Imaging
abstract
We propose a novel method for seamlessly identifying users by combining thermal and visible feet features. While it is known that users’ feet have unique characteristics, these have so far been underutilized for biometric identification, as observing those features often requires the removal of shoes and socks. As thermal cameras are becoming ubiquitous, we foresee a new form of identification, using feet features and heat traces to reconstruct the footprint even while wearing shoes or socks. We collected a dataset of users’ feet (N = 21), wearing three types of footwear (personal shoes, standard shoes, and socks) on three floor types (carpet, laminate, and linoleum). By combining visual and thermal features, an AUC between 91.1% and 98.9%, depending on floor type and shoe type can be achieved, with personal shoes on linoleum floor performing best. Our findings demonstrate the potential of thermal imaging for continuous and unobtrusive user identification.
Alia Saad, Kian Izadi, Anam Ahmad Khan, Pascal Knierim, Stefan Schneegaß, Florian Alt, Yomna Abdelrahman
CHI3
2023 Exploring Eye Expressions for Enhancing EOG-Based Interaction
Joshua Newn, Sophia Quesada, Baosheng James Hou, Anam Ahmad Khan, Florian Weidner, Hans-Werner Gellersen
INTERACT (4)4
2022 Integrating Gaze and Speech for Enabling Implicit Interactions
abstract
Gaze and speech are rich contextual sources of information that, when combined, can result in effective and rich multimodal interactions. This paper proposes a machine learning-based pipeline that leverages and combines users’ natural gaze activity, the semantic knowledge from their vocal utterances and the synchronicity between gaze and speech data to facilitate users’ interaction. We evaluated our proposed approach on an existing dataset, which involved 32 participants recording voice notes while reading an academic paper. Using a Logistic Regression classifier, we demonstrate that our proposed multimodal approach maps voice notes with accurate text passages with an average F1-Score of 0.90. Our proposed pipeline motivates the design of multimodal interfaces that combines natural gaze and speech patterns to enable robust interactions.
Anam Ahmad Khan, Joshua Newn, James Bailey 0001, Eduardo Velloso
CHI1
2022 To type or to speak? The effect of input modality on text understanding during note-taking
abstract
Though recent technological advances have enabled note-taking through different modalities (e.g., keyboard, digital ink, voice), there is still a lack of understanding of the effect of the modality choice on learning. In this paper, we compared two note-taking input modalities—keyboard and voice—to study their effects on participants’ understanding of learning content. We conducted a study with 60 participants in which they were asked to take notes using voice or keyboard on two independent digital text passages while also making a judgment about their performance on an upcoming test. We built mixed-effects models to examine the effect of the note-taking modality on learners’ text comprehension, the content of notes and their meta-comprehension judgement. Our findings suggest that taking notes using voice leads to a higher conceptual understanding of the text when compared to typing the notes. We also found that using voice triggers generative processes that result in learners taking more elaborate and comprehensive notes. The findings of the study imply that note-taking tools designed for digital learning environments could incorporate voice as an input modality to promote effective note-taking and higher conceptual understanding of the text.
Anam Ahmad Khan, Sadia Nawaz, Joshua Newn, Ryan Kelly 0001, Jason M. Lodge, James Bailey 0001, Eduardo Velloso
CHI1
2021 GAVIN: Gaze-Assisted Voice-Based Implicit Note-taking
abstract
Annotation is an effective reading strategy people often undertake while interacting with digital text. It involves highlighting pieces of text and making notes about them. Annotating while reading in a desktop environment is considered trivial but, in a mobile setting where people read while hand-holding devices, the task of highlighting and typing notes on a mobile display is challenging. In this article, we introduce GAVIN, a gaze-assisted voice note-taking application, which enables readers to seamlessly take voice notes on digital documents by implicitly anchoring them to text passages. We first conducted a contextual enquiry focusing on participants’ note-taking practices on digital documents. Using these findings, we propose a method which leverages eye-tracking and machine learning techniques to annotate voice notes with reference text passages. To evaluate our approach, we recruited 32 participants performing voice note-taking. Following, we trained a classifier on the data collected to predict text passage where participants made voice notes. Lastly, we employed the classifier to built GAVIN and conducted a user study to demonstrate the feasibility of the system. This research demonstrates the feasibility of using gaze as a resource for implicit anchoring of voice notes, enabling the design of systems that allow users to record voice notes with minimal effort and high accuracy.
Anam Ahmad Khan, Joshua Newn, Ryan Kelly 0001, Namrata Srivastava, James Bailey 0001, Eduardo Velloso
ACM Trans. Comput. Hum. Interact.1