VLDB 2026 Research / reviewers in the wild / expert
Michael Barz
dblp:176/9078
· DBLP profile ↗
20ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-6730-2466ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 15 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do You (Dis)agree With Me? Modelling Implicit User Disagreement in Human-AI Interaction Using Gaze DataabstractThe widespread use of generative AI has led to increased focus on human–AI interaction. However, AI systems can generate unexpected outputs, leading to disagreement or human–AI conflict. This paper focuses on modelling user disagreement using machine learning (ML) by observing users’ implicit viewing behaviour. We conducted a controlled study with 30 participants evaluating captions from a simulated ML image-captioning system. Participants indicated agreement or disagreement with each caption while we recorded their gaze and facial-expression data, which we used to predict (dis)agreement. We show that unimodal gaze-based personalised modelling (0.684 average balanced accuracy) outperforms generalised modelling (0.570), whereas multimodal approaches did not improve performance. Our exploratory post hoc gaze-based analysis highlights the importance of feature selection and temporal dynamics, which help guide system design and future work. We release the dataset to support reproducibility and further work. Due to the nature of this research, we also discuss the potential ethical and privacy implications of continuous passive gaze and facial monitoring. Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, Amr Gomaa, Michael Barz, Daniel Sonntag |
CHI | 4 |
| 2026 | EyeGestureLogin: Spontaneous Hands‑Free Gaze‑Based Lock Pattern Authentication for Public DisplaysabstractAs public displays become more ubiquitous, common authentication methods face security (e.g., shoulder surfing) and sanitary risks. Gaze-based systems offer a promising alternative, but their adoption is often limited by required calibration and the Midas Touch problem. We present EyeGestureLogin, a touch-free, knowledge-based method that uses gaze to enter lock patterns on a familiar 3 × 3 layout. EyeGestureLogin supports spontaneous walk-up use by replacing explicit calibration with an implicit, single-point offset estimation, and mitigates the Midas Touch using a short dwell-time trigger for fully hands-free input. The interface captures patterns by processing AOI-based fixations; we additionally evaluate an offline saccade-based detector. In a controlled study (n = 24), we achieved 91.59% accuracy (91.88% offline) with a mean entry time of 4.36 s. These results suggest that EyeGestureLogin enables fast and accurate hands-free authentication for public displays, motivating further evaluation under real-world deployment conditions. Omair Shahzad Bhatti, Abdulrahman Mohamed Selim, László Kopácsi, Maximilian Biwersi, Michael Barz, Daniel Sonntag |
ETRA | 5 |
| 2026 | Train the Spire: An ML-Driven Single Player GWAP for Image AnnotationabstractTraining image classification models requires large labelled datasets, which is particularly challenging in specialised domains where manual expert annotation remains the default, such as eye tracking. To address this challenge, we present Train the Spire, a single-player game with a purpose (GWAP) that embeds image annotation within turn-based card game mechanics for crowdsourcing data annotations. The system uses a few-shot deep learning classifier to validate player-generated labels, providing immediate feedback through rewards and penalties. The game incorporates elements, such as progression systems, a companion agent for system transparency, and balanced difficulty, to maintain player engagement while ensuring annotation accuracy. In this paper, we present the system design, implementation details, and evaluation study design for comparing Train the Spire against a baseline annotation tool using the VISUS mobile eye-tracking dataset in an online user study measuring effectiveness, usability, and enjoyment. Keno Nanninga, Abdulrahman Mohamed Selim, Sara-Jane Bittner, Pascal Lessel, Michael Barz, Daniel Sonntag |
ETRA | 5 |
| 2026 | OpenGazeLab: An Interactive Toolkit for Gaze Analysis and Event DetectionabstractEvent detection (e.g., fixations and saccades) is a prerequisite for many eye-tracking analyses. However, existing solutions often require coding experience, rely on closed-source vendor tools, or focus on narrow paradigms (e.g., reading). Additionally, most are designed for stationary eye tracking, whereas head-mounted recordings are affected by head/scene motion, making event detection more difficult. Therefore, we present OpenGazeLab, a publicly available, browser-based toolkit that unifies event extraction, parameter configuration, and data inspection for both stationary and head-mounted eye tracking. OpenGazeLab implements I-DT and I-VT, and extends them for head-mounted data using scene-motion compensation and adaptive thresholds. The toolkit also provides a timeline-based visualisation that overlays gaze and detected events on the stimulus image or scene video, enabling quick visual verification. OpenGazeLab is implemented using widely used, well-maintained Python frameworks to support reproducible, easy-to-adopt workflows, and we plan to extend it with additional event classes (e.g., smooth pursuit) and alternative detectors. Khue Minh Pham, Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, László Kopácsi, Michael Barz, Daniel Sonntag |
ETRA | 5 |
| 2025 | Eliciting Multimodal Approaches for Machine Learning-Assisted Photobook Creation
Sara-Jane Bittner, Michael Barz, Daniel Sonntag |
INTERACT (2) | 2 |
| 2025 | Gaze-Based Menu Navigation in Virtual Reality: A Comparative Study of Layouts and Interaction TechniquesabstractAbstract Integrating eye-tracking technologies in Extended Reality (XR) headsets has enabled intuitive, hands-free system interaction, such as gaze-based menu navigation. However, there is a lack of comprehensive comparisons and consensus in the literature on the optimal use of gaze-based menu navigation. This paper presents a comparative analysis of gaze-based menu navigation in virtual environments, focusing on two common menu layouts: pie and list menus, with three interaction methods: gaze-based dwell, controller-based, and a multimodal approach combining gaze and controller inputs. We conducted a 19-participant within-subject study, measuring task completion time, error rate, usability, and user preference for each condition. The results indicate that while the pie layout was statistically faster and less erroneous than the list layout, novice users tend to favour list layouts. Furthermore, we found that users preferred the multimodal interaction method, despite its lower task completion times and higher error rates compared to controller-based navigation. Based on our findings, we offer design guidelines and recommendations for implementing gaze-based menu systems. László Kopácsi, Albert Klimenko, Abdulrahman Mohamed Selim, Michael Barz, Daniel Sonntag |
INTERACT (1) | 4 |
| 2025 | How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?abstractRecent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce \textbf{GitMerge3D}, a \textbf{g}lobally \textbf{i}nformed graph \textbf{t}oken \textbf{merging} method that can reduce the token count by up to 90–95\% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at \href{https://gitmerge3d.github.io/}{https://gitmerge3d.github.io}. Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D. Doan, Roger Wattenhofer, Ngo Anh Vien, Mathias Niepert, Daniel Sonntag, Paul Swoboda |
NeurIPS | 4 |
| 2024 | Speech Imagery BCI Training Using Game with a PurposeabstractGames are used in multiple fields of brain–computer interface (BCI) research and applications to improve participants’ engagement and enjoyment during electroencephalogram (EEG) data collection. However, despite potential benefits, no current studies have reported on implemented games for Speech Imagery BCI. Imagined speech is speech produced without audible sounds or active movement of the articulatory muscles. Collecting imagined speech EEG data is a time-consuming, mentally exhausting, and cumbersome process, which requires participants to read words off a computer screen and produce them as imagined speech. To improve this process for study participants, we implemented a maze-like game where a participant navigated a virtual robot capable of performing five actions that represented our words of interest while we recorded their EEG data. The study setup was evaluated with 15 participants. Based on their feedback, the game improved their engagement and enjoyment while resulting in a 69.10% average classification accuracy using a random forest classifier. Abdulrahman Mohamed Selim, Maurice Rekrut, Michael Barz, Daniel Sonntag |
AVI | 3 |
| 2024 | HumanEYEze 2024: Workshop on Eye Tracking for Multimodal Human-Centric ComputingabstractThe HumanEYEze 2024 workshop aims to explore the role of eye tracking in developing human-centered multimodal AI systems. Over the past two decades, eye tracking has evolved from a diagnostic tool to an important input modality for real-time interactive systems, driven by advancements in hardware that have improved its affordability, availability, and performance. Initially used in specialized applications, eye tracking now significantly impacts research on gaze-based multimodal interaction. Recently, eye-based user and context modeling has emerged, utilizing eye movements to provide rich insights into user behavior and interaction contexts. The workshop aims to bring together researchers from eye tracking, multimodal human-computer interaction, and AI. It aims to enhance understanding of integrating eye tracking into multimodal human-centered computing. The expected outcomes include fostering collaborations and promoting knowledge exchange. Michael Barz, Roman Bednarik, Andreas Bulling, Cristina Conati, Daniel Sonntag |
ICMI | 1 |
| 2024 | Perceived Text Relevance Estimation Using Scanpaths and GNNsabstractA scanpath is an important concept in eye tracking that represents a person’s eye movements in a graph-like structure. Passive gaze-based interfaces, in which users do not consciously interact using their eyes, typically interpret users’ scanpaths to enable adaptive and personalised interaction. Despite the benefits of graph neural networks (GNNs) in graph processing, this technology has not been considered for that purpose. An example application is perceived relevance estimation, which still suffers from low classification performance. In this work, we investigate how and whether GNNs can be used to analyse scanpaths for readers’ perceived relevance estimation using the gazeRE dataset. This dataset contains eye tracking data from 24 participants, who rated the relevance of 12 short and 12 long documents in relation to a given query. The relevance was assigned either to an entire short document or to each paragraph within a long document, which allowed us to investigate two different GNN tasks. For comparison, we reproduced the gazeRE baseline using Random Forest and Support Vector classifiers, and an additional Convolutional Neural Network (CNN) from the literature. All models were evaluated using leave-users-out cross-validation. For short documents, the GNNs surpassed the baseline methods, with certain experiments showing an absolute balanced accuracy improvement of 7.6% and 14.3% over the CNN and gazeRE baselines, respectively. However, similar improvements were not observed in long documents. This work investigates and discusses the future potential of using GNNs as a scanpath analysis method for passive gaze-based applications, such as implicit relevance estimation. Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, Michael Barz, Daniel Sonntag |
ICMI | 3 |
| 2024 | The MASTER XR Platform for Robotics Training in ManufacturingabstractThe MASTER project introduces an open Extended Reality (XR) platform designed to enhance human-robot collaboration and train workers in robotics within manufacturing settings. It includes modules for creating safe workspaces, intuitive robot programming, and user-friendly human-robot interactions (HRI), including eye-tracking technologies. The development of the platform is supported by two open calls targeting technical SMEs and educational institutes to enhance and test its functionalities. By employing the learning-by-doing methodology and integrating effective teaching principles, the MASTER platform aims to provide a comprehensive learning environment, preparing students and professionals for the complexities of flexible and collaborative manufacturing settings. László Kopácsi, Panagiotis Karagiannis, Sotiris Makris, Johan Kildal, Andoni Rivera-Pinto, Judit Ruiz de Munain, Jesús Rosel, Maria Madarieta, Nikolaos Tseregkounis, Konstantina Salagianni, Panagiotis Aivaliotis, Michael Barz, Daniel Sonntag |
VRST | 12 |
| 2024 | GazeLock: Gaze- and Lock Pattern-Based AuthenticationabstractPassword entry is common authentication approach in Extended Reality (XR) applications for its simplicity and familiarity, but it faces challenges in public and dynamic environments due to its cumbersome nature and susceptibility to observation attacks. Manual password input can be disruptive and prone to theft through shoulder surfing or surveillance. While alternative knowledge-based approaches exist, they often require complex physical gestures and are impractical for frequent public use. We present GazeLock, an eye-tracking and lock pattern-based authentication method. This method aims to provide an easy-to-learn and efficient alternative by leveraging familiar lock patterns operated through gaze. It ensures resilience to external observation, as physical interaction is unnecessary and eyes are obscured by the headset. Its hands-free, discreet nature makes it suitable for secure public use. We demonstrate this method by simulating the unlocking of a smart lock via an XR headset, showcasing its potential applications and benefits in real-world scenarios. László Kopácsi, Tobias Sebastian Schneider, Chiara Karr, Michael Barz, Daniel Sonntag |
VRST | 4 |
| 2022 | Interactive Assessment Tool for Gaze-based Machine Learning Models in Information RetrievalabstractEye movements were shown to be an effective source of implicit relevance feedback in information retrieval tasks. They can be used to, e.g., estimate the relevance of read documents and expand search queries using machine learning. In this paper, we present the Reading Model Assessment tool (ReMA), an interactive tool for assessing gaze-based relevance estimation models. Our tool allows experimenters to easily browse recorded trials, compare the model output to a ground truth, and visualize gaze-based features at the token- and paragraph-level that serve as model input. Our goal is to facilitate the understanding of the relation between eye movements and the human relevance estimation process, to understand the strengths and weaknesses of a model at hand, and, eventually, to enable researchers to build more effective models. Pablo Valdunciel, Omair Shahzad Bhatti, Michael Barz, Daniel Sonntag |
CHIIR | 3 |
| 2021 | Explainable Automatic Evaluation of the Trail Making Test for Dementia ScreeningabstractThe Trail Making Test (TMT) is a frequently used neuropsychological test for assessing cognitive performance. The subject connects a sequence of numbered nodes by using a pen on normal paper. We present an automatic cognitive assessment tool that analyzes samples of the TMT which we record using a digital pen. This enables us to analyze digital pen features that are difficult or impossible to evaluate manually. Our system automatically measures several pen features, including the completion time which is the main performance indicator used by clinicians to score the TMT in practice. In addition, our system provides a structured report of the analysis of the test, for example indicating missed or erroneously connected nodes, thereby offering more objective, transparent and explainable results to the clinician. We evaluate our system with 40 elderly subjects from a geriatrics daycare clinic of a large hospital. Alexander Prange, Michael Barz, Anika Heimann-Steinert, Daniel Sonntag |
CHI | 2 |
| 2020 | Visual Search Target Inference in Natural Interaction Settings with Machine LearningabstractVisual search is a perceptual task in which humans aim at identifying a search target object such as a traffic sign among other objects. Search target inference subsumes computational methods for predicting this target by tracking and analyzing overt behavioral cues of that person, e.g., the human gaze and fixated visual stimuli. We present a generic approach to inferring search targets in natural scenes by predicting the class of the surrounding image segment. Our method encodes visual search sequences as histograms of fixated segment classes determined by SegNet, a deep learning image segmentation model for natural scenes. We compare our sequence encoding and model training (SVM) to a recent baseline from the literature for predicting the target segment. Also, we use a new search target inference dataset. The results show that, first, our new segmentation-based sequence encoding outperforms the method from the literature, and second, that it enables target inference in natural settings. Michael Barz, Sven Stauden, Daniel Sonntag |
ETRA | 1 |
| 2020 | Digital Pen Features Predict Task Difficulty and User Performance of Cognitive TestsabstractDigital pen signals were shown to be predictive for cognitive states, cognitive load and emotion in educational settings. We investigate whether low-level pen-based features can predict the difficulty of tasks in a cognitive test and the learner's performance in these tasks, which is inherently related to cognitive load, without a semantic content analysis. We record data for tasks of varying difficulty in a controlled study with children from elementary school. We include two versions of the Trail Making Test (TMT) and six drawing patterns from the Snijders-Oomen Non-verbal intelligence test (SON) as tasks that feature increasing levels of difficulty. We examine how accurately we can predict the task difficulty and the user performance as a measure for cognitive load using support vector machines and gradient boosted decision trees with different feature selection strategies. The results show that our correlation-based feature selection is beneficial for model training, in particular when samples from TMT and SON are concatenated for joint modelling of difficulty and time. Our findings open up opportunities for technology-enhanced adaptive learning. Michael Barz, Kristin Altmeyer, Sarah Malone, Luisa Lauer, Daniel Sonntag |
UMAP | 1 |
| 2018 | Error-aware gaze-based interfaces for robust mobile gaze interactionabstractGaze estimation error can severely hamper usability and performance of mobile gaze-based interfaces given that the error varies constantly for different interaction positions. In this work, we explore error-aware gaze-based interfaces that estimate and adapt to gaze estimation error on-the-fly. We implement a sample error-aware user interface for gaze-based selection and different error compensation methods: a naïve approach that increases component size directly proportional to the absolute error, a recent model by Feit et al. that is based on the two-dimensional error distribution, and a novel predictive model that shifts gaze by a directional error estimate. We evaluate these models in a 12-participant user study and show that our predictive model significantly outperforms the others in terms of selection rate, particularly for small gaze targets. These results underline both the feasibility and potential of next generation error-aware gaze-based user interfaces. Michael Barz, Florian Daiber, Daniel Sonntag, Andreas Bulling |
ETRA | 1 |
| 2017 | Speech-based Medical Decision Support in VR using a Deep Neural Network (Demonstration)abstractWe present a speech dialogue system that facilitates medical decision support for doctors in a virtual reality (VR) application. The therapy prediction is based on a recurrent neural network model that incorporates the examination history of patients. A central supervised patient database provides input to our predictive model and allows us, first, to add new examination reports by a pen-based mobile application on-the-fly, and second, to get therapy prediction results in real-time. This demo includes a visualisation of patient records, radiology image data, and the therapy prediction results in VR. Alexander Prange, Michael Barz, Daniel Sonntag |
IJCAI | 2 |
| 2017 | A Multimodal Dialogue System for Medical Decision Support inside Virtual RealityabstractWe present a multimodal dialogue system that allows doctors to interact with a medical decision support system in virtual reality (VR).We integrate an interactive visualization of patient records and radiology image data, as well as therapy predictions.Therapy predictions are computed in realtime using a deep learning model. Alexander Prange, Margarita Chikobava, Peter Poller, Michael Barz, Daniel Sonntag |
SIGDIAL Conference | 4 |
| 2016 | Prediction of gaze estimation error for error-aware gaze-based interfacesabstractGaze estimation error is inherent in head-mounted eye trackers and seriously impacts performance, usability, and user experience of gaze-based interfaces. Particularly in mobile settings, this error varies constantly as users move in front and look at different parts of a display. We envision a new class of gaze-based interfaces that are aware of the gaze estimation error and adapt to it in real time. As a first step towards this vision we introduce an error model that is able to predict the gaze estimation error. Our method covers major building blocks of mobile gaze estimation, specifically mapping of pupil positions to scene camera coordinates, marker-based display detection, and mapping of gaze from scene camera to on-screen coordinates. We develop our model through a series of principled measurements of a state-of-the-art head-mounted eye tracker. Michael Barz, Florian Daiber, Andreas Bulling |
ETRA | 1 |