VLDB 2026 Research / reviewers in the wild / expert
Benedikt Hosp
dblp:220/6822 · also Benedikt W. Hosp, Benedikt Werner Hosp
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-8259-5463ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FOVAL: Subject-Adaptive, Calibration-Free Fixation Depth EstimationabstractAccurate fixation depth estimation is essential for applications in extended reality (XR), robotics, and human-computer interaction. However, current methods depend heavily on user-specific calibration, limiting their scalability and usability. We introduce FOVAL, a robust calibration-free approach that combines spatiotemporal sequence modelling via Long Short-Term Memory (LSTM) networks with subject-adaptive feature engineering and normalisation. Compared to Transformers, Temporal Convolutional Networks (TCNs), and CNNs, FOVAL achieves superior performance, particularly in scenarios with limited and noisy gaze data. Evaluations across three benchmark datasets using Leave-One-Out Cross-Validation (LOOCV) and cross-dataset validation show a mean absolute error (MAE) of 9.1 cm and strong generalisation without calibration. We further analyse inter-subject variability and domain shifts, providing insight into model robustness and adaptation. FOVAL’s scalability and accuracy make it highly suitable for real-world deployment. Benedikt Hosp |
ETRA | 1 |
| 2024 | Eye tracking data set of academics making an omelette: An egg-breaking workabstractJust as there are numerous ways to cook an egg, there are numerous ways to recreate a YouTube video of cooking an omelette. We created a dataset of 10 academics replicating a viral video of making an omelette. We evaluated the saccade behavior during the varying subtasks and found differences related to the actions (whisking, sprinkling, etc.) and the objects (eggs, butter, plate, etc.). This dataset can further offer insight into eye movements in complex tasks and is potentially even applicable for task planning and intention prediction. The available data can be found at https://zenodo.org/doi/10.5281/zenodo.10875267. Yannick Sauer, Rajat Agarwala, Patrizia Lenhart, Regine Lendway, Björn Severitt, Alexander Neugebauer, Benedikt Hosp, Nora Castner, Siegfried Wahl |
ETRA | 7 |
| 2024 | Communication breakdown: Gaze-based prediction of system error for AI-assisted robotic arm simulated in VRabstractNeurological degenerative conditions can affect motor functions, making mobility daunting. Recent configurations of mobility devices that leverage artificial intelligence (AI) show its ability to handle complex information like user input. We create a virtual reality environment to measure participants’ reactions to correct and incorrect feedback from an AI-assistance system. Using gaze to evaluate these reactions, we investigate whether we can automatically predict an upcoming system error. Our results show that gaze reactions occur within 300 ms when the system highlights user input, but the delay extends to 1 second without highlighting. Subject dependent gaze behavior proved complicated for developing a generalizable model based on previous work using TCNs for online recognition of upcoming errors. Therefore, more adaptable models for individuals may be a better alternative for gaze-based accessibility systems. Björn Severitt, Patrizia Lenhart, Benedikt Hosp, Nora Castner, Siegfried Wahl |
ETRA | 3 |
| 2024 | Reflecting on Excellence: VR Simulation for Learning Indirect Vision in Complex Bi-Manual TasksabstractIndirect vision through a mirror, while bi-manually manipulating both the mirror and another tool is a relatively common way to perform operations in various types of surgery. However, learning such psychomotor skills requires extensive training; they are difficult to teach; and they can be quite costly, for instance, for dentistry schools. In order to study the effectiveness of VR simulators for learning these kinds of skills, we developed a simulator for training dental surgery procedures, which supports tracking of eye gaze and tool trajectories (mirror and drill), as well as automated outcome scoring. We carried out a pre-/post-test study in which 30 fifth-year dental students received six training sessions in the access opening stage of the root canal procedure using the simulator. In addition, six experts performed three trials using the simulator. The outcomes of drilling performed on realistic plastic teeth showed a significant learning effect due to the training sessions. Also, students with larger improvements in the simulator tended to improve more in the real-world tests. Analysis of the tracking data revealed novel relationships between several metrics w.r.t. eye gaze and mirror use, and performance and learning effectiveness: high rates of correct mirror placement during active drilling and high continuity of fixation on the tooth are associated with increased skills and increased learning effectiveness. Larger time allocation for tooth inspections using the mirror, i.e., indirect vision, and frequency of inspection are associated with increased learning effectiveness. Our findings suggest that eye tracking can provide valuable insights into student learning gains of bi-manual psychomotor skills, particularly in indirect vision environments. Maximilian Kaluschke, René Weller, Myat Su Yin, Benedikt Hosp, Farin Kulapichitr, Siriwan Suebnukarn, Peter Haddawy, Gabriel Zachmann |
VR | 4 |
| 2023 | ZING: An Eye-Tracking Experiment Software for Organization and Presentation of Omnidirectional Stimuli in Virtual RealityabstractThe growing field of eye-tracking enables many researchers to investigate human (subconscious) behavior unobtrusively, naturally, and non-invasive. For that, a highly natural, immersive, and controllable environment is essential to investigators. Currently, next to mobile eye-tracking in the wild, virtual reality is becoming state-of-the-art for such experiments, combining several eye-tracking modalities’ strengths. Next to simulations, omnidirectional video footage is massively used. 360°cameras capture these videos with resolutions of up to 16k. Afterward, they can be replayed on virtual reality glasses to learn about human behavior in a realistic, highly controlled environment. However, the pipeline from stitched video to eye-tracking experiment results depends on costly proprietary or self-developed software that lacks standardization, leading to recurrent reimplementation. This paper describes an open-source stimuli organization and presentation software implementation that enables researchers to easily organize their stimuli in a standardized way and conduct eye-tracking studies in virtual reality with a few clicks without knowledge about coding or technical details. The code is available at https://bitbucket.org/benediktwhosp/zing Benedikt Hosp, Siegfried Wahl |
ETRA | 1 |
| 2023 | ZERO: A Generic Open-Source Extended Reality Eye-Tracking Controller Interface for ScientistsabstractVirtual reality and eye-tracking technologies are nowadays standard research tools. A growing number of researchers from different disciplines are utilizing these technologies. Currently, access to eye-tracking hardware in virtual reality glasses is usually provided as APIs to call functions of the eye-tracking device. Proper implementation is device-specific and left to the user. Especially non-computer scientists are left alone with this problem, which impedes eye-tracking research in virtual reality for many scientists. This paper describes a generic open-source interface for everyone to efficiently and easily utilize common eye trackers in virtual reality. The interface is published under a friendly CC BY 4.0 license that allows for integration, modification, and extension of the code. It includes a standardized interface for several eye-tracking devices in virtual reality, is ready to be used out of the box, and allows easy addition of APIs from other manufacturers. The code is available at https://bitbucket.org/benediktwhosp/zvsl-zero Benedikt Hosp, Siegfried Wahl |
ETRA | 1 |
| 2021 | States of Confusion: Eye and Head Tracking Reveal Surgeons' Confusion during Arthroscopic SurgeryabstractDuring arthroscopic surgeries, surgeons are faced with challenges like cognitive re-projection of the 2D screen output into the 3D operating site or navigation through highly similar tissue. Training of these cognitive processes takes much time and effort for young surgeons, but is necessary and crucial for their education. In this study we want to show how to recognize states of confusion of young surgeons during an arthroscopic surgery, by looking at their eye and head movements and feeding them to a machine learning model. With an accuracy of over 94% and detection speed of 0.039 seconds, our model is a step towards online diagnostic and training systems for the perceptual-cognitive processes of surgeons during arthroscopic surgeries. Benedikt Hosp, Myat Su Yin, Peter Haddawy, Ratthaphum Watcharopas, Paphon Sa-Ngasoongsong, Enkelejda Kasneci |
ICMI | 1 |
| 2019 | Encodji: encoding gaze data into emoji space for an amusing scanpath classification approach ;)abstractTo this day, a variety of information has been obtained from human eye movements, which holds an imense potential to understand and classify cognitive processes and states - e.g., through scanpath classification. In this work, we explore the task of scanpath classification through a combination of unsupervised feature learning and convolutional neural networks. As an amusement factor, we use an Emoji space representation as feature space. This representation is achieved by training generative adversarial networks (GANs) for unpaired scanpath-to-Emoji translation with a cyclic loss. The resulting Emojis are then used to train a convolutional neural network for stimulus prediciton, showing an accuracy improvement of more than five percentual points compared to the same network trained using solely the scanpath data. As a side effect, we also obtain novel unique Emojis representing each unique scanpath. Our goal is to demonstrate the applicability and potential of unsupervised feature learning to scanpath classification in a humorous and entertaining way. Wolfgang Fuhl, Efe Bozkir, Benedikt Hosp, Nora Castner, David Geisler, Thiago Santini, Enkelejda Kasneci |
ETRA | 3 |
| 2018 | BORE: boosted-oriented edge optimization for robust, real time remote pupil center detectionabstractUndoubtedly, eye movements contain an immense amount of information, especially when looking to fast eye movements, namely time to the fixation, saccade, and micro-saccade events. While, modern cameras support recording of few thousand frames per second, to date, the majority of studies use eye trackers with the frame rates of about 120 Hz for head-mounted and 250 Hz for remote-based trackers. In this study, we aim to overcome the challenge of the pupil tracking algorithms to perform real time with high speed cameras for remote eye tracking applications. We propose an iterative pupil center detection algorithm formulated as an optimization problem. We evaluated our algorithm on more than 13,000 eye images, in which it outperforms earlier solutions both with regard to runtime and detection accuracy. Moreover, our system is capable of boosting its runtime in an unsupervised manner, thus we remove the need for manual annotation of pupil images. Wolfgang Fuhl, Shahram Eivazi, Benedikt Hosp, Anna Eivazi, Wolfgang Rosenstiel, Enkelejda Kasneci |
ETRA | 3 |