Siegfried Wahl

dblp:184/1948 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-3437-6711ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021
YearPublicationVenuePosition
2026 Gaze-informed Object Sequences for Egocentric Action Recognition using Deep Learning
abstract
We propose a gaze-informed, object-centric sequence representation pipeline for egocentric action recognition that integrates human attention signals with open-vocabulary vision-language models. This approach offers a time-saving alternative to manual AOI-labeling in dynamic scene content. Gaze fixations guide the YOLO-World model to identify objects relevant to a researcher’s application. Attended objects are encoded into temporally ordered token scanpaths representing semantic attention structure. We evaluate this approach using these textual tokens and classifying actions using a lightweight BiGRU trained on the EGTEA Gaze+ dataset. The pipeline is fully automatic with no manual intervention, emphasizing automation and accessibility. Overall, YOLO-World proves viable for object-centric sequence representation, reducing manual labeling overhead, and the full evaluation pipeline accurately recognizes the most prominent actions in the dataset. The BiGRU was able to achieve mean class accuracy of 76% for the two most labeled classes.
Nora Castner, Zhengyu Su, Siegfried Wahl
ETRA3
2025 CNN-based estimation of gaze distance in virtual reality using eye tracking and depth data
Anna-Lena von Behren, Yannick Sauer, Björn Severitt, Siegfried Wahl
ETRA4
2025 Recognition of errors in gaze-based interaction with anomaly detection
Björn Severitt, Yannick Sauer, Nora Castner, Wolfgang Fuhl, Siegfried Wahl
ETRA5
2024 Eye tracking data set of academics making an omelette: An egg-breaking work
abstract
Just as there are numerous ways to cook an egg, there are numerous ways to recreate a YouTube video of cooking an omelette. We created a dataset of 10 academics replicating a viral video of making an omelette. We evaluated the saccade behavior during the varying subtasks and found differences related to the actions (whisking, sprinkling, etc.) and the objects (eggs, butter, plate, etc.). This dataset can further offer insight into eye movements in complex tasks and is potentially even applicable for task planning and intention prediction. The available data can be found at https://zenodo.org/doi/10.5281/zenodo.10875267.
Yannick Sauer, Rajat Agarwala, Patrizia Lenhart, Regine Lendway, Björn Severitt, Alexander Neugebauer, Benedikt Hosp, Nora Castner, Siegfried Wahl
ETRA9
2024 Using mobile eye tracking for gaze- and head-contingent vision simulations
abstract
Eye tracking enables the implementation of gaze-contingent visual impairment simulations, holding promise for research, clinical uses, and demonstrations. While head-mounted displays (HMDs) are common for such simulations, screen-based implementations can be an alternative for individuals who are reluctant or unable to use HMDs. Screen-based simulations are less likely to induce virtual reality sickness and allow eye contact and social interaction with the demonstrator or examiner. To improve the immersion, we use mobile eye tracking with fiducial markers for a head- and gaze-contingent simulation of visual impairments on a screen.
Yannick Sauer, Björn Severitt, Rajat Agarwala, Siegfried Wahl
ETRA4
2024 Communication breakdown: Gaze-based prediction of system error for AI-assisted robotic arm simulated in VR
abstract
Neurological degenerative conditions can affect motor functions, making mobility daunting. Recent configurations of mobility devices that leverage artificial intelligence (AI) show its ability to handle complex information like user input. We create a virtual reality environment to measure participants’ reactions to correct and incorrect feedback from an AI-assistance system. Using gaze to evaluate these reactions, we investigate whether we can automatically predict an upcoming system error. Our results show that gaze reactions occur within 300 ms when the system highlights user input, but the delay extends to 1 second without highlighting. Subject dependent gaze behavior proved complicated for developing a generalizable model based on previous work using TCNs for online recognition of upcoming errors. Therefore, more adaptable models for individuals may be a better alternative for gaze-based accessibility systems.
Björn Severitt, Patrizia Lenhart, Benedikt Hosp, Nora Castner, Siegfried Wahl
ETRA5
2023 ZING: An Eye-Tracking Experiment Software for Organization and Presentation of Omnidirectional Stimuli in Virtual Reality
abstract
The growing field of eye-tracking enables many researchers to investigate human (subconscious) behavior unobtrusively, naturally, and non-invasive. For that, a highly natural, immersive, and controllable environment is essential to investigators. Currently, next to mobile eye-tracking in the wild, virtual reality is becoming state-of-the-art for such experiments, combining several eye-tracking modalities’ strengths. Next to simulations, omnidirectional video footage is massively used. 360°cameras capture these videos with resolutions of up to 16k. Afterward, they can be replayed on virtual reality glasses to learn about human behavior in a realistic, highly controlled environment. However, the pipeline from stitched video to eye-tracking experiment results depends on costly proprietary or self-developed software that lacks standardization, leading to recurrent reimplementation. This paper describes an open-source stimuli organization and presentation software implementation that enables researchers to easily organize their stimuli in a standardized way and conduct eye-tracking studies in virtual reality with a few clicks without knowledge about coding or technical details. The code is available at https://bitbucket.org/benediktwhosp/zing
Benedikt Hosp, Siegfried Wahl
ETRA2
2023 ZERO: A Generic Open-Source Extended Reality Eye-Tracking Controller Interface for Scientists
abstract
Virtual reality and eye-tracking technologies are nowadays standard research tools. A growing number of researchers from different disciplines are utilizing these technologies. Currently, access to eye-tracking hardware in virtual reality glasses is usually provided as APIs to call functions of the eye-tracking device. Proper implementation is device-specific and left to the user. Especially non-computer scientists are left alone with this problem, which impedes eye-tracking research in virtual reality for many scientists. This paper describes a generic open-source interface for everyone to efficiently and easily utilize common eye trackers in virtual reality. The interface is published under a friendly CC BY 4.0 license that allows for integration, modification, and extension of the code. It includes a standardized interface for several eye-tracking devices in virtual reality, is ready to be used out of the box, and allows easy addition of APIs from other manufacturers. The code is available at https://bitbucket.org/benediktwhosp/zvsl-zero
Benedikt Hosp, Siegfried Wahl
ETRA2
2023 Assessing Eye Tracking for Continuous Central Field Loss Monitoring
abstract
Eye tracking is increasingly becoming prevalent for health-related interactive systems. Eye tracking can automatically reveal the presence of Central Field Loss (CFL), a dysfunctional visual behavior requiring time-intensive medical assessments. Since CFL typically results in poor fixation stability and more frequent saccades, this work investigates the use of machine learning to estimate the likelihood of CFL based on eye-movement data. We compared random forests, support vector machines, and long-short-term memory (LSTM) neural networks for their ability to discriminate between the presence or absence of an experimentally-induced CFL. We found that the estimation accuracy increases with larger samples of eye-tracking data. However, the computational costs outweigh any increase in accuracy after classifying window sizes of 1600 msec. Here, traditional machine learning approaches outperform the LSTM neural network. We discuss implications for continuous end-user CFL monitoring and processing power to provide an outlook for gaze-based wearable health devices in human-computer interaction.
Jesse W. Grootjen, Alexandra Sipatchin, Siegfried Wahl, Tonja Machulla, Lewis L. Chuang, Thomas Kosch
MUM3
2022 Augmentation Impacts Strategy and Gaze Distribution in a Dual-Task Interleaving Scenario
abstract
When interleaving multiple tasks, people are confronted with a decision of how to distribute a finite amount of time between several tasks, which defines the task-interleaving strategy. In some challenging task interleaving scenarios where accurate timing is essential, people perform worse than they could have. With the growing advancement of technology, such as augmented reality, it became possible to impact people’s strategy and improve their performance. However, when augmenting visual input with additional visual content, the augmentation not only introduces the possible benefit but can also capture attentional resources. It is, thus, important to investigate how visual augmentation affects people’s performance in cases when otherwise people underscore in their performance. In the current study, using a psychophysics approach, it was investigated how visual augmentation impacts the task-interleaving strategy and, thus, performance in a dual-task setting with unequal task importance. In a simple dynamic 3D environment, four visual augmentations were generated aiming to prompt the user when it is more beneficial score-wise to switch from one task to another. The mean duration on one task before the task switch, as well as the resulting total performance, were evaluated in combination with the gaze direction distribution. In terms of the strategy and the total performance, all augmentations showed an advantage compared to when augmentation was not present. Furthermore, an abrupt augmentation onset based on the individual response time of the participant was more beneficial score-wise for the strategy compared to a constantly present visual augmentation. However, it affected the natural gaze direction distribution indicating the allocation of attentional resources to the augmentation. The results of this study provide an insight into potential visual augmentation designs aiming to improve user’s performance in a challenging dual-task interleaving setting.
Olga Lukashova-Sanz, Siegfried Wahl, Katharina Rifai
Int. J. Hum. Comput. Interact.2