EDBT 2026 Demo / reviewers in the wild / expert
Enkelejda Kasneci
dblp:08/1610 · also Enkelejda Tafaj
· DBLP profile ↗
134ranked-venue papers
4as first author
91since 2021 · last 2026
0000-0003-3146-4484ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 85 · 3 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 70 · 2 first-author · 40 since 2021Artificial intelligence and machine learning · 36 · 2 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 11 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PromptMirror: Visualizing LLM Use to Support STEM Student ReflectionabstractLarge language models (LLMs) are increasingly embedded in students’ academic work, yet the increasing reliance can undermine learning depth and raise integrity concerns. While reflection has long been studied in HCI to foster awareness and behavior change, little is known about how to support students in reflecting on everyday LLM use. We present PromptMirror, a student-facing dashboard that processes LLM conversation logs and visualizes four perspectives, temporal, sentiment, intent, and thematic, to encourage reflection. We informed the design of PromptMirror with two focus groups (one expert and one student with four participants each) and subsequently conducted an online think-aloud with 20 university students who uploaded their own LLM use data. Findings provide preliminary evidence that PromptMirror may support students in recognizing their LLM use estimation gap and engaging in deeper reflection on LLM reliance. Our contributions are twofold: (1) a student-centric reflection system; (2) empirical insights into reflective analytics for everyday LLM tools. Ka Hei Carrie Lau, Nada Terzimehic, Enkelejda Kasneci |
DIS | 3 |
| 2026 | Democratizing Writing Support with AI: Insights from One Year of Real-World Interactions with an Open-Access Writing Feedback ToolabstractWriting is a foundational skill for educational, professional, and civic participation, yet access to frequent and timely writing feedback remains deeply unequal. Teachers face significant workload constraints, particularly in large classes, and many learners lack alternative sources of individualized feedback. While large language models (LLMs) offer the opportunity for scalable, adaptive support, little is known about how students engage with such feedback tools in real-world, self-directed settings. We present a large-scale, year-long analysis of 23,650 voluntary interactions with an open-access AI writing feedback system used by students across diverse educational contexts and age groups, conducted in accordance with strict data protection standards. Using a clustering approach, we identify 2,800 iterative revision chains and apply a validated LLM-based multidimensional scoring framework to assess text quality over time. Our findings reveal that students who revised their texts after receiving AI feedback demonstrated statistically significant, albeit modest, improvements across both content and language-related dimensions (overall writing quality: ∆ = 0.067, p < .001, r = .17), with the greatest gains observed among initially low-performing writers. Revision frequency was positively associated with improvement, particularly in higher-order writing skills. However, engagement was uneven, with higher usage among students in academically oriented schools. These results demonstrate both the technical feasibility and social potential of deploying generative AI for educational support at scale, while highlighting the need for inclusive infrastructure, accessible design, and targeted outreach to truly democratize educational benefits. Babette Bühler, Ivo Bueno, Enkelejda Kasneci |
AAAI | 3 |
| 2026 | Should AI Ask First? Investigating the Effects of Proactive vs Reactive AI Mentoring in Self-directed Learning
Khaoula Otmani, Anna Bodonhelyi, Babette Bühler, Enkelejda Kasneci |
AIED (3) | 4 |
| 2026 | Skin-Deep Bias: How Avatar Appearances Shape Perceptions of AI HiringabstractArtificial intelligence is increasingly used in hiring, raising concerns about how applicants perceive these systems. While prior work on algorithmic fairness has emphasized technical bias mitigation, little is known about how avatar identity cues influence applicants' justice attributions in an interview context. We conducted a crowdsourcing study with 215 participants who completed an interview with photorealistic AI avatars varied in phenotypic traits (race and sex), followed by a standardized rejection. Using self-reports, sentiment analysis, and eye tracking, we measured perceptions of trust, fairness, and bias. Results show that racial mismatch heightened perceptions of ethnic bias, while partial match (sharing only one identity) reduced fairness judgments compared to both full and no match. This work extends the Computers-Are-Social-Actors paradigm by demonstrating that avatar appearances shape justice-related evaluations of AI. We contribute to HCI by revealing how identity cues influence fairness attributions and offer actionable insights for designing equitable AI interview systems. Ka Hei Carrie Lau, Philipp Stark, Efe Bozkir, Enkelejda Kasneci |
CHI | 4 |
| 2026 | Susceptibility to High-Fidelity Misinformation: An Eye-Tracking AnalysisabstractWith the rise of online misinformation and AI-generated text, understanding human perception of news truthfulness is critical. In this study, we examine visual attention and cognitive processing using eye-tracking measures as individuals read fake and real news articles sharing nearly identical structure and imagery, differing only in subtle textual changes. Using the public FakeNewsPerception dataset, we analyze advanced gaze measures, including scanpaths, AOI transitions, and luminance-corrected pupil measures, beyond basic gaze features, in relation to news truthfulness and perceived believability. Results show that, given the high fidelity of the fake news, readers exhibited comparable visual scanning patterns, attention allocation across AOIs, and cognitive load regardless of truthfulness and perceived believability of news. Interestingly, when participants believed a news item, they demonstrated lower focal attention than when uncertain or disbelieving, suggesting that prior belief reduces visual inspection. These findings, supported by behavioral analysis, may help explain human susceptibility to high-fidelity misinformation. Yasasi Abeysinghe, Gavindya Jayawardena, Enkelejda Kasneci, Sampath Jayarathna |
ETRA | 3 |
| 2026 | Night Eyes: A Reproducible Framework for Constellation-Based Corneal Reflection MatchingabstractCorneal reflection (glint) detection plays an important role in pupil-corneal reflection (P-CR) eye tracking, but in practice it is often handled as heuristics embedded within larger systems, making reproducibility difficult across hardware setups. We introduce a 2D geometry-driven, constellation-based pipeline for mulit-glint detection and matching, focusing on reproducibility and clear evaluation. Inspired by lost-in-space star identification, we treat glints as structured constellations rather than independent blobs. We propose a Similarity-Layout Alignment (SLA) procedure which adapts constellation matching to the specific constraints of multi-LED eye tracking. The framework brings together controlled over-detection, adaptive candidate fallback, appearance-aware scoring, and optional semantic layout priors while keeping detection and correspondence explicitly separated. Evaluated on a public multi-LED dataset, the system provides stable identity-preserving correspondence under noisy conditions. We release code, presets, and evaluation scripts to enable transparent replication, comparison, and dataset annotation. Virmarie Maquiling, Yasmeen Abdrabou, Enkelejda Kasneci |
ETRA | 3 |
| 2026 | As Far as Eye See: Vergence-Pupil Coupling in Near-Far Depth SwitchingabstractVergence is widely used as a proxy for depth perception and spatial attention in immersive and real-world eye-tracking studies. In this paper, we investigate how pupil size artefacts affect vergence estimates during real physical depth viewing with a head-mounted eye tracker. Using a beamsplitter setup with physically near and far targets, we elicited controlled convergent and divergent eye movements under static, luminance-modulated, and blockwise fixation conditions. Near and far targets were reliably separable in vergence angle across participants. However, pupil-vergence coupling varied substantially across individuals and conditions. Static illumination produced large inter-participant variability, while luminance modulation reduced this spread, yielding more clustered estimates. Blockwise and audio-cued recordings further showed that pupil-vergence coupling persists even without visual depth onsets. These results suggest that pupil size fluctuations can systematically influence vergence estimates, and that controlled viewing conditions can reduce—but not eliminate—this effect. Virmarie Maquiling, Yasmeen Abdrabou, Enkelejda Kasneci |
ETRA | 3 |
| 2026 | VIVA Stimuli: A Web-Based Platform for Eye Tracking StimuliabstractReproducibility in eye-tracking research is increasingly important as researchers conduct diverse experiments and seek to validate or replicate findings. However, exact replication remains challenging due to differences in laboratory practices and experimental setups. Inconsistent stimulus presentation can yield divergent metrics from identical oculomotor behavior, yet the stimulus layer remains largely unstandardized. Existing tools often require programming expertise or depend on specific hardware vendors. We introduce VIVA Stimuli, a web-based platform for standardized eye-tracking stimulus presentation. It provides configurable task types, including fixation, smooth pursuit, cognitive load, blink, slippage, content display, and questionnaires within a unified environment. The platform supports any eye-tracking technology, including wearable and screen-based VOG trackers, LFI sensors, and EOG devices. ArUco markers enable synchronization for trackers with scene cameras, while a WebSocket architecture ensures temporal synchronization for those without. A visual experiment flow editor allows protocols to be exported and shared, enabling identical stimulus replication across laboratories. Süleyman Özdel, Virmarie Maquiling, Kadir Burak Buldu, Yasmeen Abdrabou, Enkelejda Kasneci |
ETRA | 5 |
| 2026 | Understanding Password Preferences, Memorability, and Security through a Human-Centered LensabstractPasswords remain the primary authentication method, yet user-created passwords are often the weakest due to the security–usability trade-off. Although AI-based password generators are emerging, little is known about their effectiveness and user perceptions. This eye-tracking study examined how behavior during password creation, selection, and memorization relates to objective and subjective password quality. Four password models, three AI-based (DeepSeek-API, ChatGPT-API, PassGPT) and one rule-based random generator, generated suggestions from participants’ self-generated passwords across four website contexts. Eye movements were recorded throughout the experiment. Results confirm the expected trade-off between AI-generated password strength and human memorability but also reveal a novel behavioral link. Despite stronger AI-generated passwords, participants favored self-generated ones. Notably, visual attention to contextual cues was significantly correlated with higher password entropy. This suggests that security is shaped not only by the generation tool but also by users’ visual engagement with contextual cues, highlighting the potential of attention-driven security design. Duru Paker, Süleyman Özdel, Enkelejda Kasneci |
ETRA | 3 |
| 2026 | An Affordable, Wearable Stereo-Eye-Tracking PlatformabstractResearch on video-based eye-tracking has long explored stereo and glint-based methods, yet existing wearable eye trackers — both commercial and open-source — offer limited flexibility for algorithm development and comparative evaluation. We present an affordable, wearable stereo eye-tracking platform built from off-the-shelf and 3D-printable components that explicitly targets this gap. The system combines four infrared eye cameras, infrared illumination, an optional scene camera, and software support for calibration and synchronized data acquisition. By design, the platform supports multiple eye-tracking paradigms, including stereo, glint-based, and binocular approaches, within a single hardware configuration. Rather than optimizing for end-user robustness, the platform prioritizes modularity and extensibility for research use. This paper focuses on the hardware architecture and calibration pipeline and demonstrates the feasibility of the approach using a prototype implementation. All hardware designs and documentation are made openly available Alexander Zimmer, Yasmeen Abdrabou, Enkelejda Kasneci |
ETRA | 3 |
| 2026 | Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom DataabstractObservation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of classroom recordings, focusing on multimodal instructional activity and discourse recognition as a foundation for actionable feedback. Using a densely annotated dataset of 164 hours of video and 68 lesson transcripts, we design parallel, modality-specific pipelines. For video, we evaluate zero-shot multimodal LLMs, fine-tuned vision–language models, and self-supervised video transformers on 24 activity labels. For transcripts, we fine-tune a transformer-based classifier with contextualized inputs and compare it against prompting-based LLMs on 19 discourse labels. To handle class imbalance and multi-label complexity, we apply per-label thresholding, context windows, and imbalance-aware loss functions. The results show that fine-tuned models consistently outperform prompting-based approaches, achieving macro-F1 scores of 0.577 for video and 0.460 for transcripts. These results demonstrate the feasibility of automated classroom analysis and establish a foundation for scalable teacher feedback systems. Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan K. Foster, Peter A. Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
WACV | 10 |
| 2026 | CycleSL: Server-Client Cyclical Update Driven Scalable Split LearningabstractSplit learning emerges as a promising paradigm for collaborative distributed model training, akin to federated learning, by partitioning neural networks between clients and a server without raw data exchange. However, sequential split learning suffers from poor scalability, while parallel variants like parallel split learning and split federated learning often incur high server resource overhead due to model duplication and aggregation, and generally exhibit reduced model performance and convergence owing to factors like client drift and lag. To address these limitations, we introduce CycleSL, a novel aggregation-free split learning framework that enhances scalability and performance and can be seamlessly integrated with existing methods. Inspired by alternating block coordinate descent, CycleSL treats server-side training as an independent higher-level machine learning task, resampling client-extracted features (smashed data) to mitigate heterogeneity and drift. It then performs cyclical updates, namely optimizing the server model first, followed by client updates using the updated server for gradient computation. We integrate CycleSL into previous algorithms and benchmark them on five publicly available datasets with non-iid data distribution and partial client attendance. Our empirical findings highlight the effectiveness of CycleSL in enhancing model performance. Our source code is available at https://gitlab.lrz.de/hctl/CycleSL. Mengdi Wang 0002, Efe Bozkir, Enkelejda Kasneci |
WACV | 3 |
| 2026 | PerVRML: ChatGPT-Driven Personalized VR Environments for Machine Learning EducationabstractThe advent of large language models (LLMs) such as ChatGPT has demonstrated significant potential for advancing educational technologies. Recently, growing interest has emerged in integrating ChatGPT with virtual reality (VR) to provide interactive and dynamic learning environments. This study explores the effectiveness of ChatGTP-driven VR in facilitating machine learning education through PerVRML. PerVRML incorporates a ChatGPT-powered avatar that provides real-time assistance and uses LLMs to personalize learning paths based on various sensor data from VR. A between-subjects design was employed to compare two learning modes: personalized and non-personalized. Quantitative data were collected from assessments, user experience surveys, and interaction metrics. The results indicate that while both learning modes supported learning effectively, ChatGPT-powered personalization significantly improved learning outcomes and had distinct impacts on user feedback. These findings underscore the potential of ChatGPT-enhanced VR to deliver adaptive and personalized educational experiences. Hong Gao 0008, Yiyang Xie, Enkelejda Kasneci |
Int. J. Hum. Comput. Interact. | 3 |
| 2026 | Lazy or Efficient? Towards Accessible Eye-Tracking Event Detection Using LLMs ETRA024abstractGaze event detection is fundamental to vision science, human-computer interaction, and applied analytics. However, current workflows often require specialized programming knowledge and careful handling of heterogeneous raw data formats. Classical detectors such as I-VT and I-DT are effective but highly sensitive to preprocessing and parameterization, limiting their usability outside specialized laboratories. This work introduces a code-free, large language model (LLM)-driven pipeline that converts natural language instructions into an end-to-end analysis. The system (1) inspects raw eye-tracking files to infer structure and metadata, (2) generates executable routines for data cleaning and detector implementation from concise user prompts, (3) applies the generated detector to label fixations and saccades, and (4) returns results and explanatory reports, and allows users to iteratively optimize their code by editing the prompt. Evaluated on public benchmarks, the approach achieves accuracy comparable to traditional methods while substantially reducing technical overhead. The framework lowers barriers to entry for eye-tracking research, providing a flexible and accessible alternative to code-intensive workflows. Dongyang Guo, Yasmeen Abdrabou, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2026 | What Shapes Participant Data Quality? A Scoping Review and Case Study of Crowdsourced Webcam Eye Tracking in AI Interviews ETRA003abstractWebcam-based eye tracking is a cost-effective, scalable method for remote research that effectively reaches broader populations. However, uncontrolled environments and hardware diversity lead to inconsistent data quality in crowdsourcing. To assess current practices, we conducted a scoping review of crowdsourced eye-tracking from 2011–2025. The review confirms fragmented reporting and a lack of established quality benchmarks. To address this lack of predictive insight, we conducted a case study on AI fairness interviews ( N = 205) using the RealEye platform. Applying Ordered Logistic Regression (OLR) to the platform’s quality metric, we found that behavioral and technical factors significantly predict data quality. Specifically, within the RealEye platform, higher fixation counts, shorter sessions, and operating system choice yield significantly higher quality grades. Based on this review and platform-specific predictive insights, we provide actionable recommendations to enhance the reliability, transparency, and replicability of future crowdsourced webcam eye tracking in HCI and behavioral science. Ka Hei Carrie Lau, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2026 | Secure Storage and Privacy-Preserving Scanpath Comparison via Garbled Circuits in Eye Tracking ETRA008abstractWith the growing use of eye tracking on VR and mobile platforms, gaze data is increasing. While scanpath comparison is important to gaze behavior analysis, existing methods lack privacy-preserving capabilities for real-world use. We present a garbled-circuit (GC)-based approach enabling secure storage and privacy-preserving scanpath comparison under the semi-honest model. It supports two configurations: (1) a two-party setting where the data owner and processor jointly compute similarity scores without revealing their inputs, and (2) a server-assisted setting where encrypted scanpaths are stored and processed while the data owner remains offline. All decryption and comparison operations are executed inside the GC. Experiments on three eye-tracking datasets evaluate fidelity, runtime, and communication, and show secure results for MultiMatch, ScanMatch, and SubsMatch closely match plaintext outcomes, with manageable runtime and communication overhead. Tests under various network conditions indicate that the design remains feasible for real-world privacy-preserving scanpath analysis and can be extended to other GC-based behavioral algorithms. Süleyman Özdel, Amr Nader, Yasmeen Abdrabou, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | From Passive Watching to Active Learning: Empowering Proactive Participation in Digital Classrooms with AI Video Assistant
Anna Bodonhelyi, Enkeleda Thaqi, Süleyman Özdel, Efe Bozkir, Enkelejda Kasneci |
CHI | 5 |
| 2025 | TRUCE-AV: A Multimodal Dataset for Trust and Comfort Estimation in Autonomous VehiclesabstractUnderstanding and estimating driver trust and comfort are essential for the safety and widespread acceptance of autonomous vehicles. Existing works analyze user trust and comfort separately, with limited real-time assessment and insufficient multimodal data. This paper introduces a novel multimodal dataset called TRUCE-AV, focusing on trust and comfort estimation in autonomous vehicles. The dataset collects real-time trust votes and continuous comfort ratings of 31 participants during a simulator-based fully autonomous driving. Simultaneously, physiological signals, such as heart rate, gaze, and emotions, along with environmental data (e.g., vehicle speed, nearby vehicle positions, and velocity), are recorded throughout the drives. Standard pre- and post-drive questionnaires were also administered to assess participants’ trust in automation and overall well-being, enabling the correlation of subjective assessments with real-time responses. To demonstrate the utility of our dataset, we evaluated various machine learning models for trust and comfort estimation using physiological data. Our analysis showed that tree-based models like Random Forest and XGBoost and non-linear models such as KNN and MLP regressor achieved the best performance for trust classification and comfort regression. Additionally, we identified key features that contribute to these estimations by using SHAP analysis on the top-performing models. Our dataset enables the development of adaptive AV systems capable of dynamically responding to user trust and comfort levels non-invasively, ultimately enhancing safety, user experience, and human-centered vehicle design. Aditi Bhalla, Christian Hellert, Enkelejda Kasneci, Nastassja Becker |
ECAI | 3 |
| 2025 | Automated Visual Attention Detection using Mobile Eye Tracking in Behavioral Classroom Studies
Efe Bozkir, Christian Kosel, Tina Seidel, Enkelejda Kasneci |
EDM | 4 |
| 2025 | Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach
Ruikun Hou, Babette Bühler, Tim Fütterer, Efe Bozkir, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
EDM | 7 |
| 2025 | Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms
Süleyman Özdel, Can Sarpkaya, Efe Bozkir, Hong Gao 0008, Enkelejda Kasneci |
EDM | 5 |
| 2025 | From Gaze to Data: Privacy and Societal Challenges of Using Eye-tracking Data to Inform GenAI Models
Yasmeen Abdrabou, Süleyman Özdel, Virmarie Maquiling, Efe Bozkir, Enkelejda Kasneci |
ETRA | 5 |
| 2025 | Exploring promptable foundation models for high-resolution video eye tracking in the lababstractWe explore whether SAM2, a vision foundation model, can be used for accurate localization of eye image features that are used in lab-based eye tracking: corneal reflections (CRs), the pupil, and the iris. We prompted SAM2 via a typical hand annotation process that consisted of clicking on the pupil, CR, iris and sclera for only one image per participant. SAM2 was found to support better spatial precision in the resulting gaze signals for the pupil (> 44% lower RMS-S2S), but not the CR and iris, than traditional image-processing methods or two state-of-the-art deep-learning tools. Providing more frames with prompts to initialize SAM2 did not improve performance. We conclude that SAM2’s powerful zero-shot segmentation capabilities provide an interesting new avenue to explore in high-resolution lab-based eye tracking. We provide our adaptation of SAM2’s codebase that allows segmenting videos of arbitrary duration and prepending arbitrary prompting frames. Diederick Christian Niehorster, Virmarie Maquiling, Sean Anthony Byrne, Enkelejda Kasneci, Marcus Nyström |
ETRA | 4 |
| 2025 | Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
Dongyang Guo, Yasmeen Abdrabou, Enkeleda Thaqi, Enkelejda Kasneci |
ICMI | 4 |
| 2025 | Adaptive Gen-AI Guidance in Virtual Reality: A Multimodal Exploration of Engagement in Neapolitan Pizza-MakingabstractVirtual reality (VR) offers promising opportunities for procedural learning, particularly in preserving intangible cultural heritage. Advances in generative artificial intelligence (Gen-AI) further enrich these experiences by enabling adaptive learning pathways. However, evaluating such adaptive systems using traditional temporal metrics remains challenging due to the inherent variability in Gen-AI response times. To address this, our study employs multimodal behavioural metrics, including visual attention, physical exploratory behaviour, and verbal interaction, to assess user engagement in an adaptive VR environment. In a controlled experiment with (n = 54) participants, we compared three levels of adaptivity (high, moderate, and non-adaptive baseline) within a Neapolitan pizza-making VR experience. Results show that moderate adaptivity optimally enhances user engagement, significantly reducing unnecessary exploratory behaviour and increasing focused visual attention on the AI avatar. Our findings suggest that a balanced level of adaptive AI provides the most effective user support, offering practical design recommendations for future adaptive educational technologies. Ka Hei Carrie Lau, Sema Sen, Philipp Stark, Efe Bozkir, Enkelejda Kasneci |
ICMI | 5 |
| 2025 | Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoringabstractThe manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance (e.g. alignment and reliability) of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1-preview, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs’ scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The o1 model outperforms all other LLMs, achieving Spearman’s = .74 with human assessments in the Overall score, and an internal consistency of = .80, though biased towards higher scores. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency to overrate and their remaining issues to capture the content quality, the models require further refinement. Kathrin Seßler, Maurice Fürstenberg, Babette Bühler, Enkelejda Kasneci |
LAK | 4 |
| 2025 | An Explainable Machine Learning Approach for Cognitive Load Detection in Virtual Reality Using Eye Tracking DataabstractAccurate cognitive load (CL) detection during virtual reality (VR) locomotion is critical for enhancing user experience and improving interaction design. Traditional CL assessment methods, such as self-reports and physiological measures, face challenges in VR environments. Eye tracking has shown potential as a reliable indicator of CL across various human-computer interaction (HCI) tasks. It offers significant promise as a discriminative feature for predictive models in VR. This study explores the feasibility of detecting CL induced by VR locomotion using an explainable machine-learning approach along with eye-tracking techniques. A comparative user study employing a within-subjects design evaluated five unique gait-free locomotion techniques. Statistical analysis revealed distinct CL levels across these locomotion techniques. Several machine learning models were developed for CL detection using eye-tracking data, with the Light Gradient Boosting Machine (LightGBM) achieving the highest accuracy of 0.78. The SHAP approach was employed to analyze the importance of features to provide interpretability, offering insights into the machine learning model's decision-making process. Our findings highlight the potential of using eye-tracking-based machine learning techniques as a practical approach for cognitive load detection in VR, contributing to the growing research in multimedia analytics, human perception, and user intent within immersive environments. Additionally, our work demonstrates how eye-tracking data can be leveraged to improve user interactions and optimize immersive multimedia experiences based on cognitive load analysis. Hong Gao 0008, Yapeng Gao, Enkelejda Kasneci |
ICMR | 3 |
| 2025 | Imperceptible Gaze Guidance Through Ocularity in Virtual RealityabstractWe introduce to VR a novel imperceptible gaze guidance technique from a recent discovery that human gaze can be attracted to a cue that contrasts from the background in its perceptually non-distinctive ocularity, defined as the relative difference between inputs to the two eyes. This cue pops out in the saliency map in the primary visual cortex without being overtly visible. We tested this method in an odd-one-out visual search task using eye tracking with 31 participants in VR. When the target was rendered as an ocularity singleton, participants' gaze was drawn to the target faster. Conversely, when a background object served as the ocularity singleton, it distracted gaze from the target. Since ocularity is nearly imperceptible, our method maintains user immersion while guiding attention without noticeable scene alterations and can render object's depth in 3D scenes, creating new possibilities for immersive user experience across diverse VR applications. Virmarie Maquiling, Li Zhaoping, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2025 | Eye-Tracked Virtual Reality: A Comprehensive Survey on Methods and Privacy ChallengesabstractThe latest developments in computer hardware, sensor technologies, and artificial intelligence can make virtual reality (VR) and virtual spaces an important part of human everyday life. Eye tracking offers not only a hands-free way of interaction but also the possibility of a deeper understanding of human visual attention and cognitive processes in VR. Despite these possibilities, eye-tracking data also reveal users’ privacy-sensitive attributes when combined with the information about the presented stimulus. To address all, this survey first covers major works in eye tracking, VR, and privacy areas between 2012 and 2022. While eye tracking in VR part covers the computational eye-tracking pipeline from pupil detection and gaze estimation to offline data analysis, for privacy and security, we focus on eye-based authentication as well as computational methods to preserve the privacy of individuals and their eye-tracking data in VR. Later, we outline three main directions by focusing on privacy. In summary, this survey presents an extensive literature review of the utmost possibilities of eye tracking in VR and their privacy implications. Efe Bozkir, Süleyman Özdel, Mengdi Wang 0002, Brendan David-John, Hong Gao 0008, Kevin R. B. Butler, Eakta Jain, Enkelejda Kasneci |
Proc. IEEE | 8 |
| 2025 | Can You Tell Real from Fake Face Images? Perception of Computer-Generated Faces by HumansabstractWith recent advances in machine learning and big data, it is now possible to create synthetic images that look real. Face generation is often of particular interest, as faces can be used for various purposes. However, improper use of such content can lead to the dissemination of false information, such as fake news, and thus pose a threat to society. This work studies whether people believe the truthfulness of faces using eye tracking and self-reports, including free-form textual explanations when participants encounter real and computer-generated faces. We used three different datasets for our evaluations, and our experimental results show that while people are relatively better at identifying the truthfulness of real faces and faces generated by earlier machine learning algorithms with different gazing behaviors in viewing and rating phases, they perform less accurately when deciding the truthfulness of synthetic face images that are generated by newer algorithms. Our findings provide important insights for society and policymakers. Efe Bozkir, Clara Riedmiller, Athanassios N. Skodras, Gjergji Kasneci, Enkelejda Kasneci |
ACM Trans. Appl. Percept. | 5 |
| 2025 | Efficient GNN Explanation via Learning Removal-based AttributionabstractAs Graph Neural Networks (GNNs) have been widely used in real-world applications, model explanations are required not only by users but also by legal regulations. However, simultaneously achieving high fidelity and low computational costs in generating explanations has been a challenge for current methods. In this work, we propose a framework of GNN explanation named L e A rn R emoval-based A ttribution (LARA) to address this problem. Specifically, we introduce removal-based attribution and demonstrate its substantiated link to interpretability fidelity theoretically and experimentally. The explainer in LARA learns to generate removal-based attribution which enables providing explanations with high fidelity. A strategy of subgraph sampling is designed in LARA to improve the scalability of the training process. In the deployment, LARA can efficiently generate the explanation through a feed-forward pass. We benchmark our approach with other state-of-the-art GNN explanation methods on six datasets. Results highlight the effectiveness of our framework regarding both efficiency and fidelity. In particular, LARA is 3.1 \(\times\) faster and achieves higher fidelity than the state-of-the-art method on the large dataset ogbn-arxiv (more than 160K nodes and 1M edges), showing its great potential in real-world applications. Our source code is available at https://github.com/yaorong0921/LARA . Yao Rong 0001, Guanchu Wang, Qizhang Feng, Ninghao Liu 0001, Zirui Liu 0001, Enkelejda Kasneci, Xia Ben Hu |
ACM Trans. Knowl. Discov. Data | 6 |
| 2024 | I-CEE: Tailoring Explanations of Image Classification Models to User ExpertiseabstractEffectively explaining decisions of black-box machine learning models is critical to responsible deployment of AI systems that rely on them. Recognizing their importance, the field of explainable AI (XAI) provides several techniques to generate these explanations. Yet, there is relatively little emphasis on the user (the explainee) in this growing body of work and most XAI techniques generate "one-size-fits-all'' explanations. To bridge this gap and achieve a step closer towards human-centered XAI, we present I-CEE, a framework that provides Image Classification Explanations tailored to User Expertise. Informed by existing work, I-CEE explains the decisions of image classification models by providing the user with an informative subset of training data (i.e., example images), corresponding local explanations, and model decisions. However, unlike prior work, I-CEE models the informativeness of the example images to depend on user expertise, resulting in different examples for different users. We posit that by tailoring the example set to user expertise, I-CEE can better facilitate users' understanding and simulatability of the model. To evaluate our approach, we conduct detailed experiments in both simulation and with human participants (N = 100) on multiple datasets. Experiments with simulated users show that I-CEE improves users' ability to accurately predict the model's decisions (simulatability) compared to baselines, providing promising preliminary results. Experiments with human participants demonstrate that our method significantly improves user simulatability accuracy, highlighting the importance of human-centered XAI. Yao Rong 0001, Peizhu Qian, Vaibhav V. Unhelkar, Enkelejda Kasneci |
AAAI | 4 |
| 2024 | TurboSVM-FL: Boosting Federated Learning through SVM Aggregation for Lazy ClientsabstractFederated learning is a distributed collaborative machine learning paradigm that has gained strong momentum in recent years. In federated learning, a central server periodically coordinates models with clients and aggregates the models trained locally by clients without necessitating access to local data. Despite its potential, the implementation of federated learning continues to encounter several challenges, predominantly the slow convergence that is largely due to data heterogeneity. The slow convergence becomes particularly problematic in cross-device federated learning scenarios where clients may be strongly limited by computing power and storage space, and hence counteracting methods that induce additional computation or memory cost on the client side such as auxiliary objective terms and larger training iterations can be impractical. In this paper, we propose a novel federated aggregation strategy, TurboSVM-FL, that poses no additional computation burden on the client side and can significantly accelerate convergence for federated classification task, especially when clients are "lazy" and train their models solely for few epochs for next global aggregation. TurboSVM-FL extensively utilizes support vector machine to conduct selective aggregation and max-margin spread-out regularization on class embeddings. We evaluate TurboSVM-FL on multiple datasets including FEMNIST, CelebA, and Shakespeare using user-independent validation with non-iid data distribution. Our results show that TurboSVM-FL can significantly outperform existing popular algorithms on convergence rate and reduce communication rounds while delivering better test metrics including accuracy, F1 score, and MCC. Mengdi Wang 0002, Anna Bodonhelyi, Efe Bozkir, Enkelejda Kasneci |
AAAI | 4 |
| 2024 | Automated Assessment of Encouragement and Warmth in Classrooms Leveraging Multimodal Emotional Features and ChatGPT
Ruikun Hou, Tim Fütterer, Babette Bühler, Efe Bozkir, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
AIED (1) | 7 |
| 2024 | Enhancing Student Motivation Through LLM-Powered Learning Environments - A Comparative Study
Kathrin Seßler, Ozan Kepir, Enkelejda Kasneci |
EC-TEL (2) | 3 |
| 2024 | Exploring Eye Tracking as a Measure for Cognitive Load Detection in VR LocomotionabstractEye tracking data has long been recognized as a reliable indicator of user cognitive load levels during human-computer interaction (HCI) tasks. However, its potential in the context of virtual reality (VR) remains relatively unexplored. Here, we present an ongoing study aimed at investigating the feasibility of detecting cognitive load in VR, particularly during VR locomotion, using an eye-tracking-based machine-learning approach. Data were collected using a within-subjects design, with participants performing VR locomotion tasks using five locomotion techniques. Our preliminary analyses validate the effectiveness of leveraging eye-tracking data as informative features in uncovering cognitive load in VR locomotion contexts, which motivates our further explorations. Hong Gao 0008, Enkelejda Kasneci |
ETRA | 2 |
| 2024 | Exploring Communication Dynamics: Eye-tracking Analysis in Pair Programming of Computer Science EducationabstractPair programming is widely recognized as an effective educational tool in computer science that promotes collaborative learning and mirrors real-world work dynamics. However, communication breakdowns within pairs significantly challenge this learning process. In this study, we use eye-tracking data recorded during pair programming sessions to study communication dynamics between various pair programming roles across different student, expert, and mixed group cohorts containing 19 participants. By combining eye-tracking data analysis with focus group interviews and questionnaires, we provide insights into communication’s multifaceted nature in pair programming. Our findings highlight distinct eye-tracking patterns indicating changes in communication skills across group compositions, with participants prioritizing code exploration over communication, especially during challenging tasks. Further, students showed a preference for pairing with experts, emphasizing the importance of understanding group formation in pair programming scenarios. These insights emphasize the importance of understanding group dynamics and enhancing communication skills through pair programming for successful outcomes in computer science education. Wunmin Jang, Hong Gao 0008, Tilman Michaeli, Enkelejda Kasneci |
ETRA | 4 |
| 2024 | A Transformer-Based Model for the Prediction of Human Gaze Behavior on VideosabstractEye-tracking applications that utilize the human gaze in video understanding tasks have become increasingly important. To effectively automate the process of video analysis based on eye-tracking data, it is important to accurately replicate human gaze behavior. However, this task presents significant challenges due to the inherent complexity and ambiguity of human gaze patterns. In this work, we introduce a novel method for simulating human gaze behavior. Our approach uses a transformer-based reinforcement learning algorithm to train an agent that acts as a human observer, with the primary role of watching videos and simulating human gaze behavior. We employed an eye-tracking dataset gathered from videos generated by the VirtualHome simulator, with a primary focus on activity recognition. Our experimental results demonstrate the effectiveness of our gaze prediction method by highlighting its capability to replicate human gaze behavior and its applicability for downstream tasks where real human-gaze is used as input. Süleyman Özdel, Yao Rong 0001, Mert Albaba, Yen-Ling Kuo, Xi Wang 0021, Enkelejda Kasneci |
ETRA | 6 |
| 2024 | Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on IntentionabstractHumans utilize their gaze to concentrate on essential information while perceiving and interpreting intentions in videos. Incorporating human gaze into computational algorithms can significantly enhance model performance in video understanding tasks. In this work, we address a challenging and innovative task in video understanding: predicting the actions of an agent in a video based on a partial video. We introduce the Gaze-guided Action Anticipation algorithm, which establishes a visual-semantic graph from the video input. Our method utilizes a Graph Neural Network to recognize the agent’s intention and predict the action sequence to fulfill this intention. To assess the efficiency of our approach, we collect a dataset containing household activities generated in the VirtualHome environment, accompanied by human gaze data of viewing videos. Our method outperforms state-of-the-art techniques, achieving a 7% improvement in accuracy for 18-class intention recognition. This highlights the efficiency of our method in learning important features from human gaze data. Süleyman Özdel, Yao Rong 0001, Mert Albaba, Yen-Ling Kuo, Xi Wang 0021, Enkelejda Kasneci |
ETRA | 6 |
| 2024 | Using Gaze Transition Entropy to Detect Classroom Discourse in a Virtual Reality ClassroomabstractThis paper explores gaze entropy as a metric for detecting classroom discourse events in a virtual reality (VR) classroom. Using data from a laboratory experiment with N = 240 secondary school students, we distinguished between events of teacher-centered classroom discourse (question, hand raising, answer) and teacher explanation by analyzing their transition and stationary gaze entropy. Employing multi-level regression models, both entropy measures effectively discriminated between the two events and distinguished different levels of classroom participation as indicated by the degree of hand-raising by virtual students. Furthermore, using both measures in a logistic regression model, the potential of gaze entropy could be demonstrated by predicting the two events with 67% accuracy. By analyzing transition and stationary entropy, the study attempts to uncover different gaze patterns associated with learning events in a virtual classroom. The results contribute to the research and development of VR scenarios that help to simulate effective learning environments. Philipp Stark, Alexander Jonas Jung, Jens-Uwe Hahn, Enkelejda Kasneci, Richard Göllner |
ETRA | 4 |
| 2024 | SARA: Smart AI Reading Assistant for Reading ComprehensionabstractSARA integrates Eye Tracking and state-of-the-art large language models in a mixed reality framework to enhance the reading experience by providing personalized assistance in real-time. By tracking eye movements, SARA identifies the text segments that attract the user’s attention the most and potentially indicate uncertain areas and comprehension issues. The process involves these key steps: text detection and extraction, gaze tracking and alignment, and assessment of detected reading difficulty. The results are customized solutions presented directly within the user’s field of view as virtual overlays on identified difficult text areas. This support enables users to overcome challenges like unfamiliar vocabulary and complex sentences by offering additional context, rephrased solutions, and multilingual help. SARA’s innovative approach demonstrates it has the potential to transform the reading experience and improve reading proficiency. Enkeleda Thaqi, Mohamed Omar Mantawy, Enkelejda Kasneci |
ETRA | 3 |
| 2024 | Detecting Aware and Unaware Mind Wandering During Lecture Viewing: A Multimodal Machine Learning Approach Using Eye Tracking, Facial Videos and Physiological DataabstractLearners often experience aware and unaware mind wandering during educational tasks, both negatively impacting learning outcomes. Differentiating these types of task-unrelated thoughts is crucial, as they stem from different cognitive processes and warrant tailored support that addresses the specific nature of mind wandering. Automated detection of these episodes could help mitigate their adverse effects, for example, by developing adaptive, attention-aware learning environments. In this study (N = 87), we explored a novel multimodal approach, combining eye tracking, facial videos, and physiological wristbands (i.e., electrodermal activity and heart rate), to predict aware and unaware mind wandering during lecture video watching. In addition, to allow comparison to previous research, we also predicted an integrated mind-wandering category. Mind wandering was assessed using 15 two-stage thought probes to determine task-unrelated thoughts and the participants’ awareness of their mind wandering. Our findings indicate that a multimodal approach outperforms unimodal methods, utilizing the top 100 features from the fused data. Specifically, aware mind wandering was detected at 20% above chance (AUC-PR = 0.396), unaware mind wandering at 14% above chance (AUC-PR = 0.267), and the combined category at 40% above chance (AUC-PR = 0.637). Eye tracking and video features proved more predictive than physiological measures when used as standalone modalities. SHAP analysis, employed to explain the results, highlighted the significance of integrating features from all three modalities for effective detection, particularly emphasizing the role of video-based facial expressions in identifying unaware mind wandering. Going beyond the current state of the art, this study demonstrates the potential of leveraging multimodal data to enhance the precision of aware and unaware mind-wandering detection and differentiation, setting a foundation for advancing educational technologies that respond dynamically to learners’ cognitive states. Babette Bühler, Efe Bozkir, Hannah Deininger, Patricia Goldberg, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
ICMI | 7 |
| 2024 | DataliVR: Transformation of Data Literacy Education through Virtual Reality with ChatGPT-Powered EnhancementsabstractData literacy is essential in today’s data-driven world, emphasizing individuals’ abilities to effectively manage data and extract meaningful insights. However, traditional classroom-based educational approaches often struggle to fully address the multifaceted nature of data literacy. As education undergoes digital transformation, innovative technologies such as Virtual Reality (VR) offer promising avenues for immersive and engaging learning experiences. This paper introduces DataliVR, a pioneering VR application aimed at enhancing the data literacy skills of university students within a contextual and gamified virtual learning environment. By integrating Large Language Models (LLMs) like ChatGPT as a conversational artificial intelligence (AI) chatbot embodied within a virtual avatar, DataliVR provides personalized learning assistance, enriching user learning experiences. Our study employed an experimental approach, with chatbot availability as the independent variable, analyzing learning experiences and outcomes as dependent variables with a sample of thirty participants. Our approach underscores the effectiveness and user-friendliness of ChatGPT-powered DataliVR in fostering data literacy skills. Moreover, our study examines the impact of the ChatGPT-based AI chatbot on users’ learning, revealing significant effects on both learning experiences and outcomes. Our study presents a robust tool for fostering data literacy skills, contributing significantly to the digital advancement of data literacy education through cutting-edge VR and AI technologies. Moreover, our research provides valuable insights and implications for future research endeavors aiming to integrate LLMs (e.g., ChatGPT) into educational VR platforms. Hong Gao 0008, Haochuan Huai, Sena Yildiz-Degirmenci, Maria Bannert, Enkelejda Kasneci |
ISMAR | 5 |
| 2024 | DART: An automated end-to-end object detection pipeline with data Diversification, open-vocabulary bounding box Annotation, pseudo-label Review, and model TrainingabstractAccurate real-time object detection is vital across numerous industrial applications, from safety monitoring to quality control. Traditional approaches, however, are hindered by arduous manual annotation and data collection, struggling to adapt to ever-changing environments and novel target objects. To address these limitations, this paper presents DART, an innovative automated end-to-end pipeline that revolutionizes object detection workflows from data collection to model evaluation. It eliminates the need for laborious human labeling and extensive data collection while achieving outstanding accuracy across diverse scenarios. DART encompasses four key stages: (1) Data Diversification using subject-driven image generation (DreamBooth with SDXL), (2) Annotation via open-vocabulary object detection (Grounding DINO) to generate bounding box and class labels, (3) Review of generated images and pseudo-labels by large multimodal models (InternVL-1.5 and GPT-4o) to guarantee credibility, and (4) Training of real-time object detectors (YOLOv8 and YOLOv10) using the verified data. We apply DART to a self-collected dataset of construction machines named Liebherr Product, which contains over 15K high-quality images across 23 categories. The current instantiation of DART significantly increases average precision (AP) from 0.064 to 0.832. Its modular design ensures easy exchangeability and extensibility, allowing for future algorithm upgrades, seamless integration of new object categories, and adaptability to customized environments without manual labeling and additional data collection. The code and dataset are released at https://github.com/chen-xin-94/DART. Chen Xin 0002, Andreas Hartel, Enkelejda Kasneci |
Expert Syst. Appl. | 3 |
| 2024 | On Task and in Sync: Examining the Relationship between Gaze Synchrony and Self-reported Attention During Video Lecture LearningabstractSuccessful learning depends on learners' ability to sustain attention, which is particularly challenging in online education due to limited teacher interaction. A potential indicator for attention is gaze synchrony, demonstrating predictive power for learning achievements in video-based learning in controlled experiments focusing on manipulating attention. This study (N=84) examines the relationship between gaze synchronization and self-reported attention of learners, using experience sampling, during realistic online video learning. Gaze synchrony was assessed through Kullback-Leibler Divergence of gaze density maps and MultiMatch algorithm scanpath comparisons. Results indicated significantly higher gaze synchronization in attentive participants for both measures and self-reported attention significantly predicted post-test scores. In contrast, synchrony measures did not correlate with learning outcomes. While supporting the hypothesis that attentive learners exhibit similar eye movements, the direct use of synchrony as an attention indicator poses challenges, requiring further research on the interplay of attention, gaze synchrony, and video content type. Babette Bühler, Efe Bozkir, Hannah Deininger, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2024 | Privacy-preserving Scanpath Comparison for Pervasive Eye TrackingabstractAs eye tracking becomes pervasive with screen-based devices and head-mounted displays, privacy concerns regarding eye-tracking data have escalated. While state-of-the-art approaches for privacy-preserving eye tracking mostly involve differential privacy and empirical data manipulations, previous research has not focused on methods for scanpaths. We introduce a novel privacy-preserving scanpath comparison protocol designed for the widely used Needleman-Wunsch algorithm, a generalized version of the edit distance algorithm. Particularly, by incorporating the Paillier homomorphic encryption scheme, our protocol ensures that no private information is revealed. Furthermore, we introduce a random processing strategy and a multi-layered masking method to obfuscate the values while preserving the original order of encrypted editing operation costs. This minimizes communication overhead, requiring a single communication round for each iteration of the Needleman-Wunsch process. We demonstrate the efficiency and applicability of our protocol on three publicly available datasets with comprehensive computational performance analyses and make our source code publicly accessible. Süleyman Özdel, Efe Bozkir, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Towards Human-Centered Explainable AI: A Survey of User Studies for Model ExplanationsabstractExplainable AI (XAI) is widely viewed as a sine qua non for ever-expanding AI research. A better understanding of the needs of XAI users, as well as human-centered evaluations of explainable models are both a necessity and a challenge. In this paper, we explore how human-computer interaction (HCI) and AI researchers conduct user studies in XAI applications based on a systematic literature review. After identifying and thoroughly analyzing 97 core papers with human-based XAI evaluations over the past five years, we categorize them along the measured characteristics of explanatory methods, namely trust, understanding, usability, and human-AI collaboration performance. Our research shows that XAI is spreading more rapidly in certain application domains, such as recommender systems than in others, but that user evaluations are still rather sparse and incorporate hardly any insights from cognitive or social sciences. Based on a comprehensive discussion of best practices, i.e., common models, design choices, and measures in user studies, we propose practical guidelines on designing and conducting user studies for XAI researchers and practitioners. Lastly, this survey also highlights several open research directions, particularly linking psychological science and human-centered XAI. Yao Rong 0001, Tobias Leemann, Thai-trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav V. Unhelkar, Tina Seidel, Gjergji Kasneci, Enkelejda Kasneci |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | Automated Hand-Raising Detection in Classroom Videos: A View-Invariant and Occlusion-Robust Machine Learning Approach
Babette Bühler, Ruikun Hou, Efe Bozkir, Patricia Goldberg, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
AIED | 7 |
| 2023 | PEER: Empowering Writing with Large Language ModelsabstractAbstract The emerging research area of large language models (LLMs) has far-reaching implications for various aspects of our daily lives. In education, in particular, LLMs hold enormous potential for enabling personalized learning and equal opportunities for all students. In a traditional classroom environment, students often struggle to develop individual writing skills because the workload of the teachers limits their ability to provide detailed feedback on each student’s essay. To bridge this gap, we have developed a tool called PEER (Paper Evaluation and Empowerment Resource) which exploits the power of LLMs and provides students with comprehensive and engaging feedback on their essays. Our goal is to motivate each student to enhance their writing skills through positive feedback and specific suggestions for improvement. Since its launch in February 2023, PEER has received high levels of interest and demand, resulting in more than 4000 essays uploaded to the platform to date. Moreover, there has been an overwhelming response from teachers who are interested in the project since it has the potential to alleviate their workload by making the task of grading essays less tedious. By collecting a real-world data set incorporating essays of students and feedback from teachers, we will be able to refine and enhance PEER through model fine-tuning in the next steps. Our goal is to leverage LLMs to enhance personalized learning, reduce teacher workload, and ensure that every student has an equal opportunity to excel in writing. The code is available at https://github.com/Kasneci-Lab/AI-assisted-writing . Kathrin Seßler, Tao Xiang 0003, Lukas Bogenrieder, Enkelejda Kasneci |
EC-TEL | 4 |
| 2023 | Leveraging Eye Tracking in Digital Classrooms: A Step Towards Multimodal Model for Learning AssistanceabstractInstructors who teach digital literacy skills are increasingly faced with the challenges that come with larger student populations and online courses. We asked an educator how we could support student learning and better assist instructors both online and in the classroom. To address these challenges, we discuss how behavioral signals collected from eye tracking and mouse tracking can be combined to offer predictions of student performance. In our preliminary study, participants completed two image masking tasks in Adobe Photoshop based on real college-level course content. We then trained a machine learning model to predict student performance in each task based on data from other students, as a step towards offering automated student assistance and feedback to instructors. We reflect on the challenges and scalability issues to deploying such a system in-the-wild, and present some guidelines for future work. Sean Anthony Byrne, Nora Castner, Ard Kastrati, Martyna Plomecka, William Schaefer, Enkelejda Kasneci, Zoya Bylinskii |
ETRA | 6 |
| 2023 | Watch out for those bananas! Gaze Based Mario Kart Performance ClassificationabstractThis paper is about a small eye tracking study for scan path classification. Seven participants played Mario Kart while wearing a head mounted eye tracker. In total, we had 64 recordings, but one had to be removed (Only 79 gaze samples were recorded). We compared different scan path classification features to estimate the performance of the participants based on the ranking they achieved. The best performing feature was ENCODJI which incooperates saccades and the heatmap in one feature. HOV, which uses saccade angles, performed well for all tasks but was outperformed by the heatmap (HEAT) for the last two groups. Wolfgang Fuhl, Björn Severitt, Nora Castner, Babette Bühler, Johannes Meyer 0001, Daniel Weber 0003, Regine Lendway, Ruikun Hou, Enkelejda Kasneci |
ETRA | 9 |
| 2023 | Old or Modern? A Computational Model for Classifying Poem Comprehension using MicrosaccadesabstractNo abstract available. Patrizia Lenhart, Enkeleda Thaqi, Nora Castner, Enkelejda Kasneci |
ETRA | 4 |
| 2023 | Pupil Diameter during Counting Tasks as Potential Baseline for Virtual Reality ExperimentsabstractPupil diameter is a reliable indicator of mental effort, but it must be baseline corrected to account for its idiosyncratic nature. Established methods for measuring baselines cannot be applied in virtual reality (VR) experiments. To reliably measure a pupil diameter baseline in VR, we propose a short testing environment of visual arithmetic tasks. In an experiment with 66 university students, we analyzed external reliability and internal validity criteria for pupil diameter measures during counting and summation tasks. During the counting task, we found a high retest reliability between stimulus intervals. Acceptable retest reliability was found for task repetition at a second measuring time. Analyzing internal validity, we found that pupil diameter increased with task difficulty comparing both tasks. Further, a linear effect was found between the pupil diameter amplitude and luminance levels. Our findings highlight the potential of counting tasks as a pupil diameter baseline for VR experiments. Philipp Stark, Tobias Appel, Milo J. Olbrich, Enkelejda Kasneci |
ETRA | 4 |
| 2023 | Multiperspective Teaching of Unknown Objects via Shared-gaze-based Multimodal Human-Robot InteractionabstractFor successful deployment of robots in multifaceted situations, an understanding of the robot for its environment is indispensable. With advancing performance of state-of-the-art object detectors, the capability of robots to detect objects within their interaction domain is also enhancing. However, it binds the robot to a few trained classes and prevents it from adapting to unfamiliar surroundings beyond predefined scenarios. In such scenarios, humans could assist robots amidst the overwhelming number of interaction entities and impart the requisite expertise by acting as teachers. We propose a novel pipeline that effectively harnesses human gaze and augmented reality in a human-robot collaboration context to teach a robot novel objects in its surrounding environment. By intertwining gaze (to guide the robot's attention to an object of interest) with augmented reality (to convey the respective class information) we enable the robot to quickly acquire a significant amount of automatically labeled training data on its own. Training in a transfer learning fashion, we demonstrate the robot's capability to detect recently learned objects and evaluate the influence of different machine learning models and learning procedures as well as the amount of training data involved. Our multimodal approach proves to be an efficient and natural way to teach the robot novel objects based on a few instances and allows it to detect classes for which no training dataset is available. In addition, we make our dataset publicly available to the research community, which consists of RGB and depth data, intrinsic and extrinsic camera parameters, along with regions of interest. Daniel Weber 0003, Wolfgang Fuhl, Enkelejda Kasneci, Andreas Zell |
HRI | 3 |
| 2023 | Probabilistic Contrastive Learning Recovers the Correct Aleatoric Uncertainty of Ambiguous InputsabstractContrastively trained encoders have recently been proven to invert the data-generating process: they encode each input, e.g., an image, into the true latent vector that generated the image (Zimmermann et al., 2021). However, real-world observations often have inherent ambiguities. For instance, images may be blurred or only show a 2D view of a 3D object, so multiple latents could have generated them. This makes the true posterior for the latent vector probabilistic with heteroscedastic uncertainty. In this setup, we extend the common InfoNCE objective and encoders to predict latent distributions instead of points. We prove that these distributions recover the correct posteriors of the data-generating process, including its level of aleatoric uncertainty, up to a rotation of the latent space. In addition to providing calibrated uncertainty estimates, these posteriors allow the computation of credible intervals in image retrieval. They comprise images with the same latent as a given query, subject to its uncertainty. Code is at https://github.com/mkirchhof/Probabilistic_Contrastive_Learning . Michael Kirchhof 0002, Enkelejda Kasneci, Seong Joon Oh |
ICML | 2 |
| 2023 | Leveraging Saliency-Aware Gaze Heatmaps for Multiperspective Teaching of Unknown ObjectsabstractAs robots become increasingly prevalent amidst diverse environments, their ability to adapt to novel scenarios and objects is essential. Advances in modern object detection have also paved the way for robots to identify interaction entities within their immediate vicinity. One drawback is that the robot's operational domain must be known at the time of training, which hinders the robot's ability to adapt to unexpected environments outside the preselected classes. However, when encountering such challenges a human can provide support to a robot by teaching it about the new, yet unknown objects on an ad hoc basis. In this work, we merge augmented reality and human gaze in the context of multimodal human-robot interaction to compose saliency-aware gaze heatmaps leveraged by a robot to learn emerging objects of interest. Our results show that our proposed method exceeds the capabilities of the current state of the art and outperforms it in terms of commonly used object detection metrics. Daniel Weber 0003, Valentin Bolz, Andreas Zell, Enkelejda Kasneci |
IROS | 4 |
| 2023 | Detecting Teacher Expertise in an Immersive VR Classroom: Leveraging Fused Sensor Data with Explainable Machine Learning ModelsabstractCurrently, VR technology is increasingly being used in applications to enable immersive yet controlled research settings. One such area of research is expertise assessment, where novel technological approaches to collecting process data, specifically eye tracking, in combination with explainable models, can provide insights into assessing and training novices, as well as fostering expertise development. We present a machine learning approach to predict teacher expertise by leveraging data from an off-the-shelf VR device collected in a VirATec study. By fusing eye-tracking and controller-tracking data, teachers’ recognition and handling of disruptive events in the classroom are taken into account or considered. Three classification models were compared, including SVM, Random Forest, and LightGBM, with Random Forest achieving the best ROC-AUC score of 0.768 in predicting teacher expertise. The SHAP approach to model interpretation revealed informative features (e.g., fixations on identified disruptive students) for distinguishing teacher expertise. Our study serves as a pioneering effort in assessing teacher expertise using eye tracking within an interactive virtual setting, paving the way for future research and advancements in the field. Hong Gao 0008, Efe Bozkir, Philipp Stark, Patricia Goldberg, Gerrit Meixner, Enkelejda Kasneci, Richard Göllner |
ISMAR | 6 |
| 2023 | URL: A Representation Learning Benchmark for Transferable Uncertainty EstimatesabstractRepresentation learning has significantly driven the field to develop pretrained models that can act as a valuable starting point when transferring to new datasets. With the rising demand for reliable machine learning and uncertainty quantification, there is a need for pretrained models that not only provide embeddings but also transferable uncertainty estimates. To guide the development of such models, we propose the Uncertainty-aware Representation Learning (URL) benchmark. Besides the transferability of the representations, it also measures the zero-shot transferability of the uncertainty estimate using a novel metric. We apply URL to evaluate ten uncertainty quantifiers that are pretrained on ImageNet and transferred to eight downstream datasets. We find that approaches that focus on the uncertainty of the representation itself or estimate the prediction risk directly outperform those that are based on the probabilities of upstream classes. Yet, achieving transferable uncertainty quantification remains an open challenge. Our findings indicate that it is not necessarily in conflict with traditional representation learning goals. Code is available at https://github.com/mkirchhof/url. Michael Kirchhof 0002, Bálint Mucsányi, Seong Joon Oh, Enkelejda Kasneci |
NeurIPS | 4 |
| 2023 | When are post-hoc conceptual explanations identifiable?abstractInterest in understanding and factorizing learned embedding spaces through conceptual explanations is steadily growing. When no human concept labels are available, concept discovery methods search trained embedding spaces for interpretable concepts like object shape or color that can provide post-hoc explanations for decisions. Unlike previous work, we argue that concept discovery should be identifiable, meaning that a number of known concepts can be provably recovered to guarantee reliability of the explanations. As a starting point, we explicitly make the connection between concept discovery and classical methods like Principal Component Analysis and Independent Component Analysis by showing that they can recover independent concepts under non-Gaussian distributions. For dependent concepts, we propose two novel approaches that exploit functional compositionality properties of image-generating processes. Our provably identifiable concept discovery methods substantially outperform competitors on a battery of experiments including hundreds of trained models and dependent concepts, where they exhibit up to 29 % better alignment with the ground truth. Our results highlight the strict conditions under which reliable concept discovery without human labels can be guaranteed and provide a formal foundation for the domain. Our code is available online. Tobias Leemann, Michael Kirchhof 0002, Yao Rong 0001, Enkelejda Kasneci, Gjergji Kasneci |
UAI | 4 |
| 2023 | A multimodal smartwatch-based interaction concept for immersive environments
Matej Lang, Clemens Strobel, Felix Weckesser, Danielle Kathryn Langlois, Enkelejda Kasneci, Barbora Kozlíková, Michael Krone |
Comput. Graph. | 5 |
| 2023 | Deep learning-based position detection for hydraulic cylinders using scattering parameters
Chen Xin 0002, Thomas Motz, Wolfgang Fuhl, Andreas Hartel, Enkelejda Kasneci |
Expert Syst. Appl. | 5 |
| 2023 | Exploring the Effects of Scanpath Feature Engineering for Supervised Image Classification ModelsabstractImage classification models are becoming a popular method of analysis for scanpath classification. To implement these models, gaze data must first be reconfigured into a 2D image. However, this step gets relatively little attention in the literature as focus is mostly placed on model configuration. As standard model architectures have become more accessible to the wider eye-tracking community, we highlight the importance of carefully choosing feature representations within scanpath images as they may heavily affect classification accuracy. To illustrate this point, we create thirteen sets of scanpath designs incorporating different eye-tracking feature representations from data recorded during a task-based viewing experiment. We evaluate each scanpath design by passing the sets of images through a standard pre-trained deep learning model as well as a SVM image classifier. Results from our primary experiment show an average accuracy improvement of 25 percentage points between the best-performing set and one baseline set. Sean Anthony Byrne, Virmarie Maquiling, Adam Peter Frederick Reynolds, Luca Polonio, Nora Castner, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2023 | Cross-Task and Cross-Participant Classification of Cognitive Load in an Emergency Simulation GameabstractAssessment of cognitive load is a major step towards adaptive interfaces. However, non-invasive assessment is rather subjective as well as task specific and generalizes poorly, mainly due to methodological limitations. Additionally, it heavily relies on performance data like game scores or test results. In this study, we present an eye-tracking approach that circumvents these shortcomings and allows for effective generalizing across participants and tasks. First, we established classifiers for predicting cognitive load individually for a typical working memory task (n-back), which we then applied to an emergency simulation game by considering the similar ones and weighting their predictions. Standardization steps helped achieve high levels of cross-task and cross-participant classification accuracy between 63.78 and 67.25 percent for the distinction between easy and hard levels of the emergency simulation game. These very promising results could pave the way for novel adaptive computer-human interaction across domains and particularly for gaming and learning environments. Tobias Appel, Peter Gerjets, Korbinian Moeller, Manuel Ninaus, Christian Scharinger, Natalia Sevcenko, Franz Wortha, Enkelejda Kasneci |
IEEE Trans. Affect. Comput. | 9 |
| 2023 | Multimodal Engagement Analysis From Facial Videos in the ClassroomabstractStudent engagement is a key component of learning and teaching, resulting in a plethora of automated methods to measure it. Whereas most of the literature explores student engagement analysis using computer-based learning often in the lab, we focus on using classroom instruction in authentic learning environments. We collected audiovisual recordings of secondary school classes over a one and a half month period, acquired continuous engagement labeling per student (N=15) in repeated sessions, and explored computer vision methods to classify engagement from facial videos. We learned deep embeddings for attentional and affective features by training Attention-Net for head pose estimation and Affect-Net for facial expression recognition using previously-collected large-scale datasets. We used these representations to train engagement classifiers on our data, in individual and multiple channel settings, considering temporal dependencies. The best performing engagement classifiers achieved student-independent AUCs of .620 and .720 for grades 8 and 12, respectively, with attention-based features outperforming affective features. Score-level fusion either improved the engagement classifiers or was on par with the best performing modality. We also investigated the effect of personalization and found that only 60 seconds of person-specific data, selected by margin uncertainty of the base classifier, yielded an average AUC improvement of .084. Ömer Sümer, Patricia Goldberg, Sidney K. D'Mello, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
IEEE Trans. Affect. Comput. | 6 |
| 2022 | A Non-isotropic Probabilistic Take on Proxy-based Deep Metric Learning
Michael Kirchhof 0002, Karsten Roth, Zeynep Akata, Enkelejda Kasneci |
ECCV (26) | 4 |
| 2022 | Predicting Decision-Making during an Intelligence Test via Semantic Scanpath ComparisonsabstractFluid intelligence is considered to be the foundation to many aspects of human learning and performance. Individuals’ behavior while solving intelligence tests is therefore an important component in understanding problem-solving strategies and learning processes. We present preliminary results of a novel eye-tracking-based approach to predict participants’ decisions while solving a fluid intelligence test that utilizes semantic scanpath comparisons. Normalizing scanpaths and applying a knn classifier allows us to make individual predictions and combine them to predict final scores. We evaluated our proposed approach on the TüEyeQ dataset published by Kasneci et al. containing data of 315 university students, who worked on the Culture Fair Intelligence Test. Our approach was able to explain 39.207% of variance in the final score and predictions for participants’ final scores showed a correlation of τ = 0.65759 with participants’ actual scores. Overall, the proposed method has shown great potential that can be expanded on in future research. Tobias Appel, Lisa Bardach, Enkelejda Kasneci |
ETRA | 3 |
| 2022 | Regressive Saccadic Eye Movements on Fake NewsabstractWith the increasing use of the Internet, people encounter a variety of news in online media and social media every day. For digital content without fact-checking mechanisms, it is likely that people perceive fake news as real when they do not have extensive knowledge about the news topic. In this paper, we study human eye movements when reading fake news and real news. Our results suggest that people regress more with their eyes when reading fake news, while the time until the first fixation in the text area of interest is not a distinguishing factor between real and fake content. Our results show that although the truthfulness of the content is not known to people in advance, their visual behavior differs when reading such content, indicating a higher level of confusion when reading fake content. Efe Bozkir, Gjergji Kasneci, Sonja Utz, Enkelejda Kasneci |
ETRA | 4 |
| 2022 | LSTMs can distinguish dental expert saccade behavior with high "plaque-urracy"abstractMuch of the current expertise literature has found that domain specific tasks evoke different eye movements. However, research has yet to predict optimal image exploration using saccadic information and to identify and quantify differences in the search strategies between learners, intermediates, and expert practitioners. By employing LSTMs for scanpath classification, we found saccade features over time could distinguish all groups at high accuracy. The most distinguishing features were saccade velocity peak (72%), length (70%), and velocity average (68%). These findings promote the holistic theory of expert visual exploration that experts can quickly process the whole scene using longer and more rapid saccade behavior initially. The potential to integrate expertise model development from saccadic scanpath features into intelligent tutoring systems is the ultimate inspiration for our research. Additionally, this model is not confined to visual exploration in dental xrays, rather it can extend to other medical domains. Nora Castner, Jonas Frankemölle, Constanze Keutel, Fabian Hüttig, Enkelejda Kasneci |
ETRA | 5 |
| 2022 | A gaze-based study design to explore how competency evolves during a photo manipulation taskabstractShare on A gaze-based study design to explore how competency evolves during a photo manipulation task Authors: Nora Castner Human-Computer Interaction/ Wilhelm-Schickard-Institute, University of Tübingen, Germany Human-Computer Interaction/ Wilhelm-Schickard-Institute, University of Tübingen, GermanyView Profile , Bela Umlauf Human - Computer Interaction Group, University of Tübingen, Germany Human - Computer Interaction Group, University of Tübingen, GermanyView Profile , Ard Kastrati Computer Engineering and Networks Laboratory, ETH Zurich, Switzerland Computer Engineering and Networks Laboratory, ETH Zurich, SwitzerlandView Profile , Martyna Beata Płomecka Methods of Plasticity Reasearch, University of Zurich, Switzerland Methods of Plasticity Reasearch, University of Zurich, SwitzerlandView Profile , William Schaefer University of Texas at San Antonio, United States University of Texas at San Antonio, United StatesView Profile , Enkelejda Kasneci University of Tubingen, Germany University of Tubingen, GermanyView Profile , Zoya Bylinskii Adobe Research, United States Adobe Research, United StatesView Profile Authors Info & Claims ETRA '22: 2022 Symposium on Eye Tracking Research and ApplicationsJune 2022 Article No.: 37Pages 1–3https://doi.org/10.1145/3517031.3531634Online:08 June 2022Publication History 0citation30DownloadsMetricsTotal Citations0Total Downloads30Last 12 Months30Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Nora Castner, Bela Umlauf, Ard Kastrati, Martyna Plomecka, William Schaefer, Enkelejda Kasneci, Zoya Bylinskii |
ETRA | 6 |
| 2022 | HPCGen: Hierarchical K-Means Clustering and Level Based Principal Components for Scan Path GenarationabstractIn this paper, we present a new approach for decomposing scan paths and its utility for generating new scan paths. For this purpose, we use the K-Means clustering procedure to the raw gaze data and subsequently iteratively to find more clusters in the found clusters. The found clusters are grouped for each level in the hierarchy, and the most important principal components are computed from the data contained in them. Using this tree hierarchy and the principal components, new scan paths can be generated that match the human behavior of the original data. We show that this generated data is very useful for generating new data for scan path classification but can also be used to generate fake scan paths. Code can be downloaded here https://atreus.informatik.uni-tuebingen.de/seafile/d/8e2ab8c3fdd444e1a135/?p=%2FHPCGen&mode=list. Wolfgang Fuhl, Enkelejda Kasneci |
ETRA | 2 |
| 2022 | A Holographic Single-Pixel Stereo Camera Sensor for Calibration-free Eye-Tracking in Retinal Projection Augmented Reality GlassesabstractEye-tracking is a key technology for future retinal projection based AR glasses as it enables techniques such as foveated rendering or gaze-driven exit pupil steering, which both increases the system’s overall performance. However, two of the major challenges video oculography systems face are robust gaze estimation in the presence of glasses slippage, paired with the necessity of frequent sensor calibration. To overcome these challenges, we propose a novel, calibration-free eye-tracking sensor for AR glasses based on a highly transparent holographic optical element (HOE) and a laser scanner. We fabricate a segmented HOE generating two stereo images of the eye-region. A single-pixel detector in combination with our stereo reconstruction algorithm is used to precisely calculate the gaze position. In our laboratory setup we demonstrate a calibration-free accuracy of 1.35° achieved by our eye-tracking sensor; highlighting the sensor’s suitability for consumer AR glasses. Johannes Meyer 0001, Tobias Wilm, Reinhold Fiess, Thomas Schlebusch, Wilhelm Stork, Enkelejda Kasneci |
ETRA | 6 |
| 2022 | Exploiting Augmented Reality for Extrinsic Robot Calibration and Eye-based Human-Robot CollaborationabstractFor sensible human-robot interaction, it is crucial for the robot to have an awareness of its physical surroundings. In practical applications, however, the environment is manifold and possible objects for interaction are innumerable. Due to this fact, the use of robots in variable situations surrounded by unknown interaction entities is challenging and the inclusion of pre-trained object-detection neural networks not always feasible. In this work, we propose deploying augmented reality and eye tracking to flexibilize robots in non-predefined scenarios. To this end, we present and evaluate a method for extrinsic calibration of robot sensors, specifically a camera in our case, that is both fast and user-friendly, achieving competitive accuracy compared to classical approaches. By incorporating human gaze into the robot's segmentation process, we enable the 3D detection and localization of unknown objects without any training. Such an approach can facilitate interaction with objects for which training data is not available. At the same time, a visualization of the resulting 3D bounding boxes in the human's augmented reality leads to exceedingly direct feedback, providing insight into the robot's state of knowledge. Our approach thus opens the door to additional interaction possibilities, such as the subsequent initialization of actions like grasping. Daniel Weber 0003, Enkelejda Kasneci, Andreas Zell |
HRI | 2 |
| 2022 | A Consistent and Efficient Evaluation Strategy for Attribution MethodsabstractWith a variety of local feature attribution methods being proposed in recent years, follow-up work suggested several evaluation strategies. To assess the attribution quality across different attribution techniques, the most popular among these evaluation strategies in the image domain use pixel perturbations. However, recent advances discovered that different evaluation strategies produce conflicting rankings of attribution methods and can be prohibitively expensive to compute. In this work, we present an information-theoretic analysis of evaluation strategies based on pixel perturbations. Our findings reveal that the results are strongly affected by information leakage through the shape of the removed pixels as opposed to their actual values. Using our theoretical insights, we propose a novel evaluation framework termed Remove and Debias (ROAD) which offers two contributions: First, it mitigates the impact of the confounders, which entails higher consistency among evaluation strategies. Second, ROAD does not require the computationally expensive retraining step and saves up to 99% in computational costs compared to the state-of-the-art. We release our source code at https://github.com/tleemann/road_evaluation. Yao Rong 0001, Tobias Leemann, Vadim Borisov, Gjergji Kasneci, Enkelejda Kasneci |
ICML | 5 |
| 2022 | Maximum and Leaky Maximum PropagationabstractIn this work, we present an alternative to conventional residual connections, which is inspired by maxout nets. This means that instead of the addition in residual connections, our approach only propagates the maximum value or, in the leaky formulation, propagates a percentage of both. In our eval-uation, we show on different public data sets that the presented approaches are comparable to the residual connections and have other interesting properties, such as better generalization with a constant batch normalization, faster learning, and also the possibility to generalize without additional activation functions. In addition, the proposed approaches work very well if ensembles together with residual networks are formed. LinkToCodeBlind Wolfgang Fuhl, Enkelejda Kasneci |
IJCNN | 2 |
| 2022 | Evaluating the Effects of Virtual Human Animation on Students in an Immersive VR Classroom Using Eye MovementsabstractVirtual humans presented in VR learning environments have been suggested in previous research to increase immersion and further positively influence learning outcomes. However, how virtual human animations affect students’ real-time behavior during VR learning has not yet been investigated. This work examines the effects of social animations (i.e., hand raising of virtual peer learners) on students’ cognitive response and visual attention behavior during immersion in a VR classroom based on eye movement analysis. Our results show that animated peers that are designed to enhance immersion and provide companionship and social information elicit different responses in students (i.e., cognitive, visual attention, and visual search responses), as reflected in various eye movement metrics such as pupil diameter, fixations, saccades, and dwell times. Furthermore, our results show that the effects of animations on students differ significantly between conditions (20%, 35%, 65%, and 80% of virtual peer learners raising their hands). Our research provides a methodological foundation for investigating the effects of avatar animations on users, further suggesting that such effects should be considered by developers when implementing animated virtual humans in VR. Our findings have important implications for future works on the design of more effective, immersive, and authentic VR environments. Hong Gao 0008, Lisa Hasenbein, Efe Bozkir, Richard Göllner, Enkelejda Kasneci |
VRST | 5 |
| 2022 | Eye-Tracking-Based Prediction of User Experience in VR Locomotion Using Machine LearningabstractAbstract VR locomotion is one of the most important design features of VR applications and is widely studied. When evaluating locomotion techniques, user experience is usually the first consideration, as it provides direct insights into the usability of the locomotion technique and users' thoughts about it. In the literature, user experience is typically measured with post‐hoc questionnaires or surveys, while users' behavioral (i.e., eye‐tracking) data during locomotion, which can reveal deeper subconscious thoughts of users, has rarely been considered and thus remains to be explored. To this end, we investigate the feasibility of classifying users experiencing VR locomotion into L‐UE and H‐UE (i.e., low‐ and high‐user‐experience groups) based on eye‐tracking data alone. To collect data, a user study was conducted in which participants navigated a virtual environment using five locomotion techniques and their eye‐tracking data was recorded. A standard questionnaire assessing the usability and participants' perception of the locomotion technique was used to establish the ground truth of the user experience. We trained our machine learning models on the eye‐tracking features extracted from the time‐series data using a sliding window approach. The best random forest model achieved an average accuracy of over 0.7 in 50 runs. Moreover, the SHapley Additive exPlanations (SHAP) approach uncovered the underlying relationships between eye‐tracking features and user experience, and these findings were further supported by the statistical results. Our research provides a viable tool for assessing user experience with VR locomotion, which can further drive the improvement of locomotion techniques. Moreover, our research benefits not only VR locomotion, but also VR systems whose design needs to be improved to provide a good user experience. Hong Gao 0008, Enkelejda Kasneci |
Comput. Graph. Forum | 2 |
| 2022 | PACMHCI V6, ETRA, May 2022 EditorialabstractWe are delighted to present a first issue of the Proceedings of the ACM on Human-Computer Interaction to focus on contributions from the Eye Tracking Research and Applications (ETRA) community. Hans-Werner Gellersen, Enkelejda Kasneci, Krzysztof Krejtz, Daniel Weiskopf |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | U-HAR: A Convolutional Approach to Human Activity Recognition Combining Head and Eye Movements for Context-Aware Smart GlassesabstractAfter the success of smartphones and smartwatches, smart glasses are expected to be the next smart wearable. While novel display technology allows the seamlessly embedding of content into the FOV, interaction methods with glasses, requiring the user for active interaction, limiting the user experience. One way to improve this and drive immersive augmentation is to reduce user interactions to a necessary minimum by adding context awareness to smart glasses. For this, we propose an approach based on human activity recognition, which incorporates features, derived from the user's head- and eye-movement. Towards this goal, we combine an commercial eye-tracker and an IMU to capture eye- and head-movement features of 7 activities performed by 20 participants. From a methodological perspective, we introduce U-HAR, a convolutional network optimized for activity recognition. By applying a few-shot learning, our model reaches an macro-F1-score of 86.59%, allowing us to derive contextual information. Johannes Meyer 0001, Adrian Frank, Thomas Schlebusch, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | A Highly Integrated Ambient Light Robust Eye-Tracking Sensor for Retinal Projection AR Glasses Based on Laser Feedback InterferometryabstractRobust and highly integrated eye-tracking is a key technology to improve resolution of near-eye-display technologies for augmented reality (AR) glasses such as focus-free retinal projection as it enables display enhancements like foveated rendering. Furthermore, eye-tracking sensors enables novel ways to interact with user interfaces of AR glasses, improving thus the user experience compared to other wearables. In this work, we present a novel approach to track the user's eye by scanned laser feedback interferometry sensing. The main advantages over modern video-oculography (VOG) systems are the seamless integration of the eye-tracking sensor and the excellent robustness to ambient light with significantly lower power consumption. We further present an algorithm to track the bright pupil signal captured by our sensor with a significantly lower computational effort compared to VOG systems. We evaluate a prototype to prove the high robustness against ambient light and achieve a gaze accuracy of 1.62\,$^\circ$, which is comparable to other state-of-the-art scanned laser eye-tracking sensors. The outstanding robustness and high integrability of the proposed sensor will pave the way for everyday eye-tracking in consumer AR glasses. Johannes Meyer 0001, Thomas Schlebusch, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Where and What: Driver Attention-based Object DetectionabstractHuman drivers use their attentional mechanisms to focus on critical objects and make decisions while driving. As human attention can be revealed from gaze data, capturing and analyzing gaze information has emerged in recent years to benefit autonomous driving technology. Previous works in this context have primarily aimed at predicting "where" human drivers look at and lack knowledge of "what" objects drivers focus on. Our work bridges the gap between pixel-level and object-level attention prediction. Specifically, we propose to integrate an attention prediction module into a pretrained object detection framework and predict the attention in a grid-based style. Furthermore, critical objects are recognized based on predicted attended-to areas. We evaluate our proposed method on two driver attention datasets, BDD-A and DR(eye)VE. Our framework achieves competitive state-of-the-art performance in the attention prediction on both pixel-level and object-level but is far more efficient (75.3 GFLOPs less) in computation. Yao Rong 0001, Naemi-Rebecca Kassautzki, Wolfgang Fuhl, Enkelejda Kasneci |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | Comparative Analysis of Vehicle-Based and Driver-Based Features for Driver Drowsiness Monitoring by Support Vector MachinesabstractDriver drowsiness is a serious threat to road safety. Most driver monitoring systems (DMSs) already embedded in vehicles to detect drowsiness use vehicle-based features (i.e., measures) computed by outward-facing cameras for lane tracking or steering wheel angle sensors to analyze lane keeping and steering control behavior. Such DMSs are referred to as indirect DMSs as they monitor drowsiness indirectly through driving performance. In this work, we extend this classical technique by using a driver monitoring camera for tracking driver-based features associated with eye blinking behavior and head movements. We refer to DMSs based only on a driver monitoring camera as direct DMSs as they monitor drowsiness directly through observable driver-based behavioral cues. In this work, we conduct a comparative analysis between an indirect and direct DMS. We also combine vehicle-based and driver-based features to examine the potential of a so-called hybrid DMS. To this end, we use a database collected from 70 participants in driving simulator experiments. The comparative analysis is performed by means of the correlation-based feature selection technique and support vector machines independently. With a balanced accuracy of 87.1%, the direct DMS significantly outperforms the indirect DMS, which reaches a balanced accuracy of only 77.9%. The hybrid DMS achieves a slightly better balanced accuracy of 87.7%. This work motivates the development and use of direct or hybrid DMSs to detect driver drowsiness and increase road safety. Mohamed Hedi Baccour, Frauke Driewer, Tim Schäck, Enkelejda Kasneci |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Geopositioned 3D Areas of Interest for Gaze AnalysisabstractTo understand driver’s gaze behavior, the gaze is usually matched to surrounding objects or static areas of interest (AOI) at fixed positions around the car. Full surround object tracking allows for an understanding of the traffic situation. However, because it requires an extensive sensor set and a lot of processing power, it’s not yet broadly available in production cars. The use of static AOIs only requires the addition of eye tracking sensors. They are at fixed positions around the car and can’t adapt to the environment, therefore their usefulness is limited. We propose geopositioned 3D AOIs. With adaptability and the use of a small sensor set, they combine the strengths of both methods. To test 3D AOIs’ capabilities for gaze analysis, a driving simulator study with 74 participants was conducted. We show that 3D AOIs are suitable for driver’s gaze analysis and a promising tool for driver intention prediction. Jan Bickerdt, Jan Sonnenberg, Christian Gollnick, Enkelejda Kasneci |
AutomotiveUI | 4 |
| 2021 | Human Attention in Fine-grained Classification
Yao Rong 0001, Wenjia Xu, Zeynep Akata, Enkelejda Kasneci |
BMVC | 4 |
| 2021 | Digital Transformations of Classrooms in Virtual RealityabstractWith rapid developments in consumer-level head-mounted displays and computer graphics, immersive VR has the potential to take online and remote learning closer to real-world settings. However, the effects of such digital transformations on learners, particularly for VR, have not been evaluated in depth. This work investigates the interaction-related effects of sitting positions of learners, visualization styles of peer-learners and teachers, and hand-raising behaviors of virtual peer-learners on learners in an immersive VR classroom, using eye tracking data. Our results indicate that learners sitting in the back of the virtual classroom may have difficulties extracting information. Additionally, we find indications that learners engage with lectures more efficiently if virtual avatars are visualized with realistic styles. Lastly, we find different eye movement behaviors towards different performance levels of virtual peer-learners, which should be investigated further. Our findings present an important baseline for design decisions for VR classrooms. Hong Gao 0008, Efe Bozkir, Lisa Hasenbein, Jens-Uwe Hahn, Richard Göllner, Enkelejda Kasneci |
CHI | 6 |
| 2021 | Reinforcement Learning for the Privacy Preservation and Manipulation of Eye Tracking Data
Wolfgang Fuhl, Efe Bozkir, Enkelejda Kasneci |
ICANN (4) | 3 |
| 2021 | Weight and Gradient Centralization in Deep Neural Networks
Wolfgang Fuhl, Enkelejda Kasneci |
ICANN (4) | 2 |
| 2021 | States of Confusion: Eye and Head Tracking Reveal Surgeons' Confusion during Arthroscopic SurgeryabstractDuring arthroscopic surgeries, surgeons are faced with challenges like cognitive re-projection of the 2D screen output into the 3D operating site or navigation through highly similar tissue. Training of these cognitive processes takes much time and effort for young surgeons, but is necessary and crucial for their education. In this study we want to show how to recognize states of confusion of young surgeons during an arthroscopic surgery, by looking at their eye and head movements and feeding them to a machine learning model. With an accuracy of over 94% and detection speed of 0.039 seconds, our model is a step towards online diagnostic and training systems for the perceptual-cognitive processes of surgeons during arthroscopic surgeries. Benedikt Hosp, Myat Su Yin, Peter Haddawy, Ratthaphum Watcharopas, Paphon Sa-Ngasoongsong, Enkelejda Kasneci |
ICMI | 6 |
| 2021 | Rotated Ring, Radial and Depth Wise Separable Radial ConvolutionsabstractSimple image rotations significantly reduce the accuracy of deep neural networks. Moreover, training with all possible rotations increases the data set, which also increases the training duration. In this work, we address trainable rotation invariant convolutions as well as the construction of nets, since fully connected layers can only be rotation invariant with a one-dimensional input. On the one hand, we show that our approach is rotationally invariant for different models and on different public data sets. We also discuss the influence of purely rotational invariant features on accuracy. The rotationally adaptive convolution models presented in this work are more computationally intensive than normal convolution models. Therefore, we also present a depth wise separable approach with radial convolution. Wolfgang Fuhl, Enkelejda Kasneci |
IJCNN | 2 |
| 2021 | TEyeD: Over 20 Million Real-World Eye Images with Pupil, Eyelid, and Iris 2D and 3D Segmentations, 2D and 3D Landmarks, 3D Eyeball, Gaze Vector, and Eye Movement TypesabstractWe present TEyeD, the world’s largest unified public data set of eye images taken with head-mounted devices. TEyeD was acquired with seven different head-mounted eye trackers. Among them, two eye trackers were integrated into virtual reality (VR) or augmented reality (AR) devices. The images in TEyeD were obtained from various tasks, including car rides, simulator rides, outdoor sports activities, and daily indoor activities. The data set includes 2D&3D landmarks, semantic segmentation, 3D eyeball annotation and the gaze vector and eye movement types for all images. Landmarks and semantic segmentation are provided for the pupil, iris and eyelids. Video lengths vary from a few minutes to several hours. With more than 20 million carefully annotated images, TEyeD provides a unique, coherent resource and a valuable foundation for advancing research in the field of computer vision, eye tracking and gaze estimation in modern VR and AR applications. Data and code at DOWNLOAD LINK. Wolfgang Fuhl, Gjergji Kasneci, Enkelejda Kasneci |
ISMAR | 3 |
| 2021 | Exploiting Object-of-Interest Information to Understand Attention in VR ClassroomsabstractRecent developments in computer graphics and hardware technology enable easy access to virtual reality headsets along with integrated eye trackers, leading to mass usage of such devices. The immersive experience provided by virtual reality and the possibility to control environmental factors in virtual setups may soon help to create realistic digital alternatives to conventional classrooms. The importance of such settings has become especially evident during the COVID-19 pandemic, forcing many schools and universities to provide the digital teaching. Researchers foresee that such transformations will continue in the future with virtual worlds becoming an integral part of education. Until now, however, students' behaviors in immersive virtual environments have not been investigated in depth. In this work, we study students' attention by exploiting object-of-interests using eye tracking in different classroom manipulations. More specifically, we varied sitting positions of students, visualization styles of virtual avatars, and hand-raising percentages of peer-learners. Our empirical evidence shows that such manipulations play an important role in students' attention towards virtual peer-learners, instructors, and lecture material. This research may contribute to understanding of how visual attention relates to social dynamics in the virtual classroom, including significant considerations for the design of virtual learning spaces. Efe Bozkir, Philipp Stark, Hong Gao 0008, Lisa Hasenbein, Jens-Uwe Hahn, Enkelejda Kasneci, Richard Göllner |
VR | 6 |
| 2021 | Predicting visual perceivability of scene objects through spatio-temporal modeling of retinal receptive fields
David Geisler, Andrew T. Duchowski, Enkelejda Kasneci |
Neurocomputing | 3 |
| 2020 | Training Decision Trees as Replacement for Convolution LayersabstractWe present an alternative layer to convolution layers in convolutional neural networks (CNNs). Our approach reduces the complexity of convolutions by replacing it with binary decisions. Those binary decisions are used as indexes to conditional distributions where each weight represents a leaf in a decision tree. This means that only the indices to the weights need to be determined once, thus reducing the complexity of convolutions by the depth of the output tensor. Index computation is performed by simple binary decisions that require fewer cycles compared to conventionally used multiplications. In addition, we show how convolutions can be replaced by binary decisions. These binary decisions form indices in the conditional distributions and we show how they are used to replace 2D weight matrices as well as 3D weight tensors. These new layers can be trained like convolution layers in CNNs based on the backpropagation algorithm, for which we provide a formalization. Our results on multiple publicly available data sets show that our approach performs similar to conventional neuronal networks. Beyond the formalized reduction of complexity and the improved qualitative performance, we show the runtime improvement empirically compared to convolution layers. Wolfgang Fuhl, Gjergji Kasneci, Wolfgang Rosenstiel, Enkelejda Kasneci |
AAAI | 4 |
| 2020 | Deep semantic gaze embedding and scanpath comparison for expertise classification during OPT viewingabstractModeling eye movement indicative of expertise behavior is decisive in user evaluation. However, it is indisputable that task semantics affect gaze behavior. We present a novel approach to gaze scanpath comparison that incorporates convolutional neural networks (CNN) to process scene information at the fixation level. Image patches linked to respective fixations are used as input for a CNN and the resulting feature vectors provide the temporal and spatial gaze information necessary for scanpath similarity comparison. We evaluated our proposed approach on gaze data from expert and novice dentists interpreting dental radiographs using a local alignment similarity score. Our approach was capable of distinguishing experts from novices with 93% accuracy while incorporating the image semantics. Moreover, our scanpath comparison using image patch features has the potential to incorporate task semantics from a variety of tasks. Nora Castner, Thomas C. Kübler, Katharina Scheiter, Juliane Richter, Thérése Eder, Fabian Hüttig, Constanze Keutel, Enkelejda Kasneci |
ETRA | 8 |
| 2020 | A MinHash approach for fast scanpath classificationabstractThe visual scanpath describes the shift of visual attention over time. Characteristic patterns in the attention shifts allow inferences about cognitive processes, performed tasks, intention, or expertise. To analyse such patterns, the scanpath is often represented as a sequence of symbols that can be used to calculate a similarity score to other scanpaths. However, as the length of the scanpath or the number of possible symbols increases, established methods for scanpath similarity become inefficient, both in terms of runtime and memory consumption. We present a MinHash approach for efficient scanpath similarity calculation. Our approach shows competitive results in clustering and classification of scanpaths compared to established methods such as Needleman-Wunsch, but at a fraction of the required runtime. Furthermore, with time complexity of and constant memory consumption, our approach is ideally suited for real-time operation or analyzing large amounts of data. David Geisler, Nora Castner, Gjergji Kasneci, Enkelejda Kasneci |
ETRA | 4 |
| 2020 | Fully Convolutional Neural Networks for Raw Eye Tracking Data Segmentation, Generation, and ReconstructionabstractIn this paper, we use fully convolutional neural networks for the semantic segmentation of eye tracking data. We also use these networks for reconstruction, and in conjunction with a variational auto-encoder to generate eye movement data. The first improvement of our approach is that no input window is necessary, due to the use of fully convolutional networks and therefore any input size can be processed directly. The second improvement is that the used and generated data is raw eye tracking data (position X, Y and time) without preprocessing. This is achieved by pre-initializing the filters in the first layer and by building the input tensor along the z axis. We evaluated our approach on three publicly available datasets and compare the results to the state of the art. Wolfgang Fuhl, Yao Rong 0001, Enkelejda Kasneci |
ICPR | 3 |
| 2020 | Explainable Online Validation of Machine Learning Models for Practical ApplicationsabstractWe present a reformulation of the regression and classification, which aims to validate the result of a machine learning algorithm. Our reformulation simplifies the original problem and validates the result of the machine learning algorithm using the training data. Since the validation of machine learning algorithms must always be explainable, we perform our experiments with the kNN algorithm as well as with an algorithm based on conditional probabilities, which is proposed in this work. For the evaluation of our approach, three publicly available data sets were used and three classification and two regression problems were evaluated. The presented algorithm based on conditional probabilities is also online capable and requires only a fraction of memory compared to the kNN algorithm. Wolfgang Fuhl, Yao Rong 0001, Thomas Motz, Michael Scheidt, Andreas Hartel, Enkelejda Kasneci |
ICPR | 7 |
| 2020 | Distilling Location Proposals of Unknown Objects through Gaze Information for Human-Robot InteractionabstractSuccessful and meaningful human-robot interaction requires robots to have knowledge about the interaction context - e.g., which objects should be interacted with. Unfortunately, the corpora of interactive objects is - for all practical purposes - infinite. This fact hinders the deployment of robots with pre-trained object-detection neural networks other than in pre-defined scenarios. A more flexible alternative to pre-training is to let a human teach the robot about new objects after deployment. However, doing so manually presents significant usability issues as the user must manipulate the object and communicate the object's boundaries to the robot. In this work, we propose streamlining this process by using automatic object location proposal methods in combination with human gaze to distill pertinent object location proposals. Experiments show that the proposed method 1) increased the precision by a factor of approximately 21 compared to location proposal alone, 2) is able to locate objects sufficiently similar to a state-of-the-art pre-trained deep-learning method (FCOS) without any training, and 3) detected objects that were completely missed by FCOS. Furthermore, the method is able to locate objects for which FCOS was not trained on, which are undetectable for FCOS by definition. Daniel Weber 0003, Thiago Santini, Andreas Zell, Enkelejda Kasneci |
IROS | 4 |
| 2020 | Towards expert gaze modeling and recognition of a user's attention in realtimeabstractOne of the appealing areas of expertise research is devoted to measuring the effectiveness of training programs for novices. With recent progress in eye tracking, gaze-based interaction systems recognize a user’s attention and can direct it accordingly. Moreover, dynamic visualization of an expert gaze model facilitates novice training by guiding the gaze to relevant areas. In addition, the system should be aware of realtime attention to remove an overlay that could occlude relevant information. We use an implementation of subtle gaze direction (SGD) and the simplified scanpath of a dentist to train naive participants in finding anomalies in dental radiographs. We were able to effectively direct user gaze to relevant image features without occluding the area when attention was recognized. Additionally, participants reported that the intervention was helpful for image inspection. The results of the model intervention show minimal improvements in anomaly detection, which is expected of naive subjects. We advocate that the system has the potential to be highly effective for advanced students and trainees with a certain foundation of conceptual knowledge. Nora Castner, Lea Geßler, David Geisler, Fabian Hüttig, Enkelejda Kasneci |
KES | 5 |
| 2020 | Camera-based Driver Drowsiness State Classification Using Logistic Regression ModelsabstractDrowsiness at the wheel is a major problem for traffic road safety. A drowsy driver suffers from decreased vigilance, increased reaction time and degraded decision-making ability, all of which have a huge impact on the driving performance. A driver monitoring system that warns the driver of his or her critical drowsiness state is a worthwhile contribution to traffic road safety. A drowsy driver typically exhibits some observable behaviors, such as eye blinking and head movements, that can be tracked using a camera. In this study, we analyze the potential of eye closure and head rotation signals, provided by a driver camera, to classify the driver's drowsiness state using logistic regression models. This analysis is based on a large dataset collected from 71 subjects in driving simulator experiments. A reliable and independent reference for drowsiness, however, is required in order to perform this analysis. For this purpose, we devise a methodology that merges several drowsiness monitoring approaches to construct a reliable reference for drowsiness. Furthermore, we describe our approach to extract eye blink and head rotation features. Ultimately, we design logistic regression classifiers and combine them using the one-vs-one binarization technique. Our approach achieves a global balanced validation accuracy of 72.7% on a three-class classification problem (awake, questionable and drowsy) by adopting a strict and rigorous evaluation scheme (i.e., leave-one-drive-out cross-validation). Mohamed Hedi Baccour, Frauke Driewer, Tim Schäck, Enkelejda Kasneci |
SMC | 4 |
| 2020 | Attention Flow: End-to-End Joint Attention EstimationabstractThis paper addresses the problem of understanding joint attention in third-person social scene videos. Joint attention is the shared gaze behaviour of two or more individuals on an object or an area of interest and has a wide range of applications such as human-computer interaction, educational assessment, treatment of patients with attention disorders, and many more. Our method, Attention Flow, learns joint attention in an end-to-end fashion by using saliency-augmented attention maps and two novel convolutional attention mechanisms that determine to select relevant features and improve joint attention localization. We compare the effect of saliency maps and attention mechanisms and report quantitative and qualitative results on the detection and localization of joint attention in the VideoCoAtt dataset, which contains complex social scenes. Ömer Sümer, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci |
WACV | 4 |
| 2019 | Assessment of Driver Attention during a Safety Critical Situation in VR to Generate VR-based TrainingabstractCrashes involving pedestrians on urban roads can be fatal. In order to prevent such crashes and provide safer driving experience, adaptive pedestrian warning cues can help to detect risky pedestrians. However, it is difficult to test such systems in the wild, and train drivers using these systems in safety critical situations. This work investigates whether low-cost virtual reality (VR) setups, along with gaze-aware warning cues, could be used for driver training by analyzing driver attention during an unexpected pedestrian crossing on an urban road. Our analyses show significant differences in distances to crossing pedestrians, pupil diameters, and driver accelerator inputs when the warning cues were provided. Overall, there is a strong indication that VR and Head-Mounted-Displays (HMDs) could be used for generating attention increasing driver training packages for safety critical situations. Efe Bozkir, David Geisler, Enkelejda Kasneci |
SAP | 3 |
| 2019 | 500, 000 Images Closer to Eyelid and Pupil Segmentation
Wolfgang Fuhl, Wolfgang Rosenstiel, Enkelejda Kasneci |
CAIP (1) | 3 |
| 2019 | Encodji: encoding gaze data into emoji space for an amusing scanpath classification approach ;)abstractTo this day, a variety of information has been obtained from human eye movements, which holds an imense potential to understand and classify cognitive processes and states - e.g., through scanpath classification. In this work, we explore the task of scanpath classification through a combination of unsupervised feature learning and convolutional neural networks. As an amusement factor, we use an Emoji space representation as feature space. This representation is achieved by training generative adversarial networks (GANs) for unpaired scanpath-to-Emoji translation with a cyclic loss. The resulting Emojis are then used to train a convolutional neural network for stimulus prediciton, showing an accuracy improvement of more than five percentual points compared to the same network trained using solely the scanpath data. As a side effect, we also obtain novel unique Emojis representing each unique scanpath. Our goal is to demonstrate the applicability and potential of unsupervised feature learning to scanpath classification in a humorous and entertaining way. Wolfgang Fuhl, Efe Bozkir, Benedikt Hosp, Nora Castner, David Geisler, Thiago Santini, Enkelejda Kasneci |
ETRA | 7 |
| 2019 | Ferns for area of interest free scanpath classificationabstractScanpath classification can offer insight into the visual strategies of groups such as experts and novices. We propose to use random ferns in combination with saccade angle successions to compare scanpaths. One advantage of our method is that it does not require areas of interest to be computed or annotated. The conditional distribution in random ferns additionally allows for learning angle successions, which do not have to be entirely present in a scanpath. We evaluated our approach on two publicly available datasets and improved the classification accuracy by ≈ 10 and ≈ 20 percent. Wolfgang Fuhl, Nora Castner, Thomas C. Kübler, Rene Alexander Lotz, Wolfgang Rosenstiel, Enkelejda Kasneci |
ETRA | 6 |
| 2019 | Get a grip: slippage-robust and glint-free gaze estimation for real-time pervasive head-mounted eye trackingabstractA key assumption conventionally made by flexible head-mounted eye-tracking systems is often invalid: The eye center does not remain stationary w.r.t. the eye camera due to slippage. For instance, eye-tracker slippage might happen due to head acceleration or explicit adjustments by the user. As a result, gaze estimation accuracy can be significantly reduced. In this work, we propose Grip, a novel gaze estimation method capable of instantaneously compensating for eye-tracker slippage without additional hardware requirements such as glints or stereo eye camera setups. Grip was evaluated using previously collected data from a large scale unconstrained pervasive eye-tracking study. Our results indicate significant slippage compensation potential, decreasing average participant median angular offset by more than 43% w.r.t. a non-slippage-robust gaze estimation method. A reference implementation of Grip was integrated into EyeRecToo, an open-source hardware-agnostic eye-tracking software, thus making it readily accessible for multiple eye trackers (Available at: www.ti.uni-tuebingen.de/perception). Thiago Santini, Diederick Christian Niehorster, Enkelejda Kasneci |
ETRA | 3 |
| 2019 | Predicting Cognitive Load in an Emergency Simulation Based on Behavioral and Physiological MeasuresabstractThe reliable estimation of cognitive load is an integral step towards real-time adaptivity of learning or gaming environments. We introduce a novel and robust machine learning method for cognitive load assessment based on behavioral and physiological measures in a combined within- and cross-participant approach. 47 participants completed different scenarios of a commercially available emergency personnel simulation game realizing several levels of difficulty based on cognitive load. Using interaction metrics, pupil dilation, eye-fixation behavior, and heart rate data, we trained individual, participant-specific forests of extremely randomized trees differentiating between low and high cognitive load. We achieved an average classification accuracy of 72%. We then apply these participant-specific classifiers in a novel way, using similarity between participants, normalization, and relative importance of individual features to successfully achieve the same level of classification accuracy in cross-participant classification. These results indicate that a combination of behavioral and physiological indicators allows for reliable prediction of cognitive load in an emergency simulation game, opening up new avenues for adaptivity and interaction. Tobias Appel, Natalia Sevcenko, Franz Wortha, Katerina Tsarava, Korbinian Moeller, Manuel Ninaus, Enkelejda Kasneci, Peter Gerjets |
ICMI | 7 |
| 2019 | Learning to validate the quality of detected landmarksabstractWe present a new loss function for the validation of image landmarks detected via Convolutional Neural Networks (CNN). The network learns to estimate how accurate its landmark estimation is. This loss function is applicable to all regression-based location estimations and allows the exclusion of unreliable landmarks from further processing. In addition, we formulate a novel batch balancing approach which weights the importance of samples based on their produced loss. This is done by computing a probability distribution mapping on an interval from which samples can be selected using a uniform random selection scheme. We conducted experiments on the 300W, AFLW, and WFLW facial landmark datasets. In the first experiments, the influence of our batch balancing approach is evaluated by comparing it against uniform sampling. In addition, we evaluated the impact of the validation loss on the landmark accuracy based on uniform sampling. The last experiments evaluate the correlation of the validation signal with the landmark accuracy. All experiments were performed for all three datasets. Wolfgang Fuhl, Enkelejda Kasneci |
ICMV | 2 |
| 2019 | Camera-Based Eye Blink Detection Algorithm for Assessing Driver DrowsinessabstractThis paper presents an adaptive camera-based eye blink detection algorithm for assessing the level of drowsiness during driving. The data used in this study were collected from driving simulator experiments using a remote camera. Eye blink detection in the automotive context and for different driver states typically encounters some difficulties. It may be challenging to reliably distinguish between eye blink events and gaze-related eyelid closures, particularly the glances at the dashboard, since both exhibit a similar eyelid movement pattern. In addition, it is difficult to find comparable thresholds due to high inter-individual differences in the palpebral aperture. Furthermore, the blinking behavior is impacted by drowsiness and the blink patterns vary widely, which requires an adaptive algorithm to deal with this intra-individual variability of the blinks. These challenges are considered in the design of the presented blink detection algorithm. This algorithm is based essentially on a threshold for the maximum velocity of the eyelids. This threshold is determined using k-means clustering (k=2)and updated every five minutes of the drive. The accuracy of the algorithm is evaluated based on video labeling. The detection rates demonstrate that the algorithm performs very reliably in both awake and drowsy phases of the driving experiments. Mohamed Hedi Baccour, Frauke Driewer, Enkelejda Kasneci, Wolfgang Rosenstiel |
IV | 3 |
| 2019 | Person Independent, Privacy Preserving, and Real Time Assessment of Cognitive Load using Eye Tracking in a Virtual Reality SetupabstractEye tracking is handled as key enabling technology to VR and AR for multiple reasons, since it not only can help to massively reduce computational costs through gaze-based optimization of graphics and rendering, but also offers a unique opportunity to design gaze-based personalized interfaces and applications. Additionally, the analysis of eye tracking data allows to assess the cognitive load, intentions and actions of the user. In this work, we propose a person-independent, privacy-preserving and gaze-based cognitive load recognition scheme for drivers under critical situations based on previously collected driving data from a driving experiment in VR including a safety critical situation. Based on carefully annotated ground-truth information, we used pupillary information and performance measures (inputs on accelerator, brake, and steering wheel) to train multiple classifiers with the aim of assessing the cognitive load of the driver. Our results show that incorporating eye tracking data into the VR setup allows to predict the cognitive load of the user at a high accuracy above 80%. Beyond the specific setup, the proposed framework can be used in any adaptive and intelligent VR/AR application. Efe Bozkir, David Geisler, Enkelejda Kasneci |
VR | 3 |
| 2018 | Cross-subject workload classification using pupil-related measuresabstractReal-time evaluation of a person's cognitive load can be desirable in many situations. It can be employed to automatically assess or adjust the difficulty of a task, as a safety measure, or in psychological research. Eye-related measures, such as the pupil diameter or blink rate, provide a non-intrusive way to assess the cognitive load of a subject and have therefore been used in a variety of applications. Usually, workload classifiers trained on these measures are highly subject-dependent and transfer poorly to other subjects. We present a novel method to generalize from a set of trained classifiers to new and unknown subjects. We use normalized features and a similarity function to match a new subject with similar subjects, for which classifiers have been previously trained. These classifiers are then used in a weighted voting system to detect workload for an unknown subject. For real-time workload classification, our methods performs at 70.4% accuracy. Higher accuracy of 76.8% can be achieved in an offline classification setting. Tobias Appel, Christian Scharinger, Peter Gerjets, Enkelejda Kasneci |
ETRA | 4 |
| 2018 | Scanpath comparison in medical image reading skills of dental students: distinguishing stages of expertise developmentabstractA popular topic in eye tracking is the difference between novices and experts and their domain-specific eye movement behaviors. However, very little is researched regarding how expertise develops, and more specifically, the developmental stages of eye movement behaviors. Our work compares the scanpaths of five semesters of dental students viewing orthopantomograms (OPTs) with classifiers to distinguish sixth semester through tenth semester students. We used the analysis algorithm SubsMatch 2.0 and the Needleman-Wunsch algorithm. Overall, both classifiers were able distinguish the stages of expertise in medical image reading above chance level. Specifically, it was able to accurately determine sixth semester students with no prior training as well as sixth semester students after training. Ultimately, using scanpath models to recognize gaze patterns characteristic of learning stages, we can provide more adaptive, gaze-based training for students. Nora Castner, Enkelejda Kasneci, Thomas C. Kübler, Katharina Scheiter, Juliane Richter, Thérése Eder, Fabian Hüttig, Constanze Keutel |
ETRA | 2 |
| 2018 | An inconspicuous and modular head-mounted eye trackerabstractState of the art head mounted eye trackers employ glasses like frames, making their usage uncomfortable or even impossible for prescription eyewear users. Nonetheless, these users represent a notable portion of the population (e.g. the Prevent Blindness America organization reports that about half of the USA population use corrective eyewear for refractive errors alone). Thus, making eye tracking accessible for eyewear users is paramount to not only improve usability, but is also key for the ecological validity of eye tracking studies. In this work, we report on a novel approach for eye tracker design in the form of a modular and inconspicuous device that can be easily attached to glasses; for users without glasses, we also provide a 3D printable frame blueprint. Our prototypes include both low cost Commerical Out of The Shelf (COTS) and more expensive Original Equipment manufacturer (OEM) cameras, with sampling rates ranging between 30 and 120 fps and multiple pixel resolutions. Shaharam Eivazi, Thomas C. Kübler, Thiago Santini, Enkelejda Kasneci |
ETRA | 4 |
| 2018 | BORE: boosted-oriented edge optimization for robust, real time remote pupil center detectionabstractUndoubtedly, eye movements contain an immense amount of information, especially when looking to fast eye movements, namely time to the fixation, saccade, and micro-saccade events. While, modern cameras support recording of few thousand frames per second, to date, the majority of studies use eye trackers with the frame rates of about 120 Hz for head-mounted and 250 Hz for remote-based trackers. In this study, we aim to overcome the challenge of the pupil tracking algorithms to perform real time with high speed cameras for remote eye tracking applications. We propose an iterative pupil center detection algorithm formulated as an optimization problem. We evaluated our algorithm on more than 13,000 eye images, in which it outperforms earlier solutions both with regard to runtime and detection accuracy. Moreover, our system is capable of boosting its runtime in an unsupervised manner, thus we remove the need for manual annotation of pupil images. Wolfgang Fuhl, Shahram Eivazi, Benedikt Hosp, Anna Eivazi, Wolfgang Rosenstiel, Enkelejda Kasneci |
ETRA | 6 |
| 2018 | CBF: circular binary features for robust and real-time pupil center detectionabstractModern eye tracking systems rely on fast and robust pupil detection, and several algorithms have been proposed for eye tracking under real world conditions. In this work, we propose a novel binary feature selection approach that is trained by computing conditional distributions. These features are scalable and rotatable, allowing for distinct image resolutions, and consist of simple intensity comparisons, making the approach robust to different illumination conditions as well as rapid illumination changes. The proposed method was evaluated on multiple publicly available data sets, considerably outperforming state-of-the-art methods, and being real-time capable for very high frame rates. Moreover, our method is designed to be able to sustain pupil center estimation even when typical edge-detection-based approaches fail - e.g., when the pupil outline is not visible due to occlusions from reflections or eye lids / lashes. As a consequece, it does not attempt to provide an estimate for the pupil outline. Nevertheless, the pupil center suffices for gaze estimation - e.g., by regressing the relationship between pupil center and gaze point during calibration. Wolfgang Fuhl, David Geisler, Thiago Santini, Tobias Appel, Wolfgang Rosenstiel, Enkelejda Kasneci |
ETRA | 6 |
| 2018 | Development and evaluation of a gaze feedback system integrated into eyetraceabstractA growing field of studies in eye-tracking is the use of gaze data for realtime feedback to the subject. In this work, we present a software system for such experiments and validate it with a visual search task experiment. This system was integrated into an eye tracking analysis tool. Our aim was to improve subject performance in this task by employing saliency features for gaze guidance. This realtime feedback system can be applicable within many realms, such as learning interventions, computer entertainment, or virtual reality. Kai Otto, Nora Castner, David Geisler, Enkelejda Kasneci |
ETRA | 4 |
| 2018 | PuReST: robust pupil tracking for real-time pervasive eye trackingabstractPervasive eye-tracking applications such as gaze-based human computer interaction and advanced driver assistance require real-time, accurate, and robust pupil detection. However, automated pupil detection has proved to be an intricate task in real-world scenarios due to a large mixture of challenges - for instance, quickly changing illumination and occlusions. In this work, we introduce the Pupil Reconstructor with Subsequent Tracking (PuReST), a novel method for fast and robust pupil tracking. The proposed method was evaluated on over 266,000 realistic and challenging images acquired with three distinct head-mounted eye tracking devices, increasing pupil detection rate by 5.44 and 29.92 percentage points while reducing average run time by a factor of 2.74 and 1.1. w.r.t. state-of-the-art 1) pupil detectors and 2) vendor provided pupil trackers, respectively. Overall, PuReST outperformed other methods in 81.82% of use cases. Thiago Santini, Wolfgang Fuhl, Enkelejda Kasneci |
ETRA | 3 |
| 2018 | Eye-Hand Behavior in Human-Robot Shared ManipulationabstractShared autonomy systems enhance people's abilities to perform activities of daily living using robotic manipulators. Recent systems succeed by first identifying their operators' intentions, typically by analyzing the user's joystick input. To enhance this recognition, it is useful to characterize people's behavior while performing such a task. Furthermore, eye gaze is a rich source of information for understanding operator intention. The goal of this paper is to provide novel insights into the dynamics of control behavior and eye gaze in human-robot shared manipulation tasks. To achieve this goal, we conduct a data collection study that uses an eye tracker to record eye gaze during a human-robot shared manipulation activity, both with and without shared autonomy assistance. We process the gaze signals from the study to extract gaze features like saccades, fixations, smooth pursuits, and scan paths. We analyze those features to identify novel patterns of gaze behaviors and highlight where these patterns are similar to and different from previous findings about eye gaze in human-only manipulation tasks. The work described in this paper lays a foundation for a model of natural human eye gaze in human-robot shared manipulation. Reuben M. Aronson, Thiago Santini, Thomas C. Kübler, Enkelejda Kasneci, Siddhartha S. Srinivasa, Henny Admoni |
HRI | 4 |
| 2018 | Modeling Cognitive Processes from Multimodal SignalsabstractMultimodal signals allow us to gain insights into internal cognitive processes of a person, for example: speech and gesture analysis yields cues about hesitations, knowledgeability, or alertness, eye tracking yields information about a person's focus of attention, task, or cognitive state, EEG yields information about a person's cognitive load or information appraisal. Capturing cognitive processes is an important research tool to understand human behavior as well as a crucial part of a user model to an adaptive interactive system such as a robot or a tutoring system. As cognitive processes are often multifaceted, a comprehensive model requires the combination of multiple complementary signals. In this workshop at the ACM International Conference on Multimodal Interfaces (ICMI) conference in Boulder, Colorado, USA, we discussed the state-of-the-art in monitoring and modeling cognitive processes from multi-modal signals. Felix Putze, Jutta Hild, Akane Sano, Enkelejda Kasneci, Erin Treacy Solovey, Tanja Schultz |
ICMI | 4 |
| 2018 | Real-time 3D Glint Detection in Remote Eye Tracking Based on Bayesian InferenceabstractAs human gaze provides information on our cognitive states, actions, and intentions, gaze-based interaction has the potential to enable a fluent and natural human-robot collaboration. In this work, we focus on reliable gaze estimation in remote eye tracking based on calibration-free methods. Although these methods work well in controlled settings, they fail when illumination conditions change or other objects induce noise. We propose a novel, adaptive method based on a probabilistic model, which reliably detects glints from stereo images and evaluate our method using a data set that contains different challenges with regarding to light and reflections. David Geisler, Dieter Fox, Enkelejda Kasneci |
ICRA | 3 |
| 2018 | PuRe: Robust pupil detection for real-time pervasive eye tracking
Thiago Santini, Wolfgang Fuhl, Enkelejda Kasneci |
Comput. Vis. Image Underst. | 3 |
| 2017 | CalibMe: Fast and Unsupervised Eye Tracker Calibration for Gaze-Based Pervasive Human-Computer InteractionabstractAs devices around us become smart, our gaze is poised to become the next frontier of human-computer interaction (HCI). State-of-the-art mobile eye tracker systems typically rely on eye-model-based gaze estimation approaches, which do not require a calibration. However, such approaches require specialized hardware (e.g., multiple cameras and glint points), can be significantly affected by glasses, and, thus, are not fit for ubiquitous gaze-based HCI. In contrast, regression-based gaze estimations are straightforward approaches requiring solely one eye and one scene camera but necessitate a calibration. Therefore, a fast and accurate calibration is a key development to enable ubiquitous gaze-based HCI. In this paper, we introduce CalibMe, a novel method that exploits collection markers (automatically detected fiducial markers) to allow eye tracker users to gather a large array of calibration points, remove outliers, and automatically reserve evaluation points in a fast and unsupervised manner. The proposed approach is evaluated against a nine-point calibration method, which is typically used due to its relatively short calibration time and adequate accuracy. CalibMe reached a mean angular error of 0.59 (0=0.23) in contrast to 0.82 (0=0.15) for a nine-point calibration, attesting for the efficacy of the method. Moreover, users are able to calibrate the eye tracker anywhere and independently in - 10 s using a cellphone to display the collection marker. Thiago Santini, Wolfgang Fuhl, Enkelejda Kasneci |
CHI | 3 |
| 2017 | Fast and Robust Eyelid Outline and Aperture Detection in Real-World ScenariosabstractThe correct identification of the eyelids and its aperture provide essential data to infer a subject's mental state (e.g., vigilance, fatigue, and drowsiness) and to validate or reduce the search space of other eye features (e.g., pupil, and iris). This knowledge can be used not only to improve many applications, such as eye tracking and iris recognition, but also to derive information about the user (such as, the take-over readiness of the driver in the automated driving context). In this paper, we propose a computervision-based approach to eyelids identification and aperture estimation. Evaluation was performed on an existing data set from the literature as well as on a new data set introduced in this work. The new data set contains 4000 hand-labeled eye images from 11 subjects driving in a city, these contain several challenges such as reflections, makeup, wrinkles, blinks, and changing illumination. The proposed method outperformed state-of-the-art methods by up to 16.11 percentage points in terms of average similarity to the hand-labeled eyelid outline (from 34px to 12px) and 21.7 pixels (or 7.53% of the eye image height) in terms of average eyelid aperture estimation error. The proposed method implementation runs in real time even on a single core (7ms) and is available, together with the new data set, at http://www.ti.uni-tuebingen.de/Eyelid-detection.2007.0.html. Wolfgang Fuhl, Thiago Santini, Enkelejda Kasneci |
WACV | 3 |
| 2016 | On the necessity of adaptive eye movement classification in conditionally automated driving scenariosabstractAlgorithms for eye movement classification are separated into threshold-based and probabilistic methods. While the parameters of static threshold-based algorithms usually need to be chosen for the particular task (task-individual), the probabilistic methods were introduced to meet the challenge of adjusting automatically to multiple individuals with different viewing behaviors (inter-individual). In the context of conditionally automated driving, especially while the driver is performing various secondary tasks, these two requirements of task- and inter-individuality fuse to an even greater challenge. This paper shows how the combination of task- and inter-individual differences influences the viewing behavior of a driver during conditionally automated drives and that state-of-the-art algorithms are not able to sufficiently adapt to these variances. To approach this challenge, an extended version of a Bayesian online learning algorithm is introduced, which is not only able to adapt its parameters to upcoming variances in the viewing behavior, but also has real-time capability and lower computational overhead. The proposed approach is applied to a large-scale driving simulator study with 74 subjects performing secondary tasks while driving in an automated setting. The results show that the eye movement behavior of drivers performing different secondary tasks varies significantly while remaining approximately consistent for idle drivers. Furthermore, the data shows that only a few of the parameters used for describing the eye movement behavior are responsible for these significant variations indicating that it is not necessary to learn all parameters in an online-fashion. Christian Braunagel, David Geisler, Wolfgang Stolzmann, Wolfgang Rosenstiel, Enkelejda Kasneci |
ETRA | 5 |
| 2016 | ElSe: ellipse selection for robust pupil detection in real-world environmentsabstractFast and robust pupil detection is an essential prerequisite for video-based eye-tracking in real-world settings. Several algorithms for image-based pupil detection have been proposed in the past, their applicability, however, is mostly limited to laboratory conditions. In real-world scenarios, automated pupil detection has to face various challenges, such as illumination changes, reflections (on glasses), make-up, non-centered eye recording, and physiological eye characteristics. We propose ElSe, a novel algorithm based on ellipse evaluation of a filtered edge image. We aim at a robust, inexpensive approach that can be integrated in embedded architectures, e.g., driving. The proposed algorithm was evaluated against four state-of-the-art methods on over 93,000 hand-labeled images from which 55,000 are new eye images contributed by this work. On average, the proposed method achieved a 14.53% improvement on the detection rate relative to the best state-of-the-art performer. Algorithm and data sets are available for download: ftp://[email protected] (password:eyedata). Wolfgang Fuhl, Thiago Santini, Thomas C. Kübler, Enkelejda Kasneci |
ETRA | 4 |
| 2016 | Rendering refraction and reflection of eyeglasses for synthetic eye tracker imagesabstractWhile for the evaluation of robustness of eye tracking algorithms the use of real-world data is essential, there are many applications where simulated, synthetic eye images are of advantage. They can generate labelled ground-truth data for appearance based gaze estimation algorithms or enable the development of model based gaze estimation techniques by showing the influence on gaze estimation error of different model factors that can then be simplified or extended. We extend the generation of synthetic eye images by a simulation of refraction and reflection for eyeglasses. On the one hand this allows for the testing of pupil and glint detection algorithms under different illumination and reflection conditions, on the other hand the error of gaze estimation routines can be estimated in conjunction with different eyeglasses. We show how a polynomial function fitting calibration performs equally well with and without eyeglasses, and how a geometrical eye model behaves when exposed to glasses. Thomas C. Kübler, Tobias Rittig, Enkelejda Kasneci, Judith Ungewiss, Christina Krauss |
ETRA | 3 |
| 2016 | Bayesian identification of fixations, saccades, and smooth pursuitsabstractSmooth pursuit eye movements provide meaningful insights and information on subject's behavior and health and may, in particular situations, disturb the performance of typical fixation/saccade classification algorithms. Thus, an automatic and efficient algorithm to identify these eye movements is paramount for eye-tracking research involving dynamic stimuli. In this paper, we propose the Bayesian Decision Theory Identification (I-BDT) algorithm, a novel algorithm for ternary classification of eye movements that is able to reliably separate fixations, saccades, and smooth pursuits in an online fashion, even for low-resolution eye trackers. The proposed algorithm is evaluated on four datasets with distinct mixtures of eye movements, including fixations, saccades, as well as straight and circular smooth pursuits; data was collected with a sample rate of 30 Hz from six subjects, totaling 24 evaluation datasets. The algorithm exhibits high and consistent performance across all datasets and movements relative to a manual annotation by a domain expert (recall: μ = 91.42%, σ = 9.52%; precision: μ = 95.60%, σ = 5.29%; specificity μ = 95.41%, σ = 7.02%) and displays a significant improvement when compared to I-VDT, an state-of-the-art algorithm (recall: μ = 87.67%, σ = 14.73%; precision: μ = 89.57%, σ = 8.05%; specificity μ = 92.10%, σ = 11.21%). Algorithm implementation and annotated datasets are openly available at www.ti.uni-tuebingen.de/perception Thiago Santini, Wolfgang Fuhl, Thomas C. Kübler, Enkelejda Kasneci |
ETRA | 4 |
| 2016 | Pupil detection for head-mounted eye tracking in the wild: an evaluation of the state of the art
Wolfgang Fuhl, Marc Tonsen, Andreas Bulling, Enkelejda Kasneci |
Mach. Vis. Appl. | 4 |
| 2015 | ExCuSe: Robust Pupil Detection in Real-World Scenarios
Wolfgang Fuhl, Thomas C. Kübler, Katrin Sippel, Wolfgang Rosenstiel, Enkelejda Kasneci |
CAIP (1) | 5 |
| 2014 | The applicability of probabilistic methods to the online recognition of fixations and saccades in dynamic scenesabstractIn many applications involving scanpath analysis, especially when dynamic scenes are viewed, consecutive fixations and saccades, have to be identified and extracted from raw eye-tracking data in an online fashion. Since probabilistic methods can adapt not only to the individual viewing behavior, but also to changes in the scene, they are best suited for such tasks. Enkelejda Kasneci, Gjergji Kasneci, Thomas C. Kübler, Wolfgang Rosenstiel |
ETRA | 1 |
| 2014 | SubsMatch: scanpath similarity in dynamic scenes based on subsequence frequenciesabstractThe analysis of visual scanpaths, i.e., series of fixations and saccades, in complex dynamic scenarios is highly challenging and usually performed manually. We propose SubsMatch, a scanpath comparison algorithm for dynamic, interactive scenarios based on the frequency of repeated gaze patterns. Instead of measuring the gaze duration towards a semantic target object (which would be hard to label in dynamic scenes), we examine the frequency of attention shifts and exploratory eye movements. SubsMatch was evaluated on highly dynamic data from a driving experiment to identify differences between scanpaths of subjects who failed a driving test and subjects who passed. Thomas C. Kübler, Enkelejda Kasneci, Wolfgang Rosenstiel |
ETRA | 2 |
| 2014 | Gaze guidance for the visually impairedabstractVisual perception is perhaps the most important sensory input. During driving, about 90% of the relevant information is related to the visual input [Taylor 1982]. However, the quality of visual perception decreases with age, mainly related to a reduce in the visual acuity or in consequence of diseases affecting the visual system. Amongst the most severe types of visual impairments are visual field defects (areas of reduced perception in the visual field), which occur as a consequence of diseases affecting the brain, e.g., stroke, brain injury, trauma, or diseases affecting the optic nerve, e.g., glaucoma. Due to demographic aging, the number of people with such visual impairments is expected to rise [Kasneci 2013]. Since persons suffering from visual impairments may overlook hazardous objects, they are prohibited from driving. This, however, leads to a decrease in quality of life, mobility, and participation in social life. Several studies have shown that some patients show a safe driving behavior despite their visual impairment by performing effective visual exploration, i.e., adequate eye and head movements (e.g., towards their visual field defect [Kasneci et al. 2014b]). Thus, a better understanding of visual perception mechanisms, i.e., of why and how we attend certain parts of our environment while "ignoring" others, is a key question to helping visually impaired persons in complex, real-life tasks, such as driving a car. Thomas C. Kübler, Enkelejda Kasneci, Wolfgang Rosenstiel |
ETRA | 2 |
| 2013 | Online Classification of Eye Tracking Data for Automated Analysis of Traffic Hazard Perception
Enkelejda Kasneci, Thomas C. Kübler, Gjergji Kasneci, Wolfgang Rosenstiel, Martin Bogdan |
ICANN | 1 |
| 2012 | Bayesian online clustering of eye movement dataabstractThe task of automatically tracking the visual attention in dynamic visual scenes is highly challenging. To approach it, we propose a Bayesian online learning algorithm. As the visual scene changes and new objects appear, based on a mixture model, the algorithm can identify and tell visual saccades (transitions) from visual fixation clusters (regions of interest). The approach is evaluated on real-world data, collected from eye-tracking experiments in driving sessions. Enkelejda Kasneci, Gjergji Kasneci, Wolfgang Rosenstiel, Martin Bogdan |
ETRA | 1 |
| 2011 | Vishnoo - An open-source software for vision researchabstractThe visual input is perhaps the most important sensory information. Understanding its mechanisms as well as the way visual attention arises could be highly beneficial for many tasks involving the analysis of users' interaction with their environment. We present Vishnoo (Visual Search Examination Tool), an integrated framework that combines configurable search tasks with gaze tracking capabilities, thus enabling the analysis of both, the visual field and the visual attention. Our user studies underpin the viability of such a platform. Vishnoo is an open-source software and is available for download at http://www.vishnoo.de/ Enkelejda Kasneci, Thomas C. Kübler, Jörg Peter 0001, Wolfgang Rosenstiel, Martin Bogdan |
CBMS | 1 |