Gabriel J. Diaz

dblp:176/9131 · also Gabriel Jacob Diaz · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-1812-017XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
abstract
Eye feature extraction from event-based data streams can be performed efficiently and with low energy consumption, offering great utility to real-world eye tracking pipelines. However, few eye feature extractors are designed to handle sudden changes in event density caused by the changes between gaze behaviors that vary in their kinematics, leading to degraded prediction performance. In this work, we address this problem by introducing the adaptive inference state space model (AISSM), a novel architecture for feature extraction that is capable of dynamically adjusting the relative weight placed on current versus recent information. This relative weighting is determined via estimates of the signal-to-noise ratio and event density produced by a complementary dynamic confidence network. Lastly, we craft and evaluate a novel learning technique that improves training efficiency. Experimental results demonstrate that the AISSM system outperforms state-of-the-art models for event-based eye feature extraction.
Viet Dung Nguyen, Mobina Ghorbaninejad, Chengyi Ma, Reynold J. Bailey, Gabriel J. Diaz, Alexander Fix, Ryan J. Suess, Alexander Ororbia
ETRA5
2026 PACMHCI V10, N3, June 2026 Editorial ETRA000
Nora Castner, Brendan David-John, Gabriel J. Diaz, Carlos Hitoshi Morimoto
Proc. ACM Hum. Comput. Interact.3
2022 EllSeg-Gen, towards Domain Generalization for Head-Mounted Eyetracking
abstract
The study of human gaze behavior in natural contexts requires algorithms for gaze estimation that are robust to a wide range of imaging conditions. However, algorithms often fail to identify features such as the iris and pupil centroid in the presence of reflective artifacts and occlusions. Previous work has shown that convolutional networks excel at extracting gaze features despite the presence of such artifacts. However, these networks often perform poorly on data unseen during training. This work follows the intuition that jointly training a convolutional network with multiple datasets learns a generalized representation of eye parts. We compare the performance of a single model trained with multiple datasets against a pool of models trained on individual datasets. Results indicate that models tested on datasets in which eye images exhibit higher appearance variability benefit from multiset training. In contrast, dataset-specific models generalize better onto eye images with lower appearance variability.
Rakshit Sunil Kothari, Reynold J. Bailey, Christopher Kanan, Jeff B. Pelz, Gabriel J. Diaz
Proc. ACM Hum. Comput. Interact.5
2022 Temporal RIT-Eyes: From real infrared eye-images to synthetic sequences of gaze behavior
abstract
Current methods for segmenting eye imagery into skin, sclera, pupil, and iris cannot leverage information about eye motion. This is because the datasets on which models are trained are limited to temporally non-contiguous frames. We present Temporal RIT-Eyes, a Blender pipeline that draws data from real eye videos for the rendering of synthetic imagery depicting natural gaze dynamics. These sequences are accompanied by ground-truth segmentation maps that may be used for training image-segmentation networks. Temporal RIT-Eyes relies on a novel method for the extraction of 3D eyelid pose (top and bottom apex of eyelids/eyeball boundary) from raw eye images for the rendering of gaze-dependent eyelid pose and blink behavior. The pipeline is parameterized to vary in appearance, eye/head/camera/illuminant geometry, and environment settings (indoor/outdoor). We present two open-source datasets of synthetic eye imagery: sGiW is a set of synthetic-image sequences whose dynamics are modeled on those of the Gaze in Wild dataset, and sOpenEDS2 is a series of temporally non-contiguous eye images that approximate the OpenEDS-2019 dataset. We also analyze and demonstrate the quality of the rendered dataset qualitatively and show significant overlap between latent-space representations of the source and the rendered datasets.
Aayush K. Chaudhary, Nitinraj Nair, Reynold J. Bailey, Jeff B. Pelz, Sachin S. Talathi, Gabriel J. Diaz
IEEE Trans. Vis. Comput. Graph.6
2021 EllSeg: An Ellipse Segmentation Framework for Robust Gaze Tracking
abstract
Ellipse fitting, an essential component in pupil or iris tracking based video oculography, is performed on previously segmented eye parts generated using various computer vision techniques. Several factors, such as occlusions due to eyelid shape, camera position or eyelashes, frequently break ellipse fitting algorithms that rely on well-defined pupil or iris edge segments. In this work, we propose training a convolutional neural network to directly segment entire elliptical structures and demonstrate that such a framework is robust to occlusions and offers superior pupil and iris tracking performance (at least 10% and 24% increase in pupil and iris center detection rate respectively within a two-pixel error margin) compared to using standard eye parts segmentation for multiple publicly available synthetic segmentation datasets.
Rakshit Sunil Kothari, Aayush K. Chaudhary, Reynold J. Bailey, Jeff B. Pelz, Gabriel J. Diaz
IEEE Trans. Vis. Comput. Graph.5
2020 RIT-Eyes: Rendering of near-eye images for eye-tracking applications
abstract
Deep neural networks for video-based eye tracking have demonstrated resilience to noisy environments, stray reflections, and low resolution. However, to train these networks, a large number of manually annotated images are required. To alleviate the cumbersome process of manual labeling, computer graphics rendering is employed to automatically generate a large corpus of annotated eye images under various conditions. In this work, we introduce a synthetic eye image generation platform that improves upon previous work by adding features such as an active deformable iris, an aspherical cornea, retinal retro-reflection, gaze-coordinated eye-lid deformations, and blinks. To demonstrate the utility of our platform, we render images reflecting the represented gaze distributions inherent in two publicly available datasets, NVGaze and OpenEDS. We also report on the performance of two semantic segmentation architectures (SegNet and RITnet) trained on rendered images and tested on the original datasets.
Nitinraj Nair, Rakshit Sunil Kothari, Aayush K. Chaudhary, Zhizhuo Yang, Gabriel J. Diaz, Jeff B. Pelz, Reynold J. Bailey
SAP5
2018 Characterizing the Temporal Dynamics of Information in Visually Guided Predictive Control Using LSTM Recurrent Neural Networks
Kamran Binaee, Anna Starynska, Jeff B. Pelz, Christopher Kanan, Gabriel J. Diaz
CogSci5
2016 Binocular eye tracking calibration during a virtual ball catching task using head mounted display
abstract
When tracking the eye movements of an active observer, the quality of the tracking data is continuously affected by physical shifts of the eye-tracker on an observers head. This is especially true for eye-trackers integrated within virtual-reality (VR) helmets. These configurations modify the weight and inertia distribution well beyond that of the eye-tracker alone. Despite the continuous nature of this degradation, it is common practice for calibration procedures to establish eye-to-screen mappings, fixed over the time-course of an experiment. Even with periodic recalibration, data quality can quickly suffer due to head motion. Here, we present a novel post-hoc calibration method that allows for continuous temporal interpolation between discrete calibration events. Analysis focuses on the comparison of fixed vs. continuous calibration schemes and their effects upon the quality of a binocular gaze data to virtual targets, especially with respect to depth. Calibration results were applied to binocular eye tracking data from a VR ball catching task and improved the tracking accuracy especially in the dynamic case.
Kamran Binaee, Gabriel J. Diaz, Jeff B. Pelz, Flip Phillips
SAP2
2016 Novel apparatus for investigation of eye movements when walking in the presence of 3D projected obstacles
abstract
The human gait cycle is incredibly efficient and stable largely because of the use of advance visual information to make intelligent selections of heading direction, foot placement, gait dynamics, and posture when faced with terrain complexity [Patla and Vickers 1997; Patla and Vickers 2003; Matthis and Fajen 2013; Matthis and Hayhoe 2015]. This is behaviorally demonstrated by a coupling between saccades and foot placement.
Rakshit Sunil Kothari, Kamran Binaee, Jonathan S. Matthis, Reynold J. Bailey, Gabriel J. Diaz
ETRA5
2016 3D gaze point localization and visualization using LiDAR-based 3D reconstructions
abstract
We present a novel pipeline for localizing a free roaming eye tracker within a LiDAR-based 3D reconstructed scene with high levels of accuracy. By utilizing a combination of reconstruction algorithms that leverage the strengths of global versus local capture methods and user-assisted refinement, we reduce drift errors associated with Dense-SLAM techniques. Our framework supports region-of-interest (ROI) annotation and gaze statistics generation and the ability to visualize gaze in 3D from an immersive first person or third person perspective. This approach gives unique insights into viewers' problem solving and search task strategies and has high applicability in complex static environments such as crime scenes.
James Pieszala, Gabriel J. Diaz, Jeff B. Pelz, Jacqueline Speir, Reynold J. Bailey
ETRA2