Xucong Zhang

dblp:119/8485 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
7since 2021 · last 2026
0000-0002-8368-3542ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Face, body and person analysis · 60% Generative modeling · 11% 3D vision · 11%
Human-computer interaction and pervasive computing
6 papers
Wearable and physiological sensing · 72% Ubiquitous computing and smart environments · 10% Interaction techniques and input · 9%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 38% Computer animation and physical simulation · 32% Rendering · 18%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
gaze estimation
3.292023
GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields · CVPR 2023
Gaze Estimation by Exploring Two-Eye Asymmetry · IEEE Trans. Image Process. 2020
Self-Learning Transformations for Improving Gaze and Head Redirection · NeurIPS 2020
Computer vision › Face, body and person analysis › gaze estimation
appearance-based gaze estimation
1.142019
MPIIGaze: Real-World Dataset and Deep Appearance-Based Gaze Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Appearance-Based Gaze Estimation via Evaluation-Guided Asymmetric Regression · ECCV (14) 2018
Rendering of Eyes for Eye-Shape Registration and Gaze Estimation · ICCV 2015
Wearable and physiological sensing › eye tracking
gaze-based interaction
1.032019
Evaluation of Appearance-Based Methods and Implications for Gaze-Based Applications · CHI 2019
Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices · CHI 2018
Everyday Eye Contact Detection Using Unsupervised Gaze Target Discovery · UIST 2017
Computer vision › Face, body and person analysis › gaze analysis
gaze target detection
0.912025
GazeHTA: End-to-End Gaze Target Detection with Head-Target Association · ICRA 2025
Computer animation and physical simulation
motion synthesis
0.912025
SpeechCAT: Cross-Attentive Transformer for Audio to Motion Generation · HRI 2025
Computer vision › Face, body and person analysis › gaze estimation
gaze redirection
0.712023
GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields · CVPR 2023
Computer vision › 3D vision
neural radiance field
0.712023
GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields · CVPR 2023
Wearable and physiological sensing
eye tracking
0.632020
Towards End-to-End Video-Based Eye-Tracking · ECCV (12) 2020
Evaluation of Appearance-Based Methods and Implications for Gaze-Based Applications · CHI 2019
AggreGaze: Collective Estimation of Audience Attention on Public Displays · UIST 2016
Computer vision › 3D vision › pose estimation
3d hand pose estimation
0.512021
Self-Supervised 3D Hand Pose Estimation from monocular RGB via Contrastive Learning · ICCV 2021
Machine learning › Representation and self-supervised learning
contrastive learning
0.512021
Self-Supervised 3D Hand Pose Estimation from monocular RGB via Contrastive Learning · ICCV 2021
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
equivariant contrastive learning
0.512021
Self-Supervised 3D Hand Pose Estimation from monocular RGB via Contrastive Learning · ICCV 2021
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.512021
Self-Supervised 3D Hand Pose Estimation from monocular RGB via Contrastive Learning · ICCV 2021
Machine learning › Generative modeling › face synthesis
controllable face generation
0.412020
Self-Learning Transformations for Improving Gaze and Head Redirection · NeurIPS 2020
Machine learning › Generative modeling › image generation
conditional image synthesis
0.412019
Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networks · ICCV 2019
Machine learning › Generative modeling
generative adversarial network
0.412019
Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networks · ICCV 2019
Visual content generation and editing
face editing
0.412019
Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networks · ICCV 2019
Visual content generation and editing › face editing
gaze redirection
0.412019
Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networks · ICCV 2019
Wearable and physiological sensing › eye tracking › gaze estimation
appearance-based gaze estimation
0.412019
Evaluation of Appearance-Based Methods and Implications for Gaze-Based Applications · CHI 2019
Wearable and physiological sensing › eye tracking
gaze estimation
0.322025
GazeHTA: End-to-End Gaze Target Detection with Head-Target Association · ICRA 2025
AggreGaze: Collective Estimation of Audience Attention on Public Displays · UIST 2016
Interaction techniques and input
eye contact detection
0.312017
Everyday Eye Contact Detection Using Unsupervised Gaze Target Discovery · UIST 2017
Visual content generation and editing
avatar generation
0.312025
SpeechCAT: Cross-Attentive Transformer for Audio to Motion Generation · HRI 2025
Human-robot interaction › human behavior modeling
attention estimation
0.312025
GazeHTA: End-to-End Gaze Target Detection with Head-Target Association · ICRA 2025
Virtual and augmented reality
eye tracking
0.212016
Combining eye tracking with optimizations for lens astigmatism in modern wide-angle HMDs · VR 2016
Rendering › perceptual rendering
foveated rendering
0.212016
Combining eye tracking with optimizations for lens astigmatism in modern wide-angle HMDs · VR 2016
Rendering
rendering optimization
0.212016
Combining eye tracking with optimizations for lens astigmatism in modern wide-angle HMDs · VR 2016
Ubiquitous computing and smart environments
public displays
0.212016
AggreGaze: Collective Estimation of Audience Attention on Public Displays · UIST 2016
Machine learning › Learning paradigms
multi-task learning
0.212013
Robust Multi-resolution Pedestrian Detection in Traffic Scenes · CVPR 2013
Computer vision › Image recognition and object detection
pedestrian detection
0.212013
Robust Multi-resolution Pedestrian Detection in Traffic Scenes · CVPR 2013
Computer vision › Face, body and person analysis
head pose estimation
0.112020
ETH-XGaze: A Large Scale Dataset for Gaze Estimation Under Extreme Head Pose and Gaze Variation · ECCV (5) 2020
Machine learning › Deep learning architectures and training
convolutional neural network
0.112018
Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices · CHI 2018

Methods — techniques the papers use, named apart from their topics

end-to-end learning · 2.6transformer · 0.9encoder-decoder · 0.9cross-attention · 0.9asymmetric regression · 0.8volume compositing · 0.7two-stream architecture · 0.7appearance-based gaze estimation · 0.5self-supervised learning · 0.5contrastive learning · 0.5self-learning · 0.4evaluation network · 0.4disentanglement · 0.4deep learning · 0.4perceptual loss · 0.4generative adversarial network · 0.4cycle consistency loss · 0.4cross-device training · 0.3
YearPublicationVenuePosition
2026 UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training
abstract
Despite decades of research on data collection and model architectures, current gaze estimation models encounter significant challenges in generalizing across diverse data domains. Recent advances in self-supervised pre-training have demonstrated remarkable generalization across various vision tasks. However, their effectiveness in gaze estimation remains unexplored. We propose UniGaze, which for the first time leverages large-scale in-the-wild facial datasets for gaze estimation through self-supervised pre-training. Through systematic investigation, we clarify critical factors essential for effective pre-training in gaze estimation. Our experiments reveal that self-supervised approaches designed for semantic tasks fail when applied to gaze estimation, while our carefully designed pre-training pipeline consistently improves cross-domain performance. Through comprehensive experiments of the challenging cross-dataset evaluations and novel protocols, including leave-one-dataset-out and joint-dataset settings, we demonstrate that UniGaze significantly improves generalization across multiple data domains while minimizing reliance on costly labeled data. Source code and model are available at https://github.com/ut-vision/UniGaze.
Jiawei Qin, Xucong Zhang, Yusuke Sugano
WACV2
2025 SpeechCAT: Cross-Attentive Transformer for Audio to Motion Generation
abstract
Audio-to-motion generation is an important task with applications in virtual avatar creation for XR systems and intelligent robot control in daily life scenarios. However, most existing motion generation methods rely on a single encoder-decoder architecture to model all body parts simultaneously, which limits their ability to capture the diverse and complex motions exhibited by humans. In this paper, we propose a novel method, SpeechCAT, that employs three separate encoder-decoder modules to individually model the motions of the face, body, and hands. To capture the relationships and synchronization among these body parts, we introduce a cross-attention mechanism to effectively learn their correlations. SpeechCAT ensures sufficient capacity to model the unique characteristics of each body part while preserving the coherence between them. Our experimental results demonstrate the superiority of SpeechCAT over baseline methods, highlighting its effectiveness in generating diverse, realistic, and synchronized motions with face, body, and hand parts.
Sebastian Deaconu, Xiangwei Shi, Thomas Markhorst, Jouh Yeong Chew, Xucong Zhang
HRI5
2025 GazeHTA: End-to-End Gaze Target Detection with Head-Target Association
Zhiyi Lin 0003, Jouh Yeong Chew, Jan C. van Gemert, Xucong Zhang
ICRA4
2025 Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
abstract
Human intention detection with hand motion prediction is critical to drive the upper-extremity assistive robots in neurorehabilitation applications. However, the traditional methods relying on physiological signal measurement are restrictive and often lack environmental context. We propose a novel approach that predicts future sequences of both hand poses and joint positions. This method integrates gaze information, historical hand motion sequences, and environmental object data, adapting dynamically to the assistive needs of the patient without prior knowledge of the intended object for grasping. Specifically, we use a vector-quantized variational autoencoder for robust hand pose encoding with an autoregressive generative transformer for effective hand motion sequence prediction. We demonstrate the usability of these novel techniques in a pilot study with healthy subjects. To train and evaluate the proposed method, we collect a dataset consisting of various types of grasp actions on different objects from multiple subjects. Through extensive experiments, we demonstrate that the proposed method can successfully predict sequential hand movement. Especially, the gaze information shows significant enhancements in prediction capabilities, particularly with fewer input frames, highlighting the potential of the proposed method for real-world applications.
Xucong Zhang, Arno H. A. Stienen
IROS2
2025 Resource-efficient Gaze Estimation via Frequency-domain Multi-task Contrastive Learning
abstract
Gaze estimation is of great importance to many scientific fields and daily applications, ranging from fundamental research in cognitive psychology to attention-aware systems. While recent advancements in deep learning have led to highly accurate gaze estimation systems, these solutions often come with high computational costs and depend on large-scale labeled gaze data for supervised learning, posing significant practical challenges. To move beyond these limitations, we present EfficientGaze, a resource-efficient framework for gaze representation learning. We introduce the frequency-domain gaze estimation, which exploits the feature extraction capability and the spectral compaction property of discrete cosine transform to substantially reduce the computational cost of gaze estimation systems for both calibration and inference. Moreover, to overcome the data labeling hurdle, we design a novel multi-task gaze-aware contrastive learning framework to learn gaze representations that are generic across subjects in an unsupervised manner. Our evaluation on two gaze estimation datasets demonstrates that EfficientGaze achieves comparable gaze estimation performance to existing supervised learning-based approaches, while enabling up to 6.80 times and 1.67 times speedup in system calibration and gaze estimation, respectively.
Lingyu Du, Xucong Zhang, Guohao Lan
ACM Trans. Sens. Networks2
2023 GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields
abstract
We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eyeballs are separate 3D structures that move in a coordinated yet independent fashion. Our method leverages recent advancements in conditional image-based neural radiance fields and proposes a two-stream architecture that predicts volumetric features for the face and eye regions separately. Rigidly transforming the eye features via a 3D rotation matrix provides fine-grained control over the desired gaze angle. The final, redirected image is then attained via differentiable volume compositing. Our experiments show that this architecture outperforms naively conditioned NeRF baselines as well as previous state-of-the-art 2D gaze redirection methods in terms of redirection accuracy and identity preservation. Code and models will be released for research purposes.
Alessandro Ruzzi, Xiangwei Shi, Xi Wang 0021, Gengyan Li 0001, Shalini De Mello, Hyung Jin Chang, Xucong Zhang, Otmar Hilliges
CVPR7
2021 Self-Supervised 3D Hand Pose Estimation from monocular RGB via Contrastive Learning
abstract
Encouraged by the success of contrastive learning on image classification tasks, we propose a new self-supervised method for the structured regression task of 3D hand pose estimation. Contrastive learning makes use of unlabeled data for the purpose of representation learning via a loss formulation that encourages the learned feature representations to be invariant under any image transformation. For 3D hand pose estimation, it too is desirable to have invariance to appearance transformation such as color jitter. However, the task requires equivariance under affine transformations, such as rotation and translation. To address this issue, we propose an equivariant contrastive objective and demonstrate its effectiveness in the context of 3D hand pose estimation. We experimentally investigate the impact of invariant and equivariant contrastive objectives and show that learning equivariant features leads to better representations for the task of 3D hand pose estimation. Furthermore, we show that standard ResNets with sufficient depth, trained on additional unlabeled data, attain improvements of up to 14.5% in PA-EPE on FreiHAND and thus achieves state-of-the-art performance without any task specific, specialized architectures. Code and models are available at https://ait.ethz.ch/projects/2021/PeCLR/
Adrian Spurr, Aneesh Dahiya, Xi Wang 0021, Xucong Zhang, Otmar Hilliges
ICCV4
2020 Learning-based Region Selection for End-to-End Gaze Estimation
Xucong Zhang, Yusuke Sugano, Andreas Bulling, Otmar Hilliges
BMVC1
2020 Towards End-to-End Video-Based Eye-Tracking
Seonwook Park, Emre Aksan, Xucong Zhang, Otmar Hilliges
ECCV (12)3
2020 ETH-XGaze: A Large Scale Dataset for Gaze Estimation Under Extreme Head Pose and Gaze Variation
Xucong Zhang, Seonwook Park, Thabo Beeler, Derek Bradley, Siyu Tang 0001, Otmar Hilliges
ECCV (5)1
2020 Self-Learning Transformations for Improving Gaze and Head Redirection
abstract
Many computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be used for supervision of downstream tasks remains challenging. In this paper we propose a novel generative model for images of faces, that is capable of producing high-quality images under fine-grained control over eye gaze and head orientation angles. This requires the disentangling of many appearance related factors including gaze and head orientation but also lighting, hue etc. We propose a novel architecture which learns to discover, disentangle and encode these extraneous variations in a self-learned manner. We further show that explicitly disentangling task-irrelevant factors results in more accurate modelling of gaze and head orientation. A novel evaluation scheme shows that our method improves upon the state-of-the-art in redirection accuracy and disentanglement between gaze direction and head orientation changes. Furthermore, we show that in the presence of limited amounts of real-world training data, our method allows for improvements in the downstream task of semi-supervised cross-dataset gaze estimation. Please check our project page at: https://ait.ethz.ch/projects/2020/STED-gaze/
Seonwook Park, Xucong Zhang, Shalini De Mello, Otmar Hilliges
NeurIPS3
2020 Gaze Estimation by Exploring Two-Eye Asymmetry
abstract
Eye gaze estimation is increasingly demanded by recent intelligent systems to facilitate a range of interactive applications. Unfortunately, learning the highly complicated regression from a single eye image to the gaze direction is not trivial. Thus, the problem is yet to be solved efficiently. Inspired by the two-eye asymmetry as two eyes of the same person may appear uneven, we propose the face-based asymmetric regression-evaluation network (FARE-Net) to optimize the gaze estimation results by considering the difference between left and right eyes. The proposed method includes one face-based asymmetric regression network (FAR-Net) and one evaluation network (E-Net). The FAR-Net predicts 3D gaze directions for both eyes and is trained with the asymmetric mechanism, which asymmetrically weights and sums the loss generated by two-eye gaze directions. With the asymmetric mechanism, the FAR-Net utilizes the eyes that can achieve high performance to optimize network. The E-Net learns the reliabilities of two eyes to balance the learning of the asymmetric mechanism and symmetric mechanism. Our FARENet achieves leading performances on MPIIGaze, EyeDiap and RT-Gene datasets. Additionally, we investigate the effectiveness of FARE-Net by analyzing the distribution of errors and ablation study.
Yihua Cheng, Xucong Zhang, Feng Lu 0005, Yoichi Sato 0001
IEEE Trans. Image Process.2
2019 Evaluation of Appearance-Based Methods and Implications for Gaze-Based Applications
abstract
Appearance-based gaze estimation methods that only require an off-the-shelf camera have significantly improved but they are still not yet widely used in the human-computer interaction (HCI) community. This is partly because it remains unclear how they perform compared to model-based approaches as well as dominant, special-purpose eye tracking equipment. To address this limitation, we evaluate the performance of state-of-the-art appearance-based gaze estimation for interaction scenarios with and without personal calibration, indoors and outdoors, for different sensing distances, as well as for users with and without glasses. We discuss the obtained findings and their implications for the most important gaze-based applications, namely explicit eye input, attentive user interfaces, gaze-based user modelling, and passive eye monitoring. To democratise the use of appearance-based gaze estimation and interaction in HCI, we finally present OpenGaze (www.opengaze.org), the first software toolkit for appearance-based gaze estimation and interaction.
Xucong Zhang, Yusuke Sugano, Andreas Bulling
CHI1
2019 Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networks
abstract
Gaze redirection is the task of changing the gaze to a desired direction for a given monocular eye patch image. Many applications such as videoconferencing, films, games, and generation of training data for gaze estimation require redirecting the gaze, without distorting the appearance of the area surrounding the eye and while producing photo-realistic images. Existing methods lack the ability to generate perceptually plausible images. In this work, we present a novel method to alleviate this problem by leveraging generative adversarial training to synthesize an eye image conditioned on a target gaze direction. Our method ensures perceptual similarity and consistency of synthesized images to the real images. Furthermore, a gaze estimation loss is used to control the gaze direction accurately. To attain high-quality images, we incorporate perceptual and cycle consistency losses into our architecture. In extensive evaluations we show that the proposed method outperforms state-of-the-art approaches in terms of both image quality and redirection precision. Finally, we show that generated images can bring significant improvement for the gaze estimation task if used to augment real training data.
Zhe He 0004, Adrian Spurr, Xucong Zhang, Otmar Hilliges
ICCV3
2019 MPIIGaze: Real-World Dataset and Deep Appearance-Based Gaze Estimation
abstract
Learning-based methods are believed to work well for unconstrained gaze estimation, i.e. gaze estimation from a monocular RGB camera without assumptions regarding user, environment, or camera. However, current gaze datasets were collected under laboratory conditions and methods were not evaluated across multiple datasets. Our work makes three contributions towards addressing these limitations. First, we present the MPIIGaze dataset, which contains 213,659 full face images and corresponding ground-truth gaze positions collected from 15 users during everyday laptop use over several months. An experience sampling approach ensured continuous gaze and head poses and realistic variation in eye appearance and illumination. To facilitate cross-dataset evaluations, 37,667 images were manually annotated with eye corners, mouth corners, and pupil centres. Second, we present an extensive evaluation of state-of-the-art gaze estimation methods on three current datasets, including MPIIGaze. We study key challenges including target gaze range, illumination conditions, and facial appearance variation. We show that image resolution and the use of both eyes affect gaze estimation performance, while head pose and pupil centre information are less informative. Finally, we propose GazeNet, the first deep appearance-based gaze estimation method. GazeNet improves on the state of the art by 22 percent (from a mean error of 13.9 degrees to 10.8 degrees) for the most challenging cross-dataset evaluation.
Xucong Zhang, Yusuke Sugano, Mario Fritz, Andreas Bulling
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices
abstract
Learning-based gaze estimation has significant potential to enable attentive user interfaces and gaze-based interaction on the billions of camera-equipped handheld devices and ambient displays. While training accurate person- and device-independent gaze estimators remains challenging, person-specific training is feasible but requires tedious data collection for each target device. To address these limitations, we present the first method to train person-specific gaze estimators across multiple devices. At the core of our method is a single convolutional neural network with shared feature extraction layers and device-specific branches that we train from face images and corresponding on-screen gaze locations. Detailed evaluations on a new dataset of interactions with five common devices (mobile phone, tablet, laptop, desktop computer, smart TV) and three common applications (mobile game, text editing, media center) demonstrate the significant potential of cross-device training. We further explore training with gaze locations derived from natural interactions, such as mouse or touch input.
Xucong Zhang, Michael Xuelin Huang, Yusuke Sugano, Andreas Bulling
CHI1
2018 Appearance-Based Gaze Estimation via Evaluation-Guided Asymmetric Regression
Yihua Cheng, Feng Lu 0005, Xucong Zhang
ECCV (14)3
2018 Robust eye contact detection in natural multi-person interactions using gaze and speaking behaviour
abstract
Eye contact is one of the most important non-verbal social cues and fundamental to human interactions. However, detecting eye contact without specialised eye tracking equipment poses significant challenges, particularly for multiple people in real-world settings. We present a novel method to robustly detect eye contact in natural three- and four-person interactions using off-the-shelf ambient cameras. Our method exploits that, during conversations, people tend to look at the person who is currently speaking. Harnessing the correlation between people's gaze and speaking behaviour therefore allows our method to automatically acquire training data during deployment and adaptively train eye contact detectors for each target user. We empirically evaluate the performance of our method on a recent dataset of natural group interactions and demonstrate that it achieves a relative improvement over the state-of-the-art method of more than 60%, and also improves over a head pose based baseline.
Philipp Müller 0001, Michael Xuelin Huang, Xucong Zhang, Andreas Bulling
ETRA3
2018 Learning to find eye region landmarks for remote gaze estimation in unconstrained settings
abstract
Conventional feature-based and model-based gaze estimation methods have proven to perform well in settings with controlled illumination and specialized cameras. In unconstrained real-world settings, however, such methods are surpassed by recent appearance-based methods due to difficulties in modeling factors such as illumination changes and other visual artifacts. We present a novel learning-based method for eye region landmark localization that enables conventional methods to be competitive to latest appearance-based methods. Despite having been trained exclusively on synthetic data, our method exceeds the state of the art for iris localization and eye shape registration on real-world imagery. We then use the detected landmarks as input to iterative model-fitting and lightweight learning-based gaze estimation methods. Our approach outperforms existing model-fitting and appearance-based methods in the context of person-independent and personalized gaze estimation.
Seonwook Park, Xucong Zhang, Andreas Bulling, Otmar Hilliges
ETRA2
2018 Revisiting data normalization for appearance-based gaze estimation
abstract
Appearance-based gaze estimation is promising for unconstrained real-world settings, but the significant variability in head pose and user-camera distance poses significant challenges for training generic gaze estimators. Data normalization was proposed to cancel out this geometric variability by mapping input images and gaze labels to a normalized space. Although used successfully in prior works, the role and importance of data normalization remains unclear. To fill this gap, we study data normalization for the first time using principled evaluations on both simulated and real data. We propose a modification to the current data normalization formulation by removing the scaling factor and show that our new formulation performs significantly better (between 9.5% and 32.7%) in the different evaluation settings. Using images synthesized from a 3D face model, we demonstrate the benefit of data normalization for the efficiency of the model training. Experiments on real-world images confirm the advantages of data normalization in terms of gaze estimation performance.
Xucong Zhang, Yusuke Sugano, Andreas Bulling
ETRA1
2017 Everyday Eye Contact Detection Using Unsupervised Gaze Target Discovery
abstract
Eye contact is an important non-verbal cue in social signal processing and promising as a measure of overt attention in human-object interactions and attentive user interfaces. However, robust detection of eye contact across different users, gaze targets, camera positions, and illumination conditions is notoriously challenging. We present a novel method for eye contact detection that combines a state-of-the-art appearance-based gaze estimator with a novel approach for unsupervised gaze target discovery, i.e. without the need for tedious and time-consuming manual data annotation. We evaluate our method in two real-world scenarios: detecting eye contact at the workplace, including on the main work display, from cameras mounted to target objects, as well as during everyday social interactions with the wearer of a head-mounted egocentric camera. We empirically evaluate the performance of our method in both scenarios and demonstrate its effectiveness for detecting eye contact independent of target object type and size, camera position, and user and recording environment.
Xucong Zhang, Yusuke Sugano, Andreas Bulling
UIST1
2016 Labelled pupils in the wild: a dataset for studying pupil detection in unconstrained environments
abstract
We present labelled pupils in the wild (LPW), a novel dataset of 66 high-quality, high-speed eye region videos for the development and evaluation of pupil detection algorithms. The videos in our dataset were recorded from 22 participants in everyday locations at about 95 FPS using a state-of-the-art dark-pupil head-mounted eye tracker. They cover people of different ethnicities and a diverse set of everyday indoor and outdoor illumination environments, as well as natural gaze direction distributions. The dataset also includes participants wearing glasses, contact lenses, and make-up. We benchmark five state-of-the-art pupil detection algorithms on our dataset with respect to robustness and accuracy. We further study the influence of image resolution and vision aids as well as recording location (indoor, outdoor) on pupil detection performance. Our evaluations provide valuable insights into the general pupil detection problem and allow us to identify key challenges for robust pupil detection on head-mounted eye trackers.
Marc Tonsen, Xucong Zhang, Yusuke Sugano, Andreas Bulling
ETRA2
2016 AggreGaze: Collective Estimation of Audience Attention on Public Displays
abstract
Gaze is frequently explored in public display research given its importance for monitoring and analysing audience attention. However, current gaze-enabled public display interfaces require either special-purpose eye tracking equipment or explicit personal calibration for each individual user. We present AggreGaze, a novel method for estimating spatio-temporal audience attention on public displays. Our method requires only a single off-the-shelf camera attached to the display, does not require any personal calibration, and provides visual attention estimates across the full display. We achieve this by 1) compensating for errors of state-of-the-art appearance-based gaze estimation methods through on-site training data collection, and by 2) aggregating uncalibrated and thus inaccurate gaze estimates of multiple users into joint attention estimates. We propose different visual stimuli for this compensation: a standard 9-point calibration, moving targets, text and visual stimuli embedded into the display content, as well as normal video content. Based on a two-week deployment in a public space, we demonstrate the effectiveness of our method for estimating attention maps that closely resemble ground-truth audience gaze distributions.
Yusuke Sugano, Xucong Zhang, Andreas Bulling
UIST2
2016 Combining eye tracking with optimizations for lens astigmatism in modern wide-angle HMDs
abstract
Virtual Reality has hit the consumer market with affordable head-mounted displays. When using these, it quickly becomes apparent that the resolution of the built-in display panels still needs to be highly increased. To overcome the resulting higher performance demands, eye tracking can be used for foveated rendering. However, as there are lens distortions in HMDs, there are more possibilities to increase the performance with smarter rendering approaches. We present a new system using optimizations for rendering considering lens astigmatism and combining this with foveated rendering through eye tracking. Depending on the current eye gaze, this delivers a rendering speed-up of up to 20%.
Daniel Pohl, Xucong Zhang, Andreas Bulling
VR2
2016 Concept for using eye tracking in a head-mounted display to adapt rendering to the user's current visual field
abstract
With increasing spatial and temporal resolution in head-mounted displays (HMDs), using eye trackers to adapt rendering to the user is getting important to handle the rendering workload. Besides using methods like foveated rendering, we propose to use the current visual field for rendering, depending on the eye gaze. We use two effects for performance optimizations. First, we noticed a lens defect in HMDs, where depending on the distance of the eye gaze to the center, certain parts of the screen towards the edges are not visible anymore. Second, if the user looks up, he cannot see the lower parts of the screen anymore. For the invisible areas, we propose to skip rendering and to reuse the pixels colors from the previous frame. We provide a calibration routine to measure these two effects. We apply the current visual field to a renderer and get up to 2x speed-ups.
Daniel Pohl, Xucong Zhang, Andreas Bulling, Oliver Grau
VRST2
2015 Appearance-based gaze estimation in the wild
abstract
Appearance-based gaze estimation is believed to work well in real-world settings, but existing datasets have been collected under controlled laboratory conditions and methods have been not evaluated across multiple datasets. In this work we study appearance-based gaze estimation in the wild. We present the MPIIGaze dataset that contains 213,659 images we collected from 15 participants during natural everyday laptop use over more than three months. Our dataset is significantly more variable than existing ones with respect to appearance and illumination. We also present a method for in-the-wild appearance-based gaze estimation using multimodal convolutional neural networks that significantly outperforms state-of-the art methods in the most challenging cross-dataset evaluation. We present an extensive evaluation of several state-of-the-art image-based gaze estimation algorithms on three current datasets, including our own. This evaluation provides clear insights and allows us to identify key research challenges of gaze estimation in the wild.
Xucong Zhang, Yusuke Sugano, Mario Fritz, Andreas Bulling
CVPR1
2015 Rendering of Eyes for Eye-Shape Registration and Gaze Estimation
abstract
Images of the eye are key in several computer vision problems, such as shape registration and gaze estimation. Recent large-scale supervised methods for these problems require time-consuming data collection and manual annotation, which can be unreliable. We propose synthesizing perfectly labelled photo-realistic training data in a fraction of the time. We used computer graphics techniques to build a collection of dynamic eye-region models from head scan geometry. These were randomly posed to synthesize close-up eye images for a wide range of head poses, gaze directions, and illumination conditions. We used our model's controllability to verify the importance of realistic illumination and shape variations in eye-region training data. Finally, we demonstrate the benefits of our synthesized training data (SynthesEyes) by out-performing state-of-the-art methods for eye-shape registration as well as cross-dataset appearance-based gaze estimation in the wild.
Erroll Wood, Tadas Baltrusaitis, Xucong Zhang, Yusuke Sugano, Peter Robinson 0001, Andreas Bulling
ICCV3
2014 Face detection by structural models
Xucong Zhang, Zhen Lei 0001, Stan Z. Li
Image Vis. Comput.2
2013 Robust Multi-resolution Pedestrian Detection in Traffic Scenes
abstract
The serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection techniques. In this paper, we take pedestrian detection in different resolutions as different but related problems, and propose a Multi-Task model to jointly consider their commonness and differences. The model contains resolution aware transformations to map pedestrians in different resolutions to a common space, where a shared detector is constructed to distinguish pedestrians from background. For model learning, we present a coordinate descent procedure to learn the resolution aware transformations and deformable part model (DPM) based detector iteratively. In traffic scenes, there are many false positives located around vehicles, therefore, we further build a context model to suppress them according to the pedestrian-vehicle relationship. The context model can be learned automatically even when the vehicle annotations are not available. Our method reduces the mean miss rate to 60% for pedestrians taller than 30 pixels on the Caltech Pedestrian Benchmark, which noticeably outperforms previous state-of-the-art (71%).
Xucong Zhang, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
CVPR2
2012 Water Filling: Unsupervised People Counting via Vertical Kinect Sensor
abstract
People counting is one of the key components in video surveillance applications, however, due to occlusion, illumination, color and texture variation, the problem is far from being solved. Different from traditional visible camera based systems, we construct a novel system that uses vertical Kinect sensor for people counting, where the depth information is used to remove the affect of the appearance variation. Since the head is always closer to the Kinect sensor than other parts of the body, people counting task equals to find the suitable local minimum regions. According to the particularity of the depth map, we propose a novel unsupervised water filling method that can find these regions with the property of robustness, locality and scale-invariance. Experimental comparisons with mean shift and random forest on two databases validate the superiority of our water filling algorithm in people counting.
Xucong Zhang, Shikun Feng, Zhen Lei 0001, Dong Yi, Stan Z. Li
AVSS1