EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Bai
dblp:293/4391
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-0708-3772ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Face, body and person analysis · 76% Representation and self-supervised learning · 18% Deep learning architectures and training · 6% | |
| Human-computer interaction and pervasive computing
3 papers |
Wearable and physiological sensing · 81% Immersive interaction · 12% Usability and user experience research · 8% | |
| Computer graphics and multimedia
1 paper |
Virtual and augmented reality · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Wearable and physiological sensing
eye tracking |
2.2 | 3 | 2025 | PVEye: A Large Posture-Variant Eye Tracking Dataset for Head-Mounted AR Devices · IEEE Trans. Vis. Comput. Graph. 2025 Real-Time Gaze Tracking via Head-Eye Cues on Head Mounted Devices · IEEE Trans. Mob. Comput. 2024 Real-time Gaze Tracking with Head-eye Coordination for Head-mounted Displays · ISMAR 2022 |
Computer vision › Face, body and person analysis
gaze estimation |
1.9 | 2 | 2026 | KDB-Gaze: Keypoint-Guided Dual-Branch Learning for Gaze Estimation · IEEE Trans. Mob. Comput. 2026 De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimation · CVPR 2025 |
Computer vision › Face, body and person analysis
eye tracking |
1.0 | 1 | 2026 | KDB-Gaze: Keypoint-Guided Dual-Branch Learning for Gaze Estimation · IEEE Trans. Mob. Comput. 2026 |
Computer vision › Face, body and person analysis › gaze estimation
3d gaze estimation |
0.9 | 1 | 2025 | De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimation · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.9 | 1 | 2025 | De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimation · CVPR 2025 |
Wearable and physiological sensing
eye-head coordination |
0.6 | 1 | 2022 | Real-time Gaze Tracking with Head-eye Coordination for Head-mounted Displays · ISMAR 2022 |
Machine learning › Deep learning architectures and training › multi-scale representation
multi-scale feature extraction |
0.3 | 1 | 2026 | KDB-Gaze: Keypoint-Guided Dual-Branch Learning for Gaze Estimation · IEEE Trans. Mob. Comput. 2026 |
Usability and user experience research
user study |
0.3 | 1 | 2025 | PVEye: A Large Posture-Variant Eye Tracking Dataset for Head-Mounted AR Devices · IEEE Trans. Vis. Comput. Graph. 2025 |
Immersive interaction
head-mounted display |
0.2 | 1 | 2024 | Real-Time Gaze Tracking via Head-Eye Cues on Head Mounted Devices · IEEE Trans. Mob. Comput. 2024 |
Immersive interaction › head-mounted display
augmented reality head-mounted display |
0.2 | 1 | 2022 | Real-time Gaze Tracking with Head-eye Coordination for Head-mounted Displays · ISMAR 2022 |
Methods — techniques the papers use, named apart from their topics
feature-based method · 1.7NVGaze model · 1.7multi-task learning · 1.0graph attention network · 1.0feature pyramid network · 1.0deformable DETR · 1.0dual-branch decoding · 0.9deformable sparse attention · 0.9multimodal learning · 0.8hierarchical neural network · 0.8multi-modal network · 0.6dataset construction · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KDB-Gaze: Keypoint-Guided Dual-Branch Learning for Gaze EstimationabstractNear-eye gaze estimation has been a core technology for natural interaction in smart head-mounted devices. Existing gaze estimation approaches have the following limitations: (1) Data-driven methods learn nonlinear mappings directly from images to gaze direction but ignore the physiological characteristics of the iris–pupil system. (2) Model-driven methods rely on hand-designed modeling and perform poorly in complex, noisy scenes. (3) Hybrid-driven methods attempt to combine both paradigms but are restricted by simplified model designs. To address these issues, we propose KDB-Gaze, aKeypoint-GuidedDual-Branch learning for near-eyeGazeestimation. First, a backbone and a 3-Pass path aggregation feature pyramid network effectively capture multi-scale ocular features. Then, two parallel branches refine the features: an implicit spatial modeling branch employs deformable DETR to capture nonlinear features, while an explicit topological constraint branch uses a graph attention network to provide stable physiological guidance. Finally, the dual-branch features are fused and optimized jointly with a multi-task loss for precise prediction. Extensive experiments on the TEyeD, LPW, and NVGaze datasets verify the state-of-the-art performance of the proposed method. Xiaowei Bai, Liang Xie 0012, Qining Wang, Erwei Yin |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimationabstract3D Gaze estimation is a challenging task due to two main issues. First, existing methods focus on analyzing dense features (e.g., large pixel regions), which are sensitive to local noise (e.g., light spots, blurs) and result in increased computational complexity. Second, an eyeball model can correspond multiple gaze directions, and the entangled representation between gazes and models increases the learning difficulty. To address these issues, we propose De2Gaze, a lightweight and accurate model-aware 3D gaze estimation method. In De2Gaze, we introduce two key innovations for deformable and decoupled representation learning. Specifically, first, we propose a deformable sparse attention mechanism that can adapt sparse sampling points to attention areas to avoid local noise influences. Second, we propose a spatial decoupling network with a dual-branch decoding architecture to disentangle invariant (e.g., eyeball radius, position) and variable (e.g., gaze, pupil, iris) features from the latent space. Compared to existing methods, De2Gaze requires fewer sparse features, and achieves faster convergence speed, lower computational complexity, and higher accuracy in 3D gaze estimation. Qualitative and quantitative experiments demonstrate that De2Gaze achieves state-of-the-art accuracy and high-quality semantic segmentation for 3D gaze estimation on the TEyeD dataset. Yunfeng Xiao, Xiaowei Bai, Baojun Chen, Liang Xie 0012, Erwei Yin |
CVPR | 2 |
| 2025 | PVEye: A Large Posture-Variant Eye Tracking Dataset for Head-Mounted AR DevicesabstractEye tracking technology, essential for enhancing user experience in virtual reality (VR) and augmented reality (AR) devices, has been widely incorporated into advanced head-mounted devices like the Apple Vision Pro and PICO 4 Pro, becoming a standard feature. However, dedicated eye tracking datasets for such devices are severely lacking, with existing datasets commonly facing issues like camera skew and low resolution, particularly failing to adequately consider the diversity in wearing postures. To address this gap, we have developed the Posture-Variant Eye Tracking Dataset (PVEye), which includes 11,044,800 high-resolution near-eye images from 104 participants, showcasing a rich variety of wearing postures. This dataset aims to advance the development and application of appearance-based eye tracking methods. Utilizing this dataset, our evaluations demonstrate that the appearance-based method, particularly the NVGaze model, provides improved accuracy and robustness compared to the traditional feature-based method. Crucially, our experiments indicate that variations in wearing posture can significantly impact eye tracking performance, with posture-related errors contributing approximately 45% to the overall error variance. Moreover, the study delves into the specific impact of calibration and other critical factors on eye tracking performance, offering insights for further optimization of tracking effectiveness. Xiaowei Bai, Liang Xie 0012, Yingxi Li, Qining Wang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Two-Stream Vision Swin Transformer for Video-based Eye Movement Detection
Xiaowei Bai, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
CogSci | 2 |
| 2024 | Complex-Valued Gabor-Attention Residual Fusion Network for Iris RecognitionabstractIris recognition has gained significant attention in identity verification due to the unique, stable texture patterns in iris. Successfully extracting these patterns is essential for quick and precise identification. Although deep learning methods have automated the iris recognition, they predominantly rely on real-valued networks that overlook the complex-valued representation of iris texture. This means they cannot effectively process phase and amplitude information, and fail to integrate domain-specific knowledge of iris, thereby not fully capturing the intricate details of the iris texture. Inspired by classical manual methods that efficiently harness the complex-valued representation of the iris to extract both amplitude and phase information. We integrate Gabor filters with complex-valued neural networks, propose a Complex-Valued Gabor-Attention Residual Fusion Network (GRFN) tailored for iris recognition, aiming to comprehensively capture the iris texture’s multi-scale and multi-orientation phase and amplitude features. The GRFN incorporates adaptive Gabor Complex-Valued Convolution Kernels (GCVK) to introduce a Gabor attention mechanism focused on iris biometric characteristics. Furthermore, we propose a novel residual feature fusion approach that selects and merges local and global features across multiple directions and scales, mitigating model degradation and enhancing the network’s ability to extract iris texture features effectively. Extensive experiments show that the proposed network outperforms the state-of-the-art performance on two benchmark datasets. Zhuoru Li, Xiaowei Bai, Yingxi Li, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ECAI | 3 |
| 2024 | Real-Time Gaze Tracking via Head-Eye Cues on Head Mounted DevicesabstractGaze is a crucial element in human-computer interaction and plays an increasingly vital role in promoting the adoption of head-mounted devices (HMDs). Existing gaze tracking methods for HMDs either demand user calibration or face challenges in balancing accuracy and speed, compromising the overall user experience. In this paper, we introduce a novel strategy for real-time, calibration-free gaze tracking using joint head-eye cues on HMDs. Initially, we create a multimodal gaze tracking dataset named HE-Gaze, encompassing synchronized eye images and 6DoF head movement data, addressing a gap in the current data landscape. Statistical analyses unveil the correlation between head movements and gaze positions. Building on these insights, we introduce the hierarchical head-eye coordinated gaze tracking model (HHE-Tracker), which incorporates two lightweight branches to encode input eye images and head sequences efficiently. It combines encoded head velocity and posture features with eye features across various scales to infer gaze position. HHE-Tracker was implemented on a commercial HMD, and its performance was assessed in unconstrained scenarios. The results demonstrate the HHE-Tracker's capability to accurately estimate gaze positions in real-time. In comparison to the state-of-the-art gaze tracking algorithm, HHE-Tracker exhibits commendable accuracy (3.47$^{\circ }$) and a 40-fold speedup (81FPSon a Snapdragon 845 SoC). Yingxi Li, Xiaowei Bai, Liang Xie 0012, Feng Lu 0005, Feitian Zhang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | Real-time Gaze Tracking with Head-eye Coordination for Head-mounted DisplaysabstractHigh-accuracy, low-latency gaze tracking is becoming one of the indispensable features in augmented reality (AR) head-mounted devices (HMDs). Researchers have proposed different approaches to predict gaze positions from eye images. However, since only the eye modality is focused, these appearance-based algorithms are still struggle to trade off the accuracy and running speed in HMDs. In this paper, we propose a lightweight multi-modal network (HE-Tracker) to regress gaze positions. By fusing head-movement features with eye features, HE-Tracker achieves comparable accuracy (3.655° in all subjects) and $27 \times$ speedup (48 fps in the specialized AR HMD) compared to the state-of-the-art gaze tracking algorithm. We further demonstrate that when applying our head-eye coordination strategy to other baseline models, all these models achieve at least 6.36% performance improvement without a pronounced effect on running speed. Moreover, we construct HE-Gaze, the first multi-modal dataset with eye images and head-movement data for near-eye gaze tracking. This dataset is currently made of 757,360 frames and 15 persons, providing an opportunity to foster research in multi-modal gaze tracking approaches. Our dataset is available at DOWNLOAD LINK1. Lingling Chen, Yingxi Li, Xiaowei Bai, Yongqiang Hu, Mingwu Song, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ISMAR | 3 |