VLDB 2026 Research / reviewers in the wild / expert
Haofei Wang 0001
dblp:50/10382-1
· DBLP profile ↗
18ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-1199-9206ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FaceMamba: Geometry-Aware Mamba for Efficient Speech-Driven 3D Facial Animation
Yifan Ge, Zhiqiang Ren, Yazhan Zhang, Haofei Wang 0001 |
Comput. Graph. Forum | 4 |
| 2026 | EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent DiffusionabstractSpeech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies. Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, Hong Qin 0001, Yang Gao 0032 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032 |
Vis. Comput. | 4 |
| 2025 | Gam360: sensing gaze activities of multi-persons in 360 degrees
Zhuojiang Cai, Haofei Wang 0001, Yuhao Niu, Feng Lu 0005 |
CCF Trans. Pervasive Comput. Interact. | 2 |
| 2025 | From Gaze Jitter to Domain Adaptation: Generalizing Gaze Estimation by Manipulating High-Frequency Components
Ruicong Liu, Haofei Wang 0001, Feng Lu 0005 |
Int. J. Comput. Vis. | 2 |
| 2024 | Enhancing Text Entry in Mixed Reality with Tangible Feedback
Haofei Wang 0001, Feng Lu 0005 |
ICXR | 1 |
| 2024 | User Engagement Correlates Better with Behavioral than Physiological Measures in a Virtual Reality Robotic Rehabilitation SystemabstractRobotic systems to assist with movement rehabil-itation are transitioning from providing fixed pre-programmed assistance towards adaptive challenge-oriented strategies that present patients with tasks that are demanding yet achiev-able. This promotes active engagement, which is crucial for stimulating neural plasticity and promoting recovery. While it has been well established that varying the challenge level can affect user engagement, measuring engagement during task performance has received less attention. To investigate this issue, we developed a virtual reality (VR) robotic system for upper limb rehabilitation using a line-tracing task that measures physiological and behavioral signals. Challenge level can be modulated by introducing force noise disturbance. We con-ducted a preliminary study on 12 participants, measuring user engagement and physiological/behavioral signals at different noise (challenge) levels. Our findings align with the predictions of flow channel theory. Engagement peaks at an intermediate challenge level. While past work considered only physiological measures, our results reveal that behavioral measures are better correlated with user engagement. Physiological measures correlate better with arousal. This work takes a step toward systems that dynamically adapt task parameters to optimize user engagement. Haofei Wang 0001, Bertram E. Shi |
SMC | 2 |
| 2024 | Appearance-Based Gaze Estimation With Deep Learning: A Review and BenchmarkabstractHuman gaze provides valuable information on human focus and intentions, making it a crucial area of research. Recently, deep learning has revolutionized appearance-based gaze estimation. However, due to the unique features of gaze estimation research, such as the unfair comparison between 2D gaze positions and 3D gaze vectors and the different pre-processing and post-processing methods, there is a lack of a definitive guideline for developing deep learning-based gaze estimation algorithms. In this paper, we present a systematic review of the appearance-based gaze estimation methods using deep learning. First, we survey the existing gaze estimation algorithms along the typical gaze estimation pipeline: deep feature extraction, deep learning model design, personal calibration and platforms. Second, to fairly compare the performance of different approaches, we summarize the data pre-processing and post-processing methods, including face/eye detection, data rectification, 2D/3D gaze conversion and gaze origin conversion. Finally, we set up a comprehensive benchmark for deep learning-based gaze estimation. We characterize all the public datasets and provide the source code of typical gaze estimation algorithms. This paper serves not only as a reference to develop deep learning-based gaze estimation methods, but also a guideline for future gaze estimation research. Yihua Cheng, Haofei Wang 0001, Yiwei Bao, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | PnP-GA+: Plug-and-Play Domain Adaptation for Gaze Estimation Using Model VariantsabstractAppearance-based gaze estimation has garnered increasing attention in recent years. However, deep learning-based gaze estimation models still suffer from suboptimal performance when deployed in new domains, e.g., unseen environments or individuals. In our previous work, we took this challenge for the first time by introducing a plug-and-play method (PnP-GA) to adapt the gaze estimation model to new domains. The core concept of PnP-GA is to leverage the diversity brought by a group of model variants to enhance the adaptability to diverse environments. In this article, we propose the PnP-GA+ by extending our approach to explore the impact of assembling model variants using three additional perspectives: color space, data augmentation, and model structure. Moreover, we propose an intra-group attention module that dynamically optimizes pseudo-labeling during adaptation. Experimental results demonstrate that by directly plugging several existing gaze estimation networks into the PnP-GA+ framework, it outperforms state-of-the-art domain adaptation approaches on four standard gaze domain adaptation tasks on public datasets. Our method consistently enhances cross-domain performance, and its versatility is improved through various ways of assembling the model group. Ruicong Liu, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | cbPPGGAN: A Generic Enhancement Framework for Unpaired Pulse Waveforms in Camera-Based PhotoplethysmographyabstractCamera-based photoplethysmography (cbP PG) is a non-contact technique that measures cardiac-related blood volume alterations in skin surface vessels through the analysis of facial videos. While traditional approaches can estimate heart rate (HR) under different illuminations, their accuracy can be affected by motion artifacts, leading to poor waveform fidelity and hindering further analysis of heart rate variability (HRV); deep learning-based approaches reconstruct high-quality pulse waveform, yet their performance significantly degrades under illumination variations. In this work, we aim to leverage the strength of these two methods and propose a framework that possesses favorable generalization capabilities while maintaining waveform fidelity. For this purpose, we propose the cbPPGGAN, an enhancement framework for cbPPG that enables the flexible incorporation of both unpaired and paired data sources in the training process. Based on the waveforms extracted by traditional approaches, the cbPPGGAN reconstructs high-quality waveforms that enable accurate HR estimation and HRV analysis. In addition, to address the lack of paired training data problems in real-world applications, we propose a cycle consistency loss that guarantees the time-frequency consistency before/after mapping. The method enhances the waveform quality of traditional POS approaches in different illumination tests (BH-rPPG) and cross-datasets (UBFC-rPPG) with mean absolute error (MAE) values of 1.34 bpm and 1.65 bpm, and average beat-to-beat (AVBB) values of 27.46 ms and 45.28 ms, respectively. Experimental results demonstrate that the cbPPGGAN enhances cbPPG signal quality and outperforms the state-of-the-art approaches in HR estimation and HRV analysis. The proposed framework opens a new pathway toward accurate HR estimation in an unconstrained environment. Haofei Wang 0001, Bo Liu 0112, Feng Lu 0005 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Learning a Generalized Gaze Estimator from Gaze-Consistent FeatureabstractGaze estimator computes the gaze direction based on face images. Most existing gaze estimation methods perform well under within-dataset settings, but can not generalize to unseen domains. In particular, the ground-truth labels in unseen domain are often unavailable. In this paper, we propose a new domain generalization method based on gaze-consistent features. Our idea is to consider the gaze-irrelevant factors as unfavorable interference and disturb the training data against them, so that the model cannot fit to these gaze-irrelevant factors, instead, only fits to the gaze-consistent features. To this end, we first disturb the training data via adversarial attack or data augmentation based on the gaze-irrelevant factors, i.e., identity, expression, illumination and tone. Then we extract the gaze-consistent features by aligning the gaze features from disturbed data with non-disturbed gaze features. Experimental results show that our proposed method achieves state-of-the-art performance on gaze domain generalization task. Furthermore, our proposed method also improves domain adaption performance on gaze estimation. Our work provides new insight on gaze domain generalization task. Mingjie Xu, Haofei Wang 0001, Feng Lu 0005 |
AAAI | 2 |
| 2022 | Generalizing Gaze Estimation with Rotation ConsistencyabstractRecent advances of deep learning-based approaches have achieved remarkable performance on appearance-based gaze estimation. However, due to the shortage of target domain data and absence of target labels, generalizing gaze estimation algorithm to unseen environments is still challenging. In this paper, we discover the rotation-consistency property in gaze estimation and introduce the ‘sub-label’ for unsupervised domain adaptation. Consequently, we propose the Rotation-enhanced Unsupervised Domain Adaptation (RUDA) for gaze estimation. First, we rotate the original images with different angles for training. Then we conduct domain adaptation under the constraint of rotation consistency. The target domain images are assigned with sub-labels, derived from relative rotation angles rather than untouchable real labels. With such sub-labels, we propose a novel distribution loss that facilitates the domain adaptation. We evaluate the RUDA framework on four cross-domain gaze estimation tasks. Experimental results demonstrate that it improves the performance over the baselines with gains ranging from 12.2% to 30.5%. Our framework has the potential to be used in other computer vision tasks with physical constraints. Yiwei Bao, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
CVPR | 3 |
| 2022 | Assessment of Deep Learning-Based Heart Rate Estimation Using Remote Photoplethysmography Under Different IlluminationsabstractRemote photoplethysmography (rPPG) monitors heart rate (HR) without requiring physical contact, which has applications. Deep learning based rPPG has demonstrated superior performance over the traditional approaches in controlled context. However, the lighting situation in indoor space is typically complex, with uneven light distribution and frequent variations in illumination. It lacks a fair comparison of different methods under different illuminations using the same dataset. In this article, we present a public dataset, namely the BeiHang University remote photoplethysmography (BH-rPPG) dataset, which contains data from 35 subjects under three illuminations: 1) low; 2) medium; and 3) high illumination. We also provide the ground truth HR measured by an oximeter. We evaluate the performance of three deep learning-based methods (Deepphys, rPPGNet, and Physnet) to that of four traditional methods (CHROM, GREEN, ICA, and POS) using two public datasets: 1) UBFC-rPPG; 2) the BH-rPPG. The experimental results demonstrate that traditional methods are more resistant to fluctuating illuminations. We found that the Physnet achieves lowest mean absolute error among deep learning based method under medium illumination, whereas the CHROM achieves 1.04 beats per minute, outperforming the Physnet by 80$\%$. Additionally, we investigate potential methods for improving performance of deep learning based methods. We find that brightness augmentation make model more robust to variation illumination. These findings suggest that while developing deep learning based HR estimation algorithms, illumination variation should be taken into account. This work serves as a benchmark for rPPG performance evaluation and it opens a pathway for future investigation into deep learning based rPPG under illumination variations. Haofei Wang 0001, Feng Lu 0005 |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2021 | Separating Content and Style for Unsupervised Image-to-Image Translation
Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
BMVC | 2 |
| 2021 | Generalizing Gaze Estimation with Outlier-guided Collaborative AdaptationabstractDeep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation framework (PnP-GA), which is an ensemble of networks that learn collaboratively with the guidance of outliers. Since our proposed framework does not require ground-truth labels in the target domain, the existing gaze estimation networks can be directly plugged into PnPGA and generalize the algorithms to new domains. We test PnP-GA on four gaze domain adaptation tasks, ETH-to-MPII, ETH-to-EyeDiap, Gaze360-to-MPII, and Gaze360to-EyeDiap. The experimental results demonstrate that the PnP-GA framework achieves considerable performance improvements of 36.9%, 31.6%, 19.4%, and 11.8% over the baseline system. The proposed framework also outperforms the state-of-the-art domain adaptation approaches on gaze domain adaptation tasks. Code has been released at https://github.com/DreamtaleCore/PnP-GA. Yunfei Liu 0001, Ruicong Liu, Haofei Wang 0001, Feng Lu 0005 |
ICCV | 3 |
| 2021 | Interaction With Gaze, Gesture, and Speech in a Flexibly Configurable Augmented Reality SystemabstractMultimodal interaction has become a recent research focus since it offers better user experience in augmented reality (AR) systems. However, most existing works only combine two modalities at a time, e.g., gesture and speech. Multimodal interactive system integrating gaze cue has rarely been investigated. In this article, we propose a multimodal interactive system that integrates gaze, gesture, and speech in a flexibly configurable AR system. Our lightweight head-mounted device supports accurate gaze tracking, hand gesture recognition, and speech recognition simultaneously. The system can be easily configured into various modality combinations, which enables us to investigate the effects of different interaction techniques. We evaluate the efficiency of these modalities using two tasks: the lamp brightness adjustment task and the cube manipulation task. We also collect subjective feedback when using such systems. The experimental results demonstrate that theGaze+Gesture+Speechmodality is superior in terms of efficiency, and theGesture+Speechmodality is more preferred by users. Our system opens the pathway toward a multimodal interactive AR system that enables flexible configuration. Zhimin Wang 0001, Haofei Wang 0001, Huangyue Yu, Feng Lu 0005 |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2019 | Gaze awareness improves collaboration efficiency in a collaborative assembly taskabstractIn building human robot interaction systems, it would be helpful to understand how humans collaborate, and in particular, how humans use others' gaze behavior to estimate their intent. Here we studied the use of gaze in a collaborative assembly task, where a human user assembled an object with the assistance of a human helper. We found that the being aware of the partner's gaze significantly improved collaboration efficiency. Task completion times were much shorter when gaze communication was available, than when it was blocked. In addition, we found that the user's gaze was more likely to lie on the object of interest in the gaze-aware case than the gaze-blocked case. In the context of human-robot collaboration systems, our results suggest that gaze data in the period surrounding verbal requests will be more informative and can be used to predict the target object. Haofei Wang 0001, Bertram E. Shi |
ETRA | 1 |
| 2018 | SLAM-based localization of 3D gaze using a mobile eye trackerabstractPast work in eye tracking has focused on estimating gaze targets in two dimensions (2D), e.g. on a computer screen or scene camera image. Three-dimensional (3D) gaze estimates would be extremely useful when humans are mobile and interacting with the real 3D environment. We describe a system for estimating the 3D locations of gaze using a mobile eye tracker. The system integrates estimates of the user's gaze vector from a mobile eye tracker, estimates of the eye tracker pose from a visual-inertial simultaneous localization and mapping (SLAM) algorithm, a 3D point cloud map of the environment from a RGB-D sensor. Experimental results indicate that our system produces accurate estimates of 3D gaze over a much larger range than remote eye trackers. Our system will enable applications, such as the analysis of 3D human attention and more anticipative human robot interfaces. Haofei Wang 0001, Jimin Pi, Tong Qin 0001, Shaojie Shen, Bertram E. Shi |
ETRA | 1 |