VLDB 2026 Research / reviewers in the wild / expert
Pengwei Yin
dblp:212/7731
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3443-976XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gaze Label Alignment: Alleviating Domain Shift for Gaze EstimationabstractGaze estimation methods encounter significant performance deterioration when being evaluated across different domains, because of the domain gap between the testing and training data. Existing methods try to solve this issue by reducing the deviation of data distribution, however, they ignore the existence of label deviation in the data due to the acquisition mechanism of the gaze label and the individual physiological differences. In this paper, we first point out that the influence brought by the label deviation cannot be ignored, and propose a gaze label alignment algorithm (GLA) to eliminate the label distribution deviation. Specifically, we first train the feature extractor on all domains to get domain invariant features, and then select an anchor domain to train the gaze regressor. We predict the gaze label on remaining domains and use a mapping function to align the labels. Finally, these aligned labels can be used to train gaze estimation models. Therefore, our method can be combined with any existing method. Experimental results show that our GLA method can effectively alleviate the label distribution shift, and SOTA gaze estimation methods can be further improved obviously. Guanzhong Zeng, Zefu Xu, Pengwei Yin, Wenqi Ren, Di Xie |
AAAI | 4 |
| 2025 | Test Time Prompt Tuning for Domain Adaptive Gaze EstimationabstractAlthough current gaze estimation methods achieve promising results in within-domain evaluations, they suffer from significant degradation when tested on real-world scenarios due to the presence of distribution shift. As collecting labeled samples entails significant costs, some methods leverage unsupervised domain adaptation (UDA) techniques to solve this problem. However, they need the source domain data during training, which may not be accessible due to privacy concerns, and tune all parameters, which may not be practical to conduct on edge devices at the test time due to computational constraints. Therefore, it is desirable to design an efficient and accurate gaze estimation method without source domain data. To achieve this goal, we propose a test time prompt tuning framework for efficient source-free domain adaptive gaze estimation. Specifically, we only learn a negligible number of parameters as prompts to adjust the gaze feature extracted from the final layer of the gaze encoder without perturbing original network. To learn meaningful prompts, an unsupervised learning loss is designed which aligns the feature distribution of target domain with the source domain and makes the estimator predict confident labels on the target domain data. The proposed method is 2.9 times faster in terms of adaptation speed than the most recent efficient method with only 5% of its updating parameters. Extensive experiments on four cross-dataset validations demonstrate the effectiveness of the proposed method. Pengwei Yin |
ICASSP | 2 |
| 2025 | Denoising Diffusion Models are Good General Gaze Feature LearnersabstractSince the collection of labeled gaze data is laborious and time-consuming, methods which can learn generalizable features by leveraging large-scale available unlabeled data are desirable. In recent years, we have witnessed the tremendous capabilities of diffusion models in generating images as well as their potential in feature representation learning. In this paper, we investigate whether they can acquire discriminative representations for gaze estimation via generative pre-training. To achieve this goal, we propose a self-supervised learning framework with diffusion models for gaze estimation, called GazeDiff. Specifically, we utilize a conditional diffusion model to generate target image with gaze direction specified by the reference image as the pre-training task. To facilitate the diffusion model to learn gaze related features as condition, we propose a disentangling feature learning strategy, which first learns appearance feature, head pose feature, and eye direction feature respectively, and then combines them as the conditional features. Extensive experiments demonstrate denoising diffusion models are also good general gaze feature learners. Guanzhong Zeng, Pengwei Yin, Zefu Xu |
IJCAI | 3 |
| 2024 | CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic ModelabstractGaze estimation methods often experience significant performance degradation when evaluated across different domains, due to the domain gap between the testing and training data. Existing methods try to address this issue using various domain generalization approaches, but with little success because of the limited diversity of gaze datasets, such as appearance, wearable, and image quality. To overcome these limitations, we propose a novel framework called CLIP-Gaze that utilizes a pre-trained vision-language model to leverage its transferable knowledge. Our framework is the first to leverage the vision-and-language cross-modality approach for gaze estimation task. Specifically, we extract gaze-relevant feature by pushing it away from gaze-irrelevant features which can be flexibly constructed via language descriptions. To learn more suitable prompts, we propose a personalized context optimization method for text prompt tuning. Furthermore, we utilize the relationship among gaze samples to refine the distribution of gaze-relevant features, thereby improving the generalization capability of the gaze estimation model. Extensive experiments demonstrate the excellent performance of CLIP-Gaze over existing methods on four cross-domain evaluations. Pengwei Yin, Guanzhong Zeng, Di Xie |
AAAI | 1 |
| 2024 | LG-Gaze: Learning Geometry-Aware Continuous Prompts for Language-Guided Gaze Estimation
Pengwei Yin, Guanzhong Zeng, Di Xie |
ECCV (83) | 1 |
| 2024 | NERF-GAZE: A Head-Eye Redirection Parametric Model for Gaze EstimationabstractGaze estimation is a fundamental aspect of many visual tasks. However, the high cost of acquiring gaze datasets with 3D annotations hinders the optimization and application of gaze estimation models. In this work, we propose a novel Head-Eye redirection parametric model based on Neural Radiance Field. This model allows for dense gaze data generation with view consistency and accurate gaze direction. Furthermore, our head-eye redirection parametric model can decouple the face and eyes for separate neural rendering, which enables us to separately control the attributes of the face, identity, illumination, and eye gaze direction. As a result, diverse 3D-aware gaze datasets can be obtained by manipulating the latent code belonging to different face attributes in an unsupervised manner. Our method has achieved state-of-the-art performance in image quality and accuracy gaze annotations compared with existing gaze data synthesis methods. Extensive experiments on several benchmarks demonstrate that our method can effectively improve domain generalization and domain adaptation in the gaze estimation task. Pengwei Yin, Jiawu Dai |
ICASSP | 1 |