Zhenfeng Fan

dblp:183/5218 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-2299-7596ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Learning Person-Specific Animatable Face Models from In-the-Wild Images via a Shared Base Model
abstract
Training a generic 3D face reconstruction model in a self-supervised manner using large-scale, in-the-wild 2D face image datasets enhances robustness to varying lighting conditions and occlusions while allowing the model to capture animatable wrinkle details across diverse facial expressions. However, a generic model often fails to adequately represent the unique characteristics of specific individuals. In this paper, we propose a method to train a generic base model and then transfer it to yield person-specific models by integrating lightweight adapters within the large-parameter ViT-MAE base model. These person-specific models excel at capturing individual facial shapes and detailed features while preserving the robustness and prior knowledge of detail variations from the base model. During training, we introduce a silhouette vertex re-projection loss to address boundary "landmark marching" issues on the 3D face caused by pose variations. Additionally, we employ an innovative teacher-student loss to leverage the inherent strengths of UNet in feature boundary localization for training our detail MAE. Quantitative and qualitative experiments demonstrate that our approach achieves state-of-the-art performance in face alignment, detail accuracy, and richness. The source code is available at https://github.com/danielmao2000/person-specific-animatable-face.
Yuxiang Mao, Zhenfeng Fan, ZhiJie Zhang, Shihong Xia
CVPR2
2025 AIcomposer: Any Style and Content Image Composition via Feature Integration
Zhenfeng Fan, Zhang Wen, Zhengzhou Zhu, Yunjin Li
ICCV2
2024 Mining collaborative spatio-temporal clues for face forgery detection
Bo Ding 0006, Zhenfeng Fan, Zejun Zhao, Shihong Xia
Multim. Tools Appl.2
2023 Unpaired Multi-domain Attribute Translation of 3D Facial Shapes with a Square and Symmetric Geometric Map
abstract
While impressive progress has recently been made in image-oriented facial attribute translation, shape-oriented 3D facial attribute translation remains an unsolved issue. This is primarily limited by the lack of 3D generative models and ineffective usage of 3D facial data. We propose a learning framework for 3D facial attribute translation to relieve these limitations. Firstly, we customize a novel geometric map for 3D shape representation and embed it in an end-to-end generative adversarial network. The geometric map represents 3D shapes symmetrically on a square image grid, while preserving the neighboring relationship of 3D vertices in a local least-square sense. This enables effective learning for the latent representation of data with different attributes. Secondly, we employ a unified and unpaired learning framework for multi-domain attribute translation. It not only makes effective usage of data correlation from multiple domains, but also mitigates the constraint for hardly accessible paired data. Finally, we propose a hierarchical architecture for the discriminator to guarantee robust results against both global and local artifacts. We conduct extensive experiments to demonstrate the advantage of the proposed framework over the state-of-the-art in generating high-fidelity facial shapes. Given an input 3D facial shape, the proposed framework is able to synthesize novel shapes of different attributes, which covers some downstream applications, such as expression transfer, gender translation, and aging. Code at https://github.com/NaughtyZZ/3D_facial_shape_attribute_translation_ssgmap.
Zhenfeng Fan, Chongyang Zhong, Min Cao 0005, Shihong Xia
ICCV1
2023 RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search
abstract
Text-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including two novel tasks: Relation-Aware learning (RA) and Sensitivity-Aware learning (SA). For one thing, existing methods cluster representations of all positive pairs without distinction and overlook the noise problem caused by the weak positive pairs where the text and the paired image have noise correspondences, thus leading to overfitting learning. RA offsets the overfitting risk by introducing a novel positive relation detection task (i.e., learning to distinguish strong and weak positive pairs). For another thing, learning invariant representation under data augmentation (i.e., being insensitive to some transformations) is a general practice for improving representation's robustness in existing methods. Beyond that, we encourage the representation to perceive the sensitive transformation by SA (i.e., learning to detect the replaced words), thus promoting the representation's robustness. Experiments demonstrate that RaSa outperforms existing state-of-the-art methods by 6.94%, 4.45% and 15.35% in terms of Rank@1 on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets, respectively. Code is available at: https://github.com/Flame-Chasers/RaSa.
Min Cao 0005, Daming Gao, Ziqiang Cao, Chen Chen 0036, Zhenfeng Fan, Liqiang Nie, Min Zhang 0005
IJCAI6
2023 Towards Fine-Grained Optimal 3D Face Dense Registration: An Iterative Dividing and Diffusing Method
Zhenfeng Fan, Silong Peng, Shihong Xia
Int. J. Comput. Vis.1
2023 A landmark-free approach for automatic, dense and robust correspondence of 3D faces
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Xiaolian Wang, Silong Peng
Pattern Recognit.1
2022 Learning to Detect 3D Facial Landmarks via Heatmap Regression with Graph Convolutional Network
abstract
3D facial landmark detection is extensively used in many research fields such as face registration, facial shape analysis, and face recognition. Most existing methods involve traditional features and 3D face models for the detection of landmarks, and their performances are limited by the hand-crafted intermediate process. In this paper, we propose a novel 3D facial landmark detection method, which directly locates the coordinates of landmarks from 3D point cloud with a well-customized graph convolutional network. The graph convolutional network learns geometric features adaptively for 3D facial landmark detection with the assistance of constructed 3D heatmaps, which are Gaussian functions of distances to each landmark on a 3D face. On this basis, we further develop a local surface unfolding and registration module to predict 3D landmarks from the heatmaps. The proposed method forms the first baseline of deep point cloud learning method for 3D facial landmark detection. We demonstrate experimentally that the proposed method exceeds the existing approaches by a clear margin on BU-3DFE and FRGC datasets for landmark localization accuracy and stability, and also achieves high-precision results on a recent large-scale dataset.
Zhenfeng Fan, Silong Peng
AAAI3
2022 GELibRec: Third-Party Libraries Recommendation Using Graph Neural Network
Chengming Zou, Zhenfeng Fan
DASFAA (2)2
2021 Towards effective learning for face super-resolution with shape and pose perturbations
Xiyuan Hu, Zhenfeng Fan, Xu Jia 0012, Xuyun Zhang, Lianyong Qi, Zuxing Xuan
Knowl. Based Syst.2
2020 Illuminating Vehicles With Motion Priors For Surveillance Vehicle Detection
abstract
Vehicle detection in traffic surveillance videos is a special subtask in object detection, where desired objects are vehicles moving on the road while the background is still within a sequence. The disparity of speed within each frame, i.e. moving and static, is consistent with the vehicle and background semantic to some extent, thus motions can be extracted to enhance the appearance of foreground. In this paper, we propose a motion prior embedded parallel architecture for vehicle detection, aiming at illuminating vehicles and suppressing false positives in the background. We further implement extensive experiments on the UA-DETRAC dataset to validate the effectiveness of our approach, and achieve promising performance in both accuracy and speed.
Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Zhenfeng Fan, Silong Peng
ICIP4
2019 Boosting Local Shape Matching for Dense 3D Face Correspondence
abstract
Dense 3D face correspondence is a fundamental and challenging issue in the literature of 3D face analysis. Correspondence between two 3D faces can be viewed as a non-rigid registration problem that one deforms into the other, which is commonly guided by a few facial landmarks in many existing works. However, the current works seldom consider the problem of incoherent deformation caused by landmarks. In this paper, we explicitly formulate the deformation as locally rigid motions guided by some seed points, and the formulated deformation satisfies coherent local motions everywhere on a face. The seed points are initialized by a few landmarks, and are then augmented to boost shape matching between the template and the target face step by step, to finally achieve dense correspondence. In each step, we employ a hierarchical scheme for local shape registration, together with a Gaussian reweighting strategy for accurate matching of local features around the seed points. In our experiments, we evaluate the proposed method extensively on several datasets, including two publicly available ones: FRGC v2.0 and BU-3DFE. The experimental results demonstrate that our method can achieve accurate feature correspondence, coherent local shape motion, and compact data representation. These merits actually settle some important issues for practical applications, such as expressions, noise, and partial data.
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Silong Peng
CVPR1
2019 Improving Object Detection with Consistent Negative Sample Mining
Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Zhenfeng Fan, Silong Peng
ICONIP (2)4
2018 Dense Semantic and Topological Correspondence of 3D Faces without Landmarks
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Silong Peng
ECCV (16)1
2016 An Elaborately Designed Virtual Frame to Level Aeromagnetic Data
abstract
Aeromagnetic data are usually contaminated by line-to-line errors, which are most often visible as stripe patterns in a 2-D data map. Leveling is a critical step to eliminate these errors in data processing and interpretation. The conventional tie-line leveling technique involves control lines (tie lines) which are perpendicular to flight lines to extract the leveling errors. By contrast, there are also some other methods to level without the need for tie lines, which can save considerable survey cost. However, a main setback when leveling without tie lines is to distinguish the cross-line gradients from the leveling errors, making the leveled results inferior. In this letter, we revisit tie-line leveling as a nonlinear spatial-filtering technique and analyze the inadequacy of it, based on which a new approach is developed to extract the leveling errors without additional tie lines. The approach includes two steps: designing a virtual frame and leveling based on this frame. The virtual frame is elaborately designed to ensure that each extracted point can represent the local field-intensity level of a flight line and each virtual cross line can bypass the localized areas of magnetic anomalies; hence, some difficulties when leveling without tie lines are overcome. Afterward, two real field data examples are tested to demonstrate the effectiveness of this approach.
Zhenfeng Fan, Ling Huang 0007, Xiaojuan Zhang 0001, Guangyou Fang
IEEE Geosci. Remote. Sens. Lett.1