VLDB 2026 Research / reviewers in the wild / expert
Hongyu Yang 0001
dblp:57/5473-1
· DBLP profile ↗
34ranked-venue papers
5as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 21 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry FusionabstractGenerating high-quality, controllable, and structurally consistent 3D scenes in complex multi-object environments remains a fundamental challenge. We present SceneGenesis, a unified framework that synthesizes 3D scenes by combining semantic structural priors with mesh-guided video–geometry fusion. SceneGenesis first employs large language models to convert textual descriptions into category-aware object specifications, which are transformed into structured meshes using procedural approximations and pretrained asset generators, enabling precise layout control and scalable scene construction. To obtain rich and style-controllable appearances, SceneGenesis generates multi-view video representations conditioned on the initialized structure. A mesh-guided video–geometry fusion module then consolidates video evidence with mesh priors through mesh-conditioned fragment initialization, progressive geometric refinement, and structure-aware optimization, substantially improving global geometric fidelity and visual realism. Experiments demonstrate that SceneGenesis supports flexible style variation and object-level editing while achieving strong controllability, scalability, and structural quality. Yueming Zhao, Hongyu Yang 0001, Di Huang 0001 |
AAAI | 2 |
| 2026 | Learning storage-efficient 3D Gaussian head avatars from monocular videos via parametric adaptation and material decomposition
Guohao Li 0010, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
Frontiers Comput. Sci. | 2 |
| 2025 | Micro-macro Wavelet-based Gaussian Splatting for 3D Reconstruction from Unconstrained Imagesabstract3D reconstruction from unconstrained image collections presents substantial challenges due to varying appearances and transient occlusions. In this paper, we introduce Micro-macro Wavelet-based Gaussian Splatting (MW-GS), a novel approach designed to enhance 3D reconstruction by disentangling scene representations into global, refined, and intrinsic components. The proposed method features two key innovations: Micro-macro Projection, which allows Gaussian points to capture details from feature maps across multiple scales with enhanced diversity; and Wavelet-based Sampling, which leverages frequency domain information to refine feature representations and significantly improve the modeling of scene appearances. Additionally, we incorporate a Hierarchical Residual Fusion Network to seamlessly integrate these features. Extensive experiments demonstrate that MW-GS delivers state-of-the-art rendering performance, surpassing existing methods. Chengxin Lv, Hongyu Yang 0001, Di Huang 0001 |
AAAI | 3 |
| 2025 | 3D²-Actor: Learning Pose-Conditioned 3D-Aware Denoiser for Realistic Gaussian Avatar ModelingabstractAdvancements in neural implicit representations and differentiable rendering have markedly improved the ability to learn animatable 3D avatars from sparse multi-view RGB videos. However, current methods that map observation space to canonical space often face challenges in capturing pose-dependent details and generalizing to novel poses. While diffusion models have demonstrated remarkable zero-shot capabilities in 2D image generation, their potential for creating animatable 3D avatars from 2D inputs remains underexplored. In this work, we introduce 3D²-Actor, a novel approach featuring a pose-conditioned 3D-aware human modeling pipeline that integrates iterative 2D denoising and 3D rectifying steps. The 2D denoiser, guided by pose cues, generates detailed multi-view images that provide the rich feature set necessary for high-fidelity 3D reconstruction and pose rendering. Complementing this, our Gaussian-based 3D rectifier renders images with enhanced 3D consistency through a two-stage projection strategy and a novel local coordinate representation. Additionally, we propose an innovative sampling strategy to ensure smooth temporal continuity across frames in video synthesis. Our method effectively addresses the limitations of traditional numerical solutions in handling ill-posed mappings, producing realistic and animatable 3D human avatars. Experimental results demonstrate that 3D²-Actor excels in high-fidelity avatar modeling and robustly generalizes to novel poses. Zichen Tang, Hongyu Yang 0001, Hanchen Zhang, Jiaxin Chen 0002, Di Huang 0001 |
AAAI | 2 |
| 2025 | GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion PriorabstractText-guided 3D human generation has advanced with the development of efficient 3D representations and 2D-lifting methods like Score Distillation Sampling (SDS). However, current methods suffer from prolonged training times and often produce results that lack fine facial and garment details. In this paper, we propose GaussianIP, an effective two-stage framework for generating identity-preserving realistic 3D humans from text and image prompts. Our core insight is to leverage human-centric knowledge to facilitate the generation process. In stage 1, we propose a novel Adaptive Human Distillation Sampling (AHDS) method to rapidly generate a 3D human that maintains high identity consistency with the image prompt and achieves a realistic appearance. Compared to traditional SDS methods, AHDS better aligns with the human-centric generation process, enhancing visual quality with notably fewer training steps. To further improve the visual quality of the face and clothes regions, we design a View-Consistent Refinement (VCR) strategy in stage 2. Specifically, it produces detail-enhanced results of the multi-view images from stage 1 iteratively, ensuring the 3D texture consistency across views via mutual attention and distance-guided attention fusion. Then a polished version of the 3D human can be achieved by directly perform reconstruction with the refined images. Extensive experiments demonstrate that GaussianIP outperforms existing methods in both visual quality and training efficiency, particularly in generating identity-preserving results. Our code is available at: https://github.com/silence-tang/GaussianIP. Zichen Tang, Miaomiao Cui, Liefeng Bo, Hongyu Yang 0001 |
CVPR | 5 |
| 2025 | Generating Editable Head Avatars with 3D Gaussian GANsabstractGenerating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields (NeRF), achieve photo-realistic and view-consistent 3D head synthesis. However, these methods face limitations in deformation flexibility and editability, hindering the creation of lifelike and easily modifiable 3D heads. We propose a novel approach that enhances the editability and animation control of 3D head avatars by incorporating 3D Gaussian Splatting (3DGS) as an explicit 3D representation. This method enables easier illumination control and improved editability. Central to our approach is the Editable Gaussian Head (EG-Head) model, which combines a 3D Morphable Model (3DMM) with texture maps, allowing precise expression control and flexible texture editing for accurate animation while preserving identity. To capture complex non-facial geometries like hair, we use an auxiliary set of 3DGS and tri-plane features. Extensive experiments demonstrate that our approach delivers high-quality 3D-aware synthesis with state-of-the-art controllability. Our code and models are available at https://github.com/liguohao96/EGG3D. Guohao Li 0010, Hongyu Yang 0001, Yifang Men, Di Huang 0001, Weixin Li 0001, Ruijie Yang, Yunhong Wang 0001 |
ICASSP | 2 |
| 2025 | DreamScape: 3D Scene Creation via Gaussian Splatting joint Correlation ModelingabstractRecent advances in text-to-3D creation integrate the potent prior of Diffusion Models from text-to-image generation into 3D domain. Nevertheless, generating 3D scenes with multiple objects remains challenging. Therefore, we present DreamScape, a method for generating 3D scenes from text. Utilizing Gaussian Splatting for 3D representation, DreamScape introduces 3D Gaussian Guide that encodes semantic primitives, spatial transformations and relationships from text using LLMs, enabling local-to-global optimization. Progressive scale control is tailored during local object generation, addressing training instability issue arising from simple blending in the global optimization stage. Collision relationships between objects are modeled at the global level to mitigate biases in LLMs priors, ensuring physical correctness. Additionally, to generate pervasive objects like rain and snow distributed extensively across the scene, we design specialized sparse initialization and densification strategy. Experiments demonstrate that DreamScape achieves state-of-the-art performance, enabling high-fidelity, controllable 3D scene generation. Yueming Zhao, Xuening Yuan, Hongyu Yang 0001, Di Huang 0001 |
ICME | 3 |
| 2025 | ImFace++: A Sophisticated Nonlinear 3D Morphable Face Model With Implicit Neural RepresentationsabstractAccurate representations of 3D faces are of paramount importance in various computer vision and graphics applications. However, the challenges persist due to the limitations imposed by data discretization and model linearity, which hinder the precise capture of identity and expression clues in current studies. This paper presents a novel 3D morphable face model, named ImFace++, to learn a sophisticated and continuous space with implicit neural representations. ImFace++ first constructs two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, which simultaneously facilitate automatic learning of point-to-point correspondences across diverse facial shapes. To capture more sophisticated facial details, a refinement displacement field within the template space is further incorporated, enabling fine-grained learning of individual-specific facial details. Furthermore, a Neural Blend-Field is designed to reinforce the representation capabilities through adaptive blending of an array of local fields. In addition to ImFace++, we devise an improved learning strategy to extend expression embeddings, allowing for a broader range of expression variations. Comprehensive qualitative and quantitative evaluation demonstrates that ImFace++ significantly advances the state-of-the-art in terms of both face reconstruction fidelity and correspondence accuracy. Mingwu Zheng, Hongyu Yang 0001, Liming Chen 0002, Di Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | CoupleFER: Dynamic Cross-Modal Fusion via Prompt Learning for Improved 2D+3D FERabstractThe integration of 2D texture information and 3D geometric data has shown great promise in advancing the accuracy and robustness of 2D+3D facial expression recognition (FER) systems. Traditional methods in this domain often rely on projecting 3D data onto 2D maps, which limits the effective utilization of critical 3D features. To address this, we introduce CoupleFER, a novel approach that utilizes a cross-modal fusion strategy by combining image-based and point cloud-based networks. Unlike conventional multi-modal fusion methods, CoupleFER introduces the Cross-Modal Prompt Fusion (CouPle) module, enabling dynamic and interactive fusion between the two branches at every layer. This allows 2D texture information to serve as a guiding prompt, thereby enhancing the performance of the 3D FER branch. To further boost robustness and generalization, we propose a dual-level supervision mechanism, which imposes constraints at both the cluster and sample levels during training. Extensive experiments on the widely used BU-3DFE and Bosphorus datasets demonstrate that CoupleFER outperforms state-of-the-art methods, achieving superior recognition accuracy. Ablation studies validate the importance of each key component of the framework, underscoring its potential to significantly improve the performance of 2D + 3D FER systems, and robustness tests demonstrate its stability. Hebeizi Li, Hongyu Yang 0001, Di Huang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Initno: Boosting Text-to-Image Diffusion Models via Initial Noise OptimizationabstractRecent strides in the development of diffusion models, ex-emplified by advancements such as Stable Diffusion, have underscored their remarkable prowess in generating visu-ally compelling images. However, the imperative of achieving a seamless alignment between the generated image and the provided prompt persists as a formidable challenge. This paper traces the root of these difficulties to invalid initial noise, and proposes a solution in the form of Initial Noise Optimization (INITNO), a paradigm that refines this noise. Considering text prompts, not all random noises are effective in synthesizing semantically-faithful images. We design the cross-attention response score and the selfattention conflict score to evaluate the initial noise, bifurcating the initial latent space into valid and invalid sectors. A strategically crafted noise optimization pipeline is developed to guide the initial noise towards valid regions. Our method, validated through rigorous experimentation, shows a commendable proficiency in generating images in strict accordance with text prompts. Our code is available at https://github.com/xiefan-guo/initno. Xiefan Guo, Jinlin Liu, Miaomiao Cui, Jiankai Li, Hongyu Yang 0001, Di Huang 0001 |
CVPR | 5 |
| 2024 | DrFER: Learning Disentangled Representations for 3D Facial Expression RecognitionabstractFacial Expression Recognition (FER) has consistently been a focal point in the field of facial analysis. In the context of existing methodologies for 3D FER or 2D+3D FER, the extraction of expression features often gets entangled with identity information, compromising the distinctiveness of these features. To tackle this challenge, we introduce the innovative DrFER method, which brings the concept of disentangled representation learning to the field of 3D FER. DrFER employs a dual-branch framework to effectively disentangle expression information from identity information. Diverging from prior disentanglement endeavors in the 3D facial domain, we have carefully reconfigured both the loss functions and network structure to make the overall framework adaptable to point cloud data. This adaptation enhances the capability of the framework in recognizing facial expressions, even in cases involving varying head poses. Extensive evaluations conducted on the BU-3DFE and Bosphorus datasets substantiate that DrFER surpasses the performance of other 3D FER methods. Hebeizi Li, Hongyu Yang 0001, Di Huang 0001 |
FG | 2 |
| 2024 | 3D Face Modeling via Weakly-Supervised Disentanglement Network Joint Identity-Consistency PriorabstractGenerative 3D face models featuring disentangled controlling factors hold immense potential for diverse applications in computer vision and computer graphics. However, previous 3D face modeling methods face a challenge as they demand specific labels to effectively disentangle these factors. This becomes particularly problematic when integrating multiple 3D face datasets to improve the generalization of the model. Addressing this issue, this paper introduces a Weakly-Supervised Disentanglement Framework, denoted as WSDF, to facilitate the training of controllable 3D face models without an overly stringent labeling requirement. Adhering to the paradigm of Variational Autoencoders (VAEs), the proposed model achieves disentanglement of identity and expression controlling factors through a two-branch encoder equipped with dedicated identity-consistency prior. It then faithfully re-entangles these factors via a tensor-based combination mechanism. Notably, the introduction of the Neutral Bank allows precise acquisition of subject-specific information using only identity labels, thereby averting degeneration due to insufficient supervision. Additionally, the framework incorporates a label-free second-order loss function for the expression factor to regulate deformation space and eliminate extraneous information, resulting in enhanced disentanglement. Extensive experiments have been conducted to substantiate the superior performance of WSDF. Our code is available at https://github.com/liguoha096/WSDF. Guohao Li 0010, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
FG | 2 |
| 2024 | Progressive Self-supervised Representation Learning for 3D Facial Expression RecognitionabstractFacial expression recognition (FER) is a critical area of research in face analysis. While 2D data has been extensively used, 3D data offers inherent advantages, such as increased resilience to illumination and pose variations. However, the limited size of current 3D FER datasets significantly constrains the performance of 3D FER methods. To overcome this challenge, we propose a novel self-supervised pre-training scheme by leveraging large-scale external 3D data, followed by fine-tuning on 3D FER datasets. Our approach starts with self-supervised learning on a large-scale 3D point cloud object dataset, specifically ShapeNet. We then move on to the FaceScape dataset, which is primarily used for morphable face prediction. To enhance robustness, we integrate synthetic data before fine-tuning on specific FER datasets. This multi-stage process allows the model to progressively learn 3D facial expression representations from coarse to fine. For this purpose, we utilize Point-MAE, a leading self-supervised model for representation learning. To enhance its ability for FER task, we further incorporate facial priors in the masking and point sampling steps, leveraging the distinctive characteristics of facial data. Our method achieves state-of-the-art performance on both BU-3DFE and Bosphorus datasets, matching or surpassing results achieved by other 2D+3D FER techniques. Hebeizi Li, Hongyu Yang 0001, Di Huang 0001 |
IJCB | 2 |
| 2024 | ArtNeRF: A Stylized Neural Field for 3D-Aware Artistic Face Synthesis
Zichen Tang, Hongyu Yang 0001 |
ICPR (25) | 2 |
| 2024 | SA3WT: Adaptive Wavelet-Based Transformer with Self-Paced Auto Augmentation for Face Forgery Detection
Hongyu Yang 0001, Binghui Chen, Di Huang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Deep Common Feature Mining for Efficient Video Semantic SegmentationabstractRecent advancements in video semantic segmentation have made substantial progress by exploiting temporal correlations. Nevertheless, persistent challenges, including redundant computation and the reliability of the feature propagation process, underscore the need for further innovation. In response, we present Deep Common Feature Mining (DCFM), a novel approach strategically designed to address these challenges by leveraging the concept of feature sharing. DCFM explicitly decomposes features into two complementary components. The common representation extracted from a key-frame furnishes essential high-level information to neighboring non-key frames, allowing for direct re-utilization without feature propagation. Simultaneously, the independent feature, derived from each video frame, captures rapidly changing information, providing frame-specific clues crucial for segmentation. To achieve such decomposition, we employ a symmetric training strategy tailored for sparsely annotated data, empowering the backbone to learn a robust high-level representation enriched with common information. Additionally, we incorporate a self-supervised loss function to reinforce intra-class feature similarity and enhance temporal consistency. Experimental evaluations on the VSPW and Cityscapes datasets demonstrate the effectiveness of our method, showing a superior balance between accuracy and efficiency. The implementation is available athttps://github.com/BUAAHugeGun/DCFM. Yaoyan Zheng, Hongyu Yang 0001, Di Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Learning Polysemantic Spoof Trace: A Multi-Modal Disentanglement Network for Face Anti-spoofingabstractAlong with the widespread use of face recognition systems, their vulnerability has become highlighted. While existing face anti-spoofing methods can be generalized between attack types, generic solutions are still challenging due to the diversity of spoof characteristics. Recently, the spoof trace disentanglement framework has shown great potential for coping with both seen and unseen spoof scenarios, but the performance is largely restricted by the single-modal input. This paper focuses on this issue and presents a multi-modal disentanglement model which targetedly learns polysemantic spoof traces for more accurate and robust generic attack detection. In particular, based on the adversarial learning mechanism, a two-stream disentangling network is designed to estimate spoof patterns from the RGB and depth inputs, respectively. In this case, it captures complementary spoofing clues inhering in different attacks. Furthermore, a fusion module is exploited, which recalibrates both representations at multiple stages to promote the disentanglement in each individual modality. It then performs cross-modality aggregation to deliver a more comprehensive spoof trace representation for prediction. Extensive evaluations are conducted on multiple benchmarks, demonstrating that learning polysemantic spoof traces favorably contributes to anti-spoofing with more perceptible and interpretable results. Kaicheng Li, Hongyu Yang 0001, Binghui Chen, Di Huang 0001 |
AAAI | 2 |
| 2023 | NeuFace: Realistic 3D Neural Face Rendering from Multi-View ImagesabstractRealistic face rendering from multi-view images is beneficial to various computer vision and graphics applications. Due to complex spatially-varying reflectance properties and geometry characteristics of faces, however, it remains challenging to recover 3D facial representations both faithfully and efficiently in the current studies. This paper presents a novel 3D face rendering model, namely NeuFace, to learn accurate and physically-meaningful underlying 3D representations by neural rendering techniques. It naturally in-corporates the neural BRDFs into physically based rendering, capturing sophisticated facial geometry and appearance clues in a collaborative manner. Specifically, we introduce an approximated BRDF integration and a simple yet new low-rank prior, which effectively lower the ambiguities and boost the performance of the facial BRDFs. Extensive experiments are performed to demonstrate the superiority of NeuFace in human face rendering, along with a decent generalization ability to common objects. Code is released at NeuFace. Mingwu Zheng, Hongyu Yang 0001, Di Huang 0001 |
CVPR | 3 |
| 2023 | Denoising Diffusion Autoencoders are Unified Self-supervised LearnersabstractInspired by recent advances in diffusion models, which are reminiscent of denoising autoencoders, we investigate whether they can acquire discriminative representations for classification via generative pre-training. This paper shows that the networks in diffusion models, namely denoising diffusion autoencoders (DDAE), are unified self-supervised learners: by pre-training on unconditional image generation, DDAE has already learned strongly linear-separable representations within its intermediate layers without auxiliary encoders, thus making diffusion pre-training emerge as a general approach for generative-and-discriminative dual learning. To validate this, we conduct linear probe and finetuning evaluations. Our diffusion-based approach achieves 95.9% and 50.0% linear evaluation accuracies on CIFAR-10 and Tiny-ImageNet, respectively, and is comparable to contrastive learning and masked autoencoders for the first time. Transfer learning from ImageNet also confirms the suitability of DDAE for Vision Transformers, suggesting the potential to scale DDAEs as unified foundation models. Code is available at github.com/FutureXiang/ddae. Weilai Xiang, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
ICCV | 2 |
| 2022 | ABPN: Adaptive Blend Pyramid Network for Real-Time Local Retouching of Ultra High-Resolution PhotoabstractPhoto retouching finds many applications in various fields. However, most existing methods are designed for global retouching and seldom pay attention to the local region, while the latter is actually much more tedious and time-consuming in photography pipelines. In this paper, we propose a novel adaptive blend pyramid network, which aims to achieve fast local retouching on ultra high-resolution photos. The network is mainly composed of two components: a context-aware local retouching layer (LRL) and an adaptive blend pyramid layer (BPL). The LRL is designed to implement local retouching on low-resolution images, giving full consideration of the global context and local texture information, and the BPL is then developed to progressively expand the low-resolution results to the higher ones, with the help of the proposed adaptive blend module and refining module. Our method outperforms the existing methods by a large margin on two local photo retouching tasks and exhibits excellent performance in terms of running speed, achieving real-time inference on 4K images with a single NVIDIA Tesla P100 GPU. Moreover, we introduce the first high-definition cloth retouching dataset CRHD-3K to promote the research on local photo retouching. The dataset is available at https://github.com/youngLbw/crhd-3K. Biwen Lei, Xiefan Guo, Hongyu Yang 0001, Miaomiao Cui, Xuansong Xie, Di Huang 0001 |
CVPR | 3 |
| 2022 | ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural RepresentationsabstractPrecise representations of 3D faces are beneficial to various computer vision and graphics applications. Due to the data discretization and model linearity, however, it remains challenging to capture accurate identity and expression clues in current studies. This paper presents a novel 3D morphable face model, namely ImFace, to learn a nonlinear and continuous space with implicit neural representations. It builds two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, and designs an improved learning strategy to extend embeddings of expressions to allow more diverse changes. We further introduce a Neural Blend-Field to learn sophisticated details by adaptively blending a series of local fields. In addition to ImFace, an effective pre-processing pipeline is proposed to address the issue of watertight input requirement in implicit representations, enabling them to work with common facial surfaces for the first time. Extensive experiments are performed to demonstrate the superiority of ImFace. Mingwu Zheng, Hongyu Yang 0001, Di Huang 0001, Liming Chen 0002 |
CVPR | 2 |
| 2022 | Multi-view Gait Video SynthesisabstractThis paper investigates a new fine-grained video generation task, namely Multi-view Gait Video Synthesis, where the generation model works on a video of a walking human of arbitrary viewpoint and creates multi-view renderings of the subject. This task is particularly challenging, as it requires synthesizing visually plausible results, while simultaneously preserving discriminative gait cues subject to identification. To tackle the challenge caused by the entanglement of viewpoint, texture, and body structure, we present a network with two collaborative branches to decouple the novel view rendering process into two streams for human appearances (texture) and silhouettes (structure), respectively. Additionally, the prior knowledge of person re-identification and gait recognition is incorporated into the training loss for more adequate and accurate dynamic details. Experimental results show that the presented method is able to achieve promising success rates when attacking state-of-the-art gait recognition models. Furthermore, the method can improve gait recognition systems by effective data augmentation. To the best of our knowledge, this is the first task to manipulate views for human videos with person-specific behavioral constraints. Weilai Xiang, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
ACM Multimedia | 2 |
| 2021 | Boundary Guided Context Aggregation for Semantic Segmentation
Hongyu Yang 0001, Di Huang 0001 |
BMVC | 2 |
| 2021 | Image Inpainting via Conditional Texture and Structure Dual GenerationabstractDeep generative approaches have recently made considerable progress in image inpainting by introducing structure priors. Due to the lack of proper interaction with image texture during structure reconstruction, however, current solutions are incompetent in handling the cases with large corruptions, and they generally suffer from distorted results. In this paper, we propose a novel two-stream network for image inpainting, which models the structure-constrained texture synthesis and texture-guided structure reconstruction in a coupled manner so that they better leverage each other for more plausible generation. Furthermore, to enhance the global consistency, a Bi-directional Gated Feature Fusion (Bi-GFF) module is designed to exchange and combine the structure and texture information and a Contextual Feature Aggregation (CFA) module is developed to refine the generated contents by region affinity learning and multi-scale feature aggregation. Qualitative and quantitative experiments on the CelebA, Paris StreetView and Places2 datasets demonstrate the superiority of the proposed method. Our code is available at https://github.com/Xiefan-Guo/CTSDG. Xiefan Guo, Hongyu Yang 0001, Di Huang 0001 |
ICCV | 2 |
| 2021 | Intensity enhancement via GAN for multimodal face expression recognition
Hongyu Yang 0001, Kangkang Zhu, Di Huang 0001, Hebeizi Li, Yunhong Wang 0001, Liming Chen 0002 |
Neurocomputing | 1 |
| 2021 | Learning Continuous Face Age Progression: A Pyramid of GANsabstractThe two underlying requirements of face age progression, i.e., aging accuracy and identity permanence, are not well studied in the literature. This paper presents a novel generative adversarial network based approach to address the issues in a coupled manner. It separately models the constraints for the intrinsic subject-specific characteristics and the age-specific facial changes with respect to the elapsed time, ensuring that the generated faces present desired aging effects while keeping personalized properties stable. To render photo-realistic facial details, high-level age-specific features conveyed by the synthesized face are estimated by a pyramidal adversarial discriminator at multiple scales, which simulates the aging effects in a finer way. Further, an adversarial learning scheme is introduced to simultaneously train a single generator and multiple parallel discriminators, resulting in smooth continuous face aging sequences. The proposed method is applicable even in the presence of variations in pose, expression, makeup, etc., achieving remarkably vivid aging effects. Quantitative evaluations by a COTS face recognition system demonstrate that the target age distributions are accurately recovered, and 99.88 and 99.98 percent age progressed faces can be correctly verified at 0.001 percent FAR after age transformations of approximately 28 and 23 years elapsed time on the MORPH and CACD databases, respectively. Both visual and quantitative assessments show that the approach advances the state-of-the-art. Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | 3D Face Mask Anti-spoofing via Deep Fusion of Dynamic Texture and Shape CluesabstractFace anti-spoofing has recently become more important to the wide application of Face Recognition (FR) techniques. Compared to Spoofing Attacks (SAs) of printed photos and replayed videos, 3D masks bring more challenges to FR systems. This paper proposes a novel approach to 3D face mask anti-spoofing, namely Multi-Modal Dynamics Fusion Network (MM-DFN), and different from the overwhelming majority of the methods in the literature that only employ RGB data, it highlights the credit of the geometry information delivered by depth sensors or reconstructed from RGB images. Dynamic texture and shape clues are respectively encoded by a two-branch deep CNN model at different rates so that discriminative details are sufficiently captured, and they are combined at intervals for more comprehensive description. Moreover, a 3D model guided data augmentation method is applied to generate a diversity of samples with various poses, which further enhances the anti-spoofing model. The proposed method is extensively evaluated on three public benchmarks, i.e. 3DMAD, HKBU-MARs V1 and SMAD, and the results achieved are state-of-the-art, demonstrating its effectiveness for this issue. Weixin Li 0001, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
FG | 3 |
| 2020 | Pixel Sampling for Style Preserving Face Pose EditingabstractThe existing auto-encoder based face pose editing methods primarily focus on modeling the identity preserving ability during pose synthesis, but are less able to preserve the image style properly, which refers to the color, brightness, saturation, etc. In this paper, we take advantage of the well-known frontal/profile optical illusion and present a novel two-stage approach to solve the aforementioned dilemma, where the task of face pose manipulation is cast into face inpainting. By selectively sampling pixels from the input face and slightly adjust their relative locations with the proposed “Pixel Attention Sampling” module, the face editing result faithfully keeps the identity information as well as the image style unchanged. By leveraging high-dimensional embedding at the inpainting stage, finer details are generated. Further, with the 3D facial landmarks as guidance, our method is able to manipulate face pose in three degrees of freedom, i.e., yaw, pitch, and roll, resulting in more flexible face pose editing than merely controlling the yaw angle as usually achieved by the current state-of-the-art. Both the qualitative and quantitative evaluations validate the superiority of the proposed approach. Xiangnan Yin, Di Huang 0001, Hongyu Yang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002 |
IJCB | 3 |
| 2020 | Intensity Enhancement Via Gan for Multimodal Facial Expression RecognitionabstractFace expression recognition (FER) on low intensity is not well studied in the literature. This paper investigates this new problem and presents a novel Generative Adversarial Network (GAN) based multimodal approach to it. The method models the tasks of intensity enhancement and expression recognition jointly, ensuring that the synthesize faces not only present expression of high intensity, but also truly contribute to promoting the performance of FER. Extensive experiments are conducted on the BU-3DFE and BU-4DFE datasets. State-of-the-art FER performance clearly validates the effectiveness of the proposed method. Kangkang Zhu, Yunhong Wang 0001, Hongyu Yang 0001, Di Huang 0001, Liming Chen 0002 |
ICIP | 3 |
| 2019 | Attacking Gait Recognition Systems via Silhouette Guided GANsabstractThis paper investigates a new attack method to gait recognition systems. Different from typical spoofing attacks that require impostors to mimic certain clothing or walking styles, it proposes to intercept the video stream captured by the on-site camera and replace it with synthesized samples. To this end, we present a novel Generative Adversarial Network (GAN) based approach, which is able to render a faked video from the source walking sequence of a specified subject and the target scene image with both good visual effects and sufficient discriminative details. A new generator architecture is built, where the features of the source foreground sequence and the target background image are combined at multiple scales, making the synthesized video vivid. To fool recognition systems, the silhouette-conditioned losses are specially designed to constrain the static and dynamic consistency between the subjects in the source and generated videos. The person re-identification similarity based triplet loss is exploited to guide the generator, which keeps the personalized appearance properties stable. The edge and flow-related losses further regulate the generation of the attacking video. Two state-of-the-art gait recognition systems are used for evaluation, namely GaitSet and CNN-Gait, and we analyze their performance under attacking. Both the visual fidelity and attacking ability of the generated videos validate the effectiveness of the proposed method. Meijuan Jia, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
ACM Multimedia | 2 |
| 2018 | Learning Face Age Progression: A Pyramid Architecture of GANsabstractThe two underlying requirements of face age progression, i.e. aging accuracy and identity permanence, are not well studied in the literature. In this paper, we present a novel generative adversarial network based approach. It separately models the constraints for the intrinsic subject-specific characteristics and the age-specific facial changes with respect to the elapsed time, ensuring that the generated faces present desired aging effects while simultaneously keeping personalized properties stable. Further, to generate more lifelike facial details, high-level age-specific features conveyed by the synthesized face are estimated by a pyramidal adversarial discriminator at multiple scales, which simulates the aging effects in a finer manner. The proposed method is applicable to diverse face samples in the presence of variations in pose, expression, makeup, etc., and remarkably vivid aging effects are achieved. Both visual fidelity and quantitative evaluations show that the approach advances the state-of-the-art. Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Anil K. Jain 0001 |
CVPR | 1 |
| 2017 | Facial aging simulation via tensor completion and metric learningabstractFacial aging simulation is one of the most challenging issues in automatic machine based face analysis, where the most essential requirements are (i) human identity should remain stable in texture synthesis and (ii) the texture synthesised is expected to accord with human cognitive perception in aging. In this study, the authors propose a tensor completion based method to transform the simulation task to a standard matrix completion one. To protect human dependent characteristics during texture synthesis, the proposed method processes the two major components, i.e. identity and age, in different channels. Furthermore, they incorporate prior information in such a process, assuming that the textures of different subjects in the same age group are similar and similar looking people tend to age in similar ways, and the metric learning technique is adopted to measure the similarity between identities so that the faces that have the highest similarities with the one in the test image are assigned bigger weights in texture generation. In addition, shape deformation is also considered to make the synthesised images more natural. Experimental results achieved on the FG‐NET database demonstrate the effectiveness of the proposed method. Di Huang 0001, Yunhong Wang 0001, Hongyu Yang 0001 |
IET Comput. Vis. | 4 |
| 2016 | Face Aging Effect Simulation Using Hidden Factor Analysis Joint Sparse RepresentationabstractFace aging simulation has received rising investigations nowadays, whereas it still remains a challenge to generate convincing and natural age-progressed face images. In this paper, we present a novel approach to such an issue using hidden factor analysis joint sparse representation. In contrast to the majority of tasks in the literature that integrally handle the facial texture, the proposed aging approach separately models the person-specific facial properties that tend to be stable in a relatively long period and the age-specific clues that gradually change over time. It then transforms the age component to a target age group via sparse reconstruction, yielding aging effects, which is finally combined with the identity component to achieve the aged face. Experiments are carried out on three face aging databases, and the results achieved clearly demonstrate the effectiveness and robustness of the proposed method in rendering a face with aging effects. In addition, a series of evaluations prove its validity with respect to identity preservation and aging effect generation. Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 1 |
| 2014 | Age invariant face recognition based on texture embedded discriminative graph modelabstractIn an automatic face recognition system, it still remains a challenge to improve the robustness to aging. In this paper, we present a novel approach to address age invariant face recognition, by formulating it as a graph matching problem. In contrast to the majority of tasks in the literature that only make use of robust texture features, this method generates a graph from a set of fiducial landmarks of each face, which captures the texture clues that tend to be stable in a period as well as the common facial geometry configuration. The nodes of the graph denote the texture of a face area around a landmark, and the edges correspond to the geometry topology of the face. For each area, the age invariant texture information is extracted by a discriminative and compact feature encoded in the Local Gabor Binary Pattern Histogram Sequence (LGBPHS) projected in an LDA subspace. An objective function is then designed to match graphs for the purpose of registration and identification. Experiments are carried out on the FG-NET Aging database, and the results achieved outperform the state of the art ones, which clearly demonstrate the effectiveness and robustness of the proposed method in face recognition across age variations. Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001 |
IJCB | 1 |