VLDB 2026 Research / reviewers in the wild / expert
Chenyang Wang 0002
dblp:163/7308-2
· DBLP profile ↗
25ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-1156-9193ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion ModelabstractWe propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants while also embedding identity-decoupled style into generated gestures that enhance realism and expressiveness. To ensure precise synchronization between interlocutors, DialoGen adopts an interactive dual-diffusion model with mutual interaction estimation, which integrates interaction correlation into the diffusion process. More importantly, by leveraging supervised contrastive learning, we develop the identity-decoupled style guidance to adaptively decompose the identity-specific style of interlocutors into latent space, enabling multi-style dialog gesture generation. Extensive experimental results demonstrate that our model significantly outperforms existing methods in generating realistic, speech-aligned, identity-specific gestures, offering a high-quality solution for various dialog scenarios. Weiyu Zhao, Chenyang Wang 0002, Liangxiao Hu, Zonglin Li 0004, Wei Yu 0004, Shengping Zhang |
AAAI | 2 |
| 2026 | A Natural Language Guided Approach for Blind Face Restoration: Methodology and DatasetabstractBlind Face Restoration (BFR) aims to reconstruct high-quality face images from low-quality inputs without any prior knowledge of the specific degradation types or levels. In recent years, remarkable progress has been achieved, particularly through GAN- and diffusion-based approaches, which have greatly improved perceptual realism and reconstruction fidelity. However, existing approaches typically rely solely on visual cues from degraded images. This often results in inaccurate reconstruction of facial details and noticeable identity distortion, particularly under severe or complex degradations. To address these limitations, we incorporate auxiliary textual information into BFR to enable the recovery of subtle facial attributes, such as wrinkles, moles, and skin marks that are often overlooked or hard to reconstruct by conventional visual priors. To support this idea, we first construct a large-scale dataset containing 30,000 detailed textual descriptions paired with CelebA-HQ face images, explicitly designed to capture fine-grained facial semantics. To effectively bridge the gap between visual data and natural language, we further propose FaceCLIP, a fine-tuned vision-language model specifically tailored to the human face. FaceCLIP enables more accurate alignment between face images and their corresponding textual descriptions by effectively capturing nuanced semantic cues critical for faithful face reconstruction. Built upon these foundations, we propose Text-guided Blind Face Restoration (TBFR), a novel diffusion-based framework that explicitly integrates textual guidance into the face restoration pipeline. Within TBFR, a text-guided hybrid attention block is designed to effectively fuse visual and textual features, while a text-aware loss is employed to enforce semantic consistency between the generated images and their associated textual descriptions. Extensive experimental results show that TBFR outperforms state-of-the-art BFR methods in terms of both quantitative metrics and subjective perceptual quality, establishing a new benchmark for BFR tasks. Wenjie An, Chenyang Wang 0002, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Federated domain generalization via data-centric flatness optimization
Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005, Xiangyang Ji |
Pattern Recognit. | 1 |
| 2026 | D3BSR: Blind Super-Resolution via Diffusion-Based Disentangled Degradation RepresentationabstractExisting Blind Super-Resolution (BSR) methods are mostly trained on artificial synthetic degradation data pairs or rely on specific degradation priors, which lead to poor performance due to the trained degradation mismatch between other unknown complex degradations in real-world scenarios. To tackle this problem, we propose a novel Diffusion-based Disentangled Degradation representation method for BSR, dubbed D3BSR, which disentangles arbitrary unknown degradation into structure and texture degradations to enhance perception and fidelity quality individually. Specifically, the structure degradation is optimized by degradation distribution transition with a self-supervised collaborative learning strategy to recursively minimize the perception error. The texture degradation is restored through posterior sampling controlled by a fidelity coefficient to leverage rich texture priors encapsulated in a pre-trained diffusion model for preserving fidelity. The degraded image is super-resolved using an analytical solution with the pseudo inverse of the structural and texture degradation, which achieves a controllable trade-off between perception and fidelity and does not rely on any degradation priors or extra-supervised training. Extensive experiments on the nine heavily degraded synthetic and real-world natural and face datasets demonstrate that our D3BSR outperforms SOTA methods on the diverse metrics in reconstruction faithfulness and perceptual quality. Wei Yu 0004, Qinglin Liu, Quanling Meng, Chenyang Wang 0002, Xin Sun 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | Balancing Task-Invariant Interaction and Task-Specific Adaptation for Unified Image FusionabstractUnified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem facilitates task-invariant knowledge sharing, it often overlooks task-specific characteristics, thereby limiting the overall performance. Existing general image fusion methods incorporate explicit task identification to enable adaptation to different fusion tasks. However, this dependence during inference restricts the model's generalization to unseen fusion tasks. To address these issues, we propose a novel unified image fusion framework named "TITA", which dynamically balances both Task-invariant Interaction and Task-specific Adaptation. For task-invariant interaction, we introduce the Interaction-enhanced Pixel Attention (IPA) module to enhance pixel-wise interactions for better multi-source complementary information extraction. For task-specific adaptation, the Operation-based Adaptive Fusion (OAF) module dynamically adjusts operation weights based on task properties. Additionally, we incorporate the Fast Adaptive Multitask Optimization (FAMO) strategy to mitigate the impact of gradient conflicts across tasks during joint training. Extensive experiments demonstrate that TITA not only achieves competitive performance compared to specialized methods across three image fusion scenarios but also exhibits strong generalization to unseen fusion tasks. The source codes are released at https://github.com/huxingyuabc/TITA. Junjun Jiang, Chenyang Wang 0002, Kui Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ICCV | 3 |
| 2025 | REA-Listener: Real-Time Listening Head Generation with Dynamic Emotion Modeling and Flexible Modality AdaptationabstractListening head generation aims to synthesize realistic and responsive non-verbal listener head motions that respond to speakers in conversational scenarios. Existing methods typically rely on fixed audio-visual input modalities and predefined emotion labels, limiting their adaptability and expressiveness in real-world scenarios. In this paper, we propose a novel real-time framework, REA-Listener, to generate high-fidelity listening head videos with flexible modality adaptation and dynamic emotion modeling. Specifically, we first propose a Modality-Adaptive Mixture of Experts (MA-MoE) module to encode arbitrary combinations of speaker audio and visual signals into a unified embedding space, ensuring robustness under partial modality conditions. To further enhance the temporal consistency of listener emotion, we present a lightweight emotional head dynamics generator with a multi-modal emotion predictor, which infers listener emotions dynamically from speaker context alongside head motion coefficient prediction. Finally, we employ a 3D-aware renderer based on 3D Gaussian Splatting to produce high-quality listener head videos in real time. With these components, our approach achieves efficient head motion generation at 30fps on a single NVIDIA RTX 3090 GPU, supporting real-time interaction. Extensive evaluations and applications demonstrate that our method outperforms state-of-the-art methods in listening head generation. Sizhe Zhao, Chenyang Wang 0002, Weiyu Zhao, Zonglin Li 0004, Ming Li 0042, Shengping Zhang |
ACM Multimedia | 2 |
| 2025 | Enhancing consistency and mitigating bias: A data replay approach for incremental learning
Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005, Xiangyang Ji |
Neural Networks | 1 |
| 2025 | FusionINV: A Diffusion-Based Approach for Multimodal Image FusionabstractInfrared images exhibit a significantly different appearance compared to visible counterparts. Existing infrared and visible image fusion (IVF) methods fuse features from both infrared and visible images, producing a new "image" appearance not inherently captured by any existing device. From an appearance perspective, infrared, visible, and fused images belong to different data domains. This difference makes it challenging to apply fused images because their domain-specific appearance may be difficult for downstream systems, e.g., pre-trained segmentation models. Therefore, accurately assessing the quality of the fused image is challenging. To address those problem, we propose a novel IVF method, FusionINV, which produces fused images with an appearance similar to visible images. FusionINV employs the pre-trained Stable Diffusion (SD) model to invert infrared images into the noise feature space. To inject visible-style appearance information into the infrared features, we leverage the inverted features from visible images to guide this inversion process. In this way, we can embed all the information of infrared and visible images in the noise feature space, and then use the prior of the pre-trained SD model to generate visually friendly images that align more closely with the RGB distribution. Specially, to generate the fused image, we design a tailored fusion rule within the denoising process that iteratively fuses visible-style infrared and visible features. In this way, the fused image falls into the visible domain and can be directly applied to existing downstream machine systems. Thanks to advancements in image inversion, FusionINV can directly produce fused images in a training-free manner. Extensive experiments demonstrate that FusionINV achieves outstanding performance in both human visual evaluation and machine perception tasks. The code is available at https://github.com/erfect2020/FusionINV. Pengwei Liang, Junjun Jiang, Chenyang Wang 0002, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Low-Light Face Super-resolution via Illumination, Structure, and Texture Associated RepresentationabstractHuman face captured at night or in dimly lit environments has become a common practice, accompanied by complex low-light and low-resolution degradations. However, the existing face super-resolution (FSR) technologies and derived cascaded schemes are inadequate to recover credible textures. In this paper, we propose a novel approach that decomposes the restoration task into face structural fidelity maintaining and texture consistency learning. The former aims to enhance the quality of face images while improving the structural fidelity, while the latter focuses on eliminating perturbations and artifacts caused by low-light degradation and reconstruction. Based on this, we develop a novel low-light low-resolution face super-resolution framework. Our method consists of two steps: an illumination correction face super-resolution network (IC-FSRNet) for lighting the face and recovering the structural information, and a detail enhancement model (DENet) for improving facial details, thus making them more visually appealing and easier to analyze. As the relighted regions could provide complementary information to boost face super-resolution and vice versa, we introduce the mutual learning to harness the informative components from relighted regions and reconstruction, and achieve the iterative refinement. In addition, DENet equipped with diffusion probabilistic model is built to further improve face image visual quality. Experiments demonstrate that the proposed joint optimization framework achieves significant improvements in reconstruction quality and perceptual quality over existing two-stage sequential solutions. Code is available at https://github.com/wcy-cs/IC-FSRDENet. Chenyang Wang 0002, Junjun Jiang, Kui Jiang, Xianming Liu 0005 |
AAAI | 1 |
| 2024 | Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and DeblurringabstractExisting joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data, which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models, they typically rely on the assumed degradation process and cannot handle unknown real-world degradations well. To address these problems, we propose a novel zero-shot framework, FourierDiff, which embeds Fourier priors into a pre-trained diffusion model to harmoniously handle the joint degradation of luminance and structures. FourierDiff is appealing in its relaxed requirements on paired training data and degradation assumptions. The key zero-shot insight is motivated by image characteristics in the Fourier domain: most luminance information concentrates on amplitudes while structure and content information are closely related to phases. Based on this observation, we decompose the sampled results of the reverse diffusion process in the Fourier domain and take advantage of the amplitude of the generative prior to align the enhanced brightness with the distribution of natural images. To yield a sharp and content-consistent enhanced result, we further design a spatial-frequency alternating optimization strategy to progressively refine the phase of the input. Extensive experiments demonstrate the superior effectiveness of the proposed method, especially in real-world scenes. The code is available at https://github.com/aipixel/FourierDiff. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Yichen Zheng, Bineng Zhong 0001, Chongyi Li, Liqiang Nie |
CVPR | 3 |
| 2024 | DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationabstractExisting diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper, we propose a novel framework, DiffPerformer, to synthesize high-fidelity and temporally consistent human video. Without complex architecture modification or costly training, DiffPerformer finetunes a pre-trained diffusion model on a single video of the target character and introduces an implicit video representation as a proxy to learn temporally consistent guidance for the diffusion model. The guidance is encoded into VAE latent space and an iterative optimization loop is constructed between the implicit video representation and the diffusion model, allowing to harness the smooth property of the implicit video representation and the generative capabilities of the diffusion model in a mutually beneficial way. Moreover, we propose 3D-aware human flow as a temporal constraint during the optimization to explicitly model the correspondence between driving poses and human appearance. This alleviates the mis-alignment between driving poses and target performer and therefore maintains the appearance coherence under various motions. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The code is available at https://github.com/aipixel/DiffPerformer. Chenyang Wang 0002, Zerong Zheng, Tao Yu 0007, Xiaoqian Lv, Bineng Zhong 0001, Shengping Zhang, Liqiang Nie |
CVPR | 1 |
| 2024 | Shape-Guided Clothing Warping for Virtual Try-OnabstractImage-based virtual try-on aims to seamlessly fit in-shop clothing to a person image while maintaining pose consistency. Existing methods commonly employ the thin plate spline (TPS) transformation or appearance flow to deform in-shop clothing for aligning with the person's body. Despite their promising performance, these methods often lack precise control over fine details, leading to inconsistencies in shape between clothing and the person's body as well as distortions in exposed limb regions. To tackle these challenges, we propose a novel shape-guided clothing warping method for virtual try-on, dubbed SCW-VTON, which incorporates global shape constraints and additional limb textures to enhance the realism and consistency of the warped clothing and try-on results. To integrate global shape constraints for clothing warping, we devise a dual-path clothing warping module comprising a shape path and a flow path. The former path captures the clothing shape aligned with the person's body, while the latter path leverages the mapping between the pre- and post-deformation of the clothing shape to guide the estimation of appearance flow. Furthermore, to alleviate distortions in limb regions of try-on results, we integrate detailed limb guidance by developing a limb reconstruction network based on masked image modeling. Through the utilization of SCW-VTON, we are able to generate try-on results with enhanced clothing shape consistency and precise control over details. Extensive experiments demonstrate the superiority of our approach over state-of-the-art methods both qualitatively and quantitatively. Shunyuan Zheng, Zonglin Li 0004, Chenyang Wang 0002, Xin Sun 0003, Quanling Meng |
ACM Multimedia | 4 |
| 2024 | Incrementally Adapting Pretrained Model Using Network Prior for Multi-Focus Image FusionabstractMulti-focus image fusion can fuse the clear parts of two or more source images captured at the same scene with different focal lengths into an all-in-focus image. On the one hand, previous supervised learning-based multi-focus image fusion methods relying on synthetic datasets have a clear distribution shift with real scenarios. On the other hand, unsupervised learning-based multi-focus image fusion methods can well adapt to the observed images but lack the general knowledge of defocus blur that can be learned from paired data. To avoid the problems of existing methods, this paper presents a novel multi-focus image fusion model by considering both the general knowledge brought by the supervised pretrained backbone and the extrinsic priors optimized on specific testing sample to improve the performance of image fusion. To be specific, the Incremental Network Prior Adaptation (INPA) framework is proposed to incrementally integrate features extracted from the pretrained strong baselines into a tiny prior network (6.9% parameters of the backbone network) to boost the performance for test samples. We evaluate our method on both synthetic and real-world public datasets (Lytro, MFI-WHU, and Real-MFF) and show that our method outperforms existing supervised learning-based methods and unsupervised learning based methods. Junjun Jiang, Chenyang Wang 0002, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | ReSmooth: Detecting and Utilizing OOD Samples When Training With Data AugmentationabstractData augmentation (DA) is a widely used technique for enhancing the training of deep neural networks. Recent DA techniques which achieve state-of-the-art performance always meet the need for diversity in augmented training samples. However, an augmentation strategy that has a high diversity usually introduces out-of-distribution (OOD) augmented samples and these samples consequently impair the performance. To alleviate this issue, we propose ReSmooth, a framework that first detects OOD samples in augmented samples and then leverages them. To be specific, we first use a Gaussian mixture model (GMM) to fit the loss distribution of both the original and augmented samples and accordingly split these samples into in-distribution (ID) samples and OOD samples. Then we start a new training where ID and OOD samples are incorporated with different smooth labels. By treating ID samples and OOD samples unequally, we can make better use of the diverse augmented data. Furthermore, we incorporate our ReSmooth framework with negative DA (NDA) strategies. By properly handling their intentionally created OOD samples, the classification performance of NDAs is largely ameliorated. Experiments on several classification benchmarks show that ReSmooth can be easily extended to the existing augmentation strategies [such as RandAugment (RA), rotate, and jigsaw] and improve on them. Our code is available at https://github.com/Chenyang4/ReSmooth. Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Spatial-Frequency Mutual Learning for Face Super-ResolutionabstractFace super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from the low-resolution (LR) ones. With the advent of deep learning, the FSR technique has achieved significant breakthroughs. However, existing FSR methods either have a fixed receptive field or fail to maintain facial structure, limiting the FSRperformance. To circumvent this problem, Fourier transform is introduced, which can capture global facial structure information and achieve image-size receptive field. Relying on the Fourier transform, we devise a spatial-frequency mutual network (SFMNet) for FSR, which is the first FSR method to explore the correlations between spatial and frequency domains as far as we know. To be specific, our SFMNet is a two-branch network equipped with a spatial branch and a frequency branch. Benefiting from the property of Fourier transform, the frequency branch can achieve image-size receptive field and capture global dependency while the spatial branch can extract local dependency. Considering that these dependencies are complementary and both favorable for FSR, we further develop a frequency-spatial interaction block (FSIB) which mutually amalgamates the complementary spatial and frequency information to enhance the capability of the model. Quantitative and qualitative experimental results show that the proposed method out-performs state-of-the-art FSR methods in recovering face images. The implementation and model will be released at https://github.com/wcy-cs/SFMNet. Chenyang Wang 0002, Junjun Jiang, Zhiwei Zhong 0001, Xianming Liu 0005 |
CVPR | 1 |
| 2023 | Learning Enriched Hop-Aware Correlation for Robust 3D Human Pose Estimation
Shengping Zhang, Chenyang Wang 0002, Liqiang Nie, Hongxun Yao, Qingming Huang, Qi Tian 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Correction to: Learning Enriched Hop-Aware Correlation for Robust 3D Human Pose Estimation
Shengping Zhang, Chenyang Wang 0002, Liqiang Nie, Hongxun Yao, Qingming Huang, Qi Tian 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Unsupervised Low-Light Video Enhancement With Spatial-Temporal Co-Attention TransformerabstractExisting low-light video enhancement methods are dominated by Convolution Neural Networks (CNNs) that are trained in a supervised manner. Due to the difficulty of collecting paired dynamic low/normal-light videos in real-world scenes, they are usually trained on synthetic, static, and uniform motion videos, which undermines their generalization to real-world scenes. Additionally, these methods typically suffer from temporal inconsistency (e.g., flickering artifacts and motion blurs) when handling large-scale motions since the local perception property of CNNs limits them to model long-range dependencies in both spatial and temporal domains. To address these problems, we propose the first unsupervised method for low-light video enhancement to our best knowledge, named LightenFormer, which models long-range intra- and inter-frame dependencies with a spatial-temporal co-attention transformer to enhance brightness while maintaining temporal consistency. Specifically, an effective but lightweight S-curve Estimation Network (SCENet) is first proposed to estimate pixel-wise S-shaped non-linear curves (S-curves) to adaptively adjust the dynamic range of an input video. Next, to model the temporal consistency of the video, we present a Spatial-Temporal Refinement Network (STRNet) to refine the enhanced video. The core module of STRNet is a novel Spatial-Temporal Co-attention Transformer (STCAT), which exploits multi-scale self- and cross-attention interactions to capture long-range correlations in both spatial and temporal domains among frames for implicit motion estimation. To achieve unsupervised training, we further propose two non-reference loss functions based on the invertibility of the S-curve and the noise independence among frames. Extensive experiments on the SDSD and LLIV-Phone datasets demonstrate that our LightenFormer outperforms state-of-the-art methods. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Weigang Zhang, Hongxun Yao, Qingming Huang |
IEEE Trans. Image Process. | 3 |
| 2023 | Automatic Shadow Generation via Exposure FusionabstractShadow generation aims to generate a plausible shadow for the inserted foreground object in a composite image. Besides the composite image and the associated mask of the inserted foreground object, existing methods also require a mask of all background objects as well as their shadows as an auxiliary input, which is laborious in practical applications. Meanwhile, most existing methods use a linear illumination transformation to darken the shadow region, which is prone to produce unrealistic shadows especially when background illumination is complex. To address these problems, this paper proposes an automatic shadow generation method, which avoids the laborious acquisition of the background object masks while harmonizing the shadow region to achieve plausible shadow effects. Specifically, to implicitly exploit background illumination to infer the shadow shape of the inserted foreground object, we first propose a Hierarchy Attention U-Net (HAU-Net) to sequentially build global interactions between the foreground object and background across spatial and channel dimensions. Since the spatial-variant property of the shadow, we formulate shadow harmonization as an exposure fusion problem and propose an Illumination-Aware Fusion Network (IFNet), which uses an improved illumination model with a double linear transformation to produce multiple under-exposure images of the shadow region. IFNet then learns pixel-wise fusion kernels that consider the local smoothness of the shadow to fuse the composite image with these under-exposure images to generate the realistic shadow of the foreground object. Extensive experiments on the DESOBA and Shadow-AR datasets demonstrate that our method achieves state-of-the-art performance for shadow generation on both the BOS and BOS-free test images. Quanling Meng, Shengping Zhang, Zonglin Li 0004, Chenyang Wang 0002, Weigang Zhang, Qingming Huang |
IEEE Trans. Multim. | 4 |
| 2022 | Natural Image Matting with Shifted Window Self-AttentionabstractNatural image matting is a challenging and significant task in computer vision. Recently, image matting achieves fantastic development by introducing deep learning methods. To the best of our knowledge, there is no image matting method using the Transformer. Compared with CNNs, the Transformer pays more attention to the interest points and the relationships of content, which is beneficial to the image matting task. In this paper, we first present a novel Transformer-based image matting method with Shifted Window self-Attention. Specifically, our method contains two encoders, an alpha encoder and a context encoder. The former leverages the Transformer with Shifted Window self-Attention to extract features of details, such as hairs, feathers and porous parts of foreground objects. Shifted Window self-Attention focuses on patches with the size of the window and connections of adjacent patches. With this, the Transformer is capable of dealing with high-resolution images. The context encoder, which takes rescaled images as input, aims to extract the whole structure information of foreground objects. Then, we propose a novel Hierarchical Pyramid Pooling Module (HPPM) which enables the network to have the flexibility to extract features at various resolutions. Experiments show that our method achieves competitive performance on the Composition-1K dataset. Yang Liu 0119, Zonglin Li 0004, Chenyang Wang 0002, Shengping Zhang |
ICIP | 4 |
| 2022 | Progressive Limb-Aware Virtual Try-OnabstractExisting image-based virtual try-on methods directly transfer specific clothing to a human image without utilizing clothing attributes to refine the transferred clothing geometry and textures, which causes incomplete and blurred clothing appearances. In addition, these methods usually mask the limb textures of the input for the clothing-agnostic person representation, which results in inaccurate predictions for human limb regions (i.e., the exposed arm skin), especially when transforming between long-sleeved and short-sleeved garments. To address these problems, we present a progressive virtual try-on framework, named PL-VTON, which performs pixel-level clothing warping based on multiple attributes of clothing and embeds explicit limb-aware features to generate photo-realistic try-on results. Specifically, we design a Multi-attribute Clothing Warping (MCW) module that adopts a two-stage alignment strategy based on multiple attributes to progressively estimate pixel-level clothing displacements. A Human Parsing Estimator (HPE) is then introduced to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and limb regions. Finally, we propose a Limb-aware Texture Fusion (LTF) module to estimate high-quality details in limb regions by fusing textures of the clothing and the human body with the guidance of explicit limb-aware features. Extensive experiments demonstrate that our proposed method outperforms the state-of-the-art virtual try-on methods both qualitatively and quantitatively. Shengping Zhang, Qinglin Liu, Zonglin Li 0004, Chenyang Wang 0002 |
ACM Multimedia | 5 |
| 2022 | Propagating Facial Prior Knowledge for Multitask Learning in Face Super-ResolutionabstractExisting face hallucination methods always achieve improved performance through regularizing the model with facial prior. Most of them always estimate facial prior information first and then leverage it to help the prediction of the target high-resolution face image. However, the accuracy of prior estimation is difficult to guarantee, especially for the low-resolution face image. Once the estimated prior is inaccurate or wrong, the following face super-resolution performance is unavoidably influenced. A natural question that arises: how to incorporate facial prior effectively and efficiently without prior estimation? To achieve this goal, we propose to learn facial prior knowledge at training stage, but test only with low-resolution face image, which can overcome the difficulty of estimating accurate prior. In addition, instead of estimating facial prior, we directly explore the potential of high-quality facial prior in the training phase and progressively propagate the facial prior knowledge from the teacher network (trained with the low-resolution face/high-quality facial prior and high-resolution face image pairs) to the student network (trained with the low-resolution face and high-resolution face image pairs). Quantitative and qualitative comparisons on benchmark face datasets demonstrate that our method outperforms the state-of-the-art face super-resolution methods. The source codes of the proposed method will be available athttps://github.com/wcy-cs/KDFSRNet. Chenyang Wang 0002, Junjun Jiang, Zhiwei Zhong 0001, Xianming Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Learning with Noisy Labels via Sparse RegularizationabstractLearning with noisy labels is an important and challenging task for training accurate deep neural networks. Some commonly-used loss functions, such as Cross Entropy (CE), suffer from severe overfitting to noisy labels. Robust loss functions that satisfy the symmetric condition were tailored to remedy this problem, which however encounter the underfitting effect. In this paper, we theoretically prove that any loss can be made robust to noisy labels by restricting the network output to the set of permutations over a fixed vector. When the fixed vector is one-hot, we only need to constrain the output to be one-hot, which however produces zero gradients almost everywhere and thus makes gradient-based optimization difficult. In this work, we introduce the sparse regularization strategy to approximate the one-hot constraint, which is composed of network output sharpening operation that enforces the output distribution of a net-work to be sharp and the ℓp-norm (p ≤ 1) regularization that promotes the network output to be sparse. This simple approach guarantees the robustness of arbitrary loss functions while not hindering the fitting ability. Experimental results demonstrate that our method can significantly improve the performance of commonly-used loss functions in the presence of noisy labels and class imbalance, and out-perform the state-of-the-art methods. The code is available at https://github.com/hitcszx/lnl_sr. Xianming Liu 0005, Chenyang Wang 0002, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICCV | 3 |
| 2021 | Heatmap-Aware Pyramid Face HallucinationabstractRecent deep-learning-based face hallucination methods have achieved great success. Due to the parameter sharing characteristics of convolutional neural network, most existing deep-learning-based methods essentially use the same kernel for different regions of the entire face image in a convolution layer. This scheme of treating the face image as a whole will lead to the neglect of important facial details. To address this problem, we design a novel heatmap-aware convolution with spatially variant kernels rather than a spatially sharing kernel in the standard convolution to recover different regions. Based on this, we propose a heatmap-aware pyramid face super-resolution network (HaPSR) that embeds our heatmap-aware convolution into a two-branch network for both face super-resolution and facial heatmap estimation. The facial heatmap estimation branch can not only be used as an auxiliary to regularize face super-resolution reconstruction, but also provide an important basis for spatially variant kernels. Quantitative and qualitative experimental results demonstrate that our method outperforms state-of-the-arts. Chenyang Wang 0002, Junjun Jiang, Xianming Liu 0005 |
ICME | 1 |
| 2020 | Parsing Map Guided Multi-Scale Attention Network For Face HallucinationabstractFace hallucination that aims to transform a low-resolution (LR) face image to a high-resolution (HR) one is an active domain-specific image super-resolution problem. The performance of existing methods is usually not satisfactory, especially when the upscaling factor is large, such as 8×. In this paper, we propose an effective two- step face hallucination method based on a deep neural network with multi-scale channel and spatial attention mechanism. Specifically, we develop a ParsingNet to extract the prior knowledge of an input LR face, which is then fed into a carefully designed FishSRNet to recover the target HR face. Experimental results demonstrate that our method outperforms the state-of-the-arts in terms of quantitative metrics and visual quality. Chenyang Wang 0002, Zhiwei Zhong 0001, Junjun Jiang, Deming Zhai, Xianming Liu 0005 |
ICASSP | 1 |