Jie Cao 0002

dblp:39/6191-2 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0001-6368-4495ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 27 · 5 first-author · 16 since 2021Security and privacy · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoGrad3D: Spatially-Coupled Timestep Optimization with Orthogonal Gradient Fusion for 3D Generation
abstract
Score Distillation Sampling has driven recent advances in text-to-3D generation. However, current approaches often fail to produce 3D assets that are both rich in detail and consistent across viewpoints. These limitations primarily arise from imbalanced guidance on fine-grained details and an overdependence on single-view optimization—issues exacerbated by the excessive randomness in selecting diffusion timesteps and camera configurations. Such deficiencies commonly lead to blurry textures and inter-view inconsistencies, which degrade visual realism and hinder practical deployment. To tackle these challenges, we introduce CoGrad3D, a unified generative refinement framework that adopts a continuously adaptive optimization strategy. By dynamically modulating the optimization focus based on real-time convergence signals, CoGrad3D ensures balanced progress toward both geometric completeness and high-fidelity detail. Concretely, we propose an adaptive region sampling strategy that emphasizes under-converged viewing areas, promoting stable and uniform optimization. To facilitate the transition from coarse geometry to fine-grained reconstruction, we develop a region-aware temporal scheduling scheme that integrates global training dynamics with local convergence feedback. Furthermore, we introduce a gradient fusion mechanism that consolidates historical gradients from adjacent viewpoints, mitigating view-specific artifacts and promoting the emergence of coherent 3D structures. Extensive experiments demonstrate that CoGrad3D substantially surpasses existing methods in both geometric consistency and texture fidelity, enabling the generation of high-quality, view-consistent 3D models from textual descriptions.
Haoyang Tong, Jin Liu 0040, Jie Cao 0002, Ran He 0001
AAAI5
2026 Uncertainty-Aware Source-Free Adaptive Image Restoration with State Space Augmentation
Yuang Ai, Jie Cao 0002, Ran He 0001, Huaibo Huang
Int. J. Comput. Vis.2
2025 DM-DPR: Diffusion and Mamba-based Degradation Prediction for Blind Face Restoration
abstract
Blind Face Restoration (BFR), which involves converting low-quality facial images with unknown and varied degradation into high-quality counterparts, suffers from issues of sub-optimal restoration and over-correction due to inconsistent degradation levels. To rectify the above issue, inspired by the State Space model, especially the improved version Mamba’s enhanced long-range dependencies modeling ability and Stable Diffusion’s ability in integrating multi-modal prompts, we introduce a novel approach, Diff-Mamba Degradation Prediction Restoration (DM-DPR), to leverage the combination of a Mamba prompt generation framework with Stable Diffusion-based image restoration. Its core lies in two primary components: a Mamba-based multi-modal prompt generator that quantifies the degradation severity, generating corresponding textual and visual prompts, additionally with a multi-modal prompt driven Stable Diffusion process that adjusts restoration efforts based on the estimated degradation level. Derived from the CelebA-Test, we create degraded datasets exhibiting a wide range of degradation severity. Extensive experimental evaluations demonstrate that DM-DPR substantially surpasses existing state-of-the-art methods, thereby robustly establishing its enhanced capability to manage varying degrees of image degradation.
Guorong Yuan, Huaibo Huang, Jie Cao 0002, Yuang Ai, Ran He 0001
FG4
2025 Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection
abstract
With the advancement of generative AI, distinguishing real and AI-generated faces in videos has become increasingly challenging. However, traditional methods struggle to capture local details and temporal dynamics simultaneously, making it difficult to achieve high detection accuracy while maintaining low computational overhead. To address this problem, we propose a Dual-branch SpatioTemporal-Planar Network (Dual-PST) based on the selective state-space model. It is capable of extracting image features and temporal relations simultaneously, while maintaining linear computational consumption. Specifically, we design a Multi-Selective State-Space module (MS3) that can extract global features from image typography consisting of consecutive video frames by scanning them in multiple sequences. To further enhance temporal modeling capabilities, we propose a Sequential Tri-frame Local module, which captures inter-frame temporal relationships and local features by temporally splicing single-frame features. These features are first extracted using MS3 and then further enhanced through inter-frame masking operations. Experimental results show that Dual-PST significantly improves detection accuracy while maintaining low computational complexity and strong model robustness.
Junxian Duan, Jie Cao 0002, Aihua Zheng
ICASSP4
2025 Growing to Detect: A Dynamic Prototype Tree with Structured Replay for Incremental Deepfake Detection
abstract
The rapid advancement of deepfake technology poses significant threats to social trust. Recent research has improved detectors by adapting to emerging deepfakes using a limited number of samples through incremental learning. However, these approaches often overlook the scarcity of novel samples, resulting in insufficient learning of forgery patterns. To overcome this challenge, we propose a Replay-based Dynamic Prototype Network that integrates two key modules: the Dynamic Prototype Tree (DPT) module and the Similarity Subtree Replay (SSR) strategy. The DPT module dynamically introduces prototypes through a hierarchical tree structure to effectively adapt to new deepfakes. It expands prototypes based on similarity, thereby retaining the knowledge learned from previous prototypes while learning new forgery patterns. The SSR strategy mitigates catastrophic forgetting by stabilizing learned features through the replay of relevant subtrees. Experimental results demonstrate that our approach outperforms existing methods across five datasets, particularly on high-quality face swap samples generated by diffusion-based methods, achieving an AUC of 85.86% on the cross-dataset task from FaceForensics++ to DiffSwap.
Junxian Duan, Jie Cao 0002, Aihua Zheng, Ran He 0001
IJCB3
2025 Towards Robust Defense Against Customization via Protective Perturbation Resistant to Diffusion-based Purification
abstract
Diffusion models like Stable Diffusion have become prominent in visual synthesis tasks due to their powerful customization capabilities, which also introduce significant security risks, including deepfakes and copyright infringement. In response, a class of methods known as protective perturbation emerged, which mitigates image misuse by injecting imperceptible adversarial noise. However, purification can remove protective perturbations, thereby exposing images again to the risk of malicious forgery. In this work, we formalize the anti-purification task, highlighting challenges that hinder existing approaches, and propose a simple diagnostic protective perturbation named AntiPure. AntiPure exposes vulnerabilities of purification within the "purification-customization" workflow, owing to two guidance mechanisms: 1) Patch-wise Frequency Guidance, which reduces the model's influence over high-frequency components in the purified image, and 2) Erroneous Timestep Guidance, which disrupts the model's denoising strategy across different timesteps. With additional guidance, AntiPure embeds imperceptible perturbations that persist under representative purification settings, achieving effective post-customization distortion. Experiments show that, as a stress test for purification, AntiPure achieves minimal perceptual discrepancy and maximal distortion, outperforming other protective perturbation methods within the purification-customization workflow.
Wenkui Yang, Jie Cao 0002, Junxian Duan, Ran He 0001
ICCV2
2025 Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent Debate
abstract
Large Language Models (LLMs) have seen significant progress but continue to struggle with persistent reasoning mistakes. Previous methods of *self-reflection* have been proven limited due to the models’ inherent fixed thinking patterns. While Multi-Agent Debate (MAD) attempts to mitigate this by incorporating multiple agents, it often employs the same reasoning methods, even though assigning different personas to models. This leads to a "fixed mental set", where models rely on homogeneous thought processes without exploring alternative perspectives. In this paper, we introduce Diverse Multi-Agent Debate (DMAD), a method that encourages agents to think with distinct reasoning approaches. By leveraging diverse problem-solving strategies, each agent can gain insights from different perspectives, refining its responses through discussion and collectively arriving at the optimal solution. DMAD effectively breaks the limitations of fixed mental sets. We evaluate DMAD against various prompting techniques, including *self-reflection* and traditional MAD, across multiple benchmarks using both LLMs and Multimodal LLMs. Our experiments show that DMAD consistently outperforms other methods, delivering better results than MAD in fewer rounds. Code is available at https://github.com/MraDonkey/DMAD.
Yexiang Liu, Jie Cao 0002, Zekun Li 0001, Ran He 0001, Tieniu Tan
ICLR2
2025 Degradation-Aware Multi-Task Image Restoration with State Space Models
abstract
Image restoration (IR) has made significant strides, evolving from basic pixel-wise restoration to more advanced techniques capable of handling diverse degradations. However, current all-in-one IR methods face challenges in effectively generalizing across complex real-world degradation scenarios. To address this challenge, we propose multi-task Restoration Mamba (ReMamba), a novel framework that leverages the power of state space models for long-sequence modeling to extract fine-grained degradation features from low-quality data. Guided by degradation-type prediction, which helps the model accurately identify and differentiate between various degradation types in the input, ReMamba employs prompt learning across both textual and visual modalities and enhances the restoration capabilities of downstream stable diffusion networks by injecting prompts as prior knowledge. Extensive experiments on both synthetic and real-world datasets, covering six distinct IR tasks, demonstrate the adaptability, generalizability, and robustness of ReMamba, highlighting its effectiveness in addressing a wide range of degradation types under challenging conditions.
Purui Bai, Huaibo Huang, Jie Cao 0002, Yuang Ai, Ran He 0001
ICME4
2025 MTSD: Simple Yet Effective Self-Distillation for Generalizable Deepfake Detection
abstract
The rapid advancement of Deepfake technology necessitates detection systems with strong generalization capabilities. Existing methods often depend on architectural modifications or dataset-specific prior knowledge, which limits their scalability and practical deployment in real-world scenarios. We propose Multi-Teacher Self-Distillation (MTSD), a simple yet effective and generalizable strategy to enhance model generalization. MTSD comprises two key steps. First, diverse teacher generation leverages independently trained teacher models with varying dataset sampling sequences to capture complementary decision boundaries. Second, self-distillation feature fusion integrates these diverse features using a cross-attention mechanism, allowing the student model to approximate an ideal feature distribution for improved generalization. This strategy avoids architectural changes and dataset-specific adjustments, ensuring simplicity in implementation and deployment. Moreover, the multi-teacher generation and feature fusion steps are discarded after training, preserving computational efficiency during inference. Experimental results demonstrate that MTSD significantly improves model generalization, offering a practical and scalable solution for Deepfake detection.
Dexu Zhu, Jie Cao 0002, Jiangnan Shao, Junxian Duan, Ran He 0001
ICME2
2025 Straighter Flow Matching via a Diffusion-Based Coupling Prior
Siyu Xing, Jie Cao 0002, Huaibo Huang, Haichao Shi, Xiaoyu Zhang 0002
PRCV (8)2
2025 Test-time Forgery Detection with Spatial-Frequency Prompt Learning
Junxian Duan, Yuang Ai, Shenyuan Huang, Huaibo Huang, Jie Cao 0002, Ran He 0001
Int. J. Comput. Vis.6
2024 PortraitDAE: Line-Drawing Portraits Style Transfer from Photos via Diffusion Autoencoder with Meaningful Encoded Noise
abstract
The line-drawing portrait is a kind of highly abstract art that contains a sparse set of continuous graphical elements such as lines to capture a person's facial features. Due to their abstract artistic form, common style transfer methods fail to synthesize high-quality line-drawing portraits from photos. Previous works mostly concentrate on GANs, often requiring pre-calculated landmarks acquired by other models and using extra classifiers with complicated structures to capture local facial features. We propose a novel idea without these extra operations based on diffusion models, which is more flexible and stable than GAN-based methods. We utilize the diffusion-based decoder in the Diffusion Autoencoder to encode the input image to an encoded noise that contains much meaningful stochastic information by running the deterministic generative process backward. By fully utilizing the encoded noise, our method can effectively preserve the identity information and better capture facial details. We also improve the loss function to alleviate the interference of the background color. Several experiments show that our method can produce better samples with smoother lines that look more like the corresponding person, outperforming state-of-the-art methods both qualitatively and quantitatively. Our method can also be generalized to other styles such as sketch.
Yexiang Liu, Jin Liu 0040, Jie Cao 0002, Junxian Duan, Ran He 0001
FG3
2024 Semantic-Aware Detail Enhancement for Blind Face Restoration
abstract
The goal of Blind Face Restoration is to recover high-quality images from low-quality images suffering from unknown degradations, posing a significantly challenging problem. In recent years, numerous BFR methods have been proposed, achieving significant success. However, faces possess a unique facial topology, and subtle differences in texture, slight structural imbalances, and minimal asymmetry are easily perceptible in the restored face images. Previous methods often struggle to generate realistically high-quality images from real-world low-quality images and fail to preserve fine features. To more effectively restore image details and textures, providing a more natural and realistic restoration effect, we integrate facial semantic information as prior knowledge into the blind face restoration task. We employ a multi-head cross-attention mechanism to simultaneously consider facial semantic information and context information for modeling. Additionally, we introduce a local detail enhancement module specifically designed to enhance the processing capability of details around the eyes and mouth. Experimental results indicate that our proposed method recovers facial images on synthetic and real datasets more realistically and with higher fidelity.
Xiaoqiang Zhou, Jie Cao 0002, Huaibo Huang, Aihua Zheng, Ran He 0001
FG3
2024 ZePo: Zero-Shot Portrait Stylization with Faster Sampling
abstract
Diffusion-based text-to-image generation models have significantly advanced the field of art content synthesis. However, current portrait stylization methods generally require either model fine-tuning based on examples or the employment of DDIM Inversion to revert images to noise space, both of which substantially decelerate the image generation process. To overcome these limitations, this paper presents an inversion-free portrait stylization framework based on diffusion models that accomplishes content and style feature fusion in merely four sampling steps. We observed that Latent Consistency Models employing consistency distillation can effectively extract representative Consistency Features from noisy images. To blend the Consistency Features extracted from both content and style images, we introduce a Style Enhancement Attention Control technique that meticulously merges content and style features within the attention space of the target image. Moreover, we propose a feature merging strategy to amalgamate redundant features in Consistency Features, thereby reducing the computational load of attention control. Extensive experiments have validated the effectiveness of our proposed framework in enhancing stylization efficiency and fidelity. The code is available at \url{https://github.com/liujin112/ZePo}.
Jin Liu 0040, Huaibo Huang, Jie Cao 0002, Ran He 0001
ACM Multimedia3
2024 Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
abstract
Recent advancements in 3D content generation have been significant, primarily due to the visual priors provided by pretrained diffusion models. However, large 2D visual models exhibit spatial perception hallucinations, leading to multi-view inconsistency in 3D content generated through Score Distillation Sampling (SDS). This phenomenon, characterized by overfitting to specific views, is referred to as the "Janus Problem". In this work, we investigate the hallucination issues of pretrained models and find that large multimodal models without geometric constraints possess the capability to infer geometric structures, which can be utilized to mitigate multi-view inconsistency. Building on this, we propose a novel tuning-free method. We represent the multimodal inconsistency query information to detect specific hallucinations in 3D content, using this as an enhanced prompt to re-consist the 2D renderings of 3D and jointly optimize the structure and appearance across different views. Our approach does not require 3D training data and can be implemented plug-and-play within existing frameworks. Extensive experiments demonstrate that our method significantly improves the consistency of 3D content generation and specifically mitigates hallucinations caused by pretrained large models, achieving state-of-the-art performance compared to other optimization methods.
Jie Cao 0002, Jin Liu 0040, Xiaoqiang Zhou, Huaibo Huang, Ran He 0001
NeurIPS2
2024 Learning Fine-Grained and Semantically Aware Mamba Representations for Tampered Text Detection in Images
Jie Cao 0002, Huaibo Huang
PRCV (7)2
2024 TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
Wenkui Yang, Xiaoqiang Zhou, Junxian Duan, Jie Cao 0002
PRCV (11)5
2024 r-FACE: Reference guided face component editing
Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Qi Li 0005, Zhenan Sun
Pattern Recognit.2
2023 Modify: Model-Driven Face Stylization Without Style Images
abstract
Existing face stylization methods always acquire the presence of the target (style) domain during the translation process, which violates privacy regulations and limits their applicability in real-world systems. To address this issue, we propose a new method called MODel-drIven Face stYlization (MODIFY), which relies on the generative model to bypass the dependence of the target images. Briefly, MODIFY first trains a generative model in the target domain and then translates a source input to the target domain via the provided style model. To preserve the multimodal style information, MODIFY further introduces an additional remapping network, mapping a known continuous distribution into the encoder’s embedding space. During translation in the source domain, MODIFY fine-tunes the encoder module within the target style-persevering model to capture the content of the source input as precisely as possible. Our method is extremely simple and satisfies versatile training modes for face stylization. Experimental results on several different datasets validate the effectiveness of MODIFY for unsupervised face stylization. Code will be released at https://github.com/YuheD/MODIFY.
Yuhe Ding, Jian Liang 0001, Jie Cao 0002, Aihua Zheng, Ran He 0001
ICASSP3
2023 Where to Focus: Central Attention-Based Face Forgery Detection
Jinghui Sun, Yuhe Ding, Jie Cao 0002, Junxian Duan, Aihua Zheng
PRCV (5)3
2023 ScoreMix: A Scalable Augmentation Strategy for Training GANs With Limited Data
abstract
Generative Adversarial Networks (GANs) typically suffer from overfitting when limited training data is available. To facilitate GAN training, current methods propose to use data-specific augmentation techniques. Despite the effectiveness, it is difficult for these methods to scale to practical applications. In this article, we present ScoreMix, a novel and scalable data augmentation approach for various image synthesis tasks. We first produce augmented samples using the convex combinations of the real samples. Then, we optimize the augmented samples by minimizing the norms of the data scores, i.e., the gradients of the log-density functions. This procedure enforces the augmented samples close to the data manifold. To estimate the scores, we train a deep estimation network with multi-scale score matching. For different image synthesis tasks, we train the score estimation network using different data. We do not require the tuning of the hyperparameters or modifications to the network architecture. The ScoreMix method effectively increases the diversity of data and reduces the overfitting problem. Moreover, it can be easily incorporated into existing GAN models with minor modifications. Experimental results on numerous tasks demonstrate that GAN models equipped with the ScoreMix method achieve significant improvements.
Jie Cao 0002, Mandi Luo, Junchi Yu, Ming-Hsuan Yang 0001, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Semantic-Aware Noise Driven Portrait Synthesis and Manipulation
abstract
Semantic portrait synthesis has drawn consistent attention and has made significant progress, yet achieving style diversity and semantic controllability simultaneously is still a challenge. Existing methods either 1) directly take a semantic label map as input, ignoring various possibilities of semantic styles, or 2) sample global noise as input, ignoring controllability of local semantics. To fill this gap, we propose semantic-aware noise, a simple but effective input that tackles both issues and shows improved results over baselines. Semantic-aware noise introduces semantic information into noise, and each semantic is sampled from the noise separately, combining the semantic controllability and the noise sampling diversity. To further expand and manipulate real images, we propose a novel ternary network structure, allowing simultaneous diverse semantic image synthesis and real image manipulation in a unified framework. Extensive experiments demonstrate that the proposed method achieves quantitatively superior and perceptually pleasing results compared to state-of-the-art methods. We also analyze the performance of our method with respect to different noise structures and real-life applications in diverse synthesis, interactive manipulation, and extreme pose scenarios.
Qiyao Deng, Qi Li 0005, Jie Cao 0002, Yunfan Liu 0001, Zhenan Sun
IEEE Trans. Multim.3
2022 Improving Subgraph Recognition with Variational Graph Information Bottleneck
abstract
Subgraph recognition aims at discovering a compressed substructure of a graph that is most informative to the graph property. It can be formulated by optimizing Graph Information Bottleneck (GIB) with a mutual information estimator. However, GIB suffers from training instability and degenerated results due to its intrinsic optimization process. To tackle these issues, we reformulate the subgraph recognition problem into two steps: graph perturbation and subgraph selection, leading to a novel Variational Graph Information Bottleneck (VGIB) framework. VGIB first employs the noise injection to modulate the information flow from the input graph to the perturbed graph. Then, the perturbed graph is encouraged to be informative to the graph property. VGIB further obtains the desired subgraph by filtering out the noise in the perturbed graph. With the customized noise prior for each input, the VGIB objective is endowed with a tractable variational upper bound, leading to a superior empirical performance as well as theoretical properties. Extensive experiments on graph interpretation, explainability of Graph Neural Networks, and graph classification show that VGIB finds better subgraphs than existing methods11Code is avaliable on https://github.com/Samyu0304/VGIB.
Junchi Yu, Jie Cao 0002, Ran He 0001
CVPR2
2022 Adaptive Transformer-Based Conditioned Variational Autoencoder for Incomplete Social Event Classification
abstract
With the rapid development of the Internet and the expanding scale of social media, incomplete social event classification has increasingly become a challenging task. The key for incomplete social event classification is to accurately leverage the image-level and text-level information. However, most of the existing approaches may suffer from the following limitations: (1) Most Generative Models use the available features to generate the incomplete modality features for social events classification while ignoring the rich semantic label information. (2) The majority of existing multi-modal methods just simply concatenate the coarse-grained image features and text features of the event to get the multi-modal features to classify social events, which ignores the irrelevant multi-modal features and limits their modeling capabilities. To tackle these challenges, in this paper, we propose an Adaptive Transformer-Based Conditioned Variational Autoencoder Network (AT-CVAE) for incomplete social event classification. In the AT-CVAE, we propose a novel Transformer-based Conditioned Variational Autoencoder to jointly model the textual information, visual information and label information into a unified deep model, which can generate more discriminative latent features and enhance the performance of incomplete social event classification. Furthermore, the Mixture-of-Experts Mechanism is utilized to dynamically acquire the weights of each multi-modal information, which can better filter out the irrelevant multi-modal information and capture the vitally important information. Extensive experiments are conducted on two public event datasets, demonstrating the superior performance of our AT-CVAE method.
Zhangming Li 0002, Shengsheng Qian, Jie Cao 0002, Quan Fang, Changsheng Xu
ACM Multimedia3
2022 Order-aware Human Interaction Manipulation
abstract
The majority of current techniques for pose transfer disregard the interactions between the transferred person and the surrounding instances, resulting in context inconsistency when applied to complicated situations. To tackle this issue, we propose InterOrderNet, a novel framework to perform order-aware interaction learning. The proposed InterOrderNet learns the relative order on the direction of the z-axis among instances to describe instance-level occlusions. Not only does learning this order guarantee the context consistency of human pose transfer, but it also enhances its generalization to natural scenes. Additionally, we present a novel unsupervised method, named Imitative Contrastive Learning, which sidesteps the requirements of order annotations. Existing pose transfer methods are easy to be integrated into the proposed InterOrderNet. Extensive experiments demonstrate that InterOrderNet enables these methods to perform interaction manipulation.
Mandi Luo, Jie Cao 0002, Ran He 0001
ACM Multimedia2
2022 Learning 3D Human Shape and Pose From Dense Body Parts
abstract
Reconstructing 3D human shape and pose from monocular images is challenging despite the promising results achieved by the most recent learning-based methods. The commonly occurred misalignment comes from the facts that the mapping from images to the model space is highly non-linear and the rotation-based pose representation of the body model is prone to result in the drift of joint positions. In this work, we investigate learning 3D human shape and pose from dense correspondences of body parts and propose a Decompose-and-aggregate Network (DaNet) to address these issues. DaNet adopts the dense correspondence maps, which densely build a bridge between 2D pixels and 3D vertexes, as intermediate representations to facilitate the learning of 2D-to-3D mapping. The prediction modules of DaNet are decomposed into one global stream and multiple local streams to enable global and fine-grained perceptions for the shape and pose predictions, respectively. Messages from local streams are further aggregated to enhance the robust prediction of the rotation-based poses, where a position-aided rotation feature refinement strategy is proposed to exploit spatial relationships between body joints. Moreover, a Part-based Dropout (PartDrop) strategy is introduced to drop out dense information from intermediate representations during training, encouraging the network to focus on more complementary body parts as well as neighboring position features. The efficacy of the proposed method is validated on both indoor and real-world datasets including Human3.6M, UP3D, COCO, and 3DPW, showing that our method could significantly improve the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://hongwenzhang.github.io/dense2mesh.
Hongwen Zhang 0001, Jie Cao 0002, Guo Lu, Wanli Ouyang, Zhenan Sun
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 ReMix: Towards Image-to-Image Translation With Limited Data
abstract
Image-to-image (I2I) translation methods based on generative adversarial networks (GANs) typically suffer from overfitting when limited training data is available. In this work, we propose a data augmentation method (ReMix) to tackle this issue. We interpolate training samples at the feature level and propose a novel content loss based on the perceptual relations among samples. The generator learns to translate the in-between samples rather than memorizing the training set, and thereby forces the discriminator to generalize. The proposed approach effectively reduces the ambiguity of generation and renders content-preserving results. The ReMix method can be easily incorporated into existing GAN models with minor modifications. Experimental results on numerous tasks demonstrate that GAN models equipped with the ReMix method achieve significant improvements.
Jie Cao 0002, Luanxuan Hou, Ming-Hsuan Yang 0001, Ran He 0001, Zhenan Sun
CVPR1
2021 FaceInpainter: High Fidelity Face Adaptation to Heterogeneous Domains
abstract
In this work, we propose a novel two-stage framework named FaceInpainter to implement controllable Identity-Guided Face Inpainting (IGFI) under heterogeneous domains. Concretely, by explicitly disentangling foreground and background of the target face, the first stage focuses on adaptive face fitting to the fixed background via a Styled Face Inpainting Network (SFI-Net), with 3D priors and texture code of the target, as well as identity factor of the source face. It is challenging to deal with the inconsistency between the new identity of the source and the original background of the target, concerning the face shape and appearance on the fused boundary. The second stage consists of a Joint Refinement Network (JR-Net) to refine the swapped face. It leverages AdaIN considering identity and multi-scale texture codes, for feature transformation of the decoded face from SFI-Net with facial occlusions. We adopt the contextual loss to implicitly preserve the attributes, encouraging face deformation and fewer texture distortions. Experimental results demonstrate that our approach handles high-quality identity adaptation to heterogeneous domains, exhibiting the competitive performance compared with state-of-the-art methods concerning both attribute and identity fidelity.
Jia Li 0044, Jie Cao 0002, Xingguang Song, Ran He 0001
CVPR3
2021 Global Relation-Aware Attention Network for Image-Text Retrieval
abstract
The cross-modal image-text retrieval has attracted extensive attention in recent years, which contributes to the development of search engine. Fine-grained features and cross-attention have been widely used in past researches to reach the goal of cross-modal image-text matching. Although cross-related methods have achieved remarkable results, the features must be encoded again in evaluation phase due to the interaction of the two modalities, which is unsuitable for actual scenarios of search engine development. In addition, the aggregated feature does not contain sufficient semantics since it is merely obtained by simple mean pooling. Furthermore, connecting weights of self-attention blocks are target position invariant, which lacks the expected adaptability. To tackle these limitations, in this paper, we propose a novel Global Relation-aware Attention Network (GRAN) for image-text retrieval by designing Global Attention Module (GAM) and Relation-aware Attention Module (RAM) which play an important role in modeling the global feature and the relationships of local fragments. Firstly, we propose Global Attention Module (GAM) followed the fine-grained features to obtain meaningful global feature. Secondly, we use several stacked transformer encoders to further encode features separately. Finally, we propose Relation-aware Attention Module (RAM) to generate a vector which represents the relation information to infer the attention intensity of pairwise fragments. The local features, the global feature, and their relations are considered jointly to conduct an efficient image-text retrieval. Extensive experiments are conducted on the benchmark datasets of Flickr30K and MSCOCO, demonstrating the superiority of our method. On the Flickr30K, compared to the state-of-the-art method TERAN, we improve [email protected](K=1) metric by 5.8% and 4.0 on the image and text retrieval tasks, respectively.
Jie Cao 0002, Shengsheng Qian, Huaiwen Zhang, Quan Fang, Changsheng Xu
ICMR1
2021 Controllable Multi-Attribute Editing of High-Resolution Face Images
abstract
In recent years, significant progress has been achieved in face image editing due to the success of Generative Adversarial Network (GAN). However, state-of-the-art face editing methods mainly suffer from the following two limitations: 1) they are only applicable to face images with relative low-resolutions and 2) multi-attribute face editing may generate uncontrollable changes in non-target face attribute categories. To solve these problems, we propose a novel High-Quality Generative Adversarial Network (HQ-GAN) for controllable editing of multiple face attributes in high-resolution images. HQ-GAN has two novel ideas to break the limitations of resolution and controllability correspondingly: 1) fine-grained textures and realistic details of high-resolution face images are better preserved with the aid of textural features extracted by the wavelet transform module and 2) desired multi-attribute targets of face editing are emphasized using a weighted binary cross-entropy (BCE) loss so that the influence on non-target attributes is greatly reduced. To the best of our knowledge, HQ-GAN is the first attempt to achieve continuous editing of multiple face attributes on high-resolution images of the CelebA-HQ using only 28 000 training samples. Extensive qualitative results demonstrate the superiority of the proposed method in rendering realistic high-resolution face images with accurate attribute modification, and comprehensive quantitative results show that the proposed method significantly outperforms state-of-the-art face editing methods.
Qiyao Deng, Qi Li 0005, Jie Cao 0002, Yunfan Liu 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2021 FA-GAN: Face Augmentation GAN for Deformation-Invariant Face Recognition
abstract
Substantial improvements have been achieved in the field of face recognition due to the successful application of deep neural networks. However, existing methods are sensitive to both the quality and quantity of the training data. Despite the availability of large-scale datasets, the long tail data distribution induces strong biases in model learning. In this paper, we present a Face Augmentation Generative Adversarial Network (FA-GAN) to reduce the influence of imbalanced deformation attribute distributions. We propose to decouple these attributes from the identity representation with a novel hierarchical disentanglement module. Moreover, Graph Convolutional Networks (GCNs) are applied to recover geometric information by exploring the interrelations among local regions to guarantee the preservation of identities in face data augmentation. Extensive experiments on face reconstruction, face manipulation, and face recognition demonstrate the effectiveness and generalization ability of the proposed method.
Mandi Luo, Jie Cao 0002, Xin Ma 0031, Xiaoyu Zhang 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.2
2021 Partial NIR-VIS Heterogeneous Face Recognition With Automatic Saliency Search
abstract
Near-infrared-visual (NIR-VIS) heterogeneous face recognition (HFR) aims to match NIR face images with the corresponding VIS ones. It is a challenging task due to the sensing gaps among different modalities. Occlusions in the input face images make the task extremely complex. To tackle these problems, we present a Saliency Search Network (SSN) to extract domain-invariant identity features. We propose to automatically search the efficient parts of face images in a modality-aware manner, and remove redundant information. Moreover, the searching process is guided by an information bottleneck network, which mitigates the overfitting problems caused by small datasets. Extensive experiments on both complete and partial NIR-VIS HFR on multiple datasets demonstrate the effectiveness and robustness of the proposed method to modality discrepancy and occlusions.
Mandi Luo, Xin Ma 0031, Zhihang Li, Jie Cao 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.4
2020 PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer
abstract
In this paper, we address the makeup transfer task, which aims to transfer the makeup from a reference image to a source image. Existing methods have achieved promising progress in constrained scenarios, but transferring between images with large pose and expression differences is still challenging. Besides, they cannot realize customizable transfer that allows a controllable shade of makeup or specifies the part to transfer, which limits their applications. To address these issues, we propose Pose and expression robust Spatial-aware GAN (PSGAN). It first utilizes Makeup Distill Network to disentangle the makeup of the reference image as two spatial-aware makeup matrices. Then, Attentive Makeup Morphing module is introduced to specify how the makeup of a pixel in the source image is morphed from the reference image. With the makeup matrices and the source image, Makeup Apply Network is used to perform makeup transfer. Our PSGAN not only achieves state-of-the-art results even when large pose and expression differences exist but also is able to perform partial and shade-controllable makeup transfer. Both the code and a newly collected dataset containing facial images with various poses and expressions will be available at https://github.com/wtjiang98/PSGAN.
Si Liu 0001, Chen Gao 0005, Jie Cao 0002, Ran He 0001, Jiashi Feng, Shuicheng Yan
CVPR4
2020 Informative Sample Mining Network for Multi-domain Image-to-Image Translation
Jie Cao 0002, Huaibo Huang, Yi Li 0018, Ran He 0001, Zhenan Sun
ECCV (19)1
2020 $P^{2}$ Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
abstract
The target of human pose estimation is to determine the body parts and joint locations of persons in the image. Angular changes, motion blur and occlusion in the natural scenes make this task challenging, while some joints are more difficult to be detected than others. In this paper, we propose an augmented Parallel-Pyramid Net ( P2Net) with feature refinement by dilated bottleneck and attention module. During data preprocessing, we proposed a differentiable auto data augmentation ( DA2) method. We formulate the problem of searching data augmentaion policy in a differentiable form, so that the optimal policy setting can be easily updated by back propagation during training. DA2improves the training efficiency. A parallel-pyramid structure is followed to compensate the information loss introduced by the network. We innovate two fusion structures, i.e. Parallel Fusion and Progressive Fusion, to process pyramid features from backbone network. Both fusion structures leverage the advantages of spatial information affluence at high resolution and semantic comprehension at low resolution effectively. We propose a refinement stage for the pyramid features to further boost the accuracy of our network. By introducing dilated bottleneck and attention module, we increase the receptive field for the features with limited complexity and tune the importance to different feature channels. To further refine the feature maps after completion of feature extraction stage, an Attention Module ( AM) is defined to extract weighted features from different scale feature maps generated by the parallel-pyramid structure. Compared with the traditional up-sampling refining, AM can better capture the relationship between channels. Experiments corroborate the effectiveness of our proposed method. Notably, our method achieves the best performance on the challenging MSCOCO and MPII datasets.
Luanxuan Hou, Jie Cao 0002, Haifeng Shen, Jian Tang 0008, Ran He 0001
ICPR2
2020 Exemplar Guided Cross-Spectral Face Hallucination via Mutual Information Disentanglement
abstract
Recently, many Near infrared-visible (NIR-VIS) heterogeneous face recognition (HFR) methods have been proposed in the community. But it remains a challenging problem because of the sensing gap along with large pose variations. In this paper, we propose an Exemplar Guided Cross-Spectral Face Hallucination (EGCH) to reduce the domain discrepancy through disentangled representation learning. For each modality, EGCH contains a spectral encoder as well as a structure encoder to disentangle spectral and structure representation, respectively. It also contains a traditional generator that reconstructs the input from the above two representations, and a structure generator that predicts the facial parsing map from the structure representation. Besides, mutual information minimization and maximization are conducted to boost disentanglement and make representations adequately expressed. Then the translation is built on structure representations between two modalities. Provided with the transformed NIR structure representation and original VIS spectral representation, EGCH is capable to produce high-fidelity VIS images that preserve the topology structure of the input NIR while transfer the spectral information of an arbitrary VIS exemplar. Extensive experiments demonstrate that the proposed method achieves more promising results both qualitatively and quantitatively than the state-of-the-art NIR-VIS methods.
Haoxue Wu, Huaibo Huang, Aijing Yu, Jie Cao 0002, Zhen Lei 0001, Ran He 0001
ICPR4
2020 Reference Guided Face Component Editing
abstract
Face portrait editing has achieved great progress in recent years. However, previous methods either 1) operate on pre-defined face attributes, lacking the flexibility of controlling shapes of high-level semantic facial components (e.g., eyes, nose, mouth), or 2) take manually edited mask or sketch as an intermediate representation for observable changes, but such additional input usually requires extra efforts to obtain. To break the limitations (e.g. shape, mask or sketch) of the existing methods, we propose a novel framework termed r FACE (Reference Guided FAce Component Editing) for diverse and controllable face component editing with geometric changes. Specifically, r-FACE takes an image inpainting model as the backbone, utilizing reference images as conditions for controlling the shape of face components. In order to encourage the framework to concentrate on the target face components, an example-guided attention module is designed to fuse attention features and the target face component features extracted from the reference image. Through extensive experimental validation and comparisons, we verify the effectiveness of the proposed framework.
Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Zhenhua Chai, Qi Li 0005, Zhenan Sun
IJCAI2
2020 InteractGAN: Learning to Generate Human-Object Interaction
abstract
Compared with the widely studied Human-Object Interaction DE-Tection (HOI-DET), no effort has been devoted to its inverse problem, i.e. to generate an HOI scene image according to the given relationship triplet , to our best knowledge. We term this new task "Human-Object Interaction Image Generation" (HOI-IG). HOI-IG is a research-worthy task with great application prospects, such as online shopping, film production and interactive entertainment. In this work, we introduce an Interact-GAN to solve this challenging task. Our method is composed of two stages: (1) manipulating the posture of a given human image conditioned on a predicate. (2) merging the transformed human image and object image to one realistic scene image while satisfying the ir expected relative position and ratio. Besides, to address the large spatial misalignment issue caused by fusing two images content with reasonable spatial layout, we propose a Relation-based Spatial Transformer Network (RSTN) to adaptively process the images conditioned on their interaction. Extensive experiments on two challenging datasets demonstrate the effectiveness and superiority of our approach. We advocate for the image generation community to draw more attention to the new Human-Object Interaction Image Generation problem. To facilitate future research, our project will be released at: http://colalab.org/projects/InteractGAN.
Chen Gao 0005, Si Liu 0001, Defa Zhu, Jie Cao 0002, Haoqian He, Ran He 0001, Shuicheng Yan
ACM Multimedia5
2020 Towards High Fidelity Face Frontalization in the Wild
Jie Cao 0002, Yibo Hu 0001, Hongwen Zhang 0001, Ran He 0001, Zhenan Sun
Int. J. Comput. Vis.1
2020 Disentangled Representation Learning of Makeup Portraits in the Wild
Yi Li 0018, Huaibo Huang, Jie Cao 0002, Ran He 0001, Tieniu Tan
Int. J. Comput. Vis.3
2020 Adversarial Cross-Spectral Face Completion for NIR-VIS Face Recognition
abstract
Near infrared-visible (NIR-VIS) heterogeneous face recognition refers to the process of matching NIR to VIS face images. Current heterogeneous methods try to extend VIS face recognition methods to the NIR spectrum by synthesizing VIS images from NIR images. However, due to the self-occlusion and sensing gap, NIR face images lose some visible lighting contents so that they are always incomplete compared to VIS face images. This paper models high-resolution heterogeneous face synthesis as a complementary combination of two components: a texture inpainting component and a pose correction component. The inpainting component synthesizes and inpaints VIS image textures from NIR image textures. The correction component maps any pose in NIR images to a frontal pose in VIS images, resulting in paired NIR and VIS textures. A warping procedure is developed to integrate the two components into an end-to-end deep network. A fine-grained discriminator and a wavelet-based discriminator are designed to improve visual quality. A novel 3D-based pose correction loss, two adversarial losses, and a pixel loss are imposed to ensure synthesis results. We demonstrate that by attaching the correction component, we can simplify heterogeneous face synthesis from one-to-many unpaired image translation to one-to-one paired image translation, and minimize the spectral and pose discrepancy during heterogeneous recognition. Extensive experimental results show that our network not only generates high-resolution VIS face images but also facilitates the accuracy improvement of heterogeneous face recognition.
Ran He 0001, Jie Cao 0002, Lingxiao Song, Zhenan Sun, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Geometry-Aware Face Completion and Editing
abstract
Face completion is a challenging generation task because it requires generating visually pleasing new pixels that are semantically consistent with the unmasked face region. This paper proposes a geometry-aware Face Completion and Editing NETwork (FCENet) by systematically studying facial geometry from the unmasked region. Firstly, a facial geometry estimator is learned to estimate facial landmark heatmaps and parsing maps from the unmasked face image. Then, an encoder-decoder structure generator serves to complete a face image and disentangle its mask areas conditioned on both the masked face image and the estimated facial geometry images. Besides, since low-rank property exists in manually labeled masks, a low-rank regularization term is imposed on the disentangled masks, enforcing our completion network to manage occlusion area with various shape and size. Furthermore, our network can generate diverse results from the same masked input by modifying estimated facial geometry, which provides a flexible mean to edit the completed face appearance. Extensive experimental results qualitatively and quantitatively demonstrate that our network is able to generate visually pleasing face completion results and edit face attributes as well.
Linsen Song, Jie Cao 0002, Lingxiao Song, Yibo Hu 0001, Ran He 0001
AAAI2
2019 Pose-preserving Cross Spectral Face Hallucination
abstract
To narrow the inherent sensing gap in heterogeneous face recognition (HFR), recent methods have resorted to generative models and explored the ?recognition via generation? framework. Even though, it remains a very challenging task to synthesize photo-realistic visible faces (VIS) from near-infrared (NIR) images especially when paired training data are unavailable. We present an approach to avert the data misalignment problem and faithfully preserve pose, expression and identity information during cross-spectral face hallucination. At the pixel level, we introduce an unsupervised attention mechanism to warping that is jointly learned with the generator to derive pixel-wise correspondence from unaligned data. At the image level, an auxiliary generator is employed to facilitate the learning of mapping from NIR to VIS domain. At the domain level, we first apply the mutual information constraint to explicitly measure the correlation between domains and thus benefit synthesis. Extensive experiments on three heterogeneous face datasets demonstrate that our approach not only outperforms current state-of-the-art HFR methods but also produce visually appealing results at a high resolution.
Junchi Yu, Jie Cao 0002, Yi Li 0018, Xiaofei Jia, Ran He 0001
IJCAI2
2019 DaNet: Decompose-and-aggregate Network for 3D Human Shape and Pose Estimation
abstract
Reconstructing 3D human shape and pose from a monocular image is challenging despite the promising results achieved by most recent learning based methods. The commonly occurred misalignment comes from the facts that the mapping from image to model space is highly non-linear and the rotation-based pose representation of the body model is prone to result in drift of joint positions. In this work, we present the Decompose-and-aggregate Network (DaNet) to address these issues. DaNet includes three new designs, namely UVI guided learning, decomposition for fine-grained perception, and aggregation for robust prediction. First, we adopt the UVI maps, which densely build a bridge between 2D pixels and 3D vertexes, as an intermediate representation to facilitate the learning of image-to-model mapping. Second, we decompose the prediction task into one global stream and multiple local streams so that the network not only provides global perception for the camera and shape prediction, but also has detailed perception for part pose prediction. Lastly, we aggregate the message from local streams to enhance the robustness of part pose prediction, where a position-aided rotation feature refinement strategy is proposed to exploit the spatial relationship between body parts. Such a refinement strategy is more efficient since the correlations between position features are stronger than that in the original rotation feature space. The effectiveness of our method is validated on the Human3.6M and UP-3D datasets. Experimental results show that the proposed method significantly improves the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://github.com/HongwenZhang/DaNet-3DHumanReconstrution .
Hongwen Zhang 0001, Jie Cao 0002, Guo Lu, Wanli Ouyang, Zhenan Sun
ACM Multimedia2
2019 3D Aided Duet GANs for Multi-View Face Image Synthesis
abstract
Multi-view face synthesis from a single image is an ill-posed computer vision problem. It often suffers from appearance distortions if it is not well-defined. Producing photo-realistic and identity preserving multi-view results is still a not well-defined synthesis problem. This paper proposes 3D aided duet generative adversarial networks (AD-GAN) to precisely rotate the yaw angle of an input face image to any specified angle. AD-GAN decomposes the challenging synthesis problem into two well-constrained subtasks that correspond to a face normalizer and a face editor. The normalizer first frontalizes an input image, and then the editor rotates the frontalized image to a desired pose guided by a remote code. In the meantime, the face normalizer is designed to estimate a novel dense UV correspondence field, making our model aware of 3D face geometry information. In order to generate photo-realistic local details and accelerate convergence process, the normalizer and the editor are trained in a two-stage manner and regulated by a conditional self-cycle loss and a perceptual loss. Exhaustive experiments on both controlled and uncontrolled environments demonstrate that the proposed method not only improves the visual realism of multi-view synthetic images but also preserves identity information well.
Jie Cao 0002, Yibo Hu 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.1
2018 Learning a High Fidelity Pose Invariant Model for High-resolution Face Frontalization
abstract
Face frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a High Fidelity Pose Invariant Model (HF-PIM) to produce photographic and identity-preserving results. HF-PIM frontalizes the profiles through a novel texture warping procedure and leverages a dense correspondence field to bind the 2D and 3D surface spaces. We decompose the prerequisite of warping into dense correspondence field estimation and facial texture map recovering, which are both well addressed by deep networks. Different from those reconstruction methods relying on 3D data, we also propose Adversarial Residual Dictionary Learning (ARDL) to supervise facial texture map recovering with only monocular images. Exhaustive experiments on both controlled and uncontrolled environments demonstrate that the proposed method not only boosts the performance of pose-invariant face recognition but also dramatically improves high-resolution frontalization appearances.
Jie Cao 0002, Yibo Hu 0001, Hongwen Zhang 0001, Ran He 0001, Zhenan Sun
NeurIPS1