Nannan Wang 0001

dblp:10/8359-1 · DBLP profile ↗
← Back
353ranked-venue papers
12as first author
272since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 197 · 4 first-author · 153 since 2021Artificial intelligence and machine learning · 179 · 8 first-author · 132 since 2021Security and privacy · 25 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 10 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning
abstract
Continual learning (CL) aims to equip models with the ability to learn from a stream of tasks without forgetting previous knowledge. With the progress of vision-language models like Contrastive Language-Image Pre-training (CLIP), their promise for CL has attracted increasing attention due to their strong generalizability. However, the potential of rich textual semantic priors in CLIP in addressing the stability–plasticity dilemma remains underexplored. During backbone training, most approaches transfer past knowledge without considering semantic relevance, leading to interference from unrelated tasks that disrupt the balance between stability and plasticity. Besides, while text-based classifiers provide strong generalization, they suffer from limited plasticity due to the inherent modality gap in CLIP. Visual classifiers help bridge this gap, but their prototypes lack rich and precise semantics. To address these challenges, we propose Semantic-Enriched Continual Adaptation (SECA), a unified framework that harnesses the anti-forgetting and structured nature of textual priors to guide semantic-aware knowledge transfer in the backbone and reinforce the semantic structure of the visual classifier. Specifically, a Semantic-Guided Adaptive Knowledge Transfer (SG-AKT) module is proposed to assess new images' relevance to diverse historical visual knowledge via textual cues, and aggregate relevant knowledge in an instance-adaptive manner as distillation signals. Moreover, a Semantic-Enhanced Visual Prototype Refinement (SE-VPR) module is introduced to refine visual prototypes using inter-class semantic relations captured in class-wise textual embeddings. Extensive experiments on multiple benchmarks validate the effectiveness of our approach.
De Cheng, Di Xu 0010, Huaijie Wang, Nannan Wang 0001
AAAI5
2026 Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
abstract
The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-tuning pre-trained diffusion models through Low-Rank Adaptation (LoRA) module to reconstruct high-resolution (HR) images. However, these dense Real-ISR models are limited in their ability to adaptively capture the heterogeneous characteristics of complex real-world degraded samples or enable knowledge sharing between inputs under equivalent computational budgets. To address this, we investigate the integration of sparse MoE into Real-ISR and propose a Mixture-of-Ranks (MoR) architecture for single-step image super-resolution. We introduce a fine-grained expert partitioning strategy that treats each rank in LoRA as an independent expert. This design enables flexible knowledge recombination while isolating fixed-position ranks as shared experts to preserve common-sense features and minimize routing redundancy. Furthermore, we develop a degradation estimation module leveraging CLIP embeddings and predefined positive-negative text pairs to compute relative degradation scores, dynamically guiding expert activation. To better accommodate varying sample complexities, we incorporate zero-expert slots and propose a degradation-aware load-balancing loss, which dynamically adjusts the number of active experts based on degradation severity, ensuring optimal computational resource allocation. Comprehensive experiments validate our framework's effectiveness and state-of-the-art performance.
Xiao He 0014, Zhijun Tu, Mingrui Zhu, Jie Hu 0021, Nannan Wang 0001, Xinbo Gao 0001
AAAI6
2026 Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud Removal
abstract
Cloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct structured information beneath cloud-covered areas. In this work, we propose a Visibility-guided Semantic Estimation and Reconstruction network for cloud removal (VISER-CR), which reformulates CR as a structure-guided completion problem. Specifically, VISER-CR explicitly models cloud interference via spatial masking, encouraging the model to reason beyond pixel-level appearance and enhance scene-level structural understanding. Moreover, to further improve the representation of structural information, we introduce Patch Saliency Encoding, a self-guided mechanism that implicitly models structural alignment among patches, significantly enhancing clustering consistency and semantic separability in the latent space. This adaptive mechanism guides the network to focus on learning and reconstructing structurally important regions, thereby reducing redundancy and improving overall cloud removal performance. Extensive experiments on multiple benchmark datasets demonstrate the superior effectiveness of our method.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001
AAAI4
2026 A Multi-Granularity Scene-Aware Graph Convolution Method for Weakly Supervised Person Search
De Cheng, Haichun Tai, Nannan Wang 0001, Xiangqian Zhao, Jie Li 0001, Xinbo Gao 0001
Int. J. Comput. Vis.3
2026 EKPC: Elastic Knowledge Preservation and Compensation for Class-Incremental Learning
Huaijie Wang, De Cheng, Yan Li 0125, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.6
2026 One-step diffusion-based real-world image super-resolution with visual perception distillation
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neurocomputing6
2026 Symmetrical bidirectional knowledge alignment for zero-shot sketch-based image retrieval
abstract
This paper studies the problem of zero-shot sketch-based image retrieval (ZS-SBIR), which aims to use sketches from unseen categories as queries to match the images of the same category. Due to the large cross-modality discrepancy, ZS-SBIR is still a challenging task and mimics realistic zero-shot scenarios. The key is to leverage transferable knowledge from the pre-trained model to improve generalizability. Existing researchers often utilize the simple fine-tuning training strategy or knowledge distillation from a teacher model with fixed parameters, lacking efficient bidirectional knowledge alignment between student and teacher models simultaneously for better generalization. In this paper, we propose a novel Symmetrical Bidirectional Knowledge Alignment for zero-shot sketch-based image retrieval (SBKA). The symmetrical bidirectional knowledge alignment learning framework is designed to effectively learn mutual rich discriminative information between teacher and student models to achieve the goal of knowledge alignment. Instead of the former one-to-one cross-modality matching in the testing stage, a one-to-many cluster cross-modality matching method is proposed to leverage the inherent relationship of intra-class images to reduce the adverse effects of the existing modality gap. Experiments on several representative ZS-SBIR datasets (Sketchy Ext dataset, TU-Berlin Ext dataset and QuickDraw Ext dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is publicly available at https://github.com/zermatt-luo/SBKA.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
Neural Networks4
2026 AdaAlign: A unified solution for traditional and modern zero-shot sketch-based image retrieval
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks4
2026 Isolating Interference Factors for Robust Cloth-Changing Person Re-Identification
abstract
Cloth-Changing Person Re-Identification (CC-ReID) aims to recognize individuals across camera views despite clothing variations, a crucial task for surveillance and security systems. Existing methods typically frame it as a cross-modal alignment problem but often overlook explicit modeling of interference factors such as clothing, viewpoints, and pedestrian actions. This oversight can distort their impact, compromising the extraction of robust identity features. To address these challenges, we propose a novel framework that systematically disentangles interference factors from identity features while ensuring the robustness and discriminative power of identity representations. Our approach consists of two key components. First, a dual-stream identity feature learning framework leverages a raw image stream and a cloth-isolated stream, to extract identity representations independent of clothing textures. An adaptive cloth-irrelevant contrastive objective is introduced to mitigate identity feature variations caused by clothing differences. Second, we propose a Text-Driven Conditional Generative Adversarial Interference Disentanglement Network (T-CGAIDN), to further suppress interference factors beyond clothing textures, such as finer clothing patterns, viewpoint, background, and lighting conditions. This network incorporates a multi-granularity interference recognition branch to learn interference-related features, a conditional adversarial module for bidirectional transformation between identity and interference feature spaces, and an interference decoupling objective to eliminate interference dependencies in identity learning. Extensive experiments on public benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, highlighting its effectiveness in CC-ReID.
De Cheng, Chaowei Fang, Shizhou Zhang, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
abstract
Domain Generalization (DG) seeks to develop models that perform well on unseen target domains by learning domain-invariant representations. Recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have shown strong potential for enhancing DG through prompt tuning. However, existing VFM-based prompt tuning methods often focus on task-specific adaptation rather than disentangling domain-invariant features, leaving cross-domain generalization insufficiently explored. In this paper, we address this challenge by fully leveraging the controllable and flexible language prompt in VFMs. Observing that the text modality is inherently rich in semantics and easier to disentangle, we propose a novel framework termed Prompt Disentanglement via Language Guidance and Representation Alignment (PADG). PADG first employs a large language model (LLM) to disentangle textual prompts into domain-invariant and domain-specific components, which then guide the learning of domain-invariant visual representations. To complement the limitations of text-only guidance, we further introduce the Worst Explicit Representation Alignment (WERA) module, which enhances visual invariance by simulating bounded domain shifts through learnable stylization prompts and aligning representations between original and perturbed samples. Extensive experiments on mainstream DG benchmarks, including PACS, VLCS, OfficeHome, DomainNet, and TerraInc, demonstrate that PADG consistently outperforms existing state-of-the-art methods, validating its effectiveness in robust domain-invariant representation learning.
De Cheng, Xinyang Jiang, Dongsheng Li 0002, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 An Enhanced Adaptive Confidence Margin for Semi-Supervised Facial Expression Recognition
abstract
Semi-supervised learning (SSL) provides a practical framework for leveraging massive unlabeled samples, especially when labels are expensive for facial expression recognition (FER). Typical SSL methods like FixMatch select unlabeled samples with confidence scores above a fixed threshold for training. However, these methods face two primary limitations: failing to consider the varying confidence across facial expression categories and failing to utilize unlabeled facial expression samples efficiently. To address these challenges, we propose an Enhanced Adaptive Confidence Margin (EACM), consisting of dynamic thresholds for different categories, to fully learn unlabeled samples. Specifically, we employ the predictions on labeled samples at each training iteration to learn an EACM. It then partitions unlabeled samples into two subsets: (1) subset I, including samples whose confidence scores are no less than the margin; (2) subset II, including samples whose confidence scores are less than the margin. For samples in subset I, we constrain their predictions on strongly-augmented versions to match the pseudo-labels derived from the predictions on weakly-augmented versions. Meanwhile, we introduce a feature-level contrastive objective to enhance the similarity between two weakly-augmented features of a sample in subset II. We extensively evaluate EACM on image-based and video-based facial expression datasets, showing that our method achieves superior performance, significantly surpassing fully-supervised baselines in a semi-supervised manner. Additionally, our EACM is promising to leverage cross-dataset unlabeled samples for practical training to boost fully-supervised performance.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 SSD: Making Face Forgery Clues Evident Again With Self-Steganographic Detection
abstract
The rapid development of generative AI techniques enables the synthesis of highly realistic facial images, posing significant challenges for the accurate detection of face forgeries. In contrast to solely elevating detector awareness, proactively reducing the intrinsic difficulty of forgery detection can streamline detector complexity while improving both generalization and robustness. This insight motivates our defense strategy to make face forgery clues more evident. Specifically, a novel proactive approach dubbed Self-Steganographic Detection (SSD) is proposed to imperceptibly embed facial images into themselves as a form of detection evidence. The recovery process is designed to remain robust under normal manipulations while exhibiting deliberate degradation under malicious manipulations, thereby clearly revealing potential forgeries. Unlike embedding bit-level vectors, pixel-level images are informative to ensure the generalization of our approach. Due to the similarity between the protected and embedded images, SSD performs detection without storing any embedded information in advance. To support practical deployment, our approach incorporates a dual detection scheme that aims to identify unprotected images and determine the authenticity of protected images. Extensive experiments using 8 face forgery techniques demonstrate the effectiveness of our approach compared to state-of-the-art methods.
Ruiyang Xia, Dawei Zhou 0004, Lin Yuan 0002, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Strength-Adaptive Adversarial Training
abstract
Adversarial training (AT) has been shown to effectively enhance a network's resilience against adversarial attack. However, conventional AT, which relies on a fixed pre-specified perturbation budget, suffers from several limitations when training robust models. First, enforcing the same perturbation budget across networks with different capacities leads to varying levels of robustness disparity between natural and robust accuracies, which deviates from the desired outcome of a robust network. Second, because the perturbation budget is fixed throughout training, the attack strength fails to scale adaptively with the evolving robustness of the model. This mismatch often results in robust overfitting and further degradation of adversarial robustness. To address these limitations, we propose a novel technique called Strength-Adaptive Adversarial Training (SAAT). In SAAT, the adversary incorporates an adversarial-loss constraint to guide the generation of adversarial training data. This constraint allows the perturbation budget to adapt dynamically based on the current training state, which effectively mitigates robust overfitting. Moreover, by explicitly regulating the attack strength through the adversarial loss, SAAT enables precise control over the robustness disparity between natural accuracy and adversarial robustness. Extensive experiments demonstrate that SAAT substantially improves adversarial robustness over standard AT.
Chaojian Yu, Dawei Zhou 0004, Li Shen 0008, Jun Yu 0001, Bo Han 0003, Mingming Gong, Nannan Wang 0001, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 VPT-NSP2++: Importance-Aware Visual Prompt Tuning in Null Space for Continual Learning
abstract
Continual learning (CL) enables AI models to adapt to evolving environments while mitigating catastrophic forgetting, which is a critical capability for dynamic real-world applications. With the growing popularity of pre-trained Vision Transformer (ViT) models and visual prompt tuning (VPT) technique in CL, this work explores a CL method on top of the ViT-based foundation model, through VPT mechanism with theoretical guarantees. Inspired by the orthogonal projection method, we aim to leverage this approach for VPT to enhance CL performance, particularly in long-term scenarios. However, since the orthogonal projection is originally designed for linear operations in CNNs, applying it to ViTs poses challenges induced by the non-linear self-attention mechanism and the distribution drift within LayerNorm. To address these issues, we deduced two orthogonality conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of maintaining stability. Considering the strict orthogonal constraints can diminish model capacity and reduce plasticity, we further propose an importance-aware orthogonal regularization framework. By applying varying degrees of orthogonal constraints to different parameters based on their importance to old and new tasks, the framework adaptively enhances model capacity and thereby promotes long-sequence CL while improving the stability-plasticity trade-off. To implement the proposed approach, a null-space-based approximation solution is employed to efficiently achieve the prompt gradient orthogonal projection. Extensive experiments on various class-incremental learning benchmarks demonstrate that our method achieves state-of-the-art performance across diverse CL scenarios.
Shizhou Zhang, Yue Lu 0008, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 TransFA: Transformer-based representation for face attribute evaluation
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
Pattern Recognit.4
2026 IDCFace: Identity Consistent Face anonymization for secure recognition
Ruiying Lu, Shuang Wan, Zimin Miao, Nannan Wang 0001, Chunlei Peng
Pattern Recognit.4
2026 FST: Improving adversarial robustness via feature similarity-based targeted adversarial training
Dawei Zhou 0004, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.4
2026 Condense loss: Exploiting vector magnitude during person Re-identification training process
Xi Yang 0011, Wenjiao Dong, Yingzhi Tang, Gu Zheng, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.5
2026 IMEVSI: Online Adaptive Video Stream Interpolation via Inertia-Aware Motion Estimation
abstract
Recent video frame interpolation (VFI) methods rely on computationally heavy modules (e.g., global attention module) to handle large motions, incurring prohibitive costs which hinders their practical real-time deployment. In this work, we revisit the core objective of VFI: enhancing the temporal resolution of videos. We identify that previous VFI’s frame-isolated processing ignores continuous temporal modeling, introduces computational redundancy in video streaming scenarios. To address this, we propose an online learning recurrent net with inertia-aware motion estimation(IMEVSI). It consists of implicit motion propagation( IMP), explicit motion propagation(EMP) and adaptive online learning strategy(AOL). IMP and EMP are used to high order inter-frame motion modeling considering motion inertia, AOL are proposed to bridge the motion domain gap between training and deployment. For IMP, we initiate from explicit physical motion modeling, progressively integrating learnable parameters into inertia-ware motion extraction and finally unify motion propagation and extraction within our recurrent motion propagation Transformer(RMPT). For EMP, we directly inject adjacent motion into current flow estimation recognizing its inertia contribution. For AOL,we leverage cycle consistency to dynamically adjust intermediate flow estimator and maintains an adaptive threshold to control parameter update. Extensive experiments demonstrate that our method outperforms state-of-the-art (SOTA) approaches on regular, large-motion, and high-resolution benchmarks while achieving excellent inference speed and FLOPs.
Keyi Chen 0015, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Adaptive Consensus Multi-Teacher Distillation for Generalizable Face Forgery Detection
abstract
Recent advancements in deep learning have significantly lowered the cost of generating and processing facial images, but this progress has also presented a growing challenge for face forgery detection. Existing detection methods, however, often struggle to generalize effectively to forged samples generated by previously unseen forgery techniques. Additionally, transferring knowledge through knowledge distillation to enhance generalization remains a persistent challenge. To address these issues, this paper proposes Adaptive Consensus Multi-teacher Knowledge Distillation (ACMD), a novel framework aimed at improving the generalization capabilities of face forgery detection models. ACMD leverages multiple teacher models with varying network architectures and introduces a consensus mechanism to resolve knowledge conflicts arising from the diversity of teacher models. It further adapts the assignment of sample-specific distillation weights, based on each teacher model’s prediction confidence. By combining this with feature-based knowledge distillation, the student model gains a deeper understanding of the teacher’s knowledge. Extensive experiments on the DeepfakeBench benchmark demonstrate that ACMD not only outperforms a single teacher model but also achieves state-of-the-art performance in dataset evaluations. Specifically, in a cross-dataset evaluation from FaceForensics++ to CDFv2 and DFDC, ACMD achieves frame-level AUCs of 84.22% and 74.43%, respectively, surpassing all baseline models. Ablation studies further validate the effectiveness of each component and highlight their complementary contributions to the overall detection performance.
Jiuyao Jing, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 DeepFidelity: Perceptual Forgery Fidelity Assessment for Deepfake Detection
abstract
Deepfake detection refers to detecting artificially generated or edited faces in images or videos, which plays an essential role in visual information security. Despite promising progress in recent years, Deepfake detection remains a challenging problem due to the complexity and variability of face forgery techniques. Existing Deepfake detection methods are often devoted to extracting features by designing sophisticated networks but ignore the influence of perceptual quality of faces. Considering the complexity of the quality distribution of real and fake faces, we propose a deepfake detection framework called DeepFidelity, which mines the perceptual forgery fidelity of face images and introduces a quality-aware scoring mechanism to distinguish real and fake faces of different image qualities. Specifically, we improve the model’s ability to identify complex samples by mapping real and fake face data of different qualities to different scores to distinguish them in a more detailed way. In addition, we propose a network structure called Symmetric Spatial Attention Augmentation based vision Transformer (SSAAFormer), which uses the symmetry of face images to promote the network to model the geographic long-distance relationship at the shallow level and augment local features. Extensive experiments on multiple benchmark datasets demonstrate the superiority of the proposed method over state-of-the-art methods. The code is available athttps://github.com/shimmer-ghq/DeepFidelity.
Chunlei Peng, Huiqing Guo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Probabilistic Distribution Alignment for Text-Based Person Retrieval
Xi Yang 0011, Chenghuan Qi, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Nearest Neighbor Sample Constraint and ODE Guided Feature Reconstruction for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification aims to retrieve a given pedestrian image from unlabeled data. The method of clustering and assigning pseudo-labels has become mainstream, but there are still some problems that will reduce recognition accuracy. On the one hand, in the process of clustering, poor classification of hard samples between neighboring classes leads to inadequate clustering accuracy, which affects the quality of pseudo-labels. On the other hand, the representational capacity of features extracted by the backbone network is also crucial for the model’s performance. To this end, this paper proposes an unsupervised person re-identification method based on nearest neighbor sample constraint and ordinary differential equation guided feature reconstruction (NNSC-FR) to improve the clustering accuracy and pseudo-label quality while enhancing the representation of features. Specifically, we present a novel nearest neighbor sample constraint (NNSC) after neighbor sample mining for each instance sample to recognize the hard samples’ fine classification between classes. To further improve clustering accuracy, an inter-class balance loss (CB loss) is introduced to better identify the hard samples between the nearest neighbor classes. In addition, guided by the third-order adam solution of the Ordinary Differential Equation, we design a Feature Reconstruction (ODE-FR) module with residual structure to improve the model representation ability. Extensive experimental results on Market-1501, DukeMTMC-reID, and MSMT17 demonstrate that our proposed method is superior to the state-of-the-art methods.
Xi Yang 0011, Wenjiao Dong, Gu Zheng, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 AdaNoise: Cycle-Consistent Image Translation With Domain-Adaptive Noise Perturbation
abstract
Image-to-image (I2I) translation aims to transform an input image into a target domain while preserving its structural details. Recent advances in diffusion-based generative models have significantly improved the perceptual quality of generated images; however, these approaches still face challenges in controllability and consistency, largely due to the inherent randomness introduced by stochastic noise during the generation process. Specifically, directly manipulating noise distributions without semantic alignment can lead to mode collapse, texture distortion, or loss of domain-specific features. To overcome these challenges, we propose the Adaptive Noise Framework (AdaNoise), a novel and cycle-consistent I2I translation approach guided by domain-adaptive noise modulation. AdaNoise introduces a Domain-Adaptive Noise Perturbation (DANP) module, which adaptively learns structured noise patterns aligned with the target domain distribution, enhancing both the expressiveness and reliability of the translation process. Through integration with a Cycle-Consistent Dual Diffusion (CDD) architecture, AdaNoise ensures faithful content reconstruction while allowing semantically meaningful domain shifts. The framework is designed to maintain a balance between generation quality and controllability, enabling more faithful and flexible image translation. Extensive experiments on two tasks, including SAR-to-optical image translation and low-light image enhancement, validate that AdaNoise not only surpasses existing state-of-the-art methods in terms of image fidelity and semantic preservation, but also achieves more controllable and diverse outputs across varying conditions, thus offering a robust and scalable solution for cross-domain visual generation.
Xi Yang 0011, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 CSHNet: A Novel Information Asymmetric Image Translation Method
abstract
Despite the considerable advancements in cross-domain image translation, a significant challenge remains in addressing information asymmetric translation tasks such as SAR-to-Optical and Sketch-to-Instance conversions. These tasks involve transforming data from a domain with limited information into one with more detailed and richer content. Traditional CNN-based methods, while effective at capturing intricate details, often struggle to grasp the overall structural composition of the image, leading to unintended blending or merging of distinct regions within the generated images. In light of these limitations, research has increasingly turned toward Transformers. Though Transformers excel at capturing global structures, they often lack the ability to preserve fine-grained details. Recognizing the importance of both detailed features and structural relationships in information asymmetric translation tasks, we introduce the CNN-Swin Hybrid Network (CSHNet). This network employs a novel bottleneck architecture featuring two key modules: Swin Embedded CNN (SEC) and CNN Embedded Swin (CES), which together form the SEC-CES-Bottleneck (SCB). Within this structure, SEC capitalizes on CNN’s capability for detailed feature extraction while incorporating the Swin Transformer’s inherent structural bias. In contrast, CES preserves the Swin Transformer’s strength in maintaining global structural integrity, while compensating for CNN’s tendency to emphasize detail. In addition to the SCB architecture, CSHNet integrates two essential components designed to improve cross-domain information retention and ensure structural consistency. The Interactive Guided Connection (IGC) fosters dynamic information exchange between SEC and CES, encouraging a deeper understanding of image details. At the same time, Adaptive Edge Perception Loss (AEPL) is implemented to preserve well-defined structural boundaries throughout the translation process. Experimental evaluations demonstrate that CSHNet surpasses current state-of-the-art methods, achieving superior results in both visualization and performance metrics across scene-level and instance-level datasets. Our code is available at: https://github.com/XduShi/CSHNet.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Semantic-Interactive Clustering Optimization With SAM for Weakly Supervised Person Search
abstract
Weakly-supervised person search presents significant challenges when relying solely on bounding-box annotations, particularly due to inter-class confusion from clothing similarity and intra-class variations caused by illumination changes, which severely degrade cross-view matching accuracy. Existing clustering-based methods, constrained by their heavy dependence on color features, frequently produce unreliable pseudo-labels that ultimately limit model performance. To overcome these limitations, we present Segment Anything Model-based Semantic-Interactive Clustering Optimization (SAM-SICO), a novel framework that integrates the Segment Anything Model’s semantic segmentation capability with adaptive clustering optimization for weakly-supervised person search. Our framework harnesses the representational power of the Segment Anything Model (SAM) to enable detector-free semantic feature learning while significantly improving clustering precision. The proposed solution makes three key advances: the Semantic Contour Embedding (SCE) module leverages SAM’s zero-shot segmentation capability to produce highly accurate human body masks; the Relation-driven Semantic Feature Interaction (RSFI) mechanism effectively mitigates clothing-color bias through innovative dynamic affinity matrix construction across multiscale semantic masks and visual features; and the Adaptive Clustering Optimization (ACO) algorithm introduces parameter adaptation to optimize intra-class compactness and inter-class separation metrics. Experimental results show that our method outperforms existing state-of-the-art approaches on the PRW and CUHK-SYSU datasets. The source code is available at https://github.com//HawlsonZ/SAM-SICO.
Xi Yang 0011, Hexun Zhou, De Cheng, Menghui Tian, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 FPAD: Fuzzy-Prototype-Guided Adversarial Attack and Defense for Deep Cross-Modal Hashing
abstract
Deep cross-modal hashing models generally inherit the vulnerabilities of deep neural networks, making them susceptible to adversarial attacks and thus posing a serious security risk during real-world deployment. Current adversarial attack or defense strategies often establish a weak correlation between the hashing codes and the targeted semantic representations, and there is still a lack of related works that simultaneously consider the attack and defense for deep cross-modal hashing. To alleviate these concerns, we propose a Fuzzy-Prototype-guided Adversarial Attack and Defense (FPAD) framework to enhance the adversarial robustness of deep cross-modal hashing models. First, an adaptive fuzzy-prototype learning network (FpNet) is efficiently presented to extract a set of fuzzy-prototypes, aiming to encode the underlying semantic structure of the heterogeneous modalities in both feature and Hamming spaces. Then, these derived prototypical hash codes are heuristically employed to supervise the generation of high-quality adversarial examples, while a fuzzy-prototype rectification scheme is simultaneously designed to preserve the latent semantic consistency between the adversarial and benign examples. By mixing the adversarial samples with the original training samples as the augmented inputs, an efficient fuzzy-prototype-guided adversarial learning framework is proposed to execute the collaborative adversarial training and generate robust cross-modal hash codes with high adversarial defense capabilities, therefore resisting various attacks and benefiting various challenging cross-modal hashing tasks. Extensive experiments evaluated on benchmark datasets show that the proposed FPAD framework not only produces high-quality adversarial samples to enhance the adversarial training process, but also shows its high adversarial defense capability to benefit various cross-modal hashing tasks. The code is available at: https://github.com/yzq131/FPAD.
Zhongqing Yu, Xin Liu 0011, Yiu-Ming Cheung, Lei Zhu 0002, Xing Xu 0001, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Fine-Detailed Facial Sketch-to-Photo Synthesis With Detail-Enhanced Codebook Priors
abstract
Generating high-quality facial photos from fine-detailed sketches is a long-standing research topic that remains unsolved. The scarcity of large-scale paired data due to the cost of acquiring hand-drawn sketches poses a major challenge. Existing methods either lose identity information with oversimplified representations, or rely on costly inversion and strict alignment when using StyleGAN-based priors, limiting their practical applicability. Our primary finding in this work is that the discrete codebook and decoder trained through self-reconstruction in the photo domain can learn rich priors, helping to reduce ambiguity in cross-domain mapping even with current small-scale paired datasets. Based on this, a cross-domain mapping network can be directly constructed. However, empirical findings indicate that using the discrete codebook for cross-domain mapping often results in unrealistic textures and distorted spatial layouts. Therefore, we propose a Hierarchical Adaptive Texture-Spatial Correction (HATSC) module to correct the flaws in texture and spatial layouts. Besides, we introduce a Saliency-based Key Details Enhancement (SKDE) module to further enhance the synthesis quality. Overall, we present a “reconstruct-cross-enhance” pipeline for synthesizing facial photos from fine-detailed sketches. Experiments demonstrate that our method generates high-quality facial photos and significantly outperforms previous approaches across a wide range of challenging benchmarks. The code is publicly available at: https://github.com/Gardenia-chen/DECP.
Mingrui Zhu, Jianhang Chen, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Component-Specific Prompt Tuning for Deepfake Detection
abstract
With the development of deep learning technology, the facial images generated by deepfake technology have reached a level of authenticity that is difficult to distinguish, posing a serious threat to personal privacy and data security. Therefore, it is of great significance to develop efficient and reliable deepfake detection technology. In recent years, Visual Language Models (VLM) have been applied to deepfake detection tasks due to their powerful multimodal understanding capabilities. However, the existing VLM have not been specifically optimized for deepfake detection tasks. When directly applied to this task, there are problems such as insufficient model accuracy and insufficient feature extraction, especially when dealing with complex forgery scenes. In response to these challenges, this paper proposes an innovative deepfake face detection method based on VLM and component-specific prompt tuning. We transform the deepfake detection task into a Visual Question Answering (VQA) task, making full use of the multimodal understanding capabilities of VLM and the flexibility of prompt tuning technology. This method uses a local prompt strategy to customize specific prompt questions for key facial components such as eyes, nose, and mouth, guiding the model to focus on the local features of these areas, thereby accurately capturing forgery traces. In addition, we introduced a feature extraction module Q-Former based on instructions, which can flexibly adjust the focus area of visual features according to prompts, significantly improving the model’s perception of locally forged features. By fusing these local features extracted by Q-Former and combining them with the language model to judge the authenticity of the overall face image, we can finally generate accurate prediction results. A large number of experimental results show that our method is significantly better than existing technologies in terms of detection accuracy and robustness.
Yinyin Chen, Huiqing Guo, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2026 Sparse VMamba: Robust Spatio-Temporal Information Modeling for Event Camera Person Re-Identification
abstract
Event camera-based person re-identification (Re-ID) effectively addresses the challenges faced by traditional Re-ID systems, such as privacy leakage, low-light imaging degradation, and motion blur. However, traditional Convolutional Neural Networks (CNNs) struggle to model long-range spatio-temporal dependencies, while the Transformer architecture encounters fundamental conflicts with second-order computational complexity and the high temporal resolution of event streams. Additionally, sparse data leads to wasted computational resources and diluted effective data. In contrast, the Mamba architecture, with its long-term modeling capability and linear complexity, is better suited for event stream data. Therefore, we innovatively explore the potential of VMamba in event camera-based person Re-ID; however, directly using VMamba does not fully leverage the temporal asynchronicity and spatial sparsity inherent in event data. To address this, we design a novel Sparse VMamba framework to construct a more robust spatio-temporal information extraction mechanism. First, we develop a Spatio-Temporal Information Modeling (STIM) module that simultaneously employs CNNs and Gated Recurrent Units (GRUs) for modeling spatial and temporal information. Then, we enhance the robustness of sparse data feature extraction using two strategies: on one hand, we utilize Anti-Noise Contour Enhancement (ANCE) module to improve motion contour features and mitigate sensor pulse noise; on the other hand, we implement Direction-Aware Sparse Perception (DASP) module to encourage the model to extract robust person descriptors. Results on the Event-ReID-v1 and Event-ReID-v2 datasets validate the effectiveness of our approach.
Wenjiao Dong, Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Inf. Forensics Secur.3
2026 ALIGNER: Learning Fine-Grained Cross-Modal Alignment for Text-Based Person Retrieval
Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Inf. Forensics Secur.3
2026 Video Frame Interpolation via Appearance-Based Intermediate Flow Estimation
abstract
Intermediate flow estimation is an important part of video frame interpolation (VFI). Most previous works use interpolation to derive the intermediate flow assuming localized linear motion. However, this method is not effective when dealing with extreme motions. In this work, we assume that the motion trajectory of an object is determined by the appearance characteristics of this object. Based on this assumption, we propose a new intermediate flow estimation method, which obtains the motion features of intermediate frames from image appearance and inter-frame motion features. In addition, in order to fully extract the inter-frame features, we rethink the difference of VFI and previous works on using Swin-Transformer and compute the appearance features and motion features within the adaptive neighborhood by cyclically shifting the window. Experimental results show that our method achieves state-of-the-art performance on different datasets for both fixed-time and arbitrary-time interpolation. Moreover, our proposed method outperforms models that require inputting a sequence of four frames when handling videos with extremely large motion. The source code is available from https://github.com/chen12304/IFE-VFI.
Keyi Chen 0015, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2026 Toward Semantically Enhanced Representation Learning for Text-Based Person Retrieval
abstract
Text-Based Person Retrieval (TBPR), which is a pivotal technology in the intelligent surveillance field, is aimed at retrieving target pedestrians based on free-form textual descriptions. While the existing methods attempt to align cross-modal features via multigranular interactions, their performance remains fundamentally limited by two core challenges: cross-modal semantic inconsistency and cross-modal semantic discriminability. To address these issues, we propose DSEE (Diversity Semantic Embedding Expansion), a novel framework for semantically enhanced representation learning. Unlike approaches that rely on constructing larger or more detailed datasets, DSEE establishes identity-centric cross-modal consistency through contrastive learning and generative synergy. The framework consists of two key modules: Bidirectional-guided Semantic Modeling (BSM) and Generative-driven Semantic Enhancement (GSE) modules. The BSM module constructs novel semantic embeddings by modeling similarity-based interactions between the image and text modalities. Specifically, it emphasizes identity-level similarity to guide the generation of enriched, discriminative semantic representations, thereby enhancing their semantic expressiveness and cross-modal alignment. The GSE module provides enriched semantic diversity through a generative text augmentation scheme based on visual inputs, while refining the semantic precision of the method via a dual-path attention mechanism that performs both intramodal refinement and cross-modal alignment. Extensive experiments demonstrate that DSEE achieves state-of-the-art performance on major benchmarks across diverse scenarios. Our work provides an effective paradigm for advancing TBPR applications in real-world settings.
Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Image Process.3
2026 One Step Diffusion-Based Super-Resolution With Time-Aware Distillation
abstract
iffusion-based image super-resolution (SR) has shown strong potential in recovering high-fidelity details from low-resolution inputs. However, the need for tens or hundreds of sampling steps leads to substantial inference latency. Recent works attempt to accelerate this process via knowledge distillation, but often rely solely on pixel-level loss or overlook the fact that diffusion models capture different information across time steps. To address this, we propose TAD-SR, a time-aware diffusion distillation framework. Specifically, we introduce a novel score distillation strategy to align the score functions between the outputs of the student and teacher models after minor noise perturbation. This distillation strategy eliminates the inherent bias in score distillation sampling (SDS) and enables the student models to focus more on high-frequency image details by sampling at smaller time steps. We further introduce a time-aware discriminator that exploits the teacher’s knowledge to differentiate real and synthetic samples across different noise scales, using explicit temporal conditioning. Extensive experiments on SR tasks demonstrate that TAD-SR outperforms existing singl-estep diffusion methods and achieves performance on par with multi-step state-of-the-art models.iffusion-based image super-resolution (SR) has shown strong potential in recovering highfidelity details from low-resolution inputs. However, the need for tens or hundreds of sampling steps leads to substantial inference latency. Recent works attempt to accelerate this process via knowledge distillation, but often rely solely on pixel-level loss or overlook the fact that diffusion models capture different information across time steps. To address this, we propose TADSR, a time-aware diffusion distillation framework. Specifically, we introduce a novel score distillation strategy to align the score functions between the outputs of the student and teacher models after minor noise perturbation. This distillation strategy eliminates the inherent bias in score distillation sampling (SDS) and enables the student models to focus more on highf-requency image details by sampling at smaller time steps. We further introduce a time-aware discriminator that exploits the teacher’s knowledge to differentiate real and synthetic samples across different noise scales, using explicit temporal conditioning. Extensive experiments on SR tasks demonstrate that TAD-SR outperforms existing single-step diffusion methods and achieves performance on par with multi-step state-of-the-art models D.
Xiao He 0014, Huaao Tang, Zhijun Tu, Hanting Chen, Mingrui Zhu, Jie Hu 0021, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.10
2026 Toward Universal Semantic Communication via Matchable Semantic Subspace Transmission
abstract
Semantic communication targets reliable task execution at the receiver under stringent bandwidth and channel constraints. However, existing communication paradigms either focus on bit-level signal reconstruction, impeding the balance between task efficacy and bandwidth efficiency, or are limited by fixed vocabularies and lack generalization when facing unknown categories and open scenarios. To this end, we propose Universal Semantic Communication (UniSC), an open-vocabulary semantic communication framework that formulates transmission as a Matchable Semantic Subspace Transmission (MSST) problem. In this work, "universal" refers to the ability to handle arbitrary text-defined semantic categories beyond fixed vocabularies, rather than universality across all vision tasks. The transmitted representation is explicitly constrained to preserve cross-modal matchability after noisy transmission, rather than merely supporting latent recovery or closed-set inference. Concretely, UniSC comprises a Visual Semantic Engine (VSE), a Semantic Squeeze Network (SSN), a Noise-Adaptive Semantic Re-expansion (NASR) module, and a VLM-based Decoder. VSE and SSN project images into a compact semantic subspace for transmission. This subspace is optimized to preserve both robustness and cross-modal matchability under channel corruption. NASR denoises and lifts the received features back into a semantically complete visual space, from which the VLM-based Decoder performs open-category inference by matching arbitrary text queries rather than relying on a fixed classifier head. The VLM-based Decoder employs a Text Semantic Engine (TSE) to map natural language to text embeddings and, via a learnable Text-Visual Bridge (TVB), aligns them with the reconstructed visual structure for cross-modal matching. To improve cross-modal alignment and transmission robustness, a two-stage training strategy first establishes cross-modal anchors and then optimizes end-to-end robustness and compactness. Extensive experiments on semantic segmentation benchmarks demonstrate that UniSC achieves strong generalization and state-of-the-art performance under harsh channel conditions, outperforming existing methods in both low-SNR and extreme-compression regimes.
Xi Yang 0011, Songsong Duan, Nannan Wang 0001
IEEE Trans. Image Process.4
2026 Hierarchical Identity Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised visible-infrared person re-identification (USVI-ReID) aims to learn modality-invariant image features from unlabeled cross-modal person datasets by reducing the modality gap while minimizing reliance on costly manual annotations. Existing methods typically address USVI-ReID using cluster-based contrastive learning, which represents a person by a single cluster center. However, they primarily focus on the commonality of images within each cluster while neglecting the finer-grained differences among them. To address the limitation, we propose a Hierarchical Identity Learning (HIL) framework. Since each cluster may contain several smaller sub-clusters that reflect fine-grained variations among images, we generate multiple memories for each existing coarse-grained cluster via a secondary clustering. Additionally, we propose Multi-Center Contrastive Learning (MCCL) to refine representations for enhancing intra-modal clustering and minimizing cross-modal discrepancies. To further improve cross-modal matching quality, we design a Bidirectional Reverse Selection Transmission (BRST) mechanism, which establishes reliable cross-modal correspondences by performing bidirectional matching of pseudo-labels. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate that the proposed method outperforms existing approaches. The source code is available at: https://github.com/haonanshi0125/HIL.
De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.5
2026 Interpretable General Image Fusion via Scalable Autoregressive Modeling
abstract
Existing image fusion methods have developed increasingly sophisticated network architectures for exploiting modality-shared and modality-specific features. However, despite these advancements in feature extraction, most methods ultimately rely on relatively simple implicit or explicit fusion strategies, which can compromise interpretability and limit fusion accuracy. In this paper, we incorporate visual autoregressive modeling to bridge the gap between implicit feature extraction and explicit modality fusion. First, the proposed approach conducts a low-to-high resolution autoregressive objective with modality-specific features, introducing a scalable feature autoregressive mechanism. It aggregates local and global contextual dependencies while enhancing implicit cross-scale interaction. Furthermore, to promote the consistency and complementarity across modalities, we embed an explicit high-order fusion strategy within the progressive modality-specific feature extraction process. This integration facilitates a next-scale synergistic relationship between implicit learning and explicit fusion. Our High-order Feature AutoRegressive Fusion framework (HFARFusion) provides a robust and interpretable solution for general image fusion tasks, effectively balancing fusion performance and transparency through the strengths of autoregressive learning. Extensive experiments demonstrate the outstanding performance of the proposed method in several classical fusion tasks, including infrared-visible, medical, multi-focus, and multi-exposure image fusion. Our code is available at https://github.com/happysbn/HFARFusion.
Jingwei Xin, Boneng Shi, Zhen Li 0026, Xuehao Song, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.6
2026 Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) images offer unique advantages in all-weather, all-day remote sensing, but the high acquisition costs and time-consuming annotation processes limit their widespread implementation. Semi-supervised domain adaptation leverages abundant annotated optical images and a small number of labeled SAR images to achieve great performance on SAR images. However, existing semi-supervised domain adaptation object detection methods typically select SAR domain labeled samples randomly, making it difficult to fully exploit the valuable information and distinctive features inherent in the target domain data. Moreover, there is a significant style and content gap between optical and SAR images, and previous methods have not adapted to them in a task-specific manner. To this end, this paper proposes an active style-content dual-branch domain adaptation method specifically designed for semi-supervised object detection in SAR images. The proposed approach employs Task-aware Active Sampling (TAS) module to select the most valuable SAR samples, addressing inefficiencies in random sampling. Also, we employ a dual-branch framework to address the style and content gaps between optical and SAR images. Multi-layer Feature Alignment (MFA) module ensures style alignment by maintaining consistent feature representations across different visual styles, while Gaussian-SAM Image Fusion (G-SIF) module is employed to integrate content from the source domain into the target domain, effectively bridging the gap between optical and SAR images. Extensive experiments on multiple ship and aircraft datasets demonstrate the exceptional generalization capabilities of our proposed model.
Xi Yang 0011, Quantao Xie, Yirong Yang, Nannan Wang 0001
IEEE Trans. Image Process.4
2026 Overcoming Dual Incremental Challenges in Continual Person Search via Adapter and Prototype
abstract
The advancement of continual person search techniques has seen significant progress in recent years due to its practical applications in the real world. However, continual learning for person search presents significant challenges as it combines both person detection and re-identification (Re-ID) tasks, resulting in issues of domain and class incremental learning. To address these challenges, we propose a novel framework that uses an adapter-based Swin Transformer backbone, and incorporates two key components: Domain Aware Adapter (DAA) blocks and Virtual Prototype Replay-Online Instance Matching (VPR-OIM). Specifically, to solve the domain incremental problem in object detection, we introduce parallel DAA blocks to handle multiple domains, while a Domain Prototype Router (DPR) mechanism is used to dynamically route the feature to the domain-specific adapter. Additionally, for class incremental Re-ID, we extend the OIM loss with virtual prototype replay, which generates Gaussian distribution-based virtual features derived from historical prototypes, effectively enabling the model to preserve knowledge of previous identities while accommodating new identity categories. Overall, our proposed DAA and VPR-OIM simultaneously address the dual incremental challenges of continual person search. Experimental results demonstrate that our method significantly improves both person detection and Re-ID performance in continual learning settings, achieving state-of-the-art (SOTA) performance.
Xi Yang 0011, Hexun Zhou, De Cheng, Nannan Wang 0001
IEEE Trans. Image Process.4
2026 Distribution-Aware Prompt Learning for Vision-Language Models With Dynamic Boundary Prototype
abstract
Prompt learning has emerged as an effective strategy for adapting vision-language models (VLMs) which injects learnable semantic prompts into VLMs to guide the alignment between visual and textual representations. Although existing methods have shown strong performance across various tasks, they usually focus on the representative class-level samples and overlook the atypical and hard samples in visual feature space, which hinders generalization of VLMs. To address this issue, we propose the concept of dynamic boundary prototype, which highlights ambiguous samples that are far from the class centroid and is updated at each epoch. Accordingly, we propose a Distribution-Aware Prompt Learning (DAPL) framework to calibrate the distribution of visual feature space via the definition, optimization, and updating of dynamic boundary prototypes. Firstly, we introduce Boundary-Centroid Pulling to optimize the intra-class distribution by progressively reducing the distance between boundary and centroid prototypes, thereby enhancing structural consistency within each class. Secondly, to further enhance inter-class separability, a distance-weighted contrastive loss that places greater emphasis on distinguishing adjacent classes is designed, facilitating more effective fine-grained discrimination. Thirdly, we apply Low-Rank Adaptation Fine-Tuning to adapt the vision encoder through targeted modifications to its self-attention layers. Additionally, we adopt a progressive training strategy for stable optimization. DAPL is compatible with mainstream prompt learning methods such as CoOp, CoCoOp and PromptKD, and consistently improves their average performance across 11 benchmark datasets.
Xi Yang 0011, Xinyue Zhong, Nannan Wang 0001
IEEE Trans. Image Process.3
2026 Prompt-Driven Knowledge Distillation for Remote Sensing Object Detection
abstract
Remote sensing object detection requires precise identification of multi-scale and multi-directional targets in com-plex backgrounds, demanding the model that achieves both high accuracy and real-time performance. While knowledge distillation proves effective for compressing natural image models, it exhibits limitations in more realistic remote sensing scenarios, including inadequate adaptability, biases from long-tail data distributions, and the propagation of errors from the teacher model. To address these challenges, we propose a Prompt Driven Knowledge Distillation (PDKD) framework for remote sensing object detection. This framework leverages prompt-based mechanisms to guide the student model in effectively acquiring and assimilating the teacher's knowledge, which integrates three core components: (1) Scale-Decoupled Feature Prompting (SDFP) module dynamically adjusts feature representation capabilities through scale decoupling, enabling differentiated distillation for targets of varying scales; (2) Semantic Visual Co-Prompting (SVCP) module, based on CLIP's multimodal prior knowledge, constructs category-specific semantic prompt vectors to enhance the focus on features of long-tail categories; (3) Self-Correcting Prompting (SCP) module that suppresses error propagation through a cross self-distillation mechanism. The experiments on the DOTA dataset show that with a 1x training schedule, the model achieves a 49.0% $mAP$ . Source codes are available at https://github.com/Ningsui/PDKD.git.
Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Image Process.4
2026 Generating Imperceptible Perturbations to Attack Human Pose Estimation Networks
abstract
Adversarial attacks on deep networks have received significant attention recently. However, most existing research focuses on classification tasks, with limited exploration of adversarial attacks on human pose estimation networks. To bridge this gap, we propose a novel attack framework specifically tailored for human pose estimation by exploiting unique HPE characteristics including heatmap-sensitive localization, joint-influence imbalance, and keypoint-focused perception. Our objective is to substantially diminish object keypoint similarity while introducing minimal perturbations to the image. We have devised a two-stage framework to implement the attack. The first stage involves a gradient attack framework that induces deviation in the adversarial heatmap from the original heatmap by exploiting heatmap-sensitive localization. In the second stage, the perturbations are optimized and restricted to the vicinity of keypoints to make the attacks imperceptible by exploiting joint-influence imbalance and keypoint-focused perception. To achieve this, we incorporate low-frequency constraints to limit the perturbations to high-frequency components and utilize a perceptual color distance metric to control the perturbation's magnitude. Extensive experimental results on the COCO and MPII datasets demonstrate that our attack can generate adversarial examples with high strength and low detectability.
Junlong Mu, Lin Zhao 0003, Di Wang 0011, Chen Gong 0002, Nannan Wang 0001
IEEE Trans. Multim.5
2025 QuARF: Quality-Adaptive Receptive Fields for Degraded Image Perception
abstract
Advanced Deep Neural Networks (DNNs) perform well for high-quality images, but their performance dramatically decreases for degraded images. Data augmentation is commonly used to alleviate this problem, but using too much perturbed data might seriously decrease the performance on pristine images. To tackle this challenge, we take our cue from the assumption of spatial coincidence in human visual perception, i.e. multiscale and varying receptive fields are required for understanding pristine and degraded images. Correspondingly, we propose a novel plug-and-play network architecture, dubbed Quality-Adaptive Receptive Fields (QuARF), to automatically select the optimal receptive fields based on the quality of the input image. To this end, we first design a multi-kernel convolutional block, which comprises multiscale continuous receptive fields. Afterward, we design a quality-adaptive routing network to predict the significance of each kernel, based on the quality features extracted from the input image. In this way, QuARF automatically selects the optimal inference route for each image. To further boost efficiency and effectiveness, the input feature map is split into multiple groups, with each group independently learning its quality-adaptive routing parameters. We apply QuARF to a variety of DNNs and conduct experiments in both discriminative and generation tasks, including semantic segmentation, image translation, and restoration. Thorough experimental results show that QuARF significantly and robustly improves the performance for degraded images, and outperforms data augmentation in most cases.
Fei Gao 0006, Ziyun Li 0002, Wenwang Han, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
AAAI8
2025 Effective Diffusion Transformer Architecture for Image Super-Resolution
abstract
Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image generation. In this work, we design an effective diffusion transformer for image super resolution (DiT-SR) that achieves the visual quality of prior-based methods, but through a training-from-scratch manner. In practice, DiT-SR leverages an overall U-shaped architecture, and adopts uniform isotropic design for all the transformer blocks across different stages. The former facilitates multi-scale hierarchical feature extraction, while the latter reallocate the computational resources to critical layers to further enhance performance. Moreover, we thoroughly analyze the limitation of the widely used AdaLN, and present a frequency-adaptive time-step conditioning module, enhancing the model's capacity to process distinct frequency information at different time steps. Extensive experiments demonstrate that DiT-SR outperforms the existing training-from-scratch diffusion-based SR methods significantly, and even beats some of the prior-based methods on pretrained Stable Diffusion, proving the superiority of diffusion transformer in image super resolution.
Zhijun Tu, Xiao He 0014, Liyu Chen, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Hu 0021
AAAI8
2025 Asymmetric Reinforcing Against Multi-Modal Representation Bias
abstract
The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning.
Xiyuan Gao, Bing Cao 0002, Pengfei Zhu 0001, Nannan Wang 0001, Qinghua Hu
AAAI4
2025 Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations
abstract
Due to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fair Forgery Detection (FairFD) dataset, where we prove the racial bias of public state-of-the-art (SOTA) methods. Different from existing forgery detection datasets, the self-constructed FairFD dataset contains a balanced racial ratio and diverse forgery generation images with the largest-scale subjects. Additionally, we identify the problems with naive fairness metrics when benchmarking forgery detection models. To comprehensively evaluate fairness, we design novel metrics including Approach Averaged Metric and Utility Regularized Metric, which can avoid deceptive results. We also present an effective and robust post-processing technique, Bias Pruning with Fair Activations (BPFA), which improves fairness without requiring retraining or weight updates. Extensive experiments conducted with 12 representative forgery detection models demonstrate the value of the proposed dataset and the reasonability of the designed fairness metrics. By applying the BPFA to the existing fairest detector, we achieve a new SOTA. Furthermore, we conduct more in-depth analyses to offer more insights to inspire researchers in the community.
Decheng Liu, Zongqi Wang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
AAAI4
2025 Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning
abstract
Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach.
Yue Lu 0008, Shizhou Zhang, De Cheng, Guoqiang Liang 0001, Yinghui Xing, Nannan Wang 0001, Yanning Zhang 0001
AAAI6
2025 Motion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion Model
abstract
Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (frequency domain) are not well considered, which limits their applications in the clinical field. To address these issues, we propose a novel unsupervised purification method which leverages pixel-frequency information of noisy MRI images to guide a pre-trained diffusion model to recover clean MRI images. Specifically, considering that motion artifacts are mainly concentrated in high-frequency components in k-space, we utilize the low-frequency components as the guide to ensure correct tissue textures. Additionally, given that high-frequency and pixel information are helpful for recovering shape and detail textures, we design alternate complementary masks to simultaneously destroy the artifact structure and exploit useful information. Quantitative experiments are performed on datasets from different tissues and show that our method achieves superior performance on several metrics. Qualitative evaluations with radiologists also show that our method provides better clinical feedback.
Dawei Zhou 0004, Lei Hu 0002, Feng Yang 0015, Zaiyi Liu, Nannan Wang 0001, Xinbo Gao 0001
AAAI7
2025 Mitigating Feature Gap for Adversarial Robustness by Feature Disentanglement
abstract
Adversarial fine-tuning methods enhance adversarial robustness via fine-tuning the pre-trained model in an adversarial training manner. However, we identify that some specific latent features of adversarial samples are confused by adversarial perturbation and lead to an unexpectedly increasing gap between features in the last hidden layer of natural and adversarial samples. To address this issue, we propose a disentanglement-based approach to explicitly model and further remove the specific latent features. We introduce a feature disentangler to separate out the specific latent features from the features of the adversarial samples, thereby boosting robustness by eliminating the specific latent features. Besides, we align clean features in the pre-trained model with features of adversarial samples in the fine-tuned model, to benefit from the intrinsic features of natural samples. Empirical evaluations on three benchmark datasets demonstrate that our approach surpasses existing adversarial fine-tuning methods and adversarial training baselines.
Nuoyan Zhou, Dawei Zhou 0004, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
AAAI4
2025 Optimizing Label Assignment for Weakly Supervised Person Search
abstract
Weakly supervised person search aims to detect and match individuals using only bounding box annotations jointly. The existing methods mainly alternate between the clustering stage and the training stage, where the former is responsible for instance level label allocation tasks and the latter needs to undertake proposal level label allocation tasks. In the clustering phase, the conventional use of the DBSCAN algorithm for clustering pedestrian instance features often neglects key contextual information such as scene context and relative positioning of individuals. During the training phase, the Region Proposal Network assigns labels based on the MaxIoU, which tends to produce locally ambiguous labels. Finally, the proposals updated to the memory bank with extensive background information tend to interfere with the task of pseudo-label generation. To address these issues, this paper proposes an Optimizing Label Assignment (OLA) for weakly supervised person search. Firstly, in the clustering phase, Context Aware Clustering is introduced to integrate contextual information and constraints, enhancing the accuracy of clustering. Secondly, in the training phase, we adopt Prototype Matching based on Optimal Transport theory to optimize label distribution from a global perspective. Furthermore, we propose Dual Memory Bank Enhancement that effectively enhances the accuracy of label assignment. Extensive experiments conducted on the CUHK-SYSU and PRW datasets demonstrate that our method achieves state-of-the-art performance in weakly supervised person search.
Xi Yang 0011, Nannan Wang 0001
AAAI3
2025 Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation
abstract
Existing Weakly Supervised Semantic Segmentation (WSSS) relies on the CNN-based Class Activation Map (CAM) and Transformer-based self-attention map to generate class-specific masks for semantic segmentation. However, CAM and self-attention maps usually cause incomplete segmentation due to classification bias issue. To address this issue, we propose a Multi-Label Prototype Visual Spatial Search (MuP-VSS) method with a spatial query mechanism. Specifically, MuP-VSS consists of two key components: multi-label prototype representation and multi-label prototype optimization. The former designs a global embedding to learn the global tokens from the images, and then proposes a Prototype Embedding Module (PEM) to interact with patch tokens to understand the local semantic information. The latter utilizes the exclusivity and consistency principles of the multi-label prototypes to design three prototype losses to optimize them, which contain cross-class prototype (CCP) contrastive loss, cross-image prototype (CIP) contrastive loss, and patch-to-prototype (P2P) consistency loss. CCP loss models exclusivity of multi-label prototypes learned from a single image to enhance the discriminative properties of each class better. CCP loss learns the consistency of the same class-specific prototypes extracted from multiple images to enhance the semantic consistency. P2P loss is proposed to control the semantic response of the prototype to the image patches. Experimental results on Pascal VOC 2012 and MS COCO show that MuP-VSS significantly outperforms recent methods and achieves state-of-the-art performance.
Songsong Duan, Xi Yang 0011, Nannan Wang 0001
CVPR3
2025 DynPose: Largely Improving the Efficiency of Human Pose Estimation by a Simple Dynamic Framework
abstract
Top-down approaches for human pose estimation (HPE) have reached a high level of sophistication, exemplified by models such as HRNet and ViTPose. Nonetheless, the low efficiency of top-down methods is a recognized issue that has not been sufficiently explored in current research. Our analysis suggests that the primary cause of inefficiency stems from the substantial diversity found in pose samples. On one hand, simple poses can be accurately estimated without requiring the computational resources of larger models. On the other hand, a more prominent issue arises from the abundance of bounding boxes, which remain excessive even after NMS. In this paper, we present a straightforward yet effective dynamic framework called DynPose, designed to match diverse pose samples with the most appropriate models, thereby ensuring optimal performance and high efficiency. Specifically, the framework contains a lightweight router and two pre-trained HPE models: one small and one large. The router is optimized to classify samples and dynamically determine the appropriate inference paths. Extensive experiments demonstrate the effectiveness of the framework. For example, using ResNet-50 and HRNet-W32 as the pretrained models, our DynPose achieves an almost 50% increase in speed over HRNet-W32 while maintaining the same-level accuracy. More importantly, the framework can be generalized to other pre-trained models and datasets without re-training or fine-tuning. Code is available at https://github.com/Aritoria/DynPose.
Yalong Xu, Lin Zhao 0003, Chen Gong 0002, Di Wang 0011, Nannan Wang 0001
CVPR6
2025 Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
abstract
Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data through data augmentation or image generation. Given the rapid advancements in AI-generated content (AIGC), this paper is the first to propose leveraging powerful pre-trained text-to-image (T2I) foundation models to create the training data. However, manually designing textual prompts to generate images for all possible domains is often impractical, and some domain characteristics may be too abstract to describe with words. To address these challenges, we propose a novel Progressive Adversarial Prompt Tuning (PAPT) framework for pre-trained diffusion models. Instead of relying on static textual domains, our approach learns two sets of abstract prompts as conditions for the diffusion model: one that captures domain-invariant category information and another that models domain-specific styles. This adversarial learning mechanism enables the T2I model to generate images in various domain styles while preserving key categorical features. Extensive experiments demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art single-domain generalization approaches.
De Cheng, Xinyang Jiang, Nannan Wang 0001, Dongsheng Li 0002, Xinbo Gao 0001
CVPR4
2025 ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learning
abstract
Can we accurately identify the true correspondences from multimodal datasets containing mismatched data pairs? Existing methods primarily emphasize the similarity matching between the representations of objects across modalities, potentially neglecting the crucial relation consistency within modalities that are particularly important for distinguishing the true and false correspondences. Such an omission often runs the risk of misidentifying negatives as positives, thus leading to unanticipated performance degradation. To address this problem, we propose a general Relation Consistency learning framework, namely ReCon, to accurately discriminate the true correspondences among the multimodal data and thus effectively mitigate the adverse impact caused by mismatches. Specifically, ReCon leverages a novel relation consistency learning to ensure the dual-alignment, respectively of, the cross-modal relation consistency between different modalities and the intra-modal relation consistency within modalities. Thanks to such dual constrains on relations, ReCon significantly enhances its effectiveness for true correspondence discrimination and therefore reliably filters out the mismatched pairs to mitigate the risks of wrong supervisions. Extensive experiments on three widely-used benchmark datasets, including Flickr30K, MS-COCO, and Conceptual Captions, are conducted to demonstrate the effectiveness and superiority of ReCon compared with other SOTAs. The code is available at: https://github.com/qxzha/ReCon.
Quanxing Zha, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001
CVPR6
2025 풟ℐℋ-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic Segmentation
Songsong Duan, Xi Yang 0011, Nannan Wang 0001
ICCV3
2025 Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object Detection
De Cheng, Xi Yang 0011, Nannan Wang 0001
ICCV4
2025 Q-Norm: Robust Representation Learning via Quality-Adaptive Normalization
Lanning Zhang, Fei Gao 0006, Ziyun Li 0002, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
ICCV7
2025 Mixture-of-Modality-Experts for Unified Image Aesthetic Assessment with Multi-Level Adaptation
abstract
Multi-modal image aesthetic assessment (MIAA) has gained significant progress, by predicting aesthetic based on both an image and its text comments. However, most MIAA methods are not applicable, when there are no text comments available. To combat this challenge, we propose a unified image aesthetic assessment (IAA) framework, termed AesFormer, by using mixtures of vision-language Transformers. Specially, AesFormer first learns aligned image-text representations through contrastive learning, and uses a vision-language head for MIAA prediction. Afterward, we propose a multi-level adaptation (MLA) method to adapt the learned MIAA model to the case without text comments, and use another vision head for vison-only IAA (VIAA) prediction. Extensive experimental results show that AesFormer significantly outperforms previous methods in both MIAA and VIAA tasks, on diverse benchmarking datasets. Our code has been released at: https://github.com/AiArt-Gao/AesFormer
Fei Gao 0006, Xiaodan Zhang 0005, Lihuo He, Nannan Wang 0001
ICME6
2025 ReCLIP: Reconstruction-Refined Zero-/Few-Shot Anomaly Classification and Segmentation
abstract
Recent advancements in zero-/few-shot anomaly detection have demonstrated the efficiency of contrastive learning approaches. However, existing methods struggle with imprecise perception of anomaly details and often lack focus on identifying anomaly types. To address these challenges, we propose a flexible reconstruction-refined framework based on contrastive learning, which comprises three core components: a cross-modal alignment network, a reconstruction module, and a dual-attention refinement module. The reconstruction module captures fine-grained anomaly embeddings and fidelity scores, enabling flexible module switching based on task requirements. The attention module directs the cross-modal alignment network to focus on fine-grained anomaly information for accurate segmentation. Our framework leverages the strengths of current anomaly detection algorithms, and excels in zero-/few-shot tasks across industrial datasets, recognizing 43 anomaly types without separate training for each category, and significantly outperforms existing methods, particularly on the challenging MPDD dataset. Our code has been released at https://github.com/AiArt-Gao/ReCLIP.
Lanning Zhang, Yali Shi, Shujie Lan, Fei Gao 0006, Hao Qin 0001, Nannan Wang 0001
ICME6
2025 Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language Models
abstract
Prompt learning is a cutting-edge parameter-efficient fine-tuning technique for pre-trained vision-language models (VLMs). Instead of learning a single text prompt, recent works have revealed that learning diverse text prompts can effectively boost the performances on downstream tasks, as the diverse prompted text features can comprehensively depict the visual concepts from different perspectives. However, diverse prompt learning demands enormous computational resources. This efficiency issue still remains unexplored. To achieve efficient and diverse prompt learning, this paper proposes a novel Surrogate Prompt Learning (SurPL) framework. Instead of learning diverse text prompts, SurPL directly generates the desired prompted text features via a lightweight Surrogate Feature Generator (SFG), thereby avoiding the complex gradient computation procedure of conventional diverse prompt learning. Concretely, based on a basic prompted text feature, SFG can directly and efficiently generate diverse prompted features according to different pre-defined conditional signals. Extensive experiments indicate the effectiveness of the surrogate prompted text features, and show compelling performances and efficiency of SurPL on various benchmarks.
Liangchen Liu 0001, Nannan Wang 0001, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
ICML2
2025 Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts
abstract
Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformers with Mixture-of-Experts to exploit both temporarily adaptability and spatial flexibility. Our design incorporates expert-specific timestep conditioning, allowing each expert to process different spatial tokens while adapting to the generative stage, to dynamically allocate resources based on both the temporal and spatial characteristics of the generative task. Additionally, we propose a globally-aware feature recalibration mechanism that amplifies the representational capacity of expert modules by dynamically adjusting feature contributions based on input relevance. Extensive experiments on image generation benchmarks demonstrate that Diff-MoE significantly outperforms state-of-the-art methods. Our work demonstrates the potential of integrating diffusion models with expert-based designs, offering a scalable and effective framework for advanced generative modeling.
Xiao He 0014, Zhijun Tu, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Hu 0021
ICML6
2025 Phase and Amplitude-aware Prompting for Enhancing Adversarial Robustness
abstract
Deep neural networks are found to be vulnerable to adversarial perturbations. The prompt-based defense has been increasingly studied due to its high efficiency. However, existing prompt-based defenses mainly exploited mixed prompt patterns, where critical patterns closely related to object semantics lack sufficient focus. The phase and amplitude spectra have been proven to be highly related to specific semantic patterns and crucial for robustness. To this end, in this paper, we propose a Phase and Amplitude-aware Prompting (PAP) defense. Specifically, we construct phase-level and amplitude-level prompts for each class, and adjust weights for prompting according to the model’s robust performance under these prompts during training. During testing, we select prompts for each image using its predicted label to obtain the prompted image, which is inputted to the model to get the final prediction. Experimental results demonstrate the effectiveness of our method.
Dawei Zhou 0004, Decheng Liu, Nannan Wang 0001
ICML4
2025 Towards Regularized Mixture of Predictions for Class-Imbalanced Semi-Supervised Facial Expression Recognition
abstract
Semi-supervised facial expression recognition (SSFER) effectively assigns pseudo-labels to confident unlabeled samples when only limited emotional annotations are available. Existing SSFER methods are typically built upon an assumption of the class-balanced distribution. However, they are far from real-world applications due to biased pseudo-labels caused by class imbalance. To alleviate this issue, we propose Regularized Mixture of Predictions (ReMoP), a simple yet effective method to generate high-quality pseudo-labels for imbalanced samples. Specifically, we first integrate feature similarity into the linear prediction to learn a mixture of predictions. Furthermore, we introduce a class regularization term that constrains the feature geometry to mitigate imbalance bias. Being practically simple, our method can be integrated with existing semi-supervised learning and SSFER methods to tackle the challenge associated with class-imbalanced SSFER effectively. Extensive experiments on four facial expression datasets demonstrate the effectiveness of the proposed method across various imbalanced conditions. The source code is made publicly available at https://github.com/hangyu94/ReMoP.
Hangyu Li 0001, Jiangchao Yao, Nannan Wang 0001, Bo Han 0003
IJCAI4
2025 CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable Quantization
abstract
Face cartoonization remains a challenging task due to significant geometric deformations between facial photos and cartoons, as well as the absence of paired training data for supervised learning. Existing methods struggle to generate high-quality cartoonized avatars with attribute consistency. To address this challenge, this paper proposes an unsupervised facial cartoonization method based on cross-domain aligned and deformable vector quantization (CADQ). Firstly, we construct textual descriptions with facial attributes for both photo datasets and cartoon collections. Attribute consistency during transformation is enforced through individually contrastive learning between image-text cross-modal features and globally distribution alignment across photo-cartoon domains. Secondly, a deformable Transformer with dual attention is introduced during the transformation process, which queries corresponding cartoon codebook entries based on image features to simulate cross-domain geometric deformations. Experimental results demonstrate that the proposed method can convert facial photos into high-quality cartoons with attribute consistency, outperforming existing state-of-the-art approaches. Furthermore, the method can be effectively extended to unsupervised cross-domain generation of other artistic portrait styles, achieving superior or highly competitive performance. Our code has been released at: https://github.com/IIP-Lab-XDU/CADQ.
Yongjie Hu, Ziyun Li 0002, Fei Gao 0006, Henrik Boström, Nannan Wang 0001
ACM Multimedia6
2025 Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID
De Cheng, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001
Int. J. Comput. Vis.3
2025 Exploring Homogeneous and Heterogeneous Consistent Label Associations for Unsupervised Visible-Infrared Person ReID
De Cheng, Nannan Wang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.3
2025 Defending Against Adversarial Examples Via Modeling Adversarial Noise
Dawei Zhou 0004, Nannan Wang 0001, Bo Han 0003, Tongliang Liu, Xinbo Gao 0001
Int. J. Comput. Vis.2
2025 Fooling human detectors via robust and visually natural adversarial patches
Dawei Zhou 0004, Hongbin Qu, Nannan Wang 0001, Chunlei Peng, Zhuoqi Ma, Xi Yang 0011, Xinbo Gao 0001
Neurocomputing3
2025 Revisiting face forgery detection towards generalization
Chunlei Peng, Decheng Liu, Huiqing Guo, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks5
2025 FairForensics: mitigating attribute bias in deepfake detection by integrating texture and attribute features
Chunlei Peng, Yinyin Chen, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks4
2025 Frequency-Based Comprehensive Prompt Learning for Vision-Language Models
abstract
This paper targets to learn multiple comprehensive text prompts that can describe the visual concepts from coarse to fine, thereby endowing pre-trained VLMs with better transfer ability to various downstream tasks. We focus on exploring this idea on transformer-based VLMs since this kind of architecture achieves more compelling performances than CNN-based ones. Unfortunately, unlike CNNs, the transformer-based visual encoder of pre-trained VLMs cannot naturally provide discriminative and representative local visual information. To solve this problem, we propose Frequency-based Comprehensive Prompt Learning (FCPrompt) to excavate representative local visual information from the redundant output features of the visual encoder. FCPrompt transforms these features into frequency domain via Discrete Cosine Transform (DCT). Taking the advantages of energy concentration and information orthogonality of DCT, we can obtain compact, informative and disentangled local visual information by leveraging specific frequency components of the transformed frequency features. To better fit with transformer architectures, FCPrompt further adopts and optimizes different text prompts to respectively align with the global and frequency-based local visual information via a dual-branch framework. Finally, the learned text prompts can thus describe the entire visual concepts from coarse to fine comprehensively. Extensive experiments indicate that FCPrompt achieves the state-of-the-art performances on various benchmarks.
Liangchen Liu 0001, Nannan Wang 0001, Chen Chen 0128, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Unknown-Aware Bilateral Dependency Optimization for Defending Against Model Inversion Attacks
abstract
By abusing access to a well-trained classifier, model inversion (MI) attacks pose a significant threat as they can recover the original training data, leading to privacy leakage. Previous studies mitigated MI attacks by imposing regularization to reduce the dependency between input features and outputs during classifier training, a strategy known as unilateral dependency optimization. However, this strategy contradicts the objective of minimizing the supervised classification loss, which inherently seeks to maximize the dependency between input features and outputs. Consequently, there is a trade-off between improving the model's robustness against MI attacks and maintaining its classification performance. To address this issue, we propose the bilateral dependency optimization strategy (BiDO), a dual-objective approach that minimizes the dependency between input features and latent representations, while simultaneously maximizing the dependency between latent representations and labels. BiDO is remarkable for its privacy-preserving capabilities. However, models trained with BiDO exhibit diminished capabilities in out-of-distribution (OOD) detection compared to models trained with standard classification supervision. Given the open-world nature of deep learning systems, this limitation could lead to significant security risks, as encountering OOD inputs-whose label spaces do not overlap with the in-distribution (ID) data used during training-is inevitable. To address this, we leverage readily available auxiliary OOD data to enhance the OOD detection performance of models trained with BiDO. This leads to the introduction of an upgraded framework, unknown-aware BiDO (BiDO+), which mitigates both privacy and security concerns. As a highlight, with comparable model utility, BiDO-HSIC+ reduces the FPR95 by 55.02% and enhances the AUCROC by 9.52% compared to BiDO-HSIC, while also providing superior MI robustness.
Xiong Peng, Feng Liu 0003, Nannan Wang 0001, Long Lan, Tongliang Liu, Yiu-Ming Cheung, Bo Han 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Improving Adversarial Training From the Perspective of Class-Flipping Distribution
abstract
Adversarial training has been proposed and widely recognized as a very effective method to defend against adversarial noise. However, the label flipping pattern on different classes still need deeper exploration to identify potential problems and assist in further enhancing robustness. In this work, we model the class-flipping distribution via statistical investigations and find this distribution reveals two shortcomings: the highly misleading category is present in the model's predictions for data in each class, and the trend in class flipping are significantly different across classes. Based on these observations, we propose a Class-Flipping-aware Adversarial Training (CFAT) method. On the one hand, we obtain the most misleading categories for the data in each class by counting the samples flipped to different wrong categories, and utilize them as the target to construct corresponding targeted adversarial samples, respectively. On the other hand, we take the proportions of samples flipped to the most misleading category as factors to scale the perturbation budgets of adversarial training samples for the data with corresponding classes. Experimental results on datasets with different class number validate the effectiveness of the proposed method.
Dawei Zhou 0004, Nannan Wang 0001, Tongliang Liu, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Consistency-driven feature scoring and regularization network for visible-infrared person re-identification
Xueting Chen, Yan Yan 0001, Jing-Hao Xue, Nannan Wang 0001, Hanzi Wang
Pattern Recognit.4
2025 Bidirectional modality information interaction for Visible-Infrared Person Re-identification
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.3
2025 Associative graph convolution network for point cloud analysis
Xi Yang 0011, Xingyilang Yin, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.3
2025 iFADIT: Invertible Face Anonymization via Disentangled Identity Transform
Lin Yuan 0002, Tao Wu 0003, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.5
2025 Achieving Plasticity-Stability Trade-Off in Continual Learning Through Adaptive Orthogonal Projection
abstract
Catastrophic forgetting is the crucial challenge for continual learning. One of the state-of-the-art approaches is the orthogonal projection, which aims to learn each task by updating model parameters in the direction orthogonal to the subspace spanned by the previous task input. Although such strict orthogonal weight constraints ensure no interference with tasks that have been learned to achieve model stability, they greatly sacrifice model plasticity. In this paper, we propose an adaptive balanced orthogonal projection (AdaBOP) method, to search for the optimal network parameter updating direction to address the plasticity-stability dilemma in continual learning. The proposed AdaBOP method can adaptively adjust its tendency towards plasticity-stability trade-off based on the layer-wise feature space correlations of the model between old and new tasks. To further improve the training efficiency, we also implement the AdaBOP method in the uncentered covariance matrix space of the previous tasks, and finally achieve a better stability-plasticity trade-off in continual learning efficiently. Experimental results greatly demonstrate the effectiveness of the proposed method, which achieves superior performances to state-of-the-art continual learning approaches. The code is available athttps://github.com/hyscn/AdaBOP.
De Cheng, Yusong Hu, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Progressive Feature-Attribute Matching via Bi-Directional Generation for Transductive Zero-Shot Learning
abstract
Transductive zero-shot learning (TZSL) has been proposed to address the domain shift problem by leveraging additional unlabeled unseen data to enhance the generalization ability from seen classes to unseen target classes. Existing TZSL methods primarily focus on mitigating the distribution bias problem by incorporating these unlabeled samples into the generative models. Although these methods have achieved great success, they do not fully exploit the potential of these unlabeled target data. In this paper, we propose a bidirectional weakly guided conditional generative modeling approach, which utilizes the attribute regressor and the visual generator to synthesize paired training data of unseen classes for each other, thus converting unlabeled target data into matched feature-attribute pairs. Additionally, on top of the generative modeling, we also propose to progressively estimate the associations between visual features and attributes among the unlabeled target data through a semi-supervised pseudo-labeling approach, so as to further facilitate the generative model and enhance the learning of target distributions. Extensive experimental results on four benchmark datasets demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods. Our source code is released in https://github.com/LevisWei/semi-zero-master.
De Cheng, Chaowei Fang, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Unsupervised Face Super-Resolution via Integrating Faithful 3D Facial Priors
abstract
Recently, unsupervised face super-resolution (FSR) has attracted significant attention due to its remarkable generalization performance. However, existing methods neglect the incorporation of facial priors, which can effectively guide the restoration of face images. The root cause of this issue lies in the significant challenges associated with incorporating facial priors into unsupervised frameworks. First, unsupervised methods often face the challenge of real-world low-quality (LQ) images that are severely corrupted, making it unrealistic to extract reliable prior information from them. Second, the estimation of facial priors exponentially increases the model’s parameters and computational complexity, contradicting the purpose of unsupervised methods for practical deployment. In this work, we fundamentally address the aforementioned challenges and proposeFaith3D-FSR, a novel approach that incorporates faithful 3D facial priors into unsupervised FSR. Specifically, we introduceFaith3Dmechanism for faithful prior integration, which deconstructs super-resolution images into 3D elements and uses the 3D priors from real high-quality (HQ) images as reference for calibration solely during the training phase. This strategy enables more precise guidance on the super-resolution in a high-dimensional space, without requiring additional prior estimation during inference. It successfully overcomes the aforementioned challenges, making it more suitable for real-world applications, and offers a plug-and-play solution for incorporating 3D priors into unsupervised FSR. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on multiple benchmark datasets and across a range of evaluation metrics. The code is available here.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Perspectives of Calibrated Adaptation for Few-Shot Cross-Domain Classification
abstract
Current few-shot learning techniques predominantly leverage amortization techniques based on meta-learning frameworks, which effectively adapt to unknown tasks with limited examples. However, these approaches face significant challenges in cross-domain scenarios, where the data distributions between the source domain (training data) and the target domain (testing data) differ substantially. This domain shift can lead to models that overfit the global discriminative model while underfitting their local amortization on the adaptable few-shot structure. To mitigate this problem, our proposal makes an upgrade on Conditional Neural Adaptive Processes, reformulating its conditioning mechanism to better handle cross-domain adaptation. This results in calibrated amortization of task-specific feature extractors and the construction of a robust non-parametric classifier. In our implementation, we first employ generative modeling or deterministic self-attention to all labeled context features, establishing a strong task-level alignment that adapts the extractor across domains. Additionally, we introduce a novel channel-wise normalization to further enhance the adaptation process. Our experiments on the Meta-dataset benchmark demonstrate an average$6.9\sim 9$% improvement in out-of-distribution tasks, underscoring the effectiveness of exploiting calibrated adaptation in few-shot cross-domain classification.
Dechen Kong, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Devil in Shadow: Attacking NIR-VIS Heterogeneous Face Recognition via Adversarial Shadow
abstract
Near infrared-visible (NIR-VIS) heterogeneous face recognition aims to match face identities in cross-modality settings, which has achieved significant development recently. The work on adversarial attack and security issues of the heterogeneous face recognition task is still lacking. Existing adversarial face generation methods can’t deploy directly because of the inevitable large modality discrepancy. Besides, the ideal adversarial attacking generated images should maintain both high capabilities and low detectability. Considering the properties of near-infrared face images, our basic idea is to construct adversarial shadows for good stealthiness and high attack capability. In this paper, we propose a novel face adversarial shadow generation framework for NIR-VIS heterogeneous face recognition, which can synthesize fine-crafted lighting conditions containing strong identity attacking ability. Specifically, we design the variance consistency-based symmetric face attacking loss to improve the attacking generalization and the synthesized image quality. Extensive qualitative and quantitative experiments on the public large-scale NIR-VIS heterogeneous face dataset prove the proposed method achieves superior performance compared with the state-of-the-art methods. The source code is publicly available athttps://github.com/GEaMU/Devil-in-Shadow.
Decheng Liu, Rong Sheng, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Boosting Semi-Supervised Facial Attribute Recognition With Dynamic Threshold Pairs
abstract
Semi-supervised learning (SSL) has proven effective in assigning a pseudo-label to a confident sample whose largest class probability is above a fixed threshold. However, in the context of semi-supervised facial attribute recognition (SSFAR), where a sample is associated with multiple presence and absence pseudo-labels, directly applying existing SSL methods is challenging due to two issues: 1) the lack of a clear boundary between presence and absence predictions for an attribute makes it difficult to distinguish them using a single threshold; 2) the learning difficulty varies across attributes, so the fixed strategy fails to adaptively learn different attributes. To address these challenges, we propose Dynamic thrEShold Pairs (DESP), a simple yet effective method to handle the SSFAR problem. Specifically, during each training stage, we derive two sets for each attribute from labeled samples, which contain the predicted probabilities of presence and absence, respectively. We then compute the mid-ranges of the two sets as paired presence and absence thresholds. Finally, we assign a presence or absence pseudo-label for the attribute to an unlabeled sample when its prediction exceeds the presence threshold or falls below the absence threshold. Extensive experiments on the CelebA and LFWA datasets demonstrate that DESP achieves superior performance compared to state-of-the-art methods, especially in the case of scarce labeled samples. Also, DESP performs well on multi-label datasets such as Pascal VOC and MS-COCO. The code will be publicly available athttps://github.com/yihanxxu/DESP.
Hangyu Li 0001, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 CatVersion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization
abstract
We propose CatVersion, an inversion-based method that learns the personalized concept through a handful of examples. Subsequently, users can utilize text prompts to generate images that embody the personalized concept, thereby achieving text-to-image personalization. In contrast to existing approaches that emphasize word embedding learning or parameter fine-tuning for the diffusion model, which potentially causes concept dilution or overfitting, our method concatenates embeddings on the feature-dense space of the text encoder in the diffusion model to learn the gap between the personalized concept and its base class, aiming to maximize the preservation of prior knowledge in diffusion models while restoring the personalized concepts. To this end, we first dissect the text encoder’s integration in the image generation process to identify the feature-dense space of the encoder. Afterward, we concatenate embeddings on the Keys and Values in this space to learn the gap between the personalized concept and its base class. In this way, the concatenated embeddings ultimately manifest as a residual on the original attention output. To more accurately and unbiasedly quantify the results of personalized image generation, we improve the CLIP image alignment score based on masks. Qualitatively and quantitatively, CatVersion helps to restore personalization concepts more faithfully and enables more robust editing.
Mingrui Zhu, Shiyin Dong, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 A Knowledge-Guided Adversarial Defense for Resisting Malicious Visual Manipulation
abstract
Malicious applications of visual manipulation have raised serious threats to the security and reputation of users in many fields. To alleviate these issues, adversarial noise-based defenses have been enthusiastically studied in recent years. However, “data-only” methods tend to distort fake samples in the low-level feature space rather than the high-level semantic space, leading to limitations in resisting malicious manipulation. Frontier research has shown that integrating knowledge in deep learning can produce reliable and generalizable solutions. Inspired by these, we propose aknowledge-guided adversarial defense (KGAD) to actively force malicious manipulation models to output semantically confusing samples. Specifically, in the process of generating protective adversarial noise, we focus on constructing significant semantic confusions at the domain-specific knowledge level, and exploit a metric closely related to visual perception to replace the general pixel-wise metrics. The generated adversarial noise can actively interfere with the malicious manipulation model by triggering knowledge-guided and perception-related disruptions in the fake samples. To validate the effectiveness of the proposed method, we conduct qualitative and quantitative experiments on human perception and visual quality assessment. The results on two different tasks both show our defense achieves competitive performances and generalizability, indicating that it can effectively resist malicious visual manipulation.
Dawei Zhou 0004, Zhigang Su, Decheng Liu, Tongliang Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Dependable Secur. Comput.5
2025 Masked Text Adversarial Training for Cloth-Changing Person Re-Identification
Chengrui Hao, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Attention Consistency Refined Masked Frequency Forgery Representation for Generalizing Face Forgery Detection
abstract
Due to the successful development of deep image generation technology, visual data forgery detection would play a more important role in social and economic security. Existing forgery detection methods suffer from unsatisfactory generalization ability to determine the authenticity in the unseen domain. In this paper, we propose a novel Attention Consistency Refined masked frequency forgery representation model toward a generalizing face forgery detection algorithm (ACMF). Most forgery technologies always bring in high-frequency aware cues, which make it easy to distinguish source authenticity but difficult to generalize to unseen artifact types. The masked frequency forgery representation module is designed to explore robust forgery cues by randomly discarding high-frequency information. In addition, we find that the forgery saliency map inconsistency through the detection network could affect the generalizability. Thus, the forgery attention consistency is introduced to force detectors to focus on similar attention regions for better generalization ability. Experiment results on several public face forgery datasets (FaceForensic++, DFD, Celeb-DF, WDF and DFDC datasets) demonstrate the superior performance of the proposed method compared with the state-of-the-art methods. The source code and models are publicly available athttps://github.com/chenboluo/ACMF.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Improving Adversarial Robustness via Decoupled Visual Representation Masking
abstract
Deep neural networks are proven to be vulnerable to finely designed adversarial examples, and adversarial defense algorithms draw more and more attention nowadays. Pre-processing based defense is a major strategy, as well as learning robust feature representation, has been proven an effective way to boost generalization. However, existing defense works lack considering different depth-level visual features in the training process. In this paper, we first highlight two novel properties of robust features from the feature distribution perspective: 1) Diversity (robust features within the same class should maintain appropriate variety). 2) Discriminability (robust features from different classes should be sufficiently separated). We find that state-of-the-art defense methods aim to address both of these mentioned issues well. It motivates us to increase intra-class variance and decrease inter-class discrepancy simultaneously in adversarial training. Specifically, we propose a simple but effective defense based on decoupled visual representation masking. The designed Decoupled Visual Feature Masking (DFM) block can adaptively disentangle visual discriminative features and non-visual features with diverse mask strategies, while the suitable discarding information can disrupt adversarial noise to improve robustness. Our work provides a generic and easy-to-plugin block unit for any former adversarial training algorithm to achieve better protection integrally. Extensive experimental results prove that the proposed method can achieve superior performance compared with state-of-the-art defense approaches. The code is publicly available at https://github.com/chenboluo/Adversarial-defense.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Toward Fair Adversarial Defense via Class Encourage-Suppress Robust Learning
abstract
Deep Neural Networks with natural training are quite vulnerable to adversarial attacks, so it’s necessary to defend these attacks with effective defense methods like adversarial training. However, while defending against adversarial attacks, adversarially trained models’ robustness between classes shows severe unfairness. To mitigate the disparity, a lot of methods have been proposed, while many of them sacrifice the overall accuracy to leverage the worst-class accuracy. Inspired by previous works, we propose a new fairness methodology, and we name it Class Encourage-suppress Robust Learning (CRL). Based on the overall accuracy and the class-wise accuracies of the dataset, we introduce a new measurement named Class Diversity Ratio to adjust the weights of different classes in the loss function. Additionally, we propose a new learning strategy called the Competitor Encourage-suppress Strategy to mitigate the disparity between diverse classes, which is simple but effective. Experimental results on representative datasets show that our method outperforms state-of-the-art (SOTA) methods.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Semantic Token Transformer for Face Forgery Detection
abstract
In the era of digital media, the proliferation of forged images and videos poses a significant threat to societal stability. With the rapid advancement of deep learning, the generation of realistic fake images has become increasingly simple, presenting unprecedented challenges in discerning the authenticity of images. While some existing methods have shown promising results in forgery detection, they often underutilize facial semantic information. To address this issue, this paper introduces the Semantic Token Transformer for Face Forgery Detection. By incorporating facial semantic information with a transformer network, the input tokens of the transformer are transformed into tokens of varying shapes and sizes based on their importance, thereby enhancing the accuracy of the detector. To achieve this objective, we first employ an image processing stage to manipulate the image based on facial semantic information. Subsequently, we introduce a scoring network, guided by prior knowledge, which adaptively categorizes tokens into different clusters based on their importance and relevance to the results of the preprocessing stage. Finally, we merge the tokens within the clusters using an attention mechanism and input them into the detector for forgery detection. Through experiments conducted on multiple datasets and cross-dataset evaluations, we demonstrate that our approach outperforms state-of-the-art detection methods.
Chunlei Peng, Xiaoyi Luo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Within 3DMM Space: Exploring Inherent 3D Artifact for Video Forgery Detection
abstract
Recently, the breathtaking development and potential misuse of deepfake technology has raised numerous privacy and security concerns, triggering widespread apprehension. Existing deepfake detection methods focus on the analysis of local regions for faces, such as mouth movement, eye blinking frequency, etc., which, however, are limited in their ability to capture the global inconsistencies present in forged faces. Some researchers attempt to seize 3D artifacts related to facial global information, but typically treat the 3D information as mere input, lacking the in-depth analysis. To address these shortcomings and mine the inherent and delicate 3D artifacts in the forged faces, this paper innovatively proposes the 3D Artifact Detector (3DAD) method, which leverages the spatio-temporal inconsistency on the 3D semantic space in the forgery videos to uncover the deepfake clues. Specifically, we employ 3D Analysis Unit (3DAU) to pre-train the face reconstruction task within 3D Morphable Model (3DMM) space, thereby obtaining the high-level inherent 3d representation. Concurrently, for the multi-levels of information in the face, we utilize the Texture Perception Unit (TPU) to extract the texture information in the low-level semantic space of the images. Ultimately we feed the two distinct modalities into the spatiotemporal fusion model for final detection. Through extensive intra- and cross-dataset experiments on publicly available datasets, we demonstrate the effectiveness and generalizability of the proposed method. The source code is available at https://github.com/Cookie-XT/3DAD.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Granularity-Aware Hyperbolic Representation for Text-Based Person Search
abstract
Text-based person search aims to identify specific target person from the database according to the given text description. Early work adopted separately pretrained encoders to extract visual and textual features, but benefit from the bloom of visual language pre-training, recent work uses unified pretrained visual language models such as CLIP as backbone. However, visual language models are generally pretrained from coarse-grained image-text pairs, while image-text pairs in text-based person search are more fine-grained to distinguish different persons. In addition, visual and linguistic concepts naturally organize themselves in a hierarchy, which is not explicitly captured by current large-scale vision and language models such as CLIP. To bridge this gap, we propose a novel Granularity-Aware Hyperbolic Representation learning method for mining granularity and capturing semantic hierarchy. Notably, we consider both token-level and instance-level granularity. For token-granularity alignment, we present a Bidirectional Attention Interaction module to explicitly learn the matching between fine-grained visual tokens and text tokens. For instance-granularity alignment, we equip the contrastive learning loss with Semantic Margin Softmax so that image-text pairs can perceive the similarity granularity of different samples during training. Besides, the global features of images and texts are mapped into hyperbolic space through Hyperbolic Representation Learning to embed tree-like data to capture semantic hierarchy. Extensive experiments verify the effectiveness of our proposed modules and show that our method achieves state-of-the-art results on the three widely acknowledged benchmarks, namely CUHK-PEDES, ICFG-PEDES, and RST-PReID. Our code is available at https://github.com/7chQ/GAHR.
Chenghuan Qi, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2025 Disentangle Before Anonymize: A Two-Stage Framework for Attribute-Preserved and Occlusion-Robust De-Identification
abstract
In an era where personal photos are easily leaked and collected, face de-identification is a crucial method for protecting identity privacy. However, current face de-identification techniques face challenges in preserving attribute details and often produce anonymized results with reduced realistic. These shortcomings are particularly evident when handling occlusions, frequently resulting in noticeable editing artifacts. Our primary finding in this work is that simultaneous training of identity disentanglement and anonymization hinders their respective effectiveness. Therefore, we propose “Disentangle Before Anonymize”, a novel two-stage Framework (DBAF) designed for attribute-preserved and occlusion-robust de-identification. This framework includes a Contrastive Identity Disentanglement (CID) module and a Key-authorized Reversible Identity Anonymization (KRIA) module, achieving faithful attribute preservation and high-quality identity anonymization edits. Additionally, we introduce a Multi-scale Attentional Attribute Retention (MAAR) module to address the issue of reduced anonymization quality under occlusions. Extensive experiments demonstrate that our method outperforms state-of-the-art de-identification approaches, delivering superior quality, enhanced detail fidelity, improved attribute preservation performance, and greater robustness to occlusions. The code ispublicly available at: https://github.com/mrzhu-cool/DBAF.
Mingrui Zhu, Dongxin Chen, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Escaping Modal Interactions: An Efficient DESANet for Multi-Modal Object Re-Identification
abstract
Multi-modal object Re-ID aims to leverage the complementary information provided by multiple modalities to overcome challenging conditions and achieve high-quality object matching. However, existing multi-modal methods typically rely on various modality interaction modules for information fusion, which can reduce the efficiency of real-time monitoring systems. Additionally, practical challenges such as low-quality multi-modal data or missing modalities further complicate the application of object Re-ID. To address these issues, we propose the Complementary Data Enhancement and Modal-Aware Soft Alignment Network (DESANet), which is designed to be independent of interactive networks and adaptable to scenarios with missing modalities. This approach ensures a simple-yet-effective, and efficient multi-modal object Re-ID. DESANet consists of three key components: Firstly, the Dual-Color Space Data Enhancement (DCDE) module, which enhances multi-modal data by performing patch rotation in the RGB space and improving image quality in the HSV space. Secondly, the Salient Feature ReConstruction (SFRC) module, which addresses the issue of missing modalities by reconstructing features from one modality using the other two. Thirdly, the Modal-Aware Soft Alignment (MASA) module, which integrates multi-source data to avoid the blind fusion of features and prevents the propagation of noise from reconstructed modalities. Our approach achieves state-of-the-art performances on both person and vehicle datasets. Source code is available at https://github.com/DWJ11/DESANet.
Wenjiao Dong, Xi Yang 0011, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 Lightweight RGB-D Salient Object Detection From a Speed-Accuracy Tradeoff Perspective
abstract
Current RGB-D methods usually leverage large-scale backbones to improve accuracy but sacrifice efficiency. Meanwhile, several existing lightweight methods are difficult to achieve high-precision performance. To balance the efficiency and performance, we propose a Speed-Accuracy Tradeoff Network (SATNet) for Lightweight RGB-D SOD from three fundamental perspectives: depth quality, modality fusion, and feature representation. Concerning depth quality, we introduce the Depth Anything Model to generate high-quality depth maps,which effectively alleviates the multi-modal gaps in the current datasets. For modality fusion, we propose a Decoupled Attention Module (DAM) to explore the consistency within and between modalities. Here, the multi-modal features are decoupled into dual-view feature vectors to project discriminable information of feature maps. For feature representation, we develop a Dual Information Representation Module (DIRM) with a bi-directional inverted framework to enlarge the limited feature space generated by the lightweight backbones. DIRM models texture features and saliency features to enrich feature space, and employ two-way prediction heads to optimal its parameters through a bi-directional backpropagation. Finally, we design a Dual Feature Aggregation Module (DFAM) in the decoder to aggregate texture and saliency features. Extensive experiments on five public RGB-D SOD datasets indicate that the proposed SATNet excels state-of-the-art (SOTA) CNN-based heavyweight models and achieves a lightweight framework with 5.2 M parameters and 415 FPS. The code is available at https://github.com/duan-song/SATNet.
Songsong Duan, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 PrivacyHFR: Visual Privacy Preserving for Heterogeneous Face Recognition
abstract
Face recognition has achieved remarkable progress and is widely deployed in real-world scenarios. Recently more and more attention has been given to individual privacy protection, due to unauthorized sensitive image leakage by malicious attackers. Multi-modality face images captured by diverse sensors, also called heterogeneous faces, bring in more challenges in face privacy protection while lacking related research. In this paper, we propose a novel visual Privacy preserving method for Heterogeneous Face Recognition (Privacy-HFR) to protect perceptual visual information and maintain essential identity information in multi-modality face analysis scenarios. Frequency domain analysis is a vital strategy to bridge the inevitable modality gap for heterogeneous face images. Meanwhile, recent theoretical insights also inspire us to design a suitable frequency component adjustment to balance human visual sensitivity and identity discriminative information. In addition, the ability to defend against recovery attacks has emerged as an essential criterion for privacy preserving face recognition. Noting that there seems to exist a dilemma that reducing accessible information by the attack model will affect the extracted identity information for recognition. It is because these two kinds of information are mutually blended in the frequency domain, which makes it a challenge to simultaneously maintain visual privacy and identity distinguishability. Thus, we provide a novel perspective to leverage the randomly optimal solutions and design the specific adversarial perturbations against the recovery attack. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ, Tufts Face and CUFSF datasets) prove that the proposed method outperforms existing privacy-preserving face recognition methods in terms of recognition accuracy and privacy protection capability. The code is available in https://github.com/xiyin11/Privacy-HFR.
Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 SketchAging: Face Photo-Sketch Synthesis and Aging With Multi-Scale Feature Extraction
abstract
With the rapid development of Generative Adversarial Networks (GANs), facial sketch generation and age transformation have advanced considerably. These technologies show great potential in digital media, entertainment, and forensic applications, particularly in helping law enforcement reconstruct the appearance of long-term fugitives. However, current methodologies exhibit notable limitations: existing approaches typically specialize in either facial sketch generation or age progression independently, lacking an effective integration for cross-domain synthesis. Moreover, preserving identity information while ensuring high-quality image generation remains a challenge. This paper proposes Multi-Scale Feature Extraction Networks (MSFE), an image-to-image translation framework that enables continuous age transformation while maintaining the stylistic characteristics of sketch domains. The core of the MSFS network uses a Dual Conditional Normalization Attention (DCNA) architecture to extract sketch features and encode facial images into the latent space of a pre-trained StyleGAN based on the desired age change. Experimental results on public datasets demonstrate that our approach outperforms existing methods, achieving superior facial photo-sketch synthesis with enhanced realism, identity preservation, and age accuracy.
Chunlei Peng, Zhuang Tang, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 Face Forgery Detection With CLIP-Enhanced Multi-Encoder Distillation
abstract
With the development of face forgery technology, fake faces are rampant, threatening the security and authenticity of many fields. Therefore, it is of great significance to study face forgery detection. At present, existing detection methods have deficiencies in the comprehensiveness of feature extraction and model adaptability, and it is difficult to accurately deal with complex and changeable forgery scenarios. However, the rise of multimodal models provides new insights for current forgery detection methods. At present, most methods use relatively simple text prompts to describe the difference between real and fake faces. However, these researchers ignore that the CLIP model itself does not have the relevant knowledge of forgery detection. Therefore, our paper proposes a face forgery detection method based on multi-encoder fusion and cross-modal knowledge distillation. On the one hand, the prior knowledge of the CLIP model and the forgery model is fused. On the other hand, through the alignment distillation, the student model can learn the visual abnormal patterns and semantic features of the forged samples captured by the teacher model. Specifically, our paper extracts the features of face photos by fusing the CLIP text encoder and the CLIP image encoder, and uses the dataset in the field of forgery detection to pretrain and fine-tune the Deepfake-V2-Model to enhance the detection ability, which are regarded as the teacher model. At the same time, the visual and language patterns of the teacher model are aligned with the visual patterns of the pretrained student model, and the aligned representations are refined to the student model. This not only combines the rich representation of the CLIP image encoder and the excellent generalization ability of text embedding, but also enables the original model to effectively acquire relevant knowledge for forgery detection. Experiments show that our method effectively improves the performance on face forgery detection.
Chunlei Peng, Tianzhe Yan, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 MVFusion: Generative Representation Learning With Masked Variational Autoencoders for Multi-Modality Image Fusion
abstract
Creating a comprehensively representative image while maintaining the merits of various modalities is a key focus of current Multi-Modality Image Fusion research. Existing unified methods often struggle to handle varying types of degradation while extracting modality-shared and modality-specific information from source images, leading to limitations in their generative or representation capabilities under different conditions. To address the challenge, we propose MVFusion, a novel self-supervised masked variational autoencoder framework that simultaneously enhances generative training and representation learning. It is designed to cope with varying image quality and dataset composition with a unified framework while ensuring effective fusion of modality information. Specifically, MVFusion employs a self-supervised masked autoencoder to reduce the impact of redundancy and degradation in the source images, and thus learns the latent distribution of degraded input images in the generative training stage. In addition, we incorporate variational feature learning to further preserve the distinctive modality features in the representation learning stage. Extensive experiments demonstrate that our model achieves promising results in several classical fusion tasks, including infrared-visible, multi-focus, multi-exposure, and medical image fusion. The code is available at https://github.com/shiboneng/MVFusion.
Jingwei Xin, Boneng Shi, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 FA-Net: A Feature Alignment Network for Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based visible-infrared person re-identification (VVI-ReID) aims to match target pedestrians between visible and infrared videos, which is significantly applied in 24-hour surveillance systems. The key of VVI-ReID is to learn modality invariant and spatio-temporal invariant sequence-level representation to solve the challenges such as modality differences, spatio-temporal misalignment, and domain shift noise. However, existing methods predominantly emphasize on reducing modality discrepancy while relatively neglect temporal misalignment and domain shift noise reduction. To this end, this paper proposes a VVI-ReID framework called Feature Alignment Network (FA-Net) from the perspective of feature alignment, aiming to mitigate temporal misalignment. FA-Net comprises two main alignment modules: Spatial-Temporal Alignment Module (STAM) and Modality Distribution Constraint (MDC). STAM integrates global and local features to ensure individuals' spatial representation alignment. Additionally, STAM also establishes temporal relationships by exploring inter-frame features to address cross-frame person feature matching. Furthermore, we introduce the Modality Distribution Constraint (MDC), which utilizes a symmetric distribution loss to align the distributions of features from different modalities. Besides, the SAM Guidance Augmentation (SAM-GA) strategy is designed to transform the image space of RGB and IR frames to provide more informative and less noisy frame information. Extensive experimental results demonstrate the effectiveness of the proposed method, surpassing existing state-of-the-art methods. Our code will be available at: https://github.com/code/FANet.
Xi Yang 0011, Wenjiao Dong, De Cheng, Nannan Wang 0001
IEEE Trans. Image Process.5
2025 IDENet: An Inter-Domain Equilibrium Network for Unsupervised Cross-Domain Person Re-Identification
abstract
Unsupervised person re-identification aims to retrieve a given pedestrian image from unlabeled data. For training on the unlabeled data, the method of clustering and assigning pseudo-labels has become mainstream, but the pseudo-labels themselves are noisy and will reduce the accuracy. To overcome this problem, several pseudo-label improvement methods have been proposed. But on the one hand, they only use target domain data for fine-tuning and do not make sufficient use of high-quality labeled data in the source domain. On the other hand, they ignore the critical fine-grained features of pedestrians and overfitting problems in the later training period. In this paper, we propose a novel unsupervised cross-domain person re-identification network (IDENet) based on an inter-domain equilibrium structure to improve the quality of pseudo-labels. Specifically, we make full use of both source domain and target domain information and construct a small learning network to equalize label allocation between the two domains. Based on it, we also develop a dynamic neural network with adaptive convolution kernels to generate adaptive residuals for adapting domain-agnostic deep fine-grained features. In addition, we design the network structure based on ordinary differential equations and embed modules to solve the problem of network overfitting. Extensive cross-domain experimental results on Market1501, PersonX, and MSMT17 prove that our proposed method outperforms the state-of-the-art methods.
Xi Yang 0011, Wenjiao Dong, Gu Zheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 Hyperbolic Insights With Knowledge Distillation for Cross-Domain Few-Shot Learning
abstract
Cross-domain few-shot learning aims to achieve swift generalization between a source domain and a target domain using a limited number of images. Current research predominantly relies on generalized feature embeddings, employing metric classifiers in Euclidean space for classification. However, due to existing disparities among different data domains, attaining generalized features in the embedding becomes challenging. Additionally, the rise in data domains leads to high-dimensional Euclidean spaces. To address the above problems, we introduce a cross-domain few-shot learning method named Hyperbolic Insights with Knowledge Distillation (HIKD). By integrating knowledge distillation, it enhances the model's generalization performance, thereby significantly improving task performance. Hyperbolic space, in comparison to Euclidean space, offers a larger capacity and supports the learning of hierarchical structures among images, which can aid generalized learning across different data domains. So we map the Euclidean space features to the hyperbolic space via hyperbolic embedding and utilize hyperbolic fitting distillation method in the meta-training phase to obtain multi-domain unified generalization representation. In the meta-testing phase, accounting for biases between the source and target domains, we present a hyperbolic adaptive module to adjust embedded features and eliminate inter-domain gap. Experiments on the Meta-Dataset demonstrate that HIKD outperforms state-of-the-arts methods with the average accuracy of 80.6%.
Xi Yang 0011, Dechen Kong, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 Dense Information Learning Based Semi-Supervised Object Detection
abstract
Semi-Supervised Object Detection (SSOD) aims to improve the utilization of unlabeled data, and various methods, such as adaptive threshold techniques, have been extensively studied to increase exploitable information. However, these methods are passive, relying solely on the original image data. Additionally, existing approaches prioritize the predicted categories of the teacher model while overlooking the relationships between different categories in the prediction. In this paper, we introduce a novel approach called Dense Information Learning (DIL), which actively generates unlabeled data containing densely exploitable information and forces the network to have relation consistency under different perturbations. Specifically, Dense Information Augmentation (DIA) leverages the prior information of the network to create a foreground bank and actively incorporates exploitable information into the unlabeled data. DIA automatically performs information enhancement and filters noise. Furthermore, to encourage the network to maintain consistency at the manifold level under various perturbations, we introduce Relation Consistency Regularization (RCR). It considers both feature-level and image-level perturbations, guiding the network to focus on more discriminative features. Extensive experiments conducted on multiple datasets validate the effectiveness of our approach in leveraging information from unlabeled images. The proposed DIL improves the mAP by 12.6% and 10.0% relative to the supervised baseline method when utilizing 5% and 10% of labeled data on the MS-COCO dataset, respectively.
Xi Yang 0011, Qiubai Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 Uncertainty Quantification for Semi-Supervised Object Detection in Remote Sensing Images
abstract
Semi-supervised object detection (SSOD) aims to solve the data annotation challenge in object detection and can achieve remarkable progress in natural scenes; however, it remains unexplored in horizontal bounding box (HBB)-based remote sensing imagery where annotation tasks pose greater challenges. In remote sensing scenarios, objects exhibit arbitrary orientations, small scales, and dense distributions, leading to pseudoboxes with fuzzy boundaries and class imbalance issues. Therefore, we propose UNCertainty quantification (UNC) for SSOD in remote sensing images. UNC uses uncertainty to guide the network from both regression and classification perspectives: Semantic alignment SAM calibration (SASC) uses pseudoboxes as box prompts for the input of the segment anything model (SAM), achieving more precise boundaries. Subsequently, boundaries with lower regression uncertainty are selected as the final pseudoboxes, ensuring better alignment between the pseudoboxes and the ground truth. Dynamic uncertainty weighting (DUW) calculates class uncertainty and determines its correlation with the availability of instances per class. High uncertainty implies limited availability of instances, necessitating greater emphasis on instances of that class. Furthermore, we set a percentage uncertainty threshold to avoid overemphasis caused by individual classes. Extensive experiments conducted on the DIOR and DOTA HBB-based datasets demonstrate the effectiveness of our method in leveraging unlabeled image information. Specifically, compared with the supervised baseline method, the UNC method improves mAP by 12.4% and 8.6% when 5% and 10% of labeled data on DIOR, respectively.
Xi Yang 0011, Qiubai Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 S3OIL: Semi-Supervised SAR-to-Optical Image Translation via Multi-Scale and Cross-Set Matching
abstract
Image-to-image translation has achieved great success, but still faces the significant challenge of limited paired data, particularly in translatingSynthetic Aperture Radar(SAR) images to optical images. Furthermore, most existing semi-supervised methods place limited emphasis on leveraging the data distribution. To address those challenges, we propose aSemi-Supervised SAR-to-Optical Image Translation(S3OIL) method that achieves high-quality image generation using minimal paired data and extensive unpaired data while strategically exploiting the data distribution. To this end, we first introduce aCross-Set Alignment Matching(CAM) mechanism to create local correspondences between the generated results of paired and unpaired data, ensuring cross-set consistency. In addition, for unpaired data, we apply weak and strong perturbations and establish intra-setMulti-Scale Matching(MSM) constraints. For paired data, intra-modal semantic consistency (ISC) is presented to ensure alignment with the ground truth. Finally, we propose local and global cross-modal semantic consistency (CSC) to boost structural identity during translation. We conduct extensive experiments on SAR-to-optical datasets and another sketch-to-anime task, demonstrating that S3OIL delivers competitive performance compared to state-of-the-art unsupervised, supervised, and semi-supervised methods, both quantitatively and qualitatively. Ablation studies further reveal that S3OIL can ensure the preservation of both semantic content and structural integrity of the generated images. Our code is available at: https://github.com/XduShi/SOIL.
Xi Yang 0011, Ziyun Li 0002, Maoying Qiao, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Image Process.6
2025 Toward Generalizable Prompt Learning via Multi-Regularization Guided Knowledge Distillation
abstract
Prompt learning has made significant progress in vision-language models (VLMs), enabling pre-trained models like CLIP to perform cross-domain tasks with few-shot or even zero-shot learning. However, existing methods tend to overfit the training data after fine-tuning on the target domain, leading to a decline in generalization ability and limiting their performance on unseen categories.To address these challenges, we propose a multi-regularization guided knowledge distillation towards generalizable prompt learning. This approach enhances the model's adaptability and generalization through different stages of regularization while mitigating performance degradation caused by target domain training. Specifically, within the image encoder of CLIP, we introduce Residual Regularization, which binds additional residual connections to certain transformer blocks. This design provides greater flexibility, allowing the model to adjust to new data distributions when adapting to the target domain.Furthermore, during training, we impose Self-distillation Regularization to ensure that while adapting to the target domain, the model preserves its prior generalization knowledge. Specifically, we regularize the intermediate layer outputs of Transformer Blocks to prevent the model from excessively favoring target domain data. Additionally, we employ an unsupervised knowledge distillation strategy to enforce multi-level alignment between the teacher and student models by Direction Distillation Regularization. This ensures that both models maintain consistent visual feature orientations under the same textual features, thereby enhancing overall model stability and cross-domain adaptability.Experimental results demonstrate that our method achieves more stable classification performance in both cross-domain few-shot classification and domain adaptation settings.
Xi Yang 0011, Xinyue Zhong, Dechen Kong, Nannan Wang 0001
IEEE Trans. Image Process.4
2025 UCPM: Uncertainty-Guided Cross-Modal Retrieval With Partially Mismatched Pairs
abstract
The manual annotation of perfectly aligned labels for cross-modal retrieval (CMR) is incredibly labor-intensive. As an alternative, the collection of co-occurring data pairs from the Internet is a remarkably cost-effective way, but which, inevitably induces the Partially Mismatched Pairs (PMPs) and therefore significantly degrades the retrieval performance without particular treatment. Previous efforts often utilize the pair-wise similarity to filter out the mismatched pairs, and such operation is highly sensitive to mismatched or ambiguous data and thus leads to sub-optimal performance. To alleviate these concerns, we propose an efficient approach, termed UCPM, i.e., Uncertainty-guided Cross-modal retrieval with Partially Mismatched pairs, which can significantly reduce the adverse impact of mismatched data pairs. Specifically, a novel Uncertainty Guided Division (UGD) strategy is sophisticatedly designed to divide the corrupted training data into confident matched (clean), easily-identifiable mismatched (noisy) and hardly-determined hard subsets, and the derived uncertainty can simultaneously guide the informative pair learning while reducing the negative impact of potential mismatched pairs. Meanwhile, an effective Uncertainty Self-Correction (USC) mechanism is concurrently presented to accurately identify and rectify the fluctuated uncertainty during the training process, which further improves the stability and reliability of the estimated uncertainty. Besides, a Trusted Margin Loss (TML) is newly designed to enhance the discriminability between those hard pairs, by dynamically adjusting their soft margins to amplify the positive contributions of matched pairs while suppressing the negative impacts of mismatched pairs. Extensive experiments on three widely-used benchmark datasets, verify the effectiveness and reliability of UCPM compared with the existing SOTA approaches, and significantly improve the robustness in both synthetic and real-world PMPs. The code is available at: https://github.com/qxzha/UCPM.
Quanxing Zha, Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Xing Xu 0001, Nannan Wang 0001
IEEE Trans. Image Process.6
2025 ETC: Temporal Boundary Expand Then Clarify for Weakly Supervised Video Grounding With Multimodal Large Language Model
abstract
Early weakly supervised video grounding (WSVG) methods often struggle with incomplete boundary detection due to the absence of temporal boundary annotations. To bridge the gap between video-level and boundary-level annotations, explicit supervision methods (i.e., generating pseudo-temporal boundaries for training) have achieved great success. However, data augmentation in these methods might disrupt critical temporal information, yielding poor pseudo-temporal boundaries. In this paper, we propose a new perspective that maintains the integrity of the original temporal content while introducing more valuable information for expanding the incomplete boundaries. To this end, we proposeETC(ExpandthenClarify), first using the additional information to expand the initial incomplete pseudo-temporal boundaries, and subsequently refining these expanded ones to achieve precise boundaries. Motivated by video continuity, i.e., visual similarity across adjacent frames, we use powerful multi-modal large language models (MLLMs) to annotate each frame within the initial pseudo-temporal boundaries, yielding more comprehensive descriptions for expanded boundaries. To further clarify the noise in expanded boundaries, we combine mutual learning with a tailored proposal-level contrastive objective to use a learnable approach to harmonize a balance between incomplete yet clean (initial) and comprehensive yet noisy (expanded) boundaries for more precise ones. Experiments demonstrate the superiority of our method on two challenging WSVG datasets.
Guozhang Li, Xinpeng Ding, De Cheng, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2025 Knowledge-Enhanced Facial Expression Recognition With Emotional-to-Neutral Transformation
abstract
Existing facial expression recognition (FER) methods typically fine-tune a pre-trained visual encoder using discrete labels. However, this form of supervision limits to specify the emotional concept of different facial expressions. In this paper, we observe that the rich knowledge in text embeddings, generated by vision-language models, is a promising alternative for learning discriminative facial expression representations. Inspired by this, we propose a novel knowledge-enhanced FER method with an emotional-to-neutral transformation. Specifically, we formulate the FER problem as a process to match the similarity between a facial expression representation and text embeddings. Then, we transform the facial expression representation to a neutral representation by simulating the difference in text embeddings from textual facial expression to textual neutral. Finally, a self-contrast objective is introduced to pull the facial expression representation closer to the textual facial expression, while pushing it farther from the neutral representation. We conduct evaluation with diverse pre-trained visual encoders including ResNet-18 and Swin-T on four challenging facial expression datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art FER methods. The code is made publicly available athttps://github.com/hangyu94/KE2NT.
Hangyu Li 0001, Jiangchao Yao, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003
IEEE Trans. Multim.4
2025 Generalizable Prompt Learning via Gradient Constrained Sharpness-Aware Minimization
abstract
This paper targets a novel trade-off problem in generalizable prompt learning for vision-language models (VLM), i.e., improving the performance on unseen classes while maintaining the performance on seen classes. Comparing with existing generalizable methods that neglect the seen classes degradation, the setting of this problem is stricter and fits more closely with practical applications. To solve this problem, we start from the optimization perspective, and leverage the relationship between loss landscape geometry and model generalization ability. By analyzing the loss landscapes of the state-of-the-art method and vanilla Sharpness-aware Minimization (SAM) based method, we conclude that the trade-off performance correlates to bothloss valueandloss sharpness, while each of them is indispensable. However, we find the optimizing gradient of existing methods cannot maintain high relevance to both loss value and loss sharpness during optimization, which severely affects their trade-off performance. To this end, we propose a novel SAM-based method for prompt learning, denoted as Gradient Constrained Sharpness-aware Context Optimization (GCSCoOp), to dynamically constrain the optimizing gradient, thus achieving above two-fold optimization objective simultaneously. Extensive experiments verify the effectiveness of GCSCoOp in the trade-off problem.
Liangchen Liu 0001, Nannan Wang 0001, Dawei Zhou 0004, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Multim.2
2025 Masked Attribute Description Embedding for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to match persons who change clothes over long periods. The key challenge in CC-ReID is to extract cloth-irrelated features, such as face, hairstyle, body shape, and gait. Current research mainly focuses on modeling body shape using multi-modal biological features (such as silhouettes and sketches). However, it does not fully leverage the personal description information hidden in the original RGB image. Considering that there are certain attribute descriptions that remain unchanged after the changing of cloth, we propose a Masked Attribute Description Embedding (MADE) method that unifies personal visual appearance and attribute description for CC-ReID. Specifically, handling variable cloth-sensitive information, such as color and type, is challenging for effective modeling. To address this, we mask the clothes type and color information (upper body type, upper body color, lower body type, and lower body color) in the personal attribute description extracted through an attribute detection model. The masked attribute description is then connected and embedded into Transformer blocks at various levels, fusing it with the low-level to high-level features of the image. This approach compels the model to discard cloth information. Experiments are conducted on several CC-ReID benchmarks, including PRCC, LTCC, Celeb-reID-light, and LaST. Results demonstrate that MADE effectively utilizes attribute description, enhancing cloth-changing person re-identification performance, and compares favorably with state-of-the-art methods.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Multim.4
2025 Cluster Assumption-Guided Timestamp-Supervised Temporal Action Segmentation
abstract
Current timestamp-supervised temporal action segmentation (TS-TAS) methods typically follow a two-phase pipeline: initializing the model with timestamp labels and refining it with pseudo-labels. However, limited by the sparsity of timestamp annotations, current methods' performance is sub-optimal. Specifically, initializing the model with only timestamp annotations may cause overfitting to labeled frames. Additionally, sparse timestamp annotations cannot capture the diverse action representations throughout the whole instance, especially those near the ambiguous action boundaries, leading to pseudo-label noise. Inspired by the cluster assumption of semi-supervised learning (SSL) that points within the same manifold likely share the same label, we here model TS-TAS as an SSL problem. Specifically, we propose a Temporal Embedding Consistency (TEC) strategy to mitigate the excessive focus on annotated frames. The TEC strategy encourages frames with similar representations within the video to have similar classification probability distributions, thereby propagating labeled frames' information to implicit ones. Besides, we design a TS-Mix strategy to further leverage unlabeled data to mitigate the influence of pseudo-label noise in a consistency regularization manner. The TS-Mix strategy includes intra-mix, which adds linear interpolation of two adjacent timestamps to every frame between them, and inter-mix, which mixes frames from two different untrimmed videos frame-by-frame. Then the mixed video is trained with the correspondingly mixed pseudo-labels. Comprehensive experimental results on different benchmarks show that we achieve new state-of-the-art performances. Furthermore, the proposed method can seamlessly enhance existing methods, significantly improving their performances.
Ziyou Ren, Guozhang Li, Nan Cheng 0001, Anqi Wu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2025 Foodfusion: A Novel Approach for Food Image Composition via Diffusion Models
abstract
Food image composition requires the use of existing dish images and background images to synthesize a natural new image, while diffusion models have made significant advancements in image generation, enabling the construction of end-to-end architectures that yield promising results. However, existing diffusion models face challenges in processing and fusing information from multiple images and lack access to high-quality publicly available datasets, which prevents the application of diffusion models in food image composition. In this paper, we introduce a large-scale, high-quality food image composite dataset,FC22 k, which comprises 22,000 foreground, background, and ground truth ternary image pairs. Additionally, we propose a novel food image composition method,Foodfusion, which leverages the capabilities of the pre-trained diffusion models and incorporates a Fusion Module for processing and integrating foreground and background information. This fused information aligns the foreground features with the background structure by merging the global structural information at the cross-attention layer of the denoising UNet. To further enhance the content and structure of the background, we also integrate a Content-Structure Control Module. Extensive experiments demonstrate the effectiveness and scalability of our proposed method.
Chaohua Shi, Xuan Wang 0009, Xule Wang, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.6
2025 Progressive Prompt-Driven Low-Light Image Enhancement With Frequency Aware Learning
abstract
Low-light Image Enhancement (LLIE) aims to rectify inadequate illumination conditions and achieve superior visual quality in images, which plays a pivotal role in the domain of low-level computer vision. Due to poor illumination in images, many high-frequency details are obscured, which leads to an uneven distribution of low- and high-frequency information. However, most existing LLIE methods do not pay special attention to the restoration of high-frequency detail information and some challenging-to-recover areas in images. To address this issue, we propose a novel progressive prompt-driven LLIE framework with frequency aware learning, through a two-stage coarse-to-fine learning mechanism. Specifically, the proposed method fully utilizes both the specially designed brightness-aware prompt and detail-aware prompt on the prior trained model, to achieve an excellent enhanced image that exhibits more natural brightness and richer detail information. Furthermore, the proposed frequency aware learning objective can adaptively adjust the contribution of individual pixels for image reconstruction based on the statistics of high- and low-frequency features, which enables the network to focus on learning intricate details and other challenging areas in low-light images. Extensive experimental results demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods on representative real-world and synthetic datasets. Our source code is available athttps://github.com/MSL502/PPFAL.
De Cheng, Yan Li 0125, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001, Jiande Sun 0001
IEEE Trans. Multim.4
2025 Dual Semantic Reconstruction Network for Weakly Supervised Temporal Sentence Grounding
abstract
Weakly supervised temporal sentence grounding aims to identify semantically relevant video moments in an untrimmed video corresponding to a given sentence query without exact timestamps. Neuropsychology research indicates that the way the human brain handles information varies based on the grammatical categories of words, highlighting the importance of separately considering nouns and verbs. However, current methodologies primarily utilize pre-extracted video features to reconstruct randomly masked queries, neglecting the distinction between grammatical classes. This oversight could hinder forming meaningful connections between linguistic elements and the corresponding components in the video. To address this limitation, this paper introduces the dual semantic reconstruction network (DSRN) model. DSRN processes video features by distinctly correlating object features with nouns and motion features with verbs, thereby mimicking the human brain's parsing mechanism. It begins with a feature disentanglement module that separately extracts object-aware and motion-aware features from video content. Then, in a dual-branch structure, these disentangled features are used to generate separate proposals for objects and motions through two dedicated proposal generation modules. A consistency constraint is proposed to ensure a high level of agreement between the boundaries of object-related and motion-related proposals. Subsequently, the DSRN independently reconstructs masked nouns and verbs from the sentence queries using the generated proposals. Finally, an integration block is applied to synthesize the two types of proposals, distinguishing between positive and negative instances through contrastive learning. Experiments on the Charades-STA and ActivityNet Captions datasets demonstrate that the proposed method achieves state-of-the-art performance.
Kefan Tang, Lihuo He, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2025 Knowledge Distillation Meets Label Noise Learning: Ambiguity-Guided Mutual Label Refinery
abstract
Knowledge distillation (KD), which aims at transferring the knowledge from a complex network (a teacher) to a simpler and smaller network (a student), has received considerable attention in recent years. Typically, most existing KD methods work on well-labeled data. Unfortunately, real-world data often inevitably involve noisy labels, thus leading to performance deterioration of these methods. In this article, we study a little-explored but important issue, i.e., KD with noisy labels. To this end, we propose a novel KD method, called ambiguity-guided mutual label refinery KD (AML-KD), to train the student model in the presence of noisy labels. Specifically, based on the pretrained teacher model, a two-stage label refinery framework is innovatively introduced to refine labels gradually. In the first stage, we perform label propagation (LP) with small-loss selection guided by the teacher model, improving the learning capability of the student model. In the second stage, we perform mutual LP between the teacher and student models in a mutual-benefit way. During the label refinery, an ambiguity-aware weight estimation (AWE) module is developed to address the problem of ambiguous samples, avoiding overfitting these samples. One distinct advantage of AML-KD is that it is capable of learning a high-accuracy and low-cost student model with label noise. The experimental results on synthetic and real-world noisy datasets show the effectiveness of our AML-KD against state-of-the-art KD methods and label noise learning (LNL) methods. Code is available at https://github.com/Runqing-forMost/ AML-KD.
Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Nannan Wang 0001, Hanzi Wang
IEEE Trans. Neural Networks Learn. Syst.5
2025 PStyle-3D: Example-Based 3-D-Aware Portrait Style Domain Adaptation
abstract
The creation of high-quality artistic portraits is a critical and desirable task in the field of computer vision. While recent3-D generative models have achieved impressive results in generating images with view consistency and intricate 3-D shapes, their application for generating artistic portraits is often more challenging than 2-D generative models due to the potentially destructive impact of 3-D structures on human faces. This article introduces a novel approach that leverages a meticulously designed domain feature extraction module to extract the specific feature information from both the source natural face domain and the target artistic portrait domain. These extracted features are seamlessly integrated into a 3-D representation, generating multiview consistent 3-D artistic portraits. To fuse the features of the source and target domains better, we propose a new module for domain adaptation. This module adds a path to the style path established by StyleGAN to introduce the artistic portrait domain information and regulate the target domain's feature information in $\mathcal {S}$ space. Our domain adaptation module is implemented in each StyleBlock of the 3-D representation generator to integrate the target domain information with the original facial information. Experimental results demonstrate that our approach generates high-quality 3-D artistic portraits that outperform existing approaches in preserving 3-D geometric information and multiview consistency.
Chaohua Shi, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 I²NQ: Inter and Intra Nonuniform Quantization for Single Image Super-Resolution
abstract
Quantizing neural network is an efficient model compression technique that converts weights and activations from floating-point to integer. However, existing model quantization methods are primarily designed for high-level visual tasks. They do not sufficiently consider the unique characteristics of feature distribution in image super-resolution (SR) reconstruction models. On the one hand, the objective of SR is to restore high-frequency and fine-detail information while preserving the overall feature distribution. Therefore, the regularization techniques are removed to maintain the original distribution. However, vanilla quantization methods often employ regularization techniques to normalize the features for stable network training, which destroys the inherent information of the feature distribution. On the other hand, the feature distribution in SR models exhibits a nonuniform bell-shaped form. Common quantization methods adopt a uniform quantization strategy with equal quantization intervals. This fails to effectively capture the nonuniform feature distribution in SR. To address the above issue, we propose a novel method named Inter and Intra Nonuniform Quantization, which takes into account the specific characteristics of the feature distribution in the context of SR reconstruction models. Additionally, we propose a weight adjustment method called flex-scale-weight-adjust (FSWA). It can maintain the diversity of weight information and reduce quantization errors. Extensive experiments demonstrate that our proposed method surpasses other quantization methods in both the evaluation of reconstruction metrics and visual reconstruction performance.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Rectified Binary Network for Single-Image Super-Resolution
abstract
Binary neural network (BNN) is an effective approach to reduce the memory usage and the computational complexity of full-precision convolutional neural networks (CNNs), which has been widely used in the field of deep learning. However, there are different properties between BNNs and real-valued models, making it difficult to draw on the experience of CNN composition to develop BNN. In this article, we study the application of binary network to the single-image super-resolution (SISR) task in which the network is trained for restoring original high-resolution (HR) images. Generally, the distribution of features in the network for SISR is more complex than those in recognition models for preserving the abundant image information, e.g., texture, color, and details. To enhance the representation ability of BNN, we explore a novel activation-rectified inference (ARI) module that achieves a more complete representation of features by combining observations from different quantitative perspectives. The activations are divided into several parts with different quantification intervals and are inferred independently. This allows the binary activations to retain more image detail and yield finer inference. In addition, we further propose an adaptive approximation estimator (AAE) for gradually learning the accurate gradient estimation interval in each layer to alleviate the optimization difficulty. Experiments conducted on several benchmarks show that our approach is able to learn a binary SISR model with superior performance over the state-of-the-art methods. The code will be released at https://github.com/jwxintt/Rectified-BSR.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 TIENet: A Tri-Interaction Enhancement Network for Multimodal Person Reidentification
abstract
Multimodal person reidentification (ReID), which aims to learn modality-complementary information by utilizing multimodal images simultaneously for person retrieval, is crucial for achieving all-time and all-weather monitoring. Existing methods try to address this issue through modality fusion to absorb complementary information. However, most of these methods are limited to the spatial domain only and usually overlook the intra-/intermodal interactions during feature fusion, resulting in insufficient learning of modality-specific and complementary information. To address these issues, we propose a tri-interaction enhancement network (TIENet), which contains three modules: spatial-frequency interaction (SFI), intermodal mask interaction (IMMI), and intramodal feature fusion (IMFF). Specifically, the SFI boosts the modality-specific representation by integrating the amplitude-guided attention mechanism into the phase space, combined with spatial-domain convolution to achieve fine-grained information learning. Meanwhile, the IMMI enhances the richness of the feature descriptors by embedding the intermodal relationships to preserve complementary information. Finally, the IMFF module considers the structure of the human body and integrates intramodal contextual information. Extensive experimental results demonstrate the effectiveness of our method, achieving superior performances on RGBNT201 and MARKET1501_RGBNT datasets.
Xi Yang 0011, Wenjiao Dong, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Few-Shot Face Stylization via GAN Prior Distillation
abstract
Face stylization has made notable progress in recent years. However, when training on limited data, the performance of existing approaches significantly declines. Although some studies have attempted to tackle this problem, they either failed to achieve the few-shot setting (less than 10) or can only get suboptimal results. In this article, we propose GAN Prior Distillation (GPD) to enable effective few-shot face stylization. GPD contains two models: a teacher network with GAN Prior and a student network that fulfills end-to-end translation. Specifically, we adapt the teacher network trained on large-scale data in the source domain to the target domain using a handful of samples, where it can learn the target domain's knowledge. Then, we can achieve few-shot augmentation by generating source domain and target domain images simultaneously with the same latent codes. We propose an anchor-based knowledge distillation module that can fully use the difference between the training and the augmented data to distill the knowledge of the teacher network into the student network. The trained student network achieves excellent generalization performance with the absorption of additional knowledge. Qualitative and quantitative experiments demonstrate that our method achieves superior results than state-of-the-art approaches in a few-shot setting.
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation
abstract
High-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. However, this dense high-resolution representation brings a significant computational burden. In this paper, we address the following question: "Only sparse human keypoint locations are detected for human pose estimation, is it really necessary to describe the whole image in a dense, high-resolution manner?" Based on dynamic transformer models, we propose a framework that only uses Sparse High-resolution Representations for human Pose estimation (SHaRPose). In detail, SHaRPose consists of two stages. At the coarse stage, the relations between image regions and keypoints are dynamically mined while a coarse estimation is generated. Then, a quality predictor is applied to decide whether the coarse estimation results should be refined. At the fine stage, SHaRPose builds sparse high-resolution representations only on the regions related to the keypoints and provides refined high-precision human pose estimations. Extensive experiments demonstrate the outstanding performance of the proposed method. Specifically, compared to the state-of-the-art method ViTPose, our model SHaRPose-Base achieves 77.4 AP (+0.5 AP) on the COCO validation set and 76.7 AP (+0.5 AP) on the COCO test-dev set, and infers at a speed of 1.4x faster than ViTPose-Base. Code is available at https://github.com/AnxQ/sharpose.
Xiaoqi An, Lin Zhao 0003, Chen Gong 0002, Nannan Wang 0001, Di Wang 0011, Jian Yang 0003
AAAI4
2024 Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object Tracking
abstract
The global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature poses storage space requirements that challenge algorithm handling long videos. Currently, commonly used methods are still generated trajectories by building one-forward associations across frames. Such matches produced under the guidance of first-order similarity information may not be optimal from a longer-time perspective. Moreover, they often lack an end-to-end scheme for correcting mismatches. This paper proposes the Composite Node Message Passing Network (CoNo-Link), a multi-scene generalized framework for modeling ultra-long frames information for association. CoNo-Link's solution is a low-storage overhead method for building constrained connected graphs. In addition to the previous method of treating objects as nodes, the network innovatively treats object trajectories as nodes for information interaction, improving the graph neural network's feature representation capability. Specifically, we formulate the graph-building problem as a top-k selection task for some reliable objects or trajectories. Our model can learn better predictions on longer-time scales by adding composite nodes. As a result, our method outperforms the state-of-the-art in several commonly used datasets.
Yan Gao 0025, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
AAAI4
2024 Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model
abstract
Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low transferability and high detectability. In this paper, we propose a unified framework Adv-Diffusion that can generate imperceptible adversarial identity perturbations in the latent space but not the raw pixel space, which utilizes strong inpainting capabilities of the latent diffusion model to generate realistic adversarial images. Specifically, we propose the identity-sensitive conditioned diffusion generative model to generate semantic perturbations in the surroundings. The designed adaptive strength-based adversarial perturbation algorithm can ensure both attack transferability and stealthiness. Extensive qualitative and quantitative experiments on the public FFHQ and CelebA-HQ datasets prove the proposed method achieves superior performance compared with the state-of-the-art methods without an extra generative model training process. The source code is available at https://github.com/kopper-xdu/Adv-Diffusion.
Decheng Liu, Xijun Wang 0005, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
AAAI4
2024 Point Deformable Network with Enhanced Normal Embedding for Point Cloud Analysis
abstract
Recently MLP-based methods have shown strong performance in point cloud analysis. Simple MLP architectures are able to learn geometric features in local point groups yet fail to model long-range dependencies directly. In this paper, we propose Point Deformable Network (PDNet), a concise MLP-based network that can capture long-range relations with strong representation ability. Specifically, we put forward Point Deformable Aggregation Module (PDAM) to improve representation capability in both long-range dependency and adaptive aggregation among points. For each query point, PDAM aggregates information from deformable reference points rather than points in limited local areas. The deformable reference points are generated data-dependent, and we initialize them according to the input point positions. Additional offsets and modulation scalars are learned on the whole point features, which shift the deformable reference points to the regions of interest. We also suggest estimating the normal vector for point clouds and applying Enhanced Normal Embedding (ENE) to the geometric extractors to improve the representation ability of single-point. Extensive experiments and ablation studies on various benchmarks demonstrate the effectiveness and superiority of our PDNet.
Xingyilang Yin, Xi Yang 0011, Liangchen Liu 0001, Nannan Wang 0001, Xinbo Gao 0001
AAAI4
2024 Generating Handwritten Mathematical Expressions From Symbol Graphs: An End-to-End Pipeline
abstract
In this paper, we explore a novel challenging generation task, i.e. Handwritten Mathematical Expression Generation (HMEG) from symbolic sequences. Since symbolic sequences are naturally graph-structured data, we formulate HMEG as a graph-to-image (G2I) generation problem. Unlike the generation of natural images, HMEG requires critic layout clarity for synthesizing correct and recognizable formulas, but has no real masks available to supervise the learning process. To alleviate this challenge, we propose a novel end-to-end G2I generation pipeline (i.e. graph → layout →mask →image), which requires no real masks or nondifferentiable alignment between layouts and masks. Technically, to boost the capacity of predicting detailed relations among adjacent symbols, we propose a Less-is-More (LiM) learning strategy. In addition, we design a differentiable layout refinement module, which maps bounding boxes to pixel-level soft masks, so as to further alleviate ambiguous layout areas. Our whole model, including layout prediction, mask refinement, and image generation, can be jointly optimized in an end-to-end manner. Experimental results show that, our model can generate highquality HME images, and outperforms previous generative methods. Besides, a series of ablations study demonstrate effectiveness of the proposed techniques. Finally, we validate that our generated images promisingly boosts the performance of HME recognition models, through data augmentation. Our code and results are available at: https://github.com/AiArt-HDU/HMEG.
Yu Chen 0003, Fei Gao 0006, Yanguang Zhang, Maoying Qiao, Nannan Wang 0001
CVPR5
2024 Disentangled Prompt Representation for Domain Generalization
abstract
Domain Generalization (DG) aims to develop a versatile model capable of performing well on unseen target domains. Recent advancements in pre-trained Visual Foundation Models (VFMs), such as CLIP, show significant potential in enhancing the generalization abilities of deep models. Although there is a growing focus on VFM-based domain prompt tuning for DG, effectively learning prompts that disentangle invariant features across all domains remains a major challenge. In this paper, we propose addressing this challenge by leveraging the controllable and flexible language prompt of the VFM. Observing that the text modality of VFMs is inherently easier to disentangle, we introduce a novel text feature guided visual prompt tuning framework. This framework first automatically disentangles the text prompt using a large language model (LLM) and then learns domain-invariant visual representation guided by the disentangled text feature. Moreover, we also devise domain-specific prototype learning to fully exploit domain-specific information to combine with the invariant feature prediction. Extensive experiments on mainstream DG datasets, namely PACS, VLCS, OfficeHome, DomainNet and TerraInc, demonstrate that the proposed method achieves superior performances to state-of-the-art DG methods.
De Cheng, Xinyang Jiang, Nannan Wang 0001, Dongsheng Li 0002, Xinbo Gao 0001
CVPR4
2024 Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization
Xi Yang 0011, Songsong Duan, Nannan Wang 0001, Xinbo Gao 0001
ECCV (69)3
2024 On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection
abstract
Image-to-image (I2I) translation is vital in computer vision tasks like style transfer and domain adaptation. While recent advances in GAN have enabled high-quality sample generation, real-world challenges such as noise and distortion remain significant obstacles. Although Gaussian noise injection during training has been utilized, its theoretical underpinnings have been unclear. This work provides a robust theoretical framework elucidating the role of Gaussian noise injection in I2I translation models. We address critical questions on the influence of noise variance on distribution divergence, resilience to unseen noise types, and optimal noise intensity selection. Our contributions include connecting $f$-divergence and score matching, unveiling insights into the impact of Gaussian noise on aligning probability distributions, and demonstrating generalized robustness implications. We also explore choosing an optimal training noise level for consistent performance in noisy environments. Extensive experiments validate our theoretical findings, showing substantial improvements over various I2I baseline models in noisy settings. Our research rigorously grounds Gaussian noise injection for I2I translation, offering a sophisticated theoretical understanding beyond heuristic applications.
Chaohua Shi, Lu Gan 0002, Hongqing Liu 0001, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
ICLR6
2024 Robust Training of Federated Models with Extremely Label Deficiency
abstract
Federated semi-supervised learning (FSSL) has emerged as a powerful paradigm for collaboratively training machine learning models using distributed data with label deficiency. Advanced FSSL methods predominantly focus on training a single model on each client. However, this approach could lead to a discrepancy between the objective functions of labeled and unlabeled data, resulting in gradient conflicts. To alleviate gradient conflict, we propose a novel twin-model paradigm, called **Twinsight**, designed to enhance mutual guidance by providing insights from different perspectives of labeled and unlabeled data. In particular, Twinsight concurrently trains a supervised model with a supervised objective function while training an unsupervised model using an unsupervised objective function. To enhance the synergy between these two models, Twinsight introduces a neighborhood-preserving constraint, which encourages the preservation of the neighborhood relationship among data features extracted by both models. Our comprehensive experiments on four benchmark datasets provide substantial evidence that Twinsight can significantly outperform state-of-the-art methods across various experimental settings, demonstrating the efficacy of the proposed Twinsight.
Yonggang Zhang 0003, Zhiqin Yang, Xinmei Tian 0001, Nannan Wang 0001, Tongliang Liu, Bo Han 0003
ICLR4
2024 Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual Learning
abstract
Continual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis—one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task—the recent progress has achieved forget-free CL. Although addressing the CF issue well, such methods would encounter serious under-fitting in long-term CL, in which the learning process will continue for a long time and the number of new tasks involved will be much higher. To solve this problem, this paper partitions the network into three parts—with a new part for exploring the knowledge sharing between the old and new tasks. With the shared knowledge, this part of network can be learnt to simultaneously consolidate the old tasks and fit to the new task. To achieve this goal, we propose a task-aware Orthogonal Sparse Network (OSN), which contains shared knowledge induced network partition and sharpness-aware orthogonal sparse network learning. The former partitions the network to select shared parameters, while the latter guides the exploration of shared knowledge through shared parameters. Qualitative and quantitative analyses, show that the proposed OSN induces minimum to no interference with past tasks, i.e., approximately no forgetting, while greatly improves the model plasticity and capacity, and finally achieves the state-of-the-art performances.
Yusong Hu, De Cheng, Dingwen Zhang, Nannan Wang 0001, Tongliang Liu, Xinbo Gao 0001
ICML4
2024 Human-Robot Interactive Creation of Artistic Portrait Drawings
abstract
In this paper, we present a novel system for Human-Robot Interactive Creation of Artworks (HRICA). Different from previous robot painters, HRICA allows a human user and a robot to alternately draw strokes on a canvas, to collaboratively create a portrait drawing through frequent interactions. The key is to enable the robot to understand human intentions, during the interactive creation process. We here formulate this as a mask-free image inpainting problem, and propose a novel method to estimate the complete version of a portrait drawing, after the human user has drawn some initial strokes. In this way, the robot can select some complementary strokes and draw them on the canvas. To train and evaluate our inpainting method, we construct a novel large-scale portrait drawing dataset, CelebLine, which composes of high-quality portrait line-drawings, with dense labels of both 2D semantic parsing masks and 3D depth maps. Finally, we develop a human-robot interactive drawing system with low-cost hardware, user-friendly interface, and interesting creation experience. Experiments show that our robot can stably cooperate with human users to create diverse styles of portrait drawings. In addition, our portrait drawing inpainting method significantly outperforms previous advanced methods. The code and dataset have been released at: https://github.com/fei-aiart/HRICA.
Fei Gao 0006, Lingna Dai, Jingjie Zhu, Mei Du, Maoying Qiao, Chenghao Xia, Nannan Wang 0001, Peng Li 0031
ICRA8
2024 Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
Shiyin Dong, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IJCAI4
2024 Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search
Haichun Tai, De Cheng, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IJCAI4
2024 AesMamba: Universal Image Aesthetic Assessment with State Space Models
abstract
Image Aesthetic Assessment (IAA) aims to objectively predict the generic or personalized evaluations, of the aesthetic or fine-grained multi-attributes, based on visual or multimodal inputs. Previously, researchers have designed diverse and specialized methods, for specific IAA tasks, based on different input-output situations. Is it possible to design a universal IAA framework applicable for the whole IAA task taxonomy? In this paper, we explore this issue, and propose a modular IAA framework, dubbed AesMamba. Specially, we use the Visual State Space Model (VMamba), instead of CNNs or ViTs, to learn comprehensive representations of aesthetic-related attributes; because VMamba can efficiently achieve both global and local effective receptive fields. Afterward, a modal-adaptive module is used to automatically produce the integrated representations, conditioned on the type of input. In the prediction module, we propose a Multitask Balanced Adaptation (MBA) module, to boost task-specific features, with emphasis on the tail instances. Finally, we formulate the personalized IAA task as a multimodal learning problem, by converting a user's anonymous subject characters to a text prompt. This prompting strategy effectively employs the semantics of flexibly selected characters, for inferring individual preferences. AesMamba can be applied to diverse IAA tasks, through flexible combination of these modules. Extensive experiments on numerous datasets, demonstrate that AesMamba consistently achieves superior or competitive performance, on all IAA tasks, in comparison with previous SOTA methods. The code has been released at https://github.com/AiArt-Gao/AesMamba Github.
Fei Gao 0006, Maoying Qiao, Nannan Wang 0001
ACM Multimedia5
2024 Disentangling Identity Features from Interference Factors for Cloth-Changing Person Re-identification
abstract
Cloth-Changing Person Re-Identification (CC-ReID) aims to accurately identify a target person in the more realistic surveillance scenario where clothes of the pedestrian may change drastically, which is critical in public security systems for tracking down disguised criminal suspects. Existing methods mainly transform the CC-ReID problem into cross-modality feature alignment from the data-driven perspective, without modelling the interference factors such as clothes and camera view changes meticulously. This may lead to over-consideration or under-consideration of the influence of these factors on the extraction of robust and discriminative identity features. This paper proposes a novel algorithm for thoroughly disentangling identity features from interference factors brought by clothes and camera view changes while ensuring the robustness and discriminability. It adopts a dual-stream identity feature learning framework consisting of a raw image stream and a cloth-erasing stream, to explore discriminative and cloth-irrelevant identity feature representations. Specifically, an adaptive cloth-irrelevant contrastive objective is introduced to contrast features extracted by the two streams, aiming to suppress the fluctuation caused by clothes textures in the identity feature space. Moreover, we innovatively mitigate the influence of the interference factors through a generative adversarial interference factor decoupling network. This network is targeted at capturing identity-related information residing in the interference factors and disentangling the identity features from such information. Extensive experimental results demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods.
De Cheng, Chaowei Fang, Changzhe Jiao, Nannan Wang 0001, Xinbo Gao 0001
ACM Multimedia5
2024 Advancing Generalized Deepfake Detector with Forgery Perception Guidance
abstract
One of the serious impacts brought by artificial intelligence is the abuse of deepfake techniques. Despite the proliferation of deepfake detection methods aimed at safeguarding the authenticity of media across the Internet, they mainly consider the improvement of detector architecture or the synthesis of forgery samples. The forgery perceptions, including the feature responses and prediction scores for forgery samples, have not been well considered. As a result, the generalization across multiple deepfake techniques always comes with complicated detector structures and expensive training costs. In this paper, we shift the focus to real-time perception analysis in the training process and generalize deepfake detectors through an efficient method dubbed Forgery Perception Guidance (FPG). In particular, after investigating the deficiencies of forgery perceptions, FPG adopts a sample refinement strategy to pertinently train the detector, thereby elevating the generalization efficiently. Moreover, FPG introduces more sample information as explicit optimizations, which makes the detector further adapt the sample diversities. Experiments demonstrate that FPG improves the generality of deepfake detectors with small training costs, minor detector modifications, and the acquirement of real data only. In particular, our approach not only outperforms the state-of-the-art on both the cross-dataset and cross-manipulation evaluation but also surpasses the baseline that needs more than 3× training time.
Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Lin Yuan 0002, Shuodi Wang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
ACM Multimedia7
2024 Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VIReID) is widely used in fields such as video surveillance and intelligent transportation, imposing higher demands on model security. In practice, the adversarial attacks based on VIReID aim to disrupt output ranking and quantify the security risks of models. Although numerous studies have been emerged on adversarial attacks and defenses in fields such as face recognition, person re-identification, and pedestrian detection, there is currently a lack of research on the security of VIReID systems. To this end, we propose to explore the vulnerabilities of VIReID systems and prevent potential serious losses due to insecurity. Compared to research on single-modality ReID, adversarial feature alignment and modality differences need to be particularly emphasized. Thus, we advocate for feature-level adversarial attacks to disrupt the output rankings of VIReID systems. To obtain adversarial features, we introduce \textit{Universal Adversarial Perturbations} (UAP) to simulate common disturbances in real-world environments. Additionally, we employ a \textit{Frequency-Spatial Attention Module} (FSAM), integrating frequency information extraction and spatial focusing mechanisms, and further emphasize important regional features from different domains on the shared features. This ensures that adversarial features maintain consistency within the feature space. Finally, we employ an \textit{Auxiliary Quadruple Adversarial Loss} to amplify the differences between modalities, thereby improving the distinction and recognition of features between visible and infrared images, which causes the system to output incorrect rankings. Extensive experiments on two VIReID benchmarks (i.e., SYSU-MM01, RegDB) and different systems validate the effectiveness of our method.
Xi Yang 0011, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
NeurIPS4
2024 Visual Prompt Tuning in Null Space for Continual Learning
abstract
Existing prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace spanned by previous tasks' features, so as to ensure no interference on tasks that have been learned to overcome catastrophic forgetting in CL. However, different from the orthogonal projection in the traditional CNN architecture, the prompt gradient orthogonal projection in the ViT architecture shows completely different and greater challenges, i.e., 1) the high-order and non-linear self-attention operation; 2) the drift of prompt distribution brought by the LayerNorm in the transformer block. Theoretically, we have finally deduced two consistency conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of eliminating interference on previously learned knowledge via the self-attention mechanism in visual prompt tuning. In practice, an effective null-space-based approximation solution has been proposed to implement the prompt gradient orthogonal projection. Extensive experimental results demonstrate the effectiveness of anti-forgetting on four class-incremental benchmarks with diverse pre-trained baseline models, and our approach achieves superior performances to state-of-the-art methods. Our code is available at https://github.com/zugexiaodui/VPTinNSforCL
Yue Lu 0008, Shizhou Zhang, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001
NeurIPS5
2024 Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection
abstract
Unsupervised out-of-distribution (OOD) detection aims to identify out-of-domain data by learning only from unlabeled In-Distribution (ID) training samples, which is crucial for developing a safe real-world machine learning system. Current reconstruction-based method provides a good alternative approach, by measuring the reconstruction error between the input and its corresponding generative counterpart in the pixel/feature space. However, such generative methods face the key dilemma, $i.e.$, improving the reconstruction power of the generative model, while keeping compact representation of the ID data. To address this issue, we propose the diffusion-based layer-wise semantic reconstruction approach for unsupervised OOD detection. The innovation of our approach is that we leverage the diffusion model's intrinsic data reconstruction ability to distinguish ID samples from OOD samples in the latent feature space. Moreover, to set up a comprehensive and discriminative feature representation, we devise a multi-layer semantic feature extraction strategy. Through distorting the extracted features with Gaussian noises and applying the diffusion model for feature reconstruction, the separation of ID and OOD samples is implemented according to the reconstruction errors. Extensive experimental results on multiple benchmarks built upon various datasets demonstrate that our method achieves state-of-the-art performance in terms of detection accuracy and speed.
Ying Yang 0020, De Cheng, Chaowei Fang, Yubiao Wang, Changzhe Jiao, Lechao Cheng, Nannan Wang 0001, Xinbo Gao 0001
NeurIPS7
2024 Spatial-Frequency Dual-Stream Reconstruction for Deepfake Detection
Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001
PRCV (11)5
2024 UGNCL: Uncertainty-Guided Noisy Correspondence Learning for Efficient Cross-Modal Matching
Quanxing Zha, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001, Jianjia Cao
SIGIR5
2024 Pyramid-resolution person restoration for cross-resolution person re-identification
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Sci. China Inf. Sci.4
2024 GazeForensics: DeepFake detection via gaze-guided spatial inconsistency learning
Qinlin He, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks4
2024 Local artifacts amplification for deepfakes augmentation
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks4
2024 Unconstrained Facial Expression Recognition With No-Reference De-Elements Learning
abstract
Most unconstrained facial expression recognition (FER) methods take original facial images as inputs to learn discriminative features by well-designed loss functions, which cannot reflect important visual information in faces. Although existing methods have explored the visual information of constrained facial expressions, there is no explicit modeling of what visual information is important for unconstrained FER. To find out valuable information of unconstrained facial expressions, we pose a new problem of no-reference de-elements learning: we decompose any unconstrained facial image into the facial expression element and a neutral face without the reference of corresponding neutral faces. Importantly, the element provides visualization results to understand important facial expression information and improves the discriminative power of features. Moreover, we propose a simple yet effectiveDe-ElementsNetwork (DENet) to learn the element and introduce appropriate constraints to overcome no ground truth of corresponding neutral faces during the de-elements learning. We extensively evaluate the proposed method on in-the-wild FER datasets including RAF-DB, AffectNet, SFEW and FERPlus. The comparable results show that our method is promising to improve classification performance and achieves equivalent performance compared with state-of-the-art methods. Also, we demonstrate the strong generalization performance on realistic occlusion and pose variation datasets and the cross-dataset evaluation.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Affect. Comput.2
2024 Invertible Image Obfuscation for Facial Privacy Protection via Secure Flow
abstract
This paper presents a fresh paradigm for protecting facial privacy via an invertible image obfuscation framework that incorporates multiple characteristics including anonymity, diversity, reversibility, security, and lightweight all at once. We name the framework PRO-Face S, an acronym for Privacy-preserving Reversible Obfuscation of Face images via Secure flow. The core of the proposed framework is a flow-based generative model (or invertible neural network), which takes as input a face image along with its pre-obfuscated form, and outputs the privacy-protected image that visually mirrors the pre-obfuscated one. The pre-obfuscation applied can be in various forms with different types and strengths. The invertibility of the flow-based model ensures that the original image can be easily recovered from the protected image in high fidelity. An elaborate secret key mechanism is devised to securely guide the mutual transformations of privacy protection and image recovery, such that the correct recovery is only possible upon the availability of the correct secret, pre-specified by the user in the protection stage. Two modes of wrong recovery are investigated to deal with malicious recovery attempts in different scenarios. Finally, extensive experiments conducted on multiple image datasets demonstrate the superiority of the proposed framework over state-of-the-art methods.
Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Jiaxu Leng, Tao Wu 0003, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Efficient Statistical Sampling Adaptation for Exemplar-Free Class Incremental Learning
abstract
Deep learning systems typically suffer from catastrophic forgetting of old knowledge when learning from new data continually. Recently, various class incremental learning (CIL) methods have been proposed to address this issue, and some approaches achieve promising performances by relying on rehearsing the training data of previous tasks. However, storing data from previous tasks would encounter data privacy and memory issues in real-world applications. In this paper, we propose a statistical sampling adaptation method for efficient Exemplar-Free Class-Incremental Learning (EFCIL). Here, instead of preserving the images/features themselves of previous tasks/classes, we store image feature statistics from previous classes to maintain the decision boundary, which is memory-efficient and much semantic-representative. When utilizing the old-class feature statistics, we build a statistical feature adaptation network (SFAN) with a manifold consistency regularization and then train it in a transductive learning paradigm, which can map the outdated statistics onto the current feature space to facilitate a compatible and balanced classifier training subsequently. In this way, the final classifier can be jointly optimized with all the old-class features projected by SFAN and current new-class features, thus alleviating the classification bias problem in EFCIL. Experimental results greatly demonstrate the effectiveness of the proposed method, achieving superior performances than state-of-the-art approaches. Our source code is released inhttps://github.com/yxzhcv/ESSA-EFCIL.
De Cheng, Nannan Wang 0001, Guozhang Li, Dingwen Zhang, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Diff-Privacy: Diffusion-Based Face Privacy Protection
abstract
Privacy protection has become a top priority due to the widespread collection and misuse of personal data. Anonymization and visual identity information hiding are two crucial tasks in face privacy protection, both striving to alter identifying characteristics from face images to prevent privacy information leakage. However, the goals of the two are not entirely the same. Consequently, training a model to simultaneously perform both tasks proves challenging. In this paper, we propose Diff-Privacy, a novel face privacy protection method based on diffusion models that unifies the task of anonymization and visual identity information hiding. Specifically, we present a Multi-Scale image Inversion module (MSI) that, through training, generates a set of Stable Diffusion (SD) format conditional embeddings for the original image. With these conditional embeddings, we design corresponding embedding scheduling strategies and formulate distinct energy functions during the inference process to achieve anonymization and visual identity information hiding, respectively. Extensive experiments demonstrate the effectiveness of the proposed method in protecting face privacy.
Xiao He 0014, Mingrui Zhu, Dongxin Chen, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Few-Shot Font Generation by Learning Style Difference and Similarity
abstract
Few-shot font generation (FFG) aims to preserve the underlying global structure of the original character while generating target fonts by referring to a few samples. It has been applied to font library creation, a personalized signature, and other scenarios. Existing FFG methods explicitly disentangle content and style of reference glyphs universally or component-wisely. However, they ignore the difference between glyphs in different styles and the similarity of glyphs in the same style, which results in artifacts such as local distortions and style inconsistency. To address this issue, we propose a novel font generation approach by learning the Difference between different styles and the Similarity of the same style (DS-Font). We introduce contrastive learning to consider the positive and negative relationship between styles. Specifically, we propose a multi-layer style projector (MSP) for style encoding and realize a distinctive style representation via our proposed Cluster-level Contrastive Style (CCS) loss. The MSP module is employed to assist the generator during training to enhance the style consistency between the generated glyph and the reference glyphs. In addition, we design a glyph-independent patch discriminator, which comprehensively considers different areas of the image and ensures that each style can be distinguished independently. We conduct qualitative and quantitative evaluations comprehensively to demonstrate that our approach achieves significantly better results than state-of-the-art methods.
Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 MRLReID: Unconstrained Cross-Resolution Person Re-Identification With Multi-Task Resolution Learning
abstract
Cross-resolution person re-identification (ReID) is a challenging task that addresses the issue of matching individuals across different resolution conditions. Traditional person ReID methods often assume that images have sufficiently high resolution and overlook the practical scenarios involving low-resolution or blurry images. Existing cross-resolution ReID approaches either utilize image super-resolution techniques to improve the quality of low-resolution images or extract and learn resolution invariant features for person representation. Although multi-task learning has been applied in ReID to integrate auxiliary tasks including attribute recognition, image super-resolution, and so on, how to incorporate the vital resolution learning task into cross-resolution ReID has rarely explored before. Therefore, we propose a novel multi-task resolution learning based ReID network named MRLReID. Our approach treats ross-resolution person ReID as the primary task and the resolution estimation as an auxiliary task. Our network simultaneously learns the resolution information and person identity information of images, aiming to improve cross-resolution person ReID performance. Considering that existing similuated cross-resolution datasets are too simple to mimic unconstrained scenario, we further employ image degradation technique to simulate more realistic cross-resolution ReID datasets. We evaluate our method on two real-world cross-resolution datasets and two newly simulated cross-resolution datasets, and both intra-dataset and cross-dataset evaluations demonstrate the effectiveness and superiority of our method in cross-resolution person ReID. The codes and datasets are available at https://github.com/amateurbo/MRLReID.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Unleashing the Feature Hierarchy Potential: An Efficient Tri-Hybrid Person Search Model
abstract
Person search aims to locate target pedestrians from scene images, involving detection and re-identification. The former seeks to separate the background and focus on the commonality between pedestrians, while the latter aims to identify the target and focus on the difference between pedestrians. To address the paradox of detection and re-identification in search tasks, we propose an efficient Tri-Hybrid person search model utilizing the feature hierarchy design. Our model introduces three feature hybrid models for various feature levels. Before the RoI-Align, we present “Spatial-Channel Hybrid” (SCH) and “Token-Channel Hybrid” (TCH). SCH perceives the boundary frame of pedestrians at multiple scales, thereby enhancing the information disparity between pedestrians and the background and refining the accuracy of the detection frame. TCH uses multi-layer perceptrons (MLP) and blends token and channel features, emphasizing detecting fine-grained semantic information for pedestrians. The interaction of multi-scale perception and fine-grained semantic information enhances the details of detected pedestrians, making them more suitable for similarity measurement in pedestrian matching. After the RoI-Align, we design the “CNN-Transformer Hybrid” to amalgamate global and local features to extract more comprehensive detailed features. Extensive experimental results on CUHK-SYSU and PRW demonstrate the effectiveness of the proposed method over the state-of-the-art performance. Specifically, our method achieves comparable performance on two benchmark datasets, CUHK-SYSU and PRW, with mAP scores of 94.62% and 57.84%, respectively.
Xi Yang 0011, Menghui Tian, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 An Adaptive Region Proposal Network With Progressive Attention Propagation for Tiny Person Detection From UAV Images
abstract
Two-stage detectors, which consist of the multi-scale feature representations and the prediction of region proposal boxes, have been recognized as an effective paradigm for tiny object detection in Unmanned Aerial Vehicle (UAV) images. Although most previous methods primarily concentrated on developing efficient feature fusion strategies within the feature pyramid network (FPN), few studies elaborated on improving the performance of region proposal network (RPN). Conventional RPNs exhibit two key weaknesses in the majority of existing two-stage object detection approaches. Firstly, the quality of proposal boxes generated by the RPN is heavily reliant on rich feature representations extracted from the FPN backbone. Secondly, the fixed number of generated proposal boxes limits adaptability to the distribution of tiny person objects. To mitigate the aforementioned problems, in this paper we propose a novel adaptive region proposal network (ARPN) to improve the quality of the proposal boxes and generate particularly compact yet accurate proposal boxes. On one hand, a progressive attention mechanism is devised to make the ARPN focus more on prospective object regions, where a series of multi-scale front attention modules (FAM) are applied to coarsely filter out most of irrelevant background areas and a group of top-to-bottom back attention modules (BAM) aid the ARPN to finely pinpoint tiny objects of interest in a coarse-to-fine manner. On the other hand, a mini-density map, which is inspired by the philosophy of crowd counting, is elaborately designed to adaptively determine the number of region proposal boxes. This approach significantly reduces redundancy while maintaining high-quality proposal boxes. Extensive experiments verify the superiority of proposed ARPN and show obvious improvement over other competitors in terms of two performance indicators of average precision (AP) and average recall (AR). The code will be available at https://github.com/kbzhang0505/ARPN.
Youjiang Yu, Kaibing Zhang, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Learning Relationship-Enhanced Semantic Graph for Fine-Grained Image-Text Matching
abstract
Image-text matching of natural scenes has been a popular research topic in both computer vision and natural language processing communities. Recently, fine-grained image-text matching has shown its significant advance in inferring the high-level semantic correspondence by aggregating pairwise region-word similarity, but it remains challenging mainly due to insufficient representation of high-order semantic concepts and their explicit connections in one modality as its matched in another modality. To tackle this issue, we propose a relationship-enhanced semantic graph (ReSG) model, which can improve the image-text representations by learning their locally discriminative semantic concepts and then organizing their relationships in a contextual order. To be specific, two tailored graph encoders, visual relationship-enhanced graph (VReG) and textual relationship-enhanced graph (TReG), are respectively exploited to encode the high-level semantic concepts of corresponding instances and their semantic relationships. Meanwhile, the representations of each graph node are optimized by aggregating semantically contextual information to enhance the node-level semantic correspondence. Further, the hard-negative triplet ranking loss, center hinge loss, and positive-negative margin loss are jointly leveraged to learn the fine-grained correspondence between the ReSG representations of image and text, whereby the discriminative cross-modal embeddings can be explicitly obtained to benefit various image-text matching tasks in a more interpretable way. Extensive experiments verify the advantages of the proposed fine-grained graph matching approach, by achieving the state-of-the-art image-text matching results on public benchmark datasets.
Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001
IEEE Trans. Cybern.5
2024 Neighbor Consistency and Global-Local Interaction: A Novel Pseudo-Label Refinement Approach for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) aims at learning discriminative identity features for person retrieval without any annotations. Recent advances accomplish this task by leveraging clustering-based pseudo labels, but these pseudo labels are inevitably noisy, which deteriorates model performance. In this paper, we propose a Neighbour Consistency guided Pseudo Label Refinement (NCPLR) framework, which can be regarded as a transductive form of label propagation under the assumption that the prediction of each example should be similar to its nearest neighbours’. Specifically, the refined label for each training instance can be obtained from the original clustering result and a weighted ensemble of its neighbours’ predictions, with weights determined according to their similarities in the feature space. Furthermore, we also explore building a unified global-local NCPLR mechanism through a global-local label interaction module to achieve mutual label refinement. Such a strategy promotes efficient complementary learning while mitigating some unreliable information, finally improving the quality of the refined pseudo labels for each global-local region. Extensive experimental results demonstrate the effectiveness of the proposed method, showing superior performance to state-of-the-art methods by a large margin. Our source code is released inhttps://github.com/haichuntai/NCPLR-ReID.
De Cheng, Haichun Tai, Nannan Wang 0001, Chaowei Fang, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Universal Heterogeneous Face Analysis via Multi-Domain Feature Disentanglement
abstract
Heterogeneous face analysis is an important and challenge problem in face recognition community, because of the large modality discrepancy between heterogeneous face images. Existing methods either focus on transforming heterogeneous faces into the same style via face synthesis process, or intend to directly recognize heterogeneous face via modality invariant descriptors. However, the tasks of cross modality face synthesis and face recognition share a common purpose, which is to disentangle an inherent explainable representation. To this end, we propose a novel universal heterogenous face analysis method via multi-domain feature disentanglement, which does not need any face domain label. The proposed method explores to disentangle factors of variations of cross modality faces in an unsupervised manner. Then we could translate cross modality faces through modifying semantic factors, and the extracted inherent explainable representation still maintains being discriminative for heterogeneous face recognition. Experimental results on multiple cross modality face databases demonstrate the effectiveness of the proposed method. These experimental results also inspire us that the unsupervised disentangled module could help to analyze the interpretability of heterogenous face representation.
Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Where Deepfakes Gaze at? Spatial-Temporal Gaze Inconsistency Analysis for Video Face Forgery Detection
abstract
With the continuous development of generative models on face generation, how to distinguish the real and fake face has become an important problem for security. Because of the continuous improvement on the detection accuracy by facial physiological signals, video face forgery detection based on facial physiological signal analysis has received more and more attention, which has become an important research branch in the field of face forgery detection. Currently, most of the research on forgery detection based on physiological signal analysis use biometric features such as blinking patterns, head swings, heart rate signals, and lip movements. However, there hasn’t been much exploration on the usage of gaze features in face forgery detection. Through the analysis of gaze directions in face videos, we have observed differences in the distribution of gaze direction pattern between the real and forged videos. Specifically, real videos tend to have more concentrated gaze distribution within a short period of time, while forged videos have more dispersed gaze distributions. In this paper, we present a novel Deepfake gaze analysis method named DFGaze, to explore spatial-temporal gaze inconsistency for video face forgery detection. Our method uses the gaze analysis model (GAM) to analyze the gaze features of face video frames, and then applies a spatial-temporal feature aggregator to realize authenticity classification based on gaze features. In order to better mine the authenticity clues in the videos, we further use the texture analysis model (TAM) and attribute analysis model (AAM) to improve the representation ability of spatial-temporal feature differences between real and forged faces. Extensive experiments show that our method can achieve state-of-the-art performance with the help of gaze analysis. The source code is available at https://github.com/ziminMIAO/DFGaze.
Chunlei Peng, Zimin Miao, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Dual-Adversarial Representation Disentanglement for Visible Infrared Person Re-Identification
abstract
Heterogeneous pedestrian images are captured by visible and infrared cameras with different spectrums, which play an important role in night-time video surveillance. However, visible infrared person re-identification (VI-REID) is still a challenging problem due to the considerable cross-modality discrepancies. To extract modality-invariant features which are discriminative for the person identity, recent studies are inclined to regard modality-specific features as noise and discard them. Actually, the modality-specific characteristics containing background and color information are indispensable for learning modality-shared features. In this paper, we propose a novel Dual-Adversarial Representation Disentanglement (DARD) model to separate modality-specific features from tangled pedestrian representations and effectively learn the robust modality-invariant representations. Specifically, our method employs dual-adversarial learning, incorporating image-level channel exchange and feature-level magnitude change to introduce variations in modality-specific representations. This deliberate perturbation raises the learning difficulty for the model to learn modality-shared features. Simultaneously, to control the changing scope of modality-specific features, bi-constrained noise alleviation is introduced during adversarial learning, keeping the balance of feature generation and adversary. The proposed dual-adversarial learning methodology enhances the robustness against cross-modality visual discrepancy and strengthens the discriminative power of the learned modality-shared representations without introducing additional network parameters. This improvement further elevates the retrieval performance of VI-REID. Extensive experiments with insightful analysis on two cross-modality re-identification datasets verify the effectiveness and superiority of the proposed DARD method.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2024 MMNet: Multi-Collaboration and Multi-Supervision Network for Sequential Deepfake Detection
abstract
Advanced manipulation techniques have provided criminals with opportunities to make social panic or gain illicit profits through the generation of deceptive media, such as forgery face images. In response, various deepfake detection methods have been proposed to assess image authenticity. Sequential deepfake detection, which is an extension of deepfake detection, aims to identify forged facial regions with the correct sequence for recovery. Nonetheless, due to the different combinations of spatial and sequential manipulations, forgery face images exhibit substantial discrepancies that severely impact detection performance. Additionally, the recovery of forged images requires knowledge of the manipulation model to implement inverse transformations, which is difficult to ascertain as relevant techniques are often concealed by attackers. To address these issues, we propose Multi-Collaboration and Multi-Supervision Network (MMNet) that handles various spatial scales and sequential permutations in forgery face images and achieve recovery without requiring knowledge of the corresponding manipulation method. Furthermore, existing evaluation metrics only consider detection accuracy at a single inferring step, without accounting for the matching degree with ground-truth under continuous multiple steps. To overcome this limitation, we propose a novel evaluation metric called Complete Sequence Matching (CSM), which considers the detection accuracy at multiple inferring steps, reflecting the ability to detect integrally forged sequences. Extensive experiments on several typical datasets demonstrate that MMNet achieves state-of-the-art detection performance and independent recovery performance. Code will be available at https://github.com/xarryon/MMNet.
Ruiyang Xia, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Large Pose Face Recognition via Facial Representation Learning
abstract
Overcoming image acquisition perspectives and face pose variations is a key problem in unconstrained face recognition tasks. One of the practical approaches is by reconstructing the face with extreme pose into a version that is more easily recognized by the discriminator, such as a frontal face. Often, existing methods attempt to balance the accuracy of downstream tasks with human visual perception, but ignore the differences in propensity between the two. Besides, large-scale datasets of profile-frontal paired face images are absent, which further hinders the training of models. In this work, we investigate a variety of face reconstruction approaches and propose a very simple, but very effective method to match face images across different scenes, named facial representation learning (FRL). The core idea of FRL is to introduce a representation generator in front of a pre-trained face recognition model, which can extract face representations from arbitrary faces that are more suitable for recognition model discrimination. In particular, the representation generator reconstructs the facial representation by minimising identity differences from the frontal face and adds pixel-level and adversarial constraints to cater for discriminator preferences. Extensive benchmark experiments show that the proposed method not only achieves better performance than state-of-the-art methods, but also can further squeeze the inference potential of existing face recognition models.
Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Quantization Aware Attack: Enhancing Transferable Adversarial Attacks by Model Quantization
abstract
Quantized neural networks (QNNs) have received increasing attention in resource-constrained scenarios due to their exceptional generalizability. However, their robustness against realistic black-box adversarial attacks has not been extensively studied. In this scenario, adversarial transferability is pursued across QNNs with different quantization bitwidths, which particularly involve unknown architectures and defense methods. Previous studies claim that transferability is difficult to achieve across QNNs with different bitwidths on the condition that they share the same architecture. However, we discover that under different architectures, transferability can be largely improved by using a QNN quantized with an extremely low bitwidth as the substitute model. We further improve the attack transferability by proposingquantization aware attack(QAA), which fine-tunes a QNN substitute model with a multiple-bitwidth training objective. In particular, we demonstrate that QAA addresses the two issues that are commonly known to hinder transferability: 1) quantization shifts and 2) gradient misalignments. Extensive experimental results validate the high transferability of the QAA to diverse target models. For instance, when adopting the ResNet-34 substitute model on ImageNet, QAA outperforms the current best attack in attacking standardly trained DNNs, adversarially trained DNNs, and QNNs with varied bitwidths by 4.6% ~ 20.9%, 8.8% ~ 13.4%, and 2.6% ~ 11.8% (absolute), respectively. In addition, QAA is efficient since it only takes one epoch for fine-tuning. In the end, we empirically explain the effectiveness of QAA from the view of the loss landscape. Our code is available at https://github.com/yyl-github-1896/QAA/.
Yulong Yang 0002, Chenhao Lin, Qian Li 0024, Zhengyu Zhao 0001, Haoran Fan, Dawei Zhou 0004, Nannan Wang 0001, Tongliang Liu, Chao Shen 0001
IEEE Trans. Inf. Forensics Secur.7
2024 Neighbor-Guided Pseudo-Label Generation and Refinement for Single-Frame Supervised Temporal Action Localization
abstract
Due to the sparse single-frame annotations, current Single-Frame Temporal Action Localization (SF-TAL) methods generally employ threshold-based pseudo-label generation strategies. However, these approaches suffer from inefficient data utilization, as only parts of unlabeled frames with confidence scores surpassing a predefined threshold are selected for training. Moreover, the variability of single-frame annotations and unreliable model predictions introduce pseudo-label noise. To address these challenges, we propose two strategies by using the relationship of the video segments with their neighbors': 1) temporal neighbor-guided soft pseudo-label generation (TNPG); and 2) semantic neighbor-guided pseudo-label refinement (SNPR). TNPG utilizes a local-global self-attention mechanism in a transformer encoder to capture temporal neighbor information while focusing on the whole video. Then the generated self-attention map is multiplied by the network predictions to propagate information between labeled and unlabeled frames, and produce soft pseudo-label for all segments. Despite this, label noise persists due to unreliable model predictions. To mitigate this, SNPR refines pseudo-labels based on the assumption that predictions should resemble their semantic nearest neighbors'. Specifically, we search for semantic nearest neighbors of each video segment by cosine similarity in the feature space. Then the refined soft pseudo-labels can be obtained by a weight combination of the original pseudo-label and the semantic nearest neighbors'. Finally, the model can be trained with the refined pseudo-labels, and the performance has been greatly improved. Comprehensive experimental results on different benchmarks show that we achieve state-of-the-art performances on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets.
Guozhang Li, De Cheng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2024 Semi-Supervised Learning With Heterogeneous Distribution Consistency for Visible Infrared Person Re-Identification
abstract
Visible infrared person re-identification (VI-ReID) exposes considerable challenges because of the modality gaps between the person images captured by daytime visible cameras and nighttime infrared cameras. Several fully-supervised VI-ReID methods have improved the performance with extensive labeled heterogeneous images. However, the identity of the person is difficult to obtain in real-world situations, especially at night. Limited known identities and large modality discrepancies impede the effectiveness of the model to a great extent. In this paper, we propose a novel Semi-Supervised Learning framework with Heterogeneous Distribution Consistency (HDC-SSL) for VI-ReID. Specifically, through investigating the confidence distribution of heterogeneous images, we introduce a Gaussian Mixture Model-based Pseudo Labeling (GMM-PL) method, which adaptively adjusts different thresholds for each modality to label the identity. Moreover, to facilitate the representation learning of unutilized data whose prediction is lower than the threshold, Modality Consistency Regularization (MCR) is proposed to ensure the prediction consistency of the cross-modality pedestrian images and handle the modality variance. Extensive experiments with different label settings on two VI-ReID datasets demonstrate the effectiveness of our method. Particularly, HDC-SSL achieves competitive performance with state-of-the-art fully-supervised VI-ReID methods on RegDB dataset with only 1 visible label and 1 infrared label per class.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2024 Inspector for Face Forgery Detection: Defending Against Adversarial Attacks From Coarse to Fine
abstract
The emergence of face forgery has raised global concerns on social security, thereby facilitating the research on automatic forgery detection. Although current forgery detectors have demonstrated promising performance in determining authenticity, their susceptibility to adversarial perturbations remains insufficiently addressed. Given the nuanced discrepancies between real and fake instances are essential in forgery detection, previous defensive paradigms based on input processing and adversarial training tend to disrupt these discrepancies. For the detectors, the learning difficulty is thus increased, and the natural accuracy is dramatically decreased. To achieve adversarial defense without changing the instances as well as the detectors, a novel defensive paradigm called Inspector is designed specifically for face forgery detectors. Specifically, Inspector defends against adversarial attacks in a coarse-to-fine manner. In the coarse defense stage, adversarial instances with evident perturbations are directly identified and filtered out. Subsequently, in the fine defense stage, the threats from adversarial instances with imperceptible perturbations are further detected and eliminated. Experimental results across different types of face forgery datasets and detectors demonstrate that our method achieves state-of-the-art performances against various types of adversarial perturbations while better preserving natural accuracy. Code is available on https://github.com/xarryon/Inspector.
Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.6
2024 Adapting Few-Shot Classification via In-Process Defense
abstract
Most few-shot learning methods employ either adaptive approaches or parameter amortization techniques. However, their reliance on pre-trained models presents a significant vulnerability. When an attacker's trigger activates a hidden backdoor, it may result in the misclassification of images, profoundly affecting the model's performance. In our research, we explore adaptive defenses against backdoor attacks for few-shot learning. We introduce a specialized stochastic process tailored to task characteristics that safeguards the classification model against attack-induced incorrect feature extraction. This process functions during forward propagation and is thus termed an "in-process defense." Our method employs an adaptive strategy, effectively generating task-level representations, enabling rapid adaptation to pre-trained models, and proving effective in few-shot classification scenarios for countering backdoor attacks. We apply latent stochastic processes to approximate task distributions and derive task-level representations from the support set. This task-level representation guides feature extraction, leading to backdoor trigger mismatching and forming the foundation of our parameter defense strategy. Benchmark tests on Meta-Dataset reveal that our approach not only withstands backdoor attacks but also shows an improved adaptation in addressing few-shot classification tasks.
Xi Yang 0011, Dechen Kong, Ren Lin, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2024 Image-Level Adaptive Adversarial Ranking for Person Re-Identification
abstract
The potential vulnerability of deep neural networks and the complexity of pedestrian images, greatly limits the application of person re-identification techniques in the field of smart security. Current attack methods often focus on generating carefully crafted adversarial samples or only disrupting the metric distances between targets and similar pedestrians. However, both aspects are crucial for evaluating the security of methods adapted for person re-identification tasks. For this reason, we propose an image-level adaptive adversarial ranking method that comprehensively considers two aspects to adapt to changes in pedestrians in the real world and effectively evaluate the robustness of models in adversarial environments. To generate more refined adversarial samples, our image representation enhancement module leverages channel-wise information entropy, assigning varying weights to different channels to produce images with richer information content, along with a generative adversarial network to create adversarial samples. Subsequently, for adaptive perturbation of ranking, the adaptive weight confusion ranking loss is presented to calculate the weights of distances between positive or negative samples and query samples. It endeavors to push positive samples away from query samples and bring negative samples closer, thereby interfering with the ranking of system. Notably, this method requires no additional hyperparameter tuning or extra data training, making it an adaptive attack strategy. Experimental results on large-scale datasets such as Market1501, CUHK03, and DukeMTMC demonstrate the effectiveness of our method in attacking ReID systems.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2024 Protecting Prostate Cancer Classification From Rectal Artifacts via Targeted Adversarial Training
abstract
Magnetic resonance imaging (MRI)-based deep neural networks (DNN) have been widely developed to perform prostate cancer (PCa) classification. However, in real-world clinical situations, prostate MRIs can be easily impacted by rectal artifacts, which have been found to lead to incorrect PCa classification. Existing DNN-based methods typically do not consider the interference of rectal artifacts on PCa classification, and do not design specific strategy to address this problem. In this study, we proposed a novel Targeted adversarial training with Proprietary Adversarial Samples (TPAS) strategy to defend the PCa classification model against the influence of rectal artifacts. Specifically, based on clinical prior knowledge, we generated proprietary adversarial samples with rectal artifact-pattern adversarial noise, which can severely mislead PCa classification models optimized by the ordinary training strategy. We then jointly exploited the generated proprietary adversarial samples and original samples to train the models. To demonstrate the effectiveness of our strategy, we conducted analytical experiments on multiple PCa classification models. Compared with ordinary training strategy, TPAS can effectively improve the single- and multi-parametric PCa classification at patient, slice and lesion level, and bring substantial gains to recent advanced models. In conclusion, TPAS strategy can be identified as a valuable way to mitigate the influence of rectal artifacts on deep learning models for PCa classification.
Lei Hu 0002, Dawei Zhou 0004, Cheng Lu 0001, Chu Han, Zhenwei Shi 0002, Qikui Zhu, Xinbo Gao 0001, Nannan Wang 0001, Zaiyi Liu
IEEE J. Biomed. Health Informatics9
2024 Continual All-in-One Adverse Weather Removal With Knowledge Replay on a Unified Network Structure
abstract
In real-world applications, image degeneration caused by adverse weather is always complex and changes with different weather conditions from days and seasons. Systems in real-world environments constantly encounter adverse weather conditions that are not previously observed. Therefore, it practically requires adverse weather removal models to continually learn from incrementally collected data reflecting various degeneration types. Existing adverse weather removal approaches, for either single or multiple adverse weathers, are mainly designed for a static learning paradigm, which assumes that the data of all types of degenerations to handle can be finely collected at one time before a single-phase learning process. They thus cannot directly handle the incremental learning requirements. To address this issue, we made the earliest effort to investigate the continual all-in-one adverse weather removal task, in a setting closer to real-world applications. Specifically, we develop a novel continual learning framework with effective knowledge replay (KR) on a unified network structure. Equipped with a principal component projection and an effective knowledge distillation mechanism, the proposed KR techniques are tailored for the all-in-one weather removal task. It considers the characteristics of the image restoration task with multiple degenerations in continual learning, and the knowledge for different degenerations can be shared and accumulated in the unified network structure. Extensive experimental results demonstrate the effectiveness of the proposed method to deal with this challenging task, which performs competitively to existing dedicated or joint training image restoration methods. Our code is available athttps://github.com/xiaojihh/CL_all-in-one.
De Cheng, Yanling Ji, Dong Gong, Yan Li 0125, Nannan Wang 0001, Junwei Han 0001, Dingwen Zhang
IEEE Trans. Multim.5
2024 Progressive Negative Enhancing Contrastive Learning for Image Dehazing and Beyond
abstract
Image dehazing is a pivotal preliminary step in the advancement of robust intelligent surveillance system. However, it is an extremely challenging ill-posed problem, as it faces severe information degradation when accurately restoring the clean image from its haze-polluted counterpart. This paper proposes a novel Progressive Negative Enhancing (PNE) contrastive learning mechanism to fully exploit various types of negative information, thereby facilitating the traditional positive-oriented objective function for image dehazing. The proposed method can progressively update the negative samples during model training, to steadily squeeze the restored image towards its desired clean target from various directions. Furthermore, considering the image dehazing task as a many-to-one feature mapping problem, we also make an early effort to enhance the robustness of the dehazing model under variational haze densities. Specifically, a novel density-variational dehazing network is proposed to be optimized under the consistency-regularized framework using the proposed PNE learning mechanism. The consistency regularization ensures consistent output given multi-level degraded hazy images, thereby significantly enhancing the robustness of the model in dealing with various hazy scenarios. Extensive experiments demonstrate that the proposed method exhibits superior performance over existing state-of-the-art methods. It achieves average PSNR boosts of 0.60dB, 0.28dB and 0.82dB on dehazing, deraining and desnowing tasks, respectively. The source code is available athttps://github.com/YanLi-LY/PNE-Net.
De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Jiande Sun 0001, Xinbo Gao 0001
IEEE Trans. Multim.4
2024 Towards Specific Domain Prompt Learning via Improved Text Label Optimization
abstract
Prompt learning has emerged as a thriving parameter-efficient fine-tuning technique for adapting pre-trained vision-language models (VLMs) to various downstream tasks. However, existing prompt learning approaches still exhibit limited capability for adapting foundational VLMs to specific domains that require specialized and expert-level knowledge. Since this kind of specific knowledge is primarily embedded in the pre-defined text labels, we infer that foundational VLMs cannot directly interpret semantic meaningful information from these specific text labels, which causes the above limitation. From this perspective, this paper additionally models text labels with learnable tokens and casts this operation into traditional prompt learning framework. By optimizing label tokens, semantic meaningful text labels are automatically learned for each class. Nevertheless, directly optimizing text label still remains two critical problems, i.e., insufficient optimization and biased optimization. We further address these problems by proposing Modality Interaction Text Label Optimization (MITLOp) and Color-based Consistency Augmentation (CCAug) respectively, thereby effectively improving the quality of the optimized text labels. Extensive experiments indicate that our proposed method achieves significant improvements in VLM adaptation on specific domains.
Liangchen Liu 0001, Nannan Wang 0001, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Multim.2
2024 Hierarchical Forgery Classifier on Multi-Modality Face Forgery Clues
abstract
Face forgery detection plays an important role in personal privacy and social security. With the development of adversarial generative models, high-quality forgery images become more and more indistinguishable from real to humans. Existing methods always regard as forgery detection task as the common binary or multi-label classification, and ignore exploring diverse multi-modality forgery image types, e.g. visible light spectrum and near-infrared scenarios. In this article, we propose a novelHierarchicalForgeryClassifier forMulti-modalityFaceForgeryDetection(HFC-MFFD), which could effectively learn robust patches-based hybrid domain representation to enhance forgery authentication in multiple modality scenarios. The local hybrid domain representation is designed to explore strong discriminative forgery clues both in the image and frequency domain with the intra-attention mechanism. Furthermore, the specific hierarchical face forgery classifier is designed through the authenticity feedback strategy to integrate diverse discriminative clues. Experimental results on representative multi-modality face forgery datasets demonstrate the superior performance of the proposed HFC-MFFD compared with state-of-the-art algorithms.
Decheng Liu, Zeyang Zheng, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2024 Disguised Heterogeneous Face Generation With Iterative-Adversarial Style Unification
abstract
Heterogeneous face recognition (HFR), which refers to matching face images with different modalities, is essential to public safety. Although HFR has made promising progress in recent years, disguised faces in HFR scenarios still remain a major challenge for the following reasons. First, most existing HFR methods focus on traditional scenarios without disguised accessories, and the performance degrades when dealing directly with disguised faces. Second, there is a need for disguised heterogeneous face datasets, which is essential for developing the related research community. Third, colorful accessories are distinct from heterogeneous face images in terms of their modalities, and their direct combination results in style inconsistency and poor quality. Therefore, we propose a disguised heterogeneous face generation method based on an iterative-adversarial style unification framework. Our approach aims to gradually learn frame textures to detail textures in multiple confrontation iterations, resulting in style unification for disguised accessories and heterogeneous faces. We also construct a disguised heterogeneous face dataset, which contains a disguised NIR-VIS subset and a disguised sketch-photo subset. Moreover, we provide benchmark evaluations conducted on our proposed dataset with face recognition and image quality assessment, demonstrating the superiority of our method over direct addition and two representative disguised face generation techniques.
Chunlei Peng, Zimo Kong, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.4
2024 Cooperative Separation of Modality Shared-Specific Features for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging task because the different imaging principles of visible and infrared images bring about huge modality discrepancy. Existing methods primarily address this issue by generating intermediate images to align modality features and establish connections between the visible and infrared modalities. However, the quality of these generated images is often unstable, limiting the effectiveness of such approaches. To overcome this limitation, we propose a novel method called modality shared-specific features cooperative separation. It consists of two key modules: the saliency response module and the cooperative separation module, aimed at alleviating the modality gap. The saliency response module incorporates a location attention mechanism and local features to construct contextual connections and extract local saliency information. Then, the cooperative separation module employs a more concise dual-MLPs as generator to effectively separate shared-specific features. Additionally, we introduce a shared feature refinement mechanism in both the generator and discriminator. By coordinating the shared-specific features, our method achieves secondary separation and extracts purer modality-shared features without specific information. Extensive experiments conducted on the SYSU-MM01 and RegDB public datasets demonstrate that our proposed method performs excellently in VI-ReID.
Xi Yang 0011, Wenjiao Dong, Meijie Li, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2024 SSRR: Structural Semantic Representation Reconstruction for Visible-Infrared Person Re-Identification
abstract
Visible-infrared Person Re-identification (VI-ReID) aims to retrieve the images of pedestrian with the same identity from different modalities and cameras given a pedestrian image. To reduce modality discrepancy, existing methods often perform hard partitioning to mine more detail. However, these methods employ only uniform partitioning, without considering pedestrian structure, and lose a lot of pedestrian semantic information. To this end, this paper proposes a structural semantic representation reconstruction (SSRR) method to capture pedestrian semantic information by focusing on pedestrian structure. Specifically, based on the fine-grained features obtained by hard partitioning, we carry out structural reconstruction to obtain the reconstructed features containing semantic information. By adopting the direct link reconstruction structure, the reciprocal learning of fine-grained features and semantic features is ensured. Semantic features are reconstructed based on fine-grained features, and semantic information is beneficial to fine-grained features to better capture pedestrian-related details. In addition, local consistency loss is introduced to ensure the consistency of fine-grained features in the same component location, further enhancing the discriminant of the learned reconstructed representation. Extensive experiments confirm the superiority of our method on two public datasets SYSU-MM01 and RegDB.
Xi Yang 0011, Menghui Tian, Meijie Li, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.6
2024 STFE: A Comprehensive Video-Based Person Re-Identification Network Based on Spatio-Temporal Feature Enhancement
abstract
Video-based person re-identification (Re-ID) is designed to retrieve target pedestrians in video sequences under non-overlapping cameras. At present, mainstream approaches post-process the feature map extracted by the convolutional neural network backbone to obtain a global representation or a fine-grained local representation for higher accuracy. However, they still suffer from challenges, such as information loss for global-based methods and spatio-temporal feature fragmentation for local-based methods. To alleviate these problems, this article proposes a Spatio-Temporal Feature Enhancement (STFE) network from a spatio-temporal comprehensive perspective, combining the advantages of the above methods to obtain more comprehensive information from video tracklets. STFE consists of two main modules: Feature Space Projection Module (FSPM) and Global Low-frequency Enhancement Module (GLEM). FSPM mathematically converts continuous video information into a discrete feature space and selectively retains more useful information, thus avoiding spatio-temporal information loss. Meanwhile, FSPM applies global features instead of dividing feature maps spatially, thereby avoiding spatio-temporal feature fragmentation. In addition, GLEM which is based on transformer, acts as a broadband low-pass filter to mine richer global comprehensive information. Finally, by combining FSPM with GLEM, STFE can obtain spatio-temporal comprehensive video representation. Extensive experiments were conducted on two widely-used video Re-ID datasets. The experimental results verify our idea and demonstrate the effectiveness of the proposed STFE with 95.5% Rank-1 accuracy on MARS benchmarks, which surpasses previousstate-of-the-artsby a large margin of +4%.
Xi Yang 0011, Liangchen Liu 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.4
2024 Elaborate Teacher: Improved Semi-Supervised Object Detection With Rich Image Exploiting
abstract
Semi-Supervised Object Detection (SSOD) has shown remarkable results by leveraging image pairs with a teacher-student framework. An excellent strong augmentation method can generate richer images and alleviate the influence of noise in pseudo-labels. However, existing data augmentation methods for SSOD do not consider instance-level information, thus, they cannot make full use of unlabeled data. Besides, the current teacher-student framework in SSOD solely relies on pseudo-labeling techniques, which may disregard some uncertain information. In this article, we introduce a new method called Elaborate Teacher which generates and exploits image pairs in a more refined manner. To enrich strongly augmented images, a novel data augmentation method called Information-Aware Mixup Representation (IAMR) is proposed. IAMR utilizes the teacher model's predictions as prior information and considers instance-level information, which can be seamlessly integrated with existing SSOD data augmentation methods. Furthermore, to fully exploit the information in unlabeled data, we propose the Enhanced Scale Consistency Regularization (ESCR), which considers the consistency from both semantic space and feature space. Elaborate Teacher introduces a fresh data augmentation method, complemented by consistency regularization, which boosts the performance of semi-supervised object detectors. Extensive experiments on thePASCAL VOCandMS-COCOdatasets demonstrate the effectiveness of our method in leveraging unlabeled image information. Our method consistently outperforms the baseline method and improves mAP by 11.6% and 9.0% relative to the supervised baseline method when using 5% and 10% of labeled data onMS-COCO, respectively.
Xi Yang 0011, Qiubai Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2024 Toward Pixel-Level Precision for Binary Super-Resolution With Mixed Binary Representation
abstract
Binary neural network (BNN) is an effective method for reducing model computational and memory cost, which has achieved much progress in the super-resolution (SR) field. However, there is still a noticeable performance gap between a binary SR network and its full-precision counterpart. Considering that the information density in quantization features is far lower than full-precision features, we aim to improve the precision of quantization features to produce rich-enough output activations for SR task. First, we make several observations that a multibit value could be approximated by multiple 1-bit values, and the computation power of binary convolution could be improved by approximating the multibit convolution process. Then, we propose a mixed binary representation set to approximate multibit activations, which is effective in compensating the quantization precision loss. Finally, we present a new precision-driven binary convolution (PDBC) module, which increases the convolution precision and protects image detail information without extra computation. Compared with normal binary convolution, our method could largely reduce the information loss caused by binarization. In experiments, our methods consistently show superior performance over the baseline models and can surpass state-of-the-art methods in terms of peak signal to noise ratio (PSNR) and visual quality.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Weakly Supervised Temporal Action Localization With Bidirectional Semantic Consistency Constraint
abstract
Weakly supervised temporal action localization (WTAL) aims to classify and localize temporal boundaries of actions for the video, given only video-level category labels in the training datasets. Due to the lack of boundary information during training, existing approaches formulate WTAL as a classification problem, i.e., generating the temporal class activation map (T-CAM) for localization. However, with only classification loss, the model would be suboptimized, i.e., the action-related scenes are enough to distinguish different class labels. Regarding other actions in the action-related scene (i.e., the scene same as positive actions) as co-scene actions, this suboptimized model would misclassify the co-scene actions as positive actions. To address this misclassification, we propose a simple yet efficient method, named bidirectional semantic consistency constraint (Bi-SCC), to discriminate the positive actions from co-scene actions. The proposed Bi-SCC first adopts a temporal context augmentation to generate an augmented video that breaks the correlation between positive actions and their co-scene actions in the inter-video. Then, a semantic consistency constraint (SCC) is used to enforce the predictions of the original video and augmented video to be consistent, hence suppressing the co-scene actions. However, we find that this augmented video would destroy the original temporal context. Simply applying the consistency constraint would affect the completeness of localized positive actions. Hence, we boost the SCC in a bidirectional way to suppress co-scene actions while ensuring the integrity of positive actions, by cross-supervising the original and augmented videos. Finally, our proposed Bi-SCC can be applied to current WTAL approaches and improve their performance. Experimental results show that our approach outperforms the state-of-the-art methods on THUMOS14 and ActivityNet. The code is available at https://github.com/lgzlIlIlI/BiSCC.
Guozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Local Means Binary Networks for Image Super-Resolution
abstract
The success of modern single image super-resolution (SISR) algorithms is inspired by the development of deep convolutional neural networks (CNNs). However, these CNN-based methods require considerable computation and complexity, making it impossible for these methods to perform real-time calculations in edge devices. Thus, lightweight model design has become a development trend in the super-resolution field, including pruning, quantization, and other methods. The 1-bit quantization is an extreme lightweight method which can reduce the calculation amount of the model in an extreme manner and is friendly to hardware such as edge devices. Most existing binary quantization approaches lead to a large information loss during forward propagation, especially in detailed color information (e.g., edge, texture, and contrast). The loss of color information makes modern binary methods unsuitable for SISR tasks. We think the loss occurs because these methods typically utilize a uniform threshold to quantize the weights and activations. Thus, in this article, we thoroughly analyze the difference between normal classification tasks and SISR tasks, and present a binarization scheme based on local means. The proposed method can maintain more detailed information in feature maps using dynamic thresholds during quantization. Specifically, each value in the full precision activations has a corresponding threshold during the quantization process, and those thresholds are determined by the full precision values of the surroundings. In addition, a gradient approximator is introduced to adaptively optimize the gradient for updating binary weights. We then verify the effectiveness of our method for training binary networks on several SISR benchmarks including VDSR and SRResNet. Experimental results show that the proposed method can outperform the state-of-the-art algorithms to obtain binary networks for image super-resolution with better peak signal-to-noise ratio (PSNR) values and visual quality.
Nannan Wang 0001, Jingwei Xin, Jie Li 0001, Xinbo Gao 0001, Kai Han 0002, Yunhe Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 PMSGAN: Parallel Multistage GANs for Face Image Translation
abstract
In this article, we address the face image translation task, which aims to translate a face image of a source domain to a target domain. Although significant progress has been made by recent studies, face image translation is still a challenging task because it has more strict requirements for texture details: even a few artifacts will greatly affect the impression of generated face images. Targeting to synthesize high-quality face images with admirable visual appearance, we revisit the coarse-to-fine strategy and propose a novel p arallel m ultistage architecture on the basis of g enerative a dversarial n etworks (PMSGAN). More specifically, PMSGAN progressively learns the translation function by disintegrating the general synthesis process into multiple parallel stages that take images with gradually decreasing spatial resolution as inputs. To prompt the information exchange between various stages, a cross-stage atrous spatial pyramid (CSASP) structure is specially designed to receive and fuse the contextual information from other stages. At the end of the parallel model, we introduce a novel attention-based module that leverages multistage decoded outputs as in situ supervised attention to refine the final activations and yield the target image. Extensive experiments on several face image translation benchmarks show that PMSGAN performs considerably better than state-of-the-art approaches.
Changcheng Liang, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Toward Learning Joint Inference Tasks for IASS-MTS Using Dual Attention Memory With Stochastic Generative Imputation
abstract
Irregularly, asynchronously and sparsely sampled multivariate time series (IASS-MTS) are characterized by sparse and uneven time intervals and nonsynchronous sampling rates, posing significant challenges for machine learning models to learn complex relationships within and beyond IASS-MTS to support various inference tasks. The existing methods typically either focus solely on single-task forecasting or simply concatenate them through a separate preprocessing imputation procedure for the subsequent classification application. However, these methods often ignore valuable annotated labels or fail to discover meaningful patterns from unlabeled data. Moreover, the approach of separate prefilling may introduce errors due to the noise in raw records, and thus degrade the downstream prediction performance. To overcome these challenges, we propose the time-aware dual attention and memory-augmented network (DAMA) with stochastic generative imputation (SGI). Our model constructs a joint task learning architecture that unifies imputation and classification tasks collaboratively. First, we design a new time-aware DAMA that accounts for irregular sampling rates, inherent data nonalignment, and sparse values in IASS-MTS data. The proposed network integrates both attention and memory to effectively analyze complex interactions within and across IASS-MTS for the classification task. Second, we develop the stochastic generative imputation (SGI) network that uses auxiliary information from sequence data for inferring the time series missing observations. By balancing joint tasks, our model facilitates interaction between them, leading to improved performance on both classification and imputation tasks. Third, we evaluate our model on real-world datasets and demonstrate its superior performance in terms of imputation accuracy and classification results, outperforming the baselines.
Zhen Wang 0037, Yang Zhang 0042, Nannan Wang 0001, Mohamed Jaward Bah, Ke Li 0044, Ji Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Address the Unseen Relationships: Attribute Correlations in Text Attribute Person Search
abstract
Text attribute person search aims to identify the particular pedestrian by textual attribute information. Compared to person re- identification tasks which requires imagery samples as its query, text attribute person search is more useful under the circumstance where only witness is available. Most existing text attribute person search methods focus on improving the matching correlation and alignments by learning better representations of person-attribute instance pairs, with few consideration of the latent correlations between attributes. In this work, we propose a graph convolutional network (GCN) and pseudo-label-based text attribute person search method. Concretely, the model directly constructs the attribute correlations by label co- occurrence probability, in which the nodes are represented by attribute embedding and edges are by the filtered correlation matrix of attribute labels. In order to obtain better representations, we combine the cross-attention module (CAM) and the GCN. Furthermore, to address the unseen attribute relationships, we update the edge information through the instances through testing set with high predicted probability thus to better adapt the attribute distribution. Extensive experiments illustrate that our model outperforms the existing state-of-the-art methods on publicly available person search benchmarks: Market-1501 and PETA.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Cross-Modality Person Re-identification with Memory-Based Contrastive Embedding
abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve the person images of the same identity from the RGB to infrared image space, which is very important for real-world surveillance system. In practice, VI-ReID is more challenging due to the heterogeneous modality discrepancy, which further aggravates the challenges of traditional single-modality person ReID problem, i.e., inter-class confusion and intra-class variations. In this paper, we propose an aggregated memory-based cross-modality deep metric learning framework, which benefits from the increasing number of learned modality-aware and modality-agnostic centroid proxies for cluster contrast and mutual information learning. Furthermore, to suppress the modality discrepancy, the proposed cross-modality alignment objective simultaneously utilizes both historical and up-to-date learned cluster proxies for enhanced cross-modality association. Such training mechanism helps to obtain hard positive references through increased diversity of learned cluster proxies, and finally achieves stronger ``pulling close'' effect between cross-modality image features. Extensive experiment results demonstrate the effectiveness of the proposed method, surpassing state-of-the-art works significantly by a large margin on the commonly used VI-ReID datasets.
De Cheng, Nannan Wang 0001, Zhen Wang 0037, Xiaoyu Wang 0002, Xinbo Gao 0001
AAAI3
2023 Masked and Adaptive Transformer for Exemplar Based Image Translation
abstract
We present a novel framework for exemplar based image translation. Recent advanced methods for this task mainly focus on establishing cross-domain semantic correspondence, which sequentially dominates image generation in the manner of local style control. Unfortunately, cross-domain semantic matching is challenging; and matching errors ultimately degrade the quality of generated images. To overcome this challenge, we improve the accuracy of matching on the one hand, and diminish the role of matching in image generation on the other hand. To achieve the former, we propose a masked and adaptive transformer (MAT) for learning accurate cross-domain correspondence, and executing context-aware feature augmentation. To achieve the latter, we use source features of the input and global style codes of the exemplar, as sup-plementary information, for decoding an image. Besides, we devise a novel contrastive style learning method, for acquire quality-discriminative style representations, which in turn benefit high-quality image generation. Experimen-tal results show that our method, dubbed MATEBIT, performs considerably better than state-of-the-art methods, in diverse image translation tasks. The codes are available at https://github.com/AiArt-HDU/MATEBIT.
Fei Gao 0006, Nannan Wang 0001, Gang Xu 0001
CVPR5
2023 Boosting Weakly-Supervised Temporal Action Localization with Text Information
abstract
Due to the lack of temporal annotation, current Weakly-supervised Temporal Action Localization (WTAL) methods are generally stuck into over-complete or incomplete localization. In this paper, we aim to leverage the text information to boost WTAL from two aspects, i.e., (a) the discriminative objective to enlarge the inter-class difference, thus reducing the over-complete; (b) the generative objective to enhance the intra-class integrity, thus finding more complete temporal boundaries. For the discriminative objective, we propose a Text-Segment Mining (TSM) mechanism, which constructs a text description based on the action class label, and regards the text as the query to mine all class-related segments. Without the temporal annotation of actions, TSM compares the text query with the entire videos across the dataset to mine the best matching segments while ignoring irrelevant ones. Due to the shared sub-actions in different categories of videos, merely applying TSM is too strict to neglect the semantic-related segments, which results in incomplete localization. We further introduce a generative objective named Video-text Language Completion (VLC), which focuses on all semantic-related segments from videos to complete the text sentence. We achieve the state-of-the-art performance on THUMOS14 and ActivityNetl.3. Surprisingly, we also find our proposed method can be seamlessly applied to existing methods, and improve their performances with a clear margin. The code is available at https://github.com/lgzlIlIlI/Boosting-WTAL.
Guozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
CVPR4
2023 NAR-Former: Neural Architecture Representation Learning Towards Holistic Attributes Prediction
abstract
With the wide and deep adoption of deep learning models in real applications, there is an increasing need to model and learn the representations of the neural networks themselves. These models can be used to estimate attributes of different neural network architectures such as the accuracy and latency, without running the actual training or inference tasks. In this paper, we propose a neural architecture representation model that can be used to estimate these attributes holistically. Specifically, we first propose a simple and effective tokenizer to encode both the operation and topology information of a neural network into a single sequence. Then, we design a multi-stage fusion transformer to build a compact vector representation from the converted sequence. For efficient model training, we further propose an information flow consistency augmentation and correspondingly design an architecture consistency loss, which brings more benefits with less augmentation samples compared with previous random augmentation strategies. Experiment results on NAS-Bench-101, NAS-Bench-201, DARTS search space and NNLQP show that our proposed framework can be used to predict the aforementioned latency and accuracy attributes of both cell architectures and whole deep neural networks, and achieves promising performance. Code is available at https://github.com/yuny220/NAR-Former.
Yun Yi, Haokui Zhang, Wenze Hu, Nannan Wang 0001, Xiaoyu Wang 0002
CVPR4
2023 Human-Inspired Facial Sketch Synthesis with Dynamic Adaptation
abstract
Facial sketch synthesis (FSS) aims to generate a vivid sketch portrait from a given facial photo. Existing FSS methods merely rely on 2D representations of facial semantic or appearance. However, professional human artists usually use outlines or shadings to covey 3D geometry. Thus facial 3D geometry (e.g. depth map) is extremely important for FSS. Besides, different artists may use diverse drawing techniques and create multiple styles of sketches; but the style is globally consistent in a sketch. Inspired by such observations, in this paper, we propose a novel Human-Inspired Dynamic Adaptation (HIDA) method. Specially, we propose to dynamically modulate neuron activations based on a joint consideration of both facial 3D geometry and 2D appearance, as well as globally consistent style control. Besides, we use deformable convolutions at coarse-scales to align deep features, for generating abstract and distinct outlines. Experiments show that HIDA can generate high-quality sketches in multiple styles, and significantly outperforms previous methods, over a large range of challenging faces. Besides, HIDA allows precise style control of the synthesized sketch, and generalizes well to natural scenes and other artistic styles. Our code and results have been released online at: https://github.com/AiArt-HDU/HIDA.
Fei Gao 0006, Nannan Wang 0001
ICCV4
2023 Hiding Visual Information via Obfuscating Adversarial Perturbations
abstract
Growing leakage and misuse of visual information raise security and privacy concerns, which promotes the development of information protection. Existing adversarial perturbations-based methods mainly focus on the de-identification against deep learning models. However, the inherent visual information of the data has not been well protected. In this work, inspired by the Type-I adversarial attack, we propose an Adversarial Visual Information Hiding (AVIH) method to protect the visual privacy of data. Specifically, the method generates obfuscating adversarial perturbations to obscure the visual information of the data. Meanwhile, it maintains the hidden objectives to be correctly predicted by models. In addition, our method does not modify the parameters of the applied model, which makes it flexible for different scenarios. Experimental results on the recognition and classification tasks demonstrate that the proposed method can effectively hide visual information and hardly affect the performances of models. The code is available at https://github.com/suzhigangssz/AVIH.
Zhigang Su, Dawei Zhou 0004, Nannan Wang 0001, Decheng Liu, Zhen Wang 0037, Xinbo Gao 0001
ICCV3
2023 All-to-key Attention for Arbitrary Style Transfer
abstract
Attention-based arbitrary style transfer studies have shown promising performance in synthesizing vivid local style details. They typically use the all-to-all attention mechanism—each position of content features is fully matched to all positions of style features. However, all-to-all attention tends to generate distorted style patterns and has quadratic complexity, limiting the effectiveness and efficiency of arbitrary style transfer. In this paper, we propose a novel all-to-key attention mechanism—each position of content features is matched to stable key positions of style features—that is more in line with the characteristics of style transfer. Specifically, it integrates two newly proposed attention forms: distributed and progressive attention. Distributed attention assigns attention to key style representations that depict the style distribution of local regions; Progressive attention pays attention from coarse-grained regions to fine-grained key positions. The resultant module, dubbed StyA2K, shows extraordinary performance in preserving the semantic structure and rendering consistent style patterns. Qualitative and quantitative comparisons with state-of-the-art methods demonstrate the superior performance of our approach. Codes and models are available on https://github.com/LearningHx/StyA2K.
Mingrui Zhu, Xiao He 0014, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
ICCV3
2023 Phase-aware Adversarial Defense for Improving Adversarial Robustness
abstract
Deep neural networks have been found to be vulnerable to adversarial noise. Recent works show that exploring the impact of adversarial noise on intrinsic components of data can help improve adversarial robustness. However, the pattern closely related to human perception has not been deeply studied. In this paper, inspired by the cognitive science, we investigate the interference of adversarial noise from the perspective of image phase, and find ordinarily-trained models lack enough robustness against phase-level perturbations. Motivated by this, we propose a joint adversarial defense method: a *phase-level adversarial training mechanism* to enhance the adversarial robustness on the phase pattern; an *amplitude-based pre-processing operation* to mitigate the adversarial perturbation in the amplitude pattern. Experimental results show that the proposed method can significantly improve the robust accuracy against multiple attacks and even adaptive attacks. In addition, ablation studies demonstrate the effectiveness of our defense strategy.
Dawei Zhou 0004, Nannan Wang 0001, Xinbo Gao 0001, Tongliang Liu
ICML2
2023 Eliminating Adversarial Noise via Information Discard and Robust Representation Restoration
abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Denoising model-based defense is a major protection strategy. However, denoising models may fail and induce negative effects in fully white-box scenarios. In this work, we start from the latent inherent properties of adversarial samples to break the limitations. Unlike solely learning a mapping from adversarial samples to natural samples, we aim to achieve denoising by destroying the spatial characteristics of adversarial noise and preserving the robust features of natural information. Motivated by this, we propose a defense based on information discard and robust representation restoration. Our method utilize complementary masks to disrupt adversarial noise and guided denoising models to restore robust-predictive representations from masked samples. Experimental results show that our method has competitive performance against white-box attacks and effectively reverses the negative effect of denoising models.
Dawei Zhou 0004, Nannan Wang 0001, Decheng Liu, Xinbo Gao 0001, Tongliang Liu
ICML3
2023 Semantic-Aware Generation of Multi-View Portrait Drawings
abstract
Neural radiance fields (NeRF) based methods have shown amazing performance in synthesizing 3D-consistent photographic images, but fail to generate multi-view portrait drawings. The key is that the basic assumption of these methods -- a surface point is consistent when rendered from different views -- doesn't hold for drawings. In a portrait drawing, the appearance of a facial point may changes when viewed from different angles. Besides, portrait drawings usually present little 3D information and suffer from insufficient training data. To combat this challenge, in this paper, we propose a Semantic-Aware GEnerator (SAGE) for synthesizing multi-view portrait drawings. Our motivation is that facial semantic labels are view-consistent and correlate with drawing techniques. We therefore propose to collaboratively synthesize multi-view semantic maps and the corresponding portrait drawings. To facilitate training, we design a semantic-aware domain translator, which generates portrait drawings based on features of photographic faces. In addition, use data augmentation via synthesis to mitigate collapsed results. We apply SAGE to synthesize multi-view portrait drawings in diverse artistic styles. Experimental results show that SAGE achieves significantly superior or highly competitive performance, compared to existing 3D-aware image synthesis methods. The codes are available at https://github.com/AiArt-HDU/SAGE.
Fei Gao 0006, Nannan Wang 0001, Gang Xu 0001
IJCAI4
2023 Unsupervised Visible-Infrared Person ReID by Collaborative Learning with Neighbor-Guided Label Refinement
abstract
Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims at learning modality-invariant features from unlabeled cross-modality dataset, which is crucial for practical applications in video surveillance systems. The key to essentially address the USL-VI-ReID task is to solve the cross-modality data association problem for further heterogeneous joint learning. To address this issue, we propose a Dual Optimal Transport Label Assignment (DOTLA) framework to simultaneously assign the generated labels from one modality to its counterpart modality. The proposed DOTLA mechanism formulates a mutual reinforcement and efficient solution to cross-modality data association, which could effectively reduce the side-effects of some insufficient and noisy label associations. Besides, we further propose a cross-modality neighbor consistency guided label refinement and regularization module, to eliminate the negative effects brought by the inaccurate supervised signals, under the assumption that the prediction or label distribution of each example should be similar to its nearest neighbors'. Extensive experimental results on the public SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed method, surpassing existing state-of-the-art approach by a large margin of 7.76% mAP on average, which even surpasses some supervised VI-ReID methods.
De Cheng, Xiaojian Huang, Nannan Wang 0001, Zhihui Li 0001, Xinbo Gao 0001
ACM Multimedia3
2023 Efficient Bilateral Cross-Modality Cluster Matching for Unsupervised Visible-Infrared Person ReID
abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) aims to match pedestrian images of the same identity from different modalities without annotations. Existing works mainly focus on alleviating the modality gap by aligning instance-level features of the unlabeled samples. However, the relationships between cross-modality clusters are not well explored. To this end, we propose a novel bilateral cluster matching-based learning framework to reduce the modality gap by matching cross-modality clusters. Specifically, we design a Many-to-many Bilateral Cross-Modality Cluster Matching (MBCCM) algorithm through optimizing the maximum matching problem in a bipartite graph. Then, the matched pairwise clusters utilize shared visible and infrared pseudo-labels during the model training. Under such a supervisory signal, a Modality-Specific and Modality-Agnostic (MSMA) contrastive learning framework is proposed to align features jointly at a cluster-level. Meanwhile, the cross-modality Consistency Constraint (CC) is proposed to explicitly reduce the large modality discrepancy. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed method, surpassing state-of-the-art approaches by a large margin of 8.76% mAP on average.
De Cheng, Nannan Wang 0001, Shizhou Zhang, Zhen Wang 0037, Xinbo Gao 0001
ACM Multimedia3
2023 Controllable Face Sketch-Photo Synthesis with Flexible Generative Priors
abstract
Current face sketch-photo synthesis researches generally embrace an image-to-image (I2I) translation pipeline. However, these methods ignore the one-to-many mapping problem (i.e., multiple plausible photo results can correspond to a single input sketch) in sketch-to-photo synthesis task, resulting in significant performance degradation on diverse datasets. Besides, generating high-quality images on limited data is also a challenge for this task. To address these challenges, we propose a dual-path framework that introduces generative priors to better perform cross-domain reconstruction on limited data. The coarse path uses a layer-swapped pre-trained generator to achieve coarse cross-domain reconstruction, and the refinement path further improves the structure and texture details. To align the feature maps between the two paths, we introduce a spatial feature calibration module. Despite this, our framework still struggles to handle diverse datasets. Thanks to the flexibility of generative priors, we can extend the framework to achieve exemplar-guided I2I translation by incorporating an exemplar with style mixing and a proposed semantic-aware style refinement strategy, which addresses the one-to-many mapping problem in sketch-to-photo synthesis task. Furthermore, our framework can perform cross-domain editing by employing off-the-shelf editing methods based on the latent space, achieving fine-grained control. Extensive experiments on diverse datasets demonstrate the superiority of our framework over other state-of-the-art methods.
Mingrui Zhu, Nannan Wang 0001, Guozhang Li, Xiaoyu Wang 0002, Xinbo Gao 0001
ACM Multimedia3
2023 Modality-agnostic Augmented Multi-Collaboration Representation for Semi-supervised Heterogenous Face Recognition
abstract
Heterogeneous face recognition (HFR) aims to match input face identity across different image modalities. Due to the existing large modality gap and the limited number of training data, HFR is still a challenging problem in biometrics and draws more and more attention. Existing researchers always extract modality invariant features or generate homogeneous images to decrease the modality gap, lacking abundant labeled data to avoid the overfitting problem. In this paper, we proposed a novel Modality-Agnostic Augmented Multi-Collaboration representation for Heterogeneous Face Recognition (MAMCO-HFR) in a semi-supervised manner. The modality-agnostic augmentation strategy is proposed to generate adversarial perturbations to map unlabeled faces into the modality-agnostic domain. The multi-collaboration feature constraint is designed to mine the inherent relationships between diverse layers for discriminative representation. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ and Tufts Face dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is available at https://github.com/xiyin11/Semi-HFR.
Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
ACM Multimedia4
2023 NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation Learning
abstract
As more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An effective representation can be used to predict target attributes of networks without the need for actual training and deployment procedures, facilitating efficient network design and deployment. Recently, inspired by the success of Transformer, some Transformer-based representation learning frameworks have been proposed and achieved promising performance in handling cell-structured models. However, graph neural network (GNN) based approaches still dominate the field of learning representation for the entire network. In this paper, we revisit the Transformer and compare it with GNN to analyze their different architectural characteristics. We then propose a modified Transformer-based universal neural network representation learning model NAR-Former V2. It can learn efficient representations from both cell-structured networks and entire networks. Specifically, we first take the network as a graph and design a straightforward tokenizer to encode the network into a sequence. Then, we incorporate the inductive representation learning capability of GNN into Transformer, enabling Transformer to generalize better when encountering unseen architecture. Additionally, we introduce a series of simple yet effective modifications to enhance the ability of the Transformer in learning representation from graph structures. In encoding entire networks and then predicting the latency, our proposed method surpasses the GNN-based method NNLP by a significant margin on the NNLQP dataset. Furthermore, regarding accuracy prediction on the cell-structured NASBench101 and NASBench201 datasets, our method achieves highly comparable performance to other state-of-the-art methods. The code is available at https://github.com/yuny220/NAR-Former-V2.
Yun Yi, Haokui Zhang, Rong Xiao 0003, Nannan Wang 0001, Xiaoyu Wang 0002
NeurIPS4
2023 BiTGAN: bilateral generative adversarial networks for Chinese ink wash painting style transfer
Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
Sci. China Inf. Sci.3
2023 AutoEncoder-Driven Multimodal Collaborative Learning for Medical Image Synthesis
Bing Cao 0002, Zhiwei Bi, Qinghua Hu, Han Zhang 0002, Nannan Wang 0001, Xinbo Gao 0001, Dinggang Shen
Int. J. Comput. Vis.5
2023 Advanced Binary Neural Network for Single Image Super Resolution
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
Int. J. Comput. Vis.2
2023 Face photo-sketch synthesis via intra-domain enhancement
Chunlei Peng, Congyu Zhang, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Knowl. Based Syst.4
2023 Extended $T$T: Learning With Mixed Closed-Set and Open-Set Noisy Labels
abstract
The noise transition matrix T, reflecting the probabilities that true labels flip into noisy ones, is of vital importance to model label noise and build statistically consistent classifiers. The traditional transition matrix is limited to model closed-set label noise, where noisy training data have true class labels within the noisy label set. It is unfitted to employ such a transition matrix to model open-set label noise, where some true class labels are outside the noisy label set. Therefore, when considering a more realistic situation, i.e., both closed-set and open-set label noises occur, prior works will give unbelievable solutions. Besides, the traditional transition matrix is mostly limited to model instance-independent label noise, which may not perform well in practice. In this paper, we focus on learning with the mixed closed-set and open-set noisy labels. We address the aforementioned issues by extending the traditional transition matrix to be able to model mixed label noise, and further to the cluster-dependent transition matrix to better combat the instance-dependent label noise in real-world applications. We term the proposed transition matrix as the cluster-dependent extended transition matrix. An unbiased estimator (i.e., extended T-estimator) has been designed to estimate the cluster-dependent extended transition matrix by only exploiting the noisy data. Comprehensive experiments validate that our method can better cope with realistic label noise, following its more robust performance than the prior state-of-the-art label-noise learning methods.
Xiaobo Xia, Bo Han 0003, Nannan Wang 0001, Jiankang Deng, Yinian Mao, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Discriminative and Robust Attribute Alignment for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to learn models that can recognize images of semantically related unseen categories, through transferring attribute-based knowledge learned from training data of seen classes to unseen testing data. As visual attributes play a vital role in ZSL, recent embedding-based methods usually focus on learning a compatibility function between the visual representation and the class semantic attributes. While in this work, in addition to simply learning the region embedding of different semantic attributes to maintain the generalization capability of the learned model, we further consider to improve the discrimination power of the learned visual features themselves by contrastive embedding. It exploits both the class-wise and instance-wise supervision for GZSL, under the attribute guided weakly supervised representation learning framework. To further improve the robustness of the ZSL model, we also propose to train the model under the consistency regularization constraint, through taking full advantages of self-supervised signals of the image under various perturbed augmentation situations, which could make the model robust to some occluded or un-related attribute regions. Extensive experimental results demonstrate the effectiveness of the proposed ZSL method, achieving superior performances to state-of-the-art methods on three widely-used benchmark datasets, namely CUB, SUN, and AWA2. Our source code is released athttps://github.com/KORIYN/CC-ZSL.
De Cheng, Gerong Wang, Nannan Wang 0001, Dingwen Zhang, Qiang Zhang 0020, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Spatial-Temporal Frequency Forgery Clue for Video Forgery Detection in VIS and NIR Scenario
abstract
In recent years, with the rapid development of face editing and generation, more and more fake videos are circulating on social media, which has caused extreme public concerns. Existing face forgery detection methods based on frequency domain find that the GAN forged images have obvious grid-like visual artifacts in the frequency spectrum. But for synthesized videos, these methods only confine to a single frame and pay little attention to the most discriminative part and temporal frequency clue among different frames. To take full advantage of the rich information in video sequences, this paper performs video forgery detection on both spatial and temporal frequency domains and proposes a Discrete Cosine Transform-based Forgery Clue Augmentation Network (FCAN-DCT) to achieve a more comprehensive spectrum spatial-temporal feature representation. FCAN-DCT totally consists of a backbone network and two branches: Compact Feature Extraction (CFE) module and Frequency Temporal Attention (FTA) module. We conduct thorough experimental assessments on three visible light (VIS) based datasets (i.e.,, FaceForensics++, Celeb-DF (v2), WildDeepfake), and our self-built video forgery dataset DeepfakeNIR, which is the first video forgery dataset on near-infrared (NIR) modality. The experimental results demonstrate the effectiveness and robustness of our method for detecting forgery videos in both VIS and NIR scenarios.DeepfakeNIR and code are available athttps://github.com/AEP-WYK/DeepfakeNIR.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Learning a High Fidelity Identity Representation for Face Frontalization
abstract
This paper considers the problem of face frontalization in the wild, which transforms a face image with profile views into a frontal face. Face frontalization provides an effective solution to the face recognition problem in uncontrolled scenes. However, the existing methods either focus on deep learning techniques as an end-to-end framework or combine other explicit facial prior estimation tasks, such as 3D representation, optical flow estimation and so on, where computation is highly redundant and facial identity cannot be well represented. In this paper, we focus on how to maximise the potential of the model for identity learning and representation, and propose an accurate and lightweight face frontalization approach, named identity-preserving model (IPM). IPM has a well-designed encoder-decoder architecture which restores input face to a frontal counterpart. The encoder is constructed to extract representation from the input face, where a contrastive loss function is applied that encourages representations to form compact clusters, while preserving their relationships across the corpora. Then a cross-domain rectification module is proposed to eliminate the representation differences between the recognition and reconstruction domains, thus improving the accuracy of the reconstructed face. Extensive experiments on benchmark datasets show that the proposed IPM approach not only outperforms the state-of-the-art on public datasets but also can cope with images in the uncontrolled scenes.
Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Dual Conditional Normalization Pyramid Network for Face Photo-Sketch Synthesis
abstract
Face photo-sketch synthesis has undergone remarkable progress with the rapid development of deep learning techniques. Cutting-edge methods directly learn the cross-domain mapping between photos and sketches, which ignores the available reference samples. We argue that the reference samples can provide adequate prior information on texture and content in this task and improve the visual performance of synthetic images. This paper proposes a Dual Conditional Normalization Pyramid (DCNP) network with a multi-scale pyramid structure. The core of the DCNP network is a Dual Conditional Normalization (DCN) based architecture, which can obtain prior information on different semantics from reference samples. Specifically, DCN contains two conditional normalization branches. The first branch allows for spatially-adaptive normalization of the reference image conditioned on the semantic mask of the input image. The second branch enables adaptive instance normalization of the input image conditioned on the reference image. DCN can emphasize the isolated importance of textural and spatial factors by disintegrating the entire cross-domain mapping into two branches. To avoid information redundancy and improve the final performance, we propose a Gated Channel Attention Fusion (GCAF) module to distill and fuse the helpful information of the two branches. Qualitative and quantitative experimental results demonstrate the superior performance of the proposed method over the state-of-the-art approaches in structural information preservation and realistic texture generation. The code is public inhttps://github.com/Tony0720/DCNP.
Mingrui Zhu, Zicheng Wu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Unsupervised Across Domain Consistency- Difference Network for Hyperspectral Image Super-Resolution
abstract
Without reducing the spectral resolution, hyperspectral image super-resolution has achieved remarkable progress thanks to the success of deep neural networks. However, existing methods can not fully excavate the latent high-frequency details only in the single spatial domain. Different from existing methods that only achieves the super-resolution task in spatial domain, we optimize the amplitude spectrum and phase spectrum in frequency domain to obtain high resolution hyperspectral image (HR-HSI). We propose a new unsupervised framework to reconstruct HR-HSI using only the observed low resolution HSI and HR multispectral image. Based on triple-level modeling, the encoder-decoder learns abundant features including contextual information from multiple scales. In addition, we propose iterative across domain consistency-difference (ADCD) module, which is embedded between encoder and decoder. In ADCD module, three parallel convolution streams, (amplitude spectrum adjustment branch, phase spectrum adjustment branch and spatial domain branch) are used to explore the consistency-difference between each other, which is preserved by memory units within the module. Particularly, we embed the dilated causal convolution in the frequency domain processing branch, which is convenient to flexibly adjust the receptive field and adapt to different domains. Extensive experiments are conducted on widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the proposed method.
Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 FedForgery: Generalized Face Forgery Detection With Residual Federated Learning
abstract
With the continuous development of deep learning in the field of image generation models, a large number of vivid forged faces have been generated and spread on the Internet. These high-authenticity artifacts could grow into a threat to society security. Existing face forgery detection methods directly utilize the obtained public shared or centralized data for training but ignore the personal privacy and security issues when personal data couldn’t be centralizedly shared in real-world scenarios. Additionally, different distributions caused by diverse artifact types would further bring adverse influences on the forgery detection task. To solve the mentioned problems, the paper proposes a novel generalized residual Federated learning for face Forgery detection (FedForgery). The designed variational autoencoder aims to learn robust discriminative residual feature maps to detect forgery faces (with diverse or even unknown artifact types). Furthermore, the general federated learning strategy is introduced to construct distributed detection model trained collaboratively with multiple local decentralized devices, which could further boost the representation generalization. Experiments conducted on publicly available face forgery detection datasets prove the superior performance of the proposed FedForgery. The designed novel generalized face forgery detection protocols and source code would be publicly available at https://github.com/GANG370/FedForgery.
Decheng Liu, Zhan Dang, Chunlei Peng, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.6
2023 FABNet: Frequency-Aware Binarized Network for Single Image Super-Resolution
abstract
Remarkable achievements have been obtained with binary neural networks (BNN) in real-time and energy-efficient single-image super-resolution (SISR) methods. However, existing approaches often adopt the Sign function to quantize image features while ignoring the influence of image spatial frequency. We argue that we can minimize the quantization error by considering different spatial frequency components. To achieve this, we propose a frequency-aware binarized network (FABNet) for single image super-resolution. First, we leverage the wavelet transformation to decompose the features into low-frequency and high-frequency components and then employ a "divide-and-conquer" strategy to separately process them with well-designed binary network structures. Additionally, we introduce a dynamic binarization process that incorporates learned-threshold binarization during forward propagation and dynamic approximation during backward propagation, effectively addressing the diverse spatial frequency information. Compared to existing methods, our approach is effective in reducing quantization error and recovering image textures. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed methods could surpass state-of-the-art approaches in terms of PSNR and visual quality with significantly reduced computational costs. Our codes are available at https://github.com/xrjiang527/FABNet-PyTorch.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.2
2023 Frequency Information Disentanglement Network for Video-Based Person Re-Identification
abstract
Recently, most video-based person re-identification (Re-ID) methods adopt complex model or multi-scaled information to explore more discriminative spatio-temporal clues, thus achieving better retrieval accuracy. However, we witness that these approaches involve significant higher computation costs but only improve limited performances. Therefore, the overarching goal at this stage is to solve video Re-ID on the trade-off between accuracy and efficiency, thereby boosting the application in real scenarios. Frequency transform provides advantages of simplified representation, identification of hidden information and noise filtering in signal processing. Motivated by this, we treat the complex spatio-temporal feature as signal and convert it to frequency domain. By directly analyzing frequency clues, complex feature extraction procedures can be avoided. Specifically, this paper proposes a novel paradigm by categorizing video features into low/high and spatial/temporal frequency information. Then, with the help of 3D DCT, we theoretically establish the transform equivalence relationship between spatio-temporal domain and frequency domain. Finally, this paper proposes a simple and intuitive Frequency Information Disentanglement Network (FIDN) for video Re-ID. By extracting and applying both low and high frequency spatio-temporal features from a disentangling way, FIDN achieves comprehensive and discriminative video representation. Extensive experiments indicate that FIDN reaches the state-of-the-arts with only one convolution layer addition against baseline.
Liangchen Liu 0001, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2023 HiFiSketch: High Fidelity Face Photo-Sketch Synthesis and Manipulation
abstract
With the rapid development of generative adversarial networks, face photo-sketch synthesis has achieved promising performance and playing an increasingly important role in law enforcement as well as entertainment. However, most of the existing methods only work under the condition of no interference, and lack of generalization ability in wild scenes. The fidelity of the images generated by the existing methods are insufficient, and the manipulation ability according to text description is unavailable. Directly applying existing text-based image manipulation methods on face photo-sketch scenario may lead to severe distortions due to the cross-domain challenges. Therefore, we propose a novel cross-domain face photo-sketch synthesis framework named HiFiSketch, a network that learns to adjust the weights of generators for high-fidelity synthesis and manipulation. It can realize the translation of images between the photo domain and the sketch domain, and modify results according to the text input in the meanwhile. We further propose a cross-domain loss function, which can effectively preserve facial details during face photo-sketch synthesis. Extensive experiments on four public face sketch datasets show the superiority of our method compared to existing methods. We further present text-based face photo-sketch manipulation and sequential face photo-sketch manipulation for the first time to demonstrate the effectiveness of our method on high fidelity face photo-sketch synthesis and manipulation.
Chunlei Peng, Congyu Zhang, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2023 An Efficient Transformer Based on Global and Local Self-Attention for Face Photo-Sketch Synthesis
abstract
Face photo-sketch synthesis tasks have been dominated by convolutional neural networks (CNNs), especially CNN-based generative adversarial networks (GANs), because of their strong texture modeling capabilities and thus their ability to generate more realistic face photos/sketches beyond traditional methods. However, due to CNNs' locality and spatial invariance properties, there have weaknesses in capturing the global and structural information which are extremely important for face images. Inspired by the recent phenomenal success of the Transformer in vision tasks, we propose replacing CNNs with Transformers that are able to model long-range dependencies to synthesize more structured and realistic face images. However, the existing vision Transformers are mainly designed for high-level vision tasks and lack the dense prediction ability to generate high resolution images due to the quadratic computational complexity of their self-attention mechanism. In addition, the original Transformer is not capable of modeling local correlations which is an important skill for image generation. To address these challenges, we propose two types of memory-friendly Transformer encoders, one for processing local correlations via local self-attention and another for modeling global information via global self-attention. By integrating the two proposed Transformer encoders, we present an efficient GL-Transformer for face photo-sketch synthesis, which can synthesize realistic face photo/sketch images from coarse to fine. Extensive experiments demonstrate that our model achieves a comparable or better performance beyond the state-of-the-art CNN-based methods both qualitatively and quantitatively.
Wangbo Yu, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.3
2023 Dual-Affinity Style Embedding Network for Semantic-Aligned Image Style Transfer
abstract
Image style transfer aims at synthesizing an image with the content from one image and the style from another. User studies have revealed that the semantic correspondence between style and content greatly affects subjective perception of style transfer results. While current studies have made great progress in improving the visual quality of stylized images, most methods directly transfer global style statistics without considering semantic alignment. Current semantic style transfer approaches still work in an iterative optimization fashion, which is impractically computationally expensive. Addressing these issues, we introduce a novel dual-affinity style embedding network (DaseNet) to synthesize images with style aligned at semantic region granularity. In the dual-affinity module, feature correlation and semantic correspondence between content and style images are modeled jointly for embedding local style patterns according to semantic distribution. Furthermore, the semantic-weighted style loss and the region-consistency loss are introduced to ensure semantic alignment and content preservation. With the end-to-end network architecture, DaseNet can well balance visual quality and inference efficiency for semantic style transfer. Experimental results on different scene categories have demonstrated the effectiveness of the proposed method.
Zhuoqi Ma, Xin Li 0106, Fu Li 0003, Dongliang He, Errui Ding, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.7
2022 Instance-Dependent Label-Noise Learning with Manifold-Regularized Transition Matrix Estimation
abstract
In label-noise learning, estimating the transition matrix has attracted more and more attention as the matrix plays an important role in building statistically consistent classifiers. However, it is very challenging to estimate the transition matrix T(x), where x denotes the instance, because it is unidentifiable under the instance-dependent noise (IDN). To address this problem, we have noticed that, there are psychological and physiological evidences showing that we humans are more likely to annotate instances of similar appearances to the same classes, and thus poor-quality or ambiguous instances of similar appearances are easier to be mislabeled to the correlated or same noisy classes. Therefore, we propose assumption on the geometry of T(x) that “the closer two instances are, the more similar their corresponding transition matrices should be”. More specifically, we formulate above assumption into the manifold embedding, to effectively reduce the degree of freedom of T(x) and make it stably estimable in practice. The proposed manifold-regularized technique works by directly reducing the estimation error without hurting the approximation error about the estimation problem of T(x). Experimental evaluations on four synthetic and two real-world datasets demonstrate that our method is superior to state-of-the-art approaches for label-noise learning under the challenging IDN.
De Cheng, Tongliang Liu, Yixiong Ning, Nannan Wang 0001, Bo Han 0003, Gang Niu 0001, Xinbo Gao 0001, Masashi Sugiyama
CVPR4
2022 Towards Semi-Supervised Deep Facial Expression Recognition with An Adaptive Confidence Margin
abstract
Only parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all unlabeled data. In this paper, we learn an Adaptive Confidence Margin (Ada-CM) to fully leverage all unlabeled data for semi-supervised deep facial expression recognition. All unlabeled samples are partitioned into two subsets by comparing their confidence scores with the adaptively learned confidence margin at each training epoch: (1) subset I including samples whose confidence scores are no lower than the margin; (2) subset II including samples whose confidence scores are lower than the margin. For samples in subset I, we constrain their predictions to match pseudo labels. Meanwhile, samples in subset II participate in the feature-level contrastive objective to learn effective facial expression features. We extensively evaluate Ada-CM on four challenging datasets, showing that our method achieves state-of-the-art performance, especially surpassing fully-supervised baselines in a semi-supervised manner. Ablation study further proves the effectiveness of our method. The source code is available at https://github.com/hangyu94/Ada-CM.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
CVPR2
2022 Exploring Set Similarity for Dense Self-supervised Representation Learning
abstract
By considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue, in this paper, we propose to explore set similarity (SetSim) for dense self-supervised representation learning. We generalize pixel-wise similarity learning to set-wise one to improve the robustness because sets contain more semantic and structure information. Specifically, by resorting to attentional features of views, we establish the corresponding set, thus filtering out noisy backgrounds that may cause incorrect correspondences. Meanwhile, these at-tentional features can keep the coherence of the same image across different views to alleviate semantic inconsistency. We further search the cross-view nearest neighbours of sets and employ the structured neighbourhood information to enhance the robustness. Empirical evaluations demonstrate that SetSim surpasses or is on par with state-of-the-art meth-ods on object detection, keypoint detection, instance segmen-tation, and semantic segmentation.
Zhaoqing Wang, Qiang Li 0024, Pengfei Wan 0001, Nannan Wang 0001, Mingming Gong, Tongliang Liu
CVPR6
2022 SketchCLIP: Text-based Attribute Manipulation for Face Sketch Synthesis
abstract
This paper proposes a method of modifying the face sketch with text descriptions. Face sketch is widely used in the criminal field and digital entertainment field. Forensic painters usually draw face sketches based on descriptions provided by witnesses or clients. However, drawing a face sketch often takes lots of time and effort. Existing face sketch synthesis studies have not considered text-based sketch manipulation, and we find that applying text-driven editing methods on natural images directly to face sketches causes severe distortion of generated results. Therefore, this paper proposes a novel text-based attribute manipulation method for face sketch synthesis, named SketchCLIP. Our approach adopts text-driven attribute manipulation by using the powerful Contrastive Language-Image Pre-Training (CLIP) model, which not only conforms to the current drawing process of face sketches but also does not require tedious manual operations and allows for more diverse modifications. Besides, we design an intra-modality fine-tuning module to eliminate distortion and improve the quality of the modified face sketch. Through extensive comparison experiments on public face sketch datasets, our method is demonstrated to be very excellent in the effectiveness of the face sketch processing and the quality of modified results.
Mengdi Dong, Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001
IJCB5
2022 Detach and Enhance: Learning Disentangled Cross-modal Latent Representation for Efficient Face-Voice Association and Matching
abstract
Many researches in cognitive science have shown that humans often perform face-voice association for various perception tasks, and some recent data mining works have been designed in emulating such ability intelligently. Nevertheless, most methods often suffer from the degraded performance when there exist semantically irrelevant interference factors across different modalities. To alleviate this concern, this paper presents an efficient Disentangled Cross-modal Latent Representation (DCLR) method to adaptively detach the discriminative feature attributes and enhance the face-voice association. To be specific, the proposed DCLR framework consists of two-stage cross-modal disentangling process. First, the former stage employs the supervised contrastive learning to push the representations of face-voice data from the same person closer while pulling those representations of different person away. Then, the latter stage freezes all the parameters of the former stage, and further innovates a multi-layer orthogonal decoupling scheme to learn the disentangled latent representations, while filtering out the modality-dependent irrelevant factors. Besides, the cross-modal reconstruction loss is further utilized to narrow down the semantic gap between heterogeneous feature expressions. Through the joint exploitation of the above, the proposed framework can well associate the face-voice data to benefit various kinds of cross-modal perception tasks. Extensive experiments verify the superiorities of the proposed face-voice association framework and show its competitive performances.
Zhenning Yu, Xin Liu 0011, Yiu-Ming Cheung, Minghang Zhu, Xing Xu 0001, Nannan Wang 0001, Taihao Li
ICDM6
2022 Improving Adversarial Robustness via Mutual Information Estimation
abstract
Deep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples from the perspective of information theory, and propose an adversarial defense method. Specifically, we first measure the dependence by estimating the mutual information (MI) between outputs and the natural patterns of inputs (called natural MI) and MI between outputs and the adversarial patterns of inputs (called adversarial MI), respectively. We find that adversarial samples usually have larger adversarial MI and smaller natural MI compared with those w.r.t. natural samples. Motivated by this observation, we propose to enhance the adversarial robustness by maximizing the natural MI and minimizing the adversarial MI during the training process. In this way, the target model is expected to pay more attention to the natural pattern that contains objective semantics. Empirical evaluations demonstrate that our method could effectively improve the adversarial accuracy against multiple attacks.
Dawei Zhou 0004, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003, Xiaoyu Wang 0002, Yibing Zhan, Tongliang Liu
ICML2
2022 Modeling Adversarial Noise for Adversarial Training
abstract
Deep neural networks have been demonstrated to be vulnerable to adversarial noise, promoting the development of defense against adversarial attacks. Motivated by the fact that adversarial noise contains well-generalizing features and that the relationship between adversarial data and natural data can help infer natural data and make reliable predictions, in this paper, we study to model adversarial noise by learning the transition relationship between adversarial labels (i.e. the flipped labels used to generate adversarial data) and natural labels (i.e. the ground truth labels of the natural data). Specifically, we introduce an instance-dependent transition matrix to relate adversarial labels and natural labels, which can be seamlessly embedded with the target model (enabling us to model stronger adaptive adversarial noise). Empirical evaluations demonstrate that our method could effectively improve adversarial accuracy.
Dawei Zhou 0004, Nannan Wang 0001, Bo Han 0003, Tongliang Liu
ICML2
2022 Robust Single Image Dehazing Based on Consistent and Contrast-Assisted Reconstruction
abstract
Single image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the research filed of singe image dehazing. To properly address this problem, we propose a novel density-variational learning framework to improve the robustness of the image dehzing model assisted by a variety of negative hazy images, to better deal with various complex hazy scenarios. Specifically, the dehazing network is optimized under the consistency-regularized framework with the proposed Contrast-Assisted Reconstruction Loss (CARL). The CARL can fully exploit the negative information to facilitate the traditional positive-orient dehazing objective function, by squeezing the dehazed image to its clean target from different directions. Meanwhile, the consistency regularization keeps consistent outputs given multi-level hazy images, thus improving the model robustness. Extensive experimental results on two synthetic and three real-world datasets demonstrate that our method significantly surpasses the state-of-the-art approaches.
De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001
IJCAI4
2022 Sample-Efficient Kernel Mean Estimator with Marginalized Corrupted Data
abstract
Estimating the kernel mean in a reproducing kernel Hilbert space is central to many kernel-based learning algorithms. Given a finite sample, an empirical average is used as a standard estimation of the target kernel mean. Prior works have shown that better estimators can be constructed by shrinkage methods. In this work, we propose to corrupt data examples with noise from known distributions and present a new kernel mean estimator, called the marginalized kernel mean estimator, which estimates kernel mean under the corrupted distributions. Theoretically, we justify that the marginalized kernel mean estimator introduces implicit regularization in kernel mean estimation. Empirically, on a variety of tasks, we show that the marginalized kernel mean estimator is sample-efficient and obtains much lower estimation errors than the existing estimators.
Xiaobo Xia, Mingming Gong, Nannan Wang 0001, Fei Gao 0006, Haikun Wei, Tongliang Liu
KDD4
2022 Class-Dependent Label-Noise Learning with Cycle-Consistency Regularization
abstract
In label-noise learning, estimating the transition matrix plays an important role in building statistically consistent classifier. Current state-of-the-art consistent estimator for the transition matrix has been developed under the newly proposed sufficiently scattered assumption, through incorporating the minimum volume constraint of the transition matrix T into label-noise learning. To compute the volume of T, it heavily relies on the estimated noisy class posterior. However, the estimation error of the noisy class posterior could usually be large as deep learning methods tend to easily overfit the noisy labels. Then, directly minimizing the volume of such obtained T could lead the transition matrix to be poorly estimated. Therefore, how to reduce the side-effects of the inaccurate noisy class posterior has become the bottleneck of such method. In this paper, we creatively propose to estimate the transition matrix under the forward-backward cycle-consistency regularization, of which we have greatly reduced the dependency of estimating the transition matrix T on the noisy class posterior. We show that the cycle-consistency regularization helps to minimize the volume of the transition matrix T indirectly without exploiting the estimated noisy class posterior, which could further encourage the estimated transition matrix T to converge to its optimal solution. Extensive experimental results consistently justify the effectiveness of the proposed method, on reducing the estimation error of the transition matrix and greatly boosting the classification performance.
De Cheng, Yixiong Ning, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003, Tongliang Liu
NeurIPS3
2022 VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild
abstract
We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this objective into three sequential tasks: (1) face video generation with a canonical expression; (2) audio-driven lip-sync; and (3) face enhancement for improving photo-realism. Given a talking-head video, we first modify the expression of each frame according to the same expression template using the expression editing network, resulting in a video with the canonical expression. This video, together with the given audio, is then fed into the lip-sync network to generate a lip-syncing video. Finally, we improve the photo-realism of the synthesized faces through an identity-aware face enhancement network and post-processing. We use learning-based approaches for all three steps and all our modules can be tackled in a sequential pipeline without any user intervention. Furthermore, our system is a generic approach that does not need to be retrained to a specific person. Evaluations on two widely-used datasets and in-the-wild examples demonstrate the superiority of our framework over other state-of-the-art methods in terms of lip-sync accuracy and visual quality.
Xiaodong Cun, Yong Zhang 0034, Menghan Xia, Mingrui Zhu, Xuan Wang 0009, Jue Wang 0001, Nannan Wang 0001
SIGGRAPH Asia9
2022 Single image dehazing with an independent Detail-Recovery Network
Yan Li 0125, De Cheng, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001
Knowl. Based Syst.4
2022 Face photo-sketch synthesis via full-scale identity supervision
Bing Cao 0002, Nannan Wang 0001, Jie Li 0001, Qinghua Hu, Xinbo Gao 0001
Pattern Recognit.2
2022 Spatiotemporal consistency-enhanced network for video anomaly detection
Jie Li 0001, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
Pattern Recognit.3
2022 SAR-to-optical image translation based on improved CGAN
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.4
2022 Local-Global Graph Pooling via Mutual Information Maximization for Video-Paragraph Retrieval
abstract
As a task of cross-modal retrieval between long videos and paragraphs, video-paragraph retrieval is a non-trivial task. Unlike traditional video-text retrieval, the video in video-paragraph retrieval usually contains multiple clips. Each clip corresponds to a descriptive sentence; all the sentences constitute the corresponding paragraph of the video. Previous methods for video-paragraph retrieval usually encode videos and para-graphs from segment-level (clips and sentences) and overall-level (videos and paragraphs). However, there are also contents about actions and objects that exist in the segment. Hence, we propose a Local-Global Graph Pooling Network (LGGP) via Mutual Information Maximization for video-paragraph retrieval. Our model disentangles videos and paragraphs into four levels: overall-level, segment-level, motion-level, and object-level. We construct the Hierarchical Local Graph (segment-level, motion-level, and object-level) and the Hierarchical Global Graph (overall-level, segment-level, motion-level, and object-level), respectively, for semantic interaction among different levels. Meanwhile, to obtain hierarchical pooling features with fine-grained semantic information, we design hierarchical graph pooling methods to maximize the mutual information between pooling features and corresponding graph nodes. We evaluate our model on two video-paragraph retrieval datasets with three different video features. The experimental results show that our model establishes state-of-the-art results for video-paragraph retrieval. Our code will be released athttps://github.com/PengchengZhang1997/LGGP.
Zhou Zhao 0001, Nannan Wang 0001, Jun Yu 0002, Fei Wu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Dually Distribution Pulling Network for Cross-Resolution Person Reidentification
abstract
Person reidentification (Re-ID) aims at recognizing the same identity across different camera views. However, the cross resolution of images [high resolution (HR) and low resolution (LR)] is unavoidable in a realistic scenario due to the various distances among cameras and pedestrians of interest, thus leading to cross-resolution person Re-ID problems. Recently, most cross-resolution person Re-ID methods focus on solving the resolution mismatch problem, while the distribution mismatch between HR and LR images is another factor that significantly impacts the person Re-ID performance. In this article, we propose a dually distribution pulling network (DDPN) to tackle the distribution mismatch problem. DDPN is composed of two modules, that is: 1) super-resolution module and 2) person Re-ID module. They attempt to pull the distribution of LR images closer to the distribution of HR images from image and feature aspects, respectively, through optimizing the maximum mean discrepancy losses. Extensive experiments have been conducted on three benchmark datasets and the results demonstrate the effectiveness of DDPN. Remarkably, DDPN shows a great advantage when compared to the state-of-the-art methods, for instance, we achieve rank-1 accuracy of 76.9% on VR-Market1501, which outperforms the best existing cross-resolution person Re-ID method by 10%.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Cybern.4
2022 RBDF: Reciprocal Bidirectional Framework for Visible Infrared Person Reidentification
abstract
Visible infrared person reidentification (VI-REID) plays a critical role in night-time surveillance applications. Most methods attempt to reduce the cross-modality gap by extracting the modality-shared features. However, they neglect the distinct image-level discrepancies among heterogeneous pedestrian images. In this article, we propose a reciprocal bidirectional framework (RBDF) to achieve modality unification before discriminative feature learning. The bidirectional image translation subnetworks can learn two opposite mappings between visible and infrared modality. Particularly, we investigate the characteristics of the latent space and design a novel associated loss to pull close the distribution between the intermediate representations of two mappings. Mutual interaction between two opposite mappings helps the network generate heterogeneous images that have high similarity with the real images. Hence, the concatenation of original and generated images can eliminate the modality gap. During the feature learning procedure, the attention mechanism-based feature embedding network can learn more discriminative representations with the identity classification and feature metric learning. Experimental results indicate that our method achieves state-of-the-art performance. For instance, we achieve 54.41% mAP and 57.66% rank-1 accuracy on SYSU-MM01 dataset, outperforming the existing works by a large margin.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Cybern.3
2022 Learning Deep Resonant Prior for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSISR) task has been widely studied, and significant progress has been made by leveraging the deep convolution neural network (CNN) techniques. Nevertheless, the scarcity of training images hinders the research progress of HSISR task. Moreover, the differences in imaging conditions and the number of spectral bands among different datasets, make it very difficult to construct a unified deep neural network. In this paper, we first present a non-training based HSISR method based on deep prior knowledge, which captures the image prior to restore the high resolution image by using the intrinsic characteristics of CNN. Then, we append a special network input processing module onto the HSI super-resolution network to automatically adjust the structure of the input so that the choice of network structure is no longer limited, while the network design focuses on exploiting the spatial information of hyperspectral images and the correlation between spectral bands, making the method more suitable for HSISR tasks and greatly extending its applications. Extensive experiment results on the hyperspectral image datasets illustrate the effectiveness of the proposed method, and we have got comparable results with the state-of-the-art methods while requiring no training samples.
Zhaori Gong, Nannan Wang 0001, De Cheng, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 External-Internal Attention for Hyperspectral Image Super-Resolution
abstract
In recent years, hyperspectral image (HSI) super-resolution has made significant progress by leveraging convolution neural network. Existing methods with spectral or spatial attention, which only consider the spectral similarity or pixel-pixel similarity, ignore sample-sample correlations and sparsity. Therefore, based on the fusion of HSI and multispectral image, we propose a new HSI super-resolution model with external-internal attention. Instead of considering a single sample, external attention module is employed to exploit the incorporating correlations between different samples to get a better feature representation. In addition, an internal attention module based on non-local operation is designed to explore the long-range dependencies information. Particularly, oriented to high mapping precision and low computational cost inference, spherical locality sensitive hashing is used to divide features into different hash buckets so that every query point is calculated in the hash bucket assigned to it, rather than based a weight sum of features across all positions. The sequential external-internal attention greatly improves the generalization ability and robustness of the model by modeling at the dataset level and at the sample level. Extensive experiments are conducted on five widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the method we proposed.
Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 An Enhanced SiamMask Network for Coastal Ship Tracking
abstract
Coastal ship tracking is significant for cargo transportation and route determination. However, there are few effective tracking methods for specialized ship tracking. Although Siamese networks have been commonly used for object tracking in the field of deep learning, the results for the ship are not accurate due to the lack of contour and edge information. In addition, scale variation and seawater cause unstable ship movements, which aggravates the reduction in tracking accuracy. Therefore, we propose an enhanced SiamMask network for coastal ship tracking. Compared to the previous Siamese network, our algorithm has the following three advantages. First, we apply the unity of visual object tracking and semisupervised object segmentation to the ship tracking task, which completes target tracking while outputting edge shape information. Second, we propose a refined feature pyramid network that utilizes a feature fusion module and enhanced residual module (ERM) to solve the problems of scale variation in datasets. Third, we propose an attentionwise cross correlation with a multidimension attention module (MDAM) to focus more on ship targets and suppress nontargets through autonomous learning at width, height, and channel levels, which creates a tradeoff between the accuracy and the robustness of tracking algorithms. The experimental results show that our method achieves leading performance in LMD-TShips, outperforming most of the state-of-the-art trackers. Code is available athttps://gitee.com/EnhancedSiamShipTracking/code.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 SRDN: A Unified Super-Resolution and Motion Deblurring Network for Space Image Restoration
abstract
Space target super-resolution (SR) is a domain-specific single image SR problem aiming to help distinguish the satellite and spacecrafts from numerous space debris. Compared to the other object SR problem, images for space target are always in low quality with varies of degradation condition, as a result of long distance and motion blur, which significantly reduces the manual classification reliability, especially for these small targets, e.g., satellite payloads. To address this challenge, we present an end-to-end SR and deblurring network (SRDN). Concretely, focusing on the low-resolution (LR) space target images with blind motion blur, we integrate the SR and deblur function together, improving the image quality by a unified generative adversarial network (GAN)-based framework. We implement a deblur module by using contrastive learning to extract degradation feature and add symmetrical downsampling and upsampling modules to the SR network in order to restore texture information, while shortcut connections are redesigned to maintain the global similarity. Extensive experiments on the public satellite dataset, BUAA-SID-share1.5, demonstrate that our network outperforms the state-of-the-art SR and deblur methods.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 A Robust One-Stage Detector for Multiscale Ship Detection With Complex Background in Massive SAR Images
abstract
With the development of synthetic aperture radar (SAR) imaging and deep learning, SAR ship detection based on convolutional neural networks (CNNs) has been extensively applied in the last few years. Nevertheless, there are two main obstacles in SAR ship detection: 1) the SAR images have too much noise, such as the interference from land area, making it difficult to distinguish ship objects from the surrounding background, and 2) due to the multiscale characteristics of ship objects, there are numerous false negatives in the detection results, especially for small objects. To alleviate the above problems, we propose a one-stage ship detector with strong robustness against scale changes and various interferences. First, to mitigate the disturbance from complex background, a coordinate attention module (CoAM) is introduced for obtaining more representative semantic features to accurately locate and distinguish ship objects. Second, a receptive field increased module (RFIM) is devised to capture multiscale contextual information to improve the detection performance for ships with diverse scales. Finally, we verify the robustness of our method on several public SAR datasets, i.e., SAR-Ship-Dataset, high-resolution SAR images dataset (HRSID), and SAR ship detection dataset (SSDD). The experimental results demonstrate that the proposed method has a competitive performance, exceeding other state-of-the-art methods by at least 2.6% AP50on HRSID.
Xi Yang 0011, Xin Zhang 0129, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 ForgeryNIR: Deep Face Forgery and Detection in Near-Infrared Scenario
abstract
Deep face forgery and detection is an emerging topic due to the development of GANs. Face forgery detection relies greatly on existing databases for evaluation and adequate training examples for data-hungry machine learning algorithms. However, considering the wide application of face recognition in near-infrared scenarios, there is no publicly available face forgery database that includes near-infrared modality currently. In this paper, we present an attempt at constructing a large-scale dataset for face forgery detection in the near-infrared modality and propose a new forgery detection method based on knowledge distillation named cross-modality knowledge distillation aiming to use a teacher model which is pre-trained on the visible light-based (VIS) big data to guide the student model with a small amount of near-infrared (NIR) data. The proposed near-infrared face forgery dataset, named ForgeryNIR, contains a total of over 50,000 real and fake identities. A number of perturbations are applied to help simulate real-world scenarios. All source images in ForgeryNIR are collected from CASIA NIR-VIS 2.0, and fake images are generated via multiple GAN techniques. The proposed dataset fills the gap of face forgery detection research in the near-infrared modality. A comprehensive study on six representative detection baselines is conducted to evaluate the performance of face forgery detection algorithms in the NIR domain. We further construct a hard testing set, named ForgeryNIR+, which contains forged images that have bypassed existing face forgery detection methods. The proposed datasets will be publicly available and aim to help boost further research on face forgery detection, as well as NIR face detection and recognition.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2022 Edge Aware Domain Transformation for Face Sketch Synthesis
abstract
With the development of generative adversarial networks (GAN), the field of face sketch synthesis has received extensive attention. Face sketch synthesis (FSS) has promising prospects in the fields of entertainment and law enforcement, where it plays an increasingly important role. We propose a novel generative adversarial network for synthesizing sketches with similar shapes and rich details to photos. This problem is challenging because it involves the transition between the sketch domain and the photo domain. Many methods have been used for face sketch synthesis in recent years, but existing methods cannot fully exploit the semantic information between different domains. To this end, we use a cross-domain face sketch synthesis framework based on edge-preserving filters to make the boundaries of different semantics in semantic layouts have a smooth transition. We further propose a new spatially adaptive denormalization module named edge-aware enhancement Spatially Adaptive DEnormalization (eaeSPADE), which can make full use of the semantic information in the semantic layout of faces and improve the details of the synthesized face images. Extensive experiments demonstrate that our method outperforms existing face sketch synthesis methods.
Congyu Zhang, Decheng Liu, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2022 Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-Id
abstract
Unsupervised person re-identification (Re-Id) has attracted increasing attention due to its practical application in the read-world video surveillance system. The traditional unsupervised Re-Id are mostly based on the method alternating between clustering and fine-tuning with the classification or metric learning objectives on the grouped clusters. However, since person Re-Id is an open-set problem, the clustering based methods often leave out lots of outlier instances or group the instances into the wrong clusters, thus they can not make full use of the training samples as a whole. To solve these problems, we present the hybrid dynamic cluster contrast and probability distillation algorithm. It formulates the unsupervised Re-Id problem into an unified local-to-global dynamic contrastive learning and self-supervised probability distillation framework. Specifically, the proposed method can make the best of the self-supervised signals of all the clustered and un-clustered instances, from both the instances' self-contrastive level and the probability distillation respectives, in the memory-based non-parametric manner. Besides, the proposed hybrid local-to-global contrastive learning can take full advantage of the informative and valuable training examples for effective and robust training. Extensive experiment results show that the proposed method achieves superior performances to state-of-the-art methods, under both the purely unsupervised and unsupervised domain adaptation experiment settings. Our source code is released in https://github.com/zjy2050/HDCRL-ReID.
De Cheng, Jingyu Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2022 Exploring Language Hierarchy for Video Grounding
abstract
The understanding of language plays a key role in video grounding, where a target moment is localized according to a text query. From a biological point of view, language is naturally hierarchical, with the main clause (predicate phrase) providing coarse semantics and modifiers providing detailed descriptions. In video grounding, moments described by the main clause may exist in multiple clips of a long video, including both the ground-truth and background clips. Therefore, in order to correctly discriminate the ground-truth clip from the background ones, this co-existence leads to the negligence of the main clause, and concentrate the model on the modifiers that provide discriminative information on distinguishing the target proposal from the others. We first demonstrate this phenomenon empirically, and propose a Hierarchical Language Network (HLN) that exploits the language hierarchy, as well as a new learning approach called Multi-Instance Positive-Unlabelled Learning (MI-PUL) to alleviate the above problem. Specifically, in HLN, the localization is performed on various layers of the language hierarchy, so that the attention can be paid to different parts of the sentences, rather than only discriminative ones. Furthermore, MI-PUL allows the model to localize background clips that can be possibly described by the main clause, even without manual annotations. Therefore, the union of the two proposed components enhances the learning of the main clause, which is of critical importance in video grounding. Finally, we evaluate that our proposed HLN can plug into the current methods and improve their performance. Extensive experiments on challenging datasets show HLN significantly improve the state-of-the-art methods, especially achieving 6.15% gain in terms of [Formula: see text] on the TACoS dataset.
Xinpeng Ding, Nannan Wang 0001, Shiwei Zhang 0001, Ziyuan Huang 0003, Xiaomeng Li 0001, Mingqian Tang, Tongliang Liu, Xinbo Gao 0001
IEEE Trans. Image Process.2
2022 CRS-CONT: A Well-Trained General Encoder for Facial Expression Analysis
abstract
Existing facial expression recognition (FER) methods train encoders with different large-scale training data for specific FER applications. In this paper, we propose a new task in this field. This task aims to pre-train a general encoder to extract any facial expression representations without fine-tuning. To tackle this task, we extend the self-supervised contrastive learning to pre-train a general encoder for facial expression analysis. To be specific, given a batch of facial expressions, some positive and negative pairs are firstly constructed based on coarse-grained labels and a FER-specified data augmentation strategy. Secondly, we propose the coarse-contrastive (CRS-CONT) learning, where the features of positive pairs are pulled together, while pushed away from the features of negative pairs. Moreover, one key event is that the excessive constraint on the coarse-grained feature distribution will affect fine-grained FER applications. To address this, a weight vector is designed to control the optimization of the CRS-CONT learning. As a result, a well-trained general encoder with frozen weights could preferably adapt to different facial expressions and realize the linear evaluation on any target datasets. Extensive experiments on both in- the-wild and in- the-lab FER datasets show that our method provides superior or comparable performance against state-of-the-art FER methods, especially on unseen facial expressions and cross-dataset evaluation. We hope that this work will help to reduce the training burden and develop a new solution against the fully-supervised feature learning with fine-grained labels. Code and the general encoder will be publicly available at https://github.com/hangyu94/CRS-CONT.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Image Process.2
2022 FRNet: Factorized and Regular Blocks Network for Semantic Segmentation in Road Scene
abstract
Nowadays, semantic segmentation methods for systems in road scene have a great demand. Most existing methods focus on high accuracy with low inference speed. And some approaches emphasize on speed, significantly sacrificing model accuracy. To make a trade-off between accuracy and inference speed, we propose a real-time network for semantic segmentation titled Factorized and Regular Network (FRNet), which employs an asymmetric encoder-decoder architecture with Factorized and Regular (FR) blocks. Our method achieves 70.4% mIoU on the Cityscapes test set with 1 million parameters at a speed of 127 frames per second (FPS) on a single Titan Xp at a resolution of$512\times 1024$. We evaluate FRNet on Cityscapes, Camvid, Kitti, and Gatech datasets to identify that our network stands out from other state-of-the-art networks.
Mengxu Lu, Zhenxue Chen, Q. M. Jonathan Wu, Nannan Wang 0001, Xuewen Rong, Xinghe Yan
IEEE Trans. Intell. Transp. Syst.4
2022 Towards Multi-Domain Face Synthesis Via Domain-Invariant Representations and Multi-Level Feature Parts
abstract
Cross-domain face synthesis plays a positive role in the real world. It is challenging to synthesize high-quality faces across multiple domains based on limited paired data because the multiple mappings between different domains may interfere with each other. Cognitive science investigates that the brain can recognize the same person with multiple different expressions by extracting invariant information on the face and we humans perceive instances by decomposing them into parts. Motivated by these cognition, we propose a unified semi-supervised framework for multi-domain face synthesis by extracting a domain-invariant representation and exploiting parts of multi-level features. Specifically, realized by adversarial training with additional ability to utilize domain-specific information, a encoder is trained to remove domain-specific information and extract the domain-invariant representation from multiple inputs. Then, we utilize the multi-level feature parts extracted from inputs and reconstructed faces via a pre-trained recognition model to ensure that the domain-invariant representation contains enough useful semantic information. we also utilize the feature parts extracted from inputs and limited paired data to compose pseudo features in target domain for supervising the synthesis, which makes our framework suitable for large amounts of unpaired training data. By exploiting this framework, we can achieve face synthesis between multiple domains using some paired data together with a large training database without ground truth target faces. Experimental results demonstrate our framework achieves great performances on qualitative and quantitative evaluations under both artificial and uncontrolled environments, and our framework has competitive performances in single translation compared with specialized methods for translation between two specific domains.
Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Multim.2
2022 Heterogeneous Face Interpretable Disentangled Representation for Joint Face Recognition and Synthesis
abstract
Heterogeneous faces are acquired with different sensors, which are closer to real-world scenarios and play an important role in the biometric security field. However, heterogeneous face analysis is still a challenging problem due to the large discrepancy between different modalities. Recent works either focus on designing a novel loss function or network architecture to directly extract modality-invariant features or synthesizing the same modality faces initially to decrease the modality gap. Yet, the former always lacks explicit interpretability, and the latter strategy inherently brings in synthesis bias. In this article, we explore to learn the plain interpretable representation for complex heterogeneous faces and simultaneously perform face recognition and synthesis tasks. We propose the heterogeneous face interpretable disentangled representation (HFIDR) that could explicitly interpret dimensions of face representation rather than simple mapping. Benefited from the interpretable structure, we further could extract latent identity information for cross-modality recognition and convert the modality factor to synthesize cross-modality faces. Moreover, we propose a multimodality heterogeneous face interpretable disentangled representation (M-HFIDR) to extend the basic approach suitable for the multimodality face recognition and synthesis. To evaluate the ability of generalization, we construct a novel large-scale face sketch data set. Experimental results on multiple heterogeneous face databases demonstrate the effectiveness of the proposed method.
Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Flexible Body Partition-Based Adversarial Learning for Visible Infrared Person Re-Identification
abstract
Person re-identification (Re-ID) aims to retrieve images of the same person across disjoint camera views. Most Re-ID studies focus on pedestrian images captured by visible cameras, without considering the infrared images obtained in the dark scenarios. Person retrieval between visible and infrared modalities is of great significance to public security. Current methods usually train a model to extract global feature descriptors and obtain discriminative representations for visible infrared person Re-ID (VI-REID). Nevertheless, they ignore the detailed information of heterogeneous pedestrian images, which affects the performance of Re-ID. In this article, we propose a flexible body partition (FBP) model-based adversarial learning method (FBP-AL) for VI-REID. To learn more fine-grained information, FBP model is exploited to automatically distinguish part representations according to the feature maps of pedestrian images. Specially, we design a modality classifier and introduce adversarial learning which attempts to discriminate features between visible and infrared modality. Adaptive weighting-based representation learning and threefold triplet loss-based metric learning compete with modality classification to obtain more effective modality-sharable features, thus shrinking the cross-modality gap and enhancing the feature discriminability. Extensive experimental results on two cross-modality person Re-ID data sets, i.e., SYSU-MM01 and RegDB, exhibit the superiority of the proposed method compared with the state-of-the-art solutions.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Wavelet-Based Dual Recursive Network for Image Super-Resolution
abstract
Although remarkable progress has been made on single-image super-resolution (SISR), deep learning methods cannot be easily applied to real-world applications due to the requirement of its heavy computation, especially for mobile devices. Focusing on the fewer parameters and faster inference SISR approach, we propose an efficient and time-saving wavelet transform-based network architecture, where the image super-resolution (SR) processing is carried out in the wavelet domain. Different from the existing methods that directly infer high-resolution (HR) image with the input low-resolution (LR) image, our approach first decomposes the LR image into a series of wavelet coefficients (WCs) and the network learns to predict the corresponding series of HR WCs and then reconstructs the HR image. Particularly, in order to further enhance the relationship between WCs and image deep characteristics, we propose two novel modules [wavelet feature mapping block (WFMB) and wavelet coefficients reconstruction block (WCRB)] and a dual recursive framework for joint learning strategy, thus forming a WCs prediction model to realize the efficient and accurate reconstruction of HR WCs. Experimental results show that the proposed method can outperform state-of-the-art methods with more than a 2× reduction in model parameters and computational complexity.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Heng Huang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Knowledge Distillation for Face Photo-Sketch Synthesis
abstract
Significant progress has been made with face photo-sketch synthesis in recent years due to the development of deep convolutional neural networks, particularly generative adversarial networks (GANs). However, the performance of existing methods is still limited because of the lack of training data (photo-sketch pairs). To address this challenge, we investigate the effect of knowledge distillation (KD) on training neural networks for the face photo-sketch synthesis task and propose an effective KD model to improve the performance of synthetic images. In particular, we utilize a teacher network trained on a large amount of data in a related task to separately learn knowledge of the face photo and knowledge of the face sketch and simultaneously transfer this knowledge to two student networks designed for the face photo-sketch synthesis task. In addition to assimilating the knowledge from the teacher network, the two student networks can mutually transfer their own knowledge to further enhance their learning. To further enhance the perception quality of the synthetic image, we propose a KD+ model that combines GANs with KD. The generator can produce images with more realistic textures and less noise under the guide of knowledge. Extensive experiments and a user study demonstrate the superiority of our models over the state-of-the-art methods.
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Training Binary Neural Network without Batch Normalization for Image Super-Resolution
abstract
Recently, binary neural network (BNN) based super-resolution (SR) methods have enjoyed initial success in the SR field. However, there is a noticeable performance gap between the binarized model and the full-precision one. Furthermore, the batch normalization (BN) in binary SR networks introduces floating-point calculations, which is unfriendly to low-precision hardwares. Therefore, there is still room for improvement in terms of model performance and efficiency. Focusing on this issue, in this paper, we first explore a novel binary training mechanism based on the feature distribution, allowing us to replace all BN layers with a simple training method. Then, we construct a strong baseline by combining the highlights of recent binarization methods, which already surpasses the state-of-the-arts. Next, to train highly accurate binarized SR model, we also develop a lightweight network architecture and a multi-stage knowledge distillation strategy to enhance the model representation ability. Extensive experiments demonstrate that the proposed method not only presents advantages of lower computation as compared to conventional floating-point networks but outperforms the state-of-the-art binary methods on the standard SR networks.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
AAAI2
2021 Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer
abstract
Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize complex style, especially when holistic global and local patterns exist. Inspired by the common painting process of drawing a draft and revising the details, we introduce a novel feed-forward method named Laplacian Pyramid Network (LapStyle). LapStyle first transfers global style patterns in low-resolution via a Drafting Network. It then revises the local details in high-resolution via a Revision Network, which hallucinates a residual image according to the draft and the image textures extracted by Laplacian filtering. Higher resolution details can be easily generated by stacking Revision Networks with multiple Laplacian pyramid levels. The final stylized image is obtained by aggregating outputs of all pyramid levels. Experiments demonstrate that our method can synthesize high quality stylized images in real time, where holistic style patterns are properly transferred.
Zhuoqi Ma, Fu Li 0003, Dongliang He, Xin Li 0106, Errui Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
CVPR7
2021 Support-Set Based Cross-Supervision for Video Grounding
abstract
Current approaches for video grounding propose kinds of complex architectures to capture the video-text relations, and have achieved impressive improvements. However, it is hard to learn the complicated multi-modal relations by only architecture designing in fact. In this paper, we introduce a novel Support-set Based Cross-Supervision (Sscs) module which can improve existing methods during training phase without extra inference cost. The proposed Sscs module contains two main components, i.e., discriminative contrastive objective and generative caption objective. The contrastive objective aims to learn effective representations by contrastive learning, while the caption objective can train a powerful video encoder supervised by texts. Due to the co-existence of some visual entities in both ground-truth and background intervals, i.e. mutual exclusion, naively contrastive learning is unsuitable to video grounding. We address the problem by boosting the cross-supervision with the support-set concept, which collects visual information from the whole video and eliminates the mutual exclusion of entities. Combined with the original objectives, Sscs can enhance the abilities of multi-modal relation modeling for existing approaches. We extensively evaluate Sscs on three challenging datasets, and show that our method can improve current state-of-the-art methods by large margins, especially 6.35% in terms of [email protected] on Charades-STA.
Xinpeng Ding, Nannan Wang 0001, Shiwei Zhang 0001, De Cheng, Xiaomeng Li 0001, Ziyuan Huang 0003, Mingqian Tang, Xinbo Gao 0001
ICCV2
2021 Syncretic Modality Collaborative Learning for Visible Infrared Person Re-Identification
abstract
Visible infrared person re-identification (VI-REID) aims to match pedestrian images between the daytime visible and nighttime infrared camera views. The large cross-modality discrepancies have become the bottleneck which limits the performance of VI-REID. Existing methods mainly focus on capturing cross-modality sharable representations by learning an identity classifier. However, the heterogeneous pedestrian images taken by different spectrum cameras differ significantly in image styles, resulting in inferior discriminability of feature representations. To alleviate the above problem, this paper explores the correlation between two modalities and proposes a novel syncretic modality collaborative learning (SMCL) model to bridge the cross-modality gap. A new modality that incorporates features of heterogeneous images is constructed automatically to steer the generation of modality-invariant representations. Challenge enhanced homogeneity learning (CEHL) and auxiliary distributional similarity learning (ADSL) are integrated to project heterogeneous features on a unified space and enlarge the inter-class disparity, thus strengthening the discriminative power. Extensive experiments on two cross-modality benchmarks demonstrate the effectiveness and superiority of the proposed method. Especially, on SYSU-MM01 dataset, our SMCL model achieves 67.39% rank-1 accuracy and 61.78% mAP, surpassing the cutting-edge works by a large margin.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
ICCV3
2021 Removing Adversarial Noise in Class Activation Feature Space
abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this problem, in this paper, we propose to remove adversarial noise by implementing a self-supervised adversarial training mechanism in a class activation feature space. To be specific, we first maximize the disruptions to class activation features of natural examples to craft adversarial examples. Then, we train a denoising model to minimize the distances between the adversarial examples and the natural examples in the class activation feature space. Empirical evaluations demonstrate that our method could significantly enhance adversarial robustness in comparison to previous state-of-the-art approaches, especially against unseen adversarial attacks and adaptive attacks.
Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001, Xiaoyu Wang 0002, Jun Yu 0001, Tongliang Liu
ICCV2
2021 Robust early-learning: Hindering the memorization of noisy labels
Xiaobo Xia, Tongliang Liu, Bo Han 0003, Chen Gong 0002, Nannan Wang 0001, ZongYuan Ge
ICLR5
2021 Class2Simi: A Noise Reduction Perspective on Learning with Noisy Labels
abstract
Learning with noisy labels has attracted a lot of attention in recent years, where the mainstream approaches are in \emph{pointwise} manners. Meanwhile, \emph{pairwise} manners have shown great potential in supervised metric learning and unsupervised contrastive learning. Thus, a natural question is raised: does learning in a pairwise manner \emph{mitigate} label noise? To give an affirmative answer, in this paper, we propose a framework called \emph{Class2Simi}: it transforms data points with noisy \emph{class labels} to data pairs with noisy \emph{similarity labels}, where a similarity label denotes whether a pair shares the class label or not. Through this transformation, the \emph{reduction of the noise rate} is theoretically guaranteed, and hence it is in principle easier to handle noisy similarity labels. Amazingly, DNNs that predict the \emph{clean} class labels can be trained from noisy data pairs if they are first pretrained from noisy data points. Class2Simi is \emph{computationally efficient} because not only this transformation is on-the-fly in mini-batches, but also it just changes loss computation on top of model prediction into a pairwise manner. Its effectiveness is verified by extensive experiments.
Songhua Wu, Xiaobo Xia, Tongliang Liu, Bo Han 0003, Mingming Gong, Nannan Wang 0001, Gang Niu 0001
ICML6
2021 Towards Defending against Adversarial Examples via Attack-Invariant Features
abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Their adversarial robustness can be improved by exploiting adversarial examples. However, given the continuously evolving attacks, models trained on seen types of adversarial examples generally cannot generalize well to unseen types of adversarial examples. To solve this problem, in this paper, we propose to remove adversarial noise by learning generalizable invariant features across attacks which maintain semantic classification information. Specifically, we introduce an adversarial feature learning mechanism to disentangle invariant features from adversarial noise. A normalization term has been proposed in the encoded space of the attack-invariant features to address the bias issue between the seen and unseen types of attacks. Empirical evaluations demonstrate that our method could provide better protection in comparison to previous state-of-the-art approaches, especially against unseen types of attacks and adaptive attacks.
Dawei Zhou 0004, Tongliang Liu, Bo Han 0003, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001
ICML4
2021 A Sketch-Transformer Network for Face Photo-Sketch Synthesis
abstract
We present a face photo-sketch synthesis model, which converts a face photo into an artistic face sketch or recover a photo-realistic facial image from a sketch portrait. Recent progress has been made by convolutional neural networks (CNNs) and generative adversarial networks (GANs), so that promising results can be obtained through real-time end-to-end architectures. However, convolutional architectures tend to focus on local information and neglect long-range spatial dependency, which limits the ability of existing approaches in keeping global structural information. In this paper, we propose a Sketch-Transformer network for face photo-sketch synthesis, which consists of three closely-related modules, including a multi-scale feature and position encoder for patch-level feature and position embedding, a self-attention module for capturing long-range spatial dependency, and a multi-scale spatially-adaptive de-normalization decoder for image reconstruction. Such a design enables the model to generate reasonable detail texture while maintaining global structural information. Extensive experiments show that the proposed method achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations.
Mingrui Zhu, Changcheng Liang, Nannan Wang 0001, Xiaoyu Wang 0002, Zhifeng Li 0001, Xinbo Gao 0001
IJCAI3
2021 Viewing from Frequency Domain: A DCT-based Information Enhancement Network for Video Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims to match the target pedestrians under non-overlapping camera system by video tracklets. The key issue of video Re-ID focuses on exploring effective spatio-temporal features. Generally, the spatio-temporal information of a video sequence can be divided into two aspects: the discriminative information in each frame and the shared information over the whole sequence. To make full use of the rich information in video sequences, this paper proposes a Discrete Cosine Transform based Information Enhancement Network (DCT-IEN) to achieve more comprehensive spatio-temporal representation from frequency domain. Inspired by the principle that average pooling is one of the special frequency components in DCT (the lowest frequency component), DCT-IEN first adopts discrete cosine transform to convert the extracted feature maps into frequency domain, thereby retaining more information that embedded in different frequency components. With the help of DCT frequency spectrum, two branches are adopted to learn the final video representation: Frequency Selection Module (FSM) and Lowest Frequency Enhancement Module (LFEM). FSM explores the most discriminative features in each frame by aggregating different frequency components with attention mechanism. LFEM enhances the shared feature over the whole video sequence by frame feature regularization. By fusing these two kinds of features together, DCT-IEN finally achieves comprehensive video representation. We conduct extensive experiments on two widely used datasets. The experimental results verify our idea and demonstrate the effectiveness of DCT-IEN for video-based Re-ID.
Liangchen Liu 0001, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
ACM Multimedia3
2021 CRNet: Centroid Radiation Network for Temporal Action Localization
Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
PRCV (1)2
2021 Weakly Supervised Temporal Action Localization with Segment-Level Labels
Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
PRCV (1)2
2021 Learning Deep Patch representation for Probabilistic Graphical Model-Based Face Sketch Synthesis
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.3
2021 LBAN-IL: A novel method of high discriminative representation for facial expression recognition
Hangyu Li 0001, Nannan Wang 0001, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001
Neurocomputing2
2021 FSFN: feature separation and fusion network for single image super-resolution
Zhenxue Chen, Q. M. Jonathan Wu, Nannan Wang 0001
Multim. Tools Appl.4
2021 Learning lightweight super-resolution networks with weight pruning
Nannan Wang 0001, Jingwei Xin, Xiaobo Xia, Xi Yang 0011, Xinbo Gao 0001
Neural Networks2
2021 A CenterNet++ model for ship detection in SAR images
Haoyuan Guo, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.3
2021 Iterative local re-ranking with attribute guided synthesis for face sketch recognition
Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Chunlei Peng, Jie Li 0001
Pattern Recognit.3
2021 Multi-Turn Video Question Generation via Reinforced Multi-Choice Attention Network
abstract
Video question generation is a challenging task in visual information retrieval, which generates questions given a sequence of video frames. The existing methods mainly tackle the problem of single-turn video question generation, but single-turn conversation usually can't meet the needs of video information acquisition. In this paper, we propose a new framework for single-turn VQG, which introduces attention mechanism to process inference of dialog history. And we introduce selection mechanism to choose from the candidate questions generated by each round of dialog history. In the framework, we leverage a recent video question answering model to predict the answer to the generated question and adopt the answer quality as rewards to fine-tune our model based on a reinforced learning mechanism. We also introduce a new task of multi-turn video question generation (M-VQG), which is generating multiple questions based on dialog history and video information to build conversation step by step. Our method achieves the state-of-the-art performance of the single-turn VQG task on two large-scale datasets, YouTube-Clips and TACoS-MultiLevel, and provides a baseline approach for M-VQG task.
Zhaoyu Guo, Zhou Zhao 0001, Weike Jin, Zhicheng Wei, Min Yang 0007, Nannan Wang 0001, Nicholas Jing Yuan
IEEE Trans. Circuits Syst. Video Technol.6
2021 Soft Semantic Representation for Cross-Domain Face Recognition
abstract
The problem of cross-domain face recognition aims to identify facial images obtained across different domains, which attracts increasing attentions because of its wide applications on law-enforcement identification and camera surveillance. The problem is challenging due to the huge domain discrepancy. Despite great progress achieved in recent years, existing algorithms usually fail to fully exploit the semantic information for identifying cross-domain faces, which could be a strong clue for recognition. In this article, we propose an effective algorithm for cross-domain face recognition by exploiting semantic information integrated with deep convolutional neural networks (CNN). We first introduce a soft face parsing algorithm where the boundaries of facial components are measured as probabilistic values. By taking the original face image as the guidance to improve face parsing result, each pixel may belong partially to the facial component to avoid inaccurate segmentation around component boundaries. We then propose a hierarchical soft semantic representation framework for cross-domain face recognition. Both the soft semantic level and contour level deep features obtained via CNN are computed and combined together, which could fully exploit the identical semantic clue among cross-domain faces. We provide extensive experiments to demonstrate that the proposed soft semantic representation algorithm performs superior against state-of-the-art methods.
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2021 KFC: An Efficient Framework for Semi-Supervised Temporal Action Localization
abstract
In temporal action localization (TAL), semi-supervised learning is a promising technique to mitigate the cost of precise boundary annotations. Semi-supervised approaches employing consistency regularization (CR), encouraging models to be robust to the perturbed inputs, have achieved great success in image classification problems. The success of CR is largely depended on the perturbations, where instances are perturbed to train a robust model without altering their semantic information. However, the perturbations for image or video classification tasks are not fit to apply to TAL. Since videos in TAL are too long to train the model with raw videos in an end-to-end manner. In this paper, we devise a method named K-farthest crossover to construct perturbations based on video features and apply it to TAL. Motivated by the observation that features in the same action instance become more and more similar during the training process while those in different action instances or backgrounds become more and more divergent, we add perturbations to each feature along temporal axis and adopt CR to encourage the model to retain this observation. Specifically, for a feature, we first find the top-k dissimilar features and average them to form a perturbation. Then, similar to chromosomal crossover, we select a large part of the feature and a small part of the perturbation to recombine a perturbed feature, which preserves the feature semantics yet enough discrepancy.
Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002, Tongliang Liu
IEEE Trans. Image Process.2
2021 Multi-Hierarchical Category Supervision for Weakly-Supervised Temporal Action Localization
abstract
Weakly Supervised Temporal Action Localization (WTAL) aims to localize action segments in untrimmed videos with only video-level category labels in the training phase. In WTAL, an action generally consists of a series of sub-actions, and different categories of actions may share the common sub-actions. However, to distinguish different categories of actions with only video-level class labels, current WTAL models tend to focus on discriminative sub-actions of the action, while ignoring those common sub-actions shared with different categories of actions. This negligence of common sub-actions would lead to the located action segments incomplete, i.e., only containing discriminative sub-actions. Different from current approaches of designing complex network architectures to explore more complete actions, in this paper, we introduce a novel supervision method named multi-hierarchical category supervision (MHCS) to find more sub-actions rather than only the discriminative ones. Specifically, action categories sharing similar sub-actions will be constructed as super-classes through hierarchical clustering. Hence, training with the new generated super-classes would encourage the model to pay more attention to the common sub-actions, which are ignored training with the original classes. Furthermore, our proposed MHCS is model-agnostic and non-intrusive, which can be directly applied to existing methods without changing their structures. Through extensive experiments, we verify that our supervision method can improve the performance of four state-of-the-art WTAL methods on three public datasets: THUMOS14, ActivityNet1.2, and ActivityNet1.3.
Guozhang Li, Jie Li 0001, Nannan Wang 0001, Xinpeng Ding, Zhifeng Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2021 Adaptively Learning Facial Expression Representation via C-F Labels and Distillation
abstract
Facial expression recognition is of significant importance in criminal investigation and digital entertainment. Under unconstrained conditions, existing expression datasets are highly class-imbalanced, and the similarity between expressions is high. Previous methods tend to improve the performance of facial expression recognition through deeper or wider network structures, resulting in increased storage and computing costs. In this paper, we propose a new adaptive supervised objective named AdaReg loss, re-weighting category importance coefficients to address this class imbalance and increasing the discrimination power of expression representations. Inspired by human beings' cognitive mode, an innovative coarse-fine (C-F) labels strategy is designed to guide the model from easy to difficult to classify highly similar representations. On this basis, we propose a novel training framework named the emotional education mechanism (EEM) to transfer knowledge, composed of a knowledgeable teacher network (KTN) and a self-taught student network (STSN). Specifically, KTN integrates the outputs of coarse and fine streams, learning expression representations from easy to difficult. Under the supervision of the pre-trained KTN and existing learning experience, STSN can maximize the potential performance and compress the original KTN. Extensive experiments on public benchmarks demonstrate that the proposed method achieves superior performance compared to current state-of-the-art frameworks with 88.07% on RAF-DB, 63.97% on AffectNet and 90.49% on FERPlus.
Hangyu Li 0001, Nannan Wang 0001, Xinpeng Ding, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Image Process.2
2021 A Two-Stream Dynamic Pyramid Representation Model for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) leverages rich spatio-temporal information embedded in sequence data to further improve the retrieval accuracy comparing with single image Re-ID. However, it also brings new difficulties. 1) Both spatial and temporal information should be considered simultaneously. 2) Pedestrian video data often contains redundant information and 3) suffers from data quality problems such as occlusion, background clutter. To solve the above problems, we propose a novel two-stream Dynamic Pyramid Representation Model (DPRM). DPRM mainly consists of three sub-models, i.e., Pyramidal Distribution Sampling Method (PDSM), Dynamic Pyramid Dilated Convolution (DPDC) and Pyramid Attention Pooling (PAP). PDSM is applied for more effective data pre-processing according to sequence semantic distribution. DPDC and PAP can be considered as two streams to describe the motion context and static appearance of a video sequence, respectively. By fusing the two-stream features together, we finally achieve comprehensive spatio-temporal representation. Notably, dynamic pyramid strategy is applied throughout the whole model. This strategy exploits multi-scale features under attention mechanism to maximally capture the most discriminative features and mitigate the impact of video data quality problems such as partial occlusion. Extensive experiments demonstrate the outperformance of DPRM. For instance, it achieves 83.0% mAP and 89.0% Rank-1 accuracy on MARS dataset and reaches state of the art.
Xi Yang 0011, Liangchen Liu 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2021 Estimating Human Pose Efficiently by Parallel Pyramid Networks
abstract
Good performance and high efficiency are both critical for estimating human pose in practice. Recent state-of-the-art methods have greatly boosted the pose detection accuracy through deep convolutional neural networks, however, the strong performance is typically achieved without high efficiency. In this paper, we design a novel network architecture for human pose estimation, which aims to strike a fine balance between speed and accuracy. Two essential tasks for successful pose estimation, preserving spatial location and extracting semantic information, are handled separately in the proposed architecture. Semantic knowledge of joint type is obtained through deep and wide sub-networks with low-resolution input, and high-resolution features indicating joint location are processed by shallow and narrow sub-networks. Because accurate semantic analysis mainly asks for adequate depth and width of the network and precise spatial information mostly requests preserving high-resolution features, good results can be produced by fusing the outputs of the sub-networks. Moreover, the computational cost can be considerably reduced comparing with existing networks, since the main part of the proposed network only deals with low-resolution features. We refer to the architecture as "parallel pyramid" network (PPNet), as features of different resolutions are processed at different levels of the hierarchical model. The superiority of our network is empirically demonstrated on two benchmark datasets: the MPII Human Pose dataset and the COCO keypoint detection dataset. PPNet outcompetes all recent methods by using less computation and memory to achieve better human pose estimation results.
Lin Zhao 0003, Nannan Wang 0001, Chen Gong 0002, Jian Yang 0003, Xinbo Gao 0001
IEEE Trans. Image Process.2
2020 Auto-GAN: Self-Supervised Collaborative Learning for Medical Image Synthesis
abstract
In various clinical scenarios, medical image is crucial in disease diagnosis and treatment. Different modalities of medical images provide complementary information and jointly helps doctors to make accurate clinical decision. However, due to clinical and practical restrictions, certain imaging modalities may be unavailable nor complete. To impute missing data with adequate clinical accuracy, here we propose a framework called self-supervised collaborative learning to synthesize missing modality for medical images. The proposed method comprehensively utilize all available information correlated to the target modality from multi-source-modality images to generate any missing modality in a single model. Different from the existing methods, we introduce an auto-encoder network as a novel, self-supervised constraint, which provides target-modality-specific information to guide generator training. In addition, we design a modality mask vector as the target modality label. With experiments on multiple medical image databases, we demonstrate a great generalization ability as well as specialty of our method compared with other state-of-the-arts.
Bing Cao 0002, Han Zhang 0002, Nannan Wang 0001, Xinbo Gao 0001, Dinggang Shen
AAAI3
2020 Facial Attribute Capsules for Noise Face Super Resolution
abstract
Existing face super-resolution (SR) methods mainly assume the input image to be noise-free. Their performance degrades drastically when applied to real-world scenarios where the input image is always contaminated by noise. In this paper, we propose a Facial Attribute Capsules Network (FACN) to deal with the problem of high-scale super-resolution of noisy face image. Capsule is a group of neurons whose activity vector models different properties of the same entity. Inspired by the concept of capsule, we propose an integrated representation model of facial information, which named Facial Attribute Capsule (FAC). In the SR processing, we first generated a group of FACs from the input LR face, and then reconstructed the HR face from this group of FACs. Aiming to effectively improve the robustness of FAC to noise, we generate FAC in semantic, probabilistic and facial attributes manners by means of integrated learning strategy. Each FAC can be divided into two sub-capsules: Semantic Capsule (SC) and Probabilistic Capsule (PC). Them describe an explicit facial attribute in detail from two aspects of semantic representation and probability distribution. The group of FACs model an image as a combination of facial attribute information in the semantic space and probabilistic space by an attribute-disentangling way. The diverse FACs could better combine the face prior information to generate the face images with fine-grained semantic attributes. Extensive benchmark experiments show that our method achieves superior hallucination results and outperforms state-of-the-art for very low resolution (LR) noise face image super resolution.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001, Zhifeng Li 0001
AAAI2
2020 Video Face Super-Resolution with Motion-Adaptive Feedback Cell
abstract
Video super-resolution (VSR) methods have recently achieved a remarkable success due to the development of deep convolutional neural networks (CNN). Current state-of-the-art CNN methods usually treat the VSR problem as a large number of separate multi-frame super-resolution tasks, at which a batch of low resolution (LR) frames is utilized to generate a single high resolution (HR) frame, and running a slide window to select LR frames over the entire video would obtain a series of HR frames. However, duo to the complex temporal dependency between frames, with the number of LR input frames increase, the performance of the reconstructed HR frames become worse. The reason is in that these methods lack the ability to model complex temporal dependencies and hard to give an accurate motion estimation and compensation for VSR process. Which makes the performance degrade drastically when the motion in frames is complex. In this paper, we propose a Motion-Adaptive Feedback Cell (MAFC), a simple but effective block, which can efficiently capture the motion compensation and feed it back to the network in an adaptive way. Our approach efficiently utilizes the information of the inter-frame motion, the dependence of the network on motion estimation and compensation method can be avoid. In addition, benefiting from the excellent nature of MAFC, the network can achieve better performance in the case of extremely complex motion scenarios. Extensive evaluations and comparisons validate the strengths of our approach, and the experimental results demonstrated that the proposed framework is outperform the state-of-the-art methods.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001, Zhifeng Li 0001
AAAI2
2020 Binarized Neural Network for Single Image Super Resolution
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Heng Huang 0001, Xinbo Gao 0001
ECCV (4)2
2020 ABP: Adaptive Body Partition Model For Visible Infrared Person Re-Identification
abstract
Person re-identification (Re-ID) aims to match pedestrian images cross multiple cameras. Most Re-ID studies focus on visible pedestrian images, without considering the images obtained by infrared cameras in the dark. To solve the cross-modality person Re-ID problem, current methods usually exploit global feature descriptors to obtain discriminative representations. However, they ignore the fine-grained information of heterogeneous images. In this paper, we propose an adaptive body partition (ABP) model to automatically detect and distinguish effective part representations. Instead of utilizing a two-stream convolutional neural network (CNN) to extract modality-specific information, we directly design an end-to-end one-stream CNN to simultaneously learn multi-modality sharable features and map them on a common space. Global loss, part losses and threefold triplet loss are integrated to enhance the feature discriminability and minimize the cross-modality gap. Extensive experimental results on two cross-modality Re-ID datasets exhibit the superiority of the proposed method compared with the state-of-the-art solutions.
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
ICME3
2020 Part-dependent Label Noise: Towards Instance-dependent Label Noise
abstract
Learning with the \textit{instance-dependent} label noise is challenging, because it is hard to model such real-world noise. Note that there are psychological and physiological evidences showing that we humans perceive instances by decomposing them into parts. Annotators are therefore more likely to annotate instances based on the parts rather than the whole instances, where a wrong mapping from parts to classes may cause the instance-dependent label noise. Motivated by this human cognition, in this paper, we approximate the instance-dependent label noise by exploiting \textit{part-dependent} label noise. Specifically, since instances can be approximately reconstructed by a combination of parts, we approximate the instance-dependent \textit{transition matrix} for an instance by a combination of the transition matrices for the parts of the instance. The transition matrices for parts can be learned by exploiting anchor points (i.e., data points that belong to a specific class almost surely). Empirical evaluations on synthetic and real-world datasets demonstrate our method is superior to the state-of-the-art approaches for learning from the instance-dependent label noise.
Xiaobo Xia, Tongliang Liu, Bo Han 0003, Nannan Wang 0001, Mingming Gong, Gang Niu 0001, Dacheng Tao, Masashi Sugiyama
NeurIPS4
2020 Feature Space Based Loss for Face Photo-Sketch Synthesis
Nannan Wang 0001, Xinbo Gao 0001
PRCV (1)2
2020 Image Super-Resolution via Deep Feature Recalibration Network
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
PRCV (1)3
2020 Learning Discriminative Joint Embeddings for Efficient Face and Voice Association
abstract
Many cognitive researches have shown the natural possibility of face-voice association, and such potential association has attracted much attention in biometric cross-modal retrieval domain. Nevertheless, the existing methods often fail to explicitly learn the common embeddings for challenging face-voice association tasks. In this paper, we present to learn discriminative joint embedding for face-voice association, which can seamlessly train the face subnetwork and voice subnetwork to learn their high-level semantic features, while correlating them to be compared directly and efficiently. Within the proposed approach, we introduce bi-directional ranking constraint, identity constraint and center constraint to learn the joint face-voice embedding, and adopt bi-directional training strategy to train the deep correlated face-voice model. Meanwhile, an online hard negative mining technique is utilized to discriminatively construct hard triplets in a mini-batch manner, featuring on speeding up the learning process. Accordingly, the proposed approach is adaptive to benefit various face-voice association tasks, including cross-modal verification, 1:2 matching, 1:N matching, and retrieval scenarios. Extensive experiments have shown its improved performances in comparison with the state-of-the-art ones.
Xin Liu 0011, Yiu-Ming Cheung, Nannan Wang 0001, Wentao Fan 0001
SIGIR5
2020 Image super-resolution via multi-view information fusion networks
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Yi Yu 0001, Xinbo Gao 0001
Neurocomputing2
2020 Semantic-related image style transfer with dual-consistency loss
Zhuoqi Ma, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neurocomputing3
2020 Image style transfer with collection representation space and semantic-guided reconstruction
Zhuoqi Ma, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks3
2020 Person Re-Identification with Feature Pyramid Optimization and Gradual Background Suppression
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Neural Networks3
2020 Modality adversarial neural network for visible-thermal person re-identification
Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.3
2020 HCNN-PSI: A hybrid CNN with partial semantic information for space target recognition
Xi Yang 0011, Tan Wu, Nannan Wang 0001, Yan Huang 0018, Bin Song 0001, Xinbo Gao 0001
Pattern Recognit.3
2020 A novel deformable body partition model for MMW suspicious object detection and dynamic tracking
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.3
2020 Bionic Face Sketch Generator
abstract
Face sketch synthesis is a crucial technique in digital entertainment. However, the existing face sketch synthesis approaches usually generate face sketches with coarse structures. The fine details on some facial components fail to be generated. In this paper, inspired by the artists during drawing face sketches, we propose a bionic face sketch generator. It includes three parts: 1) a coarse part; 2) a fine part; and 3) a finer part. The coarse part builds the facial structure of a sketch by a generative adversarial network in the U-Net. In the middle part, the noise produced by the coarse part is erased and the fine details on the important face components are generated via a probabilistic graphic model. To compensate for the fine sketch with distinctive edge and area of shadows and lights, we learn a mapping relationship at the high-frequency band by a convolutional neural network in the finer part. The experimental results show that the proposed bionic face sketch generator can synthesize the face sketch with more delicate and striking details, satisfy the requirement of users in the digital entertainment, and provide the students with the coarse, fine, and finer face sketch copies when learning sketches. Compared with the state-of-the-art methods, the proposed approach achieves better results in both visual effects and quantitative metrics.
Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Cybern.2
2020 A Rotational Libra R-CNN Method for Ship Detection
abstract
Recently, ship detection methods based on deep learning have attracted significant attention due to their superior accuracy over traditional methods. However, there still exist two problems affecting its robustness in practical application. 1) The size of ships in one image varies greatly, i.e., different sizes; 2) Numerous ships gather in limited field-of-view, i.e., dense distribution. To address these problems, we propose a rotational Libra R-convolutional neural network (CNN) method. Our idea is to balance the three levels of neural networks for predicting the location of ships with rotational angle information, which refers to the feature level, sample level, and objective level. First, to extract a discriminative feature and improve its robustness against the impact of different sizes of ships, the concept of balanced feature pyramid is introduced. Second, to generate reliable proposals for feature pyramid and efficiently mine hard negative samples, we employ intersection over union (IoU)-balanced sampling. Finally, to eliminate the redundant background and detect densely distributed ships, we bring in a rotational region detection branch with balanced L1 loss. In general, we develop the balanced learning with rotational region detection to achieve consistent improvement on accuracy and visualization. Experimental results on DOTA data set show that the proposed method achieves the state-of-the-art accuracy.
Haoyuan Guo, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Aurora Image Search With a Saliency-Weighted Region Network
abstract
On account of the remarkable performance of convolutional neural network (CNN) features for natural image searches, utilizing it for other images collected with the anamorphic lens has become a research hotspot. This article selects the aurora images generated from a circular fisheye lens as a typical example. By considering the imaging principle and geomagnetic information, a saliency-weighted region network (SWRN) is presented and introduced into the Mask R-CNN pipeline. Our SWRN selects salient regions with important semantic information and weights them both hierarchically and spatially. Hence, regions encompassing the search target are strengthened while uninformative regions are discarded, which benefits the suppression of background interference and reduction of computational complexity. In practice, by aggregating the outputs of SWRN with post-processing, a compact CNN feature is generated to represent the aurora image. Large-scale aurora image search experiments are conducted, and the results prove that our method performs better than the state-of-the-art methods on both accuracy and efficiency.
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Face Sketch Synthesis in the Wild via Deep Patch Representation-Based Probabilistic Graphical Model
abstract
This paper considers the problem of face sketch synthesis in the wild, which transforms a face photo into a face sketch. Face sketch synthesis is widely applied in law enforcement as well as digital entertainment fields. However, the existing methods either focus on hand-crafted techniques where prior human experience is relied on or adopt deep learning techniques as an end-to-end framework, where facial details cannot be well represented. In this paper, we propose a novel approach for face sketch synthesis in the wild via a deep patch representation-based probabilistic graphical model (DeepPGM). A Siamese network is constructed to extract deep patch representation from a raw facial patch, where the representative detail information for robust face sketch synthesis can be exploited. The generated deep patch representation and facial image patches are then optimally combined through a probabilistic graphical model. The proposed DeepPGM approach not only outperforms the state-of-the-art on public face sketch datasets but also can cope with forensic photos in the wild conditions, including varying lightings, poses, occlusions, skin colors, and ethnic origins. The superiority of the proposed method is demonstrated by extensive experiments on two public face sketch datasets and real-world forensic photos in the wild.
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2020 Group Feedback Capsule Network
abstract
In capsule networks (CapsNets), the capsule is made up of collections of neurons. Their adjacent capsule layers are connected using routing-by-agreement mechanisms in an unsupervised way. The routing-by-agreement mechanisms have two main drawbacks: a) too many parameters and high computation complexity; b) the cluster distribution assumptions of these routing mechanisms may not hold in some complex real-world data. In this paper, we propose a novel Group Feedback Capsule Network (GF-CapsNet) which adopts a supervised routing strategy called group-routing. Compared with the previous routing strategies which globally transform each capsule, Group-routing equally splits capsules into groups where capsules locally share the same transformation weights, reducing routing parameters. To address the second drawback, we devise a distance network to directly predict capsules in a supervised way without making distribution assumptions. Our proposed group-routing captures local information of low-level capsules by group-wise transformation and supervisedly predicts high-level ones in a feedback way to address two drawbacks respectively. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts.
Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002, Tongliang Liu
IEEE Trans. Image Process.2
2020 Universal Face Photo-Sketch Style Transfer via Multiview Domain Translation
abstract
Face photo-sketch style transfer aims to convert a representation of a face from the photo (or sketch) domain to the sketch (respectively, photo) domain while preserving the character of the subject. It has wide-ranging applications in law enforcement, forensic investigation and digital entertainment. However, conventional face photo-sketch synthesis methods usually require training images from both the source domain and the target domain, and are limited in that they cannot be applied to universal conditions where collecting training images in the source domain that match the style of the test image is unpractical. This problem entails two major challenges: 1) designing an effective and robust domain translation model for the universal situation in which images of the source domain needed for training are unavailable, and 2) preserving the facial character while performing a transfer to the style of an entire image collection in the target domain. To this end, we present a novel universal face photo-sketch style transfer method that does not need any image from the source domain for training. The regression relationship between an input test image and the entire training image collection in the target domain is inferred via a deep domain translation framework, in which a domain-wise adaption term and a local consistency adaption term are developed. To improve the robustness of the style transfer process, we propose a multiview domain translation method that flexibly leverages a convolutional neural network representation with hand-crafted features in an optimal way. Qualitative and quantitative comparisons are provided for universal unconstrained conditions of unavailable training images from the source domain, demonstrating the effectiveness and superiority of our method for universal face photo-sketch style transfer.
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2020 CGAN-TM: A Novel Domain-to-Domain Transferring Method for Person Re-Identification
abstract
Person re-identification (re-ID) is a technique aiming to recognize person cross different cameras. Although some supervised methods have achieved favorable performance, they are far from practical application owing to the lack of labeled data. Thus, unsupervised person re-ID methods are in urgent need. Generally, the commonly used approach in existing unsupervised methods is to first utilize the source image dataset for generating a model in supervised manner, and then transfer the source image domain to the target image domain. However, images may lose their identity information after translation, and the distributions between different domains are far away. To solve these problems, we propose an image domain-to-domain translation method by keeping pedestrian's identity information and pulling closer the domains' distributions for unsupervised person re-ID tasks. Our work exploits the CycleGAN to transfer the existing labeled image domain to the unlabeled image domain. Specially, a Self-labeled Triplet Net is proposed to maintain the pedestrian identity information, and maximum mean discrepancy is introduced to pull the domain distribution closer. Extensive experiments have been conducted and the results demonstrate that the proposed method performs superiorly than the state-ofthe- art unsupervised methods on DukeMTMC-reID and Market- 1501.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2020 A Novel Symmetry Driven Siamese Network for THz Concealed Object Verification
abstract
Security inspection aims to improve the high detection rate as well as reduce the false alarm rate. However, it still suffers from two challenges affecting its robustness. 1) Existing security inspection methods are mostly designed for natural images, which cannot reflect the uniqueness and imaging principle of THz images. 2) Existing methods is sensitive to noise interference and pose variations. This work revisits these challenges and presents a novel symmetry driven Siamese network (SDSN) for THz concealed object verification. Our idea is to employ a specially designed network architecture for THz concealed object verification. First, to reflect the uniqueness and the special property of THz images, Siamese network with Contrastive loss is used for feature extraction along with symmetrical prior information consideration, which can learn symmetrical metrics from the same person. Second, to alleviate the impact of noise interference and pose variations, the adaptive identity normalization (A-IDN) is proposed to normalize the symmetrical metrics each person. Finally, to enhance the generalization of network, an adaptive selective threshold based on Gaussian mixture model (AST-GMM) is designed, which serves as a classifier for the final classification results. Extensive experiments show that SDSN significantly improves the accuracy. Specially, SDSN outperforms the state-of-the-art methods without symmetrical prior information on THz security dataset.
Xi Yang 0011, Haoyuan Guo, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2020 Cascaded Face Sketch Synthesis Under Various Illuminations
abstract
Face sketch synthesis from a photo is of significant importance in digital entertainment. An intelligent face sketch synthesis system requires a strong robustness to lighting variations. Under uncontrolled lighting conditions in real-world settings, such a system will perform consistently well and have little restriction on the lighting conditions. However, previous face sketch synthesis methods tend to synthesize sketches under well-controlled lighting conditions. These methods are sensitive to lighting variations and produce unsatisfactory results when the lighting condition varies. In this paper, we propose a novel cascaded face sketch synthesis framework composed of a multiple feature generator and a cascaded low-rank representation. The multiple feature generator not only produces a generated sketch feature consistent with an artist's drawing style but also extracts a photo feature that is robust to various illuminations. Both features ensure that given a photo patch, the optimal sketch candidates can be selected from the database. The cascaded low-rank representation enables a gradual reduction in the gap between the synthesized face sketch and the corresponding artistdrawn sketch. Experimental results illustrate that the proposed cascaded framework generates realistic sketches on par with the current methods on the Chinese University of Hong Kong face sketch database under well-controlled illuminations. Moreover, this framework exhibits greatly improved performance compared to these methods on the extended Chinese University of Hong Kong face sketch database and Chinese celebrity face photos from the web under different illuminations. We argue that this framework paves a novel way for the implementation of computer-aided optical systems that are of essential importance in both face sketch synthesis and optical imaging.
Mingjin Zhang, Yunsong Li 0001, Nannan Wang 0001, Yuan Chi, Xinbo Gao 0001
IEEE Trans. Image Process.3
2020 Coupled Attribute Learning for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) is a challenging problem in face recognition and subject to large textural and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. In addition, limited training samples of cross-modality pairs make HFR more challenging due to the complex generation procedure of these images. Despite the great progress that has been achieved in recent years, existing works mainly focus on HFR from only cross-modality image matching. However, it is more practical to obtain both facial images and semantic descriptions about facial attributes in real-world situations, in which the semantic description clues are nearly always obtained during the process of image generation. Motivated by human cognitive mechanisms, we naturally utilize the explicit invariant semantic description, i.e., face attributes, to help address the gap among face images of different modalities. Existing facial attributes-related face recognition methods primarily regard attributes as the high-level features used to enhance recognition performance, ignoring the inherent relationship between face attributes and identities. In this article, we propose novel coupled attribute learning for the HFR (CAL-HFR) method without labeling the attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances are taken as dissimilarities of pairs. Coupled attribute guided triplet loss (CAGTL) is designed to train an end-to-end HFR network that can effectively eliminate defects of incorrectly estimated attributes. Extensive experiments on multiple heterogeneous scenarios demonstrate that the proposed method achieves superior performance compared with that of state-of-the-art methods. Furthermore, we make publicly available our generated pairwise annotated heterogeneous facial attribute database for evaluation and promoting related research.
Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Chunlei Peng
IEEE Trans. Neural Networks Learn. Syst.3
2020 Neural Probabilistic Graphical Model for Face Sketch Synthesis
abstract
Neural network learning for face sketch synthesis from photos has attracted substantial attention due to its favorable synthesis performance. However, most existing deep-learning-based face sketch synthesis models stacked only by multiple convolutional layers without structured regression often lose the common facial structures, limiting their flexibility in a wide range of practical applications, including intelligent security and digital entertainment. In this article, we introduce a neural network to a probabilistic graphical model and propose a novel face sketch synthesis framework based on the neural probabilistic graphical model (NPGM) composed of a specific structure and a common structure. In the specific structure, we investigate a neural network for mapping the direct relationship between training photos and sketches, yielding the specific information and characteristic features of a test photo. In the common structure, the fidelity between the sketch pixels generated by the specific structure and their candidates selected from the training data are considered, ensuring the preservation of the common facial structure. Experimental results on the Chinese University of Hong Kong face sketch database demonstrate, both qualitatively and quantitatively, that the proposed NPGM-based face sketch synthesis approach can more effectively capture specific features and recover common structures compared with the state-of-the-art methods. Extensive experiments in practical applications further illustrate that the proposed method achieves superior performance.
Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 HSME: Hypersphere Manifold Embedding for Visible Thermal Person Re-Identification
abstract
Person Re-identification(re-ID) has great potential to contribute to video surveillance that automatically searches and identifies people across different cameras. Heterogeneous person re-identification between thermal(infrared) and visible images is essentially a cross-modality problem and important for night-time surveillance application. Current methods usually train a model by combining classification and metric learning algorithms to obtain discriminative and robust feature representations. However, the combined loss function ignored the correlation between classification subspace and feature embedding subspace. In this paper, we use Sphere Softmax to learn a hypersphere manifold embedding and constrain the intra-modality variations and cross-modality variations on this hypersphere. We propose an end-to-end dualstream hypersphere manifold embedding network(HSMEnet) with both classification and identification constraint. Meanwhile, we design a two-stage training scheme to acquire decorrelated features, we refer the HSME with decorrelation as D-HSME. We conduct experiments on two crossmodality person re-identification datasets. Experimental results demonstrate that our method outperforms the state-of-the-art methods on two datasets. On RegDB dataset, rank-1 accuracy is improved from 33.47% to 50.85%, and mAP is improved from 31.83% to 47.00%.
Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
AAAI2
2019 Residual Attribute Attention Network for Face Image Super-Resolution
abstract
Facial prior knowledge based methods recently achieved great success on the task of face image super-resolution (SR). The combination of different type of facial knowledge could be leveraged for better super-resolving face images, e.g., facial attribute information with texture and shape information. In this paper, we present a novel deep end-to-end network for face super resolution, named Residual Attribute Attention Network (RAAN), which realizes the efficient feature fusion of various types of facial information. Specifically, we construct a multi-block cascaded structure network with dense connection. Each block has three branches: Texture Prediction Network (TPN), Shape Generation Network (SGN) and Attribute Analysis Network (AAN). We divide the task of face image reconstruction into three steps: extracting the pixel level representation information from the input very low resolution (LR) image via TPN and SGN, extracting the semantic level representation information by AAN from the input, and finally combining the pixel level and semantic level information to recover the high resolution (HR) image. Experiments on benchmark database illustrate that RAAN significantly outperforms state-of-the-arts for very low-resolution face SR problem, both quantitatively and qualitatively.
Jingwei Xin, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
AAAI2
2019 Fast Semantic Preserving Hashing for Large-Scale Cross-Modal Retrieval
abstract
Most Cross-modal hashing methods do not sufficiently exploit the discrimination power of semantic information when learning hash codes, while often involving time-consuming training procedures for large-scale dataset. To tackle these issues, we first formulate the learning of similarity-preserving hash codes in terms of orthogonally rotating the semantic data to hamming space, and then propose a novel Fast Semantic Preserving Hashing (FSePH) approach to large-scale cross-modal retrieval. Specifically, FSePH introduces an orthonormal basis to regress the targeted hash codes of training examples to their corresponding reasonably relaxed class labels, featuring significantly reducing the quantization error. Meanwhile, an effective optimization algorithm is derived for modality-specific projection function learning and an efficient closed-form solution for hash code learning, which are computationally tractable. Extensive experiments have shown that the proposed FSePH approach runs sufficiently fast, and also significantly improves the retrieval performances over the state-of-the-arts.
Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Zhikai Hu, Nannan Wang 0001
ICDM6
2019 Person re-Identification with Gradual Background Suppression
abstract
Person re-identification plays an important role in public security. However, owing to the interference of background clutters, its performance still needs to be improved. Several mask-based methods aim to solve this problem by totally removing the background clutters, but the promotion is limited because of the mask sharpening effect. In this paper, we propose a novel person re-identification method with Gradual Background Suppression (GBS). The GBS adopts several CNN branches to extract deep features of images with different weight distributions between background and human body. Thus, it can not only reduce the background clutters but also keep the smoothness of target pedestrians. Afterwards, deep features from different CNN branches are integrated with a fusion scheme, and the fused feature is capable of balancing the influence of background clutter and mask sharpening. Extensive experiments have been conducted and the results prove the superiority of the proposed GBS over the background removal approach. Comparing with the state-of-the-art methods, our method achieves remarkable performance with 6.6%, 7.58% and 8.26% improvement of mAP on dataset Market-1501, CUHK03-labeled and CUHK03-detected, respectively.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
ICME3
2019 Semi-Supervised Semantic-Preserving Hashing for Efficient Cross-Modal Retrieval
abstract
Cross-modal hashing has recently gained significant popularity to facilitate retrieval across different modalities. With limited label available, this paper presents a novel Semi-Supervised Semantic-Preserving Hashing (S3PH) for flexible cross-modal retrieval. In contrast to most semi-supervised cross-modal hashing works that need to predict the label of unlabeled data, our proposed approach groups the labeled and unlabeled data together, and integrates the relaxed latent subspace learning and semantic-preserving regularization across different modalities. Accordingly, an efficient relaxed objective function is proposed to learn the latent subspaces for both labeled and unlabeled data. Further, an orthogonal rotation matrix is efficiently learned to transform the latent subspace to hash space by minimizing the quantization error. Without sacrificing the retrieval performance, the proposed S3PH method can benefit various kinds of retrieval tasks, i.e., unsupervised, semi-supervised and supervised. Experimental results compared with several competitive algorithms show the effectiveness of the proposed method and its superiority over state-of-the-arts.
Xingzhi Wang, Xin Liu 0011, Zhikai Hu, Nannan Wang 0001, Wentao Fan 0001, Jixiang Du
ICME4
2019 T-SCNN: A Two-Stage Convolutional Neural Network for Space Target Recognition
abstract
Space target recognition plays an important role in the field of space security and exploration. With the rapid development of artificial intelligence technique and explosive increase of image dataset, object recognition based on deep learning has achieved favorable performance. However, the recognition of deep space targets in visible spectrum images still remains in the traditional manual interpretation approach, thus leading to low efficiency and inevitable subjective errors. In this paper, we propose an artificial intelligence method for space target recognition, called Two-Stage Convolutional Neural Network (T-SCNN). Our T-SCNN is composed of two stages, i.e., target locating and target recognition. In the stage of target locating, we first detect all suspected targets from the total image dataset by presenting a minimum bounding rectangle with threshold (MBRT) approach, then cut out all regions encompassing targets to generate target images for training. In the stage of target recognition, we send target images to the well-trained recognition network for identification. Additionally, data augmentation is conducted in the CNN training to satisfy its data quantity requirement. Extensive experiments are performed on our synthetic space target image dataset, and the result demonstrate that the proposed method achieves high accuracy within a short time.
Tan Wu, Xi Yang 0011, Bin Song 0001, Nannan Wang 0001, Xinbo Gao 0001, Liyang Kuang, Xiaoting Nan, Dong Yang 0012
IGARSS4
2019 Multi-Margin based Decorrelation Learning for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) refers to matching face images acquired from different domains with wide applications in security scenarios. However, HFR is still a challenging problem due to the significant cross-domain discrepancy and the lacking of sufficient training data in different domains. This paper presents a deep neural network approach namely Multi-Margin based Decorrelation Learning (MMDL) to extract decorrelation representations in a hyperspherical space for cross-domain face images. The proposed framework can be divided into two components: heterogeneous representation network and decorrelation representation learning. First, we employ a large scale of accessible visual face images to train heterogeneous representation network. The decorrelation layer projects the output of the first component into decorrelation latent subspace and obtain decorrelation representation. In addition, we design a multi-margin loss (MML), which consists of tetradmargin loss (TML) and heterogeneous angular margin loss (HAML), to constrain the proposed framework. Experimental results on two challenging heterogeneous face databases show that our approach achieves superior performance on both verification and recognition tasks, comparing with state-of-the-art methods.
Bing Cao 0002, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Zhifeng Li 0001
IJCAI2
2019 Group Reconstruction and Max-Pooling Residual Capsule Network
abstract
In capsule networks, the mapping of low-level capsules to high-level capsules is achieved by a routing-by-agreement algorithm. Since the capsule is made up of collections of neurons and the routing mechanism involves all the capsules instead of simply discarding some of the neurons like Max-Pooling, the capsule network has stronger representation ability than the traditional neural network. However, considering too much low-level capsules' information will cause its corresponding upper layer capsules to be interfered by other irrelevant information or noise capsules. Therefore, the original capsule network does not perform well on complex data structure. What's worse, computational complexity becomes a bottleneck in dealing with large data networks. In order to solve these shortcomings, this paper proposes a group reconstruction and max-pooling residual capsule network (GRMR-CapsNet). We build a block in which all capsules are divided into different groups and perform group reconstruction routing algorithm to obtain the corresponding high-level capsules. Between the lower and higher layers, Capsule Max-Pooling is adopted to prevent overfitting. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts.
Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002
IJCAI2
2019 Face Photo-Sketch Synthesis via Knowledge Transfer
abstract
Despite deep neural networks have demonstrated strong power in face photo-sketch synthesis task, their performance, however, are still limited by the lack of training data (photo-sketch pairs). Knowledge Transfer (KT), which aims at training a smaller and fast student network with the information learned from a larger and accurate teacher network, has attracted much attention recently due to its superior performance in the acceleration and compression of deep neural networks. This work has brought us great inspiration that we can train a relatively small student network on very few training data by transferring knowledge from a larger teacher model trained on enough training data for other tasks. Therefore, we propose a novel knowledge transfer framework to synthesize face photos from face sketches or synthesize face sketches from face photos. Particularly, we utilize two teacher networks trained on large amount of data in related task to learn the knowledge of face photos and face sketches separately and transfer them to two student networks simultaneously. In addition, the two student networks, one for photo ? sketch task and the other for sketch ? photo task, can transfer their knowledge mutually. With the proposed method, we can train our model which has superior performance using a small set of photo-sketch pairs. We validate the effectiveness of our method across several datasets. Quantitative and qualitative evaluations illustrate that our model outperforms other state-of-the-art methods in generating face sketches (or photos) with high visual quality and recognition ability.
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Zhifeng Li 0001
IJCAI2
2019 Triplet Fusion Network Hashing for Unpaired Cross-Modal Retrieval
abstract
With the dramatic increase of multi-media data on the Internet, cross-modal retrieval has become an important and valuable task in searching systems. The key challenge of this task is how to build the correlation between multi-modal data. Most existing approaches only focus on dealing with paired data. They use pairwise relationship of multi-modal data for exploring the correlation between them. However, in practice, unpaired data are more common on the Internet but few methods pay attention to them. To utilize both paired and unpaired data, we propose a one-stream framework triplet fusion network hashing (TFNH), which mainly consists of two parts. The first part is a triplet network which is used to handle both kinds of data, with the help of zero padding operation. The second part consists of two data classifiers, which are used to bridge the gap between paired and unpaired data. In addition, we embed manifold learning into the framework for preserving both inter and intra modal similarity, exploring the relationship between unpaired and paired data and bridging the gap between them in learning process. Extensive experiments show that the proposed approach outperforms several state-of-the-art methods on two datasets in paired scenario. We further evaluate its ability of handling unpaired scenario and robustness in regard to pairwise constraint. The results show that even we discard 50% data under the setting in [19], the performance of TFNH is still better than that of other unpaired approaches and that only 70% pairwise relationships are preserved, TFNH can still outperform almost all paired approaches.
Zhikai Hu, Xin Liu 0011, Xingzhi Wang, Yiu-Ming Cheung, Nannan Wang 0001, Yewang Chen
ICMR5
2019 Dual-alignment Feature Embedding for Cross-modality Person Re-identification
abstract
Person re-identification aims at searching pedestrians across different cameras, which is a key problem in video surveillance. With requirements in night environment, RGB-infrared person re-identification which could be regarded as a cross-modality matching problem, has gained increasing attention in recent years. Aside from cross-modality discrepancy, RGB-infrared person re-identification also suffers from human pose and view point differences. We design a dual-alignment feature embedding method to extract discriminative modality-invariant features. The concept of dual-alignment is two folds: spatial and modality alignments. We adopt the part-level features to extract fine-grained camera-invariant information. We introduce distribution loss function and correlation loss function to align the embedding features across visible and infrared modalities. Finally, we can extract modality-invariant features with robust and rich identity embeddings for cross-modality person re-identification. Experiment confirms that the proposed baseline and improvement achieves competitive results with the state-of-the-art methods on two datasets. For instance, We achieve (57.5+12.6)% rank-1 accuracy and (57.3+11.8)% mAP on the RegDB dataset.
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002
ACM Multimedia2
2019 Are Anchor Points Really Indispensable in Label-Noise Learning?
abstract
In label-noise learning, the \textit{noise transition matrix}, denoting the probabilities that clean labels flip into noisy labels, plays a central role in building \textit{statistically consistent classifiers}. Existing theories have shown that the transition matrix can be learned by exploiting \textit{anchor points} (i.e., data points that belong to a specific class almost surely). However, when there are no anchor points, the transition matrix will be poorly learned, and those previously consistent classifiers will significantly degenerate. In this paper, without employing anchor points, we propose a \textit{transition-revision} ($T$-Revision) method to effectively learn transition matrices, leading to better classifiers. Specifically, to learn a transition matrix, we first initialize it by exploiting data points that are similar to anchor points, having high \textit{noisy class posterior probabilities}. Then, we modify the initialized matrix by adding a \textit{slack variable}, which can be learned and validated together with the classifier by using noisy data. Empirical results on benchmark-simulated and real-world label-noise datasets demonstrate that without using exact anchor points, the proposed method is superior to state-of-the-art label-noise learning methods.
Xiaobo Xia, Tongliang Liu, Nannan Wang 0001, Bo Han 0003, Chen Gong 0002, Gang Niu 0001, Masashi Sugiyama
NeurIPS3
2019 BoSR: A CNN-based aurora image retrieval method
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Neural Networks2
2019 DLFace: Deep local descriptor for cross-modality face recognition
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
Pattern Recognit.2
2019 Sparse graphical representation based discriminant analysis for heterogeneous face recognition
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001
Signal Process.3
2019 CNN with spatio-temporal information for fast suspicious object detection and recognition in THz security images
Xi Yang 0011, Tan Wu, Lei Zhang 0019, Dong Yang 0012, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.5
2019 Re-Ranking High-Dimensional Deep Local Representation for NIR-VIS Face Recognition
abstract
Heterogeneous face recognition refers to matching facial images captured from different sensors or sources, which has wide applications in public security and law enforcement. Because of the great differences in sensing and creating procedure, there are huge feature gap between heterogeneous facial images. Existing methods merely focus on comparing the probe image with the gallery in feature space, while the true target may not appear at the first rank due to the appearance variations caused by different sensing patterns. In order to exploit valuable information from initial ranking result, this paper proposes to re-rank high-dimensional deep local representation for matching near-infrared (NIR) and visual (VIS) facial images, i.e. NIR-VIS face recognition. A high-dimensional deep local representation is firstly constructed by extracting and concatenating deep features on local facial patches via a convolutional neural network (CNN). The initial NIR-VIS recognition ranking results can be obtained by comparing the compressed deep features. We then propose a novel and efficient locally linear re-ranking (LLRe-Rank) technique to refine the initial ranking results, which can explore valuable information from initial ranking result. The proposed re-ranking method does not require any human interaction or data annotation, and can be served as an unsupervised post processing technique. Experimental results on the most challenging Oulu-CASIA NIR-VIS database and CASIA NIR-VIS 2.0 database demonstrate the effectiveness of our method.
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2019 Data Augmentation-Based Joint Learning for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) is the process of matching face images captured from different sources. HFR plays an important role in security scenarios. However, HFR remains a challenging problem due to the considerable discrepancies (i.e., shape, style, and color) between cross-modality images. Conventional HFR methods utilize only the information involved in heterogeneous face images, which is not effective because of the substantial differences between heterogeneous face images. To better address this issue, this paper proposes a data augmentation-based joint learning (DA-JL) approach. The proposed method mutually transforms the cross-modality differences by incorporating synthesized images into the learning process. The aggregated data augments the intraclass scale, which provides more discriminative information. However, this method also reduces the interclass diversity (i.e., discriminative information). We develop the DA-JL model to balance this dilemma. Finally, we obtain the similarity score between heterogeneous face image pairs through the log-likelihood ratio. Extensive experiments on a viewed sketch database, forensic sketch database, near-infrared image database, thermal-infrared image database, low-resolution photo database, and image with occlusion database illustrate that the proposed method achieves superior performance in comparison with the state-of-the-art methods.
Bing Cao 0002, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 Deep Latent Low-Rank Representation for Face Sketch Synthesis
abstract
Face sketch synthesis is useful and profitable in digital entertainment. Most existing face sketch synthesis methods rely on the assumption that facial photographs/sketches form a low-dimensional manifold. Once the training data are insufficient, the manifold could not characterize the identity-specific information that is included in a test photograph but excluded in the training data. Thus, the synthesized sketch would lose this information, such as glasses, earrings, hairstyles, and hairpins. To provide the sufficient data and satisfy the assumption on manifold, we propose a novel face sketch synthesis framework based on deep latent low-rank representation (DLLRR) in this paper. The DLLRR induces the hidden training sketches with the identity-specific information as the hidden data to the insufficient original training sketches as the observed data. And it searches the lowest rank representation on the candidates of a test photograph from the both hidden and observed data. For the strong representational capability of the coupled autoencoder, we leverage it to reveal the hidden data. Experiment results on face photograph-sketch database illustrate that the proposed method can successfully provide the sufficient training data with the identity-specific information. And compared to the state of the arts, the proposed method synthesizes more clean and vivid face sketches.
Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 A Deep Collaborative Framework for Face Photo-Sketch Synthesis
abstract
Great breakthroughs have been made in the accuracy and speed of face photo-sketch synthesis in recent years. Regression-based methods have gained increasing attention, which benefit from deeper and faster end-to-end convolutional neural networks. However, most of these models typically formulate the mapping from photo domain X to sketch domain Y as a unidirectional feedforward mapping, G: X → Y , and vice versa, F: Y → X ; thus, the utilization of mutual interaction between two opposite mappings is lacking. Therefore, we proposed a collaborative framework for face photo-sketch synthesis. The concept behind our model was that a middle latent domain ~Z between the photo domain X and the sketch domain Y can be learned during the learning procedure of G: X → Y and F: Y → X by introducing a collaborative loss that makes full use of two opposite mappings. This strategy can constrain the two opposite mappings and make them more symmetrical, thus making the network more suitable for the photo-sketch synthesis task and obtaining higher quality generated images. Qualitative and quantitative experiments demonstrated the superior performance of our model in comparison with the existing state-of-the-art solutions.
Mingrui Zhu, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Asymmetric Joint Learning for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) refers to matching a probe face image taken from one modality to face images acquired from another modality. It plays an important role in security scenarios. However, HFR is still a challenging problem due to great discrepancies between cross-modality images. This paper proposes an asymmetric joint learning (AJL) approach to handle this issue. The proposed method transforms the cross-modality differences mutually by incorporating the synthesized images into the learning process which provides more discriminative information. Although the aggregated data would augment the scale of intra-classes, it also reduces the diversity (i.e. discriminative information) for inter-classes. Then, we develop the AJL model to balance this dilemma. Finally, we could obtain the similarity score between two heterogeneous face images through the log-likelihood ratio. Extensive experiments on viewed sketch database, forensic sketch database and near infrared image database illustrate that the proposed AJL-HFR method achieve superior performance in comparison to state-of-the-art methods.
Bing Cao 0002, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
AAAI2
2018 Face Sketch Synthesis From Coarse to Fine
abstract
Synthesizing fine face sketches from photos is a valuable yet challenging problem in digital entertainment. Face sketches synthesized by conventional methods usually exhibit coarse structures of faces, whereas fine details are lost especially on some critical facial components. In this paper, by imitating the coarse-to-fine drawing process of artists, we propose a novel face sketch synthesis framework consisting of a coarse stage and a fine stage. In the coarse stage, a mapping relationship between face photos and sketches is learned via the convolutional neural network. It ensures that the synthesized sketches keep coarse structures of faces. Given the test photo and the coarse synthesized sketch, a probabilistic graphic model is designed to synthesize the delicate face sketch which has fine and critical details. Experimental results on public face sketch databases illustrate that our proposed framework outperforms the state-of-the-art methods in both quantitive and visual comparisons.
Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Ruxin Wang 0002, Xinbo Gao 0001
AAAI2
2018 Saliency Deep Embedding for Aurora Image Search
abstract
Deep neural networks have achieved remarkable success in the field of image search. However, the state-of-the-art algorithms are trained and tested for natural images captured with ordinary cameras. In this paper, we aim to explore a new search method for images captured with circular fisheye lens, especially the aurora images. To reduce the interference from uninformative regions and focus on the most interested regions, we propose a saliency proposal network (SPN) to replace the region proposal network (RPN) in the recent Mask R-CNN. In our SPN, the centers of the anchors are not distributed in a rectangular meshing manner, but exhibit spherical distortion. Additionally, the directions of the anchors are along the deformation lines perpendicular to the magnetic meridian, which perfectly accords with the imaging principle of circular fisheye lens. Extensive experiments are performed on the big aurora data, demonstrating the superiority of our method in both search accuracy and efficiency.
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Nannan Wang 0001, Dong Yang 0012
ICME4
2018 Deep Attribute Guided Representation for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) is a challenging problem in face recognition, subject to large texture and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. Motivated by human cognitive mechanism, we naturally utilize the explicit invariant semantic information (face attributes) to help address the gap of different modalities. Existing related face recognition methods mostly regard attributes as the high level feature integrated with other engineering features enhancing recognition performance, ignoring the inherent relationship between face attributes and identities. In this paper, we propose a novel deep attribute guided representation based heterogeneous face recognition method (DAG-HFR) without labeling attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances mean similarities of pairs. An attribute guided triplet loss (AGTL) is designed to train an end-to-end HFR network which could effectively eliminate defects of incorrectly detected attributes. Extensive experiments on multiple heterogeneous scenarios (composite sketches, resident ID cards) demonstrate that the proposed method achieves superior performances compared with state-of-the-art methods.
Decheng Liu, Nannan Wang 0001, Chunlei Peng, Jie Li 0001, Xinbo Gao 0001
IJCAI2
2018 From Reality to Perception: Genre-Based Neural Image Style Transfer
abstract
We introduce a novel thought for integrating artists’ perceptions on the real world into neural image style transfer process. Conventional approaches commonly migrate color or texture patterns from style image to content image, but the underlying design aspect of the artist always get overlooked. We want to address the in-depth genre style, that how artists perceive the real world and express their perceptions in the artwork. We collect a set of Van Gogh’s paintings and cubist artworks, and their semantically corresponding real world photos. We present a novel genre style transfer framework modeled after the mechanism of actual artwork production. The target style representation is reconstructed based on the semantic correspondence between real world photo and painting, which enable the perception guidance in style transfer. The experimental results demonstrate that our method can capture the overall style of a genre or an artist. We hope that this work provides new insight for including artists’ perceptions into neural style transfer process, and helps people to understand the underlying characters of the artist or the genre.
Zhuoqi Ma, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
IJCAI2
2018 Markov Random Neural Fields for Face Sketch Synthesis
abstract
Synthesizing face sketches with both common and specific information from photos has been recently attracting considerable attentions in digital entertainment. However, the existing approaches either make the strict similarity assumption on face sketches and photos, leading to lose some identity-specific information, or learn the direct mapping relationship from face photos to sketches by the simple neural network, resulting in the lack of some common information. In this paper, we propose a novel face sketch synthesis based on the Markov random neural fields including two structures. In the first structure, we utilize the neural network to learn the non-linear photo-sketch relationship and obtain the identity-specific information of the test photo, such as glasses, hairpins and hairstyles. In the second structure, we choose the nearest neighbors of the test photo patch and the sketch pixel synthesized in the first structure from the training data which ensure the common information of Miss or Mr Average. Experimental results on the Chinese University of Hong Kong face sketch database illustrate that our proposed framework can preserve the common structure and capture the characteristic features. Compared with the state-of-the-art methods, our method achieves better results in terms of both quantitative and qualitative experimental evaluations.
Mingjin Zhang, Nannan Wang 0001, Xinbo Gao 0001, Yunsong Li 0001
IJCAI2
2018 Composite components-based face sketch recognition
Decheng Liu, Jie Li 0001, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001
Neurocomputing3
2018 Facial feature point detection: A comprehensive survey
Nannan Wang 0001, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001
Neurocomputing1
2018 Face recognition from multiple stylistic sketches: Scenarios, datasets, and evaluation
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001
Pattern Recognit.3
2018 Random sampling for fast face sketch synthesis
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
Pattern Recognit.1
2018 Back projection: An effective postprocessing method for GAN-based face sketch synthesis
Nannan Wang 0001, Wenjin Zha, Jie Li 0001, Xinbo Gao 0001
Pattern Recognit. Lett.1
2018 Anchored Neighborhood Index for Face Sketch Synthesis
abstract
Exemplar-based face sketch synthesis has long been impeded by the difficulty of accurate neighbor selection. Given a test patch extracted from the test photograph, the K-nearest neighbor (K-NN) matching algorithm is generally performed by existing methods to find K-nearest photograph patches in the training data set, which contains some pairs of face sketches and photographs. Then, the training sketch patches corresponding to the selected nearest photograph patches are taken as the candidate to synthesize the target sketch patch. In the aforementioned neighbor selection process, training sketch patches is not taken into consideration in the process of K-nearest neighbor selection. In this paper, we proposed a simple yet effective neighbor selection algorithm, namely, anchored neighborhood index (ANI), to boost the synthesis performance by taking training sketch patches into the consideration. In addition, the proposed ANI can be conducted offline and, thus, it does not increase the computational complexity. Extensive experiments on public available database demonstrate that the proposed algorithm achieves superior performance compared with the state-of-the-art methods in terms of both objective image quality scores and face recognition accuracy.
Nannan Wang 0001, Xinbo Gao 0001, Leiyu Sun, Jie Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Compositional Model-Based Sketch Generator in Facial Entertainment
abstract
Face sketch synthesis (FSS) plays an important role in facial entertainment, which includes face sketch morphing among two styles, multiview FSS and face sketch expression manipulation. For facial entertainment, most existing FSS methods generate sketches with over-smoothing effects, i.e., fine details are suppressed more or less. In this paper, we propose a face sketch generator based on the compositional model to handle this issue. It decomposes a face into different components instead of patches as before, and each component has several candidate templates. Multilevel B-spline approximation is utilized to delicately polish the chosen templates of all components. To fuse these components, Poisson blending is employed instead of the weighted average operator. The proposed compositional method crucially reduces the high frequency loss and improves the synthesis performance in comparison to the state-of-the-art methods. Experiments on face sketch morphing, expression manipulation, and multiview FSS, make further efforts to demonstrate the effectiveness of the proposed method.
Mingjin Zhang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Cybern.3
2017 Semantic Segmentation Based Automatic Two-Tone Portrait Synthesis
Zhuoqi Ma, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
ICIG (3)2
2017 Deep Graphical Feature Learning for Face Sketch Synthesis
abstract
The exemplar-based face sketch synthesis method generally contains two steps: neighbor selection and reconstruction weight representation. Pixel intensities are widely used as features by most of the existing exemplar-based methods, which lacks of representation ability and robustness to light variations and clutter backgrounds. We present a novel face sketch synthesis method combining generative exemplar-based method and discriminatively trained deep convolutional neural networks (dCNNs) via a deep graphical feature learning framework. Our method works in both two steps by using deep discriminative representations derived from dCNNs. Instead of using it directly, we boost its representation capability by a deep graphical feature learning framework. Finally, the optimal weights of deep representations and optimal reconstruction weights for face sketch synthesis can be obtained simultaneously. With the optimal reconstruction weights, we can synthesize high quality sketches which is robust against light variations and clutter backgrounds. Extensive experiments on public face sketch databases show that our method outperforms state-of-the-art methods, in terms of both synthesis quality and recognition ability.
Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
IJCAI2
2017 Adaptive representation-based face sketch-photo synthesis
Jie Li 0001, Xinye Yu, Chunlei Peng, Nannan Wang 0001
Neurocomputing4
2017 Data-driven vs. model-driven: Fast face sketch synthesis
Nannan Wang 0001, Mingrui Zhu, Jie Li 0001, Bin Song 0001, Zan Li 0001
Neurocomputing1
2017 Graphical Representation for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) refers to matching face images acquired from different sources (i.e., different sensors or different wavelengths) for identification. HFR plays an important role in both biometrics research and industry. In spite of promising progresses achieved in recent years, HFR is still a challenging problem due to the difficulty to represent two heterogeneous images in a homogeneous manner. Existing HFR methods either represent an image ignoring the spatial information, or rely on a transformation procedure which complicates the recognition task. Considering these problems, we propose a novel graphical representation based HFR method (G-HFR) in this paper. Markov networks are employed to represent heterogeneous image patches separately, which takes the spatial compatibility between neighboring image patches into consideration. A coupled representation similarity metric (CRSM) is designed to measure the similarity between obtained graphical representations. Extensive experiments conducted on multiple HFR scenarios (viewed sketch, forensic sketch, near infrared image, and thermal infrared image) show that the proposed method outperforms state-of-the-art methods.
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Unified framework for face sketch synthesis
Nannan Wang 0001, Shengchuan Zhang, Xinbo Gao 0001, Jie Li 0001, Bin Song 0001, Zan Li 0001
Signal Process.1
2017 Superpixel-Based Face Sketch-Photo Synthesis
abstract
Face sketch-photo synthesis technique has attracted growing attention in many computer vision applications, such as law enforcement and digital entertainment. Existing methods either simply perform the face sketch-photo synthesis on the holistic image or divide the face image into regular rectangular patches ignoring the inherent structure of the face image. In view of such situations, this paper presents a novel superpixel-based face sketch-photo synthesis method by estimating the face structures through image segmentation. In our proposed method, face images are first segmented into superpixels, which are then dilated to enhance the compatibility of neighboring superpixels. Each input face image induces a specific graphical structure modeled by Markov networks. We employ a two-stage synthesis process to learn the face structures through Markov networks constructed from two scales of dilation, respectively. Experiments on several public databases demonstrate that our proposed face sketch-photo synthesis method achieves superior performance compared with the state-of-the-art methods.
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2017 Face Sketch Synthesis From a Single Photo-Sketch Pair
abstract
Face sketch synthesis is crucial in many practical applications, such as digital entertainment and law enforcement. Previous methods relying on many photo-sketch pairs have made great progress. State-of-the-art face sketch synthesis algorithms adopt Bayesian inference (BI) (e.g., Markov random fields) to select local sketch patches around corresponding position from a set of training data. However, these methods have two limitations: 1) they depend on many training photo-sketch pairs and 2) they cannot tackle nonfacial factors (e.g., hairpins, glasses, backgrounds, and image size) if these factors are excluded in training data. In this paper, we propose a novel face sketch synthesis method that is capable of handling nonfacial factors only using a single photo-sketch pair from coarse to fine. Our method proposes a cascaded image synthesis (CIS) strategy and integrates sparse representation-based greedy search (SRGS) and BI for face sketch synthesis. We first apply SRGS to select candidate sketch patches from the whole training photo-sketch pairs sampled from the only photo-sketch pair. We then employ BI to estimate an initial sketch. Afterward, the input photo and the estimated initial sketch are taken as an additional photo-sketch pair for training. Finally, we adopt CIS with the given two photo-sketch pairs to further improve the quality of the initial sketch. The experimental results on several databases demonstrate that our algorithm outperforms state-of-the-art methods.
Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2017 Bayesian Face Sketch Synthesis
abstract
Exemplar-based face sketch synthesis has been widely applied to both digital entertainment and law enforcement. In this paper, we propose a Bayesian framework for face sketch synthesis, which provides a systematic interpretation for understanding the common properties and intrinsic difference in different methods from the perspective of probabilistic graphical models. The proposed Bayesian framework consists of two parts: the neighbor selection model and the weight computation model. Within the proposed framework, we further propose a Bayesian face sketch synthesis method. The essential rationale behind the proposed Bayesian method is that we take the spatial neighboring constraint between adjacent image patches into consideration for both aforementioned models, while the state-of-the-art methods neglect the constraint either in the neighbor selection model or in the weight computation model. Extensive experiments on the Chinese University of Hong Kong face sketch database demonstrate that the proposed Bayesian method could achieve superior performance compared with the state-of-the-art methods in terms of both subjective perceptions and objective evaluations.
Nannan Wang 0001, Xinbo Gao 0001, Leiyu Sun, Jie Li 0001
IEEE Trans. Image Process.1
2016 Evaluation on synthesized face sketches
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Bin Song 0001, Zan Li 0001
Neurocomputing1
2016 SERF: A Simple, Effective, Robust, and Fast Image Super-Resolver From Cascaded Linear Regression
abstract
Example learning-based image super-resolution techniques estimate a high-resolution image from a low-resolution input image by relying on high- and low-resolution image pairs. An important issue for these techniques is how to model the relationship between high- and low-resolution image patches: most existing complex models either generalize hard to diverse natural images or require a lot of time for model training, while simple models have limited representation capability. In this paper, we propose a simple, effective, robust, and fast (SERF) image super-resolver for image super-resolution. The proposed super-resolver is based on a series of linear least squares functions, namely, cascaded linear regression. It has few parameters to control the model and is thus able to robustly adapt to different image data sets and experimental settings. The linear least square functions lead to closed form solutions and therefore achieve computationally efficient implementations. To effectively decrease these gaps, we group image patches into clusters via k-means algorithm and learn a linear regressor for each cluster at each iteration. The cascaded learning process gradually decreases the gap of high-frequency detail between the estimated high-resolution image patch and the ground truth image patch and simultaneously obtains the linear regression parameters. Experimental results show that the proposed method achieves superior performance with lower time consumption than the state-of-the-art methods.
Yanting Hu, Nannan Wang 0001, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001
IEEE Trans. Image Process.2
2016 Robust Face Sketch Style Synthesis
abstract
Heterogeneous image conversion is a critical issue in many computer vision tasks, among which example-based face sketch style synthesis provides a convenient way to make artistic effects for photos. However, existing face sketch style synthesis methods generate stylistic sketches depending on many photo-sketch pairs. This requirement limits the generalization ability of these methods to produce arbitrarily stylistic sketches. To handle such a drawback, we propose a robust face sketch style synthesis method, which can convert photos to arbitrarily stylistic sketches based on only one corresponding template sketch. In the proposed method, a sparse representation-based greedy search strategy is first applied to estimate an initial sketch. Then, multi-scale features and Euclidean distance are employed to select candidate image patches from the initial estimated sketch and the template sketch. In order to further refine the obtained candidate image patches, a multi-feature-based optimization model is introduced. Finally, by assembling the refined candidate image patches, the completed face sketch is obtained. To further enhance the quality of synthesized sketches, a cascaded regression strategy is adopted. Compared with the state-of-the-art face sketch synthesis methods, experimental results on several commonly used face sketch databases and celebrity photos demonstrate the effectiveness of the proposed method.
Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Image Process.3
2016 Multiple Representations-Based Face Sketch-Photo Synthesis
abstract
Face sketch-photo synthesis plays an important role in law enforcement and digital entertainment. Most of the existing methods only use pixel intensities as the feature. Since face images can be described using features from multiple aspects, this paper presents a novel multiple representations-based face sketch-photo-synthesis method that adaptively combines multiple representations to represent an image patch. In particular, it combines multiple features from face images processed using multiple filters and deploys Markov networks to exploit the interacting relationships between the neighboring image patches. The proposed framework could be solved using an alternating optimization strategy and it normally converges in only five outer iterations in the experiments. Our experimental results on the Chinese University of Hong Kong (CUHK) face sketch database, celebrity photos, CUHK Face Sketch FERET Database, IIIT-D Viewed Sketch Database, and forensic sketches demonstrate the effectiveness of our method for face sketch-photo synthesis. In addition, cross-database and database-dependent style-synthesis evaluations demonstrate the generalizability of this novel method and suggest promising solutions for face identification in forensic science.
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Dacheng Tao, Xuelong Li 0001, Jie Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2015 Recognition of facial sketch styles
Mingjin Zhang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neurocomputing3
2015 Face Sketch Synthesis via Sparse Representation-Based Greedy Search
abstract
Face sketch synthesis has wide applications in digital entertainment and law enforcement. Although there is much research on face sketch synthesis, most existing algorithms cannot handle some nonfacial factors, such as hair style, hairpins, and glasses if these factors are excluded in the training set. In addition, previous methods only work on well controlled conditions and fail on images with different backgrounds and sizes as the training set. To this end, this paper presents a novel method that combines both the similarity between different image patches and prior knowledge to synthesize face sketches. Given training photo-sketch pairs, the proposed method learns a photo patch feature dictionary from the training photo patches and replaces the photo patches with their sparse coefficients during the searching process. For a test photo patch, we first obtain its sparse coefficient via the learnt dictionary and then search its nearest neighbors (candidate patches) in the whole training photo patches with sparse coefficients. After purifying the nearest neighbors with prior knowledge, the final sketch corresponding to the test photo can be obtained by Bayesian inference. The contributions of this paper are as follows: 1) we relax the nearest neighbor search area from local region to the whole image without too much time consuming and 2) our method can produce nonfacial factors that are not contained in the training set and is robust against image backgrounds and can even ignore the alignment and image size aspects of test photos. Our experimental results show that the proposed method outperforms several state-of-the-arts in terms of perceptual and objective metrics.
Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Mingjin Zhang
IEEE Trans. Image Process.3
2014 A Comprehensive Survey to Face Hallucination
Nannan Wang 0001, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Jie Li 0001
Int. J. Comput. Vis.1
2013 Heterogeneous image transformation
Nannan Wang 0001, Jie Li 0001, Dacheng Tao, Xuelong Li 0001, Xinbo Gao 0001
Pattern Recognit. Lett.1
2013 Transductive Face Sketch-Photo Synthesis
abstract
Face sketch-photo synthesis plays a critical role in many applications, such as law enforcement and digital entertainment. Recently, many face sketch-photo synthesis methods have been proposed under the framework of inductive learning, and these have obtained promising performance. However, these inductive learning-based face sketch-photo synthesis methods may result in high losses for test samples, because inductive learning minimizes the empirical loss for training samples. This paper presents a novel transductive face sketch-photo synthesis method that incorporates the given test samples into the learning process and optimizes the performance on these test samples. In particular, it defines a probabilistic model to optimize both the reconstruction fidelity of the input photo (sketch) and the synthesis fidelity of the target output sketch (photo), and efficiently optimizes this probabilistic model by alternating optimization. The proposed transductive method significantly reduces the expected high loss and improves the synthesis performance for test samples. Experimental results on the Chinese University of Hong Kong face sketch data set demonstrate the effectiveness of the proposed method by comparing it with representative inductive learning-based face sketch-photo synthesis methods.
Nannan Wang 0001, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Jie Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2012 Face Sketch-Photo Synthesis and Retrieval Using Sparse Representation
abstract
Sketch-photo synthesis plays an important role in sketch-based face photo retrieval and photo-based face sketch retrieval systems. In this paper, we propose an automatic sketch-photo synthesis and retrieval algorithm based on sparse representation. The proposed sketch-photo synthesis method works at patch level and is composed of two steps: sparse neighbor selection (SNS) for an initial estimate of the pseudoimage (pseudosketch or pseudophoto) and sparse-representation-based enhancement (SRE) for further improving the quality of the synthesized image. SNS can find closely related neighbors adaptively and then generate an initial estimate for the pseudoimage. In SRE, a coupled sparse representation model is first constructed to learn the mapping between sketch patches and photo patches, and a patch-derivative-based sparse representation method is subsequently applied to enhance the quality of the synthesized photos and sketches. Finally, four retrieval modes, namely, sketch-based, photo-based, pseudosketch-based, and pseudophoto-based retrieval are proposed, and a retrieval algorithm is developed by using sparse representation. Extensive experimental results illustrate the effectiveness of the proposed face sketch-photo synthesis and retrieval algorithms.
Xinbo Gao 0001, Nannan Wang 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 Face Sketch-Photo Synthesis under Multi-dictionary Sparse Representation Framework
abstract
Sketch-photo synthesis is one of the important research issues of heterogeneous image transformation. Some available popular synthesis methods, like locally linear embedding (LLE), usually generate sketches or photos with lower definition and blurred details, which reduces the visual quality and the recognition rate across the heterogeneous images. In order to improve the quality of the synthesized images, a multi-dictionary sparse representation based face sketch-photo synthesis model is constructed. In the proposed model, LLE is used to estimate an initial sketch or photo, while the multi-dictionary sparse representation model is applied to generate the high frequency and detail information. Finally, by linear superimposing, the enhanced face sketch or photo can be obtained. Experimental results show that sketches and photos synthesized by the proposed method have higher definition and much richer detail information resulting in a higher face recognition rate between sketches and photos.
Nannan Wang 0001, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001
ICIG1
2011 Face sketch-photo synthesis based on support vector regression
abstract
The existing face sketch-photo synthesis methods trend to lose some vital details more or less. In this paper, we propose a novel sketch-photo synthesis approach based on support vector regression (SVR) to handle this difficulty. First, we utilize an existing method to acquire the initial estimate of the synthesized image. Then, the final synthesized image is obtained by combining the initial estimate and the SVR based high frequency information together to further enhance the quality of synthesized image. Experimental results on the benchmark database and our new constructed database demonstrate that the proposed method can achieve significant improvement on perceptual quality. Moreover, the synthesized face images can obtain higher recognition rate when used in retrieval system.
Jiewei Zhang, Nannan Wang 0001, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001
ICIP2