Weipeng Hu

dblp:124/5693 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0003-2886-7346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Adaptive memory refinement and perception enhancement for exo-to-ego video generation
Weipeng Hu, Jiun Tian Hoe, Ping Hu 0001, Xudong Jiang 0001, Yap-Peng Tan
Neurocomputing2
2026 HP-Gaussian: Head Prior-Guided Gaussian Splatting for Personalized Talking Head Synthesis From Few-Second Video
abstract
Gaussian Splatting-based talking head synthesis has made significant progress in recent years, yet existing methods often struggle with generalization beyond specific training identity. In this paper, we propose Head Prior guided Gaussian Splatting for personalized talking head synthesis (HP-Gaussian) that can generalize to new identities with only few training data. Unlike traditional optimization-based Gaussian Splatting methods, our approach directly predicts Gaussian parameters from multi-modal inputs, including audio and visual cues. This feed-forward design enables multiple identities pre-training, allowing the model to learn shared head priors from large-scale datasets, while supporting flexible speaker-specific adaptation. To further enhance Gaussian feature learning, we introduce a Spatial Gaussian Transformer that captures correlations among neighboring Gaussians, improving parameter estimation accuracy. Additionally, recognizing the critical importance of personalized speaking styles in the synthesis of high-quality talking videos, a two-stage training strategy is implemented. A base model is initially trained across diverse identities to establish the foundational head prior knowledge. Subsequently, we introduce the short-video personalized adaptation phase for more realistic customized talking video generation. Extensive experiments demonstrate that our HP-Gaussian can synthesize high-fidelity and personalized talking videos with remarkably few training examples, setting a new benchmark for efficiency and quality in talking head synthesis. We highly recommend viewing our demonstration video at https://youtu.be/RpjWdvikKhU for intuitive visual comparisons and qualitative results.
Shuai Shen, Wanhua Li 0001, Weipeng Hu, Jiwen Lu, Yap-Peng Tan
IEEE Trans. Image Process.4
2026 Learning Action Distribution Flow for Open-Set Temporal Action Segmentation
abstract
In this paper, we tackle the open-set temporal action segmentation task, which aims to identify unknown frames while ensuring accurate segmentation of known actions in the temporal domain. Existing open-set methods struggle with identifying unknown frames due to their indistinguishability against ambiguous known frames during action transitions, resulting in significant performance degradation. To address this, we propose the action distribution flow, which models transitions between action sequences to capture the inherent feature discrepancies between unknown and known frames. Specifically, our method first models the distributions of known actions using the training data, and then interpolates these distributions along the optimal transport path for consecutive actions in the testing videos. By evaluating the likelihood of testing frames against the modeled action distribution flow, our approach effectively identifies unknown frames without requiring additional training or prior knowledge of the unknown data. Extensive experiments on open-set versions of the GTEA, 50Salads, and Breakfast datasets demonstrate the superiority of the proposed method across all evaluation metrics.
Runzhong Zhang, Fengrui Tian, Yueqi Duan, Ziwei Wang 0010, Weipeng Hu, Peijun Bao, Suchen Wang, Yap-Peng Tan
IEEE Trans. Image Process.6
2025 E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Haifeng Hu 0001, Yap-Peng Tan
ACM Multimedia3
2025 Cascaded Dynamic Memory Refinement and Semantic Alignment for Exo-to-Ego Cross-View Video Generation
abstract
Cross-view video generation from exocentric (third-person) to egocentric (first-person) perspectives poses a challenging task, due to the significant viewpoint gap and limited overlap between these two views. Previous methods exhibit limitations in capturing long-range temporal context and overlook egocentric semantic priors, leading to degraded performance in cross-view synthesis. To address these challenges, we propose a cue-free video-based approach termed cascaded Dynamic memory Refinement and Semantic Alignment (DRSA), which integrates temporal knowledge over extended periods and learns egocentric semantic information to generate videos. The Dynamic Memory Refinement (DMR) exploits long horizon temporal dynamics to learn salient information that compensates for the limited overlap between views. Specifically, we devise a dynamic memory that serves as a knowledge repository, and utilize a sliding window to locate the corresponding long-term temporal information, which is subsequently processed with adaptive weighting and cross-attention transformer to refine feature representations. Furthermore, aware of the considerable viewpoint divergence that hinder semantic learning of target view, we propose Viewpoint-aware Semantic Alignment (VSA) with dual encoder-decoder learning and semantic alignment, which transfer egocentric semantic details from the egocentric synthesis pipeline to the exocentric synthesis pipeline. In particular, the VSA module narrows the semantic gap between views, further promoting long-range temporal modeling in DMR under alignment constraints. By extending this into a cascaded fashion, the Cascaded Alignment and Refinement (CAR) progressively aligns semantic features and performs feature refinement to facilitate viewpoint learning at different levels of granularity. To overcome the limitations of existing databases known for their limited static scenes and scarcity of interacting objects, we create a new dataset with dynamic exocentric scenes and rich interacting objects to further promote the task. Thorough experimental analysis reveals that our method surpasses current state-of-the-art techniques in terms of both quantitative metrics and qualitative evaluations.
Weipeng Hu, Jiun Tian Hoe, Haifeng Hu 0001, Xudong Jiang 0001, Yap-Peng Tan
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation
abstract
Timestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets.
Runzhong Zhang, Yueqi Duan, Weipeng Hu, Suchen Wang, Yap-Peng Tan
IEEE Trans. Circuits Syst. Video Technol.4
2025 Progressive Cross-Modal Association Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) aims to explore the cross-modal associations and learn modality-invariant representations without manual labels. The field provides flexible and economical methods for person re-identification across light and dark scenes. Existing approaches utilize cluster-level strong association methods, such as graph matching and optimal transport, to correlate modal differences, which may result in mis-linking between clusters and introduce noise. To overcome this limitation and gradually acquire reliable cross-modal associations, we propose a Progressive Cross-modal Association Learning (PCAL) method for USL-VI-ReID. Specifically, our PCAL naturally integrates Triple-modal Adversarial Learning (TAL), Cross-modal Neighbor Expansion (CNE) and Modality-invariant Contrastive Learning (MCL) into a unified framework. TAL fully utilizes the advantage of Channel Augmented (CA) technique to reduce modal differences, which facilitates subsequent mining of cross-modal associations. Furthermore, we identify the modal bias problem in existing clustering methods, which hinders the effective establishment of cross-modal associations. To address this problem, CNE is proposed to balance the contribution of cross-modal neighbor information, linking potential cross-modal neighbors as much as possible. Finally, MCL is then introduced to refine the cross-modal associations and learn modality-invariant representations. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the competitive performance of PCAL method. Code is available at https://github.com/YimingYang23/PCA USLVIReID.
Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Residual Quotient Learning for Zero-Reference Low-Light Image Enhancement
abstract
Recently, neural networks have become the dominant approach to low-light image enhancement (LLIE), with at least one-third of them adopting a Retinex-related architecture. However, through in-depth analysis, we contend that this most widely accepted LLIE structure is suboptimal, particularly when addressing the non-uniform illumination commonly observed in natural images. In this paper, we present a novel variant learning framework, termed residual quotient learning, to substantially alleviate this issue. Instead of following the existing Retinex-related decomposition-enhancement-reconstruction process, our basic idea is to explicitly reformulate the light enhancement task as adaptively predicting the latent quotient with reference to the original low-light input using a residual learning fashion. By leveraging the proposed residual quotient learning, we develop a lightweight yet effective network called ResQ-Net. This network features enhanced non-uniform illumination modeling capabilities, making it more suitable for real-world LLIE tasks. Moreover, due to its well-designed structure and reference-free loss function, ResQ-Net is flexible in training as it allows for zero-reference optimization, which further enhances the generalization and adaptability of our entire framework. Extensive experiments on various benchmark datasets demonstrate the merits and effectiveness of the proposed residual quotient learning, and our trained ResQ-Net outperforms state-of-the-art methods both qualitatively and quantitatively. Furthermore, a practical application in dark face detection is explored, and the preliminary results confirm the potential and feasibility of our method in real-world scenarios.
Linfeng Fei, Huanjie Tao, Yaocong Hu, Wei Zhou 0042, Jiun Tian Hoe, Weipeng Hu, Yap-Peng Tan
IEEE Trans. Image Process.7
2025 Snippet-Inter Difference Attention Network for Weakly-Supervised Temporal Action Localization
abstract
The purpose of weakly-supervised temporal action localization (WTAL) task is to simultaneously classify and localize action instances in untrimmed videos with only video-level labels. Previous works fail to extract multi-scale temporal features to identify action instances with different durations, and they do not fully use the temporal cues of action video to learn discriminative features. In addition, the classifiers trained by current methods usually focus on easy-to-distinguish snippets while ignoring other semantically ambiguous features, which leads to incomplete and over-complete localization. To address these issues, we introduce a new Snippet-inter Difference Attention Network (SDANet) for WTAL, which can be trained end-to-end. Specifically, our model presents three modules, with primary contributions lying in the snippet-inter difference attention (SDA) module and potential feature mining (PFM) module. Firstly, we construct a simple multi-scale temporal feature fusion (MTFF) module to generate multi-scale temporal feature representation, so as to help the model better detect short action instances. Secondly, we consider the temporal cues of video features and design SDA module based on the Transformer to capture global discriminative features for each modality based on multi-scale features. It calculates the differences between temporal neighbor snippets in each modality to explore salient-difference features, and then utilizes them to guide correlation modeling. Thirdly, after learning discriminative features, we devise PFM module to excavate potential action and background snippets from ambiguous features. By contrastive learning, potential actions are forced closer to discriminative actions and away from the background, thereby learning more accurate action boundaries. Finally, two losses (i.e., similarity loss and reconstruction loss) are further developed to constrain the consistency between two modalities and help the model retain original feature information for better localization results. Extensive experiments show that our model achieves better performance against current WTAL methods on three datasets, i.e., THUMOS14, ActivityNet1.2 and ActivityNet1.3.
Wei Zhou 0042, Kang Lin, Weipeng Hu, Haifeng Hu 0001, Yap-Peng Tan
IEEE Trans. Multim.3
2025 Dynamic Modality-Camera-Invariant Clustering for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) offers a more flexible and cost-effective alternative compared to supervised methods. This field has gained increasing attention due to its promising potential. Existing methods simply cluster modality-specific samples and employ strong association techniques to achieve instance-to-cluster or cluster-to-cluster cross-modality associations. However, they ignore cross-camera differences, leading to noticeable issues with excessive splitting of identities. Consequently, this undermines the accuracy and reliability of cross-modal associations. To address these issues, we propose a novel dynamic modality-camera-invariant clustering (DMIC) framework for USL-VI-ReID. Specifically, our DMIC naturally integrates modality-camera-invariant expansion (MIE), dynamic neighborhood clustering (DNC), and hybrid modality contrastive learning (HMCL) into a unified framework, which eliminates both the cross-modality and cross-camera discrepancies in clustering. MIE fuses intermodal and intercamera distance coding to bridge the gaps between modalities and cameras at the clustering level. DNC employs two dynamic search strategies to refine the network's optimization objective, transitioning from improving discriminability to enhancing cross-modal and cross-camera generalizability. Moreover, HMCL is designed to optimize instance- and cluster-level distributions. Memories for intramodality and intermodality training are updated using randomly selected samples, facilitating real-time exploration of modality-invariant representations. Extensive experiments have demonstrated that our DMIC addresses the limitations present in current clustering approaches and achieves competitive performance, which significantly reduces the performance gap with supervised methods.
Yiming Yang 0001, Weipeng Hu, Qiaolin He, Haifeng Hu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
abstract
Large-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions, enabling vast applications in content generation. While recent advancements have introduced control over factors such as object localization, posture, and image contours, a crucial gap remains in our ability to control the interactions between objects in the generated content. Well-controlling interactions in generated images could yield meaningful applications, such as creating realistic scenes with interacting characters. In this work, we study the problems of conditioning T2I diffusion models with Human-Object Interaction (HOI) information, consisting of a triplet label (person, action, object) and corresponding bounding boxes. We propose a pluggable interaction control model, called InteractDiffusion that extends existing pre-trained T2I diffusion models to enable them being better conditioned on interactions. Specifically, we tokenize the HOI information and learn their relationships via interaction embeddings. A conditioning self-attention layer is trained to map HOI tokens to visual tokens, thereby conditioning the visual tokens better in existing T2I diffusion models. Our model attains the ability to control the interaction and location on existing T2I diffusion models, which outperforms existing baselines by a large margin in HOI detection score, as well as fidelity in FID and KID. Project page: https://jiuntian.github.io/interactdiffusion.
Jiun Tian Hoe, Xudong Jiang 0001, Chee Seng Chan, Yap-Peng Tan, Weipeng Hu
CVPR5
2024 Unsupervised NIR-VIS Face Recognition via Homogeneous-to-Heterogeneous Learning and Residual-Invariant Enhancement
abstract
Near-Infrared and Visible light (NIR-VIS) face recognition methods have achieved remarkable success in the fields of security surveillance, criminal investigation, and multimedia information retrieval. But the existing methods heavily rely on carefully annotated labels, leading to expensive manual labelling consumption and deployment flexibility. This motivates us to design unsupervised methods to address NIR-VIS recognition without relying on label information. To this end, we propose a novel homogeneous-to-HEterogeneous learning and Residual-invariant Enhancement (HERE) network for Unsupervised NIR-VIS Heterogeneous Face Recognition (NIR-VIS-UHFR). As the name suggests, the optimization of HERE follow a ”homogeneous-to-heterogeneous learning” strategy to fully explore complementary and common semantic information across different modalities. During the homogeneous learning phase, Modality-Adversarial Contrastive Learning (MACL) leverages the collaboration of modality contrastive learning and adversarial learning. On the one hand, MACL learns compact and discriminative intra-modal representations for NIR and VIS data, respectively. On the other hand, MACL guarantees that NIR-VIS data conform to the common feature distribution in a shared feature space, effectively reducing modal differences even in the absence of identity information between modalities. In the heterogeneous learning phase, K-reciprocal-Encoding-based Cross-modal Labeling (KECL) is introduced as robust pseudo label estimation to fully explore cross-modal relationships and group cross-modal features into clusters. With the pseudo labels provided by KECL, Refined cross-modal Contrastive Learning (RCL) is developed with modality-invariant averaging initialization and dynamic focus weighting strategies to extract modality-invariant features. Finally, Residual-invariant Representations Enhancement (RRE) mines partial features under the cross-modal face for robust matching. Compared to supervised methods, our unsupervised HERE demonstrates comparable performance on multiple datasets, greater scalability and practicality in deployment by reducing data acquisition requirements and costs.
Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Inf. Forensics Secur.2
2024 Pseudo Label Association and Prototype-Based Invariant Learning for Semi-Supervised NIR-VIS Face Recognition
abstract
Remarkable success of the existing Near-InfraRed and VISible (NIR-VIS) approaches owes to sufficient labeled training data. However, collecting and tagging data from different domains is a time-consuming and expensive task. In this paper, we tackle the NIR-VIS face recognition problem in a semi-supervised manner, termed as semi-supervised NIR-VIS Heterogeneous Face Recognition (NIR-VIS-sHFR). To cope with this problem, we propose a novel pseudo Label association and Prototype-based invariant Learning (LPL), consisting of three key components, i.e., Cross-domain pseudo Label Association (CLA), Intra-domain Compact Representation learning (ICR), and Prototype-based Inter-domain Invariant learning (PII). Firstly, the CLA iteratively builds inter-domain association graphs for pseudo-label association, subsequently facilitating cross-domain model development based on the generated pseudo-labels. Furthermore, the ICR is proposed to achieve the separation of in-domain features from different clusters and the aggregation of features from the same cluster, by performing cluster adaptation learning with prototype-based initialization. Finally, with the cross-domain pseudo-label training data produced by CLA, the PII explores potential domain-invariant and identity-related features, which employs cross-domain prototypes with identity-associated momentum updating to effectively guide inter-domain instances learning. The semi-supervised LPL method achieves comparable performance to recent supervised learning methods on multiple challenging NIR-VIS datasets, which demonstrates that the LPL is capable of learning robust cross-domain representations even without identity label information.
Weipeng Hu, Yiming Yang 0001, Haifeng Hu 0001
IEEE Trans. Image Process.1
2024 Syncretic Space Learning Network for NIR-VIS Face Recognition
abstract
To overcome the technical bottleneck of face recognition in low-light scenarios, Near-InfraRed and VISible (NIR-VIS) heterogeneous face recognition is proposed for matching well-lit VIS faces with poorly lit NIR faces. Current cross-modal synthesis methods visually convert the NIR modality to the VIS modality and then perform face matching in the VIS modality. However, using a heavyweight GAN network on unpaired NIR-VIS faces may lead to high synthesis difficulty, low inference efficiency, and other problems. To alleviate the above problems, we simultaneously synthesize NIR and VIS images into modality-independent syncretic images and propose a novel syncretic space learning (SSL) model to eliminate the modal gap. First, Syncretic Modality Generator (SMG) synthesizes NIR and VIS images into syncretic images using channel-level convolution with a shallow CNN. In particular, the discriminative structural information is well preserved and the face quality can be further improved with small modal variations in a self-supervised learning manner. Second, Modality-adversarial Syncretic space Learning (MSL) projects NIR and VIS images into the syncretic space by a syncretic-modality adversarial learning strategy with syncretic pattern guided objective, so the modal gap of NIR-VIS faces can be effectively reduced. Finally, the Syncretic Distribution Consistency (SDC) constructed by NIR-syncretic, syncretic-syncretic, and VIS-syncretic consistency can enhance the intra-class compactness and learn discriminative representations. Extensive experiments on three challenging datasets demonstrate the effectiveness of the SSL method.
Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Neutral Face Learning and Progressive Fusion Synthesis Network for NIR-VIS Face Recognition
abstract
To meet the strong demand for deploying face recognition systems in low-light scenarios, the Near-InfraRed and VISible (NIR-VIS) face recognition task is receiving increasing attention. However, heterogeneous faces have the characteristics of heterogeneity and non-neutrality. Heterogeneity refers to the fact that the matching images are in different modalities, and non-neutrality means that the matching images are significantly different in pose, expression, lighting, etc. Both situations pose challenges for NIR-VIS face matching. To address this problem, we propose a novel Neutral face Learning and Progressive Fusion synthesis (NLPF) network to disentangle the latent attributes of heterogeneous faces and learn neutral face representations. Our approach naturally integrates Identity-related Neutral face Learning (INL) and Attribute Progressive Fusion (APF) into a joint framework. Firstly, INL eliminates modal variations and residual variations by guiding the network to learn homogeneous neutral face feature representations, which tackles the challenge of heterogeneity and non-neutrality by mapping cross-modal images to a common neutral representation subspace. Besides, APF is presented to perform the disentanglement and reintegration of identity-related features, modality-related features and residual features in a progressive fusion manner, which helps to further purify identity-related features. Comprehensive evaluations are carried out on three mainstream NIR-VIS datasets to verify the robustness and effectiveness of the NLPF model. In particular, NLPF has competitive recognition performance on LAMP-HQ, the most challenging NIR-VIS dataset so far.
Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Robust Cross-Domain Pseudo-Labeling and Contrastive Learning for Unsupervised Domain Adaptation NIR-VIS Face Recognition
abstract
Near-infrared and visible face recognition (NIR-VIS) is attracting increasing attention because of the need to achieve face recognition in low-light conditions to enable 24-hour secure retrieval. However, annotating identity labels for a large number of heterogeneous face images is time-consuming and expensive, which limits the application of the NIR-VIS face recognition system to larger scale real-world scenarios. In this paper, we attempt to achieve NIR-VIS face recognition in an unsupervised domain adaptation manner. To get rid of the reliance on manual annotations, we propose a novel Robust cross-domain Pseudo-labeling and Contrastive learning (RPC) network which consists of three key components, i.e., NIR cluster-based Pseudo labels Sharing (NPS), Domain-specific cluster Contrastive Learning (DCL) and Inter-domain cluster Contrastive Learning (ICL). Firstly, NPS is presented to generate pseudo labels by exploring robust NIR clusters and sharing reliable label knowledge with VIS domain. Secondly, DCL is designed to learn intra-domain compact yet discriminative representations. Finally, ICL dynamically combines and refines intrinsic identity relationships to guide the instance-level features to learn robust and domain-independent representations. Extensive experiments are conducted to verify an accuracy of over 99% in pseudo label assignment and the advanced performance of RPC network on four mainstream NIR-VIS datasets.
Yiming Yang 0001, Weipeng Hu, Haiqi Lin, Haifeng Hu 0001
IEEE Trans. Image Process.2
2022 Orthogonal Modality Disentanglement and Representation Alignment Network for NIR-VIS Face Recognition
abstract
Near-infrared and visual (NIR-VIS) face matching, as the most typical task in Heterogeneous Face Recognition (HFR), has attracted increasing attention in recent years. However, due to the large within-class discrepancies, including domain differences and residual discrepancies (i.e., lighting, expressions, occlusion, blurry, pose, etc), this is still a difficult task. Conventional NIR-VIS FR methods only focus on reducing the modality gap between cross-domain images, while neglecting to eliminate the residual variations. To better solve the above problems, this paper proposes a novel Orthogonal Modality Disentanglement and Representation Alignment (OMDRA) approach, which consists of three key components, including Modality-Invariant (MI) loss, Orthogonal Modality Disentanglement (OMD) and Deep Representation Alignment (DRA). Firstly, the MI loss is designed to learn modality-invariant and identity-discriminative representation, by increasing between-class separability and within-class compactness between NIR and VIS heterogeneous data. Secondly, the high-level Hybrid Facial Feature (HFF) layer of the backbone network is projected into two subspaces: the modality-related and identity-related subspaces. The OMD is designed to decouple modal information via an adversarial process, and we further impose Orthogonal Representation Decorrelation (ORD) to the OMD to decrease the correlation between identity representations and domain representations, as well as enhancing their representation capabilities. Finally, the DRA aims to eliminate the residual variations by performing a high-level representation alignment between non-neutral face and neutral face, which can effectively guides the network to learn discriminative and residual-invariant face representation. The joint scheme enables the disentanglement of modality variations, elimination of residual discrepancies, and the purification of identity information. Extensive experiments on challenging cross-domain databases indicate that our OMDRA method is superior to the state-of-the-art methods.
Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Adversarial Decoupling and Modality-Invariant Representation Learning for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (RGB-IR ReID) has now attracted increasing attention due to its surveillance applications under low-light environments. However, the large intra-class variations between different domains are still a challenging issue in the field of computer vision. To address the above issue, we propose a novel adversarial Decoupling and Modality-invariant Representation learning (DMiR) method to explore potential spectrum-invariant yet identity-discriminative representations for cross-modality pedestrians. Our model consists of three key components, including Domain-related Representation Disentanglement (DrRD), Modality-invariant Discriminative Representation (MiDR) and Representation Orthogonal Decorrelation (ROD). First, two subnets named Identity-Net and Domain-Net are designed to extract identity-related features and domain-related features, respectively. Given this two-stream structure, the DrRD is introduced to achieve adversarial decoupling against domain-specific features via a min-max disentanglement process. Specifically, the classification objective function on Domain-Net is minimized to extract spectrum-specific information while maximizing it to reduce domain-specific information. Second, in Identity-Net, we introduce MiDR to enhance intra-class compactness and reduce domain variations by exploring positive and negative pair variations, semantic-wise differences, and pair-wise semantic variations. Finally, the correlation between the two decomposed features, i.e., identity-related features and domain-related features, may lead to the introduction of modal information in identity representations, and vice versa. Therefore, we present the ROD constraint to make the two decomposed features unrelated to each other, which can more effectively separate the two-component features and enhance feature representations. Practically, we construct ROD at the feature-level and parameter-level, and finally select feature-level ROD as the decorrelation strategy because of its superior decorrelation performance. The whole scheme leads to disentangling spectrum-dependent information, as well as purifying identity information. Extensive experiments are carried out on two mainstream RGB-IR ReID datasets, and the results demonstrate the effectiveness of our method.
Weipeng Hu, Bohong Liu, Haitang Zeng, Yanke Hou, Haifeng Hu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Dual Face Alignment Learning Network for NIR-VIS Face Recognition
abstract
As the most important topic in Heterogeneous Face Recognition (HFR), Near-InfraRed and VISual (NIR-VIS) face recognition has attracted increasing research attentions owing to its potential application in the field of criminal detective cases and multimedia information retrieval. However, due to its dramatic intra-class variations, including modality, pose, occlusion, blurry, lighting, distance, expression, etc, it is very challenging to retain inherent identity information. To address the above issue, we propose a novel Dual Face Alignment Learning (DFAL) algorithm to explore the potential domain-invariant neutral face representations of the cross-modal images. Our model contains three effective components including Feature-level Face Alignment (FFA), Image-level Face Alignment (IFA) and Cross-domain compact Representation (CdR). Firstly, Teacher-Encoder CNNs (TeEn-CNNs) and Student-Encoder CNNs (StEn-CNNs) are designed to encode features for VIS neutral face images and non-neutral face images, and the FFA is introduced to learn neutral face representations by performing feature-level alignment between non-neutral face and VIS neutral face. Secondly, Student-Decoder CNNs (StDe-CNNs) is developed to decode features to restore face images, and the IFA is designed to reconstruct neutral face image by imposing image-level alignment. Notably, the FFA acts as the primary target to learn VIS neutral face representations for cross-view data, while the IFA plays a role in the icing on the cake, i.e., further disentangling domain and residual information through the synthesis process. Finally, the CdR dispels modality features and distills identity features by mining inter-class information, inter-domain information and inter-semantic relationship. The joint scheme enables the elimination of intra-class variations and the purification of identity information. We carry out comprehensive experiments to illustrate the effectiveness of the DFAL approach on three challenging NIR-VIS databases.
Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Domain-Private Factor Detachment Network for NIR-VIS Face Recognition
abstract
Near-InfraRed and VISual (NIR-VIS) face matching, as one of the most representative tasks in Heterogeneous Face Recognition (HFR), aims at retrieving a face image across different domains. With the development of deep learning and the growing demand for intelligent surveillance, it has aroused more and more research attention in the computer vision community. However, due to the dramatic modality gap between NIR and VIS images, the task of NIR-VIS face recognition becomes practically very challenging. In this paper, we propose a novel Domain-private Factor Detachment (DFD) network to disentangle domain-dependent factors and achieve identity information distillation. Our approach consists of three key components, including Domain-identity Representation Learning (DiRL), Cross-domain Factor Detachment (CdFD) and Cross-domain Aggregation Learning (CAL). Firstly, the proposed DiRL aims to achieve domain-specific information distillation and learn identity-related representations. Specifically, three sub-networks, i.e., NIR sub-Network (NIR-Net), VIS sub-Network (VIS-Net) and IDentity-dependent sub-Network (ID-Net) are designed to learn NIR facial representations, VIS facial representations and identity-dependent representations, respectively, and they can promote each other to facilitate the learning of identity-discriminative representations. Secondly, considering that the entangled modal components in face representations negatively affect the subsequent matching process, to reduce modality-related components, we model the cross-modal face matching problem into three parts, comprising Identity Variation (IV), Inter-Spectrum Variation (ISV) and Identity-Domain Variation (IDV). The CdFD is presented to eliminate ISV components and IDV components by introducing inter-spectrum invariant constraint and identity-domain invariant constraint, so that cross-modal face recognition can be performed under pure identity information differences without modal interference. Finally, the CAL is developed to learn modality-invariant yet discriminative representations by exploring within-class aggregation, negative pair separability and cross-domain positive pair compactness. Experimental results on multiple challenging databases demonstrate the effectiveness of the DFD approach.
Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Inf. Forensics Secur.1
2021 Domain Discrepancy Elimination and Mean Face Representation Learning for NIR-VIS Face Recognition
abstract
Due to its potential application in criminal cases, security systems and multimedia information retrieval, Near InfraRed (NIR) to VISible (VIS) face recognition has attracted increasing research attention in the field of computer vision. However, it is still a challenging task because of the large intra-class variations including spectrum, occlusion, lighting, blurry, expression and pose. To address the above problem, we propose a novel Domain discrepancy Elimination and Mean face Representation learning (DEMR) for NIR-VIS face recognition. The DEMR consists of two key components comprising Class-wise Domain Discrepancy Elimination (CDDE) and Cross-modal Mean Face Alignment (CMFA). Specifically, two-branch modality-specific networks are designed to extract features for VIS images and NIR images, respectively. Considering that distribution variations of cross-modal images will decrease recognition performance, we present CDDE to eliminate modality gap by narrowing distribution differences of VIS images and NIR images in a category-by-category manner. Moreover, to reduce the intra-class discrepancies and obtain compact feature representation, the CMFA is designed to achieve representation alignment between cross-domain images and VIS prototypes (i.e., VIS mean face representations), through optimizing a quadruplet constraint. Extensive experiments on multiple challenging NIR-VIS databases validate that our DEMR is effective for cross-modal face recognition task.
Weipeng Hu, Haifeng Hu 0001
IEEE Signal Process. Lett.1
2021 Dual Adversarial Disentanglement and Deep Representation Decorrelation for NIR-VIS Face Recognition
abstract
The task of near-infrared and visual (NIR-VIS) face recognition refers to matching face data from different modalities, which has broad application prospects in areas such as multimedia information retrieval and criminal investigation. However, it remains a challenging task due to high intra-class variations and small-scale NIR-VIS dataset. In this paper, we propose a novel approach called Dual Adversarial Disentanglement and deep Representation Decorrelation (DADRD) to solve the NIR-VIS matching problem. In order to reduce the gap between NIR-VIS images, three key components are designed for DADRD model, including Cross-modal Margin (CmM) loss, Dual Adversarial Disentangled Variations (DADV) and Deep Representation Decorrelation (DRD). Firstly, the CmM loss captures within- and between-class information of the data, and it further reduces modality difference by a center-variation item. Secondly, the Mixed Facial Representation (MFR) layer of the backbone network is divided into three parts: the identity-related layer, the modality-related layer and the residual-related layer. The DADV is designed to reduce the intra-class variations, which consists of Adversarial Disentangled Modality Variations (ADMV) and Adversarial Disentangled Residual Variations (ADRV). Specifically, the ADMV and ADRV aim at eliminating spectrum variations and residual variations (i.e., lighting, pose, expression, occlusion, etc) respectively via an adversarial mechanism. Finally, we impose a DRD on the three decomposed features to make them irrelevant to each other, which can more effectively separate the three component information and enhance feature representations. In particular, we develop a Joint Three-stage Optimization (JTsO) strategy to effectively optimize the network. The joint formulation leads to the purification of identity information and the disentanglement of within-class variation information. Extensive experiments have been carried out on three challenging datasets, and the results demonstrate the effectiveness of our method.
Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Inf. Forensics Secur.1
2021 Adversarial Disentanglement Spectrum Variations and Cross-Modality Attention Networks for NIR-VIS Face Recognition
abstract
Near-infrared and visual (NIR-VIS) matching task refers to the face recognition between the two images of different modalities, which remains a challenging task in the field of machine vision. The main problems of NIR-VIS Heterogeneous Face Recognition (HFR) tasks include two aspects: large intra-class differences caused by cross-modal data, and insufficient paired training samples. In this paper, an effective Adversarial Disentanglement spectrum variations and Cross-modality Attention Networks (ADCANs) is proposed for VIS-NIR matching task. Three key components are introduced to the ADCANs for reducing the gap of cross-modal images: Advanced Scatter Loss (ASL), Modality-adversarial Feature Learning (MaFL) and Cross-modality Attention Block (CmAB). The proposed ASL loss captures between- and within-class information of the data and embeds them to the network for more effective training, and it focuses on categories with small between-class distance and increases the distance between them. The MaFL consists of an Identity-Discriminative Feature Learning Network (IDFLN) and a Modality-Adversarial Disentanglement Network (MADN), which can enhance the identity-discriminative feature representations as well as disentangling spectrum variations via an adversarial learning. The IDFLN built by an end-to-end CNNs aims at learning identity-discriminative feature. While the MADN built by a discriminator D and a generator G focuses on removing modality-related information. Furthermore, to increase representation power as well as disentangling spectrum variations effectively, a CmAB block is introduced to the network, which sequentially applies spatial and channel attention modules to both the IDFLN and MADN. Since the channel attention module focuses on `what' features to suppress or emphasize, an orthogonality constraint is introduced to the two channel attention modules, which allows MADN and IDFLN to focus on learning modality-related features and identity-related features, respectively. In particular, the ADCANs consists of multiple CmAB blocks to learn discriminative features and disentangle spectrum variations. A large number of experiments on three challenging HFR datasets indicate that the proposed ADCANs is effective for VIS-NIR HFR task.
Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Multim.1
2020 SOAPTyping: an open-source and cross-platform tool for sequence-based typing for HLA class I and II alleles
abstract
BACKGROUND: The human leukocyte antigen (HLA) gene family plays a key role in the immune response and thus is crucial in many biomedical and clinical settings. Utilizing Sanger sequencing, the golden standard technology for HLA typing enables accurate identification of HLA alleles in high-resolution. However, only the commercial software, such as uTYPE, SBT-Assign, and SBTEngine, and very few open-source tools could be applied to perform HLA typing based on Sanger sequencing. RESULTS: We developed a user-friendly, cross-platform and open-source desktop application, known as SOAPTyping, for Sanger-based typing in HLA class I and II alleles. SOAPTyping can produce accurate results with a comprehensible protocol and featured functions. Moreover, SOAPTyping supports a more advanced group-specific sequencing primers (GSSP) module to solve the ambiguous typing results. We used SOAPTyping to analyze 36 samples with known HLA typing from the University of California Los Angeles (UCLA) International HLA DNA Exchange platform and 100 anonymous clinical samples, and the HLA typing results from SOAPTyping are identical to the golden results and 5.5 times faster than commercial software uTYPE, which shows the usability of SOAPTyping. CONCLUSIONS: We introduce the SOAPTyping as the first open-source and cross-platform HLA typing software with the capability of producing high-resolution HLA typing predictions from Sanger sequence data.
Yong Zhang 0036, Yongsheng Chen, Huixin Xu, Weipeng Hu, Xiaoqin Yang, Jia Ye, Jiayin Wang 0002, Weiqiang Sun, Jian Wang 0065, Huanming Yang
BMC Bioinform.6
2020 Disentangled Spectrum Variations Networks for NIR-VIS Face Recognition
abstract
Surveillance cameras often capture near infrared images since it provides a low-cost and effective solution to acquire high-quality images under low-light environments. However, visual versus near infrared (VIS-NIR) heterogeneous face recognition (HFR) is still a challenging issue in computer vision community due to the gap between sensing patterns of different spectrums as well as the lack of sufficient training samples. To solve the above problem, in this paper, we present an effective Disentangled Spectrum Variations Networks (DSVNs) for VISNIR HFR. Two key strategies are introduced to the DSVNs for disentangling spectrum variations between two domains: Spectrum-adversarial Discriminative Feature Learning (SaDFL) and Step-wise Spectrum Orthogonal Decomposition (SSOD). The SaDFL consists of Identity-Discriminative subnetwork (IDNet) and Auxiliary Spectrum Adversarial subnetwork (ASANet). On the one hand, the IDNet is composed of a generator GHand a discriminator DUfor extracting identity-discriminative feature. On the other hand, the ASANet is built by a generator GHand a discriminator DMfor eliminating modality-variant spectrum information under the guidance of the discriminator DM. The identity-label and modality-label HFR datasets are used to train the DSVNs with triplet loss. Both IDNet and ASANet can jointly enhance the domain-invariant feature representations via an adversarial learning. Furthermore, to disentangle spectrum variations effectively as well as making identity information and modality information unrelated to each other, we present a new topology of connection block called Disentangled Spectrum Variations (DSV). An orthogonality constraint is imposed to DSV at the convolution level for channel-wise orthogonal decomposition between the modality-invariant identity information and modalityvariant spectrum information. In particular, the SSOD is built by stacking multiple modularized mirco-block DSV, and thereby enjoys the benefits of disentangling spectrum variation step by step. Moreover, we investigate the similarity calculation method to further improve the HFR performance. To sum up, the designed DSVNs leads to a purification of identity information as well as an elimination of modality information. Extensive experiments are carried out on two challenging NIR-VIS HFR datasets CASIA NIRVIS 2.0 and Oulu-CASIA NIR-VIS, demonstrating the superiority of the proposed method.
Weipeng Hu, Haifeng Hu 0001
IEEE Trans. Multim.1
2019 Discriminant Deep Feature Learning based on joint supervision Loss and Multi-layer Feature Fusion for heterogeneous face recognition
Weipeng Hu, Haifeng Hu 0001
Comput. Vis. Image Underst.1
2019 Fine Tuning Dual Streams Deep Network with Multi-scale Pyramid Decision for Heterogeneous Face Recognition
Weipeng Hu, Haifeng Hu 0001
Neural Process. Lett.1
2016 Patch-based alignment-free generic sparse representation for pose-robust face recognition
abstract
Sparse representation based classification method has been successfully applied to face recognition in recent years. However, it is still a problem in the scenario of pose variation in face recognition with single sample per person. In this paper, we propose a novel alignment-free model, called Gabor-based Partial Face Sparse Representation (GPFSR), to solve the problem of pose variation in face recognition with single sample per person by using partial face. In our method, we firstly locate five facial landmarks in different images. Then partial face is obtained, which is used to construct Gabor-based local dictionary and compute the weights of each patch. Our classification principle is based on sparse representation. The experimental results on the Multi-PIE and FERET show that GPFSR is robust to pose variation in FR with single sample per person.
Jianquan Gu, Haifeng Hu 0001, Haoxi Li, Weipeng Hu
ICIP4