EDBT 2026 Demo / reviewers in the wild / expert
Haifeng Hu 0001
dblp:28/5938-1
· DBLP profile ↗
211ranked-venue papers
14as first author
119since 2021 · last 2026
0000-0002-4884-323XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 10 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 98 · 4 first-author · 57 since 2021Computer networks · 26 · 15 since 2021Security and privacy · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tri-Modal grouping fusion network with optimal transport learning for unaligned multimodal sentiment analysis
Aolin Xiong, Sijie Mai, Haifeng Hu 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Omniscient bottom-up double-stream symmetric network for image captioning
Jianchao Li, Wei Zhou 0042, Kai Wang 0033, Haifeng Hu 0001 |
Knowl. Based Syst. | 4 |
| 2026 | PSKNet: Lightweight kernel-aware slice network for real-time stereo depth estimation on edge devices
Bifa Liang, Haifeng Hu 0001, Dihu Chen |
Knowl. Based Syst. | 3 |
| 2026 | SPADNet: lightweight stereo disparity estimation for real-time embedded multimedia systems
Bifa Liang, Yiming Zeng 0008, Haifeng Hu 0001, Dihu Chen |
Multim. Syst. | 3 |
| 2026 | Affection-Guided Bottleneck Diffusion for Missing Modality Issue in Multimodal Affective ComputingabstractMissing modality issue in multimodal affective computing severely hinders the robustness and performance of multimodal learning, particularly in real-world scenarios. Existing methods often fail in fully exploiting the remaining modalities, leading to noisy reconstruction process for the missing modalities and yielding suboptimal results. Besides, most of these methods rely on designing sophisticated networks to handle various missing scenarios, which prevents them from taking advantage of the original pre-trained multimodal networks trained for complete multimodal inputs. To address these challenges, we propose Affection-guided Bottleneck Diffusion (ABDiff), a novel approach leveraging score-based diffusion generative encoders to reconstruct missing modalities in the latent space without modification to the pre-trained fusion models. By incorporating self- and cross-attention mechanisms inside and among the missing and remaining modalities, ABDiff captures both modality-specific dynamics and cross-modal interactions during generation. Furthermore, an affection-guided information bottleneck is introduced to filter task-unrelated noise and modality-specific redundancy, stabilizing the generation process of missing modalities. The generated representations are seamlessly integrated with the remaining modalities into the pre-trained fusion networks. Extensive experiments on four public multimodal affective computing datasets demonstrate that ABDiff surpasses previous methods under both complete and incomplete modality scenarios. The code is released inhttps://github.com/RH-Lin/ABDiff. Ronghao Lin, Qiaolin He, Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2026 | RCAENet: Residual Convolutional and Attention-Enhanced Stereo Matching for Real-Time Depth Estimation on Edge DevicesabstractAs a core technology in real-time video processing and intelligent surveillance, stereo matching provides essential depth perception capabilities for multimedia applications. However, high-precision stereo networks often come with significant computational costs, making real-time inference on power- and memory-constrained edge devices challenging. On the other hand, lightweight real-time networks still struggle with accuracy limitations. To address this challenge, we propose RCAENet, a high-performance stereo network designed for real-time and high-accuracy depth estimation on edge devices. To enhance feature extraction efficiency, we introduce the Residual Convolutional Feature Extraction (RCFE) module, which replaces conventional convolutional layers to capture more expressive features while maintaining computational efficiency. Additionally, we propose the Enhanced Adaptive Upsampling (EAU) module, which integrates channel and spatial attention mechanisms to improve feature fusion and disparity refinement. Furthermore, we design an Enhanced 3D CNN (E3DC) along with the Cost Aggregation and Residual Attention (CA-ResAgg) module for cost volume regularization. This module incorporates residual aggregation and efficient channel attention to further enhance disparity estimation accuracy. Built upon these components, RCAENet features a multi-scale architecture that effectively balances accuracy and efficiency. Extensive experiments demonstrate that these innovations enable RCAENet to achieve real-time inference on edge devices while maintaining state-of-the-art depth accuracy. Bifa Liang, Haifeng Hu 0001, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Supervised Attention Mechanism for Low-quality Multimodal DataabstractIn practical applications, multimodal data are often of low quality, with noisy modalities and missing modalities being typical forms that severely hinder model performance, robustness, and applicability.However, current studies address these issues separately.To this end, we propose a framework for multimodal affective computing that jointly addresses missing and noisy modalities to enhance model robustness in low-quality data scenarios.Specifically, we view missing modality as a special case of noisy modality, and propose a supervised attention framework.In contrast to traditional attention mechanisms that rely on main task loss to update the parameters, we design supervisory signals for the learning of attention weights, ensuring that attention mechanisms can focus on discriminative information and suppress noisy information.We further propose a ranking-based optimization strategy to compare the relative importance of different interactions by adding a ranking constraint for attention weights, avoiding training noise caused by inaccurate absolute labels.The proposed model consistently outperforms state-of-the-art baselines on multiple datasets under the settings of complete modalities, missing modalities, and noisy modalities. Sijie Mai, Shiqin Han, Haifeng Hu 0001 |
EMNLP | 3 |
| 2025 | E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Haifeng Hu 0001, Yap-Peng Tan |
ACM Multimedia | 7 |
| 2025 | CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal LearningabstractMultimodal machine learning, mimicking the human brain’s ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real‑world deployments, however, the presence of modality is highly variable and unpredictable, causing the pre-trained models in suffering significant performance drops and fail to remain robust with dynamic missing modalities circumstances. In this paper, we present a novel Cyclic INformative Learning framework (CyIN) to bridge the gap between complete and incomplete multimodal learning. Specifically, we firstly build an informative latent space by adopting token- and label-level Information Bottleneck (IB) cyclically among various modalities. Capturing task-related features with variational approximation, the informative bottleneck latents are purified for more efficient cross-modal interaction and multimodal fusion. Moreover, to supplement the missing information caused by incomplete multimodal input, we propose cross-modal cyclic translation by reconstruct the missing modalities with the remained ones through forward and reverse propagation process. With the help of the extracted and reconstructed informative latents, CyIN succeeds in jointly optimizing complete and incomplete multimodal learning in one unified model. Extensive experiments on 4 multimodal datasets demonstrate the superior performance of our method in both complete and diverse incomplete scenarios. Ronghao Lin, Qiaolin He, Sijie Mai, Aolin Xiong, Yap-Peng Tan, Haifeng Hu 0001 |
NeurIPS | 8 |
| 2025 | Learning by Comparing: Boosting Multimodal Affective Computing through Ordinal LearningabstractPrevious studies on multimodal affective computing primarily focus on approximating predictions to annotated labels, often neglecting the ordinal nature of affective states. In this paper, we address this issue by exploring ordinal learning, and a Multimodal Ordinal Affective Computing (MOAC) framework is designed to enhance the understanding of the nature of affective concepts. Specifically, we propose coarse-grained label-level ordinal learning that prompts the model to learn to compare in the label space, encouraging higher predictive values for samples annotated with larger labels over those with smaller labels. Moreover, a regularization loss is proposed to prevent the output distributions from deviating significantly from the annotated label distributions. Fine-grained feature-level ordinal learning is then performed via the feature difference operation and the neutral embedding. The former compares samples in the feature space, calculating the difference between features of different samples to generate 'new' features for a more robust training. The latter seeks to reduce the difficulty of prediction by estimating the difference between the target multimodal representations and a neutral reference. We first demonstrate MOAC in multimodal sentiment analysis, which is a regression task that aligns well with the function of ordinal learning. Then we extend MOAC to classification tasks including multimodal humor detection and sarcasm detection to evaluate its generalizability. Experiments suggest that MOAC outperforms state-of-the-art methods. Sijie Mai, Haifeng Hu 0001 |
WWW | 3 |
| 2025 | From grids to pseudo-regions: Dynamic memory augmented image captioning with dual relation transformer
Wei Zhou 0042, Weitao Jiang, Zhijie Zheng 0003, Jianchao Li, Haifeng Hu 0001 |
Expert Syst. Appl. | 6 |
| 2025 | From multi-scale grids to dynamic regions: Dual-relation enhanced transformer for image captioning
Wei Zhou 0042, Chuanle Song, Dihu Chen, Haifeng Hu 0001, Chun Shan |
Knowl. Based Syst. | 5 |
| 2025 | DRTN: Dual Relation Transformer Network with feature erasure and contrastive learning for multi-label image classification
Wei Zhou 0042, Kang Lin, Zhijie Zheng 0003, Dihu Chen, Haifeng Hu 0001 |
Neural Networks | 6 |
| 2025 | Cascaded Dynamic Memory Refinement and Semantic Alignment for Exo-to-Ego Cross-View Video GenerationabstractCross-view video generation from exocentric (third-person) to egocentric (first-person) perspectives poses a challenging task, due to the significant viewpoint gap and limited overlap between these two views. Previous methods exhibit limitations in capturing long-range temporal context and overlook egocentric semantic priors, leading to degraded performance in cross-view synthesis. To address these challenges, we propose a cue-free video-based approach termed cascaded Dynamic memory Refinement and Semantic Alignment (DRSA), which integrates temporal knowledge over extended periods and learns egocentric semantic information to generate videos. The Dynamic Memory Refinement (DMR) exploits long horizon temporal dynamics to learn salient information that compensates for the limited overlap between views. Specifically, we devise a dynamic memory that serves as a knowledge repository, and utilize a sliding window to locate the corresponding long-term temporal information, which is subsequently processed with adaptive weighting and cross-attention transformer to refine feature representations. Furthermore, aware of the considerable viewpoint divergence that hinder semantic learning of target view, we propose Viewpoint-aware Semantic Alignment (VSA) with dual encoder-decoder learning and semantic alignment, which transfer egocentric semantic details from the egocentric synthesis pipeline to the exocentric synthesis pipeline. In particular, the VSA module narrows the semantic gap between views, further promoting long-range temporal modeling in DMR under alignment constraints. By extending this into a cascaded fashion, the Cascaded Alignment and Refinement (CAR) progressively aligns semantic features and performs feature refinement to facilitate viewpoint learning at different levels of granularity. To overcome the limitations of existing databases known for their limited static scenes and scarcity of interacting objects, we create a new dataset with dynamic exocentric scenes and rich interacting objects to further promote the task. Thorough experimental analysis reveals that our method surpasses current state-of-the-art techniques in terms of both quantitative metrics and qualitative evaluations. Weipeng Hu, Jiun Tian Hoe, Haifeng Hu 0001, Xudong Jiang 0001, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Extended Cross-Modality United Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims to learn modality-invariant features from unlabeled cross-modality data. However, existing approaches lack comprehensive cross-modality clustering or excessively pursue cluster-level association, which hinders reliable learning of modality-invariant features. To address these challenges, we propose an Extended Cross-Modality United Learning (ECUL) framework, which integrates Extended Modality-Camera Clustering (EMCC) and Two-Step Memory Updating Strategy (TSMem) modules. Specifically, we design ECUL to naturally unify intra-modality clustering, inter-modality clustering, and inter-modality instance selection, establishing compact and accurate cross-modality associations while reducing the introduction of noisy labels. Moreover, EMCC captures and filters neighborhood relationships by extending the encoding vector, which further promotes the learning of modality-invariant and camera-invariant knowledge in terms of the clustering algorithm. Finally, TSMem provides accurate and generalized proxy points for contrastive learning by updating memory in stages. Comprehensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate that the proposed ECUL framework shows promising performance and even outperforms certain supervised methods. Ruixing Wu, Yiming Yang 0001, Jiakai He, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Injecting Multimodal Information Into Pre-Trained Language Model for Multimodal Sentiment AnalysisabstractWith the increasing availability of computational and data resources, numerous powerful pre-trained language models (PLMs) have emerged for natural language processing tasks. However, how to inject nonverbal modalities into PLMs to handle multimodal information remains a practical problem. In this paper, we explore the application of PLM on multimodal sentiment analysis from a different perspective. Unlike many recent methods that develop multimodal fusion layers that are sequential to attention layers, we investigate the effectiveness of cross-modal additive attention that is parallel to attention layers, which takes the language modality as dominant modality. Moreover, we devise a gating mechanism to control the flow of nonverbal information by estimating its discriminative level. In this way, we can prevent noisy multimodal information from damaging the performance of pre-trained language model. In our framework, nonverbal modalities serve as auxiliary roles to provide the model with additional information and improve the understanding of multimodal human language. Additionally, cross-modal margin and matching losses are proposed to align the distributions of various modalities and simultaneously retain modality-specific information, which to some extent address the shortcoming of contrastive learning loss. Comprehensive experiments show that our approach surpasses existing state-of-the-art methods on multimodal sentiment analysis and emotion recognition tasks. Sijie Mai, Aolin Xiong, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Hierarchical Knowledge Stripping for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) has emerged as a prominent research area that focuses on leveraging multimodal data to understand intention and sentiment signals. Despite significant progress, two major challenges remain in integrating diverse modalities: modal heterogeneity and interference information. To address these issues, we propose a novel framework called Multimodal Hierarchical Knowledge Stripping (MHKS), which enables the progressive extraction of informative knowledge. First, inspired by the information bottleneck (IB), we design a hierarchical disentanglement strategy to stepwise separate task-relevant and task-irrelevant information at the feature, attribute, and semantic levels. This enables MHKS to extract valuable knowledge in unimodal representations and eliminate interference information. Then, to mitigate the distribution gap across multiple modalities, we further design an adaptive alignment strategy based on contrastive learning. We utilize text modality as a bridge to connect other nonverbal modalities, which encourages adaptive alignment across modalities and facilitates the learning of more harmonized joint representations. Comprehensive experiments on three popular datasets demonstrate our method achieves excellent performance on MSA tasks. Aolin Xiong, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Part-Based Bi-Directional Enhancement Learning for Unsupervised Visible-Infrared Re-IdentificationabstractUnsupervised Learning Visible-Infrared Person Re-identification (USL-VI-ReID) aims to learn uniform feature representations for retrieving persons from unlabeled cross-modality data, which can accomplish 24-hour surveillance without expensive manual annotations. However, USL-VI-ReID is a cross-modality retrieval task that suffers from cross-modality label association and cross-modality feature discrepancy problems. To address these two problems, we propose a Part-based Bidirectional Enhancement (PBE) framework for learning a cross-modality uniform representation of USL-VI-ReID. The PBE consists of both forward and backward enhancements: 1) To alleviate the cross-modality label association problem, we propose a Part-based Label Forward Enhancement (PLFE) module. The PLFE module employs part features to complement global features during the label association process, thus generating higher-quality VI-associated pseudo-labels for the forward enhancement of the feature learning process. 2) To mitigate the cross-modality feature discrepancy problem, we propose a Part-based Feature Backward Enhancement (PFBE) module. The PFBE module utilizes part features to augment global features during the feature learning process, thus learning more robust uniform features for the backward enhancement of the label association process. Based on the part features, our PBE method achieves bi-directional enhancement during the label association and feature learning processes for robust recurrent training. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that the proposed PBE framework outperforms existing USL-VI-ReID methods. Code is available at https://github.com/heqlin5/PBE. Qiaolin He, Yiming Yang 0001, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Prompt-Guided Transformer and MLLM Interactive Learning for Text-Based Pedestrian SearchabstractAiming to retrieve pedestrian images based on a textual description query, Text-Based Pedestrian Search (TBPS) gains increasingly attention due to its applications in security surveillance. As a fine-grained classification task, TBPS requires identifying images of individuals with different semantic contexts yet the same identity, as well as distinguishing images of individuals who share similar appearances but distinct identities. Consequently, TBPS is challenged by semantic variations in positive pairs and appearance similarity between negative pairs. To tackle these challenges, we propose the Prompt-guided Transformer and MLLM Interactive learning (PTMI) model to learn identity-discriminative representations across different modalities. PTMI consists of three components: the Prompt-guided Transformer (Promformer), MLLM Interactive Learning (MIL) and Dual-branch Cross-modal Learning (DCL). Firstly, the Promformer is designed to handle semantic variations in positive pairs by introducing learnable prompts, composing of three types: instance-shared, instance-specific and layer-specific. Optimized by Cross-modal Intra-class Consistency (CIC) loss, these prompts minimize intra-class variations and retrieve positive images with various semantics. Secondly, the MIL component is introduced to address appearance similarity between negative pairs by focusing on key image patches and description words filtering by the local discriminator. Powered by Multimodal Large Language Model (MLLM), the local discriminator adopts soft attention to highlight important image regions and descriptive words, which preserves semantic information while emphasize discriminative details. Lastly, the DCL integrates global and local branches to bridge modality discrepancies. The global branch employs SDM loss for heterogeneous distribution alignment, while the local branch applies Anchor-Based Contrastive (ABC) loss for instance-level contrastive learning. Unlike conventional contrastive loss, ABC loss leverages MLLM features as anchors to decouple modality and semantic differences, enhancing alignment efficiency. Extensive experiments on three TBPS datasets have validated the effectiveness of PTMI. Zefeng Lu, Ronghao Lin, Yap-Peng Tan, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Progressive Cross-Modal Association Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (USL-VI-ReID) aims to explore the cross-modal associations and learn modality-invariant representations without manual labels. The field provides flexible and economical methods for person re-identification across light and dark scenes. Existing approaches utilize cluster-level strong association methods, such as graph matching and optimal transport, to correlate modal differences, which may result in mis-linking between clusters and introduce noise. To overcome this limitation and gradually acquire reliable cross-modal associations, we propose a Progressive Cross-modal Association Learning (PCAL) method for USL-VI-ReID. Specifically, our PCAL naturally integrates Triple-modal Adversarial Learning (TAL), Cross-modal Neighbor Expansion (CNE) and Modality-invariant Contrastive Learning (MCL) into a unified framework. TAL fully utilizes the advantage of Channel Augmented (CA) technique to reduce modal differences, which facilitates subsequent mining of cross-modal associations. Furthermore, we identify the modal bias problem in existing clustering methods, which hinders the effective establishment of cross-modal associations. To address this problem, CNE is proposed to balance the contribution of cross-modal neighbor information, linking potential cross-modal neighbors as much as possible. Finally, MCL is then introduced to refine the cross-modal associations and learn modality-invariant representations. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the competitive performance of PCAL method. Code is available at https://github.com/YimingYang23/PCA USLVIReID. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | RAFDet: Range View Augmented Fusion Network for Point-Based 3D Object DetectionabstractIn recent years, point-based methods have achieved promising performance on 3D object detection task. Although effective, they still suffer from the inherent sparsity of point cloud, which makes it challenging to distinguish objects with backgrounds only relying on the view of raw point. To this end, we propose a straightforward yet effective multi-view fusion network termed RAFDet to alleviate this issue. The core idea of our method lies in combining the merits of raw point and its range view to enhance the representation learning for sparse point cloud, thus mitigating the sparsity problem and boosting the detection performance. In particular, we introduce a novel bidirectional attentive fusion module to equip sparse point with interacted fine-grained semantic clues during feature learning process. Then, we devise the range-view augmented fusion module to fully exploit the supplementary relationship between different perspectives with the aim of enhancing original point-view features. In the end, a single-stage detection head is utilized to predict final 3D bounding boxes based on the enhanced semantics. We have evaluated our method on the popular KITTI Dataset, DAIR-V2X Dataset and Waymo Open Dataset. Experimental results on the above three datasets demonstrate the effectiveness and robustness of our approach in terms of detection performance and model complexity. Zhijie Zheng 0003, Kang Lin, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Multim. | 5 |
| 2025 | Snippet-Inter Difference Attention Network for Weakly-Supervised Temporal Action LocalizationabstractThe purpose of weakly-supervised temporal action localization (WTAL) task is to simultaneously classify and localize action instances in untrimmed videos with only video-level labels. Previous works fail to extract multi-scale temporal features to identify action instances with different durations, and they do not fully use the temporal cues of action video to learn discriminative features. In addition, the classifiers trained by current methods usually focus on easy-to-distinguish snippets while ignoring other semantically ambiguous features, which leads to incomplete and over-complete localization. To address these issues, we introduce a new Snippet-inter Difference Attention Network (SDANet) for WTAL, which can be trained end-to-end. Specifically, our model presents three modules, with primary contributions lying in the snippet-inter difference attention (SDA) module and potential feature mining (PFM) module. Firstly, we construct a simple multi-scale temporal feature fusion (MTFF) module to generate multi-scale temporal feature representation, so as to help the model better detect short action instances. Secondly, we consider the temporal cues of video features and design SDA module based on the Transformer to capture global discriminative features for each modality based on multi-scale features. It calculates the differences between temporal neighbor snippets in each modality to explore salient-difference features, and then utilizes them to guide correlation modeling. Thirdly, after learning discriminative features, we devise PFM module to excavate potential action and background snippets from ambiguous features. By contrastive learning, potential actions are forced closer to discriminative actions and away from the background, thereby learning more accurate action boundaries. Finally, two losses (i.e., similarity loss and reconstruction loss) are further developed to constrain the consistency between two modalities and help the model retain original feature information for better localization results. Extensive experiments show that our model achieves better performance against current WTAL methods on three datasets, i.e., THUMOS14, ActivityNet1.2 and ActivityNet1.3. Wei Zhou 0042, Kang Lin, Weipeng Hu, Haifeng Hu 0001, Yap-Peng Tan |
IEEE Trans. Multim. | 6 |
| 2025 | Disentangling Modality and Posture Factors: Memory-Attention and Orthogonal Decomposition for Visible-Infrared Person Re-IdentificationabstractStriving to match the person identities between visible (VIS) and near-infrared (NIR) images, VIS-NIR reidentification (Re-ID) has attracted increasing attention due to its wide applications in low-light scenes. However, owing to the modality and pose discrepancies exhibited in heterogeneous images, the extracted representations inevitably comprise various modality and posture factors, impacting the matching of cross-modality person identity. To solve the problem, we propose a disentangling modality and posture factors (DMPFs) model to disentangle modality and posture factors by fusing the information of features memory and pedestrian skeleton. Specifically, the DMPF comprises three modules: three-stream features extraction network (TFENet), modality factor disentanglement (MFD), and posture factor disentanglement (PFD). First, aiming to provide memory and skeleton information for modality and posture factors disentanglement, the TFENet is designed as a three-stream network to extract VIS-NIR image features and skeleton features. Second, to eliminate modality discrepancy across different batches, we maintain memory queues of previous batch features through the momentum updating mechanism and propose MFD to integrate features in the whole training set by memory-attention layers. These layers explore intramodality and intermodality relationships between features from the current batch and memory queues under the optimization of the optimal transport (OT) method, which encourages the heterogeneous features with the same identity to present higher similarity. Third, to decouple the posture factors from representations, we introduce the PFD module to learn posture-unrelated features with the assistance of the skeleton features. Besides, we perform subspace orthogonal decomposition on both image and skeleton features to separate the posture-related and identity-related information. The posture-related features are adopted to disentangle the posture factors from representations by a designed posture-features consistency (PfC) loss, while the identity-related features are concatenated to obtain more discriminative identity representations. The effectiveness of DMPF is validated through comprehensive experiments on two VIS-NIR pedestrian Re-ID datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Dynamic Modality-Camera-Invariant Clustering for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised learning visible-infrared person re-identification (USL-VI-ReID) offers a more flexible and cost-effective alternative compared to supervised methods. This field has gained increasing attention due to its promising potential. Existing methods simply cluster modality-specific samples and employ strong association techniques to achieve instance-to-cluster or cluster-to-cluster cross-modality associations. However, they ignore cross-camera differences, leading to noticeable issues with excessive splitting of identities. Consequently, this undermines the accuracy and reliability of cross-modal associations. To address these issues, we propose a novel dynamic modality-camera-invariant clustering (DMIC) framework for USL-VI-ReID. Specifically, our DMIC naturally integrates modality-camera-invariant expansion (MIE), dynamic neighborhood clustering (DNC), and hybrid modality contrastive learning (HMCL) into a unified framework, which eliminates both the cross-modality and cross-camera discrepancies in clustering. MIE fuses intermodal and intercamera distance coding to bridge the gaps between modalities and cameras at the clustering level. DNC employs two dynamic search strategies to refine the network's optimization objective, transitioning from improving discriminability to enhancing cross-modal and cross-camera generalizability. Moreover, HMCL is designed to optimize instance- and cluster-level distributions. Memories for intramodality and intermodality training are updated using randomly selected samples, facilitating real-time exploration of modality-invariant representations. Extensive experiments have demonstrated that our DMIC addresses the limitations present in current clustering approaches and achieves competitive performance, which significantly reduces the performance gap with supervised methods. Yiming Yang 0001, Weipeng Hu, Qiaolin He, Haifeng Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Unsupervised Visible-Infrared Person ReID via Modality-Camera Balance Label RefinementabstractUnsupervised Learning Visible-Infrared Person Re-Identification (USL-VI-ReID) focuses on developing a cross-modality retrieval model without the need for labels, minimizing the dependence on costly manual annotation across modalities. Recently, various approaches focus on reducing the cross-modality discrepancies. However, they ignore that USL-VI-ReID is also a task of solving discrepancies while exploring fine-grained information in hierarchical domains. In this article, we propose a hierarchical Modality-Camera Balance Label Refinement (MCBL) framework to balance the contributions of each camera-modality. Meanwhile, we explore the fine-grained features and refine the noise labels at each training stages. Specifically, our MCBL naturally combines Modality-Camera Balanced Label Mining (MBLM), Unreliable Pseudo-Label Re-align (UPR), and Hybrid Modality-Camera Contrastive Learning (HMCCL) into a unified framework, which balances the association information for each hierarchical domain through refining noise labels. Technically, MBLM filters cluster-level noise samples utilizing a modality-camera balance strategy, thereby ensuring that reliable samples are stored in memory for effective contrast learning. UPR refines the noise labels through the re-alignment methods at the instance level, thus improving the accuracy of labels and further enhancing the model’s generalization ability. Moreover, the key of HMCCL is optimizing the distribution at both the instance and cluster levels, which forces the sample to be close to its cluster proxy while being far from others in a real-time memory update phase. Extensive experiments have shown that our MCBL addresses the current limitations of camera discrepancy and achieves competitive performance. Jiakai He, Yiming Yang 0001, Haifeng Hu 0001, Ruixing Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | AMVFNet: Attentive Multi-View Fusion Network for 3D Object DetectionabstractPillar-based method is significant in the field of LiDAR-based 3D object detection which could directly make use of efficient 2D backbone and save computational resources during reference. Existing methods usually sequentially project the original point clouds into the cylindrical view or the bird-eye view for feature extraction. However, the former suffers from obscured problems and the scales of instances vary greatly with distance, while the latter leads to considerable confusion problems due to the loss of semantic information caused by the sparsity of the projected point cloud. In this article, we present a novel and efficient two-stage point-pillar hybrid architecture named Attentive Multi-View Fusion Network (AMVFNet), in which we abstract features from all cylindrical view, bird-eye view, and raw point clouds. Rather than designing more complex modules to solve the problems inherent in the single-view approach, our multi-view fusion architecture effectively combines the strengths of multiple perspectives to improve performance at a more fundamental level. Besides, to compensate for quantization distortion caused by projection operations, we propose attentive feature enhancement layers to further improve the capability of contextual information capturing. Extensive experiments on the KITTI detection benchmark illustrate that our proposed AMVFNet achieves competitive performance compared with other SOTA 3D object detectors. Haifeng Hu 0001, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Text-and-Image Learning Transformer for Cross-Modal Person Re-IdentificationabstractText-based person re-identification aims to find the target person from a large pedestrian gallery with the given natural language description. Previous works mainly focus on embedding salient textual and visual representations in a common latent space by utilizing the dual-path structure or parameter-shared network. However, they still lack the ability to effectively extract fine-grained unimodal features as well as fuse the cross-modal data, leading to the increase of misaligned cases. To settle these issues, we propose a text-and-image implicit learning Transformer (TILT) to eliminate textual anisotropy and enhance the cross-modal alignment from both domains based on the bi-direction multi-modal encoders. Specifically, we apply the pre-trained multi-modal embedding module to overcome the unimodal anisotropy problem with contrastive learning, and map fine-grained features with dual encoder in bi-directional masking. Then, we design the cross-modal interaction encoder to comprehensively mine implicit cross-modal relations by reconstructing masked tokens, and fuse rich multi-modal knowledge in a common space. In addition, the cross-modal similarity matching module is proposed to optimize the intra-domain classification and decrease the inter-domain divergence. Extensive experiments are conducted on three public benchmarks CUHK-PEDES, ICFG-PEDES, and RSTPReid to verify the effectiveness of our proposed framework. Results prove that our model outperforms state-of-the-art methods on all metrics. Tinghui Wu, Shuhe Zhang, Dihu Chen, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Triple-Stream Commonsense Circulation Transformer Network for Image Captioning
Jianchao Li, Wei Zhou 0042, Kai Wang 0033, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Relation-dependent contrastive learning with cluster sampling for inductive relation prediction
Aolin Xiong, Sijie Mai, Haifeng Hu 0001 |
Neurocomputing | 4 |
| 2024 | Multi-Task Momentum Distillation for Multimodal Sentiment AnalysisabstractIn the field of Multimodal Sentiment Analysis (MSA), the prevailing methods are devoted to developing intricate network architectures to capture the intra- and inter-modal dynamics, which necessitates numerous parameters and poses more difficulties in terms of interpretability in multimodal modeling. Besides, the heterogeneous nature of multiple modalities (text, audio, and vision) introduces significant modality gaps, thereby making multimodal representation learning an ongoing challenge. To address the aforementioned issues, by considering the learning process of modalities as multiple subtasks, we propose a novel approach named Multi-Task Momentum Distillation (MTMD) which succeeds in reducing the gap among different modalities. Specifically, according to the abundance of semantic information, we treat the subtasks of textual and multimodal representations as the teacher networks while the subtasks of acoustic and visual representations as the student ones to present knowledge distillation, which transfers the sentiment-related knowledge guided by the regression and classification subtasks. Additionally, we adopt unimodal momentum models to explore modality-specific knowledge deeply and employ adaptive momentum fusion factors to learn a robust multimodal representation. Furthermore, we provide a theoretical perspective of mutual information maximization by interpreting MTMD as generating sentiment-related views in various ways. Extensive experiments illustrate the superiority of our approach compared with the state-of-the-art methods in MSA. Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Dual Learning for Conversational Emotion Recognition and Emotional Response GenerationabstractEmotion recognition in conversation (ERC) and emotional response generation (ERG) are two important NLP tasks. ERC aims to detect the utterance-level emotion from a dialogue, while ERG focuses on expressing a desired emotion. Essentially, ERC is a classification task, with its input and output domains being the utterance text and emotion labels, respectively. On the other hand, ERG is a generation task with its input and output domains being the opposite. These two tasks are highly related, but surprisingly, they are addressed independently without making use of their duality in prior works. Therefore, in this paper, we propose to solve these two tasks in a dual learning framework. Our contributions are fourfold: (1) We propose a dual learning framework for ERC and ERG. (2) Within the proposed framework, two models can be trained jointly, so that the duality between them can be utilised. (3) Instead of a symmetric framework that deals with two tasks of the same data domain, we propose a dual learning framework that performs on a pair of asymmetric input and output spaces, i.e., the natural language space and the emotion labels. (4) Experiments are conducted on benchmark datasets to demonstrate the effectiveness of our framework. Shuhe Zhang, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | DATran: Dual Attention Transformer for Multi-Label Image ClassificationabstractMulti-label image classification is a fundamental yet challenging task, which aims to predict the labels associated with a given image. Most of previous methods directly exploit the high-level features from the last layer of convolutional neural network for classification. However, these methods cannot obtain global features due to the limited size of convolutional kernels, and they fail to extract multi-scale features to effectively recognize small-scale objects in the images. Recent studies exploit the graph convolution network to model the label correlations for boosting the classification performance. Despite substantial progress, these methods rely on manually pre-defined graph structures. Besides, they ignore the associations between semantic labels and image regions, and do not fully explore the spatial context of images. To address above issues, we propose a novel Dual Attention Transformer (DATran) model, which adopts a dual-stream architecture that simultaneously learns spatial and channel correlations from multi-label images. Firstly, in order to solve the problem that current methods are difficult to recognize small-size objects, we develop a new multi-scale feature fusion (MSFF) module to generate multi-scale feature representation by jointly integrating both high-level semantics and low-level details. Secondly, we design a prior-enhanced spatial attention (PSA) module to learn the long-range correlation between objects from different spatial positions in images to enhance the model performance. Thirdly, we devise a prior-enhanced channel attention (PCA) module to capture the inter-dependencies between different channel maps, thus effectively improving the correlation between semantic categories. It is worth noting that PSA module and PCA module complement and promote each other to further augment the feature representations. Finally, the outputs of these two attention modules are fused to obtain the final features for classification. Performance evaluation experiments are conducted on MS-COCO 2014, PASCAL VOC 2007 and VG-500 datasets, demonstrating that DATran model achieves better performance than current state-of-the-art models. Wei Zhou 0042, Zhijie Zheng 0003, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Mind the Inconsistent Semantics in Positive Pairs: Semantic Aligning and Multimodal Contrastive Learning for Text-Based Pedestrian SearchabstractAiming at retrieving pedestrian images based on a provided textual description query, Text-Based Pedestrian Search (TBPS) has gained attention due to its implications in public security tasks such as suspect tracking. Nevertheless, the modality discrepancies between textual descriptions and visual images pose a challenge in aligning semantic information between these two modalities. Moreover, the text description annotated on a particular pedestrian image may not align with the content of other images sharing the same identity, due to variations in viewpoint. These text-image pairs exhibiting inconsistent semantics, termed weak positive pairs, have a discernible impact on the model’s performance. To address these challenges, we propose a Semantic Aligning and Multimodal Contrastive learning (SAMC) model to capture cross-modality identity-invariant features, including three modules: Multi-modality Features Fusion (MFF), Semantic-aligning Optimal Transport (SOT), and Multi-modality Contrastive Learning (MCL). Firstly, the MFF is designed to fuse textual and visual information and extract identity-discriminative multimodal features using self- and cross-attention mechanisms. The multimodal features act as anchors, bridging the gap between the two modalities and enhancing the identity-invariance of unimodal features. Secondly, the SOT is designed to address the semantic misalignment issue between textual descriptions and visual images. Utilizing the Optimal Transport (OT) theory, SOT encourages high features similarity between positive samples from different modalities, thereby exploring semantic relationships between image and text data without requiring extra supervised labels. Lastly, the MCL is introduced to narrow the modality gap, compelling two unimodal features towards the identity-discriminative multimodal features through contrastive learning. Different temperature coefficients are employed for strong and weak positive pairs to mitigate the inconsistency in text-image pair correlation. The effectiveness of SAMC is validated by extensive comprehensive experiments on three TBPS datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Unsupervised NIR-VIS Face Recognition via Homogeneous-to-Heterogeneous Learning and Residual-Invariant EnhancementabstractNear-Infrared and Visible light (NIR-VIS) face recognition methods have achieved remarkable success in the fields of security surveillance, criminal investigation, and multimedia information retrieval. But the existing methods heavily rely on carefully annotated labels, leading to expensive manual labelling consumption and deployment flexibility. This motivates us to design unsupervised methods to address NIR-VIS recognition without relying on label information. To this end, we propose a novel homogeneous-to-HEterogeneous learning and Residual-invariant Enhancement (HERE) network for Unsupervised NIR-VIS Heterogeneous Face Recognition (NIR-VIS-UHFR). As the name suggests, the optimization of HERE follow a ”homogeneous-to-heterogeneous learning” strategy to fully explore complementary and common semantic information across different modalities. During the homogeneous learning phase, Modality-Adversarial Contrastive Learning (MACL) leverages the collaboration of modality contrastive learning and adversarial learning. On the one hand, MACL learns compact and discriminative intra-modal representations for NIR and VIS data, respectively. On the other hand, MACL guarantees that NIR-VIS data conform to the common feature distribution in a shared feature space, effectively reducing modal differences even in the absence of identity information between modalities. In the heterogeneous learning phase, K-reciprocal-Encoding-based Cross-modal Labeling (KECL) is introduced as robust pseudo label estimation to fully explore cross-modal relationships and group cross-modal features into clusters. With the pseudo labels provided by KECL, Refined cross-modal Contrastive Learning (RCL) is developed with modality-invariant averaging initialization and dynamic focus weighting strategies to extract modality-invariant features. Finally, Residual-invariant Representations Enhancement (RRE) mines partial features under the cross-modal face for robust matching. Compared to supervised methods, our unsupervised HERE demonstrates comparable performance on multiple datasets, greater scalability and practicality in deployment by reducing data acquisition requirements and costs. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Pseudo Label Association and Prototype-Based Invariant Learning for Semi-Supervised NIR-VIS Face RecognitionabstractRemarkable success of the existing Near-InfraRed and VISible (NIR-VIS) approaches owes to sufficient labeled training data. However, collecting and tagging data from different domains is a time-consuming and expensive task. In this paper, we tackle the NIR-VIS face recognition problem in a semi-supervised manner, termed as semi-supervised NIR-VIS Heterogeneous Face Recognition (NIR-VIS-sHFR). To cope with this problem, we propose a novel pseudo Label association and Prototype-based invariant Learning (LPL), consisting of three key components, i.e., Cross-domain pseudo Label Association (CLA), Intra-domain Compact Representation learning (ICR), and Prototype-based Inter-domain Invariant learning (PII). Firstly, the CLA iteratively builds inter-domain association graphs for pseudo-label association, subsequently facilitating cross-domain model development based on the generated pseudo-labels. Furthermore, the ICR is proposed to achieve the separation of in-domain features from different clusters and the aggregation of features from the same cluster, by performing cluster adaptation learning with prototype-based initialization. Finally, with the cross-domain pseudo-label training data produced by CLA, the PII explores potential domain-invariant and identity-related features, which employs cross-domain prototypes with identity-associated momentum updating to effectively guide inter-domain instances learning. The semi-supervised LPL method achieves comparable performance to recent supervised learning methods on multiple challenging NIR-VIS datasets, which demonstrates that the LPL is capable of learning robust cross-domain representations even without identity label information. Weipeng Hu, Yiming Yang 0001, Haifeng Hu 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Spatial and Temporal Dual-Attention for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to learn discriminative representations for person retrieval from unlabeled data. Recent research accomplishes this task with pseudo-labels and a center-level memory, but the pseudo-labels are inherently noisy and the update status of central features in memory is inconsistent, thus reducing the accuracy of Re-ID. In this paper, we propose a novel Spatial and Temporal Dual-Attention (STDA) framework to solve the above two problems. Firstly, to overcome the noisy label problem, we design an Intra-class Neighbor-based Spatial Attention (INSA) module to refine pseudo-labels by mining the spatial-level connections of positive instances. Specifically, we design a neighbor agreement as the similarity between central features and positive instance features in feature space to exploit the reliable complementary relationship. Based on the neighbor agreement, we aggregate the predictions of positive instances, thus jointly mitigating the noise in single instance feature clustering. Secondly, the central features in memory cannot be updated simultaneously, which leads to an inter-class update inconsistency problem. We introduce an Inter-class Sequence-based Temporal Attention (ISTA) module to alleviate this problem by computing a memory update factor based on the temporal-level update sequence of centrals. The weights are larger for newer updated centrals than for older updated centrals and non-updated centrals to mitigate the inter-class inconsistent update status. Finally, we combine the INSA and ISTA modules with contrastive learning for training. Extensive experimental results on Market-1501, MSMT17, and VeRi-776 show the effectiveness of the proposed method over the state-of-the-art performance. The code is available at:https://github.com/heqlin5/STDA. Qiaolin He, Zhijie Zheng 0003, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Dynamically Shifting Multimodal Representations via Hybrid-Modal Attention for Multimodal Sentiment AnalysisabstractIn the field of multimodal machine learning, multimodal sentiment analysis task has been an active area of research. The predominant approaches focus on learning efficient multimodal representations containing intra- and inter-modality information. However, the heterogeneous nature of different modalities brings great challenges to multimodal representation learning. In this article, we propose a multi-stage fusion framework to dynamically fine-tune multimodal representations via a hybrid-modal attention mechanism. Previous methods mostly only fine-tune the textual representation due to the success of large corpus pre-trained models and neglect the inconsistency problem of different modality spaces. Thus, we design a module called the Multimodal Shifting Gate (MSG) to fine-tune the three modalities by modeling inter-modality dynamics and shifting representations. We also adopt a module named Masked Bimodal Adjustment (MBA) on the textual modality to improve the inconsistency of parameter spaces and reduce the modality gap. In addition, we utilize syntactic-level and semantic-level textual features output from different layers of the Transformer model to sufficiently capture the intra-modality dynamics. Moreover, we construct a Shifting HuberLoss to robustly introduce the variation of the shifting value into the training process. Extensive experiments on the public datasets, including CMU-MOSI and CMU-MOSEI, demonstrate the efficacy of our approach. Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Tri-Level Modality-Information Disentanglement for Visible-Infrared Person Re-IdentificationabstractAiming to match the person identity between daytime VISible (VIS) and nighttime Near-InfraRed (NIR) images, VIS-NIR re-identification (Re-ID) has attracted increasing attention due to its wide applications in low-light scenes. However, dramatic modality discrepancies between VIS and NIR images lead to a considerable intra-class gap in the feature space, which impacts identity matching. To bridge the modality gap, we propose a Tri-level Modality-information Disentanglement (TMD) to disentangle modality information at the levels of raw image, features distribution and instance features. Our model consists of three key modules, including Style-Aligned Converter (SAC), Two-Steps Wasserstein Loss (TSWL) and Self-supervised Orthogonal Disentanglement (SOD) to handle the modality information at the three levels. Firstly, aiming at reducing modality discrepancy at image-level, the SAC is introduced to generate style-aligned images by the designed style converter and$\mathcal {A}$-distance learning approach. The SAC can effectively alleviate the style discrepancy between VIS and NIR images with a negligible increase in model complexity. Secondly, considering the heterogeneity of VIS and NIR feature distribution caused by the structure- and style-misaligned raw images, we propose the TSWL to decrease the VIS-NIR gap at distribution-level by two distribution alignment steps. Specifically, after generating style-consistent images, we eliminate modality-related discrepancy by aligning the distribution between structure-aligned original and generated VIS/NIR images and bridge the modality-unrelated gap by aligning the style-consistent generated VIS-NIR images. Thirdly, focusing on further reducing the modality discrepancy at instance-level, the SOD is presented to construct orthogonal constraints between the extracted modality- and identity-related features. Since the modality-related factors are disentangled from the instance features, the proposed TMD efficiently learns the modality-unrelated and identity-discriminative representations, which are productive to conduct person Re-ID task on the VIS-NIR images. Comprehensive experiments are carried out on two cross-modality pedestrian Re-ID datasets to demonstrate the effectiveness of TMD. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Multimodal Boosting: Addressing Noisy Modalities and Identifying Modality ContributionabstractIn multimodal representation learning, different modalities do not contribute equally. Especially when learning with noisy modalities that convey non-discriminative information, the prediction based on multimodal representation is often biased and even ignores the knowledge from informative modalities. In this paper, we aim to address the noisy modality problem and balance the contributions of multiple modalities dynamically in a parallel format. Specifically, we construct multiple base learners and formulate our framework as a boosting-like algorithm, where different base learners focus on different aspects of multimodal learning. To identify the contributions of individual base learners, we develop a contribution learning network that dynamically determines the contribution and noise level of each base learner. In contrast to the commonly considered attention mechanism, we define the transformation of predictive loss as the supervision signal to train the contribution learning network, which enables more accurate learning of modality importance. We derive the final prediction by incorporating the predictions of base learners based on their contributions. Notably, different from late fusion, we devise a multimodal base learner to explore the cross-modal interactions. To update the network, we design the ‘complementary update mechanism’, where for each base learner, we assign higher weights to those samples that are incorrectly predicted by other base learners. In this way, we can leverage the available information to correctly predict each sample to the utmost extent and enable different base learners to learn different aspects of multimodal information. Extensive experiments demonstrate that the proposed method achieves superior performance on multimodal sentiment analysis and emotion recognition. Sijie Mai, Aolin Xiong, Haifeng Hu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Multimodal Reaction: Information Modulation for Cross-Modal Representation LearningabstractIn multimodal machine learning, proper handling of cross-modal information is essential for obtaining an ideal joint embedding. Despite the progress made by recent fusion strategies, we hold that before the fusion stage, the unimodal representation inevitably contains noise that may hinder the correct learning of cross-modal dynamics and affect multimodal fusion. It is worthwhile to investigate how the information is being utilized and how to make the full use of it. Rethinking the process of leveraging multiple modalities for the joint embedding, multimodal learning can be regarded as achemical reactionprocess and two steps may benefit learning: 1) purification to filter impurity, and 2) catalyst to facilitate learning. In this paper, we propose aMultimodalInformationModulation (MIM) learning framework to modulate the contribution and utilization of the cross-modal information, which identifies and handles the ‘impurity’ and ‘catalyst’ in multimodal learning. Specifically, a Unimodal Purification Network (UPN) is proposed to identify and explicitly filter out the impurity within each modality before fusion, which reduces the possibility of learning incorrect cross-modal dynamics. Besides, based on the intuition that useful information has the potential in the guidance of model updating, it plays a role to facilitate learning, which is achieved by the design of the Knowledge Guidance Scheme (KGS) considering both the intra- and inter-modal scenarios. Different to a majority of works that emphasize the role of useful information in the fusion and inference stage, KGS considers its potential role in assisting the representation learning of weaker components. Besides, it fully considers the modality dominance problem and sample variations for optimization. In short, MIM manages to modulate the useless/useful information to minimize/emphasize their contribution. Experimental results verify the effectiveness of the proposed method. Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Mining Semantic Information With Dual Relation Graph Network for Multi-Label Image ClassificationabstractThe purpose of multi-label image classification is to assign multiple labels for multiple objects presented in one image. Recent research efforts exploit graph convolution network (GCN) to learn the label co-occurrence dependencies for enhancing the semantic representation. Although these methods have achieved promising results, they can not capture the intrinsic correlation between objects in images and do not consider the inter-channel relationship. In addition, the previous methods treat each single image independently and fail to explore the relationship between different images. To address the above challenges, we propose a novelDualRelationGraphNetwork (DRGN) model, which adopts a double branch structure to excavate rich semantic information from intra-image and cross-image simultaneously. Specifically, we first develop an intra-image channel-relation mining (ICM) module to mine the inter-channel relationship in features while learning the importance of different channels. Secondly, we design a new GCN-based intra-image spatial-relation exploring (ISE) module to capture the correlation between objects in individual image. Notably, ISE module and ICM module can complement and promote each other from the spatial and channel dimensions of images to improve the correlation between objects in individual image. Thirdly, we propose a novel GCN-based cross-image semantic learning (CSL) module to learn the semantic relationship between different images in the mini-batch. Through graph reasoning, our CSL module can iteratively refine input image features by acquiring common semantic information from other images in the mini-batch. Extensive experiments on the MS-COCO 2014, PASCAL VOC 2007, and VG-500 datasets demonstrate that the proposed DRGN model outperforms current state-of-the-art methods. Wei Zhou 0042, Weitao Jiang, Dihu Chen, Haifeng Hu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Inverse-Free DZNN Models for Solving Time-Dependent Linear System via High-Precision Linear Six-Step MethodabstractTime-dependent linear system (TDLS) is usually encountered in scientific research, which is the mathematical formulation of many practical applications. Different from conventional inverse-need models, by utilizing zeroing neural network (ZNN) method twice, an inverse-free continuous ZNN (CZNN) model is developed for solving TDLS. For conveniently practical use, a discrete model is naturally desired. Superior to conventional discretization methods, a general linear six-step (LSS) method with the seventh-order precision and five variable parameters is proposed for the first time. Constraints about five variable parameters are theoretically analyzed to guarantee the efficacy of the general LSS method. Within constraints, 12 specific LSS methods are further developed. Aided with the general LSS method, an inverse-free discrete ZNN (DZNN) is proposed and termed DZNN-LSS model, and its precision is greatly improved compared with conventional discrete models. For comparison, three conventional discretization methods are also utilized to generate DZNN models. Detailed theoretical analyses are provided to prove the efficacy of relevant models. In addition, a specific TDLS example is considered to show the effectiveness and superiority of the DZNN-LSS model. More than that, applications to manipulator control and sound source localization are conducted to illustrate the applicability of the DZNN-LSS model. Min Yang 0010, Yunong Zhang, Haifeng Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Feature Map is Worth a Video Frame: Rethinking Convolutional Features for Visible-Infrared Person Re-identificationabstractVisible-Infrared Person Re-identification (VI-ReID) aims to search for the identity of the same person across different spectra. The feature maps obtained from the convolutional layers are generally used for loss calculation in the later stages of the model in VI-ReID, but their role in the early and middle stages of the model remains unexplored. In this article, we propose a novel Rethinking Convolutional Features (ReCF) approach for VI-ReID. ReCF consists of two modules: Middle Feature Generation (MFG), which utilizes the feature maps in the early stage to reduce significant modality gap, and Temporal Feature Aggregation (TFA), which uses the feature maps in the middle stage to aggregate multi-level features for enlarging the receptive field. MFG generates middle modality features in the form of a learnable convolution layer as a bridge between RGB and IR modalities, which is more flexible than using fixed-parameter grayscale images and yields a better middle modality to further reduce the modality gap. TFA first treats the convolution process as a video sequence, and the feature map of each convolution layer can be considered a worthwhile video frame. Based on this, we can obtain a multi-level receptive field and a temporal refinement. In addition, we introduce a color-unrelated loss and a modality-unrelated loss to constrain the modality features for providing a common feature representation space. Experimental results on the challenging VI-ReID datasets demonstrate that our proposed method achieves state-of-the-art performance. Qiaolin He, Zhijie Zheng 0003, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Self-Supervised Consistency Based on Joint Learning for Unsupervised Person Re-identificationabstractRecently, unsupervised domain adaptive person re-identification (Re-ID) methods have been extensively studied thanks to not requiring annotations, and they have achieved excellent performance. Most of the existing methods aim to train the Re-ID model for learning a discriminative feature representation. However, they usually only consider training the model to learn a global feature of a pedestrian image, but neglecting the local feature, which restricts further improvement of model performance. To address this problem, two local branches are added to the networks, aiming to allow the model to focus on the local feature containing identity information. Furthermore, we propose a self-supervised consistency constraint to further improve robustness of the model. Specifically, the self-supervised consistency constraint uses the basic data augmentation operations without other auxiliary networks, which can improve performance of the model effectively. Then, a learnable memory matrix is designed to store the mapping vectors that maps person features into probability distributions. Finally, extensive experiments are conducted on multiple commonly used person Re-ID datasets to verify the effectiveness of the proposed generative adversarial networks fusing global and local features. Experimental results reveal that our method achieves results comparable to state-of-the-art methods. Xulei Lou, Tinghui Wu, Haifeng Hu 0001, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Syncretic Space Learning Network for NIR-VIS Face RecognitionabstractTo overcome the technical bottleneck of face recognition in low-light scenarios, Near-InfraRed and VISible (NIR-VIS) heterogeneous face recognition is proposed for matching well-lit VIS faces with poorly lit NIR faces. Current cross-modal synthesis methods visually convert the NIR modality to the VIS modality and then perform face matching in the VIS modality. However, using a heavyweight GAN network on unpaired NIR-VIS faces may lead to high synthesis difficulty, low inference efficiency, and other problems. To alleviate the above problems, we simultaneously synthesize NIR and VIS images into modality-independent syncretic images and propose a novel syncretic space learning (SSL) model to eliminate the modal gap. First, Syncretic Modality Generator (SMG) synthesizes NIR and VIS images into syncretic images using channel-level convolution with a shallow CNN. In particular, the discriminative structural information is well preserved and the face quality can be further improved with small modal variations in a self-supervised learning manner. Second, Modality-adversarial Syncretic space Learning (MSL) projects NIR and VIS images into the syncretic space by a syncretic-modality adversarial learning strategy with syncretic pattern guided objective, so the modal gap of NIR-VIS faces can be effectively reduced. Finally, the Syncretic Distribution Consistency (SDC) constructed by NIR-syncretic, syncretic-syncretic, and VIS-syncretic consistency can enhance the intra-class compactness and learn discriminative representations. Extensive experiments on three challenging datasets demonstrate the effectiveness of the SSL method. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Computer Simulations of Applying Zhang Inequation Equivalency and Solver of Neurodynamics to Redundant Manipulators at Acceleration Level
Ji Lu, Min Yang 0010, Ning Tan 0003, Haifeng Hu 0001, Yunong Zhang |
ICONIP (1) | 4 |
| 2023 | Complementary Attention Network for Weakly Supervised Temporal Action Localization
Peng Dou, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2023 | Hadamard Product Perceptron Attention for Image Captioning
Weitao Jiang, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2023 | Unsupervised Person Re-identification Using Unified Domanial Learning
Suian Zhang, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2023 | Egocentric Action Recognition by Automatic Relation ModelingabstractEgocentric videos, which record the daily activities of individuals from a first-person point of view, have attracted increasing attention during recent years because of their growing use in many popular applications, including life logging, health monitoring and virtual reality. As a fundamental problem in egocentric vision, one of the tasks of egocentric action recognition aims to recognize the actions of the camera wearers from egocentric videos. In egocentric action recognition, relation modeling is important, because the interactions between the camera wearer and the recorded persons or objects form complex relations in egocentric videos. However, only a few of existing methods model the relations between the camera wearer and the interacting persons for egocentric action recognition, and moreover they require prior knowledge or auxiliary data to localize the interacting persons. In this work, we consider modeling the relations in a weakly supervised manner, i.e., without using annotations or prior knowledge about the interacting persons or objects, for egocentric action recognition. We form a weakly supervised framework by unifying automatic interactor localization and explicit relation modeling for the purpose of automatic relation modeling. First, we learn to automatically localize the interactors, i.e., the body parts of the camera wearer and the persons or objects that the camera wearer interacts with, by learning a series of keypoints directly from video data to localize the action-relevant regions with only action labels and some constraints on these keypoints. Second, more importantly, to explicitly model the relations between the interactors, we develop an ego-relational LSTM (long short-term memory) network with several candidate connections to model the complex relations in egocentric videos, such as the temporal, interactive, and contextual relations. In particular, to reduce human efforts and manual interventions needed to construct an optimal ego-relational LSTM structure, we search for the optimal connections by employing a differentiable network architecture search mechanism, which automatically constructs the ego-relational LSTM network to explicitly model different relations for egocentric action recognition. We conduct extensive experiments on egocentric video datasets to illustrate the effectiveness of our method. Haoxin Li, Wei-Shi Zheng 0001, Jianguo Zhang 0001, Haifeng Hu 0001, Jiwen Lu, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Feature learning network with transformer for multi-label image classification
Wei Zhou 0042, Peng Dou, Haifeng Hu 0001, Zhijie Zheng 0003 |
Pattern Recognit. | 4 |
| 2023 | PSA-Det3D: Pillar set abstraction for 3D object detection
Zhijie Zheng 0003, Haifeng Hu 0001, Dihu Chen |
Pattern Recognit. Lett. | 4 |
| 2023 | DTSSD: Dual-Channel Transformer-Based Network for Point-Based 3D Object DetectionabstractIn the field of 3D object detection, previous methods mainly utilize one channel feature encoding network to extract point-wise features. Despite the effectiveness, we find that only leveraging one channel encoding network is not sufficient and impedes the detection performance. To this end, we propose a dual-channel transformer-based feature encoding network, which integrates both set abstraction layer and transformer block as backbone. It enables the model to exploit fine-grained as well as long-range contextual information of objects, thus providing complementary relationship of two methods. In addition, a centroid estimation module is introduced to obtain powerful representation of the whole object. Finally, considering the significance of point density, which is crucial for detection performance, we propose a central density-aware enhancement module to equip center features with distinct density features. Experimental results on KITTI dataset show the effectiveness of our proposed method. Zhijie Zheng 0003, Haifeng Hu 0001, Dihu Chen |
IEEE Signal Process. Lett. | 4 |
| 2023 | MissModal: Increasing Robustness to Missing Modality in Multimodal Sentiment AnalysisabstractAbstract When applying multimodal machine learning in downstream inference, both joint and coordinated multimodal representations rely on the complete presence of modalities as in training. However, modal-incomplete data, where certain modalities are missing, greatly reduces performance in Multimodal Sentiment Analysis (MSA) due to varying input forms and semantic information deficiencies. This limits the applicability of the predominant MSA methods in the real world, where the completeness of multimodal data is uncertain and variable. The generation-based methods attempt to generate the missing modality, yet they require complex hierarchical architecture with huge computational costs and struggle with the representation gaps across different modalities. Diversely, we propose a novel representation learning approach named MissModal, devoting to increasing robustness to missing modality in a classification approach. Specifically, we adopt constraints with geometric contrastive loss, distribution distance loss, and sentiment semantic loss to align the representations of modal-missing and modal-complete data, without impacting the sentiment inference for the complete modalities. Furthermore, we do not demand any changes in the multimodal fusion stage, highlighting the generality of our method in other multimodal learning systems. Extensive experiments demonstrate that the proposed method achieves superior performance with minimal computational costs in various missing modalities scenarios (flexibility), including severely missing modality (efficiency) on two public MSA datasets. Ronghao Lin, Haifeng Hu 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Hybrid Contrastive Learning of Tri-Modal Representation for Multimodal Sentiment AnalysisabstractThe wide application of smart devices enables the availability of multimodal data, which can be utilized in many tasks. In the field of multimodal sentiment analysis, most previous works focus on exploring intra- and inter-modal interactions. However, training a network with cross-modal information (language, audio and visual) is still challenging due to the modality gap. Besides, while learning dynamics within each sample draws great attention, the learning of inter-sample and inter-class relationships is neglected. Moreover, the size of datasets limits the generalization ability of the models. To address the afore-mentioned issues, we propose a novel framework HyCon for hybrid contrastive learning of tri-modal representation. Specifically, we simultaneously perform intra-/inter-modal contrastive learning and semi-contrastive learning, with which the model can fully explore cross-modal interactions, learn inter-sample and inter-class relationships, and reduce the modality gap. Besides, refinement term and modality margin are introduced to enable a better learning of unimodal pairs. Moreover, we devise pair selection mechanism to identify and assign weights to the informative negative and positive pairs. HyCon can naturally generate many training pairs for better generalization and reduce the negative effect of limited datasets. Extensive experiments demonstrate that our method outperforms baselines on multimodal sentiment analysis and emotion recognition. Sijie Mai, Shuangjia Zheng, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Learning to Learn Better Unimodal Representations via Adaptive Multimodal Meta-LearningabstractMultimodal sentiment analysis is an emerging field of artificial intelligence. The most predominant approaches have made notable progress by designing sophisticated fusion architectures, exploring inter-modal interactions between modalities. However, these works tend to utilize a uniform optimization strategy for each modality, so that only sub-optimal unimodal representations are obtained for multimodal fusion. To address this issue, we propose a novel meta-learning based paradigm that can retain the advantages of unimodal existence and further boost the performance of multimodal fusion. Specifically, we introduce the Adaptive Multimodal Meta-Learning (AMML) to meta-learn the unimodal networks and adapt them for multimodal inference. AMML can (1) effectively obtain more optimized unimodal representation via meta-training on unimodal tasks, which adaptively adjusts the learning rate and assigns a more specific optimization procedure for each modality; (2) and adapt the optimized unimodal representations for multimodal fusion via meta-testing on multimodal tasks. Considering multimodal fusion often suffers from the distributional mismatches between features of different modalities due to heterogeneous nature of the signals, we implement a distribution transformation layer on unimodal representations to regularize the unimodal distributions. In this way, distribution gaps can be reduced to achieve a better effect of fusion. Extensive experiments on two widely-used datasets demonstrate that AMML achieves state-of-the-art performance. Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Neutral Face Learning and Progressive Fusion Synthesis Network for NIR-VIS Face RecognitionabstractTo meet the strong demand for deploying face recognition systems in low-light scenarios, the Near-InfraRed and VISible (NIR-VIS) face recognition task is receiving increasing attention. However, heterogeneous faces have the characteristics of heterogeneity and non-neutrality. Heterogeneity refers to the fact that the matching images are in different modalities, and non-neutrality means that the matching images are significantly different in pose, expression, lighting, etc. Both situations pose challenges for NIR-VIS face matching. To address this problem, we propose a novel Neutral face Learning and Progressive Fusion synthesis (NLPF) network to disentangle the latent attributes of heterogeneous faces and learn neutral face representations. Our approach naturally integrates Identity-related Neutral face Learning (INL) and Attribute Progressive Fusion (APF) into a joint framework. Firstly, INL eliminates modal variations and residual variations by guiding the network to learn homogeneous neutral face feature representations, which tackles the challenge of heterogeneity and non-neutrality by mapping cross-modal images to a common neutral representation subspace. Besides, APF is presented to perform the disentanglement and reintegration of identity-related features, modality-related features and residual features in a progressive fusion manner, which helps to further purify identity-related features. Comprehensive evaluations are carried out on three mainstream NIR-VIS datasets to verify the robustness and effectiveness of the NLPF model. In particular, NLPF has competitive recognition performance on LAMP-HQ, the most challenging NIR-VIS dataset so far. Yiming Yang 0001, Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Diverse Feature Learning Network With Attention Suppression and Part Level Background Suppression for Person Re-IdentificationabstractIn this paper, we propose a Diverse Feature Learning Network with Attention Suppression and Part Level Background Suppression (DFLN) for person re-identification (ReID). DFLN includes two key components: attention suppression mechanism (ASM) and part level background suppression mechanism (PLBSM). Firstly, despite attention mechanism has made great progress in current state-of-the-art ReID methods, they can only pay attention to the most salient region but ignore other discriminative information limiting the diversity of networks, which is not optimal for ReID due to the models tend to match persons by diverse clues (e.g., legs, arms, body, logo of clothes). To tackle the limitation above, we propose the ASM to assist the network to make full use of the most salient features and capture the other sub-salient features, so as to attain diverse features to improve the network performance. Secondly, we adopt a novel PLBSM to develop the part-based method which is proved effective for enhancing the diversity of ReID network. The PLBSM consists of a part feature refined module and a background suppression loss function, and aims to attain pure part level feature by filtering background clutter. Our DFLN integrates the ASM and part-based method developed by PLBSM into an end-to-end network and is able to extract robust diversity feature representations leading to higher performance. Extensive experimental results demonstrate the effectiveness of each component and our method achieves state-of-the-art results on mainstream person re-identification datasets. Shengrong Yang, Weihong Liu, Yangbin Yu, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Explicit Linear Left-and-Right 5-Step Formulas With Zeroing Neural Network for Time-Varying ApplicationsabstractIn this article, being different from conventional time-discretization (simply called discretization) formulas, explicit linear left-and-right 5-step (ELLR5S) formulas with sixth-order precision are proposed. The general sixth-order ELLR5S formula with four variable parameters is developed first, and constraints of these four parameters are displayed to guarantee the zero stability, consistence, and convergence of the formula. Then, by choosing specific parameter values within constraints, eight specific sixth-order ELLR5S formulas are developed. The general sixth-order ELLR5S formula is further utilized to generate discrete zeroing neural network (DZNN) models for solving time-varying linear and nonlinear systems. For comparison, three conventional discretization formulas are also utilized. Theoretical analyses are presented to show the performance of ELLR5S formulas and DZNN models. Furthermore, abundant experiments, including three practical applications, that is, angle-of-arrival (AoA) localization and two redundant manipulators (PUMA560 manipulator and Kinova manipulator) control, are conducted. The synthesized results substantiate the efficacy and superiority of sixth-order ELLR5S formulas as well as the corresponding DZNN models. Min Yang 0010, Yunong Zhang, Ning Tan 0003, Haifeng Hu 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | Modality and Camera Factors Bi-Disentanglement for NIR-VIS Object Re-IdentificationabstractAiming to match object identities across different modalities of images, the challenging task namely cross-modality object re-identification (NIR-VIS object Re-ID), has attracted increasing attention due to its wide application in low-light scenes. However, dramatic modality-dependent and camera-related discrepancies between Near-InfraRed-spectrum (NIR) and VISible-spectrum (VIS) images lead to a considerable intra-class gap in feature space. To address the problem, we propose a novel Modality and Camera factors Bi-Disentanglement (MCBD) model to learn modality-independent and camera-unrelated features for NIR-VIS object Re-ID. Our model consists of three key modules, including Confused Modality Generation (CMG), Modality-independent Information Distillation (MID), and Cameras Factor Disentanglement (CFD). Firstly, aiming at aligning image style between NIR and VIS data, the CMG utilizes a designed channel-interactive generator to generate confused modality images which preserves the structure information of original images. Besides, CMG is trained with confused adversarial learning which bridges the modality gap at the image level. Nevertheless, training the model with the confused modality images discards identity-related information such as color and contrast, which is not conducive to the extraction of distinctive features. To solve this problem, the MID is presented to distill out modality-independent information by feeding the original images to the model and reconstructing corresponding NIR and VIS modalities features to the confused modality images. Finally, due to the complexity of camera-related information in images, identity representation inevitably contains camera-related elements such as background and perspective information, which may interfere the matching process. To address this issue, the CFD is introduced to disentangle camera-related factors by the designed three-streams network and two factor decoupling losses, i.e., Camera-Camera Factor loss (CCF) and Identity-Camera Factor (ICF) loss. Comprehensive experiments are carried out on two cross-modality pedestrian Re-ID datasets and a cross-modality vehicle Re-ID dataset to demonstrate that the MCBD is effective in cross-modality object Re-ID task. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Robust Cross-Domain Pseudo-Labeling and Contrastive Learning for Unsupervised Domain Adaptation NIR-VIS Face RecognitionabstractNear-infrared and visible face recognition (NIR-VIS) is attracting increasing attention because of the need to achieve face recognition in low-light conditions to enable 24-hour secure retrieval. However, annotating identity labels for a large number of heterogeneous face images is time-consuming and expensive, which limits the application of the NIR-VIS face recognition system to larger scale real-world scenarios. In this paper, we attempt to achieve NIR-VIS face recognition in an unsupervised domain adaptation manner. To get rid of the reliance on manual annotations, we propose a novel Robust cross-domain Pseudo-labeling and Contrastive learning (RPC) network which consists of three key components, i.e., NIR cluster-based Pseudo labels Sharing (NPS), Domain-specific cluster Contrastive Learning (DCL) and Inter-domain cluster Contrastive Learning (ICL). Firstly, NPS is presented to generate pseudo labels by exploring robust NIR clusters and sharing reliable label knowledge with VIS domain. Secondly, DCL is designed to learn intra-domain compact yet discriminative representations. Finally, ICL dynamically combines and refines intrinsic identity relationships to guide the instance-level features to learn robust and domain-independent representations. Extensive experiments are conducted to verify an accuracy of over 99% in pseudo label assignment and the advanced performance of RPC network on four mainstream NIR-VIS datasets. Yiming Yang 0001, Weipeng Hu, Haiqi Lin, Haifeng Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Graph-Based Progressive Fusion Network for Multi-Modality Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is a critical task in intelligent transportation, aiming to match vehicle images of the same identity captured by non-overlapping cameras. However, it is difficult to achieve satisfactory results based on RGB images alone in darkness. Therefore, it is of great importance to consider multi-modality vehicle re-identification. Currently, the proposed works deal with different modality features through direct summation and fusion based on heat map, which however ignores the relationship between them. Meanwhile, there is a huge gap between the different modalities, which needs to be reduced. In this paper, to solve the above two problems, we propose a Graph-based Progressive Fusion Network (GPFNet) using a graph convolutional network to adaptively fuse multi-modality features in an end-to-end learning framework. GPFNet consists of a CNN feature extraction module (FEM), a GCN feature fusion module (FFM), and a loss function module (LFM). Firstly, in FEM, we employ a multi-stream network architecture to extract single-modality features and common-modality features and employ a random modality substitution module to extract mixed-modality features. Secondly, in FFM, we design an efficient graph structure to associate the features of different modalities and adopt a progressive two-stage strategy to fuse them. Finally, in LFM, we use GCN-aware multi-modality loss to constrain the features. For reducing modality differences and contributing better initial mixed-modality features to FFM, we propose random modality substitution as a data enhancement method for multi-modality datasets. Extensive experiments on multi-modality vehicle Re-ID datasets RGBN300 and RGBNT100 show that our model achieves state-of-the-art performance. Qiaolin He, Zefeng Lu, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | MART: Mask-Aware Reasoning Transformer for Vehicle Re-IdentificationabstractAs a significant topic in Intelligent Transportation Systems (ITS), vehicle Re-Identification (Re-ID) has attracted increasing research attention. However, the variation of shooting scenes and the similar appearance among the vehicles with the same type and color lead to large intra-class variances and small inter-class variances, respectively. To address the problems, we propose a novel Mask-Aware Reasoning Transformer (MART) to extract the background-unrelated global features and perspective-invariant local features. The MART contains three effective modules including Foreground Global Features Extraction (FGFE), Mask-guided Local Features Extraction (MLFE) and Cross-images Local Features Reasoning (CLFR). Firstly, due to the complexity of background information in images, identity representation inevitably contains background elements, which may impact identity matching. To address this issue, we propose the FGFE to extract background-independent global features by introducing the mask semantic information to the inputs of Vision Transformer (ViT). Secondly, to fill the gap that the previous local features extraction methods cannot be directly applied to ViT, the MLFE is presented to extract distinctive local features by recombining token features according to vehicle mask. Thirdly, when local components are invisible in the image due to the occlusion problem, the corresponding local information is absent from the image, leading to unreliable local features. To solve this problem, the CLFR is proposed to reason the occluded local features by exploiting the correlation between cross-image local features. We carry out comprehensive experiments to illustrate the effectiveness of the MART on two challenging datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Mask-Aware Pseudo Label Denoising for Unsupervised Vehicle Re-IdentificationabstractAs a significant part of Intelligent Transportation System (ITS), vehicle Re-Identification (Re-ID) aims to retrieve all target vehicle images captured from non-overlapping cameras. Though the Re-ID methods based on supervised learning have achieved rapid progress, they are still difficult to be applied in real scenarios due to the domain bias between the training set and real scenarios. Recently, methods based on unsupervised learning have been proposed to address the problem of domain bias by exploring techniques of pseudo-label generation. However, these methods suffer from pseudo-label noise. To solve this problem, we propose the Mask-Aware Pseudo Label Denoising framework (MAPLD) consisting of three key components, i.e., Mask-Aware Feature Extraction (MAFE), Adaptive Threshold Neighborhood Consistency (ATNC), and Compact Loss (CL). Firstly, the MAFE is proposed to improve the distinguishability of feature representation and widen the gap in feature space among vehicles with different IDs. Next, the ATNC is introduced to filter out pseudo-label noise of hard negative samples by comparing the image ID of the samples in their neighborhood set i.e., neighborhood consistency. Moreover, the threshold of neighborhood consistency is adaptively adjusted according to feature similarity ranking, which is robust to hyper-parameter variation. Finally, consisting of regression term and compact term, the CL is designed to drive the cluster more compact and alleviate the impact of outliers of hard positive samples. Extensive experiments on VeRi-776 and VeRi-Wild datasets demonstrate that MAPLD can generate reliable pseudo-labels and achieve superior performance in unsupervised target-only and unsupervised domain adaptation tasks. Zefeng Lu, Ronghao Lin, Qiaolin He, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Subgraph-Aware Few-Shot Inductive Link Prediction Via Meta-LearningabstractLink prediction for knowledge graphs aims to predict missing connections between entities. Prevailing methods are limited to a transductive setting and hard to process unseen entities. The recently proposed subgraph-based models provide alternatives to predict links from the subgraph structure surrounding a candidate triplet. However, these methods require abundant known facts of training triplets and perform poorly on relationships that only have a few triplets. In this paper, we propose Meta-iKG, a novel subgraph-based meta-learner for few-shot inductive relation reasoning. Meta-iKG utilizes local subgraphs to transfer subgraph-specific information and to rapidly learn transferable patterns via meta-gradients. In this way, we find the model can quickly adapt to few-shot relationships using only a handful of known facts with inductive settings. Moreover, we introduce a large-shot relation updating procedure to ensure that our model can generalize well to both few-shot and large-shot relations. We evaluate Meta-iKG on inductive benchmarks sampled from the NELL and Freebase, and the results show that Meta-iKG outperforms the currently state-of-the-art methods in both few-shot scenarios and standard inductive settings. Shuangjia Zheng, Sijie Mai, Haifeng Hu 0001, Yuedong Yang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Multimodal Information Bottleneck: Learning Minimal Sufficient Unimodal and Multimodal RepresentationsabstractLearning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, and the discriminative unimodal information may be ignored, which often interferes with accurate prediction and leads to a higher risk of overfitting. Moreover, unimodal representations also contain noisy information that negatively influences the learning of cross-modal dynamics. To this end, we introduce the multimodal information bottleneck (MIB), aiming to learn a powerful and sufficient multimodal representation that is free of redundancy and to filter out noisy information in unimodal representations. Specifically, inheriting from the general information bottleneck (IB), MIB aims to learn the minimal sufficient representation for a given task by maximizing the mutual information between the representation and the target and simultaneously constraining the mutual information between the representation and the input data. Different from general IB, our MIB regularizes both the multimodal and unimodal representations, which is a comprehensive and flexible framework that is compatible with any fusion methods. We develop three MIB variants, namely, early-fusion MIB, late-fusion MIB, and complete MIB, to focus on different perspectives of information constraints. Experimental results suggest that the proposed method reaches state-of-the-art performance on the tasks of multimodal sentiment analysis and multimodal emotion recognition across three widely used datasets. The codes are available at https://github.com/TmacMai/Multimodal-Information-Bottleneck. Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Attention-Augmented Memory Network for Image Multi-Label ClassificationabstractThe purpose of image multi-label classification is to predict all the object categories presented in an image. Some recent works exploit graph convolution network to capture the correlation between labels. Although promising results have been reported, these methods cannot learn salient object features in the images and ignore the correlation between channel feature maps. In addition, the current researches only learn the feature information within individual input image, but fail to mine the contextual information of various categories from the dataset to enhance the input feature representation. To address these issues, we propose an A ttention- A ugmented M emory N etwork ( AAMN ) model for the image multi-label classification task. Specifically, we first propose a novel categorical memory module to excavate the contextual information of various categories from the dataset to augment the current input feature. Secondly, we design a new channel-relation exploration module to capture the inter-channel relationship of features, so as to enhance the correlation between objects in the images. Thirdly, we develop a spatial-relation enhancement module to model second-order statistics of features and capture long-range dependencies between pixels in feature maps, so as to learn salient object features. Experimental results on standard benchmarks, including MS-COCO 2014, PASCAL VOC 2007, and VG-500, demonstrate the effectiveness and superiority of AAMN model, which outperforms current state-of-the-art methods. Wei Zhou 0042, Yanke Hou, Dihu Chen, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Aligning Image Semantics and Label Concepts for Image Multi-Label ClassificationabstractImage multi-label classification task is mainly to correctly predict multiple object categories in the images. To capture the correlation between labels, graph convolution network based methods have to manually count the label co-occurrence probability from training data to construct a pre-defined graph as the input of graph network, which is inflexible and may degrade model generalizability. Moreover, most of the current methods cannot effectively align the learned salient object features with the label concepts, so that the predicted results of model may not be consistent with the image content. Therefore, how to learn the salient semantic features of images and capture the correlation between labels, and then effectively align them is one of the key to improve the performance of image multi-label classification task. To this end, we propose a novel image multi-label classification framework which aims to align I mage S emantics with L abel C oncepts ( ISLC ). Specifically, we propose a residual encoder to learn salient object features in the images, and exploit the self-attention layer in aligned decoder to automatically capture the correlation between labels. Then, we leverage the cross-attention layers in aligned decoder to align image semantic features with label concepts, so as to make the labels predicted by model more consistent with image content. Finally, the output features of the last layer of residual encoder and aligned decoder are fused to obtain the final output feature for classification. The proposed ISLC model achieves good performance on various prevalent multi-label image datasets such as MS-COCO 2014, PASCAL VOC 2007, VG-500, and NUS-WIDE with 87.2%, 96.9%, 39.4%, and 64.2%, respectively. Wei Zhou 0042, Zhiwu Xia, Peng Dou, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Multiple Temporal Pooling Mechanisms for Weakly Supervised Temporal Action LocalizationabstractRecent action localization works learn in a weakly supervised manner to avoid the expensive cost of human labeling. Those works are mostly based on the Multiple Instance Learning framework, where temporal pooling is an indispensable part that usually relies on the guidance of snippet-level Class Activation Sequences (CAS) . However, we observe that previous works only leverage a simple convolutional neural network for the generation of CAS, which ignores the weak discriminative foreground action segments and the background ones, and meanwhile, the relationship between different actions has not been considered. To solve this problem, we propose multiple temporal pooling mechanisms (MTP) for a more sufficient information utilization. Specifically, with the design of the Foreground Variance Branch, Dual Foreground Attention Branch and Hybrid Attention Fine-tuning Branch, MTP can leverage more effective information from different aspects and generate different CASs to guide the learning of temporal pooling. Moreover, different loss functions are designed for a better optimization of individual branches, aiming to effectively distinguish the action from the background. Our method shows excellent results on the THUMOS14 and ActivityNet1.2 datasets. Peng Dou, Zhuoqun Wang, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multimodal Graph for Unaligned Multimodal Sequence Analysis via Graph Convolution and Graph PoolingabstractMultimodal sequence analysis aims to draw inferences from visual, language, and acoustic sequences. A majority of existing works focus on the aligned fusion of three modalities to explore inter-modal interactions, which is impractical in real-world scenarios. To overcome this issue, we seek to focus on analyzing unaligned sequences, which is still relatively underexplored and also more challenging. We propose Multimodal Graph, whose novelty mainly lies in transforming the sequential learning problem into graph learning problem. The graph-based structure enables parallel computation in time dimension (as opposed to recurrent neural network) and can effectively learn longer intra- and inter-modal temporal dependency in unaligned sequences. First, we propose multiple ways to construct the adjacency matrix for sequence to perform sequence to graph transformation. To learn intra-modal dynamics, a graph convolution network is employed for each modality based on the defined adjacency matrix. To learn inter-modal dynamics, given that the unimodal sequences are unaligned, the commonly considered word-level fusion does not pertain. To this end, we innovatively devise graph pooling algorithms to automatically explore the associations between various time slices from different modalities and learn high-level graph representation hierarchically. Multimodal Graph outperforms state-of-the-art models on three datasets under the same experimental setting. Sijie Mai, Songlong Xing, Jiaxuan He, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Double Attention Based on Graph Attention Network for Image Multi-Label ClassificationabstractThe task of image multi-label classification is to accurately recognize multiple objects in an input image. Most of the recent works need to leverage the label co-occurrence matrix counted from training data to construct the graph structure, which are inflexible and may degrade model generalizability. In addition, these methods fail to capture the semantic correlation between the channel feature maps to further improve model performance. To address these issues, we propose DA-GAT (a D ouble A ttention framework based on the G raph A ttention ne T work) to effectively learn the correlation between labels from training data. First, we devise a new channel attention mechanism to enhance the semantic correlation between channel feature maps, so as to implicitly capture the correlation between labels. Second, we propose a new label attention mechanism to avoid the adverse impact of a manually constructed label co-occurrence matrix. It only needs to leverage the label embedding as the input of network, then automatically constructs the label relation matrix to explicitly establish the correlation between labels. Finally, we effectively fuse the output of these two attention mechanisms to further improve model performance. Extensive experiments are conducted on three public multi-label classification benchmarks. Our DA-GAT model achieves mean average precision of 87.1%, 96.6%, and 64.3% on MS-COCO 2014, PASCAL VOC 2007, and NUS-WIDE, respectively, and obviously outperforms other existing state-of-the-art methods. In addition, visual analysis experiments demonstrate that each attention mechanism can capture the correlation between labels well and significantly promote the model performance. Wei Zhou 0042, Zhiwu Xia, Peng Dou, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Curriculum Learning Meets Weakly Supervised Multimodal Correlation LearningabstractIn the field of multimodal sentiment analysis (MSA), a few studies have leveraged the inherent modality correlation information stored in samples for self-supervised learning.However, they feed the training pairs in a random order without consideration of difficulty.Without human annotation, the generated training pairs of self-supervised learning often contain noise.If noisy or hard pairs are used for training at the easy stage, the model might be stuck in bad local optimum.In this paper, we inject curriculum learning into weakly supervised modality correlation learning.The weakly supervised correlation learning leverages the label information to generate scores for negative pairs to learn a more discriminative embedding space, where negative pairs are defined as two unimodal embeddings from different samples.To assist the correlation learning, we feed the training pairs to the model according to difficulty by the proposed curriculum learning, which consists of elaborately designed scoring and feeding functions.The scoring function computes the difficulty of pairs using pre-trained and current correlation predictors, where the pairs with large losses are defined as hard pairs.Notably, the hardest pairs are discarded in our algorithm, which are assumed as noisy pairs.Moreover, the feeding function takes the difference of correlation losses as feedback to determine the feeding actions ('stay', 'step back', or 'step forward').The proposed method reaches state-of-the-art performance on MSA. Sijie Mai, Haifeng Hu 0001 |
EMNLP | 3 |
| 2022 | Contextual relation embedding and interpretable triplet capsule for inductive relation prediction
Sijie Mai, Haifeng Hu 0001 |
Neurocomputing | 3 |
| 2022 | Dynamic graph dropout for subgraph-based relation prediction
Sijie Mai, Shuangjia Zheng, Yuedong Yang, Haifeng Hu 0001 |
Knowl. Based Syst. | 6 |
| 2022 | Attention-Guided Multi-Clue Mining Network for Person Re-identification
Yangbin Yu, Shengrong Yang, Haifeng Hu 0001, Dihu Chen |
Neural Process. Lett. | 3 |
| 2022 | AVPL: Augmented visual perception learning for person Re-identification and beyond
Yewen Huang, Sicheng Lian, Haifeng Hu 0001 |
Pattern Recognit. | 3 |
| 2022 | Language Reinforced Superposition Multimodal Fusion for Sentiment AnalysisabstractNowadays, BERT has been effective in context extraction and widely utilized in multimodal fusion. However, many multimodal learning systems with BERT consider different modalities equally important, which ignores the huge capability difference between the pre-trained models and others. Moreover, equal multimodal fusion might introduce noise and limit the performance of these systems. To this end, we propose a superposition multimodal fusion framework to strengthen the importance of language modality with nonverbal information transmission. With the superposition strategy, language information will fuse with nonverbal information at different stages, which allows the model to fully learn language modality information. Besides, to ensure that language information is the main content of the fusion, we propose a comparison loss to help the reinforced cross-modal attention module better transfer nonverbal information to language features. Extensive experiments are conducted to compare with baselines and the results demonstrate that our method achieves the superior performance. Jiaxuan He, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | MF-BERT: Multimodal Fusion in Pre-Trained BERT for Sentiment AnalysisabstractMultimodal sentiment analysis mainly concentrates on language, acoustic and visual information. Previous work based on BERT utilizes only text (language) representation to fine-tune BERT, while ignoring the importance of nonverbal information. Due to the fact that features extracted from a single modality may contain uncertainty, it is challenging for BERT to perform well in real-world applications. In this paper, we propose a multimodal fusion BERT that can explore the time-dependent interactions among different modalities. Additionally, prior BERT-based methods tend to train the models with only one optimizer to update the parameters. However, we argue that BERT has been pre-trained with a lot of corpora so it needs to be fine-tuned slightly. Therefore, an internal updating mechanism is introduced to avoid the overfitting of the model in the training process. We set two optimizers for multimodal fusion BERT and other components of the model with different learning rates, which enables the model to attain optimal parameters. The results of experiments on public datasets demonstrate that our model is superior to the baselines and achieves the state-of-the-art. Jiaxuan He, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Two-Branch Asymmetric Model With Alternately Clustering for Unsupervised Person Re-IdentificationabstractIn the field of unsupervised person Re-identification (Re-ID), mainstream methods adopt cluster algorithm to generate pseudo labels for training. Despite the effectiveness, the cluster algorithm generates noisy labels, which are retained in further model updating and hinder higher performance. To solve this problem, we propose a Two-branch Asymmetric Model with Alternately Clustering. Specifically, the designed Alternately Clustering (AC) strategy leverages a two-branch model to cluster different pseudo labels for each branch, which prevents the continuous existence of identical noisy labels. To establish a mapping between the two label sets, pseudo label mapping constraint (PLMC) module is devised, which helps retain reliable pseudo labels. Our method improves the quality of generated pseudo labels by keeping noisy labels changing and retaining the reliable ones. Experimental results demonstrate that our proposed method outperforms the state-of-the-art methods. Yangbin Yu, Haifeng Hu 0001, Dihu Chen |
IEEE Signal Process. Lett. | 3 |
| 2022 | Multi-Fusion Residual Memory Network for Multimodal Human Sentiment ComprehensionabstractMultimodal human sentiment comprehension refers to recognizing human affection from multiple modalities. There exist two key issues for this problem. First, it is difficult to explore time-dependent interactions between modalities and focus on the important time steps. Second, processing the long fused sequence of utterances is susceptible to the forgetting problem due to the long-term temporal dependency. In this article, we introduce a hierarchical learning architecture to classify utterance-level sentiment. To address the first issue, we perform time-step level fusion to generate fused features for each time step, which explicitly models time-restricted interactions by incorporating information across modalities at the same time step. Furthermore, based on the assumption that acoustic features directly reflect emotional intensity, we pioneer emotion intensity attention to focus on the time steps where emotion changes or intense affections take place. To handle the second issue, we propose Residual Memory Network (RMN) to process the fused sequence. RMN utilizes some techniques such as directly passing the previous state into the next time step, which helps to retain the information from many time steps ago. We show that our method achieves state-of-the-art performance on multiple datasets. Results also suggest that RMN yields competitive performance on sequence modeling tasks. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Adapted Dynamic Memory Network for Emotion Recognition in ConversationabstractIn this article, we address Emotion Recognition in Conversation (ERC) where conversational data are presented in a multimodal setting. Psychological evidence shows that self and inter-speaker influence are two central factors to emotion dynamics in conversation. State-of-the-art models do not effectively synthesise these two factors. Therefore, we propose an Adapted Dynamic Memory Network (A-DMN) where self and inter-speaker influences are modelled individually and further synthesised oriented towards the current utterance. Specifically, we model the dependency of the constituent utterances in a dialogue video using a global RNN to capture inter-speaker influence. Likewise, each speaker is assigned an RNN to capture their self influence. Afterwards, an Episodic Memory Module is devised to extract contexts for self and inter-speaker influence and synthesise them to update the memory. This process repeats itself for multiple passes until a refined representation is obtained and used for final prediction. Additionally, we explore cross-modal fusion in the context of multimodal ERC, and propose a convolution-based method which proves effective in extracting local interactions and computationally efficient. Extensive experiments demonstrate that A-DMN outperforms the state-of-the-art models on benchmark datasets. Songlong Xing, Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Interpretable Multimodal Capsule FusionabstractWith the development of social networking platform, multimodal sentiment analysis has become increasingly prominent. Existing models focus on capturing intramodal and intermodal interactions to produce effective modality representations. However, they overlook the study of interpretability which reveals how modalities interact with each other and which modality contributes most to the final prediction. In this paper, we propose an interpretable model called Interpretable Multimodal Capsule Fusion (IMCF) which integrates routing mechanism of Capsule Network (CapsNet) and Long Short-Term Memory (LSTM) to produce refined modality representations and provide interpretation. By constructing features of different modalities into input sequence, we are able to obtain highly expressive representation of intermodal dynamics due to the strong ability of LSTM to produce representation of sequence. As routing mechanism is applied during modality fusion and prediction stages, the value of routing coefficient can reveal the contributions of different modalities or dynamics, which provides interpretation. Meanwhile, routing mechanism can iteratively adjust the information flows of different modalities, which makes the process of modality fusion more reasonable. The experimental results show that our model achieves competitive performance on two benchmark datasets with effective modality fusion by LSTM and interpretation provided by routing mechanism. Sijie Mai, Haifeng Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Orthogonal Modality Disentanglement and Representation Alignment Network for NIR-VIS Face RecognitionabstractNear-infrared and visual (NIR-VIS) face matching, as the most typical task in Heterogeneous Face Recognition (HFR), has attracted increasing attention in recent years. However, due to the large within-class discrepancies, including domain differences and residual discrepancies (i.e., lighting, expressions, occlusion, blurry, pose, etc), this is still a difficult task. Conventional NIR-VIS FR methods only focus on reducing the modality gap between cross-domain images, while neglecting to eliminate the residual variations. To better solve the above problems, this paper proposes a novel Orthogonal Modality Disentanglement and Representation Alignment (OMDRA) approach, which consists of three key components, including Modality-Invariant (MI) loss, Orthogonal Modality Disentanglement (OMD) and Deep Representation Alignment (DRA). Firstly, the MI loss is designed to learn modality-invariant and identity-discriminative representation, by increasing between-class separability and within-class compactness between NIR and VIS heterogeneous data. Secondly, the high-level Hybrid Facial Feature (HFF) layer of the backbone network is projected into two subspaces: the modality-related and identity-related subspaces. The OMD is designed to decouple modal information via an adversarial process, and we further impose Orthogonal Representation Decorrelation (ORD) to the OMD to decrease the correlation between identity representations and domain representations, as well as enhancing their representation capabilities. Finally, the DRA aims to eliminate the residual variations by performing a high-level representation alignment between non-neutral face and neutral face, which can effectively guides the network to learn discriminative and residual-invariant face representation. The joint scheme enables the disentanglement of modality variations, elimination of residual discrepancies, and the purification of identity information. Extensive experiments on challenging cross-domain databases indicate that our OMDRA method is superior to the state-of-the-art methods. Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Adversarial Decoupling and Modality-Invariant Representation Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (RGB-IR ReID) has now attracted increasing attention due to its surveillance applications under low-light environments. However, the large intra-class variations between different domains are still a challenging issue in the field of computer vision. To address the above issue, we propose a novel adversarial Decoupling and Modality-invariant Representation learning (DMiR) method to explore potential spectrum-invariant yet identity-discriminative representations for cross-modality pedestrians. Our model consists of three key components, including Domain-related Representation Disentanglement (DrRD), Modality-invariant Discriminative Representation (MiDR) and Representation Orthogonal Decorrelation (ROD). First, two subnets named Identity-Net and Domain-Net are designed to extract identity-related features and domain-related features, respectively. Given this two-stream structure, the DrRD is introduced to achieve adversarial decoupling against domain-specific features via a min-max disentanglement process. Specifically, the classification objective function on Domain-Net is minimized to extract spectrum-specific information while maximizing it to reduce domain-specific information. Second, in Identity-Net, we introduce MiDR to enhance intra-class compactness and reduce domain variations by exploring positive and negative pair variations, semantic-wise differences, and pair-wise semantic variations. Finally, the correlation between the two decomposed features, i.e., identity-related features and domain-related features, may lead to the introduction of modal information in identity representations, and vice versa. Therefore, we present the ROD constraint to make the two decomposed features unrelated to each other, which can more effectively separate the two-component features and enhance feature representations. Practically, we construct ROD at the feature-level and parameter-level, and finally select feature-level ROD as the decorrelation strategy because of its superior decorrelation performance. The whole scheme leads to disentangling spectrum-dependent information, as well as purifying identity information. Extensive experiments are carried out on two mainstream RGB-IR ReID datasets, and the results demonstrate the effectiveness of our method. Weipeng Hu, Bohong Liu, Haitang Zeng, Yanke Hou, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Dual Face Alignment Learning Network for NIR-VIS Face RecognitionabstractAs the most important topic in Heterogeneous Face Recognition (HFR), Near-InfraRed and VISual (NIR-VIS) face recognition has attracted increasing research attentions owing to its potential application in the field of criminal detective cases and multimedia information retrieval. However, due to its dramatic intra-class variations, including modality, pose, occlusion, blurry, lighting, distance, expression, etc, it is very challenging to retain inherent identity information. To address the above issue, we propose a novel Dual Face Alignment Learning (DFAL) algorithm to explore the potential domain-invariant neutral face representations of the cross-modal images. Our model contains three effective components including Feature-level Face Alignment (FFA), Image-level Face Alignment (IFA) and Cross-domain compact Representation (CdR). Firstly, Teacher-Encoder CNNs (TeEn-CNNs) and Student-Encoder CNNs (StEn-CNNs) are designed to encode features for VIS neutral face images and non-neutral face images, and the FFA is introduced to learn neutral face representations by performing feature-level alignment between non-neutral face and VIS neutral face. Secondly, Student-Decoder CNNs (StDe-CNNs) is developed to decode features to restore face images, and the IFA is designed to reconstruct neutral face image by imposing image-level alignment. Notably, the FFA acts as the primary target to learn VIS neutral face representations for cross-view data, while the IFA plays a role in the icing on the cake, i.e., further disentangling domain and residual information through the synthesis process. Finally, the CdR dispels modality features and distills identity features by mining inter-class information, inter-domain information and inter-semantic relationship. The joint scheme enables the elimination of intra-class variations and the purification of identity information. We carry out comprehensive experiments to illustrate the effectiveness of the DFAL approach on three challenging NIR-VIS databases. Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Double-Stream Position Learning Transformer Network for Image CaptioningabstractImage captioning has made significant achievement through developing feature extractor and model architecture. Recently, the image region features extracted by object detector prevail in most existing models. However, region features are criticized for the lacking of background and full contextual information. This problem can be remedied by providing some complementary visual information from patch features. In this paper, we propose a Double-Stream Position Learning Transformer Network (DSPLTN) which exploits the advantages of region features and patch features. Specifically, the region-stream encoder utilizes a Transformer encoder with Relative Position Learning (RPL) module to enhance the representations of region features through modeling the relationships between regions and positions respectively. As for the patch-stream encoder, we introduce convolutional neural network into the vanilla Transformer encoder and propose a novel Convolutional Position Learning (CPL) module to encode the position relationships between patches. CPL improves the ability of relationship modeling by combining the position and visual content of patches. Incorporating CPL into the Transformer encoder can synthesize the benefits of convolution in local relation modeling and self-attention in global feature fusion, thereby compensating for the information loss caused by the flattening operation of 2D feature maps to 1D patches. Furthermore, an Adaptive Fusion Attention (AFA) mechanism is proposed to balance the contribution of enhanced region and patch features. Extensive experiments on MSCOCO demonstrate the effectiveness of the double-stream encoder and CPL, and show the superior performance of DSPLTN. Weitao Jiang, Wei Zhou 0042, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Domain Adversarial Disentanglement Network With Cross-Domain Synthesis for Generalized Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) plays an increasingly important role in face recognition systems for preventing malicious attacks. Existing FAS methods usually show poor generalization performance due to the large difference of domain information (such as collection environment, collection equipment, etc.) between the training set and the testing set. Therefore, we consider training a FAS network that does not pay attention to domain information, in this way to extract relevant liveness feature for the task of FAS. Learned from adversarial learning and disentanglement learning, we design the Domain Adversarial Disentanglement Network with Cross-Domain Synthesis (DADN-CDS) for Face Anti-Spoofing, which achieves disentanglement between domain-irrelevant liveness feature and domain-related feature, so as to minimize the effect of domain information in the needed representation that is used for inference. Specifically, DADN-CDS proposes a dual-branch architecture that can explicitly model and separate out the domain information. To achieve targeted optimization for each component, a novel Task-oriented Three-step Update Strategy (TTUS) is designed to explore a better model update method. Irrelevant loss and 2N-Pair Cross-Domain Loss in TTUS further ensure the disentanglement and task-oriented optimization. Moreover, an Attention-based Cross Synthesis Module (ACSM) is elaborately devised to obtain higher-quality synthetic feature, which performs attention-based feature fusion in a channel-wise way. The design of ACSM can help verify and implicitly facilitate the disentanglement process. Extensive experiments demonstrate that our method achieves the state-of-the-art performance on public datasets and results also suggest the generalization ability of our proposed method. Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | 7-Instant Discrete-Time Synthesis Model Solving Future Different-Level Linear Matrix System via Equivalency of Zeroing Neural NetworkabstractDiffering from the common linear matrix equation, the future different-level linear matrix system is considered, which is much more interesting and challenging. Because of its complicated structure and future-computation characteristic, traditional methods for static and same-level systems may not be effective on this occasion. For solving this difficult future different-level linear matrix system, the continuous different-level linear matrix system is first considered. On the basis of the zeroing neural network (ZNN), the physical mathematical equivalency is thus proposed, which is called ZNN equivalency (ZE), and it is compared with the traditional concept of mathematical equivalence. Then, on the basis of ZE, the continuous-time synthesis (CTS) model is further developed. To satisfy the future-computation requirement of the future different-level linear matrix system, the 7-instant discrete-time synthesis (DTS) model is further attained by utilizing the high-precision 7-instant Zhang et al. discretization (ZeaD) formula. For a comparison, three different DTS models using three conventional ZeaD formulas are also presented. Meanwhile, the efficacy of the 7-instant DTS model is testified by the theoretical analyses. Finally, experimental results verify the brilliant performance of the 7-instant DTS model in solving the future different-level linear matrix system. Min Yang 0010, Yunong Zhang, Ning Tan 0003, Mingzhi Mao, Haifeng Hu 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | Domain-Private Factor Detachment Network for NIR-VIS Face RecognitionabstractNear-InfraRed and VISual (NIR-VIS) face matching, as one of the most representative tasks in Heterogeneous Face Recognition (HFR), aims at retrieving a face image across different domains. With the development of deep learning and the growing demand for intelligent surveillance, it has aroused more and more research attention in the computer vision community. However, due to the dramatic modality gap between NIR and VIS images, the task of NIR-VIS face recognition becomes practically very challenging. In this paper, we propose a novel Domain-private Factor Detachment (DFD) network to disentangle domain-dependent factors and achieve identity information distillation. Our approach consists of three key components, including Domain-identity Representation Learning (DiRL), Cross-domain Factor Detachment (CdFD) and Cross-domain Aggregation Learning (CAL). Firstly, the proposed DiRL aims to achieve domain-specific information distillation and learn identity-related representations. Specifically, three sub-networks, i.e., NIR sub-Network (NIR-Net), VIS sub-Network (VIS-Net) and IDentity-dependent sub-Network (ID-Net) are designed to learn NIR facial representations, VIS facial representations and identity-dependent representations, respectively, and they can promote each other to facilitate the learning of identity-discriminative representations. Secondly, considering that the entangled modal components in face representations negatively affect the subsequent matching process, to reduce modality-related components, we model the cross-modal face matching problem into three parts, comprising Identity Variation (IV), Inter-Spectrum Variation (ISV) and Identity-Domain Variation (IDV). The CdFD is presented to eliminate ISV components and IDV components by introducing inter-spectrum invariant constraint and identity-domain invariant constraint, so that cross-modal face recognition can be performed under pure identity information differences without modal interference. Finally, the CAL is developed to learn modality-invariant yet discriminative representations by exploring within-class aggregation, negative pair separability and cross-domain positive pair compactness. Experimental results on multiple challenging databases demonstrate the effectiveness of the DFD approach. Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Concise Discrete ZNN Controllers for End-Effector Tracking and Obstacle Avoidance of Redundant ManipulatorsabstractObstacle avoidance is usually an additional task for a redundant manipulator when the end-effector tracking task is performed, which guarantees the safety of the redundant manipulator. By formulating and combining the end-effector tracking task and the obstacle avoidance task using zeroing neural network (ZNN), in this article, a concise continuous ZNN (CZNN) controller is first proposed. To develop discrete controllers for practical control, a second-order discrete formula and a third-order discrete formula are introduced. Therefore, by utilizing two discrete formulas to discretize the CZNN controller, two concise discrete ZNN (DZNN) controllers are proposed. Detailed theoretical analyses guarantee the effectiveness of task formulations, CZNN controller, and DZNN controllers. In addition, some comparisons with existing studies are presented in details. Furthermore, three groups of simulative experiments on the basis of UR5 manipulator and two groups of physical experiments on the basis of Kinova manipulator are conducted to illustrate the effectiveness, superiority, and practicability of two DZNN controllers. Min Yang 0010, Yunong Zhang, Ning Tan 0003, Haifeng Hu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Meta PID Attention Network for Flexible and Efficient Real-World Noisy Image DenoisingabstractRecent deep convolutional neural networks for real-world noisy image denoising have shown a huge boost in performance by training a well-engineered network over external image pairs. However, most of these methods are generally trained with supervision. Once the testing data is no longer compatible with the training conditions, they can exhibit poor generalization and easily result in severe overfitting or degrading performances. To tackle this barrier, we propose a novel denoising algorithm, dubbed as Meta PID Attention Network (MPA-Net). Our MPA-Net is built based upon stacking Meta PID Attention Modules (MPAMs). In each MPAM, we utilize a second-order attention module (SAM) to exploit the channel-wise feature correlations with second-order statistics, which are then adaptively updated via a proportional-integral-derivative (PID) guided meta-learning framework. This learning framework exerts the unique property of the PID controller and meta-learning scheme to dynamically generate filter weights for beneficial update of the extracted features within a feedback control system. Moreover, the dynamic nature of the framework enables the generated weights to be flexibly tweaked according to the input at test time. Thus, MPAM not only achieves discriminative feature learning, but also facilitates a robust generalization ability on distinct noises for real images. Extensive experiments on ten datasets are conducted to inspect the effectiveness of the proposed MPA-Net quantitatively and qualitatively, which demonstrates both its superior denoising performance and promising generalization ability that goes beyond those of the state-of-the-art denoising methods. Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001, Haifeng Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Identity-Unrelated Information Decoupling Model for Vehicle Re-IdentificationabstractAs an indispensable part of intelligent transportation system (ITS), vehicle re-identification (Re-ID) aims to retrieve all target vehicle images captured from non-overlapping cameras. However, this task remains very challenging due to the variation of camera perspective and the similar appearance among the vehicles with the same type and color, i.e., large intra-class variances and small inter-class variances. Previous methods have made a great progress on vehicle Re-ID by leveraging local details and aligning local features, this issue is still far from being solved. In this work, we propose to decouple identity-unrelated information from vehicle representation, tackling the problems of camera perspective variation and vehicle appearance similarity. The keypoint of this method is to learn a distinguishable feature embedding that is independent of identity-unrelated information. Specially, the novel Identity-Unrelated Information Decoupling (IUID) paradigm is designed to learn invariant features of the vehicle with the same ID in different scenes. In our approach, identity-unrelated information can be divided into two kinds of information, i.e., camera perspective information and background information. For the former, through a feature-level camera generative adversarial module, we can decouple camera perspective information from the feature embedding after extracting invariant features across different cameras perspective. For the latter, we propose a vehicle-mask transformer to enhance the attention of the model on local details while reducing the impact of background information. Extensive experiments on two public datasets demonstrate the superiority of IUID over the current state-of-the-arts methods. Zefeng Lu, Ronghao Lin, Xulei Lou, Lifeng Zheng, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | A Unimodal Representation Learning and Recurrent Decomposition Fusion Structure for Utterance-Level Multimodal Embedding LearningabstractLearning a unified embedding for utterance-level video attracts significant attention recently due to the rapid development of social media and its broad applications. An utterance normally contains not only spoken language but also the nonverbal behaviors such as facial expressions and vocal patterns. Instead of directly learning utterance embedding based on low-level features, we firstly explore high-level representation for each modality separately via an unimodal representation learning gyroscope structure. In this way, the learnt unimodal representations are more representative and contain more abstract semantic information. In the gyroscope structure, we introduce multi-scale kernel learning, ‘channel expansion’ and ‘channel fusion’ operations to explore high-level features both spatially and channelwise. Another insight of our method lies in that we fuse representations of all modalities to obtain a unified embedding by interpreting fusion procedure as the flow of inter-modality information between various modalities, which is more specialized in terms of the information to be fused and the fusion process. Specifically, considering that each modality carries modality-specific and cross-modality interactions, we innovate to decompose unimodal representations into intra- and inter-modality dynamics using gating mechanism, and further fuse the inter-modality dynamics by passing them from previous modalities to the following one using a recurrent neural fusion architecture. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple benchmark datasets. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Multim. | 2 |
| 2022 | Cascaded Structure-Learning Network with Using Adversarial Training for Robust Facial Landmark DetectionabstractRecently, great progress has been achieved on facial landmark detection based on convolutional neural network, while it is still challenging due to partial occlusion and extreme head pose. In this paper, we propose a Cascaded Structure-Learning Network (CSLN) with using adversarial training to improve the performance of 2D facial landmark detection by taking the structure of facial landmarks into account. In the first stage, we improve the original stacked hourglass network, which applies a multi-branch module to capture different scales of features, a progressive convolution structure to compensate for the missing structural features in hourglass networks, and a pyramid inception structure to expand the receptive field. Specially, by introducing a discriminator, we use the adversarial training strategy to urge the improved hourglass network for generating more accurate heatmaps. The second stage, which is based on attention mechanism, optimizes the spatial correlations between different facial landmarks by reusing the structural features. Moreover, we propose a novel region loss, which can adaptively allocate proper weights to different regions. In this way, the network can focus more on those occluded landmarks. The experimental results on several datasets, i.e. 300W, COFW, and AFLW, show that our proposed method achieves superior performance compared with the state-of-the-art methods. Shenming Feng, Xingzhong Nong, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | 6-Step Discrete ZNN Model for Repetitive Motion Control of Redundant ManipulatorabstractIn this article, the repetitive motion control of redundant manipulators is investigated. First, a repetitive motion control scheme is presented, and a continuous zeroing neural network (CZNN) model is obtained for solving the scheme. Meanwhile, the development of a discrete zeroing neural network (DZNN) model is desired for convenient computational processing. Based on this, this article proposes a 6-step discretization formula, which has high precision. By using the 6-step discretization formula and the 4-step backward difference formula, a 6-step DZNN (6SDZNN) model is further proposed to handle the repetitive motion control scheme. Theoretical analyses verify the efficacy of the 6SDZNN model. Additionally, some discrete forms of conventional models are developed for comparison. Computer simulations on the basis of the 4-link redundant manipulator are carried out, verifying the theoretical analyses and showing the efficacy of the 6SDZNN model. Finally, physical experiments on the basis of the Kinova Jaco2manipulator substantiate the practicability of the 6SDZNN model. Min Yang 0010, Yunong Zhang, Zhijun Zhang 0003, Haifeng Hu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Communicative Message Passing for Inductive Relation ReasoningabstractRelation prediction for knowledge graphs aims at predicting missing relationships between entities. Despite the importance of inductive relation prediction, most previous works are limited to a transductive setting and cannot process previously unseen entities. The recent proposed subgraph-based relation reasoning models provided alternatives to predict links from the subgraph structure surrounding a candidate triplet inductively. However, we observe that these methods often neglect the directed nature of the extracted subgraph and weaken the role of relation information in the subgraph modeling. As a result, they fail to effectively handle the asymmetric/anti-symmetric triplets and produce insufficient embeddings for the target triplets. To this end, we introduce a Communicative Message Passing neural network for Inductive reLation rEasoning, CoMPILE, that reasons over local directed subgraph structures and has a vigorous inductive bias to process entity-independent semantic relations. In contrast to existing models, CoMPILE strengthens the message interactions between edges and entitles through a communicative kernel and enables a sufficient flow of relation information. Moreover, we demonstrate that CoMPILE can naturally handle asymmetric/anti-symmetric relations without the need for explosively increasing the number of model parameters by extracting the directed enclosing subgraphs. Extensive experiments show substantial performance gains in comparison to state-of-the-art methods on commonly used benchmark datasets with variant inductive settings. Sijie Mai, Shuangjia Zheng, Yuedong Yang, Haifeng Hu 0001 |
AAAI | 4 |
| 2021 | Graph Capsule Aggregation for Unaligned Multimodal SequencesabstractHumans express their opinions and emotions through multiple modalities which mainly consist of textual, acoustic and visual modalities. Prior works on multimodal sentiment analysis mostly apply Recurrent Neural Network (RNN) to model aligned multimodal sequences. However, it is unpractical to align multimodal sequences due to different sample rates for different modalities. Moreover, RNN is prone to the issues of gradient vanishing or exploding and it has limited capacity of learning long-range dependency which is the major obstacle to model unaligned multimodal sequences. In this paper, we introduce Graph Capsule Aggregation (GraphCAGE) to model unaligned multimodal sequences with graph-based neural model and Capsule Network. By converting sequence data into graph, the previously mentioned problems of RNN are avoided. In addition, the aggregation capability of Capsule Network and the graph-based structure enable our model to be interpretable and better solve the problem of long-range dependency. Experimental results suggest that GraphCAGE achieves state-of-the-art performance on two benchmark datasets with representations refined by Capsule Network and interpretation provided. Sijie Mai, Haifeng Hu 0001 |
ICMI | 3 |
| 2021 | Feature Matching Network for Weakly-Supervised Temporal Action Localization
Peng Dou, Wei Zhou 0042, Zhongke Liao, Haifeng Hu 0001 |
PRCV (4) | 4 |
| 2021 | LPF: A Language-Prior Feedback Objective Function for De-biased Visual Question AnsweringabstractMost existing Visual Question Answering (VQA) systems tend to overly rely on the language bias and hence fail to reason from the visual clue. To address this issue, we propose a novel Language-Prior Feedback (LPF) objective function, to re-balance the proportion of each answer's loss value in the total VQA loss. The LPF firstly calculates a modulating factor to determine the language bias using a question-only branch. Then, the LPF assigns a self-adaptive weight to each training sample in the training process. With this reweighting mechanism, the LPF ensures that the total VQA loss can be reshaped to a more balanced form. By this means, the samples that require certain visual information to predict will be efficiently used during training. Our method is simple to implement, model-agnostic, and end-to-end trainable. We conduct extensive experiments and the results show that the LPF (1) brings a significant improvement over various VQA models, (2) achieves competitive performance on the bias-sensitive VQA-CP v2 benchmark. Zujie Liang, Haifeng Hu 0001, Jiaying Zhu |
SIGIR | 2 |
| 2021 | Image translation with dual-directional generative adversarial networksabstractAbstract Image‐to‐image translation is a class of vision and graphics problems where the goal is to learn the mapping between input images and output images. However, due to the unstable training and limited training samples, many existing GAN‐based works have difficulty in producing photo‐realistic images. Herein, dual‐directional generative adversarial networks are proposed, which consist of four adversarial networks, to produce images of high perceptual quality. In this framework, self‐reconstruction strategy is used to construct auxiliary sub‐networks, which impose more effective constraints on encoder‐generator pairs. Using this idea, this model can increase the use ratio of paired data conditioned on the same dataset and obtain well‐trained encoder‐generator pairs with the help of the proposed cross‐network skip connections. Moreover, the proposed framework not only produces realistic images but also addresses the problem where condition GAN produces sharp images containing many small, hallucinated objects. Training on multiple supervised datasets, convincing evidences are shown to prove that this model can achieve compelling results by latently learning a common feature representation. Qualitative and quantitative comparisons against other methods, demonstrate the effectiveness and superiority of the method. Congcong Ruan, Liuchun Yuan, Haifeng Hu 0001, Dihu Chen |
IET Comput. Vis. | 3 |
| 2021 | Posture coordination control of two-manipulator system using projection neural network
Min Yang 0010, Yunong Zhang, Haifeng Hu 0001 |
Neurocomputing | 3 |
| 2021 | Boundary Adjusted Network Based on Cosine Similarity for Temporal Action Proposal Generation
Jingye Zheng, Dihu Chen, Haifeng Hu 0001 |
Neural Process. Lett. | 3 |
| 2021 | A Unimodal Reinforced Transformer With Time Squeeze Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis refers to inferring sentiment from language, acoustic, and visual sequences. Previous studies focus on analyzing aligned sequences, while the unaligned sequential analysis is more practical in real-world scenarios. Due to the long-time dependency hidden in the multimodal unaligned sequence and time alignment information is not provided, exploring the time-dependent interactions within unaligned sequences is more challenging. To this end, we introduce the time squeeze fusion to automatically explore the time-dependent interactions by modeling the unimodal and multimodal sequences from the perspective of compressing the time dimension. Moreover, prior methods tend to fuse unimodal features into a multimodal embedding, based on which sentiment is inferred. However, we argue that the unimodal information may be lost or the generated multimodal embedding may be redundant. Addressing this issue, we propose a unimodal reinforced Transformer to progressively attend and distill unimodal information from the multimodal embedding, which enables the multimodal embedding to highlight the discriminative unimodal information. Extensive experiments suggest that our model reaches state-of-the-art performance in terms of accuracy and F1 score on MOSEI dataset. Jiaxuan He, Sijie Mai, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Part-Relation-Aware Feature Fusion Network for Person Re-IdentificationabstractThe research of part-based methods has been proven as an effective way in person re-identification (Re-ID) task. However, in existing part-based Re-ID methods, the informative interactions and potential associations among parts are neglected, which demands further study. To fill this gap, we propose a novel Part-relation-aware Feature Fusion Network (PFFN) which achieves a part-level feature fusion and enhances the discrimination of part features by fully employing helpful information from associations among parts. More specifically, a Dual-stage Attention (DA) module, consisting of spatial and part-based channel attention, is proposed to exploit complementary benefits of two kinds of attention information, thereby facilitating model with learning more discriminative features. Furthermore, Part-relation Exploitation (PE) module is proposed to learn relation-aware part features where correlative information among parts are fully employed, thereby bringing a noticeable improvement in performance. Extensive experiments are conducted on four mainstream Re-ID datasets to verify the superiority of PFFN. Compared with baseline model, PFFN has gained rank-1 accuracy improvement of 18.2% on MSMT17-v2, 12.1% on the CUHK03-Labeled, 5.6% on DukeMTMC-reid and 1.3% on Market1501, compellingly validating its effectiveness. Yanke Hou, Sicheng Lian, Haifeng Hu 0001, Dihu Chen |
IEEE Signal Process. Lett. | 3 |
| 2021 | Domain Discrepancy Elimination and Mean Face Representation Learning for NIR-VIS Face RecognitionabstractDue to its potential application in criminal cases, security systems and multimedia information retrieval, Near InfraRed (NIR) to VISible (VIS) face recognition has attracted increasing research attention in the field of computer vision. However, it is still a challenging task because of the large intra-class variations including spectrum, occlusion, lighting, blurry, expression and pose. To address the above problem, we propose a novel Domain discrepancy Elimination and Mean face Representation learning (DEMR) for NIR-VIS face recognition. The DEMR consists of two key components comprising Class-wise Domain Discrepancy Elimination (CDDE) and Cross-modal Mean Face Alignment (CMFA). Specifically, two-branch modality-specific networks are designed to extract features for VIS images and NIR images, respectively. Considering that distribution variations of cross-modal images will decrease recognition performance, we present CDDE to eliminate modality gap by narrowing distribution differences of VIS images and NIR images in a category-by-category manner. Moreover, to reduce the intra-class discrepancies and obtain compact feature representation, the CMFA is designed to achieve representation alignment between cross-domain images and VIS prototypes (i.e., VIS mean face representations), through optimizing a quadruplet constraint. Extensive experiments on multiple challenging NIR-VIS databases validate that our DEMR is effective for cross-modal face recognition task. Weipeng Hu, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Learning to Balance the Learning Rates Between Various Modalities via Adaptive Tracking FactorabstractMultimodal networks with richer information contents should always outperform the unimodal counterparts. In our experiment, however, we observe that this is not always the case. Prior efforts on multimodal tasks mainly tend to design a uniform optimization algorithm for all modalities, and yet only obtain a sub-optimal multimodal representation with the fusion of under-optimized unimodal representations, which are still challenged by performance drop on multimodal networks caused by heterogeneity among modalities. In this work, to remove the slowdowns in performance on multimodal tasks, we decouple the learning procedures of unimodal and multimodal networks by dynamically balancing the learning rates for various modalities, so that the modality-specific optimization algorithm for each modality can be obtained. Specifically, the adaptive tracking factor (ATF) is introduced to adjust the learning rate for each modality on a real-time basis. Furthermore, adaptive convergent equalization (ACE) and bilevel directional optimization (BDO) are proposed to equalize and update the ATF, avoiding sub-optimal unimodal representations due to overfitting or underfitting. Extensive experiments on multimodal sentiment analysis demonstrate that our method achieves superior performance. Sijie Mai, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Analyzing Multimodal Sentiment Via Acoustic- and Visual-LSTM With Channel-Aware Temporal Convolution NetworkabstractThe emotion of human is always expressed in a multimodal perspective. Analyzing multimodal human sentiment remains challenging due to the difficulties of the interpretation in inter-modality dynamics. Mainstream multimodal learning architectures tend to design various fusion strategies to learn inter-modality interactions, which barely consider the fact that the language modality is far more important than the acoustic and visual modalities. In contrast, we learn inter-modality dynamics in a different perspective via acoustic- and visual-LSTMs where language features play dominant role. Specifically, inside each LSTM variant, a well-designed gating mechanism is introduced to enhance the language representation via the corresponding auxiliary modality. Furthermore, in the unimodal representation learning stage, instead of using RNNs, we introduce `channel-aware' temporal convolution network to extract high-level representations for each modality to explore both temporal and channel-wise interdependencies. Extensive experiments demonstrate that our approach achieves very competitive performance compared to the state-of-the-art methods on three widely-used benchmarks for multimodal sentiment analysis and emotion recognition. Sijie Mai, Songlong Xing, Haifeng Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Cross-Age Identity Difference Analysis Model Based on Image Pairs for Age Invariant Face VerificationabstractFace recognition (FR) is a widely studied topic in the field of computer vision research. Although promising results are achieved, FR researches still face challenges of age variations. Most existing FR networks conduct classification based on the feature similarity, which may be misled by large intra class difference under age variations. Instead of calculating feature similarity, in this paper, we derive a novel cross-age face verification framework named Cross-Age Identity Difference Analysis (CIDA) model, which analyzes the identity difference between image pairs under age variations. Specifically, our framework includes two cascading networks. Firstly, an Identity Difference Feature Extractor (IDFE) is proposed to extract the difference information between the input image pair, where the identity discriminant features are effectively extracted, while other interference factors such as age, illumination, posture are suppressed. Secondly, the Direct Cross-age Verification Network (DCVN) is proposed to directly decide whether the input image pair is from the same individual. We derive a novel loss function, where the classification loss with larger age difference is assigned larger weights, which urges the classifier to pay attention to the classification process of the samples with large age gap. Besides, the loss of DCVN are integrated with the loss function of IDFE as a feedback of the final classification performance, improving the discriminant power of IDFE. Through synchronous training of the two networks, we can finally achieve end-to-end network architecture. Compared with the existing cross-age face recognition (CAFR) methods, we do not need to consider feature similarity comparison, which provides a new insight for cross-age face recognition task. Extensive experiments have been performed on the benchmark CAFR datasets which verify the effectiveness of our model. Lingshuang Du, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Parallel Architecture of Age Adversarial Convolutional Neural Network for Cross-Age Face RecognitionabstractCross-Age Face Recognition (CAFR) remains one of the most challenging tasks in the field of face recognition, as the aging process significantly affects the facial appearance. Another limitation is the insufficient special datasets that cover a wide range of ages. The key to tackle this problem is to separate the variations caused by aging process from facial features and obtain stable person-specific features. Specifically, we proposed a novel end-to-end CNN method called Age Adversarial Convolutional Neural Network (AA-CNN) with parallel network architecture. By adversarial training in Age Discrimination Network (ADN), the features extracted by AA-CNN are invariant to age variation, while remaining identity discriminative via joint training in Identity Recognition Network (IRN). Furthermore, we adopt a pyramid architecture of feature fusion to assist the ADN in adversarial training to obtain effective age-related information. The training datasets of AA-CNN are labeled by identity or age, and there is no need to search for the datasets with both identity and age labels. Extensive experiments have been conducted on the challenging aging face datasets, including FG-NET dataset, MORPH Album 2 dataset, Cross-Age Celebrity dataset, and Cross-Age LFW dataset, which demonstrates the superiority and effectiveness of the AA-CNN model. Yangjian Huang, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Multiscale Omnibearing Attention Networks for Person Re-IdentificationabstractThe past few years in the fields of Person Re-Identification (RE-ID) have seen attention mechanism receives enormous interest as it has superior performance in obtaining discriminative feature representations. However, a wide range of state-of-the-art RE-ID attention models only focus on one-dimensional attention design method, e.g. spatial attention and channels attention, hence the produced attention maps are neither detailed enough nor discriminative enough to capture complicated interactions of visual parts. Developing multi-scale attention mechanism for RE-ID, an under-studied approach, becomes a practicable method to overcome this deficiency. Toward this goal, we propose a Multiscale Omnibearing Attention Networks (MOAN) for RE-ID which is capable of utilizing the complex fusion information acquired from the multiscale attention mechanism with features being more representative. Specifically, MOAN takes full advantage of multi-sized convolution filters to obtain discriminative holistic and local feature maps, and adaptively conducts feature information augmentation by introducing an Omnibearing Attention (OA) module. Through the OA module, spatial attention and channel attention are integrated together in a unique way where they work in a complementary way. To sum up, MOAN not only inherits the merit of two kinds of attention mechanism but also performs well in extracting comprehensive feature information. Furthermore, taking into account the robustness of model performance, we formulate a Random Drop (RD) Function to facilitate training MOAN and further increase the diversity of training model for adaptation. Furthermore, to achieve end-to-end training, we utilize trainable parameters to take place of initial fixed parameters, and the model performance is experimentally promoted. Extensive experiments have been carried out on the four mainstream RE-ID datasets. As the result shows, our method with re-ranking achieves rank-1 accuracy of 92.29% on CUHK03-NP, 97.45% on Market-1501, 93.81% on DukeMTMC-reID and 81.53% on MSMT17-V2, outperforming the state-of-the-art methods and confirming the effectiveness of our method. Yewen Huang, Sicheng Lian, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Attention-Aligned Network for Person Re-IdentificationabstractCurrently, attention mechanism receives enormous interest and has been extensively employed in the fields of Person Re-Identification (RE-ID), as it gains superior performance in learning discriminative feature representations. However, most off-the-shelf attention methods are still vulnerable to cross-view inconsistency problem. Besides, they merely exploit imprecise channel attention information and coarse-grained spatial attention of homogeneous scales, being insufficient to capture subtle differences among highly-similar individuals. To this end, we propose a novel Attention-Aligned Network (AANet) to address the aforementioned problems, in which a novel Omnibearing Foreground-aware Attention (OFA) module, Attention Alignment Mechanism (AAM) and an improved triplet loss with hard mining are proposed to learn foreground attentive features for RE-ID. Specifically, AANet firstly leverages OFA module to exploit heterogeneous-scale spatial attention and foreground-aware channel attention information. Then AANet further reduces the impact of background clutter and learns camera-invariant and background-invariant representations by virtue of AAM. Last but not least, an improved triplet loss with hard mining is also introduced to enhance the feature learning capability, which can jointly minimize the intra-class distance and maximize the inter-class distance in each triplet unit. Extensive experiments are carried out to demonstrate that the proposed method outperforms most current methods on three main RE-ID benchmarks. Sicheng Lian, Weitao Jiang, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Event-Centric Hierarchical Representation for Dense Video CaptioningabstractDense video captioning aims to localize and describe multiple events in untrimmed videos, which is a challenging task that draws attention recently in computer vision. Although existing methods have achieved impressive performance, most of them only focus on local information of event segments or very simple event-level context, overlooking the complexity of event-event relationship and the holistic scene. As a result, the coherence of captions within the same video could be damaged. In this article, we propose a novel event-centric hierarchical representation to alleviate this problem. We enhance the event-level representation by capturing rich relationship between events in terms of both temporal structure and semantic meaning. Then, a caption generator with late fusion is developed to generate surrounding-event-aware and topic-aware sentences, conditioned on the hierarchical representation of visual cues from the scene level, the event level, and the frame level. Furthermore, we propose a duplicate removal method, namely temporal-linguistic non-maximum suppression (TL-NMS) to distinguish redundancy in both localization and captioning stages. Quantitative and qualitative evaluations on the ActivityNet Captions and YouCook2 datasets demonstrate that our method improves the quality of generated captions and achieves state-of-the-art performance on most metrics. Teng Wang 0007, Huicheng Zheng, Mingjing Yu, Qian Tian, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Parallel Multi-Path Age Distinguish Network for Cross-Age Face RecognitionabstractCross Age Face Recognition (CAFR) is a challenging task in the field of face recognition. There still exist some limitations in mainstream CAFR methods. On one hand, some methods need synthesizing multiple groups of features and fusing the results under different ages, where the similarity score in some age groups may affect final classification results. On the other hand, many other methods takes a facial image as a linear combination of identity information and age information, which treats age factor as a value independent of identity information, but may be inconsistent with the aging pattern of many individuals. And these methods require both age labels and identity labels in training, which is limited by the scale of existing CAFR datasets. To address the above limitations, this work proposes the Parallel Multi-path Age Distinguish Network (PMADN) model. Specifically, our model consists of two cascading networks, an Age Distinguish Mapping Network (ADMN) and a Cross-Age Feature Recombination Network (CFRN). Firstly, the face features are mapped into different age groups by parallel multi-path full connected layers in ADMN, which can better extract the identity features in a small age span. Secondly, CFRN nonlinearly recombines the mapped features to extract the age robust features that are beneficial to identity classification, which can avoid the simple linear combination of identity factor and age factor in the existing methods. What's more, our algorithm is combined with transfer learning and only uses the age label and a pre-trained ordinary face recognition network for training, which can make use of a larger aging face dataset for training. Extensive CAFR experiments performed on the benchmark MORPH Album2, CACD-VS and Cross Age LFW databases demonstrate the effectiveness and superiority of our method. Yongbo Wu, Lingshuang Du, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Facial Expression Recognition With Two-Branch Disentangled Generative Adversarial NetworkabstractFacial Expression Recognition (FER) is a challenging task in computer vision as features extracted from expressional images are usually entangled with other facial attributes, e.g., poses or appearance variations, which are adverse to FER. To achieve a better FER performance, we propose a model named Two-branch Disentangled Generative Adversarial Network (TDGAN) for discriminative expression representation learning. Different from previous methods, TDGAN learns to disentangle expressional information from other unrelated facial attributes. To this end, we build the framework with two independent branches, which are specific for facial and expressional information processing respectively. Correspondingly, two discriminators are introduced to conduct identity and expression classification. By adversarial learning, TDGAN is able to transfer an expression to a given face. It simultaneously learns a discriminative representation that is disentangled from other facial attributes for each expression image, which is more effective for FER task. In addition, a self-supervised mechanism is proposed to improve representation learning, which enhances the power of disentangling. Quantitative and qualitative results in both in-the-lab and in-the-wild datasets demonstrate that TDGAN is competitive to the state-of-the-art methods. Siyue Xie, Haifeng Hu 0001, Yizhen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Appearance-and-Dynamic Learning With Bifurcated Convolution Neural Network for Action RecognitionabstractFor a long time, learning spatiotemporal features with deep neural networks has been a difficult task in the field of computer vision. In this paper, we present a novel deep architecture, termed as Bifurcated Convolutional Neural Network (BifurcatedNet) to learn the discriminative video representation in an end-to-end manner. In our work, the BifurcatedNet is built upon the stacking bifurcated blocks that aim at simultaneously capturing the static appearance information and the temporal dynamic from input data. Specifically, the bifurcated block is composed of two separated branch, i.e., an appearance branch and a dynamic branch. The appearance branch employs 2D convolutional operation to obtain the spatial responses of image pixels or filters of each input frame, while the design of the dynamic branch is based on the spatio-temporal convolutional operation to exploit the temporal dynamic between pixels and filter response across multiple frames. Multiple experiments are conducted on two popular action recognition benchmarks: UCF101 and HMDB51. With only RGB input, the BifurcatedNet obtains the superior performance over the existing state-of-the-art models under the same experimental setting. The proposed BifurcatedNet is also implemented in a two-stream fashion by using both RGB and optical flow input, and still achieves the state-of-the-art performance, demonstrating the effectiveness of the network design. Furthermore, in order to evaluate the generalization ability, we conduct experiments on the Chalearn LAP IsoGD dataset and find that our model works well in gesture recognition tasks. Junxuan Zhang, Haifeng Hu 0001, Zheng Liu 0024 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Dual Adversarial Disentanglement and Deep Representation Decorrelation for NIR-VIS Face RecognitionabstractThe task of near-infrared and visual (NIR-VIS) face recognition refers to matching face data from different modalities, which has broad application prospects in areas such as multimedia information retrieval and criminal investigation. However, it remains a challenging task due to high intra-class variations and small-scale NIR-VIS dataset. In this paper, we propose a novel approach called Dual Adversarial Disentanglement and deep Representation Decorrelation (DADRD) to solve the NIR-VIS matching problem. In order to reduce the gap between NIR-VIS images, three key components are designed for DADRD model, including Cross-modal Margin (CmM) loss, Dual Adversarial Disentangled Variations (DADV) and Deep Representation Decorrelation (DRD). Firstly, the CmM loss captures within- and between-class information of the data, and it further reduces modality difference by a center-variation item. Secondly, the Mixed Facial Representation (MFR) layer of the backbone network is divided into three parts: the identity-related layer, the modality-related layer and the residual-related layer. The DADV is designed to reduce the intra-class variations, which consists of Adversarial Disentangled Modality Variations (ADMV) and Adversarial Disentangled Residual Variations (ADRV). Specifically, the ADMV and ADRV aim at eliminating spectrum variations and residual variations (i.e., lighting, pose, expression, occlusion, etc) respectively via an adversarial mechanism. Finally, we impose a DRD on the three decomposed features to make them irrelevant to each other, which can more effectively separate the three component information and enhance feature representations. In particular, we develop a Joint Three-stage Optimization (JTsO) strategy to effectively optimize the network. The joint formulation leads to the purification of identity information and the disentanglement of within-class variation information. Extensive experiments have been carried out on three challenging datasets, and the results demonstrate the effectiveness of our method. Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Adversarial Disentanglement Spectrum Variations and Cross-Modality Attention Networks for NIR-VIS Face RecognitionabstractNear-infrared and visual (NIR-VIS) matching task refers to the face recognition between the two images of different modalities, which remains a challenging task in the field of machine vision. The main problems of NIR-VIS Heterogeneous Face Recognition (HFR) tasks include two aspects: large intra-class differences caused by cross-modal data, and insufficient paired training samples. In this paper, an effective Adversarial Disentanglement spectrum variations and Cross-modality Attention Networks (ADCANs) is proposed for VIS-NIR matching task. Three key components are introduced to the ADCANs for reducing the gap of cross-modal images: Advanced Scatter Loss (ASL), Modality-adversarial Feature Learning (MaFL) and Cross-modality Attention Block (CmAB). The proposed ASL loss captures between- and within-class information of the data and embeds them to the network for more effective training, and it focuses on categories with small between-class distance and increases the distance between them. The MaFL consists of an Identity-Discriminative Feature Learning Network (IDFLN) and a Modality-Adversarial Disentanglement Network (MADN), which can enhance the identity-discriminative feature representations as well as disentangling spectrum variations via an adversarial learning. The IDFLN built by an end-to-end CNNs aims at learning identity-discriminative feature. While the MADN built by a discriminator D and a generator G focuses on removing modality-related information. Furthermore, to increase representation power as well as disentangling spectrum variations effectively, a CmAB block is introduced to the network, which sequentially applies spatial and channel attention modules to both the IDFLN and MADN. Since the channel attention module focuses on `what' features to suppress or emphasize, an orthogonality constraint is introduced to the two channel attention modules, which allows MADN and IDFLN to focus on learning modality-related features and identity-related features, respectively. In particular, the ADCANs consists of multiple CmAB blocks to learn discriminative features and disentangle spectrum variations. A large number of experiments on three challenging HFR datasets indicate that the proposed ADCANs is effective for VIS-NIR HFR task. Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Y-Net: Dual-branch Joint Network for Semantic SegmentationabstractMost existing segmentation networks are built upon a “ U -shaped” encoder–decoder structure, where the multi-level features extracted by the encoder are gradually aggregated by the decoder. Although this structure has been proven to be effective in improving segmentation performance, there are two main drawbacks. On the one hand, the introduction of low-level features brings a significant increase in calculations without an obvious performance gain. On the other hand, general strategies of feature aggregation such as addition and concatenation fuse features without considering the usefulness of each feature vector, which mixes the useful information with massive noises. In this article, we abandon the traditional “ U -shaped” architecture and propose Y-Net, a dual-branch joint network for accurate semantic segmentation. Specifically, it only aggregates the high-level features with low-resolution and utilizes the global context guidance generated by the first branch to refine the second branch. The dual branches are effectively connected through a Semantic Enhancing Module, which can be regarded as the combination of spatial attention and channel attention. We also design a novel Channel-Selective Decoder (CSD) to adaptively integrate features from different receptive fields by assigning specific channelwise weights, where the weights are input-dependent. Our Y-Net is capable of breaking through the limit of singe-branch network and attaining higher performance with less computational cost than “ U -shaped” structure. The proposed CSD can better integrate useful information and suppress interference noises. Comprehensive experiments are carried out on three public datasets to evaluate the effectiveness of our method. Eventually, our Y-Net achieves state-of-the-art performance on PASCAL VOC 2012, PASCAL Person-Part, and ADE20K dataset without pre-training on extra datasets. Yizhen Chen, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Bi-Directional Co-Attention Network for Image CaptioningabstractImage Captioning, which automatically describes an image with natural language, is regarded as a fundamental challenge in computer vision. In recent years, significant advance has been made in image captioning through improving attention mechanism. However, most existing methods construct attention mechanisms based on singular visual features, such as patch features or object features, which limits the accuracy of generated captions. In this article, we propose a Bidirectional Co-Attention Network (BCAN) that combines multiple visual features to provide information from different aspects. Different features are associated with predicting different words, and there are a priori relations between these multiple visual features. Based on this, we further propose a bottom-up and top-down bi-directional co-attention mechanism to extract discriminative attention information. Furthermore, most existing methods do not exploit an effective multimodal integration strategy, generally using addition or concatenation to combine features. To solve this problem, we adopt the Multivariate Residual Module (MRM) to integrate multimodal attention features. Meanwhile, we further propose a Vertical MRM to integrate features of the same category, and a Horizontal MRM to combine features of the different categories, which can balance the contribution of the bottom-up co-attention and the top-down co-attention. In contrast to the existing methods, the BCAN is able to obtain complementary information from multiple visual features via the bi-directional co-attention strategy, and integrate multimodal information via the improved multivariate residual strategy. We conduct a series of experiments on two benchmark datasets (MSCOCO and Flickr30k), and the results indicate that the proposed BCAN achieves the superior performance. Weitao Jiang, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionabstractLearning joint embedding space for various modalities is of vital importance for multimodal fusion. Mainstream modality fusion approaches fail to achieve this goal, leaving a modality gap which heavily affects cross-modal fusion. In this paper, we propose a novel adversarial encoder-decoder-classifier framework to learn a modality-invariant embedding space. Since the distributions of various modalities vary in nature, to reduce the modality gap, we translate the distributions of source modalities into that of target modality via their respective encoders using adversarial training. Furthermore, we exert additional constraints on embedding space by introducing reconstruction loss and classification loss. Then we fuse the encoded representations using hierarchical graph neural network which explicitly explores unimodal, bimodal and trimodal interactions in multi-stage. Our method achieves state-of-the-art performance on multiple datasets. Visualization of the learned embeddings suggests that the joint embedding space learned by our method is discriminative. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
AAAI | 2 |
| 2020 | Adaptive Interaction Modeling via Graph Operations SearchabstractInteraction modeling is important for video action analysis. Recently, several works design specific structures to model interactions in videos. However, their structures are manually designed and non-adaptive, which require structures design efforts and more importantly could not model interactions adaptively. In this paper, we automate the process of structures design to learn adaptive structures for interaction modeling. We propose to search the network structures with differentiable architecture search mechanism, which learns to construct adaptive structures for different videos to facilitate adaptive interaction modeling. To this end, we first design the search space with several basic graph operations that explicitly capture different relations in videos. We experimentally demonstrate that our architecture search framework learns to construct adaptive interaction modeling structures, which provides more understanding about the relations between the structures and some interaction characteristics, and also releases the requirement of structures design efforts. Additionally, we show that the designed basic graph operations in the search space are able to model different interactions in videos. The experiments on two interaction datasets show that our method achieves competitive performance with state-of-the-arts. Haoxin Li, Wei-Shi Zheng 0001, Haifeng Hu 0001, Jian-Huang Lai |
CVPR | 4 |
| 2020 | Learning to Contrast the Counterfactual Samples for Robust Visual Question AnsweringabstractIn the task of Visual Question Answering (VQA), most state-of-the-art models tend to learn spurious correlations in the training set and achieve poor performance in out-ofdistribution test data.Some methods of generating counterfactual samples have been proposed to alleviate this problem.However, the counterfactual samples generated by most previous methods are simply added to the training data for augmentation and are not fully utilized.Therefore, we introduce a novel selfsupervised contrastive learning mechanism to learn the relationship between original samples, factual samples and counterfactual samples.With the better cross-modal joint embeddings learned from the auxiliary training objective, the reasoning capability and robustness of the VQA model are boosted significantly.We evaluate the effectiveness of our method by surpassing current state-of-the-art models on the VQA-CP dataset, a diagnostic benchmark for assessing the VQA model's robustness. Zujie Liang, Weitao Jiang, Haifeng Hu 0001, Jiaying Zhu |
EMNLP (1) | 3 |
| 2020 | Partial domain adaptation based on shared class oriented adversarial network
Wenjie Qiu 0003, Wendong Chen, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2020 | Self-adaptive weighted synthesised local directional pattern integrating with sparse autoencoder for expression recognition based on improved multiple kernel learning strategyabstractThis study presents a novel method for solving facial expression recognition (FER) tasks which uses a self‐adaptive weighted synthesised local directional pattern (SW‐SLDP) descriptor integrating sparse autoencoder (SA) features based on improved multiple kernel learning (IMKL) strategy. The authors’ work includes three parts. Firstly, the authors propose a novel SW‐SLDP feature descriptor which divides the facial images into patches and extracts sub‐block features synthetically according to both distribution information and directional intensity contrast. Then self‐adaptive weights are assigned to each sub‐block feature according to the projection error between the expressional image and neutral image of each patch, which can highlight such areas containing more expressional texture information. Secondly, to extract a discriminative high‐level feature, they introduce SA for feature representation, which extracts the hidden layer representation including more comprehensive information. Finally, to combine the above two kinds of features, an IMKL strategy is developed by effectively integrating both soft margin learning and intrinsic local constraints, which is robust to noisy condition and thus improve the classification performance. Extensive experimental results indicate their model can achieve competitive or even better performance with existing representative FER methods. Lingshuang Du, Yongbo Wu, Haifeng Hu 0001 |
IET Comput. Vis. | 3 |
| 2020 | ADN for object detectionabstractOwing to large‐scale diversity and location uncertainty in object detection, how to enrich semantic information has become an important issue that attracts a lot of concern. In this study, the authors propose a novel attentional detection network (ADN) to enrich semantic information of feature maps by adding an extra attention branch to the classic detection network. Compared to previous methods (e.g. feature pyramid network (FPN), single shot multibox detector (SSD)) that producing massive anchors in different layers of feature maps to detect objects with different scales and aspect ratios, which is very time‐consuming, their network is lightweight and do not need to produce extra anchors. Furthermore, ADN can be applied to different object detectors with little computational cost. Extensive experiments indicate that ADN has good detection performance on different datasets without bells and whistles. Jinding Wang, Haifeng Hu 0001, Xinlong Lu |
IET Comput. Vis. | 2 |
| 2020 | Discrete ZNN models of Adams-Bashforth (AB) type solving various future problems with motion control of mobile manipulator
Min Yang 0010, Yunong Zhang, Haifeng Hu 0001 |
Neurocomputing | 3 |
| 2020 | Multi-layer Adaptive Feature Fusion for Semantic Segmentation
Yizhen Chen, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2020 | Unsupervised Domain Adaptation via Discriminative Classes-Center Feature Learning in Adversarial Network
Wendong Chen, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2020 | Action Recognition with Multiple Relative Descriptors of Trajectories
Zhongke Liao, Haifeng Hu 0001, Yichu Liu |
Neural Process. Lett. | 2 |
| 2020 | Gaussian Pyramid of Conditional Generative Adversarial Network for Real-World Noisy Image Denoising
Ruijun Ma 0001, Bob Zhang 0001, Haifeng Hu 0001 |
Neural Process. Lett. | 3 |
| 2020 | Complementary Boundary Estimation Network for Temporal Action Proposal Generation
Jinding Wang, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2020 | Discriminative Face Recognition Methods with Structure and Label Information via l2-Norm Regularization
Haifeng Hu 0001, Lin Li 0032, Tundong Liu |
Neural Process. Lett. | 2 |
| 2020 | Generative attention adversarial classification network for unsupervised domain adaptation
Wendong Chen, Haifeng Hu 0001 |
Pattern Recognit. | 2 |
| 2020 | Age Factor Removal Network Based on Transfer Learning and Adversarial Learning for Cross-Age Face RecognitionabstractIt is well known that expression, pose variations, and especially age factors always affect the performance of a face recognition system in practical conditions. In this paper, we propose a novel framework called age factor removal network (AFRN) for cross-age face recognition, which combines the concepts of transfer learning and adversarial learning to enhance the performance of a pretrained face recognition network and suppress the influence of attributes, such as aging, expression, and pose variations. Similar to the generative adversarial network (GAN), our model consists of two networks: an identity feature generator G and an age discriminator D. First, the D network is trained to discriminate the age information from the feature extracted by G. Second, the G network simultaneously learns two tasks: extracting the feature through transfer learning from a pretrained face recognition network and suppressing age information through adversarial learning with D. Through the optimization of the two networks, the age factor in feature extraction process of G is removed. Besides, our network only requires the attributes' label to train and preserves the identity discriminant power through transfer learning, and thus, it does not depend on multi-label face databases. The extensive experiments have been performed on the benchmark cross-age face datasets, including MORPH Album2, CACD-VS, and Cross Age LFW, which verify the effectiveness of our model. What is more, our model can be extended to other practical face recognition tasks and decreases the influence of expressions and poses, which is verified by the experimental results on CMU Multi-PIE database. Lingshuang Du, Haifeng Hu 0001, Yongbo Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Deeply Associative Two-Stage Representations Learning Based on Labels Interval Extension Loss and Group Loss for Person Re-IdentificationabstractPerson Re-identification (ReID) aims to match people across non-overlapping camera views in a public space, which is usually regarded as an image retrieval problem to match query images with pedestrian images in the gallery. It is challenging since many difficulties exist such as pose misalignments, occlusions, similar appearance when detecting people. Existing researches on ReID mainly focus on two major problems: representation learning and metric learning. In this paper, we target at learning discriminative representations and make two contributions in total. (i) We propose a novel architecture named Deeply Associative Two-stage Representations Learning (DATRL). It contains the global re-initialization stage and fully-perceptual classification stage employing two identical CNNs associatively at the same time. On the global stage, we take on the backbone of one deep CNN e.g., dozens of layers in the front of Resnet-50 as a normal re-initialization subnetwork. Meanwhile, we apply our own proposed 3D-transpose technique into the backbone of the other CNN to form the 3D-transpose re-initialization subnetwork. The fully-perceptual stage is actually made up of the leftover layers of the original CNNs. On this stage, we take both the global representations learned at multiple hierarchies and the local representations uniformly-partitioned on the highest conv-layer into consideration, and then optimizing them separately for classification. (ii) We introduce a new joint loss function in which our proposed Labels Interval Extension loss (LIEL) and Group loss (GL) are combined to enhance the performance of gradient decent as well as increasing the distances between image features with different identities. We apply the above DATRL, LIEL and GL to ReID thus obtaining DATRL-ReID. Experimental results on four datasets CUHK03, Market-1501, DukeMTMC-reID and MSMT17-V2 demonstrate that DATRL-ReID shows excellent performance in improving recognition accuracy and is superior to state-of-the-art methods. Yewen Huang, Yi Huang 0035, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Three-Dimension Transmissible Attention Network for Person Re-IdentificationabstractIn this work, we propose a Three-Dimensional Transmissible Attention Network (3DTANet) for Person Re-Identification, which can transmit the attention information from layer to layer and attend to the person image from a three-dimensional perspective. Main contributions of the 3DTANet are: (i) A novel Transmissible Attention (TA) mechanism is introduced, which can transfer attention information between convolution layers. Different from traditional attention mechanism, not only can it convey accumulated attention information layer by layer but also guide the network to retain holistic attention information. (ii) We propose a Three-Dimension Attention (3DA) mechanism, which is capable of extracting a three-dimensional attention map. While previous researches on image attention mechanism extracts channel or spatial attention information separately, 3DA mechanism pays attention to channel and spatial information simultaneously, thereby making them play better complementary role in attention extraction. (iii) A new loss function named L2-norm Multi-labels Loss (L2ML) is applied to acquire higher recognition accuracy calculated by multi labels of same ID and corresponding feature representation. Quite different from the common loss functions, L2-norm Multi-labels Loss is specifically good at optimizing feature distance. In brief, 3DTANet gains two-fold benefit toward higher accuracy. For one thing, the attention information is informative and can be transmitted, feature being more representative. For another, our model is computationally lightweight and can be easily applied to real scenarios. We extensively conduct experiments on four Person Re-Identification benchmark datasets. Our model achieves rank-1 accuracy of 87.50% on CUHK03, 96.23% on Market-1501, 92.50% on DukeMTMC-reID and 76.60% on MSMT17-V2 respectively. The results confirm that the 3DTANet can extract more representative features and attain a higher recognition accuracy, outperforming the state-of-the-art methods. Yewen Huang, Sicheng Lian, Suian Zhang, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Cycle Age-Adversarial Model Based on Identity Preserving Network and Transfer Learning for Cross-Age Face RecognitionabstractAge variations bring a large challenge for face recognition tasks. Existing Cross-Age Face Recognition (CAFR) methods have two limitations. Firstly, many CAFR approaches require both age labels and identity labels for training. However, it is difficult to collect images under a large age span from each individual. Secondly, many works are based on the assumption that age and identity information are independent of each other, which may not satisfy various conditions. In this paper, a Cycle Age-Adversarial Model (CAAM) is proposed for CAFR, which only uses the age labels for training without considering independence hypothesis. CAAM includes two different branch networks. Firstly, the branch of Age-robust Feature Extracting Model (AFEM) is designed to adaptively learn age-invariant features by adversarial learning scheme, which includes an age discriminator network and a feature generator network. The age discriminator network is trained to discriminate the age information, and the generator extracts age-invariant features through adversarial learning with discriminator. Secondly, a branch of the Identity Preserving Network (IPN) is proposed to keep identity information, which introduces Unsupervised Identity Loss (UIL) to enlarge the inter-class distance, and decrease the loss of identity information in the learning process. Finally, the features of the two branches are cyclically optimized through minmizing Feature Consistency Loss (FCL), which integrates age invariance learning and identity discrimination learning into final feature representation. Different from existing CAFR networks, our adversarial learning strategy for age-robust feature learning can be generalized to other attributes including pose and expression. Moreover, we introduce cycle optimization strategy to merge the advantages of two branch networks, which is a novel strategy to fuse multi-task features. Extensive CAFR experiments performed on the benchmark MORPH Album2, CACD-VS and Cross Age LFW databases demonstrate the effectiveness and superiority of CAAM. Lingshuang Du, Haifeng Hu 0001, Yongbo Wu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Adaptive Discrete ZND Models for Tracking Control of Redundant ManipulatorabstractIn recent years, many models with high precision for redundant manipulator tracking control have been proposed based on precise kinematics equations. Nevertheless, without precise kinematic equations, developing a model with high precision for tracking control is meaningful. With the help of zeroing neural dynamics (ZND), a continuous ZND model with adaptive Jacobian matrix is obtained. For better computer operation and easier understanding, developing corresponding discrete ZND (DZND) model is also significant. Therefore, two DZND models (termed DZND-I model and DZND-II model) are proposed in this article on the basis of two discretization formulas, respectively. Meanwhile, theoretical analyses are conducted to ensure the efficacy of DZND-I model and DZND-II model. Finally, the efficacy of the two DZND models with adaptive Jacobian matrix is substantiated by experimental results on the basis of the four-link manipulator, UR5 manipulator, and Jaco2 manipulator, respectively. Min Yang 0010, Yunong Zhang, Zhijun Zhang 0003, Haifeng Hu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Efficient and Fast Real-World Noisy Image Denoising by Combining Pyramid Neural Network and Two-Pathway Unscented Kalman FilterabstractRecently, image prior learning has emerged as an effective tool for image denoising, which exploits prior knowledge to obtain sparse coding models and utilize them to reconstruct the clean image from the noisy one. Albeit promising, these prior-learning based methods suffer from some limitations such as lack of adaptivity and failed attempts to improve performance and efficiency simultaneously. With the purpose of addressing these problems, in this paper, we propose a Pyramid Guided Filter Network (PGF-Net) integrated with pyramid-based neural network and Two-Pathway Unscented Kalman Filter (TP-UKF). The combination of pyramid network and TP-UKF is based on the consideration that the former enables our model to better exploit hierarchical and multi-scale features, while the latter can guide the network to produce an improved (a posteriori) estimation of the denoising results with fine-scale image details. Through synthesizing the respective advantages of pyramid network and TP-UKF, our proposed architecture, in stark contrast to prior learning methods, is able to decompose the image denoising task into a series of more manageable stages and adaptively eliminate the noise on real images in an efficient manner. We conduct extensive experiments and show that our PGF-Net achieves notable improvement on visual perceptual quality and higher computational efficiency compared to state-of-the-art methods. Ruijun Ma 0001, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Image Process. | 2 |
| 2020 | Robust Facial Landmark Detection via Heatmap-Offset RegressionabstractFacial landmark detection aims at localizing multiple keypoints for a given facial image, which usually suffers from variations caused by arbitrary pose, diverse facial expression and partial occlusion. In this paper, we develop a two-stage regression network for facial landmark detection on unconstrained conditions. Our model consists of a Structural Hourglass Network (SHN) for detecting the initial locations of all facial landmarks based on heatmap generation, and a Global Constraint Network (GCN) for further refining the detected locations based on offset estimation. Specifically, SHN introduces an improved Inception-ResNet unit as basic building block, which can effectively improve the receptive field and learn contextual feature representations. In the meanwhile, a novel loss function with adaptive weight is proposed to make the whole model focus on the hard landmarks precisely. GCN attempts to explore the spatial contextual relationship between facial landmarks and refine the initial locations of facial landmarks by optimizing the global constraint. Moreover, we develop a pre-processing network to generate features with different scales, which will be transmitted to SHN and GCN for effective feature representations. Different from existing models, the proposed method realizes the heatmap-offset framework, which combines the outputs of heatmaps generated by SHN and coordinates estimated by GCN, to obtain an accurate prediction. The extensive experimental results on several challenging datasets, including 300W, COFW, AFLW, and 300-VW confirm that our method achieve competitive performance compared with the state-of-the-art algorithms. Haifeng Hu 0001, Shenming Feng |
IEEE Trans. Image Process. | 2 |
| 2020 | Disentangled Spectrum Variations Networks for NIR-VIS Face RecognitionabstractSurveillance cameras often capture near infrared images since it provides a low-cost and effective solution to acquire high-quality images under low-light environments. However, visual versus near infrared (VIS-NIR) heterogeneous face recognition (HFR) is still a challenging issue in computer vision community due to the gap between sensing patterns of different spectrums as well as the lack of sufficient training samples. To solve the above problem, in this paper, we present an effective Disentangled Spectrum Variations Networks (DSVNs) for VISNIR HFR. Two key strategies are introduced to the DSVNs for disentangling spectrum variations between two domains: Spectrum-adversarial Discriminative Feature Learning (SaDFL) and Step-wise Spectrum Orthogonal Decomposition (SSOD). The SaDFL consists of Identity-Discriminative subnetwork (IDNet) and Auxiliary Spectrum Adversarial subnetwork (ASANet). On the one hand, the IDNet is composed of a generator GHand a discriminator DUfor extracting identity-discriminative feature. On the other hand, the ASANet is built by a generator GHand a discriminator DMfor eliminating modality-variant spectrum information under the guidance of the discriminator DM. The identity-label and modality-label HFR datasets are used to train the DSVNs with triplet loss. Both IDNet and ASANet can jointly enhance the domain-invariant feature representations via an adversarial learning. Furthermore, to disentangle spectrum variations effectively as well as making identity information and modality information unrelated to each other, we present a new topology of connection block called Disentangled Spectrum Variations (DSV). An orthogonality constraint is imposed to DSV at the convolution level for channel-wise orthogonal decomposition between the modality-invariant identity information and modalityvariant spectrum information. In particular, the SSOD is built by stacking multiple modularized mirco-block DSV, and thereby enjoys the benefits of disentangling spectrum variation step by step. Moreover, we investigate the similarity calculation method to further improve the HFR performance. To sum up, the designed DSVNs leads to a purification of identity information as well as an elimination of modality information. Extensive experiments are carried out on two challenging NIR-VIS HFR datasets CASIA NIRVIS 2.0 and Oulu-CASIA NIR-VIS, demonstrating the superiority of the proposed method. Weipeng Hu, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Locally Confined Modality Fusion Network With a Global Perspective for Multimodal Human Affective ComputingabstractIn this paper, we propose a novel multimodal fusion framework, called the locally confined modality fusion network (LMFN), that contains a bidirectional multiconnected LSTM (BM-LSTM) to address the multimodal human affective computing problem. In the LMFN, we introduce a generic fusion structure that explores both local and global fusion to obtain an integral comprehension of information. Specifically, we partition the feature vector corresponding to each modality into multiple segments and learn every local interaction through a tensor fusion procedure. Global interaction is then modeled by learning the dependence between local tensors via an originally designed BM-LSTM architecture, establishing a direct connection of cells and states of local tensors that are several time steps apart. With the LMFN, we achieve advantages over other methods in the following aspects: 1) local interactions are successfully modeled using a feasible vector segmentation procedure that can explore cross-modal dynamics in a more specialized manner; 2) global interactions are modeled to obtain an integral view of multimodal information using BM-LSTM, which guarantees an adequate flow of information; and 3) our general fusion structure is highly extendable by applying other local and global fusion methods. Experiments show that the LMFN yields state-of-the-art results. Moreover, the LMFN achieves higher efficiency compared to other models by applying the outer product as the fusion method. Sijie Mai, Songlong Xing, Haifeng Hu 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | General 7-Instant DCZNN Model Solving Future Different-Level System of Nonlinear Inequality and Linear EquationabstractIn this article, a novel and challenging problem called future different-level system of nonlinear inequality and linear equation (FDLSNILE) is proposed and investigated. To solve FDLSNILE, the corresponding continuous different-level system of nonlinear inequality and linear equation (CDLSNILE) is first analyzed, and then, a continuous combined zeroing neural network (CCZNN) model for solving CDLSNILE is proposed. To obtain a discrete combined zeroing neural network (DCZNN) model for solving FDLSNILE, a high-precision general 7-instant Zhang et al. discretization (ZeaD) formula for the first-order time derivative approximation is proposed. Furthermore, by applying the general 7-instant ZeaD formula to discretize the CCZNN model, a general 7-instant DCZNN (7IDCZNN) model is thus proposed for solving FDLSNILE. For comparison, by using three conventional ZeaD formulas, three conventional DCZNN models are also developed. Meanwhile, theoretical analyses and results guarantee the efficacy and superiority of the general 7IDCZNN model compared with the other three conventional DCZNN models for solving FDLSNILE. Finally, several comparative numerical experiments, including the motion control of a 5-link redundant manipulator, are provided to substantiate the efficacy and superiority of the general 7-instant ZeaD formula and the corresponding 7IDCZNN model. Min Yang 0010, Yunong Zhang, Haifeng Hu 0001, Binbin Qiu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Learning Joint Structure for Human Pose EstimationabstractRecently, tremendous progress has been achieved on human pose estimation with the development of convolutional neural networks (CNNs). However, current methods still suffer from severe occlusion, back view, and large pose variation due to the lack of consideration of the spatial relationship between different joints, which can provide strong cues for localizing the hidden keypoints. In this work, we design a Structural Pose Network (SPN) to take full advantage of joint structure for human pose estimation under unconstrained environment. Specifically, the proposed model is composed of two subnets: Structure Residual Network (SRN) and Structure Improving Network (SIN). Given an input image, SRN first captures rich joint structure as priors through a multi-branch feature extraction module, following a hourglass network with pyramid residual units to enlarge the receptive field and further obtain structural feature representations. SIN, based on coordinate regression, can optimize the spatial relationship of different joints via the attention mechanism, thus refining the initial prediction from SRN. In addition, we propose a novel structure-consistency constraint, which can maintain the structural consistency between the joints and body parts via estimating whether the joints are located in their corresponding parts. At the same time, an online hard regions mining (OHRM) strategy is introduced to drive the network to pay corresponding attention to different body parts. The experimental results on three challenging datasets show that our method outperforms other state-of-the-art algorithms. Shenming Feng, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | FIN: Feature Integrated Network for Object DetectionabstractMulti-layer detection is a widely used method in the field of object detection. It extracts multiple feature maps with different resolutions from the backbone network to detect objects of different scales, which can effectively cope with the problem of object scale change in object detection. Although the multi-layer detection utilizes multiple detection layers to alleviate the burden of one single detection layer and can improve the detection accuracy to some extent, this method has two limitations. First, manually assigning anchor boxes of different sizes to different feature maps is too dependent on the human experience. Second, there is a semantic gap between each detection layer in multi-layer detection. The same detector needs to simultaneously process the detection layers with inconsistent semantic strength, which increases the optimization difficulty of the detector. In this article, we propose a feature integrated network (FIN) based on single layer detection to deal with the problems mentioned above. Different from the existing methods, we design a series of verification experiments based on the multi-layer detection model, which shows that the shallow high-resolution feature map has the potential to simultaneously and effectively detect objects of various scales. Considering that the semantic information of the shallow feature map is weak, we propose two modules to enhance the representation ability of the single detection layer. First, we propose a detection adaptation network (DANet) to extract powerful feature maps that are useful for object detection tasks. Second, we combine global context information and local detail information with a verified hourglass module (VHM) to generate a single feature map with high resolution and rich semantic information so that we can assign all anchor boxes to this detection layer. In our model, all the detection operations are concentrated on a high-resolution feature map whose semantic information and detailed information are enhanced as much as possible. Therefore, the proposed model can solve the problem of anchor assignment and inconsistent semantic strength between multiple detection layers mentioned above. A large number of experiments on the Pattern Analysis, Statistical Modelling and Computational Learning Visual Object Classes (PASCAL VOC) and Microsoft Common Objects in Context (MS COCO) datasets show that our model has good detection performance for objects of various sizes. The proposed model can achieve<?brk?> 81.9 mAP when the size of the input image is 300 × 300. Xiaofan Luo, Fukoeng Wong, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Constrained LSTM and Residual Attention for Image CaptioningabstractVisual structure and syntactic structure are essential in images and texts, respectively. Visual structure depicts both entities in an image and their interactions, whereas syntactic structure in texts can reflect the part-of-speech constraints between adjacent words. Most existing methods either use visual global representation to guide the language model or generate captions without considering the relationships of different entities or adjacent words. Thus, their language models lack relevance in both visual and syntactic structure. To solve this problem, we propose a model that aligns the language model to certain visual structure and also constrains it with a specific part-of-speech template. In addition, most methods exploit the latent relationship between words in a sentence and pre-extracted visual regions in an image yet ignore the effects of unextracted regions on predicted words. We develop a residual attention mechanism to simultaneously focus on the pre-extracted visual objects and unextracted regions in an image. Residual attention is capable of capturing precise regions of an image corresponding to the predicted words considering both the effects of visual objects and unextracted regions. The effectiveness of our entire framework and each proposed module are verified on two classical datasets: MSCOCO and Flickr30k. Our framework is on par with or even better than the state-of-the-art methods and achieves superior performance on COCO captioning Leaderboard. Haifeng Hu 0001, Songlong Xing, Xinlong Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Joint Stacked Hourglass Network and Salient Region Attention Refinement for Robust Face AlignmentabstractFacial landmark detection aims to locate keypoints for facial images, which typically suffer from variations caused by arbitrary pose, diverse facial expressions, and partial occlusion. In this article, we propose a coarse-to-fine framework that joins a stacked hourglass network and salient region attention refinement for robust face alignment. To achieve this goal, we first present a multi-scale region learning module to analyze the structure information at a different facial region and extract a strong discriminative deep feature. Then we employ a stacked hourglass network for heatmap regression and initial facial landmarks prediction. Specifically, the stacked hourglass network introduces an improved Inception-ResNet unit as a basic building block, which can effectively improve the receptive field and learn contextual feature representations. Meanwhile, a novel loss function takes into account global weights and local weights to make the heatmap regression more accurate. Different from existing heatmap regression models, we present a salient region attention refinement module to extract a precise feature based on the heatmap regression, and utilize the filtered feature for landmarks refinement to achieve accurate prediction. Extensive experimental results of several challenging datasets (including 300 Faces in the Wild, Caltech Occluded Faces in the Wild, and Annotated Facial Landmarks Faces in the Wild) confirm that our approach can achieve more competitive performance than the most advanced algorithms. Haifeng Hu 0001, Guobin Shen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Hierarchical Attention Network for Image CaptioningabstractRecently, attention mechanism has been successfully applied in image captioning, but the existing attention methods are only established on low-level spatial features or high-level text features, which limits richness of captions. In this paper, we propose a Hierarchical Attention Network (HAN) that enables attention to be calculated on pyramidal hierarchy of features synchronously. The pyramidal hierarchy consists of features on diverse semantic levels, which allows predicting different words according to different features. On the other hand, due to the different modalities of features, a Multivariate Residual Module (MRM) is proposed to learn the joint representations from features. The MRM is able to model projections and extract relevant relations among different features. Furthermore, we introduce a context gate to balance the contribution of different features. Compared with the existing methods, our approach applies hierarchical features and exploits several multimodal integration strategies, which can significantly improve the performance. The HAN is verified on benchmark MSCOCO dataset, and the experimental results indicate that our model outperforms the state-of-the-art methods, achieving a BLEU1 score of 80.9 and a CIDEr score of 121.7 in the Karpathy’s test split. Haifeng Hu 0001 |
AAAI | 3 |
| 2019 | Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective ComputingabstractWe propose a general strategy named 'divide, conquer and combine' for multimodal fusion.Instead of directly fusing features at holistic level, we conduct fusion hierarchically so that both local and global interactions are considered for a comprehensive interpretation of multimodal embeddings.In the 'divide' and 'conquer' stages, we conduct local fusion by exploring the interaction of a portion of the aligned feature vectors across various modalities lying within a sliding window, which ensures that each part of multimodal embeddings are explored sufficiently.On its basis, global fusion is conducted in the 'combine' stage to explore the interconnection across local interactions, via an Attentive Bi-directional Skipconnected LSTM that directly connects distant local interactions and integrates two levels of attention mechanism.In this way, local interactions can exchange information sufficiently and thus obtain an overall view of multimodal information.Our method achieves state-ofthe-art performance on multimodal affective computing with higher efficiency. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
ACL (1) | 2 |
| 2019 | Stacked Hourglass Network Joint with Salient Region Attention Refinement for Face AlignmentabstractLocalizing facial landmarks is a fundamental step in facial image analysis. However, the problem continues to be challenging in condition of large variations caused by pose disparity, illumination, expression and occlusion. In this paper, we propose a coarse-to-fine framework which joints stacked hourglass network and salient region attention refinement for robust face alignment. To achieve this, we firstly develop a multi-scale region learning module (MSL) to analyze the structure and texture information at different facial region and extract strong discriminative deep feature. Then we employ a novel convolutional neural network named stacked hourglass network (SHN) for heatmap regression and initial facial landmarks prediction. Moreover, we present a salient region attention module (SRA) to extract precise feature based on the heatmap regression, and the filtered feature is used for landmarks refinement. The extensive experimental results on two public datasets, including 300W and COFW, confirm the validity of our model. Haifeng Hu 0001 |
FG | 2 |
| 2019 | Channel and Constraint Compensation for Generative Adversarial Networks
Wei Wang 0210, Haifeng Hu 0001, Dihu Chen |
PRCV (1) | 2 |
| 2019 | Weighted Patch-based Manifold Regularization Dictionary Pair Learning model for facial expression recognition using Iterative Optimization Classification Strategy
Lingshuang Du, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | Discriminant Deep Feature Learning based on joint supervision Loss and Multi-layer Feature Fusion for heterogeneous face recognition
Weipeng Hu, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | Residual attention unit for action recognition
Zhongke Liao, Haifeng Hu 0001, Junxuan Zhang, Chang Yin |
Comput. Vis. Image Underst. | 2 |
| 2019 | Attentive matching network for few-shot learning
Sijie Mai, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | Visual Skeleton and Reparative Attention for Part-of-Speech image captioning system
Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | Adaptive learning feature pyramid for object detectionabstractInconsistent detection performance for objects of different scales lies in many state‐of‐the‐art object detection models. The feature pyramid network (FPN) alleviates this problem by fusing multi‐scale feature maps through a top‐down path. However, the features fusion strategy used in FPN lacks learning ability, which may result in suboptimal performance of the model. In this study, the authors propose a cross‐scale feature fusion network (CSFF) to fuse the low‐level location feature maps with the high‐level semantic feature maps. The CSFF first embeds a dilated convolution and deconvolution layer into the top‐down path of the FPN to enhance the learning ability of feature fusion. After that, an attention module is applied to suppress distraction and interference in the feature map. Each component of the CSFF is highly decoupled and can easily cooperate with a base network in an end‐to‐end training manner. In this study, they combine the CSFF with faster region with convolutional neural network and conduct a series of experiments on the PASCAL VOC 2007 and 2012 object detection datasets. Without any bells and whistles, the CSFF achieves a considerable detection improvement over the baseline network. Fukoeng Wong, Haifeng Hu 0001 |
IET Comput. Vis. | 2 |
| 2019 | Hierarchical extended collaborative representation based classification for single-sample face recognitionabstractCollaborative representation based classification (CRC) has been widely used and shown good performance in face recognition (FR). Afterwards, hierarchical representation based classification has recently been proposed and aims to enhance the classification performance of the CRC method. However, these methods highly depend on the over‐complete dictionary comprised of sufficient training samples, and cannot be directly applied for single‐sample FR. In this study, the authors propose a novel CRC‐based FR framework to address this issue, which is named hierarchical extended collaborative representation based classification (HECRC). Firstly, they integrate hierarchical representation based model with low‐rank constrained variation dictionaries. Secondly, they select training samples that are the nearest neighbours of test images to obtain a more discriminative training dictionary, where an adaptive scheme is introduced to select proper samples automatically instead of setting a predefined number in traditional methods. Finally, the refined training dictionary and the learned variation dictionaries are jointly utilised to represent the test sample. Moreover, they combined the proposed HECRC with deep features to further improve the recognition rate. Experiments have been conducted on the AR and FERET datasets, and the results show that the proposed method has a substantial improvement over existing algorithms for single‐sample FR. Yuelai Yuan, Dihu Chen, Haifeng Hu 0001, Lingshuang Du |
IET Comput. Vis. | 3 |
| 2019 | Nuclear norm based adapted occlusion dictionary learning for face recognition with occlusion and illumination changes
Lingshuang Du, Haifeng Hu 0001 |
Neurocomputing | 2 |
| 2019 | An Improved Method for Semantic Image Inpainting with GANs: Progressive Inpainting
Yizhen Chen, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Image Captioning with Text-Based Visual Attention
Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Fine Tuning Dual Streams Deep Network with Multi-scale Pyramid Decision for Heterogeneous Face Recognition
Weipeng Hu, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Action Recognition Using Multiple Pooling Strategies of CNN Features
Haifeng Hu 0001, Zhongke Liao |
Neural Process. Lett. | 1 |
| 2019 | c-RNN: A Fine-Grained Language Model for Image Captioning
Gengshi Huang, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Spatiotemporal Fusion Networks for Video Action Recognition
Zheng Liu 0024, Haifeng Hu 0001, Junxuan Zhang |
Neural Process. Lett. | 2 |
| 2019 | Image Captioning Using Region-Based Attention Joint with Time-Varying Attention
Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Image Caption with Endogenous-Exogenous Attention
Teng Wang 0007, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Age-Invariant Face Recognition Using Coupled Similarity Reference Coding
Yongbo Wu, Haifeng Hu 0001, Haoxi Li |
Neural Process. Lett. | 2 |
| 2019 | Adaptive Syncretic Attention for Constrained Image Captioning
Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Deep Captioning with Attention-Based Visual Concept Transfer Mechanism for Enriching Description
Junxuan Zhang, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2019 | Deep multi-path convolutional neural network joint with salient region attention for facial expression recognition
Siyue Xie, Haifeng Hu 0001, Yongbo Wu |
Pattern Recognit. | 2 |
| 2019 | Domain learning joint with semantic adaptation for human action recognition
Junxuan Zhang, Haifeng Hu 0001 |
Pattern Recognit. | 2 |
| 2019 | Face Recognition Using Simultaneous Discriminative Feature and Adaptive Weight Learning Based on Group Sparse RepresentationabstractTo better address face recognition task with occlusions and illumination changes, this letter proposes a novel framework called simultaneous discriminative feature and adaptive weight learning (SDFAWL). Sepecifically, SDFAWL uses a novel unified objective function to simultaneously learn the discriminant features, adaptive feature weights, and classification coefficients. First, a discriminant feature projection is incorporated into group sparse representation model, which reduces the intra-class distance between training samples. Second, we integrate adaptive feature weights into our model to penalize the noisy pixels, which is simultaneously learned by our unified objective function. Besides, we derive an efficient algorithm to optimize the proposed objective function, where the simultaneously learning scheme can encourage obtained parameters combining better and decrease the information loss. Extensive experiments under different conditions including occlusion, random noise, and illumination changes are conducted on the famous Aleix Martinez and Robert Benavente and ExYale B database demonstrate the effectiveness of our model. Lingshuang Du, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2019 | New Discrete-Solution Model for Solving Future Different-Level Linear Inequality and Equality With Robot Manipulator ControlabstractDifferent from general linear inequality or equality, the problem of future different-level linear inequality and equality (FDLLIE) is investigated, which is much more interesting and challenging. In order to solve this difficult FDLLIE, continuous different-level linear inequality and equality (CDLLIE) is first considered. A zeroing equivalency theorem is proposed based on the zeroing neural network method, and then a continuous solution model is, thus, obtained for CDLLIE solving. Furthermore, a new discrete-solution (NDS) model is developed for FDLLIE solving by using a proposed new 7-instant Zhang et al. discretization (ZeaD) formula to discretize the continuous solution model. Meanwhile, theoretical analyses and results are presented to show the excellent properties of the NDS model. Numerical results illustrate the effectiveness and superiority of the NDS model for solving FDLLIE. Furthermore, application experiments for motion planning of robot manipulator are conducted to substantiate the efficacy of the NDS model for FDLLIE solving. Yunong Zhang, Min Yang 0010, Huan-Chang Huang, Mengling Xiao, Haifeng Hu 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Facial Expression Recognition Using Hierarchical Features With Deep Comprehensive Multipatches Aggregation Convolutional Neural NetworksabstractFacial expression recognition (FER) has long been a challenging task in computer vision. In this paper, we propose a novel method, named deep comprehensive multipatches aggregation convolutional neural networks (CNNs), to solve the FER problem. The proposed method is a deep-based framework, which mainly consists of two branches of the CNN. One branch extracts local features from image patches while the other extracts holistic features from the whole expressional image. In the model, local features depict expressional details and holistic features characterize the high-level semantic information of an expression. We aggregate both local and holistic features before making classification. These two types of hierarchical features represent expressions in different scales. Compared with most current methods with single type of feature, the model can represent expressions more comprehensively. Additionally, in the training stage, a novel pooling strategy named expressional transformation-invariant pooling is proposed for handling nuisance variations, such as rotations, noises, etc. Extensive experiments are conducted on the famous the Extended Cohn-Kanade (CK+) dataset and the Japanese Female Facial Expression (JAFFE) database expression datasets, where the recognition results obtained. Siyue Xie, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | Image Captioning With Visual-Semantic Double AttentionabstractIn this article, we propose a novel Visual-Semantic Double Attention (VSDA) model for image captioning. In our approach, VSDA consists of two parts: a modified visual attention model is used to extract sub-region image features, then a new SEmantic Attention (SEA) model is proposed to distill semantic features. Traditional attribute-based models always neglect the distinctive importance of each attribute word and fuse all of them into recurrent neural networks, resulting in abundant irrelevant semantic features. In contrast, at each timestep, our model selects the most relevant word that aligns with current context. In other words, the real power of VSDA lies in the ability of not only leveraging semantic features but also eliminating the influence of irrelevant attribute words to make the semantic guidance more precise. Furthermore, our approach solves the problem that visual attention models cannot boost generating non-visual words. Considering that visual and semantic features are complementary to each other, our model can leverage both of them to strengthen the generations of visual and non-visual words. Extensive experiments are conducted on famous datasets: MS COCO and Flickr30k. The results show that VSDA outperforms other methods and achieves promising performance. Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Photorealistic Face Completion with Semantic Parsing and Face Identity-Preserving FeaturesabstractTremendous progress on deep learning has shown exciting potential for a variety of face completion tasks. However, most learning-based methods are limited to handle general or structure specified face images (e.g., well-aligned faces). In this article, we propose a novel face completion algorithm, called Learning and Preserving Face Completion Network (LP-FCN), which simultaneously parses face images and extracts face identity-preserving (FIP) features. By tackling these two tasks in a mutually boosting way, the LP-FCN can guide an identity preserving inference and ensure pixel faithfulness of completed faces. In addition, we adopt a global discriminator and a local discriminator to distinguish real images from synthesized ones. By training with a combined identity preserving, semantic parsing and adversarial loss, the LP-FCN encourages the completion results to be semantically valid and visually consistent for more complicated image completion tasks. Experiments show that our approach obtains similar visual quality, but achieves better performance on unaligned faces completion and fine detailed synthesis against the state-of-the-art methods. Ruijun Ma 0001, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Pseudo-3D Attention Transfer Network with Content-aware Strategy for Image CaptioningabstractIn this article, we propose a novel Pseudo-3D Attention Transfer network with Content-aware Strategy (P3DAT-CAS) for the image captioning task. Our model is composed of three parts: the Pseudo-3D Attention (P3DA) network, the P3DA-based Transfer (P3DAT) network, and the Content-aware Strategy (CAS). First, we propose P3DA to take full advantage of three-dimensional (3D) information in convolutional feature maps and capture more details. Most existing attention-based models only extract the 2D spatial representation from convolutional feature maps to decide which area should be paid more attention to. However, convolutional feature maps are 3D and different channel features can detect diverse semantic attributes associated with images. P3DA is proposed to combine 2D spatial maps with 1D semantic-channel attributes and generate more informative captions. Second, we design the transfer network to maintain and transfer the key previous attention information. The traditional attention-based approaches only utilize the current attention information to predict words directly, whereas transfer network is able to learn long-term attention dependencies and explore global modeling pattern. Finally, we present CAS to provide a more relevant and distinct caption for each image. The captioning model trained by maximum likelihood estimation may generate the captions that have a weak correlation with image contents, resulting in the cross-modal gap between vision and linguistics. However, CAS is helpful to convey the meaningful visual contents accurately. P3DAT-CAS is evaluated on Flickr30k and MSCOCO, and it achieves very competitive performance among the state-of-the-art models. Jie Wu 0030, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Moving Foreground-Aware Visual Attention and Key Volume Mining for Human Action RecognitionabstractRecently, many deep learning approaches have shown remarkable progress on human action recognition. However, it remains unclear how to extract the useful information in videos since only video-level labels are available in the training phase. To address this limitation, many efforts have been made to improve the performance of action recognition by applying the visual attention mechanism in the deep learning model. In this article, we propose a novel deep model called Moving Foreground Attention (MFA) that enhances the performance of action recognition by guiding the model to focus on the discriminative foreground targets. In our work, MFA detects the moving foreground through a proposed variance-based algorithm. Meanwhile, an unsupervised proposal is utilized to mine the action-related key volumes and generate corresponding correlation scores. Based on these scores, a newly proposed stochastic-out scheme is exploited to train the MFA. Experiment results show that action recognition performance can be significantly improved by using our proposed techniques, and our model achieves state-of-the-art performance on UCF101 and HMDB51. Junxuan Zhang, Haifeng Hu 0001, Xinlong Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Age-Puzzle FaceNet for Cross-Age Face Recognition
Yangjian Huang, Wendong Chen, Haifeng Hu 0001 |
ACCV (6) | 3 |
| 2018 | Multivariate Attention Network for Image Captioning
Haifeng Hu 0001 |
ACCV (6) | 3 |
| 2018 | Perceptual Face Completion using a Local-Global Generative Adversarial NetworkabstractFace completion is one of the most challenging problems, as the reconstruction algorithm should render the missing pixels with semantically plausible contents. Recent methods have achieved promising advances in photorealistic human face synthesis. However, these approaches are limited to deal with general or structure specified faces. In this paper, we propose a Two-Pathway Perceptual Generative Adversarial Network (TPP-GAN) for face completion by perceiving semantic representations from both global structures and local details of a face. We combine a reconstruction network and a perceptual network containing two pathway adversarial networks (local and global) into our framework to efficiently ensure the transfer of the prominent facial features to the occluded parts, which encourages a visually high-quality image completion results. Experimental results well demonstrate that our proposed framework not only generates locally semantic and globally consistent fragments, but also outperforms existing methods on unaligned faces and synthesis of part components. Ruijun Ma 0001, Haifeng Hu 0001 |
ICPR | 2 |
| 2018 | Grouped Multi-Task CNN for Facial Attribute RecognitionabstractThe main goal of facial attribute recognition is to determine various attributes of human faces, e.g. facial expressions, shapes of mouth and nose, headwears, age and race, by extracting features from the images of human faces. Facial attribute recognition has a wide range of potential application, including security surveillance and social networking. The available approaches, however, fail to consider the correlations and heterogeneities between different attributes. This paper proposes that by utilizing these correlations properly, an improvement can be achieved on the recognition of different attributes. Therefore, we propose a facial attribute recognition approach based on the grouping of different facial attribute tasks and a multi-task CNN structure. Our approach can fully utilize the correlations between attributes, and achieve a satisfactory recognition result on a large number of attributes with limited amount of parameters. Several modifications to the traditional architecture have been tested in the paper, and experiments have been conducted to examine the effectiveness of our approach. Chitung Yip, Haifeng Hu 0001 |
ICPR | 2 |
| 2018 | Nuclear Norm Based Superposed Collaborative Representation Classifier for Robust Face Recognition
Yongbo Wu, Haifeng Hu 0001 |
PRCV (3) | 2 |
| 2018 | Exemplar-based Cascaded Stacked Auto-Encoder Networks for robust face alignment
Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2018 | Enhanced multi-dataset transfer learning method for unsupervised person re-identification using co-training strategyabstractThis study proposes progressive unsupervised co‐learning for unsupervised person re‐identification by introducing a co‐training strategy in an iterative training process. The authors’ method adopts an iterative training process to improve transferred models by iterating among clustering, selection, exchange, and fine‐tuning. To solve the problem of transferring representations learned from multiple source datasets, their method utilises multiple convolutional neural network (CNN) models trained on different labelled source datasets by feeding soft labels obtained by clustering on target dataset to each other. The enhanced model can learn more discriminative person representations than the single model trained on multiple datasets. Experimental results on two large‐scale benchmark datasets (i.e. DukeMTMC‐reID and Market‐1501) demonstrate that their method can enhance transferred CNN models by using more source datasets and is competitive to the state‐of‐the‐art methods. Yuqiao Xian, Haifeng Hu 0001 |
IET Comput. Vis. | 2 |
| 2018 | Facial expression recognition using intra-class variation reduced features and manifold regularisation dictionary pair learningabstractA novel framework, named intra‐class variation reduced features‐based manifold regularisation dictionary pair learning model, is presented for solving facial expression recognition (FER) tasks. Since a query face and its corresponding image with intra‐class variations (e.g. identity and illumination) are similar in appearance, the authors generate intra‐class variation reduced features (IVRF) from the difference between a query face image and its corresponding estimated image of each expression class. IVRF can reduce negative influence from the intra‐class variations and make their model robust to intra‐class variations. Furthermore, a manifold regularisation term is incorporated into the dictionary pair learning model, which leads to a smoothly varying sparse representation. Their model fully takes advantage of the geometrical structure of data, which benefits the FER task. The experimental results on two public databases verify the effectiveness and superiority of their method and indicate its promising capability in expression discrimination. Siyue Xie, Haifeng Hu 0001, Ziyu Yin |
IET Comput. Vis. | 2 |
| 2018 | Age-Related Factor Guided Joint Task Modeling Convolutional Neural Network for Cross-Age Face RecognitionabstractCross-age face recognition has remained a popular research topic as most regular facial recognition systems have failed in dealing with facial changes through age. In order to enhance the system's capability of discriminating facial identity features in spite of age changes, this paper proposes a novel deep convolutional network method for cross-age face recognition called age-related factor guided joint task modeling convolutional neural networks, which combines an identity discrimination network with an age discrimination network that shares the same feature layers. By alternatively training the fusion networks and the combined factor model, the cross-age identity features and cross-identity age features can be effectively separated with high inter-class distension and intra-class compactness. Extensive experiments have been performed on the benchmark aging data sets, including MORPH, CACD-VS, and Cross Age LFW. The results have demonstrated the superiority and effectiveness of our model. Haoxi Li, Haifeng Hu 0001, Chitung Yip |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Image Captioning with Affective Guiding and Selective AttentionabstractImage captioning is an increasingly important problem associated with artificial intelligence, computer vision, and natural language processing. Recent works revealed that it is possible for a machine to generate meaningful and accurate sentences for images. However, most existing methods ignore latent emotional information in an image. In this article, we propose a novel image captioning model with Affective Guiding and Selective Attention Mechanism named AG-SAM. In our method, we aim to bridge the affective gap between image captioning and the emotional response elicited by the image. First, we introduce affective components that capture higher-level concepts encoded in images into AG-SAM. Hence, our language model can be adapted to generate sentences that are more passionate and emotive. In addition, a selective gate acting on the attention mechanism controls the degree of how much visual information AG-SAM needs. Experimental results have shown that our model outperforms most existing methods, clearly reflecting an association between images and emotional components that is usually ignored in existing works. Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Image Captioning via Semantic Guidance Attention and Consensus Selection StrategyabstractRecently, a series of attempts have incorporated spatial attention mechanisms into the task of image captioning, which achieves a remarkable improvement in the quality of generative captions. However, the traditional spatial attention mechanism adopts latent and delayed semantic representations to decide which area should be paid more attention to, resulting in inaccurate semantic guidance and the introduction of redundant information. In order to optimize the spatial attention mechanism, we propose the Semantic Guidance Attention (SGA) mechanism in this article. Specifically, SGA utilizes semantic word representations to provide an intuitive semantic guidance that focuses accurately on semantic-related regions. Moreover, we reduce the difficulty of generating fluent sentences by updating the attention information in time. At the same time, the beam search algorithm is widely used to predict words during sequence generation. This algorithm generates a sentence according to the probabilities of words, so it is easy to push out a generic sentence and discard some distinctive captions. In order to overcome this limitation, we design the Consensus Selection (CS) strategy to choose the most descriptive and informative caption, which is selected by the semantic similarity of captions instead of the probabilities of words. The consensus caption is determined by selecting the one with the highest cumulative semantic similarity with respect to the reference captions. Our proposed model (SGA-CS) is validated on Flickr30k and MSCOCO, which shows that SGA-CS outperforms state-of-the-art approaches. To our best knowledge, SGA-CS is the first attempt to jointly produce semantic attention guidance and select descriptive captions for image captioning tasks, achieving one of the best performance ratings among any cross-entropy training methods. Jie Wu 0030, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Joint Head Attribute Classifier and Domain-Specific Refinement Networks for Face AlignmentabstractIn this article, a two-stage refinement network is proposed for facial landmarks detection on unconstrained conditions. Our model can be divided into two modules, namely the Head Attribude Classifier (HAC) module and the Domain-Specific Refinement (DSR) module. Given an input facial image, HAC adopts multi-task learning mechanism to detect the head pose and obtain an initial shape. Based on the obtained head pose, DSR designs three different CNN-based refinement networks trained by specific domain, respectively, and automatically selects the most approximate network for the landmarks refinement. Different from existing two-stage models, HAC combines head pose prediction with facial landmarks estimation to improve the accuracy of head pose prediction, as well as obtaining a robust initial shape. Moreover, an adaptive sub-network training strategy applied in the DSR module can effectively solve the issue of traditional multi-view methods that an improperly selected sub-network may result in alignment failure. The extensive experimental results on two public datasets, AFLW and 300W, confirm the validity of our model. Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Enhanced dictionary pair learning sparse representation model for facial expression classificationabstractFacial expression recognition (FER) is a challenging task in the community of affect analysis and pattern recognition. In this paper, we propose a novel framework, namely Enhanced Dictionary Pair Learning Sparse Representation (EDPLSR), for facial expression recognition. The key idea behind our model is that it jointly learns a synthesis dictionary as well as an analysis dictionary, which require that all coding vectors should be group sparse. Furthermore, inspired by the observation that the geometrical information of the data is discriminative, a manifold regularization term is introduced to obtain smoothly vary sparse representations along the geodesics of data manifold. This is distinctive from most of the existing approaches which fail to consider the geometrical structure of data space. The experimental results demonstrate the effectiveness of our method. Jianquan Gu, Haifeng Hu 0001, Siyue Xie |
ICIP | 2 |
| 2017 | Trajectories-based motion neighborhood feature for human action recognitionabstractRecently, a common and popular method that produces competitive accuracy is to employ dense trajectories to identity human action. However, computing descriptors of dense trajectories may spend lots of time, and many trajectories which belong to the background trajectories may not be useful for the recognition. Moreover, the relationship between trajectories is always ignored. In this paper, we propose a trajectories-based motion neighborhood feature (TMNF) method for action recognition. We first select the trajectories of central particular region at the original video resolution to reduce the computation as well as the background trajectories. A new descriptor, which is referred to as TMNF, is proposed to explore the orientation and motion relationship between different trajectories. Finally, an improved vector of locally aggregated descriptors (IVLAD) method is used to represent videos and linear SVM is applied for classification. Experiments on the YouTube dataset demonstrate that our approach achieves superior performance. Haifeng Hu 0001 |
ICIP | 2 |
| 2017 | Cross-age face recognition using reference coding with kernel direct discriminant analysisabstractWhile face recognition methods have received wide application for decades, the aging process on human face could disable the original method. In this paper, we present a modification on a cross-age face recognition model which utilizes a reference set arranged in time order to eliminate the age difference of input images. In the proposed method, we utilize the identity information of the gallery set to perform discriminative analysis, which can further discriminate between persons in a subspace after the reference coding. Compared with classical cross-age reference coding method, our experiments on Cross-age Celebrity Dataset (CACD) acquire a 35% improvement in recognition rate. Haoshan Zou, Haifeng Hu 0001 |
ICIP | 2 |
| 2017 | Modified Hidden Factor Analysis for Cross-Age Face RecognitionabstractCross-age face recognition has remained a popular research topic because the sophisticated facial change across age disables regular face recognition systems. Widely applied in age-related tasks, the hidden factor analysis (HFA) model decomposes face feature into independent age and identity factors. However, the hypothesis that the identity and age factors are independent is not in accordance with the fact that aging has different appearance changes on different people's faces. To address this problem, this letter presents a novel method for cross-age face recognition, called age-identity modified HFA, which exploits a new latent factor modeled as a linear combination with the age factor and the identity factor. Hence, the cross-age identity information can be extracted and separated preferably. A maximum likelihood strategy is proposed to judge which gallery face has the same identity with the probe image, while we do not need to know what the probe identity is. Extensive experiments are performed on the benchmark aging datasets MORPH and FG-Net, and the recognition rate of our method outperforms HFA by 10.4% and 1.15%, respectively. Haoxi Li, Haoshan Zou, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | Comments on "Iterative Re-Constrained Group Sparse Face Recognition With Adaptive Weights Learning"abstractIn the above paper [1] the authors designed an IRGSC architecture for face recognition with pixel corruption and occlusion. We have discovered that the calculation of the weights vector s in the original paper is flawed. In this comment we demonstrate the correct calculations, and further analysed some of the results in the paper. Haoxi Li, Haifeng Hu 0001, Chitung Yip |
IEEE Trans. Image Process. | 2 |
| 2016 | Patch-based alignment-free generic sparse representation for pose-robust face recognitionabstractSparse representation based classification method has been successfully applied to face recognition in recent years. However, it is still a problem in the scenario of pose variation in face recognition with single sample per person. In this paper, we propose a novel alignment-free model, called Gabor-based Partial Face Sparse Representation (GPFSR), to solve the problem of pose variation in face recognition with single sample per person by using partial face. In our method, we firstly locate five facial landmarks in different images. Then partial face is obtained, which is used to construct Gabor-based local dictionary and compute the weights of each patch. Our classification principle is based on sparse representation. The experimental results on the Multi-PIE and FERET show that GPFSR is robust to pose variation in FR with single sample per person. Jianquan Gu, Haifeng Hu 0001, Haoxi Li, Weipeng Hu |
ICIP | 2 |
| 2015 | Illumination invariant face recognition based on dual-tree complex wavelet transformabstractThis study presents a new dual‐tree complex wavelet transform (DT‐CWT)‐based illumination normalisation approach for face recognition under varying lighting conditions. The method consists of three steps. First, the DT‐CWT‐based edge detection method is proposed which can obtain estimation for facial feature edges in different directionality and resolution level. Second, the DT‐CWT‐based denoising model is employed to obtain the multi‐scale illumination invariant structures in the logarithm domain. Finally, by combining the obtained illumination invariant features and edge estimation information, the enhanced facial features are obtained which have more discriminating power for variable lighting face recognition. The effectiveness of the method is validated in comparative performance against many classical illumination compensation methods using the YaleB database and the CMU PIE database. Haifeng Hu 0001 |
IET Comput. Vis. | 1 |
| 2015 | Sparse Discriminative Multimanifold Grassmannian Analysis for Face Recognition With Image SetsabstractWe propose an efficient and robust solution, called sparse discriminative multimanifold Grassmannian analysis (SDMMGA), for face recognition based on image set (FRIS), where each set contains face images belonging to the same subject and typically covering large variations. In our work, linearity constrained hierarchical agglomerative clustering (LC-HAC) method is first employed to partition each image set into several local linear models (LLMs), each depicted as a point on the Grassmannian manifold using positive definite Gaussian kernel function. In contrast to the standard discriminative learning algorithms that assume that all data are sampled from one single manifold and only one projection is derived for feature extraction, we model all the LLMs of each person as a manifold and present SDMMGA model to seek multiple projection matrices, which can uncover the geometrical information of different manifolds. Aiming to better separate manifold margins in the low-dimensional feature space, we introduce the ℓ1and ℓ2norms penalty in the SDMMGA objective function. An efficient regression method is presented for finding the most discriminative features. Comprehensive experiments on three standard data sets show that our method consistently outperforms the state of the art. Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Multiview Gait Recognition Based on Patch Distribution Features and Uncorrelated Multilinear Sparse Local Discriminant Canonical Correlation AnalysisabstractIt is well recognized that gait is an important biometric feature to identify a person at a distance, such as in video surveillance application. However, in reality, a change of viewing angle causes a significant challenge for gait recognition. In this paper, a novel approach is proposed for multiview gait recognition with the view angle of a probe gait sequence unknown. We formulate a new patch distribution feature based classification framework to estimate the view angle of each probe gait sequence. In this method, each gait energy image is represented as a set of dual-tree complex wavelet transform (DTCWT) features derived from different scales and orientations together with the x-y coordinates. Then, a two-stage Gaussian mixture model is presented that can represent each DTCWT based gait feature with a set of patch distribution parameters. A simple nearest-neighbor classifier is employed for view classification. To measure the similarity of gait sequences, we also propose a sparse local discriminant canonical correlation analysis algorithm to model the correlation of gait features from different views and use the correlation strength as similarity measure. An uncorrelated multilinear SLDCCA (UMSLDCCA) framework is further presented that aims to extract uncorrelated discriminative features directly from multidimensional gait features through solving a tensor-to-vector projection. The solution consists of sequential iterative processes based on the alternating projection method. Different from existing approaches, UMSLDCCA considers the spatial structure information within each gait sample and local geometry information among multiple gait samples. Moreover, our approach does not need explicit reconstruction and is robust against feature noise. Extensive experiments have been performed on two benchmark gait databases. The results demonstrate that our method outperforms the state-of-the-art methods in terms of accuracy and efficiency. Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Face Recognition With Image Sets Using Locally Grassmannian Discriminant AnalysisabstractWe propose an efficient and robust solution, called locally Grassmannian discriminant analysis (LGDA), for face recognition with image set, where each set contains images belonging to the same subject and typically covering large variations. In our work, by modeling each image set as a nonlinear manifold, linearity-constrained nearest neighborhood clustering is first presented for expressing a manifold by a collection of local linear models (LLMs), each depicted by a subspace. With a proper kernel function defined by canonical correlation between the subspaces, the obtained LLMs can be projected into low-dimensional LGDA embedding space using a set of locally linear transformations. Different from traditional discriminant analysis approaches, LGDA is for multiclass nonlinear discrimination and it can maximize discriminatory power while simultaneously promoting consistency between the multiple local representations of single class objects. A novel accelerated proximal gradient-based learning algorithm is proposed for finding the optimal set of local linear bases. To measure the similarities between the face image sets, three distance criterions are presented, which integrate the distance between the pairs of low-dimensional Grassmannian points from one of the involved manifolds. Comprehensive experiments on the UCSD/Honda, CMU MoBo, and YouTube Celebrities face data sets show that our method consistently outperforms the state of the art. Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Spatially consistent exemplar-based clusteringabstractExemplar-based clustering has drawn much attention in recent years as it produces state-of-the-art results on many practical clustering problems. However, spatial information is missed in the exemplar-based clustering methods, resulting in difficulties in some applications, for example in the image segmentation problem. In this paper, we investigate the issue of integrating spatial information into the exemplar-based clustering through the Markov random field formulation. Two algorithms are proposed to achieve this aim. First, based on the min-sum loopy belief propagation algorithm, a spatially consistent affinity propagation algorithm is proposed. Second, by showing the spatially consistent exemplar-based clustering energy function satisfies the regular property, an efficient minimal s-t graph cut based convergent algorithm is proposed. Experimental results on the image segmentation problem show that the spatially consistent exemplar-based clustering achieves better results than other methods. Pei Chen 0001, Yuan He 0001, Jun Sun 0004, Haifeng Hu 0001 |
ICME | 5 |
| 2013 | Enhanced Gabor Feature Based Classification Using a Regularized Locally Tensor Discriminant Model for Multiview Gait RecognitionabstractThis paper presents a novel multiview gait recognition method that combines the enhanced Gabor (EG) representation of the gait energy image and the regularized local tensor discriminant analysis (RLTDA) method. EG first derives desirable gait features characterized by spatial frequency, spatial locality, and orientation selectivity to cope with the variations due to surface, shoe types, clothing, carrying conditions, and so on. Unlike traditional Gabor transformation, which does not consider the structural characteristics of the gait features, our representation method not only considers the statistical property of the input features but also adopts a nonlinear mapping to emphasize those important feature points. The dimensionality of the derivation of EG gait feature is further reduced by using RLTDA, which directly obtains a set of locally optimal tensor eigenvectors and can capture nonlinear manifolds of gait features that exhibit appearance changes due to variable viewing angles. An aggregation scheme is adopted to combine the complementary information from differently RLTDA recognizers at the matching score level. The proposed method achieves the best average Rank-1 recognition rates for multiview gait recognition based on image sequences from the USF HumanID gait challenge database and the CASIA gait database. Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Multiscale illumination normalization for face recognition using dual-tree complex wavelet transform in logarithm domain
Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2011 | Augmented DT-CWT feature based classification using Regularized Neighborhood Projection Discriminant Analysis for face recognition
Haifeng Hu 0001 |
Pattern Recognit. | 1 |
| 2011 | Variable lighting face recognition using discrete wavelet transform
Haifeng Hu 0001 |
Pattern Recognit. Lett. | 1 |
| 2009 | Direct kernel neighborhood discriminant analysis for face recognition
Haifeng Hu 0001, Zhengming Ma |
Pattern Recognit. Lett. | 1 |
| 2008 | ICA-based neighborhood preserving analysis for face recognition
Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2008 | Orthogonal neighborhood preserving discriminant analysis for face recognition
Haifeng Hu 0001 |
Pattern Recognit. | 1 |
| 2007 | 3D Reconstruction Approach Based on Neural Network
Haifeng Hu 0001, Zhi Yang 0004 |
ISNN (2) | 1 |
| 2006 | Camera Calibration and 3D Reconstruction Using RBF Network in Stereovision System
Haifeng Hu 0001 |
ISNN (2) | 1 |