VLDB 2026 Research / reviewers in the wild / expert
Jiaxiang Wang 0001
dblp:183/1758-1 · also Jia-Xiang Wang 0001
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-3059-798XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional intervention attention network for audio-visual matching
Jiaxiang Wang 0001, Aihua Zheng, Dequan Li, Chenglong Li 0002, Wenjuan Cheng, Ran He 0001 |
Pattern Recognit. | 1 |
| 2025 | Feature Decoupling with Modality Modulation for Multimodal Sentiment Analysis
Jiaxiang Wang 0001, Aihua Zheng, Wenjuan Cheng, Xiaofei Sheng |
ICIG (2) | 2 |
| 2025 | Modality Modulation with Adaptive Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims at extracting effective information from different modalities such as text, audio, and visual to infer the speaker’s sentiment state. Due to the high heterogeneity across modalities, most existing approaches decouple modalities into specific and invariant features, which can capture effective cross-modal representations to some extent. However, in multimodal tasks, different modalities exhibit varying strengths and weaknesses, with strong modalities dominating the overall optimization direction of the network, leading to under-optimization of weak modalities. To address the modality imbalance issue, we propose a Modality Modulation Adaptive Fusion Network (MMAFNet) to optimize the learning of valuable information from each modality. Specifically, for modality-specific features, we design a specific feature gradient modulation strategy to stimulate the weak modalities learning and adaptively modulate the corresponding gradients by measuring different importance to better optimize each modality. For modality-invariant features, in terms of the distance between different modalities, we propose an invariant feature parameter reset strategy that prevents overfitting irrelevant information while enhancing the feature extraction capability of weaker modalities. Finally, we incorporate an adaptive fusion module to combine modality-specific and invariant features based on their respective weights. Overall, we analyze the characteristics of various features and propose modality modulation strategies that mitigate modality imbalance. Extensive experiments on two multimodal sentiment analysis datasets, demonstrate the superior performance of our method. Aihua Zheng, Jiaxiang Wang 0001, Xiaofei Sheng, Wenjuan Cheng |
IJCNN | 3 |
| 2025 | Modality Modulation with Adaptive Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims at extracting effective information from different modalities such as text, audio, and visual to infer the speaker’s sentiment state. Due to the high heterogeneity across modalities, most existing approaches decouple modalities into specific and invariant features, which can capture effective cross-modal representations to some extent. However, in multimodal tasks, different modalities exhibit varying strengths and weaknesses, with strong modalities dominating the overall optimization direction of the network, leading to under-optimization of weak modalities. To address the modality imbalance issue, we propose a Modality Modulation Adaptive Fusion Network (MMAFNet) to optimize the learning of valuable information from each modality. Specifically, for modality-specific features, we design a specific feature gradient modulation strategy to stimulate the weak modalities learning and adaptively modulate the corresponding gradients by measuring different importance to better optimize each modality. For modality-invariant features, in terms of the distance between different modalities, we propose an invariant feature parameter reset strategy that prevents overfitting irrelevant information while enhancing the feature extraction capability of weaker modalities. Finally, we incorporate an adaptive fusion module to combine modality-specific and invariant features based on their respective weights. Overall, we analyze the characteristics of various features and propose modality modulation strategies that mitigate modality imbalance. Extensive experiments on two multimodal sentiment analysis datasets, demonstrate the superior performance of our method. Aihua Zheng, Jiaxiang Wang 0001, Xiaofei Sheng, Wenjuan Cheng |
IJCNN | 3 |
| 2025 | Prompt-Based Cross-Modal Feature Alignment for Weakly Supervised IFERabstractInfrared Facial Expression Recognition (IFER) encounters challenges in data acquisition and annotation under low-light conditions, making fully supervised training difficult. Although pre-trained Vision-Language Models (VLMs) can enhance generalization for downstream tasks, their insufficient attention modeling in cross-domain scenarios leads to ineffective local semantic correlation. To address this, we propose a Prompt-based Cross-modal feature Alignment (PCA) method that improves weakly supervised IFER performance by leveraging RGB facial expression data. The PCA framework comprises two key components: (1) a Cross-modal Prompt Transfer (CPT) strategy that integrates category-specific information to distinguish expressions, and (2) an Image-Guided Alignment (IGA) module that achieves feature alignment using dual-domain feature banks. Experimental results on two benchmark datasets demonstrate that our method significantly outperforms current state-of-the-art approaches, confirming its effectiveness and superiority. Hanqin Shi, Xiaofeng Kang, Jiaxiang Wang 0001, Aihua Zheng, Wenjuan Cheng |
IEEE Signal Process. Lett. | 3 |
| 2025 | Adaptive Interaction and Correction Attention Network for Audio-Visual MatchingabstractAudio-visual matching techniques aim to recognize and match information across different identities by learning a similarity metric across modalities. However, modal differences arise from insufficient cross-modal correlations and noise interference, which substantially hinder the performance of traditional deep metric learning methods in audio-visual matching tasks. To address the modal differences issue, we propose a novel Adaptive Interactive and Correction Attention Network (AICANet). This network efficiently captures deep information connections, generating modality-consistent feature embeddings within a unified metric framework. The core of AICANet is its two-pronged approach to reducing modal differences. First, we propose the Adaptive Interactive Attention (AIA) module, which flexibly establishes associations among cross-modal local features using dynamically generated pseudo-labels. Second, we propose the Adaptive Correction Attention (ACA) mechanism, which employs an adaptive threshold to de-interference effectively and accurately adjust the representation of local feature associations. Notably, the ACA mechanism is suitable for both intra-modal and inter-modal refined attention correction. Additionally, we design a relative distance stretching metric loss (LRDSM), which reinforces the similarity invariance of feature embeddings in a uniform space and enhances matching accuracy. Extensive tests on the VoxCeleb and VoxCeleb2 datasets demonstrate that AICANet outperforms leading existing algorithms across several evaluation metrics, validating its superior performance. The codes can be found at https://github.com/w1018979952/AICANet. Jiaxiang Wang 0001, Aihua Zheng, Lei Liu 0049, Chenglong Li 0002, Ran He 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Attention-Guided Contrastive Masked Autoencoders for Self-supervised Cross-Modal Biometric Matching
Jiaxiang Wang 0001, Hanqin Shi, Zhenda Yu, Yin Lin |
ICDF2C (2) | 2 |
| 2024 | Public-Private Attributes-Based Variational Adversarial Network for Audio-Visual Cross-Modal MatchingabstractExisting audio-visual cross-modal matching methods focus on mitigating cross-modal heterogeneity but ignore the impact of intra-class discrepancy of the same identity in different scenarios, which might greatly limit the matching performance. To simultaneously handle both problems of intra-class discrepancy and cross-modal heterogeneity, we propose a novel public-private attributes-based variational adversarial network (P2VANet), which captures the consistency within and between classes, for audio-visual cross-modal matching. In particular,P2VANet first uses a variational auto-encoder, which captures the inherent global information in diverse scenarios from the hidden variable through reconstruction, to reduce the intra-class discrepancy. Then it integrates a public attributes guidance module to capture the consistency of audio and visual by supervision of the common high-level semantic information to mitigate cross-modal heterogeneity. In addition,P2VANet designs private attributes embedding module to enhance the discriminative features inherent in each class to decrease inter-class similarity. Extensive experiments on audio-visual cross-modal matching demonstrate the effectiveness of the proposed approach compared with the state-of-the-art methods. Aihua Zheng, Jiaxiang Wang 0001, Chao Tang 0002, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Attribute-Guided Cross-Modal Interaction and Enhancement for Audio-Visual MatchingabstractAudio-visual matching is an essential task that measures the correlation between audio clips and visual images. However, current methods rely solely on the joint embedding of global features from audio clips and face image pairs to learn semantic correlations. This approach overlooks the importance of high-confidence correlations and discrepancies of local subtle features, which are crucial for cross-modal matching. To address this issue, we propose a novel Attribute-guided Cross-modal Interaction and Enhancement Network (ACIENet), which employs multiple attributes to explore the associations of different key local subtle features. The ACIENet contains two novel modules: the Attribute-guided Interaction (AGI) module and the Attribute-guided Enhancement (AGE) module. The AGI module employs global feature alignment similarity to guide cross-modal local feature interactions, which enhances cross-modal association features for the same identity and expands cross-modal distinctive features for different identities. Additionally, the interactive features and original features are fused to ensure intra-class discriminability and inter-class correspondence. The AGE module captures subtle attribute-related features by using an attribute-driven network, thereby enhancing discrimination at the attribute level. Specifically, it strengthens the combined attribute-related features of gender and nationality. To prevent interference between multiple attribute features, we design a multi-attribute learning network as a parallel framework. Experiments conducted on a public benchmark dataset demonstrate the efficacy of the ACIENet method in different scenarios. Code and models are available at https://github.com/w1018979952/ACIENet. Jiaxiang Wang 0001, Aihua Zheng, Yan Yan 0002, Ran He 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Diverse features discovery transformer for pedestrian attribute recognition
Aihua Zheng, Jiaxiang Wang 0001, Huaibo Huang, Ran He 0001, Amir Hussain 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Looking and Hearing Into Details: Dual-Enhanced Siamese Adversarial Network for Audio-Visual MatchingabstractAudio-visual cross-modal matching aims to explore the intrinsic correspondence between face images and audio clips. Existing methods usually focus on the salient features of identities between visual images and voice clips, while neglecting their subtle differences, which are crucial to distinguishing cross-modal samples. To deal with this problem, we propose a novel Dual-enhanced Siamese Adversarial Network (DSANet), which pursues the adversarial dual enhancement to highlight both salient and subtle features for robust audio-visual cross-modal matching. First, we designed a dual enhancement mechanism to enhance potential subtle features by randomly selecting a region feature for salient feature suppression, while enhancing salient features in the corresponding region to ensure the global discriminative ability. Second, to establish the correlation of subtle features in the process of eliminating cross-modal heterogeneity, we design a siamese adversarial structure to perform modal heterogeneity elimination for both enhanced salient and subtle features in a parallel manner. Moreover, we propose an adaptive masked cross-entropy loss to force the network to focus on the feature differences among hard classes. Experiments on public benchmark datasets validate the effectiveness of the proposed algorithm. Jiaxiang Wang 0001, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Prior-Guided Multi-scale Fusion Transformer for Face Attribute Recognition
Shaoheng Song, Huaibo Huang, Jiaxiang Wang 0001, Aihua Zheng, Ran He 0001 |
PRCV (1) | 3 |
| 2021 | 3D Shape Estimation With an Enhanced Sparse Representation ApproachabstractIn this paper, an enhanced sparse representation approach is proposed to estimate the 3D shapes of objects in 2D image sequences. In the proposed method, the unknown 3D shape is estimated via a two-stage scheme, namely the main 3D shape estimation stage and the compensatory 3D shape estimation stage. Moreover, a reweighted sparse representation model is constructed to extract the shape bases for each estimation stage. In the sparse model, a reweighted constraint is enforced to enhance the coefficient sparsity of the shape bases. Experimental results on the well-known CMU image sequences demonstrate the effectiveness and feasibility of the proposed approach. Jiaxiang Wang 0001, Zhigang Zeng, Kin-Man Lam 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | A CSF-Based CNR Approach for Small-Size Image SequencesabstractFor non-rigid structure from motion (NRSFM), the performance of most traditional approaches may decrease significantly when the frame number of the image sequence is relatively small. In this letter, a column space fitting (CSF) based consensus of non-rigid (CNR) reconstruction approach is proposed to deal with the 3D structure estimation problem for small-size image sequences. In the proposed method, a set of trajectory groups are first extracted by utilizing the distance weight of the pairwise points. In order to improve the estimation accuracy, an adaptive rank selection strategy is designed to choose the approximately optimal rank parameter. Corresponding to the trajectory group, the z-coordinates of the observation matrix are estimated by the CSF algorithm due to its good performance. After obtaining the outputs of the CSF-based weak estimators, the final 3D shape is derived by combining the outputs via the alternating directional method of multipliers. Experimental results on several widely used image sequences demonstrate the effectiveness and feasibility of the proposed algorithm. Jiaxiang Wang 0001, Xia Chen 0008, Kin-Man Lam 0001, Zhigang Zeng |
IEEE Signal Process. Lett. | 1 |