Jie Wu 0033

dblp:181/2833-33 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-8896-2110ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AGPL-KEM : Attribute-guided prompt learning with knowledge experts mixture for few-shot remote sensing image classification
Chunlei Wu, Congzheng Zhu, Qinfu Xu, Yongzhen Zhang, Leiquan Wang, Jie Wu 0033
Knowl. Based Syst.7
2025 Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution Constraints
abstract
Sentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also neglected the incongruity in semantic distribution among modalities. In light of this, we introduce a novel Hierarchical Correlation Modeling Network (HCMNet), which enhances the multimodal sentiment analysis by exploring both the atomic-level correlations based on dynamic attention reasoning and the composition-level correlations through topological graph reasoning. In addition, we also alleviate the impact of distributional inconsistencies between modalities from both atomic-level and composition-level perspectives. Specifically, we first design an atomic-level contrastive loss that constrains the semantic distribution across modalities to mitigate the atomic-level inconsistency. Then, we design a graph optimal transport module that integrates transport flows with different graphs to constrain the composition-level semantic distribution, thus reducing the inconsistency of compositional nodes. Experiments on three public benchmark datasets have demonstrated the superiority of the proposed model over the state-of-the-art methods.
Qinfu Xu, Chunlei Wu, Leiquan Wang, Shaozu Yuan, Jie Wu 0033, Jing Lu 0013, Hengyang Zhou
AAAI6
2025 Multiple Feature Refining Network for Visual Emotion Distribution Learning
abstract
The significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle emotion awareness. To learn the distribution of emotions involved in images, most previous works learn coarse semantic knowledge with unbiased filtering. Consequently, they focus on the entire scene and suffer from the redundancy of semantic-irrelevant information, which diminishes the affective coherence, impeding the comprehension of emotional attributes within the treated features. In light of this, we reanalyze from the perspective of information filtering and propose a novel method called Multiple Feature Refining Network (MFRN). To minimize low-level feature redundancy, we design a wavelet-based separated frequency modeling, named Spectral Mixer, to learn invariant representations and enhance emotion saliency in low-level image features. At the higher semantic level, we design a Semantic Graph Prompt Learning for emotional semantic filtering, ensuring the purity of emotional information and providing the model with richer content semantics. Experiments conducted on three commonly used datasets have demonstrated the superiority of our MFRN model over cutting-edge methods.
Qinfu Xu, Shaozu Yuan, Jie Wu 0033, Leiquan Wang, Chunlei Wu
AAAI4
2025 Recursive bidirectional cross-modal reasoning network for vision-and-language navigation
Jie Wu 0033, Chunlei Wu, Xiuxuan Shen, Fengjiang Wu, Leiquan Wang
Expert Syst. Appl.1
2025 Adaptive Cross-Modal Experts Network with Uncertainty-Driven Fusion for Vision-Language Navigation
Jie Wu 0033, Chunlei Wu, Xiuxuan Shen, Leiquan Wang
Knowl. Based Syst.1
2025 Learning contrastive semantic decomposition for visual grounding
Jie Wu 0033, Chunlei Wu, Qinfu Xu, Faming Gong
Neural Networks1
2024 Memory Self-Calibrated Network for Visual Grounding
abstract
Visual Grounding (VG) aims to locate the most relevant object or region in an image according to a natural language query. Existing methods in VG utilize fixed image and text representations to capture cross-modal semantic consistency, which limits the flexibility in adjusting image representations according to diverse textual information and hinders performance. To handle this limitation, we propose a novel Memory Self-Calibrated Network (MSCN) by dynamically refining image representations based on the query, thereby improving the semantic consistency between texts and images for visual grounding. Specifically, we introduce two modules: Semantic Relevance Filtering Module (SRFM) and Adaptive Memory Fusion Module (AMFM), to explicitly model the relationship between image and text. SRFM focuses on filtering out image information that is irrelevant to the query, while AMFM adaptively fuses text-related representations with initial image features to enhance the understanding ability of the MSCN model. Comprehensive experiments on three datasets demonstrate the superiority of our method compared to existing approaches.
Jie Wu 0033, Chunlei Wu, Xiuxuan Shen, Leiquan Wang
ICASSP1
2024 Improving visual grounding with multi-scale discrepancy information and centralized-transformer
Jie Wu 0033, Chunlei Wu, Fuyan Wang, Leiquan Wang
Expert Syst. Appl.1
2024 Vertical-horizontal latent space with iterative memory review network for multi-class anomaly detection
Chunlei Wu, Jie Wu 0033, Huan Zhang 0015, Leiquan Wang
Knowl. Based Syst.3
2024 Towards visual emotion analysis via Multi-Perspective Prompt Learning with Residual-Enhanced Adapter
Chunlei Wu, Qinfu Xu, Shaozu Yuan, Jie Wu 0033, Leiquan Wang
Knowl. Based Syst.5
2024 Multistage Synergistic Aggregation Network for Remote Sensing Visual Grounding
abstract
Visual Grounding has a broad application prospect in the field of remote sensing. Current state-of-the-art methods predominantly are based on the transformer architecture, utilizing multi-head self-attention in multi-modal encoders to integrate visual and textual features. However, they typically rely on a single fusion approach, which may limit the model’s capacity to learn intricate correlations between textual semantics and visual information. Moreover, they did not establish a direct dependency between features and bounding box representations, thereby restricting the fusion features to conventional object detection paradigm. Consequently, the interactions between regression results and encoded features are constrained. To address these limitations, a generative paradigm is harnessed to directly generate discrete coordinates sequence in an auto-regressive manner, which explores the interaction between direct regression features and encoded multi-modal features. Meanwhile, a novel multi-stage synergistic aggregation module is proposed to facilitate the acquisition of multi-modal features at multiple scales by effectively aggregating visual and textual contexts, enhancing the overall performance. In this work, we validate our framework on the DIOR-RSVG dataset and conduct a comparative analysis with existing methods, achieving a noteworthy improvement in accuracy. The proposed approach presents a promising direction for advancing visual grounding techniques in the context of remote sensing applications. The related code and weights are available at https://github.com/waynamigo/MSAM.
Fuyan Wang, Chunlei Wu, Jie Wu 0033, Leiquan Wang, Canwei Li
IEEE Geosci. Remote. Sens. Lett.3
2023 Nested Attention Network with Graph Filtering for Visual Question and Answering
abstract
Recently, Visual Question Answering(VQA), which is required to generate the answer by understanding both visual and textual content, has attracted considerable research interest. Most existing works extract visual features with the CNN network and learn its feature embedding with an attention mechanism. However, this mechanism may ignore the interaction between entities in the image, which has a fuzzy impact on the answer generation. To better explore the relationship between different entities in the image, a novel Nested Attention Network with Graph Filtering (NANGF) is proposed. It composes of two novel designed modules: a graph filtering mechanism to mine more precise visual semantics and avoid understanding deviation and nested attention to effectively guide the integration of visual features and question features. Extensive experiments conducted on the VQA2.0 datasets demonstrate the effectiveness of the proposed method.
Jing Lu 0013, Chunlei Wu, Leiquan Wang, Shaozu Yuan, Jie Wu 0033
ICASSP5
2023 Multi-view inter-modality representation with progressive fusion for image-text matching
Jie Wu 0033, Leiquan Wang, Chenglizhao Chen, Jing Lu 0013, Chunlei Wu
Neurocomputing1
2023 Dynamic Pruning of Regions for Image-Sentence Matching
abstract
Image–sentence matching is becoming increasingly essential in the integrated understanding of vision and language. Prior approaches apply a pre-trained detection model to extract region features and explore fine-grained relationships between image and sentence by aggregating the similarities of all region–word pairs. However, all images are represented by the same number of regions, regardless of their respective semantic complexity, which results in a large number of redundant regions interfering with semantic inference and bringing additional computational burden. To address the lack of flexibility in image representation and information redundancy, a novel method named Dynamic Pruning of Regions for Image–Sentence Matching (DPRM) is proposed to efficiently capture relationships between text and image. In particular, a dynamic region pruning module is presented to dynamically select the appropriate number of regions according to the semantic complexity of each image, thus pruning redundant regions and reducing superfluous computations. Moreover, an inter-modality refinement module is designed to refine the fine-grained relationships of region–word pairs by retaining meaningful interaction features and suppressing interference from redundant alignments, which learns the more accurate semantic correspondences. Extensive experiments on MSCOCO and Flickr30K datasets prove the superiority of DPRM compared with previous approaches.
Jie Wu 0033, Weifeng Liu 0001, Leiquan Wang, Xiuxuan Shen, Chunlei Wu
Signal Process. Image Commun.1
2023 A Novel Deep Learning Model for Medical Report Generation by Inter-Intra Information Calibration
abstract
Automatic generation of medical reports can provide diagnostic assistance to doctors and reduce their workload. To improve the quality of the generated medical reports, injecting auxiliary information through knowledge graphs or templates into the model is widely adopted in previous methods. However, they suffer from two problems: 1) The injected external information is limited in amount and difficult to adequately meet the information needs of medical report generation in content. 2) The injected external information increases the complexity of model and is hard to be reasonably integrated into the generation process of medical reports. Therefore, we propose an Information Calibrated Transformer (ICT) to address the above issues. First, we design a Precursor-information Enhancement Module (PEM), which can effectively extract numerous inter-intra report features from the datasets as the auxiliary information without external injection. And the auxiliary information can be dynamically updated with the training process. Secondly, a combination mode, which consists of PEM and our proposed Information Calibration Attention Module (ICA), is designed and embedded into ICT. In this method, the auxiliary information extracted from PEM is flexibly injected into ICT and the increment of model parameters is small. The comprehensive evaluations validate that the ICT is not only superior to previous methods in the X-Ray datasets, IU-X-Ray and MIMIC-CXR, but also successfully be extended to a CT COVID-19 dataset COV-CTR.
Junsan Zhang, Xiuxuan Shen, Shaohua Wan 0001, Sotirios K. Goudos, Jie Wu 0033, Weishan Zhang
IEEE J. Biomed. Health Informatics5
2022 Region Reinforcement Network With Topic Constraint for Image-Text Matching
abstract
Image and sentence matching has attracted increasing attention since it is associated with two important modalities of vision and language. Previous methods aim to find the latent correspondences between image regions and words by aggregating the similarities of the region-word pairs. However, these approaches consider little about the relationships of diverse regions in the image and treat the similarities of all region-word pairs equally. Moreover, focusing on fine-grained alignment overly, the true meaning of the original image will be likely distorted. In this paper, a novel Region Reinforcement Network with Topic Constraint (RRTC) is proposed to explore the correspondences between images and texts. Specifically, the region reinforcement network is built to infer fine-grained correspondence by considering the relationships of regions and re-assigning region-word similarities. Meanwhile, the topic constraint module is presented to summarize the central theme of images, which constrains the original image deviation. Extensive experimental results on MSCOCO and Flickr30k datasets verify the effectiveness of our proposed RRTC.
Jie Wu 0033, Chunlei Wu, Jing Lu 0013, Leiquan Wang, Xue-rong Cui
IEEE Trans. Circuits Syst. Video Technol.1
2021 Dual-View Semantic Inference Network for image-text matching
Chunlei Wu, Jie Wu 0033, Haiwen Cao, Leiquan Wang
Neurocomputing2