VLDB 2026 Research / reviewers in the wild / expert
Xue Wang 0011
dblp:39/2811-11
· DBLP profile ↗
18ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0001-6674-8140ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EFNet: enhanced activation and fine-grained information auxiliary network for salient object detection
Xue Wang 0011, Wenbi Ma |
Appl. Intell. | 3 |
| 2026 | Heterogeneous graph knowledge tracing with novel contrastive view and meta-path contexts
Jie Liu 0056, Xue Wang 0011 |
Expert Syst. Appl. | 3 |
| 2026 | Cross-modality masked autoencoder for infrared and visible image fusion
Cong Bi, Wenhua Qian, Qiuhan Shao, Jinde Cao, Xue Wang 0011, Kaixiang Yan |
Pattern Recognit. | 5 |
| 2025 | Residual Prior-driven Frequency-aware Network for Image FusionabstractImage fusion aims to integrate complementary information across modalities to generate high-quality fused images, thereby enhancing the performance of high-level vision tasks. While global spatial modeling mechanisms show promising results, constructing long-range feature dependencies in the spatial domain incurs substantial computational costs. Additionally, the absence of ground-truth exacerbates the difficulty of capturing complementary features effectively. To tackle these challenges, we propose a Residual Prior-driven Frequency-aware Network, termed as RPFNet. Specifically, RPFNet employs a dual-branch feature extraction framework: the Residual Prior Module (RPM) extracts modality-specific difference information from residual maps, thereby providing complementary priors for fusion; the Frequency Domain Fusion Module (FDFM) achieves efficient global feature modeling and integration through frequency-domain convolution. Additionally, the Cross Promotion Module (CPM) enhances the synergistic perception of local details and global structures through bidirectional feature interaction. During training, we incorporate an auxiliary decoder and saliency structure loss to strengthen the model's sensitivity to modality-specific differences. Furthermore, a combination of adaptive weight-based frequency contrastive loss and SSIM loss effectively constrains the solution space, facilitating the joint capture of local details and global features while ensuring the retention of complementary information. Extensive experiments validate the fusion performance of RPFNet, which effectively integrates discriminative features, enhances texture details and salient objects, and can effectively facilitate the deployment of the high-level vision task. The source code can be available at https://github.com/wang-x-1997/RPFNet. Xue Wang 0011, Wenhua Qian, Peng Liu 0056, Runzhuo Ma |
ACM Multimedia | 2 |
| 2025 | FSGFuse: Feature synergy-guided multi-modality image fusionabstractThe purpose of infrared and visible image fusion (IVIF) is to integrate the complementary properties of the two modalities to obtain a fused image with a more comprehensive information representation. However, existing multi-modal image fusion (MMIF) methods mainly rely on auto-encoders to extract source features and achieve feature complementarity through fusion mechanisms. These methods fail to fully utilize the potential of the encoder in feature decoupling and information integration , resulting in difficulties in effectively capturing and fusing cross-modal features, which in turn introduces redundant information. To tackle the challenge, this study proposes a Feature Synergy-Guided Fusion (FSGFuse) network, which achieves IVIF by synergizing complementary features to preserve texture details and acquire thermal target information . In the pre-training stage, the feature synergism loss is utilized to guide the double-branch encoder to perceive the consistency and difference of cross-modal features, which in turn motivates the model to capture the more discriminatory feature representations in the source features. In the fusion stage, a Feature Complementary Fusion (FCF) module is designed to guide feature fusion . This module not only integrates complementary features, but also captures long-range contextual relationships between features, which facilitates cross-modal interaction while enriching the semantic information. Experiments on publicly available benchmark datasets show that the fusion performance of FSGFuse significantly outperforms existing state-of-the-art methods and positively affects the effectiveness of the downstream object detection task, with a mean average precision (mAP) improvement of 18.57%. This result fully demonstrates its practical application value. Qiuhan Shao, Xue Wang 0011 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Cross-discipline cold-start knowledge tracing via overlapping entities and contrastive learning
Yulong Deng, Xue Wang 0011, Jie Liu 0056 |
Knowl. Based Syst. | 4 |
| 2025 | VT2V: A benchmark of object tracking with aerial & ground cooperation and multi-modal images
Kaixiang Yan, Wenhua Qian, Cong Bi, Xue Wang 0011 |
Knowl. Based Syst. | 4 |
| 2025 | PRPSV: Parking Efficiency and Reservation Service Optimization Based on Parking Space ViewabstractDifficulty in parking leads to many issues, such as traffic congestion, and hinders the development of intelligent transportation systems. One significant reason is that the cruise parking does not fully utilize real-time information about the parking lot (e.g., availability status of the parking spaces), resulting in low parking efficiency and high parking costs. Even though reservation parking improves parking efficiency, its reservation service is coarse (e.g., could not reserve a specific parking space). Therefore, to achieve more efficient parking and optimize existing reservation services, we propose a real-time parking space view (PSV) framework (called PRPSV). PSV can reflect the position distribution and availability status of each parking space in a parking lot, which enables drivers to quickly and efficiently obtain real-time parking information and complete parking decisions. However, there have been fewer studies on PSV in recent years, and these methods are high-cost, limited scalability, and do not use PSV to optimize parking efficiency and services. Therefore, we propose a method to construct and update PSV accurately. Further, we model cruise and reservation modes in non-PSV and PSV-based scenarios to compare and analyze the impact of PSV on parking efficiency. Finally, the comprehensive qualitative comparison with related work demonstrates the innovativeness of PRPSV and the sufficient experimental results and a case study in an actual parking lot show that PRPSV can efficiently and accurately construct and update PSV, and the introduction of PSV can effectively improve parking efficiency and optimize reservation services. Jishu Wang, Xuan Zhang 0002, LinYu Li 0001, Xue Wang 0011, Shenglong Lv, Rui Zhu 0009, Tong Li 0004 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | PID Controller-Driven Network for Image FusionabstractWith its well-designed network architecture, the deep learning-based infrared and visible image fusion (IVIF) method shows its efficiency and effectiveness by realizing a fine feature extraction and fusion mechanism. However, disparities in cross-modal features often result in an imbalance between texture details and contextual information, causing detailed features to be overshadowed by prevailing contextual information. To tackle this issue, this study introduces PIDFusion, a fusion model driven by a PID controller, designed to dynamically optimize cross-modal feature fusion deviations. The core of PIDFusion is the dynamic adaptation capability of the PID controller, which facilitates real-time corrections for deviations encountered during the fusion process, thereby maintaining a harmonious balance between texture details and contextual information. Additionally, we introduced the Cyclic Self-Supervised Feature Refinement (CSSFR), which under the constraint of self-supervised loss, minimizes redundant information within the feature flow and ensures the preservation of salient feature through the cyclic input of decoupled features. Concurrently, we developed the Iterative Attention Module (IAM), utilizing the unique gating mechanism of LSTM to capture feature changes across successive iterations, thereby driving the model to cultivate more discriminative feature representations. Extensive experiments revealed that PIDFusion outperforms SOTA methods in terms of both efficiency and cost-effectiveness, through static statistics and high-level vision tasks. Our code is available athttps://github.com/wang-x-1997/PIDFusion. Xue Wang 0011, Wenhua Qian, Jinde Cao, Runzhuo Ma |
IEEE Trans. Multim. | 1 |
| 2025 | STFuse: Infrared and Visible Image Fusion via Semisupervised Transfer LearningabstractInfrared and visible image fusion (IVIF) aims to obtain an image that contains complementary information about the source images. However, it is challenging to define complementary information between source images in the lack of ground truth and without borrowing prior knowledge. Therefore, we propose a semisupervised transfer learning-based method for IVIF, termed STFuse, which aims to transfer knowledge from an informative source domain to a target domain, thus breaking the above limitations. The critical aspect of our method is to borrow supervised knowledge from the multifocus image fusion (MFIF) task and to filter out task-specific attribute knowledge by using a guidance loss , which motivates its cross-task use in IVIF tasks. Using this cross-task knowledge effectively alleviates the limitation of the lack of ground truth on fusion performance, and the complementary expression ability under the constraint of supervised knowledge is more instructive than prior knowledge. Moreover, we designed a cross-feature enhancement module (CEM) that utilizes self-attention and mutual-attention features to guide each branch to refine features and then facilitate the integration of cross-modal complementary features. Extensive experiments demonstrate that our method has good advantages in terms of visual quality and statistical metrics, as well as the docking of high-level vision tasks, compared with other state-of-the-art methods. Xue Wang 0011, Wenhua Qian, Jinde Cao, Chengchao Wang 0002, Runzhuo Ma |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Unpaired high-quality image-guided infrared and visible image fusion via generative adversarial networkabstractCurrent infrared and visible image fusion (IVIF) methods lack ground truth and require prior knowledge to guide the feature fusion process. However, in the fusion process, these features have not been placed in an equal and well-defined position, which causes the degradation of image quality. To address this challenge, this study develops a new end-to-end model, termed unpaired high-quality image-guided generative adversarial network (UHG-GAN). Specifically, we introduce the high-quality image as the reference standard of the fused image and employ a global discriminator and a local discriminator to identify the distribution difference between the high-quality image and the fused image. Through adversarial learning, the generator can generate images that are more consistent with high-quality expression. In addition, we also designed the laplacian pyramid augmentation (LPA) module in the generator, which integrates multi-scale features of source images across domains so that the generator can more fully extract the structure and texture information. Extensive experiments demonstrate that our method can effectively preserve the target information in the infrared image and the scene information in the visible image and significantly improve the image quality. Xue Wang 0011, Qiuhan Shao |
Comput. Aided Geom. Des. | 3 |
| 2024 | LightingFormer: Transformer-CNN hybrid network for low-light image enhancement
Cong Bi, Wenhua Qian, Jinde Cao, Xue Wang 0011 |
Comput. Graph. | 4 |
| 2024 | Research on reflective clothing recognition algorithm based on combining omni-dimensional dynamic convolution and partial convolution
Wenbi Ma, Xue Wang 0011, Zhuqing Zhang, Jinde Cao |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Multi-level difference information replenishment for medical image fusion
Luping Chen, Xue Wang 0011, Ya Zhu, Rencan Nie |
Appl. Intell. | 2 |
| 2023 | MCNN: Conditional focus probability learning to multi-focus image fusion via mutually coupled neural networkabstractAbstract In this paper, a novel conditional focus probability learning model, termed MCNN, is proposed for multi‐focus image fusion (MFIF). Given a pair of source images, their conditional focus probabilities can be generated by using the well‐trained MCNN, which is further converted into the binary focus masks to directly produce an all‐focus image with no postprocessing. To this end, a fully convolutional encoder is designed with two mutually coupled Siamese branches in MCNN, which include a coupling block that bridge between the two branches to provide conditional information to each other, at different layers, such that the encoder can more strongly extract conditional focus features and further encourage the decoder pixel‐wisely to give more robust conditional focus probabilities. Moreover, a hybrid loss is designed with a structural sparse fidelity loss and a structural similarity loss to force the network to learn more accurate conditional focus probabilities. Particularly, a convolutional norm with good structural group sparse is proposed, to construct the structural sparse fidelity loss. Simulation results substantiate the superiority of our MCNN over other state‐of‐the‐art, in terms of both visual perception and quantitative evaluation. Chengchao Wang 0002, Xue Wang 0011, Chaozhen Ma, Rencan Nie |
IET Image Process. | 3 |
| 2023 | Multisynchronization of Delayed Fractional-Order Neural Networks via Average Impulsive Interval
Xue Wang 0011, Xiaoshuai Ding, Jinde Cao |
Neural Process. Lett. | 1 |
| 2022 | NCDCN: multi-focus image fusion via nest connection and dilated convolution network
Xue Wang 0011, Rencan Nie, Shishuang Yu, Chengchao Wang 0002 |
Appl. Intell. | 2 |
| 2022 | CEFusion: Multi-Modal medical image fusion via cross encoderabstractAbstract Most existing deep learning‐based multi‐modal medical image fusion (MMIF) methods utilize single‐branch feature extraction strategies to achieve good fusion performance. However, for MMIF tasks, it is thought that this structure cuts off the internal connections between source images, resulting in information redundancy and degradation of fusion performance. To this end, this paper proposes a novel unsupervised network, termed CEFusion. Different from existing architecture, a cross‐encoder is designed by exploiting the complementary properties between the original image to refine source features through feature interaction and reuse. Furthermore, to force the network to learn complementary information between source images and generate the fused image with high contrast and rich textures, a hybrid loss is proposed consisting of weighted fidelity and gradient losses. Specifically, the weighted fidelity loss can not only force the fusion results to approximate the source images but also effectively preserve the luminance information of the source image through weight estimation, while the gradient loss preserves the texture information of the source image. Experimental results demonstrate the superiority of the method over the state‐of‐the‐art in terms of subjective visual effect and quantitative metrics in various datasets. Ya Zhu, Xue Wang 0011, Luping Chen, Rencan Nie |
IET Image Process. | 2 |