Xin Wang 0078

dblp:10/5630-78 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
24since 2021 · last 2025
0000-0003-2386-5405ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 14 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Novel Graph-Based Skip Connection Design for Hierarchical and Relational Feature Modeling
abstract
Accurately extracting both global context and detailed local features remains a significant hurdle in medical image segmentation. Traditional techniques often lack the capacity to effectively model hierarchical structures and pixel-level spatial dependencies, which can lead to blurred boundaries and loss of intricate details. To overcome these limitations, we introduce a novel graph-based skip connection design named SGraphNet, which combines the Parallel Hierarchical-Relational Graph (PHRG) module with the Dynamic Attention Convolution (DAC) module. The PHRG module builds Hierarchical Spatial Graphs (HSG) and Relational Position Graphs (RPG) as skip connections to independently capture multi-scale spatial features and the relative dependencies among different feature regions. Complementing this, the DAC module integrates spatial attention with dynamic convolution to further enhance the representation of features. Experiments conducted on the Synapse dataset show that SGraphNet achieves a Dice Similarity Coefficient (DSC) of 84.88%, surpassing G-CASCADE and MSA-2Net by 1.82% and 0.75%, respectively. Additionally, the model delivers competitive results on the ISIC2017 and ISIC2018 datasets across both DSC and mIoU metrics. These results underscore the promise of graph-based convolutional designs as strong contenders to CNN- and Transformer-driven approaches in medical image segmentation. You can get the code on Github https://anonymous.4open.science/r/SGraphNet
Wei Wang 0229, Xin Wang 0078
BIBM3
2025 LRD-3DSAM: A Novel SAM-Based 3D Medical Segmentation Network for Modeling Long-Range Dependencies
abstract
Although the Segment Anything Model (SAM) performs well in natural image segmentation, its application to 3D medical imaging remains challenging due to limited modeling of complex spatial structures and long-range slice dependencies, as well as high sensitivity to input prompt, which affects segmentation stability and clinical reliability. To address these issues, we propose LRD-3DSAM, a novel network tailored for 3D medical image segmentation. It incorporates an Adaptive Spatial Adapter to enhance 3D spatial context modeling and effectively capture cross-slice dependencies. On top of this, we introduce a Tumor Prototype Prompt Encoder, which aggregates tumor features across samples using structured representations from the adapter. This yields globally consistent tumor prototypes, improving semantic stability, inter-slice consistency, and robustness to annotation noise. Experiments on multiple public 3D tumor segmentation datasets show that LRD-3DSAM outperforms SAM and its variants in segmentation accuracy, inter-slice consistency, and cross-dataset generalization. On the pancreas tumor dataset, it achieves a 1.2% higher Dice score and 1.66% improvement in NSD, setting a new state-of-the-art.
Zhipeng Hu, Wei Wang 0229, Xin Wang 0078
BIBM4
2025 GBA-Net: A Method for 3D Brain Tumor Segmentation Based on Multi-scale Gaussian Boundary Attention
abstract
The complex nature of brain tumors, characterized by their individual shapes, sizes, and locations, as well as the presence of indistinct boundaries, presents a challenging task for precise automatic segmentation. While U-Net has been a top performer in medical image segmentation, it struggles with capturing multi-scale details, preserving information across layers, and focusing on critical features. To address these issues, a new 3D brain tumor segmentation network called GBA-Net is proposed, which introduces - 1) A multi-scale Gaussian boundary attention (GBA) module with the ability to automatically focus on boundary features, 2) Efficient inverted bottleneck convolution upsample (Up-IBC) and inverted bottleneck convolution downsample (Down-IBC) modules that enhance the richness of cross-scale information, 3) The high-low feature fusion (HLFF) module that mitigates information loss during the decoder’s restoration of full spatial resolution. The proposed GBA-Net achieves state-of-the-art performance on the 3D brain tumor dataset BraTS 2021. Cross-validation on the BraTS 2018 and BraTS 2019 datasets indicates that GBA-Net generalizes well on the external datasets.
Wei Wang 0229, Longrun Wang, Xin Wang 0078
ICASSP3
2025 FSISNet: Exploring Mamba and Transformer for Polyp Segmentation
Wei Wang 0229, Feng Jiang 0024, Xin Wang 0078
ICIC (28)3
2025 STINet: Spatio-Temporal Interaction Network for Remote Sensing Image Change Detection
Wei Wang 0229, Huilin Ren, Xin Wang 0078
ICIC (2)3
2025 DDECNet: Dual-Branch Difference Enhanced Network with Novel Efficient Cross-Attention for Remote Sensing Change Detection
Wei Wang 0229, Xin Wang 0078
ICIC (1)3
2025 MPCGNet: A Multiscale Feature Extraction and Progressive Feature Aggregation Network Using Coupling Gates for Polyp Segmentation
abstract
Automatic segmentation methods of polyps is crucial for assisting doctors in colorectal polyp screening and cancer diagnosis. Despite the progress made by existing methods, polyp segmentation faces several challenges: (1) small-sized polyps are prone to being missed during identification, (2) the boundaries between polyps and the surrounding environment are often ambiguous, (3) noise in colonoscopy images, caused by uneven lighting and other factors, affects segmentation results. To address these challenges, this paper introduces coupling gates as components in specific modules to filter noise and perform feature importance selection. Three modules are proposed: the coupling gates multiscale feature extraction (CGMFE) module, which effectively extracts local features and suppresses noise; the windows cross attention (WCAD) decoder module, which restores details after capturing the precise location of polyps; and the decoder feature aggregation (DFA) module, which progressively aggregates features, further extracts them, and performs feature importance selection to reduce the loss of small-sized polyps. Experimental results demonstrate that MPCGNet outperforms recent networks, with mDice scores 2.20% and 0.68% higher than the second-best network on the ETIS-LaribPolypDB and CVC-ColonDB datasets, respectively.
Wei Wang 0229, Feng Jiang 0024, Xin Wang 0078
IJCNN3
2025 MSFINet: Multi-Scale Feature Interaction Fusion Network for Skin Lesion Image Segmentation
abstract
Segmentation of lesion areas in dermoscopic images is key to diagnosing melanoma. However, it remains challenging due to the blurring boundary between the lesion area and the normal skin and the presence of interfering factors such as hair occlusion. To address the above challenges, a multi-scale feature interaction fusion network (MSFINet) is proposed. First, to unify the semantic features at different stages and capture more diverse feature information, a Channel Interaction Transformer (CIT) module and a Channel Refinement Attention (CRA) module are proposed. Then, to improve the utilization of features and help the model to resist interference, the Multi-Kernel Dual Branch Aggregation (MKDA) module is designed. This module extracts features from different perspectives to enhance the model’s capability of grasping global contextual information. In the decoder part, a Boundary Difference Optimization (BDO) module is proposed, which aggregates the boundary difference information between high-level features and low-level features to assist the decoder’s feature refinement and make the network more sensitive to boundary information. The proposed method is compared with other methods on three datasets: ISIC 2017, ISIC 2018, and PH2. The DSC scores increase by 1.71%, 1.36%, and 0.87%, respectively, compared to the suboptimal network. The code is available at: https://anonymous.4open.science/r/MSFINet-6E52.
Xin Wang 0078, Yuewen Luo, Wei Wang 0229
IJCNN1
2025 Cross-Domain Feature Interaction: A Robust Generalization Network for Cross-Category Medical Image Classification
abstract
Accurate medical image classification is crucial in early diagnosis and treatment of diseases. However, the high heterogeneity among different medical images leads to the prevalence of networks designed for single medical image classification. In contrast, universal medical image classification networks have stronger generalization and scalability in multi-source datasets. To this end, based on multi-scale frequency-spatial domain analysis, we propose a novel cross-domain feature interaction network (CDFINet) for medical image classification. Primarily, the LocalGlobal Frequency-domain Scan (LGFS) block is designed to extract fine-grained features of lesions, enhancing spatial detailed feature representation by introducing gated spectral attention. Drawing on feature extraction of the "center-local-global" multi-view, the Multi-View Bi-dimensional Attention (MVBA) block is devised to capture multi-scale features by modeling global multi-scale contextual dependencies. Finally, the Frequency-Spatial Domain Attention (FSDA) block is designed to guide the network in extracting rich semantic features by leveraging cross-domain features. To verify the generalization of CDFINet, experiments are conducted on gastrointestinal disease, skin disease, breast cancer histopathology and COVID-19 datasets. The accuracy of CDFINet on the gastrointestinal disease classification is 90.47%, which is 8.27% higher than that of SPANet (2024), 4.8% higher than that of HiFuse (2024) and 3.84% higher than that of MedMamba (2024). CDFINet achieves state-of-the-art (SOTA) performance with fewer parameters and less computational cost. Codes: https://github.com/PURSUETHESUN/CDFINet.git.
Xin Wang 0078, Wei Wang 0229, Jixing He
IJCNN1
2025 BF-UNet: A Brain Function-Inspired Model with Memory Mechanism for Medical Segmentation
abstract
Medical segmentation is a foundational task in the realm of medical image analysis. However, existing researches tend to focus on optimizing the network structure and ignore the exploration of the connection between the deep learning model and the human brain functions. Based on this question, we propose a network inspired by brain functions, BF-UNet, which combines the mechanisms of memory, attention and neurons to better capture complex textural details in medical images. Firstly, in order to make the encoder more sensitive to complex details of images, the residual memory module (RMM) inspired by the memory function of the human brain is proposed aiming to store and grasp the overall texture features unique to a dataset. Secondly, to minimise the disparity in semantic meaning between the encoder and decoder, the attention gate module (AGM) is introduced. Finally, the dendritic neuron segmentation head (DNH) imitating the dendritic neuron mechanism of the human brain can capture complex details better than the ordinary segmentation head with convolutional structure. The BF-UNet is compared with other models on three segmentation datasets: ISIC2017, ISIC2018 and BUSI. The DSC scores are as high as 89.25%, 90.11% and 79.4%, respectively, which are substantially better than other networks.
Wei Wang 0229, Xin Wang 0078, BeiBei Zhen
IJCNN3
2025 Rethinking Feature Guidance for Medical Image Segmentation
abstract
Despite the evident advantages of variants of UNet in medical image segmentation, these methods still exhibit limitations in the extraction of foreground, background, and boundary features. Based on feature guidance, we propose a new network (FG-UNet). Specifically, adjacent high-level and low-level features are used to gradually guide the network to perceive lesion features. To accommodate lesion features of different scales, the multi-order gated aggregation (MGA) block is designed based on multi-order feature interactions. Furthermore, a novel feature-guided context-aware (FGCA) block is devised to enhance the capability of FG-UNet to segment lesions by fusing boundary-enhancing features, object-enhancing features, and uncertain areas. Eventually, a bi-dimensional interaction attention (BIA) block is designed to enable the network to highlight crucial features effectively. To appraise the effectiveness of FG-UNet, experiments were conducted on Kvasir-seg, ISIC2018, and COVID-19 datasets. The experimental results illustrate that FG-UNet achieves a DSC score of 92.70% on the Kvasir-seg dataset, which is 1.15% higher than that of the latest SCUNet++, 4.70% higher than that of ACC-UNet, and 5.17% higher than that of UNet. Codes:https://github.com/PURSUETHESUN/FG-UNet.
Wei Wang 0229, Jixing He, Xin Wang 0078
IEEE Signal Process. Lett.3
2024 BEFNet: A Hybrid CNN-Mamba Architecture for Accurate Skin Lesion Image Segmentation
abstract
Accurate image segmentation of skin lesions is crucial for the detection and treatment of skin cancer. Based on the modern state space model Mamba, a novel hybrid CNN-Mamba network (BEFNet) is proposed. Specifically, BEFNet introduces Visual State Space (VSS) blocks as the main encoder for extracting and aggregating global information. To improve the expressiveness and utilization of the features, the Adjacent Feature Interaction Module (AFIM) is proposed, and the information lost during the channel interaction process is compensated by the Knowledge Retrospective Module (KRM). To compensate for Mamba’s lack of attention to local detail information, an auxiliary encoder is constructed based on CNN blocks, and a Boundary Optimization Auxiliary Module (BOA) is designed to obtain richer boundary detail information. Finally, a Refined Feature Guidance Module (RFGM) was designed to use the enhanced boundary information as a priori information to guide feature fusion. We conducted experiments on three publicly available skin cancer datasets, and the experimental results demonstrate that BEFNet has competitive performance on the task of skin lesion image segmentation.
Wei Wang 0229, Yuewen Luo, Xin Wang 0078
BIBM3
2024 Building Change Detection Based on Fully Convolutional Network in High-Resolution Remote Sensing Images
Wei Wang 0229, Luocheng Xia, Xin Wang 0078
ICIC (4)3
2024 A Multi-Scale Additive Enhanced Network for Remote Sensing Scene Classification
Wei Wang 0229, Zixin Zhou, Xin Wang 0078
ICIC (5)3
2024 RMFFNet: A Reverse Multi-Scale Feature Fusion Network for Remote Sensing Scene Classification
abstract
In recent years, convolutional neural network (CNN) and transformer have become mainstream approaches in computer vision tasks and are increasingly used in remote sensing scene classification (RSSC) tasks. In addition, hybrid architectures of CNN and Transformer show great potential, where the models can capture local and global information simultaneously. However, many hybrid networks extract features only at a single scale, which are not enough to accurately recognize remote sensing images. So a reverse multi-scale feature fusion network (RMFFNet) based on a hybrid architecture is proposed for extracting features of different object sizes in remote sensing images and enhancing the model’s contextual understanding. The model uses convolution to capture texture detail information and channel attention to supplement the features in the shallow stage of the network. It adopts a combination of local and global in the deeper stage of the network. The proposed multi-scale multihead self-attention (MMSA) is used to extract the global features. The local branch uses depthwise convolutional self-modulation followed by inter-channel interaction, and then this branch is injected into the global branch for fusion. In addition, a reverse cross-scale interaction (RCI) module is designed to fuse features from different stages, which contributes to more diverse feature representations. The proposed method achieves 97.05% accuracy on RSSCN7 dataset and 95.83% classification accuracy on AID dataset. We perform a series of comparative experiments with the proposed RMFFNet and several CNN-based, Transformer-based, and hybrid network-based models on two datasets, and our method outperforms the performance of these networks.
Wei Wang 0229, YuTing Shi, Xin Wang 0078
IJCNN3
2024 LSSNet: A Method for Colon Polyp Segmentation Based on Local Feature Supplementation and Shallow Feature Supplementation
Wei Wang 0229, Huiying Sun, Xin Wang 0078
MICCAI (7)3
2024 From Coarse to Fine: A Novel Colon Polyp Segmentation Method Like Human Observation
Wei Wang 0229, Huiying Sun, Xin Wang 0078
PRCV (14)3
2024 A Boundary Guided Cross Fusion Approach for Remote Sensing Image Segmentation
abstract
Remote sensing images have a variety of application prospects because of their rich information. Due to recent advances in deep learning methods, solid improvements have been made in the semantic segmentation of high-resolution remote sensing images. However, achieving precise segmentation of small and crowded objects remains a challenge. To tackle this challenging task, a Boundary Guided Cross Fusion module (BGCFM) is proposed. The Bidirectional Boundary Gate module (BBGM) is designed to provide reliable boundary information for BGCFM. Based on these two models, a remote sensing images real-time semantic segmentation network, boundary guided cross fusion network (BGCFNet), is designed. The effectiveness of the boundary-guided fusion method and the performance of BGCFNet were verified on the GID-5 dataset without pretraining. The application of the boundary-guided fusion method on the SOTA dual-branch real-time semantic segmentation network improves segmentation accuracy. BGCFNet achieves the best performance with a Mean Intersection over Union (mIoU) of 88.82%. Its inference speed is about 1.5 times that of other networks in the experiment, achieving an excellent balance between accuracy and speed.
Wei Wang 0229, Yingfeng Zhang, Xin Wang 0078, Ji Li 0011
IEEE Geosci. Remote. Sens. Lett.3
2024 CF-GCN: Graph Convolutional Network for Change Detection in Remote Sensing Images
abstract
The remote sensing image change detection methods based on deep learning have made great progress.However, many CNN-based methods persistently face challenges in connecting long-range semantic concepts because of their limited receptive fields. Recently, some methods that combine transformers effectively extract global information by modeling the context in the temporal and spatial domains has been proposed to solve the problem, but they still suffer from both the incorrect identification of "non-semantic changes" and the incomplete and irregular boundary extraction due to the deterioration of local feature details. In response to these inquiries, we propose a novel network, CF-GCN, based on graph convolutional structures for change detection. Specifically, in the encoder and decoder of the network, different projection strategies are employed to construct coordinate space graph convolution and feature interaction graph convolution. The Boundary Perception Module extracts spatial boundary features of shallow layers and enhances boundary perception ability during graph-based information propagation, effectively suppressing the tendency of image boundary information to gradually smooth out. At the same time, the knowledge review module is utilized to form knowledge complementarity between key layers of the network, effectively mitigating the propagation of erroneous knowledge in the deep network. On the LEVIR-CD dataset, the IoU score of CF-GCN is 83.41%, which is 0.35% and 0.39% higher than ChangeStar and DMINet, respectively. On the WHU-CD dataset, the F1 and IoU are as high as 91.83% and 84.90%, which are significantly better than other state-of-the-art networks. The experimental results show that, in addition to CNN and Transformer, the graph-convolution structure approach is expected to be another major research direction for performing fully supervised change detection. Our code and pre-trained models will be available at https://github.com/liucongcharles/CF-GCN.
Wei Wang 0229, Cong Liu 0043, Guanqun Liu 0006, Xin Wang 0078
IEEE Trans. Geosci. Remote. Sens.4
2024 From Spatial to Frequency Domain: A Pure Frequency Domain FDNet Model for the Classification of Remote Sensing Images
abstract
In remote sensing scene classification (RSSC), a common approach to reduce processing costs is to downsample images in the spatial domain to lower resolution. However, this approach can lead to the loss of texture information. Conversely, the frequency domain analysis can effectively extract frequency variations of pixel values, aiding in distinguishing texture differences in remote sensing scenes. Therefore, it is common in RSSC to use frequency domain analysis as an auxiliary branch of neural networks to extract additional information. The essence of these approaches is to extract feature information from both the spatial and frequency domains and then fuse them together. However, this combination method does not fully leverage the benefits of frequency domain feature extraction. The approach of extracting feature information solely from the frequency domain may require further research and experimentation to explore its potential advantages and applications. This study integrates wavelet transform into downsampling operations and creatively designs a network model called the frequency domain deep learning network (FDNet), which uses frequency domain methods as the backbone of a neural network and performs feature extraction solely in the frequency domain. FDNet not only solves the problem of information loss caused by spatial downsampling but also provides a new research direction of combining frequency domain methods with neural networks, that is, using frequency domain methods as the backbone of neural networks. FDNet also provides a new research direction for RSSC, which is to extract features only in the frequency domain. FDNet has been evaluated on three public remote sensing datasets: SIRI-WHU, RSSCN7, and AID. Each dataset underwent three experiments, with average classification accuracies of 97.78%, 97.62%, and 95.92%, respectively. The results show that FDNet outperforms other advanced methods participating in the comparison of classification tasks.
Wei Wang 0229, Xin Wang 0078
IEEE Trans. Geosci. Remote. Sens.3
2023 BFRNet: Bidimensional Feature Representation Network for Remote Sensing Images Classification
abstract
In recent years, convolutional neural network (CNN) and transformer, as mainstream classification methods, have made good progress in improving the classification performance of remote sensing (RS) images. Furthermore, the CNN-transformer hybrid architecture has shown a greater potential for enabling models to obtain local information and global dependency relationships. In general, numerous researches improve self-attention block of vision transformer (ViT) in the spatial dimension. Nevertheless, spatial self-attention mostly achieves a single spatial feature extraction, which cannot meet the requirement of accurate recognition of high-resolution RS images. In this work, a method for representing spatial and channel dimensions of RS images is proposed, which not only extracts global–local spatial features but also pays special attention to incorporate channel information. Specifically, bidimensional local window self-attention (BLWS) and pyramid pool self-attention are conducted to extract local–global features. Subsequently, a linear attention module will fuse local–global information in the channel dimension when computing multihead self-attention (MHSA). A bidimensional gating unit (BGU) is used to replace the traditional multilayer perceptron (MLP) of the feedforward network (FFN). The above improvements result in a bidimensional feature representation (BFR) block, and BFR network (BFRNet) is designed based on BFR blocks. BFRNet consists of four stages, and each stage repeatedly stacks BFR blocks with different layers. Experiments show that the classification accuracy of BFRNet is significantly better than the existing methods of CNN, ViTs, and CNN-transformer networks. On dataset RSSCN7, BFRNet achieves a classification accuracy of 98.75% with only 1.9G floating point operations (FLOPs), which is 8.21% higher than ViT, 3.21% higher than Resnet50, and 2.68% higher than CoAtNet, respectively.
Wei Wang 0229, Xin Wang 0078, Ji Li 0011
IEEE Trans. Geosci. Remote. Sens.3
2022 A COVID-19 CXR image recognition method based on MSA-DDCovidNet
abstract
Currently, coronavirus disease 2019 (COVID-19) has not been contained. It is a safe and effective way to detect infected persons in chest X-ray (CXR) images based on deep learning methods. To solve the above problem, the dual-path multi-scale fusion (DMFF) module and dense dilated depth-wise separable (D3S) module are used to extract shallow and deep features, respectively. Based on these two modules and multi-scale spatial attention (MSA) mechanism, a lightweight convolutional neural network model, MSA-DDCovidNet, is designed. Experimental results show that the accuracy of the MSA-DDCovidNet model on COVID-19 CXR images is as high as 97.962%, In addition, the proposed MSA-DDCovidNet has less computation complexity and fewer parameter numbers. Compared with other methods, MSA-DDCovidNet can help diagnose COVID-19 more quickly and accurately.
Wei Wang 0229, Wendi Huang, Xin Wang 0078, Peng Zhang 0079
IET Image Process.3
2022 A ViT-Based Multiscale Feature Fusion Approach for Remote Sensing Image Segmentation
abstract
Semantic segmentation plays an indispensable role in automatic analysis of remote sensing image data. However, the abundant semantic information and irregular shape patterns in remote sensing images are difficult to utilize, making it hard to segment remote sensing images only using convolution and single-scale feature maps. To achieve better segmentation performance, a multiscale feature pyramid decoder (MFPD) is proposed to fuse image features extracted by vision transformer (ViT). The decoder employs a novel 2-D-to-3-D transform method to obtain multiscale feature maps that contain rich context information and fuses the multiscale feature maps by channel concatenation. Furthermore, a dimension attention module (DAM) is designed to further aggregate the context information of the extracted remote sensing image features. This approach yields superior mean intersection over union (mIoU) on the Gaofen2-CZ dataset (60.42%) and GID-5 dataset (68.21%). Experimental results indicate that the comprehensive performance of our approach exceeds the compared segmentation methods based on convolutional neural network (CNN) and ViT.
Wei Wang 0229, Xin Wang 0078
IEEE Geosci. Remote. Sens. Lett.3
2021 JCOGIN: a programming framework for particle transport on combinatorial geometry
Baoyin Zhang, Zeyao Mo, Xin Wang 0078, Wei Wang 0229, Aiqing Zhang, Xiaolin Cao
J. Supercomput.3