Xinyu Wang 0036

dblp:68/1277-36 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2025
0009-0004-0388-0805ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 In2NeCT: Inter-class and Intra-class Neural Collapse Tuning for Semantic Segmentation of Imbalanced Remote Sensing Images
abstract
Remote sensing images (RSIs) are frequently characterized by multi-scale inter-class objects and inconsistently distributed objects due to scene limitations, which would cause a significant data imbalance challenging the corresponding semantic segmentation. Recent methods have leveraged various deep learning techniques to capture high-quality representations for RSI semantic segmentation, but are hardly capable of addressing the afore-mentioned challenge given their limited explorations towards the mechanisms behind the representations. The recently discovered Neural Collapse (NC) phenomenon in computer vision models suggests the simplex equiangular tight frame (ETF) as the optimal representation structure, which has motivated us to observe that the optimal structure of last-layer representations is disrupted and inter-class representations for minor classes tend to become closer to each other beacuse of data imbalance. To address these issues, we propose Inter-class and Intra-class Neural Collapse Tuning (In2NeCT) to optimize the representations that satisfy the simplex ETF, which facilitates the discrimination of inter-class representations and the coherence of intra-class representations. Extensive experiments on three datasets demonstrate that our In2NeCT consistently leads to significant improvements in performance and outperforms the state-of-the-art methods.
Junao Shen, Qiyun Hu, Tian Feng 0001, Xinyu Wang 0036, Hui Cui 0002, Sensen Wu, Wei Zhang 0243
AAAI4
2025 SuraGS: Toward efficient few-shot novel view synthesis via surface-aware Gaussian splatting
Junao Shen, Tian Feng 0001, Haojie Dong, Jinkang Ji, Xinyu Wang 0036, Tianjia Shao
Comput. Graph.5
2024 CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment unseen objects in a query image using a few pixel-wise annotated support images, thus expanding the capabilities of semantic segmentation. The main challenge lies in extracting sufficient information from the limited support images to guide the segmentation process. Conventional methods typically address this problem by generating single or multiple prototypes from the support images and calculating their cosine similarity to the query image. However, these methods often fail to capture meaningful information for modeling the de facto joint distribution of pixel and category. Consequently, they result in incomplete segmentation of foreground objects and mis-segmentation of the complex background. To overcome this issue, we propose the Cross Gaussian Mixture Generative Model (CGMGM), a novel Gaussian Mixture Models~(GMMs)-based FSS method, which establishes the joint distribution of pixel and category in both the support and query images. Specifically, our method initially matches the feature representations of the query image with those of the support images to generate and refine an initial segmentation mask. It then employs GMMs to accurately model the joint distribution of foreground and background using the support masks and the initial segmentation mask. Subsequently, a parametric decoder utilizes the posterior probability of pixels in the query image, by applying the Bayesian theorem, to the joint distribution, to generate the final segmentation mask. Experimental results on PASCAL-5i and COCO-20i datasets demonstrate our CGMGM's effectiveness and superior performance compared to the state-of-the-art methods.
Junao Shen, Kun Kuang 0001, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243
AAAI4
2024 SESAME: Toward Medical Image Segmentation via Foundation Model-assisted Semi-supervised Learning
abstract
Medical image segmentation is essential for diagnosis but requires expensive and time-consuming labeled data. Semi-supervised learning (SSL) mitigates this issue by using unlabeled data to improve generalization. However, current SSL methods encounter issues with inadaptive perturbations and low-quality pseudo-labels. Vision foundation models, such as SAM, have shown promise in segmentation. We propose SESAME, an SSL method integrating SAM and U-Net to improve labeling accuracy through a foundation model-assisted pipeline. In particular, we introduce a reliability score to address low-quality pseudo-labels and employ strategies for utilizing both reliable and unreliable images. Reliable images are associated with refined pseudo-labels via a conflict resolving strategy, whereas unreliable ones undergo a mutual region swapping strategy. Extensive experiments demonstrate that our SESAME outperforms representative methods for medical image segmentation.
Qiyun Hu, Junao Shen, Jinkang Ji, Xinyu Wang 0036, Tian Feng 0001, Hui Cui 0002
BIBM4
2024 Retinal Vessel Segmentation via Cross-attention Feature Fusion
abstract
Retinal vessel segmentation from fundus images is of significant importance for detecting and diagnosing common ocular diseases. Conventional deep learning-based methods for retinal vessel segmentation follow the U-Net framework with an encoder-decoder architecture and employ skip connections for the recovery of spatial information lost during downsampling. However, skip connections cannot consistently have positive contributions to segmentation performance, which is caused by the semantic incompatibility between encoder features and decoder features. Based on this observation, we propose CaFFNet, a Cross-attention Feature Fusion Network designed specifically for retinal vessel segmentation. Specifically, we improve skip connections by introducing a Cross-attention Feature Fusion (CaFF) module, which effectively mitigates the semantic gap between encoder and decoder feature maps by leveraging the cross-attention mechanism for feature fusion. Besides, we introduce a Dual-Branch Pooling Fusion (DBPF) module to address the loss of vessel spatial information during pooling and capture contextual details more effectively, so as to improve segmentation performance. Experimental results on three fundus image datasets demonstrate that our CaFFNet outperforms current representative methods for retinal vessel segmentation.
Tian Feng 0001, Junao Shen, Qiangguo Jin, Xinyu Wang 0036
ICME6
2024 WirePAuS: Auxiliary-free Single-shot Wireframe Parsing
abstract
Wireframe parsing aims to identify vectorized line segments as pairs of endpoints from an image. Conventional methods usually require field-specific knowledge for manual introduction of auxiliary processes or auxiliary learning tasks towards satisfactory performances. Such pipelines are, however, characterized by high complexity, insignificant efficiency, and limited space for further performance improvement. To address these issues, we propose WirePAuS, a novel Wireframe Parser with an Auxiliary-free Single-shot pipeline, which requires no auxiliary processes or auxiliary learning tasks. This is based on its capability to generate appropriate prior information from a prior-informed feature extractor, which incorporates frequency-domain and Hough-domain prior information on line segments in the backbone. Meanwhile, we devise a structurally compact pipeline that enables the parser to directly predict the focal midpoints as line objects and exploit their displacements for the corresponding endpoints. Our end-to-end trainable WirePAuS is capable to capture rich structural details during single-shot inference. Extensive experiments suggest that the proposed method reaches significantly improved performances for wireframe parsing and outperforms a series of state-of-the-art methods.
Jinkang Ji, Junao Shen, Xinyu Wang 0036, Tian Feng 0001, Sensen Wu
ICME3
2024 SSETPAN: Spatial-Spectral Enhanced Transformer based network for pansharpening
abstract
Pansharpening aims for effective spatial-spectral fusion of low-resolution multispectral (LR-MS) and panchromatic (PAN) images, yielding high-resolution multispectral (HR-MS) images. PAN images contain rich spatial details and LR-MS images contain abundant spectral features. However, most of the learning-based methods ignore their distinct attributes, and employ weaker fusion strategies. Besides, Transformer has recently gained considerable popularity in target feature extraction. Therefore, our paper develops a novel Transformer-based network for pansharpening, dubbed Spatial-Spectral Enhanced Transformer based network (SSETPAN), proficient in fine-grained spatial-spectral feature extraction and interaction. SSET-PAN comprises three main modules: channel-wise transformer (CTM), spatial-wise transformer (STM), and adaptive spatial-spectral feature fusion (ASSFM). CTM extracts LR-MS explicit spectral features, while STM obtains high-quality spatial features. ASSFM achieves adaptive spatial-spectral feature fusion via kernel-varied convolution combination. Extensive experiments on GaoFen-2 and WorldView-3 datasets demonstrate that SSETPAN achieves favorable performance against existing pansharpening methods.
Huanting Zhang, Mengting Ma, Xinyu Wang 0036, Wei Zhang 0243
ICME3
2024 DuCoFPan: Dual-Condition Flow-based Network for Pan-sharpening
abstract
Pan-sharpening aims to reconstruct high-resolution multi-spectral (HR-MS) images from panchromatic (PAN) images and low-resolution multi-spectral (LR-MS) images. Despite demonstrated performance, existing learning-based methods struggle to address the ill-posed problem from spectral and spatial perspectives. Generally, mitigating the ill-posed issue involves obtaining the probability distribution of HR-MS images. In this paper, we propose DuCoFPan, a novel dual-condition flow-based network for pan-sharpening, which learns the distributions of HR-MS images guided by spectral and spatial conditions, respectively. Specifically, we design a Dual-Condition Flow Module (DCFM) that adopts spectral and spatial conditions through reversible affine transformations. For condition injection, we present a Spectral-based Condition Injection Block (SPEB) capturing fine-grained spectral features in the Fourier domain and a Spatial-based Condition Injection Block (SPAB) extracting spatial features. Additionally, we devise a Feature Interaction Block (FIB) to promote information flow. Extensive experiments on QuickBird and GaoFen-2 datasets demonstrate the effectiveness and superiority of our method.
Mengjiao Zhao, Mengting Ma, Xinyu Wang 0036, Ao Gao, Wei Zhang 0243
ICME5
2024 MuMoSNet: 3D MRI-based Brain Tumor Segmentation via Multi-modal and Multi-scale Feature Fusion
abstract
MRI images contain multi-modal information, introducing complexity to brain tumor segmentation. Recent studies have incorporated the Transformer model, given its exceptional capability to model long-range dependence, into convolutional neural networks (CNNs) to address limited receptive fields. However, such a hybrid strategy often neglects the inherent multimodal characteristics of MRI images and lacks the capacity to capture modality-specific features. In this paper, we propose a multi-modal and multi-scale feature fusion network (MuMoSNet) for brain tumor segmentation from 3D MRI images. Specifically, our MuMoSNet introduces a parallel ME-Transformer encoder alongside the CNN-based encoder in 3D U-Net to separately extract modality-specific features. Besides, we devise a multi-feature fusion (MuFF) module to learn affinity relationships between cross-modality shared features and modality-specific features, maximizing the exploration of multi-modal information. Extensive experiments on both BraTS21 and BraTS20 datasets suggest that our MuMoSNet outperforms current representative methods for brain tumor segmentation.
Hui Cui 0002, Junao Shen, Xinyu Wang 0036, Tian Feng 0001
ICME6
2024 DOCNet: Dual-Domain Optimized Class-Aware Network for Remote Sensing Image Segmentation
abstract
The spatial attention mechanism has been frequently employed for the semantic segmentation of remote sensing images, given its renowned capability to model long-range dependencies. As remote sensing images often exhibit intricate backgrounds, significant intraclass variability, and a foreground-background imbalance, spatial attention mechanism-based methods somehow tend to introduce an extensive amount of background context through intensive affinity operations, causing unsatisfactory segmentation outcomes. While several class-aware methods attempt to attenuate the interference of background context by generating class representations as representative features, they still encounter challenges related to independent correlation calculation and single-confidence scale class representations. We introduce a dual-domain optimized class-aware network designed to address these challenges. In the semantic domain, we use category confidence as a scaling criterion to derive class representations at multiple confidence scales, effectively modeling pixel-class relationships. In the spatial domain, we leverage pixel-class relationships and their consensus to enhance relevant correlations while suppressing erroneous ones. Experimental results on three datasets demonstrate that the proposed method surpasses previous state-of-the-art ones for remote sensing image segmentation. Code is available athttps://github.com/xwmaxwma/rssegmentation.
Rui Che, Xinyu Wang 0036, Mengting Ma, Sensen Wu, Tian Feng 0001, Wei Zhang 0243
IEEE Geosci. Remote. Sens. Lett.3
2023 DBDAN: Dual-Branch Dynamic Attention Network for Semantic Segmentation of Remote Sensing Images
Rui Che, Tingfeng Hong, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243
PRCV (4)4
2023 MAPMaN: Multi-Stage U-Shaped Adaptive Pattern Matching Network for Semantic Segmentation of Remote Sensing Images
abstract
Abstract Remote sensing images (RSIs) often possess obvious background noises, exhibit a multi‐scale phenomenon, and are characterized by complex scenes with ground objects in diversely spatial distribution pattern, bringing challenges to the corresponding semantic segmentation. CNN‐based methods can hardly address the diverse spatial distributions of ground objects, especially their compositional relationships, while Vision Transformers (ViTs) introduce background noises and have a quadratic time complexity due to dense global matrix multiplications. In this paper, we introduce Adaptive Pattern Matching (APM), a lightweight method for long‐range adaptive weight aggregation. Our APM obtains a set of pixels belonging to the same spatial distribution pattern of each pixel, and calculates the adaptive weights according to their compositional relationships. In addition, we design a tiny U‐shaped network using the APM as a module to address the large variance of scales of ground objects in RSIs. This network is embedded after each stage in a backbone network to establish a Multi‐stage U‐shaped Adaptive Pattern Matching Network (MAPMaN), for nested multi‐scale modeling of ground objects towards semantic segmentation of RSIs. Experiments on three datasets demonstrate that our MAPMaN can outperform the state‐of‐the‐art methods in common metrics. The code can be available at https://github.com/INiid/MAPMaN .
Tingfeng Hong, Xinyu Wang 0036, Rui Che, Chenlu Hu, Tian Feng 0001, Wei Zhang 0243
Comput. Graph. Forum3