Xiaosheng Yu 0001

dblp:99/8589-1 · DBLP profile ↗
← Back
36ranked-venue papers
5as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Of-GS: Overlap-free Gaussian Splatting via visual geometry grounded transformer for few-shot novel view synthesis
Qianjun Li, Jubo Chen, Runsheng Diao, Xiaosheng Yu 0001
Comput. Graph.4
2026 SFC-net: A spatial frequency-inspired feature collaborative alignment and fusion network for weakly aligned RGB-T salient object detection
Jubo Chen, Xiaosheng Yu 0001, Liuyi Meng, Yangpu Tang, Mingli Song
Neurocomputing2
2026 Deep Fourier-Embedded Network for RGB and Thermal Salient Object Detection
abstract
The rapid development of deep learning has significantly improved salient object detection (SOD) combining both RGB and thermal (RGB-T) images. However, existing Transformer-based RGB-T SOD models with quadratic complexity are memory-intensive, limiting their application in high-resolution bimodal feature fusion. To overcome this limitation, we propose a purely Fourier Transform-based model, namely Deep Fourier-embedded Network (FreqSal), for accurate RGB-T SOD. Specifically, we leverage the efficiency of Fast Fourier Transform with linear complexity to design three key components: (1) To fuse RGB and thermal modalities, we propose Modal-coordinated Perception Attention, which aligns and enhances bimodal Fourier representation in multiple dimensions; (2) To clarify object edges and suppress noise, we design Frequency-decomposed Edge-aware Block, which deeply decomposes and filters Fourier components of low-level features; (3) To accurately decode features, we propose Fourier Residual Channel Attention Block, which prioritizes high-frequency information while aligning channel-wise global relationships. Additionally, even when converged, existing deep learning-based SOD models’ predictions still exhibit frequency gaps relative to ground-truth. To address this problem, we propose Co-focus Frequency Loss, which dynamically weights hard frequencies during edge frequency reconstruction by cross-referencing bimodal edge information in the Fourier domain. Extensive experiments on ten bimodal SOD benchmark datasets demonstrate that FreqSal outperforms twenty-nine existing state-of-the-art bimodal SOD models. Comprehensive ablation studies further validate the value and effectiveness of our newly proposed components. The code is available at https://github.com/JoshuaLPF/FreqSal.
Pengfei Lyu, Xiaosheng Yu 0001, Pak-Hei Yeung, Chengdong Wu 0001, Jagath C. Rajapakse
IEEE Trans. Circuits Syst. Video Technol.2
2026 A dual-branch RGB-T salient object detection via spatial-frequency integration
Xiaosheng Yu 0001, Ying Wang 0152, Jubo Chen
Vis. Comput.1
2025 FreqSpace-NeRF: A fourier-enhanced Neural Radiance Fields method via dual-domain contrastive learning for novel view synthesis
Xiaosheng Yu 0001, Xiaolei Tian, Jubo Chen, Ying Wang 0152
Comput. Graph.1
2025 PONet: Prototype optimization network for few-shot medical image segmentation
Xiaosheng Yu 0001, Jianning Chi, Chengdong Wu 0001, Xiujing Gao
Neurocomputing2
2025 Conditional variational underwater image enhancement with kernel decomposition and adaptive hybrid normalization
Haopeng Zhang 0018, Hongli Xu 0003, Hao Liu 0008, Xiaosheng Yu 0001, Xiangyue Zhang, Chengdong Wu 0001
Neurocomputing4
2025 RoGLSNet: An Efficient Global-Local Scene Awareness Network With Rotary Position Embedding for Remote Image Segmentation
abstract
Accurate segmentation of very high-resolution remote sensing images is vital for downstream tasks. Most semantic segmentation methods fail to fully consider the inherent characteristics of the images, such as intricate backgrounds, significant intraclass variance, and spatial interdependence of geographic object distribution. To address these challenges, we propose an efficient global–local scene awareness network with rotary position embedding (RoGLSNet). Specifically, we introduce the dynamic global filter (DGF) module to adaptively select frequency components, thereby mitigating interference from background noise. For high intraclass variance, the class center aware block (CCAB) performs class-level contextual modeling with spatial information integration. Additionally, the rotary position embedding (RoPE) is incorporated into vanilla attention to indirectly model the positional and distance relationships of geographic target objects. Extensive experimental results on two widely used datasets demonstrate that RoGLSNet outperforms the state-of-the-art (SOTA) segmentation methods. The code is available athttps://github.com/bai101315/RoGLSNet
Xiaosheng Yu 0001, Weiqi Bai, Jubo Chen, Zhuoqun Fang, Zhaokui Li
IEEE Geosci. Remote. Sens. Lett.1
2025 Efficient Fourier Filtering Network With Contrastive Learning for AAV-Based Unaligned Bimodal Salient Object Detection
abstract
Unmanned aerial vehicle (UAV)-based bi-modal salient object detection (BSOD) aims to segment salient objects in a scene utilizing complementary cues in unaligned RGB and thermal image pairs. However, the high computational expense of existing UAV-based BSOD models limits their applicability to real-world UAV devices. To address this problem, we propose an efficient Fourier filter network with contrastive learning that achieves both real-time and accurate performance. Specifically, we first design a semantic contrastive alignment loss to align the two modalities at the semantic level, which facilitates mutual refinement in a parameter-free way. Second, inspired by the fast Fourier transform that obtains global relevance in linear complexity, we propose synchronized alignment fusion, which aligns and fuses bi-modal features in the channel and spatial dimensions by a hierarchical filtering mechanism. Our proposed model, AlignSal, reduces the number of parameters by 70.0%, decreases the floating point operations by 49.4%, and increases the inference speed by 152.5% compared to the cutting-edge BSOD model (i.e., MROS). Extensive experiments on the UAV RGB-T 2400 and seven bi-modal dense prediction datasets demonstrate that AlignSal achieves both real-time inference speed and better performance and generalizability compared to nineteen state-of-the-art models across most evaluation metrics. In addition, our ablation studies further verify AlignSal’s potential in boosting the performance of existing aligned BSOD models on UAV-based unaligned data. The code is available at: https://github.com/JoshuaLPF/AlignSal.
Pengfei Lyu, Pak-Hei Yeung, Xiaosheng Yu 0001, Xiufei Cheng, Chengdong Wu 0001, Jagath C. Rajapakse
IEEE Trans. Geosci. Remote. Sens.3
2025 CDF-UIE: Leveraging Cross-Domain Fusion for Underwater Image Enhancement
abstract
Underwater image enhancement (UIE) aims to restore image quality by mitigating inherent degradations in underwater imaging systems. While existing learning-based methods show promise, they face limitations in separating and processing frequency components, effectively fusing domain information, and balancing the enhancement of structures and details. To resolve these limitations, we propose cross-domain fusion (CDF)-UIE, a novel network that leverages and fuses cross-domain information for mitigating the degradation in underwater images. CDF-UIE first performs domain decoupling of input features using the proposed spatial-frequency decoupling (SFD) block. Then, we design an innovative CDF block, which effectively bridges the spatial- and frequency-domain features through the cross-domain attention mechanism. To produce stable and detailed enhanced outputs, we exploit the coarse and fine-scale information in the image reconstruction stage. In addition, we introduce a multiscale objective function that incorporates pixel-level, structural, and perceptual constraints to guide the enhancement process. We conduct extensive experiments on six diverse real-world underwater image datasets. Comprehensive experiments and real-world application tests demonstrate that CDF-UIE significantly outperforms existing methods, offering promising future applications in various underwater scenarios. The source code is available athttps://github.com/hpzhan66/CDF-UIE.
Haopeng Zhang 0018, Hongli Xu 0003, Xiaosheng Yu 0001, Xiangyue Zhang, Xiujing Gao, Chengdong Wu 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 TwinsTNet: Broad-View Twins Transformer Network for Bi-Modal Salient Object Detection
abstract
Exploring complementary information between RGB and thermal/depth modalities is crucial for bi-modal salient object detection (BSOD). However, the distinct characteristics of different modalities often lead to large differences in information distributions. Existing models, which rely on convolutional operations or plug-and-play attention mechanisms, struggle to address this issue. To overcome this challenge, we rethink the relationship between information complementarity and long-range relevance, and propose a uniform broad-view Twins Transformer Network (TwinsTNet) for accurate BSOD. Specifically, to efficiently fuse bi-modal information, we first design the Cross-Modal Federated Attention (CMFA), which mines complementary cues across modalities through element-wise global dependency. Second, to ensure accurate modality fusion, we propose the Semantic Consistency Attention Loss, which supervises the co-attention feature in CMFA using the ground-truth-generated attention map. Additionally, existing BSOD models lack the exploration of inter-layer interactions, for which we propose the Cross-Scale Retracing Attention (CSRA), which retrieves query-relevant information from stacked features of all previous layers, enabling flexible cross-layer interactions. The cooperation between CMFA and CSRA mitigates inductive bias in both modality and layer dimensions, enhancing TwinsTNet's representational capability. Extensive experiments demonstrate that TwinsTNet outperforms twenty-two existing state-of-the-art models on ten BSOD benchmark datasets. The code is available at: https://github.com/JoshuaLPF/TwinsTNet.
Pengfei Lyu, Xiaosheng Yu 0001, Jianning Chi, Hao Wu 0064, Chengdong Wu 0001, Jagath C. Rajapakse
IEEE Trans. Image Process.2
2025 A Dual-Branch Cross-Modality-Attention Network for Thyroid Nodule Diagnosis Based on Ultrasound Images and Contrast-Enhanced Ultrasound Videos
abstract
Contrast-enhanced ultrasound (CEUS) has been extensively employed as an imaging modality in thyroid nodule diagnosis due to its capacity to visualise the distribution and circulation of micro-vessels in organs and lesions in a non-invasive manner. However, current CEUS-based thyroid nodule diagnosis methods suffered from: 1) the blurred spatial boundaries between nodules and other anatomies in CEUS videos, and 2) the insufficient representations of the local structural information of nodule tissues by the features extracted only from CEUS videos. In this paper, we propose a novel dual-branch network with a cross-modality-attention mechanism for thyroid nodule diagnosis by integrating the information from tow related modalities, i.e., CEUS videos and ultrasound image. The mechanism has two parts: US-attention-from-CEUS transformer (UAC-T) and CEUS-attention-from-US transformer (CAU-T). As such, this network imitates the manner of human radiologists by decomposing the diagnosis into two correlated tasks: 1) the spatio-temporal features extracted from CEUS are hierarchically embedded into the spatial features extracted from US with UAC-T for the nodule segmentation; 2) the US spatial features are used to guide the extraction of the CEUS spatio-temporal features with CAU-T for the nodule classification. The two tasks are intertwined in the dual-branch end-to-end network and optimized with the multi-task learning (MTL) strategy. The proposed method is evaluated on our collected thyroid US-CEUS dataset. Experimental results show that our method achieves the classification accuracy of 86.92%, specificity of 66.41%, and sensitivity of 97.01%, outperforming the state-of-the-art methods. As a general contribution in the field of multi-modality diagnosis of diseases, the proposed method has provided an effective way to combine static information with its related dynamic information, improving the quality of deep learning based diagnosis with an additional benefit of explainability.
Jianning Chi, Xiaosheng Yu 0001, Wenjun Zhang 0005
IEEE J. Biomed. Health Informatics6
2025 Low-Dose CT Image Super-Resolution With Noise Suppression Based on Prior Degradation Estimator and Self-Guidance Mechanism
abstract
The anatomies in low-dose computer tomography (LDCT) are usually distorted during the zooming-in observation process due to the small amount of quantum. Super-resolution (SR) methods have been proposed to enhance qualities of LDCT images as post-processing approaches without increasing radiation damage to patients, but suffered from incorrect prediction of degradation information and incomplete leverage of internal connections within the 3D CT volume, resulting in the imbalance between noise removal and detail sharpening in the super-resolution results. In this paper, we propose a novel LDCT SR network where the degradation information self-parsed from the LDCT slice and the 3D anatomical information captured from the LDCT volume are integrated to guide the backbone network. The prior degradation estimator (PDE) is proposed following the contrastive learning strategy to estimate the degradation features in the LDCT images without paired low-normal dose CT images. The self-guidance fusion module (SGFM) is designed to capture anatomical features with internal 3D consistencies between the squashed images along the coronal, sagittal, and axial views of the CT volume. Finally, the features representing degradation and anatomical structures are integrated to recover the CT images with higher resolutions. We apply the proposed method to the 2016 NIH-AAPM Mayo Clinic LDCT Grand Challenge dataset and our collected LDCT dataset to evaluate its ability to recover LDCT images. Experimental results illustrate the superiority of our network concerning quantitative metrics and qualitative observations, demonstrating its potential in recovering detail-sharp and noise-free CT images with higher resolutions from the practical LDCT images.
Jianning Chi, Zhiyi Sun, Liuyi Meng, Xiaosheng Yu 0001, Xiaolin Wei
IEEE Trans. Medical Imaging5
2025 Enhancing pixel-level analysis in medical imaging through visual instruction tuning: introducing PLAMi
Maocheng Bai, Xiaosheng Yu 0001, Ying Wang 0152, Jubo Chen, Pengfei Lyu
Vis. Comput.2
2025 CMT-6D: a lightweight iterative 6DoF pose estimation network based on cross-modal Transformer
Suyi Liu, Chengdong Wu 0001, Jianning Chi, Xiaosheng Yu 0001, Longxing Wei, Chuanjiang Leng
Vis. Comput.5
2024 Ensembling disentangled domain-specific prompts for domain generalization
Fangbin Xu, Shizhuo Deng, Tong Jia 0001, Xiaosheng Yu 0001, Dongyue Chen 0001
Knowl. Based Syst.4
2024 Progressive local-to-global vision transformer for occluded face hallucination
Jianning Chi, Chengdong Wu 0001, Xiaosheng Yu 0001, Hao Wu 0064
Multim. Tools Appl.4
2024 Cross-modal collaborative propagation for RGB-T saliency detection
Xiaosheng Yu 0001, Jianning Chi
Vis. Comput.1
2023 Low-Dose CT Image Super-Resolution Network with Dual-Guidance Feature Distillation and Dual-Path Content Communication
Jianning Chi, Zhiyi Sun, Tianli Zhao, Xiaosheng Yu 0001, Chengdong Wu 0001
MICCAI (10)5
2023 Cross-view information interaction and feedback network for face hallucination
Jianning Chi, Chengdong Wu 0001, Xiaosheng Yu 0001, Hao Wu 0064
J. Vis. Commun. Image Represent.4
2023 Unsupervised Multi-Subclass Saliency Classification for Salient Object Detection
abstract
Numerous bottom-up salient object detection algorithms formulate the problem as a classification task. For an input image, these methods usually utilize prior cues to select some regions as training set, and learn a classifier to classify all regions into foreground/background. However, such binary classification based approaches suffer from accuracy problems in some complex scenes. To this end, we propose a novel framework, namely Multi-Subclass Classification with Label Distribution Learning (MSCLDL). Specifically, prior knowledge is firstly employed to build a training set from input image, in which each sample is associated with one of two class labels. Previous works usually learn directly a binary classification model from training set. Different with them, we further decompose two classes into a certain number of subclasses, each sample is thus described by one of multiple subclass labels. Based on the multi-subclass training set, we learn a label distribution model to predict the subclass label of each image region. Furthermore, the saliency value of each image region could be computed via exploring the relationship class and subclass labels. The MSCLDL could overcome the limitation of existing classification-based algorithms in some challenging scenes. Finally, a novel refinement technology is presented to further refine the saliency map obtained by MSCLDL. We compare the proposed method and other state-of-the-art methods on four benchmark datasets, the superiority of our model is adequately demonstrated via the experimental results analysis.
Chengdong Wu 0001, Hao Wu 0064, Xiaosheng Yu 0001
IEEE Trans. Multim.4
2023 Over-sampling strategy-based class-imbalanced salient object detection and its application in underwater scene
Chengdong Wu 0001, Hao Wu 0064, Xiaosheng Yu 0001
Vis. Comput.4
2022 Optic disc detection based on fully convolutional neural network and structured matrix decomposition
Ying Wang 0152, Xiaosheng Yu 0001, Chengdong Wu 0001
Multim. Tools Appl.2
2022 MID-UNet: Multi-input directional UNet for COVID-19 lung infection segmentation from CT images
Jianning Chi, Xiaoying Han, Chengdong Wu 0001, Xiaosheng Yu 0001
Signal Process. Image Commun.6
2021 Brain tumor segmentation in MR images using a sparse constrained level set algorithm
Xiaoliang Lei, Xiaosheng Yu 0001, Jianning Chi, Ying Wang 0152, Jingsi Zhang, Chengdong Wu 0001
Expert Syst. Appl.2
2021 Image super-resolution using multi-granularity perception and pyramid attention networks
Chengdong Wu 0001, Jianning Chi, Xiaosheng Yu 0001, Hao Wu 0064
Neurocomputing4
2021 Underwater image super-resolution using multi-stage information distillation networks
Hao Wu 0064, Jianning Chi, Xiaosheng Yu 0001, Chengdong Wu 0001
J. Vis. Commun. Image Represent.5
2021 DCLNet: Dual Closed-loop Networks for face super-resolution
Chengdong Wu 0001, Jianning Chi, Xiaosheng Yu 0001, Hao Wu 0064
Knowl. Based Syst.5
2020 Bagging-based saliency distribution learning for visual saliency detection
Xiaosheng Yu 0001, Yunhe Wu, Chengdong Wu 0001
Signal Process. Image Commun.2
2019 Saliency detection via integrating deep learning architecture and low-level features
Jianning Chi, Chengdong Wu 0001, Xiaosheng Yu 0001, Hao Chu
Neurocomputing3
2019 Salient object detection based on novel graph model
Xiaosheng Yu 0001, Ying Wang 0152, Chengdong Wu 0001
J. Vis. Commun. Image Represent.2
2012 A New Method for Hand Detection Based on Hough Forest
Dongyue Chen 0001, Zongwen Chen, Xiaosheng Yu 0001
ISNN (2)3
2012 A New Method of Edge Detection Based on PSO
Dongyue Chen 0001, Xiaosheng Yu 0001
ISNN (2)3
2012 A Novel Method of River Detection for High Resolution Remote Sensing Image Based on Corner Feature and SVM
Ziheng Tian, Chengdong Wu 0001, Dongyue Chen 0001, Xiaosheng Yu 0001
ISNN (2)4
2012 A Remote Sensing Image Matching Algorithm Based on the Feature Extraction
Chengdong Wu 0001, Dongyue Chen 0001, Xiaosheng Yu 0001
ISNN (2)4
2012 Gradient Vector Flow Based on Anisotropic Diffusion
Xiaosheng Yu 0001, Chengdong Wu 0001, Dongyue Chen 0001, Tong Jia 0001
ISNN (2)1