Xin Shu 0001

dblp:23/4309-1 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0003-1079-8434ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LMFF-Net: Lightweight Multiscale Feature Fusion Object Detection for Underwater IoT Networks
abstract
With the advancement of Underwater Internet of Things (U-IoT) networks, underwater object detection has become crucial. However, detection performance is severely hindered by underwater image degradation, small object scales and weak textures, as well as the stringent resource limitations of deployment platforms. Meanwhile, existing detectors rely on heavy backbones and standard downsampling, which not only increases computation but also discards fine-grained spatial information, resulting in low precision for small objects and suboptimal deployability. To address these challenges, a lightweight multi-scale feature fusion network named LMFF-Net is proposed. Utilizing RepGhostNet as the feature extraction network and incorporating an enhanced bidirectional feature pyramid structure, the proposed LMFF-Net effectively reduces the number of parameters and computational cost while maintaining detection accuracy. In the feature pyramid, a multi-scale object enhancement (MSOE) module is designed, which synergistically combines spatial multi-scale convolutions with spatial-frequency attention mechanisms. This synergy enhances feature discriminability and improves small object representation. Furthermore, a channelized spatial fusion (CSF) module is constructed to achieve lossless downsampling of feature maps by leveraging the spatial-to-depth transformation principle, thereby maximizing the retention of fine-grained spatial information in deep networks. Experimental results on the DUO dataset show that LMFF-Net achieves an [email protected] of 84.7% with only 1.73 M parameters, significantly outperforming other classic object detection models. Additionally, generalization experiments on the URPC2020 and RUOD datasets also demonstrate the excellent generalization ability of the proposed model.
Xiushuai Xu, Zhibin Xie, Xin Shu 0001, Chang-Bin Shao
IEEE Internet Things J.4
2026 DSW-Net: A dual-skip connection wavelet network for underwater image enhancement
Xin Shu 0001, Chang-Bin Shao, Zhibin Xie
Knowl. Based Syst.2
2026 SADST: Style-aware dynamic style transfer for domain generalized semantic segmentation
Jingxian Shen, Xin Shu 0001, Linbin Pang, Zhongbin Zhang
Neural Networks5
2026 High-Frequency Information Supported Domain Adaptation for Cross-Domain Object Detection
Chang-Bin Shao, Zhibin Xie, Xin Shu 0001, Hualong Yu
IEEE Signal Process. Lett.4
2026 Enhancing defect detection in photovoltaic cells: a dynamic group YOLOv8 approach
Huhao Shen, Xin Shu 0001, Xiaofang Guo, Chang-Bin Shao, Zhibin Xie
Vis. Comput.2
2025 DA-Net: Deep attention network for biomedical image segmentation
Yingyan Gu, Yan Wang 0112, Xin Shu 0001
Signal Process. Image Commun.4
2025 Extended Receptive Field UDA Semantic Segmentation Based on Spatial Alignment and Knowledge Distillation
abstract
In recent years, unsupervised domain adaptation (UDA) has significantly advanced, addressing the issue of requiring large amounts of labeled data in deep learning. Some UDA strategies have effectively alleviated domain shift, but they are still struggling to tackle challenges such as the model’s inability to continuously extract features, the difficulty to achieve better segmentation boundaries in the target domain, and tend to neglect previously acquired knowledge. To address these issues, we propose an extended receptive field UDA semantic segmentation based on spatial alignment and knowledge distillation (ERF). Firstly, based on the idea of combining serial and parallel, we design a novel large continuous receptive field decoder (largeCF) to extract large continuous receptive field features. This approach alleviates the bias of model in feature extraction between objects of different sizes and simultaneously reduces model complexity. Secondly, we propose an edge consistency strategy that aligns edge features and matches the spatial arrangement between predicted and ground truth labels, improving edges segmentation accuracy of the target domain. Finally, we employ a knowledge distillation module to achieve an optimized student-teacher framework, where the teacher effectively guides the student to retain previously learned information, resulting in more accurate segmentation of the target domain. Experimental results demonstrated the effectiveness of the proposed approach, which achieved mIoU of 76.5% and 68.8% on UDA benchmark tasks GTA$\rightarrow $CityScapes and SYNTHIA$\rightarrow $CityScapes, respectively. The code is available at:https://github.com/fz-ss/ERF. Note to Practitioners—This paper focuses on the challenges of semantic segmentation in autonomous driving, particularly facing the issue of extensive manual annotation required for dense semantic labels. We propose a novel extended receptive field UDA semantic segmentation based on spatial alignment and knowledge distillation. The article begins by outlining the initial implementation process of UDA, laying the foundation for subsequent in-depth discussions. Subsequently, we conduct a theoretical analysis of the Large Continuous Decoder, Boundary Consistency Strategy, and Knowledge Distillation Scheme, which constitute the core components of our method. Finally, experiments on two UDA benchmark tasks demonstrate the feasibility of our approach. However, the performance gap between UDA and supervised semantic segmentation still exists. In future research, we will focus on reducing the feature gap between different domains. Additionally, we will strive to fully use image features and distilled features to make greater progress, thereby driving advancements in semantic segmentation for autonomous driving.
Yunna Song, Caisheng Liu, Suqin Bai, Xin Shu 0001, Yunhan Sun
IEEE Trans Autom. Sci. Eng.6
2024 CSCA U-Net: A channel and space compound attention CNN for medical image segmentation
Xin Shu 0001, Aoping Zhang
Artif. Intell. Medicine1
2024 GCCF: A lightweight and scalable network for underwater image enhancement
Chufan Liu, Xin Shu 0001
Eng. Appl. Artif. Intell.2
2024 MPFC-Net: A multi-perspective feature compensation network for medical image segmentation
Xianghu Wu, Shucheng Huang, Xin Shu 0001, Chunlong Hu, Xiaojun Wu 0001
Expert Syst. Appl.3
2024 EPM-Net: Efficient Feature Extraction, Point-Pair Feature Matching for Robust 6-D Pose Estimation
abstract
Estimating the 6-D poses of objects from RGB-D images holds great potential for several applications. However, given that the 6-D pose estimation accuracy is significantly affected by occlusion and noise between the objects in an image, this paper proposes a novel 6-D pose estimation method based on Efficient feature extraction and Point-pair feature matching. Specifically, we develop the Efficient channel attention Convolutional Neural Network (ECNN) and SO(3)-Encoder modules to extract 2-D features from the RGB image and SO(3)-equivariant features from the depth image, respectively. These features are fused in the DenseFusion module to obtain 3-D features in the camera space. Meanwhile, we exploit CAD model priors to obtain 3-D features in the model space through the model feature encoder, and then we globally regress the 3-D features in the camera and model space. According to these features, we generate oriented point clouds in each space, and then conduct point-pair feature matching to obtain pose information. Finally, we perform direct pose regression on the 3-D features in the camera and model space, and then resulting point-pair feature matching pose information is combined with the direct point-wise pose regression information to enhance pose prediction accuracy. Experimental results on three widely used benchmarking datasets demonstrate that our method achieves state-of-the-art performance, particularly for severe occluded scenes.
Danping Zou, Xin Shu 0001, Suqin Bai, Haowei Zhu, Yunhan Sun
IEEE Trans. Multim.4
2023 Face spoofing detection based on multi-scale color inversion dual-stream convolutional neural network
Xin Shu 0001
Expert Syst. Appl.1
2023 HCPSNet: heterogeneous cross-pseudo-supervision network with confidence evaluation for semi-supervised medical image segmentation
Xianhua Duan, Chaoqiang Jin, Xin Shu 0001
Multim. Syst.3
2023 RES-CapsNet: an improved capsule network for micro-expression recognition
Xin Shu 0001, Shucheng Huang
Multim. Syst.1
2022 Using global information to refine local patterns for texture representation and classification
Xin Shu 0001, Xiaoning Song, Xiaojun Wu 0001
Pattern Recognit.1
2021 Face spoofing detection based on chromatic ED-LBP texture feature
Xin Shu 0001, Shucheng Huang
Multim. Syst.1
2021 Multiple channels local binary pattern for color texture representation and classification
Xin Shu 0001, Zhigang Song, Shucheng Huang, Xiaojun Wu 0001
Signal Process. Image Commun.1
2017 Converted-face identification: using synthesized images to replace original images for recognition
Chang-Bin Shao, Xiaoning Song, Xin Shu 0001, Xiaojun Wu 0001
Multim. Tools Appl.3
2015 Multi-scale contour flexibility shape signature for Fourier descriptor
Xin Shu 0001, Lei Pan 0009, Xiaojun Wu 0001
J. Vis. Commun. Image Represent.1
2011 A novel contour descriptor for 2D shape matching and its application to image retrieval
Xin Shu 0001, Xiaojun Wu 0001
Image Vis. Comput.1