Shuai Yuan 0013

dblp:19/1243-13 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0002-1031-6239ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LPL3D: LVLM-Driven Pseudo-Labeling for 3D Object Detection
abstract
Effective 3D object detection requires large-scale annotated datasets, which are expensive and time-consuming to produce - especially in indoor environments containing dense object arrangements. To address this, we propose an Large Vision-Language Model (LVLM)-driven automatic high-quality pseudo-label generation technique for 3D object detection in single- and multi-view scenarios. We propose an Auto3DLabeler that introduces the first-ever text-to-3D Bounding Box transformation. Its pipeline employs a text-based detector, a segmenter and an LVLM to generate annotation estimates, which are further refined by our IoU-guided iterative Box Aggregator and layout-aware prompt Class Refiner modules. We also introduce a semantic-enhanced multi-modal fusion module that integrates image-level semantic information into point cloud representations for precise detections. Collectively, our contributions provide a remarkable boost to the 3D object detection state-of-the-art. Extensive experiments on SUN RGB-D and ScanNet datasets show our unsupervised detector variant outperforming existing semi-supervised detectors, and our semi-supervised variant achieving up to 28.2% absolute gain in challenging scenarios - all this while maintaining considerable compute advantage over existing label-efficient methods. Our code and models will be made public for the community. Our code and model will be made public after acceptance.
Zechuan Li, Hongshan Yu, Yihao Ding, Shuai Yuan 0013, Naveed Akhtar
IEEE Trans. Circuits Syst. Video Technol.4
2025 HDTCNet: A hybrid-dimensional convolutional network for multivariate time series classification
Yongli Gu, Hanlin Qin, Naveed Akhtar, Shuai Yuan 0013, Honghao Fu, Shuowen Yang, Ajmal Mian
Pattern Recognit.5
2025 LCIRE-Net: Lightweight Cross-Modal Information Interaction for Road Feature Extraction From Remote Sensing Images and GPS Trajectory/LiDAR
abstract
Due to obstructions such as trees and buildings, single-modal satellite or aerial images are insufficient for continuous high-precision representation of road features. To address this problem, this article proposes a lightweight cross-modal information interaction for road feature extraction (LCIRE-Net) from high-resolution remote sensing images (HRSIs) and GPS trajectory/LiDAR images. We design two parallel encoders for modality feature learning, using pairs of multimodal information as inputs to the encoders. By designing a cross-modal information dynamic interaction (CMIDI) mechanism, thresholds are used to decide whether to supplement redundant information from another modality, solving the issue of ineffective fusion calculations due to minor differences in multimodal feedback. A multimodal feature fusion module (MFFM) is proposed after the encoder output to achieve effective dual-modal fusion while addressing the interference of redundant noise generated during extraction. Subsequently, we present the feature refinement and enhancement module (FREM), which successfully captures edge features of the image using the receptive field of dilated convolution kernels. Additionally, in terms of lightweight design, we employ a novel SOTA method on D-LinkNet by replacing the original residual blocks with an enhanced ghost basic block. Extensive experiments are conducted on the BJRoad, Porto, and TLCGIS datasets, demonstrating that our network, with smaller parameters and FLOPs, outperforms other road-based semantic segmentation methods.
Yifei Duan, Dan Yang 0006, Xiaochen Qu, Lu Chao, Peilu Gan, Shuai Yuan 0013, Hanlin Qin, Junsuo Qu
IEEE Trans. Geosci. Remote. Sens.7
2025 Rethinking Generalizable Infrared Small Target Detection: A Real-Scene Benchmark and Cross-View Representation Learning
abstract
Infrared small target detection (ISTD) is highly sensitive to sensor type, observation conditions, and the intrinsic properties of the target. These factors can introduce substantial variations in the distribution of acquired infrared image data, a phenomenon known as domain shift. Such distribution discrepancies significantly hinder the generalization capability of ISTD models across diverse scenarios. To tackle this challenge, this paper introduces an ISTD framework enhanced by domain adaptation. To alleviate distribution shift between datasets and achieve cross-sample alignment, we introduce Cross-view Channel Alignment (CCA). Additionally, we propose the Cross-view Top-K Fusion strategy, which integrates target information with diverse background features, enhancing the model’s ability to extract critical data characteristics. To further mitigate the impact of noise on ISTD, we develop a Noise-guided Representation learning strategy. This approach enables the model to learn more noise-resistant feature representations, to improve its generalization capability across diverse noisy domains. Finally, we develope a dedicated infrared small target dataset, RealScene-ISTD. Compared to state-of-the-art methods, our approach demonstrates superior performance in terms of detection probability (Pd), false alarm rate (Fa), and intersection over union (IoU). The code is available at: https://github.com/luy0222/RealScene-ISTD.
Yahao Lu, Yuehui Li, Xingyuan Guo, Shuai Yuan 0013, Yukai Shi, Liang Lin 0004
IEEE Trans. Geosci. Remote. Sens.4
2025 DRPCA-Net: Make Robust PCA Great Again for Infrared Small Target Detection
abstract
Infrared small target detection plays a vital role in remote sensing, industrial monitoring, and various civilian applications. Despite recent progress powered by deep learning, many end-to-end convolutional models tend to pursue performance by stacking increasingly complex architectures, often at the expense of interpretability, parameter efficiency, and generalization. These models typically overlook the intrinsic sparsity prior of infrared small targets–an essential cue that can be explicitly modeled for both performance and efficiency gains. To address this, we revisit the model-based paradigm of Robust Principal Component Analysis (RPCA) and propose Dynamic RPCA Network (DRPCA-Net), a novel deep unfolding network that integrates the sparsity-aware prior into a learnable architecture. Unlike conventional deep unfolding methods that rely on static, globally learned parameters, DRPCA-Net introduces a dynamic unfolding mechanism via a lightweight hypernetwork. This design enables the model to adaptively generate iteration-wise parameters conditioned on the input scene, thereby enhancing its robustness and generalization across diverse backgrounds. Furthermore, we design a Dynamic Residual Group (DRG) module to better capture contextual variations within the background, leading to more accurate low-rank estimation and improved separation of small targets. Extensive experiments on multiple public infrared datasets demonstrate that DRPCA-Net significantly outperforms existing state-of-the-art methods in detection accuracy. Code is available at https://github.com/GrokCV/DRPCA-Net.
Zihao Xiong, Fei Zhou 0006, Fengyi Wu, Shuai Yuan 0013, Maixia Fu, Zhenming Peng, Jian Yang 0003, Yimian Dai
IEEE Trans. Geosci. Remote. Sens.4
2025 ASCNet: Asymmetric Sampling Correction Network for Infrared Image Destriping
abstract
In a real-world infrared (IR) imaging system, effectively learning a consistent stripe noise removal model is essential. Most existing destriping methods cannot precisely reconstruct images due to cross-level semantic gaps and insufficient characterization of the global column features. To tackle this problem, we propose a novel IR image destriping method, called asymmetric sampling correction network (ASCNet), that can effectively capture global column relationships and embed them into a U-shaped framework, providing comprehensive discriminative representation and seamless semantic connectivity. Our ASCNet consists of three core elements: residual Haar discrete wavelet transform (RHDWT), pixel shuffle (PS), and column nonuniformity correction module (CNCM). Specifically, RHDWT is a novel downsampler that employs double-branch modeling to effectively integrate stripe-directional prior knowledge and data-driven semantic interaction to enrich the feature representation. Observing the semantic patterns crosstalk of stripe noise, PS is introduced as an upsampler to prevent excessive a priori decoding and performing semantic-bias-free image reconstruction. After each sampling, CNCM captures the column relationships in long-range dependencies. By incorporating column, spatial, and self-dependence information, CNCM well establishes a global context to distinguish stripes from the scene’s vertical structures. Extensive experiments on synthetic data, real data, and IR small target detection (IRSTD) tasks demonstrate that the proposed method outperforms state-of-the-art single-image destriping methods both visually and quantitatively. The code is available athttps://github.com/xdFai/ASCNet.
Shuai Yuan 0013, Hanlin Qin, Shiqi Yang 0001, Shuowen Yang, Naveed Akhtar, Huixin Zhou
IEEE Trans. Geosci. Remote. Sens.1
2024 SCTransNet: Spatial-Channel Cross Transformer Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) has recently benefitted greatly from U-shaped neural models. However, largely overlooking effective global information modeling, existing techniques struggle when the target has high similarities with the background. We present aSpatial-channelCrossTransformerNetwork (SCTransNet) that leverages spatial-channel cross transformer blocks (SCTBs) on top of long-range skip connections to address the aforementioned challenge. In the proposed SCTBs, the outputs of all encoders are interacted with cross transformer to generate mixed features, which are redistributed to all decoders to effectively reinforce semantic differences between the target and clutter at full levels. Specifically, SCTB contains the following two key elements: (a) spatial-embedded single-head channel-cross attention (SSCA) for exchanging local spatial features and full-level global channel information to eliminate ambiguity among the encoders and facilitate high-level semantic associations of the images, and (b) a complementary feed-forward network (CFN) for enhancing the feature discriminability via a multi-scale strategy and cross-spatial-channel information interaction to promote beneficial information transfer. Our SCTransNet effectively encodes the semantic differences between targets and backgrounds to boost its internal representation for detecting small infrared targets accurately. Extensive experiments on three public datasets, NUDT-SIRST, NUAA-SIRST, and IRSTD-1K, demonstrate that the proposed SCTransNet outperforms existing IRSTD methods. Our code will be made public at https://github.com/xdFai/SCTransNet.
Shuai Yuan 0013, Hanlin Qin, Naveed Akhtar, Ajmal Mian
IEEE Trans. Geosci. Remote. Sens.1
2024 IRSTDID-800: A Benchmark Analysis of Infrared Small Target Detection-Oriented Image Destriping
abstract
Deep learning-based single-image infrared (IR) destriping has made significant advances. However, these methods are typically evaluated using “synthetic” images with specific stripe noise, making it unclear how well they handle “real” IR images. In fact, a clear and fair benchmarking of the existing destriping methods on real images, especially for the downstream IR small target detection (IRSTD) task, is currently an open gap. To tackle this problem, we introduce a novel benchmark, called IRSTD-oriented image destriping (IRSTDID-800), which thoroughly showcases the real distribution of IR small targets under stripe noise perturbation for the first time. Concretely, it consists of two subsets. (1) IRSTDID-SKY: composed of 500 real-world images afflicted with stripe noise, including unmanned aerial vehicles (UAVs) of various shapes, sizes, and contrasts. Moreover, these images are annotated with precise pixel levels for objective evaluation of IRSTD. (2) IRSTDID-GND: comprising 300 real-world images featuring common daily life objects, providing a richer scene under stripe noise. Based on the proposed IRSTDID-800, we comprehensively assess the performance of ten state-of-the-art (SOTA) destriping methods across eleven metrics, including full-reference, no-reference, and task-driven metrics with six advanced IRSTD methods. Furthermore, inspired by the correlation between image quality assessment and IRSTD, we proposed a task-oriented destriping optimization strategy. A loss function is introduced for IR image destriping, leveraging the structural properties of noise as a penalty term to strengthen image destriping and IRSTD. Overall, our analysis reveals interesting observations to guide future research in destriping and IRSTD tasks. Our dataset is available athttps://github.com/xdFai/IRSTDID-800.
Shuai Yuan 0013, Hanlin Qin, Naveed Akhtar, Shiqi Yang 0001, Shuowen Yang
IEEE Trans. Geosci. Remote. Sens.1