Siyuan Wang 0011

dblp:12/9626-11 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-5506-7451ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NSBRNet: Non-Local Spatio-Temporal Bidirectional Recurrent Network for Satellite Video Super-Resolution
abstract
In recent years, intelligent processing of satellite videos has emerged as a significant research focus within the field of remote sensing, driven by the growing demand for enhanced spatial resolution. This need has led to increased interest in satellite video super-resolution (SVSR) algorithms, which aim to improve the quality of satellite imagery. However, many existing SVSR methods tend to neglect the global dependencies among frames in satellite videos, resulting in an incomplete utilization of spatio-temporal feature information. To tackle this issue, we propose a novel non-local spatio-temporal bidirectional recurrent network specifically designed for SVSR applications. Our approach employs a gate-guided deformable alignment module that effectively enhances feature alignment and fusion using a dynamic gating mechanism. This allows the network to adaptively focus on relevant features during the reconstruction process. Furthermore, we introduce a non-local spatio-temporal fusion module that integrates both temporal and spatial relationships over long sequences of frames, ensuring a comprehensive extraction of feature information. Through extensive experiments, our proposed method demonstrates superior performance compared to state-of-the-art SVSR techniques in terms of reconstruction quality. Additionally, it demonstrates outstanding performance in downstream satellite video applications, showcasing its potential in satellite video processing tasks. The source code is publicly available at https://github.com/Yu-Wang-0801/NSBRNet.
Yu Wang 0140, Xiaolong Zuo, Tao Lu 0001, Jiaming Wang 0001, Yuankun Wang, Siyuan Wang 0011, Zhizheng Zhang 0009, Xiaojin Zhao
IEEE Trans. Circuits Syst. Video Technol.7
2025 Spatial-Frequency Multiple Feature Alignment for Cross-Domain Remote Sensing Scene Classification
abstract
Domain adaptation is a pivotal technique for improving the classification performance of remote sensing scenes impacted by data distribution shifts. Existing spatial-domain feature alignment methods are vulnerable to complex scene clutter and spectral variations. Considering the robustness of frequency representation in preserving edge details and structural patterns, this paper presents a novel spatial-frequency multiple alignment domain adaptation (SFMDA) method for remote sensing scene classification. First, a frequency-domain invariant feature learning module is introduced, which employs the Fourier transform and high-frequency mask strategy to derive frequency-domain features exhibiting enhanced inter-domain invariance. Subsequently, a spatial-frequency feature cross fusion module is developed to achieve more robust and domain-representative spatial-frequency fusion representations through dot product attention and interaction mechanisms. Finally, a multiple feature alignment strategy is devised to minimize both spatial-domain feature differences and fusion feature discrepancies across the source and target domains, thereby facilitating more effective inter-domain knowledge transfer. Experimental results on six cross-domain scenarios demonstrate that SFMDA outperforms eight state-of-the-art methods, achieving a 3.87%–17.98% accuracy improvement. Furthermore, SFMDA is compatible with existing spatial-domain learning frameworks, enabling seamless integration for further performance gains. Our code will be available at https://github.com/GeoRSAI/SFMDA.
Dongyang Hou, Siyuan Wang 0011, Xiaoguang Zhou, Wei Wang 0107
IEEE Geosci. Remote. Sens. Lett.3
2025 CANet: A Spatial Structure Constraint and Local Semantic Awareness Based Network for Weakly Supervised Building Extraction
abstract
Benefitting from the easy availability of image-level labels, weakly supervised semantic segmentation (WSSS) methods based on class activation maps (CAMs) have made significant progress in building extraction from remote sensing imagery. However, image-level labels lack precise spatial locations and boundary ranges of buildings, posing challenges in achieving comprehensive and structurally clear building extraction. Furthermore, due to the complex background interference and the diversity of building in high-resolution remote sensing imagery (HRRS), small and sparse buildings suffer from insufficient attention in CAMs. To solve the above problems, this article proposes a spatial structure constraint and local semantic awareness-based WSSS method, CANet, for extracting buildings from HRRS. Specifically, we design a spatial structure constraint module to generate CAMs with finer spatial structural details of buildings, which minimizes feature differences from patches of different granularities and the whole image. Moreover, a local semantic awareness module is designed to address the issue of insufficient coverage of CAMs on sparse and tiny building. This module first strengthens the feature extraction network by embedding discriminative suppression units to force the network to focus on more nondiscriminative regions. Subsequently, visual word learning is introduced to identify additional object categories. Finally, four WSSS datasets are constructed based on public datasets with two simple and two complex scenarios. The results demonstrate that the proposed method outperforms 11 state-of-the-art methods, improving the intersection over union by at least 3.46% and 0.29% in both simple and complex scenarios, respectively.
Siyuan Wang 0011, Dongyang Hou, Yu Wang 0140, Bowen Cai 0002
IEEE Trans. Geosci. Remote. Sens.1
2025 Revisiting the Learning Stage in Range View Representation for Autonomous Driving
abstract
LiDAR segmentation is crucial for autonomous driving perception. Range view methods have been widely adopted for these applications due to their intuitiveness and ease of implementation. However, the inherent shortcomings of the range view approach (e.g., assuming that point clouds within the same pixel of a range image have the same semantic class) make it difficult to perform accurate fine-grained segmentation tasks, thus limiting its potential in practical applications. To address these issues, we propose RangeFusion, an end-to-end framework that greatly improves the ability to learn and process LiDAR point clouds from range views by employing a multispatial learning model. A novel range-scan space (RSS) is proposed to address the inability of existing range view methods to accurately aggregate features of neighboring points. This space achieves accurate and efficient neighboring point feature aggregation with linear time complexity. In addition, a supervised label smoothing method called multilevel feature selection heads (MFSHs) is designed, which achieves more fine-grained semantic prediction by subdividing the full point cloud into multisemantic hierarchical subclouds and adaptively fusing the features with confidence filtering. The performance of the proposed method was evaluated on several benchmarks, including SemanticKITTI and nuScenes. On these two datasets, mean intersection over union (mIoU) scores of 67.9% and 80.2% were achieved, respectively. This demonstrates that the proposed approach outperforms existing range view- and multiview-based approaches while maintaining efficient performance at 26.5 FPS. In addition, real road data were collected for testing. The code is available athttps://github.com/Wansit99/RangeFusion.
Jinsheng Xiao, Siyuan Wang 0011, Jian Zhou 0011, Ziyin Zeng, Ruijia Chen
IEEE Trans. Geosci. Remote. Sens.2
2024 MF-BHNet: A Hybrid Multimodal Fusion Network for Building Height Estimation Using Sentinel-1 and Sentinel-2 Imagery
abstract
Integrated Sentinel-1 synthetic aperture radar (SAR) imagery and Sentinel-2 optical imagery have shown great promise in mapping large-scale building height. Effectively fusing the complementary features of SAR and optical imagery is a key challenge in enhancing the building height estimation performance. However, SAR imagery and optical imagery have significant heterogeneity, which makes obtaining accurate building height a challenging problem. In this article, we propose a hybrid multimodal fusion network (MF-BHNet) for building height estimation using Sentinel-1 SAR imagery and Sentinel-2 optical imagery. First, we design a hybrid multimodal encoder to mine modal-specific feature and model intermodal correlation. In particular, an intramodal encoder (IME) is designed to reconstruct valuable intramodal information, and a transformer-based cross-modal encoder (CME) is used to model intermodal correlation and capture contextual information. Then, a coarse-fine progressive multimodal fusion method is proposed to fuse SAR feature and optical feature to improve the building height estimation performance. We construct a building height dataset by introducing superior building footprints to validate our method. Experimental results demonstrate that our MF-BHNet method outperforms the compared 11 state-of-the-art methods, which achieves the lowest root-mean-square error (RMSE) of 3.6421 m. Besides, compared to the four publicly available building height products, the mapping result of the proposed method has significant advantages in terms of spatial detail and accuracy.
Siyuan Wang 0011, Bowen Cai 0002, Dongyang Hou, Jiaming Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 PCLUDA: A Pseudo-Label Consistency Learning- Based Unsupervised Domain Adaptation Method for Cross-Domain Optical Remote Sensing Image Retrieval
abstract
Recent advances in deep learning have dramatically improved the performance of content-based remote sensing image retrieval (CBRSIR) with the same distribution of training set (source domain) and test set (target domain). In fact, their distributions are inconsistent in most cases, which can lead to a dramatic decrease in retrieval performance. Currently, some unsupervised domain adaptation (DA) methods for other remote sensing applications have been proposed to eliminate the inconsistency. However, the current unsupervised DA methods do not make full use of the target domain’s distribution characteristics when delineating its decision boundary. This tends to degrade the cross-domain retrieval performance. In this article, a pseudo-label consistency learning-based unsupervised DA method (PCLUDA) is proposed for cross-domain CBRSIR. Our PCLUDA method minimizes the difference in probability distribution between the target domain and its perturbed output by a pseudo-label self-training and consistency regularization strategy, followed by adjusting the target domain’s decision boundaries to the low-density region. Besides, minimize class confusion (MCC) is introduced to reduce negative transfer caused by large intraclass variance of RSIs. Two cross-domain datasets with 12 cross-domain scenarios are constructed based on six open access datasets to measure DA methods. Experimental results show that our PCLUDA method achieves superior retrieval performances with average retrieval precision improvement by 4.9%–32.3% compared with eight state-of-the-art DA approaches in complex cross-domain scenarios. Furthermore, other experimental results indicate that our PCLUDA can also reach optimal retrieval performances in different kinds of deep learning networks [i.e., vision transformer (ViT) and convolutional neural networks (CNNs)].
Dongyang Hou, Siyuan Wang 0011, Xueqing Tian, Huaqiao Xing
IEEE Trans. Geosci. Remote. Sens.2
2023 A Self-Supervised-Driven Open-Set Unsupervised Domain Adaptation Method for Optical Remote Sensing Image Scene Classification and Retrieval
abstract
Unsupervised domain adaptation (UDA) is an important solution to reduce the bias between the labeled source domain and the unlabeled target domain. It has attracted more attention for optical remote sensing image scene classification and retrieval. Currently, most of the previous work is devoted to closed-set UDA. In fact, the target domain often contains unknown classes. Moreover, some open UDA methods mine structural information of the target domain directly from the type knowledge of the source domain, and less directly from the unlabeled data of the target domain. In this paper, we propose a new self-supervised-driven open-set UDA method combining contrastive self-supervised learning with consistency self-training for optical remote sensing scene classification and retrieval. Specifically, a contrastive self-supervised learning network is introduced to learn discriminative features from the unlabeled target domain data. Moreover, a novel open-set class learning module is developed based on two-level confidence rules and the consistency self-training strategy, which can obtain reliable unknown class samples for co-training. Finally, an open-set dataset including six cross-domain scenarios is constructed based on three public datasets and several experiments are conducted with eleven state-of-the-art domain adaptation methods. Experimental results demonstrate that our proposed method achieves superior performances on the six open-set cross-domain scenarios in both scene classification and retrieval. Especially, our method improves the overall classification accuracies by 9.72% to 24.06% and improve mean average retrieval precisions by 8.06% to 16.21% on the complex UCMD (source domain) → NWPU (target domain) scenario, compared with the other eleven state-of-the-art methods.
Siyuan Wang 0011, Dongyang Hou, Huaqiao Xing
IEEE Trans. Geosci. Remote. Sens.1