EDBT 2026 Demo / reviewers in the wild / expert
Wanshou Jiang
dblp:04/8500
· DBLP profile ↗
12ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-3162-0566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Hierarchical Local-Global-Aware Transformer With Scratch Learning Capabilities for Change DetectionabstractMost transformer-based methods rely on pretraining weights on large datasets such as Imagenet or pretraining from specific change detection (CD) datasets and then fine-tuning on the target dataset. When the target dataset significantly diverges from the dataset used for pretraining, the model’s ability to generalize to remote sensing imagery may be compromised due to the domain gap. In this letter, we propose HierFormer, which has the advantage of processing semantic features hierarchically, using simple operations for shallow features, spatial position transformation for middle-level features, and channel information interaction for high-level features. In addition, we propose a local-global-aware (LGA) attention block, which reduces the computational overhead of self-attention by sparse attention and increases the locality inductive bias (LIB) of the transformer by focusing attention on the local region and sparse part of the global region, which enables the model to be trained from scratch on small to medium-sized CD datasets. Finally, a new feature fusion decoder (FFD) is proposed to fuse the bitemporal features, which reweights the features of each channel through attention mechanism. Compared with other transformer-based or transformer-CNN-based hybrid networks, our method significantly improves F1, reaching 91.56% and 97.56% on the LEVIR-CD and CDD-CD change detection datasets. Our code is available athttps://github.com/WesternTrail/HierFormer. Wanshou Jiang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Extraction of Building Roofs and Facades Based on Axial Feature EnhancementabstractAffected by the vertical expansion of the city and non-nadir imaging, the roof of a building cannot completely overlap with its bottom footprint. Therefore, it is difficult to obtain the actual precise location of the building only by extracting the roof of the building. This paper proposes a new model framework that embeds the Axial Feature Enhancement Module in the network to represent the spatial dependence of the building roof and facade, and relies on the axial feature enhancement loss to constrain the network to learn the category distribution of the horizontal and vertical axes. As far as we know, this is the first time that the importance of building facade information is considered to carry out research on building roof and facade extraction. The experimental results also demonstrate the effectiveness of the method proposed in this article. Wanshou Jiang |
IGARSS | 2 |
| 2024 | Frequency Mining and Complementary Fusion Network for RGB-Infrared Object DetectionabstractIn recent years, object detection on visible (RGB) and infrared (IR) has gained significant attention as a promising solution for robust detection in complex scenarios, especially in low-light conditions. With the help of IR images, object detectors have become more reliable and robust in practical condition by combining the RGB and IR information. Despite significant progress in this field, current methods ignore the distinct characteristics of the two modalities when extracting features. RGB images contain detailed texture and color information, which means they have many high-frequency signals. Meanwhile, IR images have smoother textures and edges but clear shapes, indicating a significant amount of low-frequency information. We must consider the differences between the two modalities when extracting corresponding features. To address this issue, we propose a novel network architecture: the frequency mining and complementary fusion network (FMCFNet), which accounts for the intermodal variability. Our network contains two critical modules: the frequency feature extraction (FFE) module and the complementary fusion (CF) module. The FFE module utilizes filters of varying kernel and pooling sizes to extract features with diverse frequency information and then adaptively selects the most responsive frequency component. The CF module uses the similarity scores generated by cross attention to model the interactions between two modalities. Comprehensive experimental results demonstrate that our method can effectively combine RGB-IR complementary information, achieving robust detection results. Yangfeixiao Liu, Wanshou Jiang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | WHU-Stereo: A Challenging Benchmark for Stereo Matching of High-Resolution Satellite ImagesabstractStereo matching of high-resolution satellite images (HRSIs) is still a fundamental but challenging task in the field of photogrammetry and remote sensing. Recently, deep learning (DL) methods have demonstrated the potential for stereo matching on public benchmark datasets. Currently, mainstream stereo matching models depend on quantities of training data with ground truth. However, datasets for stereo matching of satellite images are scarce, which profoundly blocks the application of DL in this field. To facilitate further research, this article publishes a large-scale dataset, termed WHU-Stereo, for stereo matching DL network training and testing. This dataset is created by using airborne light detection and ranging (LiDAR) point clouds and high-resolution stereo imageries taken from the Chinese GaoFen-7 (GF-7) satellite. Occlusions should be seriously considered in the generation of ground-truth disparities. We propose an occlusion removal technique, which can adapt to point clouds with different densities and shows high potential in training data preparation. The WHU-Stereo dataset contains more than 1700 epipolar rectified image pairs, which cover six areas in China and includes various kinds of landscapes. We have assessed the accuracy of ground-truth disparity maps, and it is shown that our dataset achieves comparable precision compared with existing state-of-the-art stereo matching datasets. To verify its feasibility, in experiments, the handcrafted semiglobal matching (SGM) algorithm and recent DL networks have been tested on the dataset. Experimental results show that the WHU-Stereo dataset can serve as a challenging benchmark for stereo matching of HRSIs and performance evaluation of DL models. Our dataset is available athttps://github.com/Sheng029/WHU-Stereo. Shenhong Li, San Jiang, Wanshou Jiang, Lin Zhang 0036 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | MaskNet++: Inlier/outlier identification for two point clouds
Ruqin Zhou, Hanyun Wang, Xixing Li, Yulan Guo, Chenguang Dai, Wanshou Jiang |
Comput. Graph. | 6 |
| 2022 | Parallel Structure From Motion for UAV Images via Weighted Connected Dominating SetabstractIncremental Structure from Motion (ISfM) has been widely used for UAV image orientation. Its efficiency, however, decreases dramatically due to iterative BA (bundle adjustment). Although the divide-and-conquer strategy has been utilized for efficiency improvement, cluster merging becomes difficult or depends on seriously designed common image poses or 3D points. This paper proposes an algorithm to extract the global model for cluster merging and designs a parallel ISfM solution to achieve efficient and accurate image orientation. First, based on vocabulary tree retrieval, match pairs are selected to construct an undirected weighted match graph, whose edge weights are calculated by considering both the number and distribution of feature matches. Second, an algorithm termed weighted connected dominating set (WCDS), is designed to achieve the simplification of the match graph and build the global model, which incorporates the edge weight in the graph vertex selection and enables the successful reconstruction of the global model. Third, the match graph is simultaneously divided into compact and non-overlapped clusters. After the parallel reconstruction, cluster merging is conducted with the aid of the global model. Finally, by using three UAV datasets that are captured by classical oblique and recent optimized views photogrammetry, the validation of the proposed solution is verified through comprehensive analysis and comparison. The experimental results demonstrate that the proposed parallel ISfM can achieve 17.4 times efficiency improvement and comparative orientation accuracy. In absolute BA, the geo-referencing accuracy is approximately 2.0 and 3.0 times the GSD (Ground Sampling Distance) value in the horizontal and vertical directions, respectively. For parallel ISfM, the proposed solution is a more reliable alternative. San Jiang, Qingquan Li 0001, Wanshou Jiang, Wu Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Transformer and CNN Hybrid Deep Neural Network for Semantic Segmentation of Very-High-Resolution Remote Sensing ImageryabstractThis article presents a transformer and convolutional neural network (CNN) hybrid deep neural network for semantic segmentation of very high resolution (VHR) remote sensing imagery. The model follows an encoder–decoder structure. The encoder module uses a new universal backbone Swin transformer to extract features to achieve better long-range spatial dependencies modeling. The decoder module draws on some effective blocks and successful strategies of CNN-based models in remote sensing image segmentation. In the middle of the framework, an atrous spatial pyramid pooling block based on depthwise separable convolution (SASPP) is applied to obtain a multiscale context. A U-shaped decoder is used to gradually restore the size of the feature maps. Three skip connections are built between the encoder and decoder feature maps of the same size to maintain the transmission of local details and enhance the communication of multiscale features. A squeeze-and-excitation (SE) channel attention block is added before segmentation for feature augmentation. An auxiliary boundary detection branch is combined to provide edge constraints for semantic segmentation. Extensive ablation experiments were conducted on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam benchmarks to test the effectiveness of multiple components of the network. At the same time, the proposed method is compared with the current state-of-the-art methods on the two benchmarks. The proposed hybrid network achieved the second highest overall accuracy (OA) on both the Potsdam and Vaihingen benchmarks (code and models are available athttps://github.com/zq7734509/mmsegmentation-multilayer). Cheng Zhang 0037, Wanshou Jiang, Wei Wang 0323, Chenjie Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | SCANet: A Spatial and Channel Attention based Network for Partial-to-Partial Point Cloud Registration
Ruqin Zhou, Xixing Li, Wanshou Jiang |
Pattern Recognit. Lett. | 3 |
| 2014 | A Robust Image Fusion Method Based on Local Spectral and Spatial CorrelationabstractTo solve the potential color distortion problem of synthetic-variable-ratio method, an improved fusion method based on local spectral and spatial correlation (SSC) is presented. The proposed method, which uses SSC characteristics and local optimization strategy to simulate a low-resolution panchromatic image, can effectively reduce the spectral distortion of the fused image. QuickBird and other satellite images are used to assess the quality of the method. Visual and quantitative analysis demonstrates that the proposed approach can significantly improve the fusion quality. Huixian Wang, Wanshou Jiang, Chengqiang Lei, Shanlan Qin, Jiaolong Wang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | Multi-spectral image inter-band registration technology researchabstractMulti-spectral image refers to the use of multi-spectral sensors from the same object (target or region) to obtain the spectral range of more than one image. With the development of multi-spectral imaging technology, the existence of non-rigid mismatch among the original inter-band images requires a higher accuracy and faster registration. Here an automatic elastic registration method based on mutual information and thin-plate spline interpolation is presented, we carried out HJ-1A/B satellite Level 1 visible light inter-band images registration. This does not require other priori assumptions among multi-spectral image gray relationships, or the need for image segmentation and pre-processing (such as feature extraction, classification, etc.). Through the research about multi-spectral image inter-band registration, deviations may be controlled within an acceptable range, we can get more information about multi-spectral images. Chunxiang Cao, Wanshou Jiang, Min Xu 0007, Shilei Lu |
IGARSS | 3 |
| 2012 | Gross primary production estimation by combining MODIS products and Ameriflux data through Artificial Neural Network for croplandsabstractVegetation productivity is the basis of all the biosphere activities on the land surface that relate to global biogeochemical cycles of carbon and nitrogen. The accurate quantification of gross primary production (GPP) in crops is important for regional and global studies of carbon budgets. Many flux observation nets have been established to help us monitoring the carbon cycling. However, estimation of GPP of terrestrial ecosystems for regions, continents, or the globe can improve our understanding of the feedbacks between the terrestrial biosphere and the atmosphere in the context of global change and facilitate climate policymaking. Remote sensing is a potentially powerful technology with which to extrapolate eddy covariance-based GPP to continental scales. In this paper, we combined MODIS products and Ameriflux networks data to simulate and predict GPP at four different cropland sites, using the Artificial Neural Networks (ANN). The results were quite approving compared to MODIS GPP product and tower-based measurements, which indicated it could be an applicable approach for GPP estimation. Zhaocong Wu, Wanshou Jiang |
IGARSS | 3 |
| 2008 | A New Star Identification Algorithm based on Matching ProbabilityabstractA new star identification algorithm based on matching probability is proposed for satellite attitude determination. In this algorithm, the brightest observed star is considered as the primary star, and the radial geometry pattern is constructed by linking the primary star to the other adjacent stars in the FOV. When a link is matched with star database, the two corresponding stars in the star database are recorded. The star with the most appearance times is regarded as the correspondence of the primary star. Experiments show that the computation time and the storage requirement of the algorithm are small, and the identification rate is high, compared with improved triangle matching algorithms. Wanshou Jiang, Jianya Gong |
IGARSS (3) | 2 |