VLDB 2026 Research / reviewers in the wild / expert
Nan Su 0001
dblp:76/3472-1
· DBLP profile ↗
36ranked-venue papers
3as first author
26since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FGOD-YOLOv8: Fine-Grained Object Detection for Crops and WeedsabstractThe task of fine-grained object detection aims to accurately detect small objects in images, which has become increasingly valuable in smart agriculture applications. Existing detection methods tend to have high false detection and missed detection when dealing with small objects such as crops and weeds, resulting in low accuracy. To address this issue, we proposed FGOD-YOLOv8, which consists of a cross-scale feature fusion module, a tiny object detection head, and a multi-scale dilated attention module. The cross-scale feature fusion module ensures effective transmission of high-level semantic information and low-level detail information. The tiny object detection head captures finer-grained features. The multi-scale dilated attention module focuses on object features at different scales to further improve accuracy. According to experimental results, our proposed method demonstrates better accuracy on cropandweed datasets compared to other approaches. Maosheng Wei, Baoyu Ge, Nan Su 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | FCMMA: Fourier Conditional Mask-Based Mixed Attention Method for Hyperspectral Anomaly DetectionabstractIn recent years, reconstruction-based methods have achieved excellent detection results in the field of hyperspectral anomaly detection (HAD). These methods predominantly operate on two aspects regarding their working principles: 1) reconstructing background pixels and 2) suppressing anomalous pixels. However, most methods only tackle the HAD task from the spatial and spectral domains, making it challenging to effectively suppress anomalies. To eliminate these issues, this article proposes a Fourier conditional mask-based mixed attention (FCMMA) method. First, we propose the FCMMA method for HAD. FCMMA generates a conditional mask (CMASK) that suppresses anomalous high-frequency information and preserves background low-frequency information in the frequency domain, optimizing the anomaly detection process. In addition, to achieve fine-grained HAD, we propose the Fourier anomaly suppression filter (FASF). FASF uses Fourier techniques to manage background and anomalies, improving detection via precise frequency decoupling. Finally, a CMASK network is designed to effectively suppress anomalies. The CMASK network integrated the FASF module and the spatial-spectral multilayer perceptual (SSMLP) machine module together to enhance the transformation and representation capabilities of the generated masks, which can also help suppress anomalies. The results on five different datasets show that the proposed method is more effective and superior when compared to nine state-of-the-art methods. Shou Feng, Nan Su 0001, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Multi-Modality Feature Enhancement Method Based On Feature Disentanglement For Sar Image Target DetectionabstractSynthetic Aperture Radar (SAR) ship detection algorithms have achieved extensive development in recent years. In spite of this, the insufficient data and the non-intuitive feature of SAR images still brought certain challenges. This paper proposes a multi-modality feature enhancement (MMFE) method based on feature disentanglement for SAR image target detection. By precisely exploring modality-shared features of optical and SAR images, MMFE can optimize the SAR feature representation capability. First, we propose a feature disentanglement (FD) module to acquire transferable modality-shared knowledge, thereby effectively alleviating the modality shift phenomenon in the subsequent modality alignment. Second, we introduce a multi-granularity modality alignment (MGMA) module that further eliminates inter-modality differences, ultimately achieving effective compensation for the SAR modality. Extensive experimental results convincingly demonstrate the compelling ability of MMFE. Jiayue He, Nan Su 0001, Yanping Liao, Shou Feng, Chunhui Zhao 0003 |
ICIP | 2 |
| 2024 | Auxiliary Edge Traction Network for Vehicle Re-Identification Based on Multimodal Aerial ImagesabstractDue to their outstanding flexibility and safety, unmanned aerial vehicles (UAVs) have received widespread attention in the fields of video surveillance and target tracking. However, conventional RGB cameras may not meet the needs of dark environments and harsh scenes. The Synthetic Aperture Radar (SAR) camera, being independent of optical imaging, enables observations under a variety of weather conditions and at extended ranges. Currently, there is no complete cross-modal vehicle re-identification dataset in the field of remote sensing. We have collected and constructed a dataset named Unmanned Aerial Vehicle Multimodal Vehicle Re-identification (UAV-Mu), which includes 11 identities, 2598 RGB samples, 2621 infrared samples, and 770 SAR samples. In addition, to address the significant differences between modalities, we propose an Auxiliary Edge Traction Network (AET-Net). AET-Nat can effectively aggregate different layers of the network, thereby enhancing attention to vehicle outline details to output more discriminative features. Comprehensive experimental results demonstrate that the proposed AET-Net outperforms several other advanced methods. Fengjiao Gao, Nan Su 0001, Chunhun Zhao |
IGARSS | 5 |
| 2024 | Multi-Modal Target Detection Method Based on Adaptive Feature SearchabstractThe optical remote sensing image has a high resolution, while the infrared image provides temperature information about the detected object. These two types of information are complementary. However, optical images often suffer from spatial misalignment issues, which make feature fusion operations challenging. To address these problems, we propose a Transformer feature fusion module that captures high-quality fusion feature information. Building upon this, we design a novel two-branch backbone network that utilizes infrared image features to adaptively screen optical image features, thereby enhancing the detection performance. Experimental results demonstrate the superiority of our approach over the baseline on multi-modal data with non-alignment problems. Nan Su 0001, Minghui Sha, Chunhui Zhao 0003, Shou Feng, Yingshen Zhu |
IGARSS | 2 |
| 2024 | PNBT-CR: A Cloud Removal Method for Ship DetectionabstractIn the ship detection of the remote sensing images, cloud occlusion could blur the boundaries between ships and backgrounds, making it more challenging to distinguish them. Cloud occlusion can also result in partial or complete occlusion of target, making it difficult for models to detect ships in their entirety. Therefore, the use of cloud removal techniques is essential to enhance the accuracy and robustness of target detection. However, existing cloud removal processes are applied to entire images, providing limited improvements for specific object detection. In this letter, A Perlin Noise Based Thin Cloud Removal (PNBT-CR) Network is proposed for ship detection. The proposed algorithm introduces a Perlin noise mist mix module, which can improve the cloud removal effect of the network effectively. And it designed a Target-Oriented Structural Similarity (TOSS) loss function that enhances the network’s ability to boost the confidence of detected ships in the results. Experimental results demonstrate the efficacy of this approach in restoring texture details in ships and enhancing the accuracy of ship detection in remote sensing imagery. Moreover, the images processed using our method can have 89.86% SSIM compared to the original images, and when used for ship detection, there can be a maximum improvement of 11.7% F1-score. Yanming He, Nan Su 0001, Guangjun He |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Fractional Fourier-Based Frequency-Spatial-Spectral Prototype Network for Agricultural Hyperspectral Image Open-Set ClassificationabstractAt present, hyperspectral image classification (HSIC) technology has been warmly concerned in all walks of life, especially in agriculture. However, existing classification methods operate under the closed-set assumption, which deviates from the real world with open properties. At the same time, there are more serious phenomena of different crops with similar spectrum and same crops with different spectrum in agricultural hyperspectral data, which is also a great challenge to existing methods. In this work, a fractional Fourier based frequency-spatial-spectral prototype network is proposed to address the challenges of open-set hyperspectral image classification in agricultural scenarios. Firstly, fractional Fourier transform is introduced into the network to combine the information in the frequency domain with the spatial-spectral information, so as to expand the difference between different classes on the premise of ensuring the similarity between classes. Then, the prototype learning strategy is introduced into the network to improve the feature recognition capability of the network through prototype loss. Finally, in order to break the stubbornly closed-set property of closed-set classification method, the open-set recognition module is proposed. The difference between the prototype vector and the feature vector is used to judge the unknown class. Experiments on three agricultural hyperspectral datasets show that this method can effectively identify unknown class without sacrificing the classification accuracy of closed-set, and has satisfactory classification performance. Maoyang Chen, Shou Feng, Chunhui Zhao 0003, Bo Qu, Nan Su 0001, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | High-Resolution Remote Sensing Image Change Detection Based on Fourier Feature Interaction and Multiscale PerceptionabstractAs a significant means of Earth observation, change detection in high-resolution remote sensing images has received extensive attention. Nevertheless, the variability in imaging conditions introduces style discrepancies and a range of pseudochange regions between bitemporal image pairs. Furthermore, changing objects possess diverse morphological representations, which makes accurately identifying change areas and delineating their boundaries within complex object distributions increasingly difficult. In response to the aforementioned challenges, we propose the Fourier feature interaction and multiscale perception (FIMP) model for effective change detection. To mitigate the impact of style discrepancies, FIMP employs the Fourier transform to adaptively filter bitemporal features in the frequency domain while mining the optimized bitemporal features relevant to the change detection task. To enhance the ability to recognize multiscale changing objects, FIMP aggregates and emphasizes the change areas with the introduced temporal change enhancement module (TCEM). By utilizing the U-fusion change perception module (UCPM) to perform multilevel bidirectional fusion of change features at different scales, FIMP can further enhance the ability to delineate complex semantic change boundaries. Experiments on three public datasets show that our approach outperforms seven state-of-the-art methods. Shou Feng, Chunhui Zhao 0003, Nan Su 0001, Wei Li 0032, Ran Tao 0003, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Foreground-Driven Fusion Network for Gully Erosion Extraction Utilizing UAV Orthoimages and Digital Surface ModelsabstractUnmanned aerial vehicle (UAV) orthoimages and digital surface models (DSMs) can provide valuable insights for semantic segmentation methods in comprehending gully erosion (GE) from diverse perspectives. While the integration of these two modalities has the potential to improve the GE extraction performance, the extent of enhancement primarily depends on the quality of modality-specific features and the synergistic fusion manner employed for integrating features from both modalities. Toward this end, we propose a novel multimodal segmentation method, which is called foreground-driven fusion network (FFNet). Guided by the prototypes of foreground objects (i.e., gullies), the network effectively tackles the challenges from the modality itself and between different modalities, ultimately achieving high-quality GE extraction results. Specifically, a foreground prototype sampling (FPS) module is first devised for precisely sampling foreground prototypes related to gullies from two modalities. Then, a local-global hybrid purification (LHP) module is proposed to effectively mitigate the erroneous activation within each modality at multiple dimensions by leveraging foreground prototypes. Finally, a multimodal foreground synergy (MFS) module is introduced to further activate foreground features and facilitate full complementarity between multimodal foreground features. To validate our network, a comprehensive multimodal dataset for GE extraction is constructed based on UAV orthoimages and DSMs from northeastern China. Furthermore, a public road extraction dataset is employed to evaluate the generalizability of this network. In the experiments conducted on these two datasets, the proposed FFNet exhibits obvious superiority, outperforming the second-best method with an average improvement of 2.55% in terms of intersection over union (IoU) and 2.77% in terms of$F1$-score. These experimental results not only demonstrate the practicality of FFNet in GE extraction tasks, but also highlight its significant advantage in similar road extraction tasks. Yi Shen 0013, Nan Su 0001, Chunhui Zhao 0003, Shou Feng, Wei Xiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Building Extraction at Amodal-Instance- Segmentation Level: Datasets and FrameworkabstractThis article presents two amodal instance segmentation (AIS) datasets in the field of remote sensing: the multiview building dataset for the Zurich region (MVB-Zurich) and the multisize building dataset for the Dortmund region (MSB-Dortmund). Additionally, a new AIS network framework, named FE-RSBL-AmodalNet, is proposed, which is based on feature enhancement and remote sensing boundary loss. Instance segmentation has emerged as a popular approach for building extraction in recent years. However, a limitation of using such algorithms for building extraction is the inability to predict the invisible areas. Consequently, the extracted building contours are incomplete, which hampers certain applications relying on accurate building extraction. To address this limitation, AIS has emerged as a promising research field. Unfortunately, there is currently a lack of datasets available for developing AIS algorithms in the field of remote sensing, which poses a barrier to the widespread application of AIS in this domain. This article introduces two AIS datasets specifically designed for remote-sensing buildings. The datasets consist of images captured from tilted views using a tilt photography system, resulting in a significant presence of occluded areas within the images. Moreover, the MVB-Zurich dataset comprises aerial images captured from five different viewpoints, while the MSB-Dortmund dataset encompasses diverse buildings, including garages and residences, with varying sizes. The multiview and multisize attributes of these datasets offer enhanced research opportunities. Furthermore, a new AIS framework, specifically designed for remote sensing buildings, was proposed with the aim of accurately predicting complete building contours. Cong'an Xu, Nan Su 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | FVMD-ISRe: 3-D Reconstruction From Few-View Multidate Satellite Images Based on the Implicit Surface Representation of Neural Radiance FieldsabstractThree-dimensional reconstruction utilizing few-view satellite images can effectively reduce the cost of resources. However, methods that rely on dense stereo matching suffer severe performance degradation when confronted with differences of intersection angles and illumination in such non-standard stereo images, which are limited by the matching mechanism. Neural radiance fields (NeRF) gets rid of the matching restriction by utilizing the rendering pipeline, which has made significant progress in the synthesizing novel views and 3D reconstruction. Nevertheless, when the available images are few-view and multi-date, insufficient geometric constraints and illumination variations present great challenges in generating accurate 3D models. In this paper, we proposed a 3D reconstruction framework called FVMD-ISRe, which is based on the implicit surface representation of NeRF, for generating complete watertight mesh and digital surface model (DSM) from few-view multi-date satellite images. We introduce the positional encoding with adaptive frequency module to extract finer geometric information from few-view images. Then, the solar information is incorporated into the rendering process, allowing the network to distinguish color changes arising from illumination differences in multi-date images. Additionally, we employ some tricks like coordinate system optimization and network restructuring for enhancing network training. Advantages of FVMD-ISRe are demonstrated through both qualitative and quantitative experiments conducted on the US3D dataset. The results highlight our framework’s capacity to reconstruct accurate 3D geometry, overcome the performance of traditional methods based on stereo matching and other NeRFs in various evaluation metrics. Code and data are available at https://github.com/HEU-super-generalized-remote-sensing/FVMD-ISRe. Chi Zhang 0047, Chunhui Zhao 0003, Nan Su 0001, Weikun Zhou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | An Attention-Guided Matching Association Network for Hyperspectral and RGB Fusion TrackingabstractRGB-based trackers are prone to drift in some challenging scenarios. Hyperspectral data can provide more material information to address these challenges. Therefore, using hyperspectral information to supplement the RGB modality defects to improve tracking performance is worth exploring. However, there is almost no relevant work about this valuable issue. In addition, the two modality data in the existing hyperspectral-RGB dataset are not strictly matched and aligned, which brings significant challenges to the multi-modality tracking task. Therefore, we propose a simple but effective scheme to alleviate the problem of the difficulty of utilizing multi-modality information in tracking caused by the spatial difference between two modalities. In addition, to promote the development of the hyperspectral-RGB multi-modality tracking field, we propose a novel network that adaptively captures the relationship between the two modality information using the attention mechanism in Transformer to improve the tracking performance. Experimental results demonstrate the proposed method’s effectiveness. Hongjiao Liu, Nan Su 0001, Chunhui Zhao 0003 |
IGARSS | 2 |
| 2023 | A Multispectral-Infrared Object Detection Method Based on Cross-Modality Image Feature Filtering FusionabstractIn this paper, a multi-spectral infrared target detection method is established based on image feature filtering fusion. Multispectral(MS) images and infrared(INF) images are used for fusion. Considering the advantages of the multi-spectral image and infrared image, use the feature filtering method to extract the features of the two kinds of images, and then carry out the fusion and subsequent detection. We use FLIR datasets to evaluate the proposed method, and the accuracy has been improved. The comparison results show that the proposed multispectral-infrared object detection method based on cross-modality image feature filtering fusion over the conventional detection methods. Nan Su 0001, Chunhui Zhao 0003 |
IGARSS | 2 |
| 2023 | A Method Based on Multi-Scale Consistency Regularization and Color-Spatial Constraints for Road Segmentation with Noisy LabelsabstractRoad segmentation methods based on (Convolutional Neural Networks, CNN) commonly require accurate pixel-level labels. However, accurate labels are hard to obtain in some cases. To address this issue, a method based on multi-scale consistency regularization and color-spatial constraints is proposed for road segmentation with noisy labels. First, a multi-scale consistency regularization is devised to improve the multi-scale aggregation capability and noise immunity of model. In addition, we utilize the color and spatial information of input images to constrain the model predictions, thereby giving the model a more reliable learning objective. In the experimental section, the accuracy and visualization results obtained from two road datasets provide compelling evidence of the efficacy of our method in addressing the challenges posed by noisy labels in road segmentation tasks. Yi Shen 0013, Nan Su 0001, Chunhui Zhao 0003 |
IGARSS | 2 |
| 2023 | Scene-Level Matching Between Remote Sensing Optical and SAR ImagesabstractDue to the relative complementarity between optical images and Synthetic Aperture Radar (SAR) images, the method of SAR-optical matching is widely used in auxiliary navigation, disaster monitoring, rescue and other fields. However, there are huge geometric and radiometric differences between SAR and optical images, which pose serious challenges for multi-modal image matching. To solve this problem, this paper proposes a Refined Subdivision Processing Network (RSP-Net) for SAR-optical matching. Firstly, to extract representative features of images, we propose to employ the pseudo-Siamese network structure with dual-branch partial weight sharing in RSPNet. Then, to preserve the detailed information in the image, the method of subdividing the features to generate part-level features of the image is proposed. Finally, to remove modality-specific but task-independent information, part-level features are refined using information bottleneck methods. Experiments show that our proposed method has an excellent performance in the Scene-level matching between optical and SAR images. Haiyang Zhong, Guihua Gu, Nan Su 0001, Hongzhe Zhang |
IGARSS | 4 |
| 2023 | Multi-Scale Geo-Localization Based on Local Similarity Area Distance Measurement MethodabstractCross-view geo-localization is to match remote sensing images from different platforms. The UAV-view image is matched with the satellite-view image with geographical localization information, so as to determine the specific geographical localization of the UAV. Cross-view geo-localization can be applied in many fields. For example, cross-view geo-localization can assist traditional satellite positioning to improve accuracy or perform positioning tasks independently. The main challenge now is that there are great differences in remote sensing images obtained by satellites and UAVs, such as changes in viewpoints and differences in scales. The current method mainly studies the influence of viewpoint change, while ignoring the scale difference between cross-source images caused by different resolutions. In order to reduce the influence of scale differences, We propose a network structure called local similarity network (LSNet) based on the siamese network and multi-scale sliding windows. LSNet adopts a new distance measurement method based on the most similar area. Experiments show that our method has excellent performance in multi-scale remote sensing geo-localization. Jianzheng Zhou, Guihua Gu, Nan Su 0001 |
IGARSS | 4 |
| 2023 | GEOP-Net: Shape Reconstruction of Buildings From LiDAR Point CloudsabstractThe shape reconstruction of buildings based on LiDAR point clouds is extremely significant in remote sensing. In recent years, reconstruction methods based on the implicit network have been widely used in object-level shape reconstruction. However, the incompleteness and sparsity of airborne LiDAR scanning point clouds will lead to poor reconstruction results. To solve this problem, GEOP-Net: an implicit modeling framework embedded with high-dimensional geometric features, is proposed in this letter. Firstly, the geometric encoding module added to extract high-dimensional features enhances the feature extraction ability of the network to the detailed structures. The point clouds of buildings in Zurich are collected and used to evaluate the performance of the proposed method. The experimental results show that the proposed method have better accuracy than the existing methods, so it provides a new research idea for building reconstruction. Cong'an Xu, Nan Su 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | DSTNet: Dynamic-Static Transformer Style Network for Cross-Resolution Vehicle ReidentificationabstractVehicle ReIdentification (ReID) can be applied to multi-temporal remote sensing target-matching tasks in different locations. However, due to the uncertainty of UAV height and maneuvering target motion, a huge resolution mismatch can be expected. In the traditional cross-resolution ReID method, the Super-Resolution (SR) method is generally used. However, there is still a large data difference between the Super-Resolution Recovered (SR-Recovered) image and the High-Resolution (HR) image, which leads to a decrease in matching efficiency. Therefore, a dynamic-static TransFormer style network is proposed, which is named DSTNet. DSTNet is designed to reduce the difference between the SR-Recovered image and the HR image and to obtain the identity invariant representation of the SR-Recovered image and the HR image. Firstly, CNN and TransFormer are used to extract context information statically and dynamically, respectively, to enhance the representation of the target identity. Secondly, in order to obtain the invariant information between the SR-Recovered image and the HR image, different normalization strategies are designed in different depths of the DSTNet. Finally, to obtain a consistent representation of the SR-Recovered image and the HR image, the High-Resolution Constraint (HRC) input method is applied to the network. To the experimental results, the performance of rank-5 and mAP is improved by 3% and 3.6% respectively on datasets with large resolution differences by our method. Chunhui Zhao 0003, Nan Su 0001, Shou Feng |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | A Cross-Modality Feature Transfer Method for Target Detection in SAR ImagesabstractSynthetic aperture radar (SAR) ship detection methods have achieved remarkable progress in recent years. However, unlike RGB images, the characteristics of SAR imaging will result in non-intuitive feature representations. Furthermore, due to the insufficient data of SAR images, existing methods relying on plenty of labeled SAR images may be hard to achieve promising performance. To address the aforementioned issues, a cross-modality feature transfer (CMFT) method is proposed in this article, which enhances feature representations in the SAR modality by transferring rich knowledge in the RGB modality. First, we propose a multilevel modality alignment network (MMAN), which encourages the model to effectively learn modality-invariant features and alleviate the large cross-modality discrepancies by aligning features from multilevels (scene level, local level, global level, and instance level). Second, to address the underperformance of samples with non-intuitive features in the modality alignment, we introduce a hard-sample supervision module (HSM) in the stage of feature extraction, which can thoroughly exploit the feature of hard-to-align samples by giving more optimization energy for them. Third, to enhance the discriminability of instance-level features, a feature complementary module (FCM) is customized to fully explore the potential complementary clues between instance-level features and context information for the instance-level feature alignment. Extensive experimental results demonstrate that the CMFT outperforms the state-of-the-art detectors. Compared to the baseline model, CMFT improves the accuracy by 3.1% mean average precision (mAP) on the SSDD dataset and 3.4% mAP on the HRSID dataset, demonstrating its superior SAR ship detection performance. Jiayue He, Nan Su 0001, Cong'an Xu, Yanping Liao, Chunhui Zhao 0003, Shou Feng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Detect Larger at Once: Large-Area Remote-Sensing Image Arbitrary-Oriented Ship DetectionabstractShip detection is one of the main problems of satellite image analysis. Since ships are scattered on the sea and major ports, large-area remote-sensing images need to be processed in order to realize the detection of ships. In addition, since the satellite is a top-down view, the ship with aspect ratios cannot be covered in complex backgrounds by a horizontal bounding box very well and need a rotating bounding box to achieve this task. Although considerable progress has been made in object detection techniques, there are still challenges for fast detection of ships in large-area remote-sensing images. In this letter, an arbitrary-oriented detector for large-area remote-sensing images is proposed to quickly locate ship positions. A new feature extraction network DCNDarknet25 based on you only look once (YOLO) is designed by reducing paraments and adding deformable convolution (DCN) to improve the speed and accuracy. And the rotation detection capability without angle regression is added to the YOLO detection algorithm for the first time. Finally, thanks to the advantages of our fully convolutional lightweight network, a method for detecting large-area remote-sensing images at once is proposed. In the public dataset HRSC2016 and our own large-area remote-sensing (LARS) image dataset, it has achieved very good accuracy and several times the speed of other algorithms. Nan Su 0001, Zhibo Huang, Chunhui Zhao 0003, Shuyuan Zhou |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Cross-Dimensional Object-Level Matching Method for Buildings in Airborne Optical Image and LiDAR Point CloudabstractIn this letter, an object-level matching method was proposed to perform building matching in cross-dimensional remote-sensing data. Object-level matching of buildings is essential pre-work for 3-D shape reconstruction, to further generate productions, such as a digital building model. Optical images and LiDAR point clouds, with rich colors and precise 3-D positioning information, respectively, are mostly used alone for 3-D shape reconstruction. Obviously, it is better to combine both the 2-D image and 3-D point cloud. However, it is difficult to first perform the cross-dimensional object-level matching (COlM), since there are few descriptors which can unify the 2-D image and 3-D point cloud. To address this issue, a feature transformation framework was proposed. First, a cross-dimensional encoder module is introduced to unify the descriptor extracted from the optical image and LiDAR point cloud. Second, a spatial occupancy probability descriptor (SOPD) is employed to associate the descriptor extracted from different dimensional data with instinct geometric structure of buildings. Then, the 3-D geometric structure is transformed into a feature vector for matching. For experiments, a cross-dimensional object-level building dataset was collected and labeled to verify our method. It includes cross-view 2-D optical images and LiDAR point clouds for each building from hundreds of buildings. The results show that a high-accuracy COlM was achieved. Nan Su 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Shape Reconstruction of Object-Level Building From Single Image Based on Implicit Representation NetworkabstractThree-dimensional shape reconstruction of the object-level building (SROLB) is one of the essential issues in remote sensing. Especially, utilizing the single remote sensing image (SRSI) to perform shape reconstruction can offer better scalability and transferability, in terms of simplifying input data. Recently, the methods of shape reconstruction based on neural networks have been widely studied. However, most of them generate models with irregular surfaces and few details. Besides, complex background in SRSI leads to a poor generalization of networks and reduces the quality of generated models. To solve the above problems, an implicit representation network (IRNet) is proposed in this letter. IRNet is composed of two parts: 3-D space decoding and feature extraction. First, the signed distance function (SDF) is employed to fit implicit representation better in the decoding module. Moreover, a multistage weight loss function is designed, making the network generating models with flatter surfaces and more details. Then, a channel attention (CA) module is added to the feature extraction network. It reduces the interference of the background in the image effectively and improves the generalization of the network. Finally, our method generates mesh models of the individual buildings. The experimental results show that a better accuracy can be obtained compared with state-of-the-art methods. Chunhui Zhao 0003, Chi Zhang 0047, Nan Su 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | TFTN: A Transformer-Based Fusion Tracking Framework of Hyperspectral and RGBabstractAlthough the RGB image has a high spatial resolution, it only depicts color intensities in red, green, and blue channels, which easily leads to the failure of the tracker based on RGB modality in some challenging scenarios, for example, when the color of the object and background is similar. The hyperspectral image with rich spectral information is more robust in these difficult situations, so it is essential to explore how to effectively apply hyperspectral features to supplement RGB information in object tracking. However, there is no fusion tracking algorithm based on hyperspectral and RGB data. Based on this, we propose a novel fusion tracking framework of hyperspectral and RGB in this article, termed as Transformer-based Fusion Tracking Network (TFTN), to enhance the performance of object tracking. Within the framework, we construct a dual-branch structure based on the Siamese Network to obtain the modality-specific representations of different modality images. Besides, the framework is generic, which is suitable for the Siamese series of tracking algorithms. In addition, we design a Siamese three-dimensional convolutional neural network as the specific branch of hyperspectral modality for synchronous extraction of the spatial and spectral features of hyperspectral data, to give full play to the role of hyperspectral data in improving network tracking performance. Particularly, inspired by the structure of Transformer, we design a Transformer-based fusion module to capture the potential interaction of intra-modality and inter-modality features of different modalities. This is the first work that combines the information of hyperspectral and RGB modalities to improve tracking performance. At the same time, it is also the first time that employs the self-attention module of Transformer to combine the information of different modalities for multi-modality fusion tracking. Experimental results on the dataset composed of hyperspectral and RGB image sequences show that the proposed TFTN tracker is superior to the state-of-the-art trackers, demonstrating the effectiveness of this method. Chunhui Zhao 0003, Hongjiao Liu, Nan Su 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Hyperspectral Target Detection Method Based on Nonlocal Self-Similarity and Rank-1 TensorabstractIn recent years, many target detection methods based on tensor representation theory have been proposed and achieved good results for hyperspectral images (HSIs). However, these methods still have some deficiencies. For example, 3-D hyperspectral data are first transformed into 1-D vectors in these methods, which may destroy the spatial structure of HSI data and reduce the detection performance. Besides, when the number of training samples is small, the results of the target detection method usually become worse. To solve these problems, a hyperspectral target detection method based on nonlocal self-similarity and rank-1 tensor is proposed in this article. First, different from these traditional tensor representation-based methods, the third-order tensor data are directly used as the input of the proposed method to preserve the spatial information and structure of an HSI. Second, the tensor blocks related to the class are constructed by using the nonlocal self-similarity of HSI data. Finally, by taking advantage of rank-1 canonical decomposition attribute, the process of tensor operation can be simplified, and the number of training samples can be reduced. The proposed method is compared with six state-of-the-art hyperspectral target detection methods on four HSI data sets. The experimental results show that the proposed method can have better target detection results than other compared methods, especially in the case of fewer training samples. Chunhui Zhao 0003, Shou Feng, Nan Su 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Hyperspectral Anomaly Detection Using Bilateral-Filtered Generative Adversarial NetworksabstractWithout any prior information of anomalies or background, hyperspectral anomaly detection has received a wide attention. However, such unsupervised style brings difficulties in training and learning effective features of hyperspectral image to perform detection. This paper proposes a novel hyperspectral anomaly detection algorithm using bilateral-filtered generative adversarial networks (BFGAN). Bilateral filter can smooth images and remove anomalous points while preserving edges. With closeness weights and similarity weights, the bilateral-filtered hyperspectral image can be considered as background data, so that hyperspectral background labels are obtained. Only with one class of labels, the structure of generative adversarial networks has an ability to solve two-class problem. By using the filtered background data and their labels, generative adversarial networks are trained to improve discriminator's discriminative capability for background data in a competing style. Finally, the model discriminator can finally output big probabilities for background samples and small probabilities for anomalous samples. Experiments on two real hyperspectral images demonstrate that the proposed method outperforms other state-of-the-art competitors. Chunhui Zhao 0003, Chuang Li 0005, Shou Feng, Nan Su 0001 |
IGARSS | 4 |
| 2021 | A Complete Building Extraction Framework for Airborne Laser Scanning Point CloudabstractIn this paper we proposed a complete building extraction framework (CBEF) for airborne laser scanning point clouds. By using 3D instance segmentation to extract rough buildings, we proposed a post-processing method to optimize the extract results, and used a point cloud completion network to repair the incomplete building instances. Experimental results show that this proposed framework can better extract building instances from airborne laser scanning point cloud, and can repair incomplete building point clouds those lost facades. Chunhui Zhao 0003, Hemin Lin, Nan Su 0001, Shu Tian |
IGARSS | 4 |
| 2020 | A Distribution Controllable Simulation Method of Remote Sensing Sea-Ice ImagesabstractIn the case of sailing out in the sea-ice areas, it is instructive for route planning to research the distribution characters of ice in the target sea. Existing deep learning methods have shown their strength on sea-ice images processing like image classification. Due to the complex environment around sea-ice area, capturing large quantities of images is not easy. Besides, it's often hard to guarantee the abundance of sea-ice distribution of each different scene class, which causes unsatisfactory classification results. Therefore, it is of considerable practical value to research on sea-ice images simulation. In this paper, a distribution controllable simulation method is proposed based on generative adversarial networks for remote sensing sea-ice images. This research can help settle the problem of small sea-ice samples, as well as can provide a practical method for optical image simulation and similar type problems. Chunhui Zhao 0003, Nan Su 0001 |
IGARSS | 4 |
| 2020 | Spectral-Spatial Stacked Autoencoders Based on the Bilateral Filter for Hyperspectral Anomaly DetectionabstractTaking advantaging of the ability to extract high-level features, the algorithms based on deep learning for hyperspectral imagery (HSI) anomaly detection have drawn great attention in recent years. In this paper, we propose a method named spectral-spatial stacked autoencoders based on the bilateral filter (SSSAE-BF). First, the bilateral filter is employed to obtain the derived anomaly components and background components. Second, stacked autoencoders (SAE) are respectively utilized on the derived anomaly component and background component for deep features. Finally, the Reed and Xiaoli detector (RXD) is used on the spectral-spatial features to calculate the detection result. Experiments on two real hyperspectral images demonstrate that the proposed method outperforms the other competitors. Chunhui Zhao 0003, Chuang Li 0005, Shou Feng, Nan Su 0001 |
IGARSS | 4 |
| 2020 | Dictionary Learning Hyperspectral Target Detection Algorithm Based on Tucker Tensor Decompositionabstract1As a research hotspot, hyperspectral image target detection is more and more widely used in military and civilian fields. In order to make use of the spatial and spectrum information of hyperspectral image data at the same time, a new dictionary learning hyperspectral image target detection algorithm based on Tucker tensor decomposition is proposed in this paper. The algorithm uses Tucker tensor decomposition to extract effective local image block spatial spectrum features. A detection model based on sparse representation and collaborative representation is established, and experiments are carried out on two representative hyperspectral images data. From the visual detection results, the algorithm effectively extracts the spatial spectrum features in the complex background and strong noise environment, has a good ability to suppress the background, and the detection target is significant. Chunhui Zhao 0003, Nan Su 0001, Shou Feng |
IGARSS | 3 |
| 2020 | A Novel Building Reconstruction Framework using Single-View Remote Sensing Images Based on Convolutional Neural NetworksabstractBuilding model reconstruction is one of the remaining challenges for satellite imagery. In this article, we present a framework that leverages single-view remote sensing images to reconstruct buildings in 3D space. The framework consists of two main steps: the powerful feature learning ability in convolutional neural networks is utilized to reconstruct roof-based 3D voxel grid buildings. Another is improving the resolution and accuracy of voxel models through interpolation optimization methods with prior information. By evaluating the experimental results, our method can accurately recover the polygonal buildings structure with several different roof types. The model optimization results have significantly improved in subjective visual effects and objective evaluation. Chunhui Zhao 0003, Chi Zhang 0047, Nan Su 0001 |
IGARSS | 3 |
| 2018 | Sea-Ice Image Classification for Channel Navigation in Polar ApplicationabstractIn the paper, we proposed a novel framework for the sea-ice images classification, which will be very helpful for the future channel navigation in polar applications. Firstly, the three different types of sea ice are defined in the paper based on the large amounts of data inspection, such as thick ice, thin ice and water. Further, we present to use relative radiometric normalization method to solve the grayscale difference of different temporal sea-ice images, which can affect the classification accuracy of sea ice. Finally, the SAE (sparse auto-encoder) classifier is employed for different types of sea-ice classification. We designed several experiments to discuss and analyze the influence factors on the classification accuracy of sea-ice images. The experiment results indicated the validity of the proposed method and the rationality of sea ice type definition. Nan Su 0001, Chunhui Zhao 0003, Zhichao Tan |
IGARSS | 1 |
| 2018 | Sea-Ice Scene Classification Using Aerial Images in Arctic Based on Transfer LearningabstractArctic travel has become a significant way for commercial or scientific researches. Sea-ice is always serious threat for Arctic navigation. In this paper, a classification method of aerial images is proposed to recognize different types of sea-ice scenes, which can be further used to analyze the threat of each scene in Arctic Ocean. However, due to the diverse types and distribution of sea-ice in various scenes, the categories of sea-ice scenes are indistinct. So a grouping way is primarily introduced to define different sea-ice scenes. These scenes are grouped into six typical categories according to the distribution of sea-ice. Then a transfer learning strategy is used to fine-tune a pre-trained deep convolution neural network model. Finally, using that trained network to classify new sea ice scenes. Experimental results show that an acceptable classification accuracy can be obtained. Zhichao Tan, Nan Su 0001 |
IGARSS | 3 |
| 2017 | A Novel Deep Embedding Network for Building Shape RecognitionabstractBuilding shape, as a key structured element, plays a significant role in various urban remote sensing applications. However, because of high complexity and intraclass variations between building structures, the capability of building shape description and recognition becomes limited or even impoverished. In this letter, a novel deep embedding network is proposed for building shape recognition, which combines the strength of the unsupervised feature learning of convolutional neural networks (CNNs) and a novel triplet loss. Specifically, we take advantage of the strong discriminative power of CNNs to learn an efficient building shape representation for shape recognition. With this deep embedding network, the high-dimensional image space can be mapped into a low-dimensional feature space, and the deep features can effectively reduce the intraclass variations while increasing the interclass variation between different building shape images. Afterward, the derived deep features are exploited for the process of building shape recognition. This method consists of two stages. In the first stage, for standard building shape image queries stored in the shape primitives library and the building shape data set, two sets of deep features are extracted with the deep embedding network. In the second stage, we formulate the shape recognition task into a feature matching problem and the final building shape recognition results can be achieved by set-to-set feature matching method. Experiments on the VHR-10 and UCML data sets demonstrate the effectiveness and precision of the proposed method. Shu Tian, Ye Zhang 0008, Junping Zhang, Nan Su 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | RPC estimation via feature points for urban areasabstractRational function model (RFM), which is composed of 80 rational polynomial coefficients (RPCs), has been widely used to establish a functional relationship between the image space and the object space in photogrammetry and remote sensing. In order to gain precise RPCs, a set of ground control points (GCPs) need to be selected. And in all most applications, many GCPs selected are located in the center of one object in order to increase the measurement precision. Note that, for a special area, such as urban areas which consist of many buildings, have many feature points which are usually located in the corner. This paper proposes an improved method using feature points instead of non-feature points to solve RPCs. Evaluation of aerial data shows that the RPCs can reach a high fitting accuracy by means of these feature points as GCPs, more than those using non-feature points. An important advantage is that this method can enhance edge protection of buildings, so that it can applied in subsequent applications, such as shadow removing, target recognition and so on. Moreover, this method can remain computation accuracy of buildings even if the number of GCPs decreases. Ye Zhang 0008, Nan Su 0001 |
IGARSS | 3 |
| 2015 | A roof-contour guided multi-side interpolation method for building texture-mapping using remote sensing resourceabstractIn this paper, we proposed a novel texture-mapping method for buildings using remote sensing images and digital surface model. For generating better 3D map, boring manually or semi-automatic texture-mapping of buildings is always needed. However, only with remote sensing images and digital surface model, it is difficult to generate `sides' of buildings, and corresponding relation between triangular mesh and texture image is also hard to be found. Inspired by building extraction result, we found that roof-contour could guide interpolation of `sides' of building, and generate triangular mesh, and an interactive frame of texture mapping is introduced for processing the generated multi-sides and corresponding texture-images of buildings. Experiments show that texture-mapping of buildings using remote sensing resource is easily realized with our frame, and excellent results could be obtained. Xi Chen 0004, Fengjiao Gao, Ye Zhang 0008, Yi Shen 0001, Nan Su 0001, Shu Tian |
IGARSS | 6 |
| 2013 | A novel model for building information acquisition optimization technology of remote sensing observationabstractIt is an important problem in remote sensing that using limited observing points acquire the maximum quantity of building information. In this paper, a building information acquisition (BIA) model based on Support Vector Machine (SVM) is proposed for quantitative description of the mathematical relationship between the information quantity acquisition and the observing angles, which is optimized to obtain the maximum information quantity in the multi-temporal remote sensing observation. The main idea of the BIA model is that, to calculate information quantity at different observing angles, the target is decomposed into multiple faces whose information is described by the combined vector. Further, the modified bee colony algorithm is utilized to optimize the model to achieve the ideal maximum information quantity. The corresponding combined vector is optimal observing angles combination. The proposed model method performs well in our imaging simulation system data. Experiment results demonstrate that the proposed BIA model optimized will provide much more information quantity than observing randomly. Nan Su 0001, Ye Zhang 0008, Yanfeng Gu |
IGARSS | 1 |