EDBT 2026 Demo / reviewers in the wild / expert
Xin Su 0003
dblp:54/3643-3
· DBLP profile ↗
36ranked-venue papers
5as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 2 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCTNet: A Shallow CNN-Transformer Network With Statistics-Driven Modules for Cloud DetectionabstractExisting cloud detection methods often rely on deep neural networks, leading to excessive computational overhead. To address this, we propose a shallow convolutional neural network (CNN)-Transformer hybrid architecture that limits the maximum downsampling rate to 8×. This design preserves local details while effectively capturing global context through a lightweight Transformer branch. To enhance adaptability across diverse cloud scenes, we introduce two novel statistics-driven modules: statistics-adaptive convolution (SAC) and statistical mixing augmentation (SMA). SAC dynamically generates convolutional kernels based on input feature statistics, enabling adaptive feature extraction for varying cloud patterns. SMA improves model generalization by interpolating channel-wise statistics across training samples, increasing feature diversity. Experiments on four datasets show that the proposed method achieves state-of-the-art performance with 732K parameters and 1G multiply-accumulate operations. Code will be available at https://weix-liu.github.io/ for further research. Bin Luo 0005, Jun Liu 0072, Xin Su 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | LOSA: Learnable Online Style Adaptation for Test-Time Domain Adaptive Object DetectionabstractDomain adaptive object detection methods for remote sensing images typically rely on large-scale target domain data and multi-epoch offline adaption training. However, the wide variation in remote sensing conditions makes it difficult to gather sufficient data for every potential target domain, especially for unexpected domains. To address this challenge, we propose Learnable Online Style Adaptation (LOSA), a method that enables the source model to adapt effectively to new target domain styles with low test-time latency. Specifically, LOSA captures target domain styles using shallow feature channel statistics and predicts style shifts based on channel dependencies to recalibrate target features. Through a coarse-to-fine alignment loss between online target features and pre-computed source domain statistics, LOSA autonomously learns domain-specific style adaptation strategies. By adopting a dynamically optimized high learning rate, LOSA only requires a small number of samples for test-time training, making it suitable for real-time applications. Experimental results across various scenarios, including normal-to-corrupted, cross-band, and sim-to-real adaptation, demonstrate that the proposed method significantly improves cross-domain object detection performance. Moreover, our method can be well generalized to cross-domain image classification tasks. Bin Luo 0005, Jun Liu 0072, Xin Su 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | MTCSAR: Fusing Multitemporal Interaction With Coherence Prior for SAR Image DenoisingabstractSynthetic aperture radar (SAR) images are inherently affected by speckle noise due to their imaging principles, significantly impacting downstream research. Denoising methods based on deep learning have garnered attention and are gradually maturing, yet they face certain challenges. Denoising single SAR images lacks temporal information, often resulting in structural fitting that introduces artifacts. The existing denoising techniques do not adequately account for the differences between SAR and optical images, despite their distinct structural characteristics. In addition, mainstream deep learning approaches heavily rely on data-driven methods, with limited consideration for statistical properties. In response to these challenges, we propose a novel network for denoising SAR images fusing multitemporal interaction with coherence prior, termed MTCSAR. The proposed network leverages multitemporal interaction (MTI) to gather information from different moments in various regions. The redundancy across these temporal dimensions helps mitigate artifacts and edge blurring. To address SAR image characteristics, we design a dual-branch network combining global and local information to handle large-scale regions and fine structures. In addition, we incorporate prior coherence information from SAR images into the network, utilizing statistical properties to enhance transparency during training. Experimental results on simulated and real datasets demonstrate that injecting MTIs and coherence improves our method’s qualitative and quantitative performance, surpassing current state-of-the-art algorithms. This validates the effectiveness of the proposed MTCSAR for multitemporal SAR image denoising. Xin Su 0003, Yi Xiao 0003, Jie Li 0022, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Diffpurifier: An Optical and SAR Image Change Detection Method Based on Diffusion PurificationabstractDisasters occur continuously around the world, causing significant damage to human life and property. In practice, there are inevitably situations that require the use of heterogeneous remote sensing images, for instant optical and synthetic aperture radar (SAR) images, for change detection (CD) and disaster recognition. However, the significant differences between heterogeneous remote sensing images make direct comparison difficult. Although some methods use image translation (IT) to reduce these differences, most employ a two-stage ”translation followed by change detection” strategy, which can lead to feature degradation. Furthermore, these methods often require adversarial training or complex loss functions that are sensitive to hyperparameters. This paper proposes a change detection network for optical and SAR images, named Diffpurifier. First, optical images are translated into SAR images using pre-trained denoising diffusion probabilistic models (DDPMs) and ordinary differential equations (ODEs) while simultaneously extracting multi-scale features. Then, change detection is performed under superpixel enhancement to improve the homogeneity of the change detection maps. Diffpurifier not only integrates image translation and feature extraction, simplifying the workflow, but also maintains high accuracy, stable training, and generalization to different types of data without the need for additional translation constraints. In comparative experiments on four public datasets, Diffpurifier outperforms the second-best method by an average of approximately 5% in terms of F1-score, validating the effectiveness and robustness of the method. Yiquan Xu, Xin Su 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Hyperspectral Image Denoising Based on Hyper-Laplacian Total Variation in Spectral Gradient DomainabstractThis article introduces a novel hyperspectral image (HSI) denoising method based on Hyper-Laplacian Total Variation in Spectral Gradient (HLTVSG). Our research identifies that compared to the original HSI, the spectral gradient of HSIs exhibits both a lower rank and a more favorable structure for total variation (TV) regularization due to its lower image entropy. To capitalize on the low-rank property of HSIs, we employ a subspace representation approach within the spectral gradient domain. Beyond low rankness, HSIs also possess abundant spatial information, which is often regularized by using TV. Traditional TV inherently assumes that the$\ell _{1}$-norm of spatial gradient follows the Laplacian distribution. However, our analysis reveals that the spatial gradient of the subspace representation coefficients obeys a hyper-Laplacian distribution. Therefore, we introduce the hyper-Laplacian TV (HLTV) to better constrain the spatial local smoothness (LSS) of the representation coefficients. In addition, to eliminate the sparse noise and the stripe noise, an$\ell _{1}$-norm and low-rank constraint are also employed. Finally, an augmented Lagrange multiplier (ALM) algorithm is utilized to solve the proposed model. Extensive experimental results on simulation and real datasets indicate that the proposed method outperforms many state-of-the-art methods both visually and quantitatively. Fang Yang 0005, Qiangfu Hu, Xin Su 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | GLIFT: A Global-to-Local Invariant Feature Transformation Method for Multimodal Remote Sensing Image Matching
Shaochen Zhang, Bin Luo 0005, Jun Liu 0072, Zhitao Fu, Xin Su 0003, Shiliang Zhu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual ModelsabstractRecently, there has been a surge in interest in Large Language Models (LLMs), with ChatGPT standing out for its exceptional capabilities in language comprehension, reasoning, and interactive communication. These models have garnered attention from a diverse array of users and researchers across various disciplines. While LLMs have demonstrated remarkable proficiency in mimicking human task execution through natural language, their application in remote sensing interpretation remains largely uncharted. Furthermore, the current lack of automation in remote sensing task planning limits the accessibility of these sophisticated interpretation techniques, especially for non-specialists in the field. To bridge this gap, we introduce Remote Sensing ChatGPT, an innovative LLM-driven agent that integrates ChatGPT with a suite of AI-powered remote sensing models to tackle complex interpretation challenges. This system is designed to interpret user requests, delineate task planning based on the functionalities required, execute each subtask sequentially, and compile the final output by synthesizing the results from each stage. Given that LLMs, trained predominantly on natural language, do not inherently comprehend visual elements present in remote sensing imagery, we have devised a method to incorporate visual cues, effectively embedding the visual context of remote sensing images into the ChatGPT framework. With Remote Sensing ChatGPT, users can effortlessly submit a remote sensing image alongside their query and promptly receive detailed interpretation outcomes along with comprehensive linguistic feedback. Experiments and case studies demonstrate that our method is adept at handling a diverse range of remote sensing tasks and has the potential to be expanded to encompass an even wider array of applications with the integration of more advanced models, such as remote sensing foundation model. The code and demo of Remote Sensing ChatGPT is publicly available at https://github.com/HaonanGuo/Remote-Sensing-ChatGPT. Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 2 |
| 2024 | Knowledge-Guided Satellite Image Time Series Classification Network for Crop MappingabstractRecent studies demonstrate the effectiveness of deep learning (DL) methods in crop mapping using Satellite Image Time Series (SITS). Despite this progress, the existing crop mapping methods often neglect prior knowledge, leading to unsatisfactory performance, especially in scenarios with limited data. To address this limitation, this paper introduces the Knowledge-Guided Crop Mapping Network (KGCMNet) as an innovative solution for crop mapping. Beyond a spatiotemporal encoder and a crop type prediction decoder, KGCMNet incorporates a knowledge reconstruction (KR) module for introducing essential crop-related prior knowlege. The KR module, utilizing NDVI as the chosen knowledge, aids the encoder in learning discriminative phenology from the SITS data. Experimental results demonstrate the outperformance of KGCMNet compared to other methods, showcasing the effectiveness of KR module in improving the accuracy across various crop types. Xiaolei Qin, Xin Su 0003, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2024 | Statistic Ratio Attention-Guided Siamese U-Net for SAR Image Semantic Change DetectionabstractSemantic change detection, which aims to locate land cover changes and identify their categories using pixel-level boundaries, has promising applications in Earth vision, including precise urban planning and natural resource management. This paper proposes a novel Siamese U-Net architecture for semantic change detection in synthetic aperture radar (SAR) images, incorporating a residual network with weight-sharing as the backbone network. The network is capable of simultaneously yielding binary change detection and semantic change detection results. Additionally, we have designed a statistic ratio attention module that utilizes statistical features from the original image as spatial ratio attention, coupled with channel attention, to extract change information from the bi-temporal SAR images. Furthermore, as there is currently no existing dataset for semantic change detection in SAR images, we have constructed a dedicated dataset to facilitate model training and evaluation. Our experiment results demonstrate the superiority of our proposed model over other comparison algorithms. Shuhui Chen, Xin Su 0003, Li Zheng 0004, Qiangqiang Yuan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | PHTrack: Prompting for Hyperspectral Video TrackingabstractHyperspectral (HS) video captures continuous spectral information of objects, enhancing material identification in tracking tasks. It is expected to overcome the inherent limitations of red-green–blue (RGB) and multimodal tracking, such as finite spectral cues and cumbersome modality alignment. However, HS tracking faces challenges such as data anxiety, bandgaps, and huge volumes. In this study, inspired by prompt learning in language models, we propose the prompting for hyperspectral video tracking (PHTrack) framework. PHTrack learns prompts to adapt foundation models, mitigating data anxiety and enhancing performance and efficiency. First, the modality prompter (MOP) is proposed to capture rich spectral cues and bridge bandgaps for improved model adaptation and knowledge enhancement. In addition, the distillation prompter (DIP) is developed to refine cross-modal features. PHTrack follows feature-level fusion, effectively managing huge volumes compared to traditional decision-level fusion fashions. Extensive experiments validate the proposed framework, offering valuable insights for future research. The code and data will be available athttps://github.com/YZCU/PHTrack Yuzeng Chen, Xin Su 0003, Jie Li 0022, Yi Xiao 0003, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Building-Road Collaborative Extraction From Remote Sensing Images via Cross-Task and Cross-Scale InteractionabstractBuildings and roads are the two most basic man-made environments that carry and interconnect human society. Building and road information has important application value in the frontier fields of regional coordinated development, disaster prevention, auto-driving, etc. Mapping buildings and roads from very high-resolution (VHR) remote sensing images has become a hot research topic. However, the existing methods often extract buildings and roads with separate models, ignoring their strong spatial correlation. To fully utilize their complementary relation, we propose a method that simultaneously extracts buildings and roads from remote sensing images. The accuracy of both tasks can be improved using our proposed multi-task feature interaction and cross-scale feature interaction modules. To be specific, a multi-task interaction module is proposed to interact information across building extraction and road extraction tasks while preserving the unique information of each task. Furthermore, a cross-scale interaction module is designed to automatically learn the optimal reception field for buildings and roads under varied appearances and structures. Compared with existing methods that train individual models for each task separately, the proposed collaborative extraction method can utilize the complementary advantages between buildings and roads and reduce the inference time by half using a single model. Experiments on a wide range of urban and rural scenarios show that the proposed algorithm can achieve building-road extraction with outstanding performance and efficiency. Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Novel Rotation and Scale Equivariant Network for Optical-SAR Image Matching
Bin Luo 0005, Jun Liu 0072, Zhitao Fu, Chenjie Wang, Xin Su 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | PCDASNet: Position-Constrained Differential Attention Siamese Network for Building Damage AssessmentabstractSudden natural disasters and man-made disasters pose a threat to human life and property safety, and real-time semantic segmentation of high-resolution remote sensing images is crucial for disaster damage assessment applications. In recent years, with the wide application of high spatial resolution (HSR) remote sensing images and semantic change detection methods based on deep learning (DL), the acquisition of information on damaged areas has become more and more convenient and accurate. However, due to the black box characteristics of existing methods, the lack of interpretability and prior knowledge embedding (such as building positioning information), as well as the low utilization of damage conditions around the building, lead to automatically learned feature representations that still need to be improved. To solve these problems, we proposed the position-constrained differential attention siamese network (PCDASNet). The main idea is to merge building extraction and disaster damage assessment into a cascaded framework to improve building damage recognition results under the constraints of building positioning information. In particular, the proposed Differential Attention Module (DAM) adaptively extracts change information corresponding to buildings and surrounding environments from dual-temporal images, with interpretability and theoretical guarantee, which enables the integration of prior positioning knowledge into the design of network architecture. The objective metrics of the method on the building damage dataset show that the method achieves a test F1 score of more than 73% compared with other baseline methods, and also outperforms several state-of-the-art methods in terms of visual results. Xin Su 0003, Li Zheng 0004, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | HVL-SLAM: Hybrid Vision and LiDAR Fusion for SLAMabstractIn the field of simultaneous localization and mapping (SLAM), map-based localization has been widely used in autonomous driving, particularly for all-speed and all-road adaptive cruise, automatic parking, and other high-level functions. As a result, LiDAR sensors are frequently used in visual-based SLAM to improve the overall accuracy of ego-motion estimation and environment reconstruction. In this article, a novel tightly coupled monocular hybrid visual LiDAR SLAM (HVL-SLAM), which utilizes both visual and LiDAR measurements in tracking and mapping. First, the proposed method reduces the 3-D uncertainty of features by employing object segmentation and Delaunay triangulation. The motion between adjacent frames is then estimated using a hybrid tracking module that minimizes photometric and reprojection error. Finally, a joint optimization method for refining the pose is proposed, which incorporates visual and LiDAR measurements into optimization with dynamic weights, resulting in higher positioning accuracy and robustness. The experiments on the public KITTI odometry benchmark and real-world outdoor datasets demonstrate that HVL-SLAM outperforms state-of-the-art approaches in terms of pose estimation and mapping performance. The code is released to the community. Code available athttps://github.com/kinggreat24/hvl_slam. Wei Wang 0323, Chenjie Wang, Jun Liu 0072, Xin Su 0003, Bin Luo 0005, Cheng Zhang 0037 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | SAAN: Similarity-Aware Attention Flow Network for Change Detection With VHR Remote Sensing ImagesabstractChange detection (CD) is a fundamental and important task for monitoring the land surface dynamics in the earth observation field. Existing deep learning-based CD methods typically extract bi-temporal image features using a weight-sharing Siamese encoder network and identify change regions using a decoder network. These CD methods, however, still perform far from satisfactorily as we observe that 1) deep encoder layers focus on irrelevant background regions; and 2) the models' confidence in the change regions is inconsistent at different decoder stages. The first problem is because deep encoder layers cannot effectively learn from imbalanced change categories using the sole output supervision, while the second problem is attributed to the lack of explicit semantic consistency preservation. To address these issues, we design a novel similarity-aware attention flow network (SAAN). SAAN incorporates a similarity-guided attention flow module with deeply supervised similarity optimization to achieve effective change detection. Specifically, we counter the first issue by explicitly guiding deep encoder layers to discover semantic relations from bi-temporal input images using deeply supervised similarity optimization. The extracted features are optimized to be semantically similar in the unchanged regions and dissimilar in the changing regions. The second drawback can be alleviated by the proposed similarity-guided attention flow module, which incorporates similarity-guided attention modules and attention flow mechanisms to guide the model to focus on discriminative channels and regions. We evaluated the effectiveness and generalization ability of the proposed method by conducting experiments on a wide range of CD tasks. The experimental results demonstrate that our method achieves excellent performance on several CD tasks, with discriminative features and semantic consistency preserved. Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | An Ensemble Learning Approach With Attention Mechanism for Detecting Pavement Distress and Disaster-Induced Road DamageabstractRoad damage presents a significant risk to traffic safety, including pavement distress and disaster-induced damage. Thanks to their high efficiency, computer vision-based methods for pavement distress detection have been widely developed. In disaster scenarios, the automatic extraction of road damage information from extensive social media images plays a critical role in rescue efforts. However, few existing studies have focused on detecting object-level disaster-induced road damage. To fill the gap, this paper presents a Social media image dataset of Object detection for Disaster-induced Road damage (SODR), including 1,552 images and two categories (i.e., collapses and blockages). Additionally, this paper proposes an ensemble learning approach with attention mechanisms based on YOLOv5 (You Only Look Once) network. Initially, attention modules are employed to create two distinct detectors for ensemble learning. Subsequently, one standard YOLOv5 and two variant networks are trained with consistent settings, and test time augmentation is applied during the inference phase. The proposed method has been implemented across five scales of YOLOv5, offering alternatives for balancing accuracy and computational cost. To demonstrate the validity, comprehensive experiments were conducted on two datasets. Compared with some mainstream detectors and ensemble learning methods, our approach achieved competitive results with a fewer number of parameters and a simpler training and testing process. The SODR dataset and source code are available at (https://github.com/nonondayo/yolov5_SODRv1). Shouxing Wang, Hongzan Jiao, Xin Su 0003, Qiangqiang Yuan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Decoupling Semantic and Edge Representations for Building Footprint Extraction From Remote Sensing ImagesabstractVery high-resolution (VHR) earth observation systems provide an ideal data source for man-made structure detection such as building footprint extraction. Manually delineating building footprints from the remotely sensed VHR images, however, is laborious and time-intensive; thus, automation is needed in the building extraction process to increase productivity. Recently, many researchers have focused on developing building extraction algorithms based on the encoder-decoder architecture of convolutional neural networks. However, we observe that this widely adopted architecture cannot well preserve the precise boundaries and integrity of the extracted buildings. Moreover, features obtained by shallow convolutional layers contain irrelevant background noises that degrade building feature representations. This paper addresses these problems by presenting a feature decoupling network (FD-Net) that exploits two essential building information from the input image, including semantic information that concerns building integrity and edge information that improves building boundaries. The proposed FD-Net improves the existing encoder-decoder framework by decoupling image features into the edge subspace and the semantic subspace; the decoupled features are then integrated by a supervision-guided fusion process considering the heterogeneity between edge and semantic features. Furthermore, a lightweight and effective global context attention module is introduced to capture contextual building information and thus enhance feature representations. Comprehensive experimental results on three real-world datasets confirm the effectiveness of FD-Net in large-scale building mapping. We applied the proposed method to various encoder-decoder variants to verify the generalizability of the proposed framework. Experimental results show remarkable accuracy improvements with less computational cost. Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Fast Hyperspectral Image Denoising and Destriping Method Based on Graph Laplacian RegularizationabstractHyperspectral images (HSIs) contain rich spatial and spectral information about the earth, and are widely used in the remote sensing field. However, an HSI is frequently corrupted by various types of noise, such as Gaussian noise, sparse noise, stripe noise, and so on, which severely limits the subsequent application of the HSI. In this paper, we propose a graph Laplacian regularizer (GLR) to exploit the low-rank information across the bands of the HSI. Compared with the traditional low-rank regularization, our graph Laplacian regularization can achieve equivalent or better performance with less time consumption. Besides, a sparse constraint and a low-rank constraint are employed to remove the sparse and stripe noise. In addition, the augmented Lagrangian multiplier is used to solve each component to restore a clean image. Finally, we have carried out experiments on the simulated and real noisy HSI data. The results show the superiority of the proposed method over state-of-the-art methods, in terms of PSNR and SSIM, time cost and visual effect. The MATLAB code is available through: https://github.com/zzhang-99/FGLR. Xin Su 0003, Zhi Zhang 0011, Fang Yang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Multilevel Attention Siamese Network for Keypoint Detection in Optical and SAR ImagesabstractOptical and synthetic aperture radar (SAR) image keypoint detection is an important foundation for multimodal remote sensing image matching. The influence of nonlinear radiometric differences and geometric deformation between optical and SAR images leads to low repeatability of existing keypoint detection methods. To address the problem that existing keypoint detection methods cannot provide the required homonymous points for heterogenous image matching, we propose a keypoint detection method (SKD-Net) for optical and SAR images, and improve it in terms of both network structure and network optimization. First, we propose a multilevel attention Siamese network, which is composed of multiple convolutional modules and transformer modules with shared weights to extract common features at different levels for keypoint detection. We introduce a transformer module in the keypoint detection pipeline and fuse shallow and deep features to obtain more spatial and rich semantic information to facilitate heterogeneous image keypoint detection. Then, to ensure that the detected keypoints have more homonymous points and localization accuracy, we propose a position consistent loss. Unlike previous loss functions, our designed position-consistent loss function takes the differences between heterogeneous image score maps into account, and it autonomously selects the optimized correct point pairs to enable the network to perform correct learning. Finally, extensive experiments show that our detection method outperforms the current state-of-the-art keypoint detection methods in terms of repeatability, localization accuracy, and matching performance. Our source code is available at https://github.com/zhangschen/ SKD-Net. Shaochen Zhang, Zhitao Fu, Jun Liu 0072, Xin Su 0003, Bin Luo 0005, Bo-Hui Tang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | An Ensemble Learning Approach with Multi-depth Attention Mechanism for Road Damage DetectionabstractRoad damage detection is significant for road maintenance. Traditional manual visual inspection methods consume lots of time and labor. Developments in the field of computer vision create opportunities for automated and efficient image-based road damage detection. Through deep convolution neural networks, road damage localization and classification can be achieved simultaneously. This paper proposes an ensemble model with test time augmentation based on the You Only Look Once (YOLOv5) network and attention modules. The approach utilizes a state-of-the-art object detector known as YOLOv5. To focus more on the road in images, five improved YOLOv5 models with attention modules are proposed. Moreover, ensemble learning and test time augmentation are adopted to improve model generalization and detection performance. The proposed method was evaluated through the IEEE Big Data Crowdsensing-based Road Damage Detection Challenge 2022. Different ensemble models achieved an average F1-score of 0.65177 on the five test datasets. Shouxing Wang, Xusi Liao, Haoliang Feng, Hongzan Jiao, Xin Su 0003, Qiangqiang Yuan |
IEEE Big Data | 7 |
| 2022 | Learning an Intrinsic Graph Neural Network for Sartellite Video Super-ResolutionabstractExisting video super-resolution (VSR) methods usually merge the redundant temporal information along frames to achieve information enhancement, which naturally discards the spatial redundancy information. This paper proposes an intrinsic Graph Neural Network (GNN) framework for satellite VSR to fully explore the internal spatial prior while considering the temporal information in the video frame sequence. Firstly, a Multi-Scale Deformable convolution (MSD) is adopted to accurately model the spatial-temporal relationship between frames. Then, we search for k-nearest neighbors to construct the spatial graph and profoundly excavate the prior spatial information brought by patch recurrence. Finally, the spatial-temporal redundant information is integrated and complementary. Experiments on Jilin-1 satellite video demonstrate the effectiveness of our framework. Yi Xiao 0003, Xin Su 0003, Qiangqiang Yuan |
IGARSS | 2 |
| 2022 | New Building Detection Using SAR Images with Different ResolutionsabstractIn this paper, Synthetic Aperture Radar (SAR) images including high resolution Gaofen-3 images and medium resolution Sentinel-1 images are used to detect the new buildings, and the detection results of different resolutions were compared and analyzed. The SAR image change detection algorithm based on likelihood ratio change matrix clustering was used, and the new buildings were extracted by threshold from the detected changes. The detection results show that images acquired by the two sensors have good performance on new building detection. However, the noise and geometric distortion in urban area of Gaofen-3 image reduce the detection accuracy, while the resolution of Sentinel-1 is insufficient for the detection of new building with small size. Chengyu Zou, Lingli Zhao, Peilei Sun, Bingxue Fu, Xin Su 0003 |
IGARSS | 5 |
| 2022 | Spatial-Temporal Gray-Level Co-Occurrence Aware CNN for SAR Image Change DetectionabstractDeep learning-based synthetic aperture radar (SAR) image change detection has recently achieved remarkable success due to its great potential for extracting abstract features. However, the existing methods still have room for improvement in dealing with the speckle of SAR images. In this letter, a deep spatial–temporal gray-level co-occurrence aware convolutional neural network (STGCNet) is proposed, which can effectively mine the spatial–temporal information of the bitemporal images and obtain the speckle-robust results by introducing the 3-D gray-level co-occurrence matrix (3-D-GLCM) as auxiliary feature. Specifically, representative features are extracted from original image pairs and their corresponding 3-D-GLCM through two-stream network, followed by an adaptive fusion module to balance the contribution of each branch. Then, the final binary change detection results are obtained by a fully connected layer. The training process relies on reliable labels generated by unsupervised models rather than manually annotated data, and therefore, the proposed STGCNet is practical in reality. Experiments on synthesized and real SAR data sets demonstrate the robustness and competitiveness of the proposed method compared with the state-of-the-art algorithms. Xin Su 0003, Qiangqiang Yuan, Qing Wang 0046 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Multivehicle Object Tracking in Satellite Video Enhanced by Slow Features and Motion FeaturesabstractWith the development of video satellites, multimoving object tracking in satellite video is possible and has become a new challenging task. The difficulties are mainly caused by the characteristics of satellite videos: 1) small objects; 2) low contrast between objects and background; and 3) background in a state of continuous motion. These characteristics make it difficult for the advanced multiobject tracking algorithms in the natural video to give full play to their advantages, resulting in vast false alarms, missed objects, ID switches, and low-confidence bounding boxes. To tackle these problems, a novel multimoving object tracking method considering slow features (SFs) and motion features has been proposed in this research, named SF and motion feature-guided multiobject tracking (SFMFMOT), which realizes the continuous tracking of moving vehicles in satellite videos. A nonmaximum suppression (NMS) module guided by bounding box proposals based on SFs is designed to assist the object detection part by utilizing the sensitivity of SF analysis to the changed pixels. While removing a large number of static false alarms and supplementing missed objects, it improves the recall rate by increasing the confidence score of the correctly detected object bounding boxes. In order to improve the tracking performance, a set of optimization strategies based on motion features and time accumulation information are proposed to smooth the trajectory, remove static false alarms, and duplicate bounding boxes. The proposed method is evaluated in three satellite videos and its superiority is demonstrated. Jialian Wu, Xin Su 0003, Qiangqiang Yuan, Huanfeng Shen, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Satellite Video Super-Resolution via Multiscale Deformable Convolution Alignment and Temporal Grouping ProjectionabstractAs a new earth observation tool, satellite video has been widely used in remote-sensing field for dynamic analysis. Video super-resolution (VSR) technique has thus attracted increasing attention due to its improvement to spatial resolution of satellite video. However, the difficulty of remote-sensing image alignment and the low efficiency of spatial–temporal information fusion make poor generalization of the conventional VSR methods applied to satellite videos. In this article, a novel fusion strategy of temporal grouping projection and an accurate alignment module are proposed for satellite VSR. First, we propose a deformable convolution alignment module with a multiscale residual block to alleviate the alignment difficulties caused by scarce motion and various scales of moving objects in remote-sensing images. Second, a temporal grouping projection fusion strategy is proposed, which can reduce the complexity of projection and make the spatial features of reference frames play a continuous guiding role in spatial–temporal information fusion. Finally, a temporal attention module is designed to adaptively learn the different contributions of temporal information extracted from each group. Extensive experiments on Jilin-1 satellite video demonstrate that our method is superior to current state-of-the-art VSR methods. Yi Xiao 0003, Xin Su 0003, Qiangqiang Yuan, Denghong Liu, Huanfeng Shen, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Self-supervised Hyperspectral and Multispectral Image Fusion in Deep Neural Network
Jianhao Gao, Jie Li 0022, Qiangqiang Yuan, Xin Su 0003 |
ICIG (3) | 5 |
| 2021 | A Recurrent Refinement Network for Satellite Video Super-ResolutionabstractDeep learning-based methods have shown superior performance in VSR tasks. However, satellite video frames are characterized by large width, low resolution, and lack of features. Consequently, the conventional VSR method is not suitable for satellite video. In this paper, a recurrent refinement network is proposed. Considering that the vast majority of remote sensing images belong to the static background, a single-image SR (SISR) method is first used to obtain high-resolution features for a specific target frame. To further complement the missing details, the network learns the complementary information enhanced by an Encoder-Decoder structure from adjacent frames to refine the results of SISR. To measure the contribution of different adjacent frames to the recovery of the target frame, a temporal attention mechanism is introduced in the final fusion stage. The experiment on the video data of Jilin-1 demonstrates the effectiveness of our method. Yi Xiao 0003, Xin Su 0003, Qiangqiang Yuan |
IGARSS | 2 |
| 2020 | Geometry-Aware Graph Transforms for Light Field Compact RepresentationabstractThe paper addresses the problem of energy compaction of dense 4D light fields by designing geometry-aware local graph-based transforms. Local graphs are constructed on super-rays that can be seen as a grouping of spatially and geometry-dependent angularly correlated pixels. Both non separable and separable transforms are considered. Despite the local support of limited size defined by the super-rays, the Laplacian matrix of the non separable graph remains of high dimension and its diagonalization to compute the transform eigen vectors remains computationally expensive. To solve this problem, we then perform the local spatio-angular transform in a separable manner. We show that when the shape of corresponding super-pixels in the different views is not isometric, the basis functions of the spatial transforms are not coherent, resulting in decreased correlation between spatial transform coefficients. We hence propose a novel transform optimization method that aims at preserving angular correlation even when the shapes of the super-pixels are not isometric. Experimental results show the benefit of the approach in terms of energy compaction. A coding scheme is also described to assess the rate-distortion perfomances of the proposed transforms and is compared to state of the art encoders namely HEVC-lozenge [1], JPEG pleno 1.1 [2], HEVC-pseudo [3] and HLRA [4]. Mira Rizkallah, Xin Su 0003, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2018 | Graph-based Transforms for Predictive Light Field Compression based on Super-PixelsabstractIn this paper, we explore the use of graph-based transforms to capture correlation in light fields. We consider a scheme in which view synthesis is used as a first step to exploit inter-view correlation. Local graph-based transforms (GT) are then considered for energy compaction of the residue signals. The structure of the local graphs is derived from a coherent super-pixel over-segmentation of the different views. The GT is computed and applied in a separable manner with a first spatial unweighted transform followed by an inter-view GT. For the inter-view GT, both unweighted and weighted GT have been considered. The use of separable instead of non separable transforms allows us to limit the complexity inherent to the computation of the basis functions. A dedicated simple coding scheme is then described for the proposed GT based light field decomposition. Experimental results show a significant improvement with our method compared to the CNN view synthesis method and to the HEVC direct coding of the light field views. Mira Rizkallah, Xin Su 0003, Thomas Maugey, Christine Guillemot |
ICASSP | 2 |
| 2017 | Graph-based light fields representation and coding using geometry informationabstractThis paper describes a graph-based coding scheme for light fields (LF). It first adapts graph-based representations (GBR) to describe color and geometry information of LF. Graph connections describing scene geometry capture inter-view dependencies. They are used as the support of a weighted Graph Fourier Transform (wGFT) to encode disoccluded pixels. The quality of the LF reconstructed from the graph is enhanced by adding extra color information to the representation for a sub-set of sub-aperture images. Experiments show that the proposed scheme yields rate-distortion gains compared with HEVC based compression (directly compressing the LF as a video sequence by HEVC). Xin Su 0003, Mira Rizkallah, Thomas Maugey, Christine Guillemot |
ICIP | 1 |
| 2017 | Rate-Distortion Optimized Graph-Based Representation for Multiview Images With Complex Camera ConfigurationsabstractGraph-based representation (GBR) has recently been proposed for describing color and geometry of multiview video content. The graph vertices represent the color information, while the edges represent the geometry information, i.e., the disparity, by connecting corresponding pixels in two camera views. In this paper, we generalize the GBR to multiview images with complex camera configurations. Compared with the existing GBR, the proposed representation can handle not only horizontal displacements of the cameras but also forward/backward translations, rotations, etc. However, contrary to the usual disparity that is a 2-D vector (denoting horizontal and vertical displacements), each edge in GBR is represented by a 1-D disparity. This quantity can be seen as the disparity along an epipolar segment. In order to have a sparse (i.e., easy to code) graph structure, we propose a rate-distortion model to select the most meaningful edges. Hence the graph is constructed with "just enough" information for rendering the given predicted view. The experiments show that the proposed GBR allows high reconstruction quality with lower or equivalent coding rate than traditional depth-based representations. Xin Su 0003, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 1 |
| 2016 | Graph-based representation for multiview images with complex camera configurationsabstractInstead of lossily coding depth images resulting in undesirable geometric distortion, graph-based representation (GBR) describes disparity information as a graph with a controllable accuracy. In this paper, we propose a more compact graphical representation called GBR-plus to code both disparity and color information of a target view given a reference view. Specifically, first we differentiate between disocclusion holes (occluded spatial regions in the reference view) and rounding holes (insufficiently sampled regions in the reference view) in the synthesized target view, so that the decoder can optionally complete rounding holes via signal interpolation without coding overhead. Second, we use a compact graphical representation to delimit disparity-shifted boundaries of objects in the target view, which is coded losslessly. Finally, color pixels in disocclusion holes are predicted using adjacent background pixels as predictors, and prediction residuals in a local neighborhood are coded using Graph Fourier Transform (GFT). Experimental results show that GBR-plus outperforms previous GBR, and has comparable performance as HEVC at mid to high bitrates with lower encoder complexity. Xin Su 0003, Thomas Maugey, Christine Guillemot |
ICIP | 1 |
| 2015 | Local Topographic Shape Patterns for Texture DescriptionabstractThis letter introduces a new image descriptor named Local Topographic Shape Pattern (LTSP) for texture description by relying on complete shape-based image representation. Firstly, a texture image is decomposed into a tree of shapes, i.e., topographic map, to obtain a multi-scale and contrast-invariant representation. Secondly, the shape of each node in the tree is handled with four shape patterns which are proposed as spatial codes of the shape. Finally, statistical histograms of shape patterns are used as texture descriptors. Contrast experiments of the proposed method show satisfactory performance in terms of both texture retrieval and classification tasks when applied to the Brodatz and UMD texture datasets. Chu He, Tong Zhuo, Xin Su 0003, Feng Tu |
IEEE Signal Process. Lett. | 3 |
| 2013 | The algorithm of building area extraction based on boundary prior and conditional random field for SAR imageabstractIn this paper, an algorithm applied for building area extraction on SAR image is proposed, which is based on conditional random model, then a boundary prior relation is introduced to strengthen the description of prior item around the edge of building area, aiming at improving the classification performance nearby the boundary lines encompass building area. Firstly, pre-segmentation and boundary lines extraction can be accomplished respectively rely on mean shift algorithm and ratio of average edge detection. After that a combination term of the distances between the boundary lines and pixels around them and the pixels' label information can help to improve the prior item in CRF and build the boundary prior-CRF model. Finally, several experimental results on TerraSAR-X images prove that the proposed approach significantly improves the extraction accuracy and classification performance when compared to CRF. Chu He, Yu Zhang 0019, Xin Su 0003, Wen Yang 0001, Xin Xu 0005 |
IGARSS | 4 |
| 2013 | Target detection on high-resolution SAR image using Part-based CFAR ModelabstractThis letter proposed a Part-based CFAR Model for object detection of power tower on high-resolution SAR images. Firstly, Part-based Model is used to describe the structure feature of the target, then Compressing Sensing approach is added to reduce the speckle by means of rebuilding background clutter, next, CFAR method is used to extract local shape and scale parameters, at last, Part-based CFAR Model combines these procedures together to form the finally algorithm, not only includes the distribution features, but also considers the structure relationship in the proposed approach. The algorithm is tested on TerraSAR-X data set with the resolution of 1m and 3m. Experiments show that unlike the CFAR method can only gives the high-light points of the targets; Part-based CFAR Model illuminates the target and its local components by plotting the bounding boxes around them. Chu He, Yu Zhang 0019, Xin Su 0003, Xin Xu 0005, Mingsheng Liao |
IGARSS | 3 |
| 2011 | A Supervised Classification Method Based on Conditional Random Fields With Multiscale Region Connection Calculus Model for SAR ImageabstractThis letter presents a supervised classification method for synthetic aperture radar (SAR) images based on multiscale region connection calculus (RCC) and conditional random fields (CRF). Using this method, first, a SAR image is oversegmented into multisuperpixels via the image pyramid. We then use the multiscale RCC model to describe the spatial logic relationships among these superpixels. To complete the process, multiscale RCC relationships are learned and reasoned under the CRF reasoning framework. This method employs iteration strategy for CRF reasoning to get better details in the classification results as well. We illustrate the proposed method by experiments conducted on DLR ESAR image. The results reveal efficient performance. Xin Su 0003, Chu He, Xinping Deng |
IEEE Geosci. Remote. Sens. Lett. | 1 |