EDBT 2026 Demo / reviewers in the wild / expert
Liang Chen 0004
dblp:01/5394-4
· DBLP profile ↗
52ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0003-3273-5522ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 45 · 2 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Style-adaptive detection transformer for single-source domain generalized object detection
Jianhong Han, Yupei Wang, Liang Chen 0004 |
Neurocomputing | 3 |
| 2025 | Boosting Domain Generalization in Remote Sensing Image Segmentation via Style Mapping and General Prototypical Contrast
Yupei Wang, Xiaoxing Hu, Yongkang Hu, Shanghang Zhang, Liang Chen 0004 |
Int. J. Comput. Vis. | 6 |
| 2025 | High-Throughput and Energy-Efficient FPGA-Based Accelerator for All Adder Neural NetworksabstractNeural networks have been extensively applied across various Internet of Things (IoT) applications, such as drone- and satellite-based remote sensing and autonomous driving. With the increasing resolution and amount of data captured by sensors, the demand for real-time response in IoT applications is markedly increasing. However, it is difficult for existing convolutional neural network (CNN) accelerators for IoT applications on field-programmable gate array (FPGA) platforms to achieve high throughput because of the inherent dense multiplication operations of CNNs, memory bandwidth limitations and inefficient mapping mechanisms. In this article, a high-throughput and energy-efficient all adder neural network (A2NN) accelerator for IoT applications on FPGA platform is proposed to solve this problem. First, a series of hardware-oriented algorithm optimization methods are proposed to simplify the processing flow of A2NN and further minimize its deployment overhead. Second, a novel hardware architecture based on the idea of near-memory computation (NMC) is proposed to eliminate off-chip memory access completely and accelerate the reconstructed A2NN in the pipeline. Third, a set of quantitative analysis methods for the proposed accelerator is presented to balance throughput and energy consumption, allowing the accelerator to adapt to the varying demands of different IoT application scenarios. Extensive experimental results on the AMD-Xilinx VC709 board demonstrate that the proposed accelerator achieves state-of-the-art performance in terms of throughput, energy efficiency, and throughput efficiency. Moreover, experiments on the AMD-Xilinx KV260 board highlight the architecture’s exceptional scalability and energy efficiency, enabling a balance between speed and power consumption tailored to the specific requirements of IoT application scenarios. Ning Zhang 0042, Shuo Ni, Liang Chen 0004, He Chen 0004 |
IEEE Internet Things J. | 3 |
| 2025 | Toward Effective Knowledge Distillation for Fine-Grained Object Recognition in Remote SensingabstractWith advancements in on-board computing devices deployed on remote sensing platforms, the demand for efficiently processing remote sensing imagery has become increasingly prominent. Knowledge distillation, as an effective lightweight method, has been introduced into this domain. Intuitively, distillation from a larger teacher model is expected to yield better performance. However, in our investigation of fine-grained object recognition in remote sensing imagery, we observed a counter-intuitive phenomenon: as the size of the teacher model increases, the performance of the student model initially improves but then degrades. This capacity gap issue hinders the effective utilization of stronger teacher models. To address this issue, we propose a novel distillation framework named BL-KD. It integrates two tailored components: the Class-level Learnable Orthogonal Projection (CLOP) module and the Object Re-Balance (ORB) module, which are jointly optimized to mitigate the negative impact of the capacity gap while effectively adapting to the unique distributional patterns and challenges inherent in remote sensing imagery. Experiments conducted on multiple fine-grained object recognition tasks in remote sensing demonstrate that our method consistently improves student performance, particularly in scenarios involving large teacher-student gaps, and outperforms several widely used distillation baselines. Yangte Gao, Chenwei Deng, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Soft-Guided Open-Vocabulary Semantic Segmentation of Remote Sensing ImagesabstractOpen-vocabulary remote sensing semantic segmentation strives to assign both seen and unseen class labels to individual pixels in remote sensing images. Existing models follow the “fine-tune” paradigm based on Vision-Language Models (VLMs). However, as VLMs are predominantly tailored to natural scenes, these directly fine-tuned models often collapse into the seen categories and show insensitivity in perceiving remote sensing semantic cues. This critical issue of model collapse is closely related with the miss-alignment between image and text, making them struggle with the unique challenges of remote sensing images, such as complex and diverse scenes, and objects with significant scale differences. To this end, we propose a soft-guided open-vocabulary remote sensing semantic segmentation framework, which is the first to explore how to softly adapt VLMs to the downstream task of semantic segmentation for remote sensing images. Concretely, instead of directly fine-tuning, we introduce a generalization compensation strategy, which employs an additional frozen VLM encoder to provide implicit semantic guidance for dynamic optimization of visual representation. By introducing prior knowledge from frozen encoder, this soft strategy compensates potential losses incurred during fine-tuning, thus enhancing the model’s pixel-level perceptual alignment while avoiding model collapse. Afterwards, to optimize the sensitivity of VLMs’ textual and visual embeddings to remote sensing semantic information, bias-guided image-text collaborative optimization is presented to achieve a bilateral interaction of semantic information with the guidance of RS-Bias. Finally, an improved upsampling decoder is employed to obtain the progressive refinement and calibration of cost map through the integration of multi-scale information and textual embeddings. Extensive experiments demonstrate that our method achieves state-of-the-art performance on widely used challenging benchmarks. Code is available at https://github.com/H1NATA111/SGSeg.git. Yupei Wang, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | PGMNet: A Prototype-Guided Multimodal Network for Ship Recognition in SAR ImagesabstractShip recognition in synthetic aperture radar (SAR) images has extensive applications across various fields. However, the substantial intra-class variability and inter-class similarity present inherent challenges to achieving high-precision recognition. Speckle noise in the background reduces the signal-to-noise ratio, complicating the extraction of discriminative characteristics. Additionally, traditional convolutional neural networks-based methods, which rely solely on image processing frameworks without leveraging additional modality information, struggle to accurately represent the features of SAR targets with bright scatters. To address these issues, we propose a prototype-guided multimodal network (PGMNet), marking a pioneering effort to introduce an image-text multimodal fusion processing paradigm into SAR ship recognition. First, a generation strategy of the prototype and the key area is designed to improve the distinguishability between targets and backgrounds. Besides, a prototype-guided alignment module (PGAM) is implemented to assist the network in characterizing key area information, enhancing intra-class feature consistency. Furthermore, a text feature processing branch is incorporated to precisely describe ship size information and effectively integrate image-text multimodal features, reducing intra-class feature distance while enlarging inter-class feature distance. Extensive experiments on the OpenSARShip and FUSARShip datasets demonstrate that the proposed PGMNet achieves state-of-the-art (SOTA) performance. Notably, the accuracy of PGMNet is at least 11% higher than the current SOTA algorithms on the OpenSARShip-VI dataset. Liang Chen 0004, Honghu Zhong, Hao Shi 0006, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | DOGAN: DINO-Based Optical-Prior-Driven GAN for SAR-to-Optical Image TranslationabstractTo leverage the complementary advantages of SAR’s all-weather and all-day imaging capability and optical imagery’s intuitive visualization, SAR-to-optical image translation (S2OIT) has emerged as a promising solution to mitigate the interpretability challenges posed by SAR’s speckle noise and geometric distortions. However, the scale of high-quality registered SAR-optical data is limited, where incorporating priors is a viable solution. What’s more, the digging out of optical prior is insufficient among the existing methods, leading to inadequate synthesis of optical-like texture in translated optical images. To address these challenges, we propose DOGAN, a DINO-based optical-prior-driven GAN framework that integrates ample optical priors extracted from a pretrained DINO model into the S2OIT process. Specifically, to fully exploit the tremendous optical prior preserved in pretrained DINO and extract multiscale optical prior, a DINO-based Optical-prior Extraction (DOE) module is proposed. Furthermore, to elevate the domain adaptability of optical prior, a lightweight Stacked Optical-Aware (SOA) adapter is proposed to fine-tune DINO for remote sensing data with minimal trainable parameters. To instill the extracted affluent optical prior into the S2OIT pipeline stably, the SAR-optical Multi-scale Domain Alignment (SO-MDA) module is proposed, which employs L1 and Multi-kernel Maximum Mean Discrepancy (MK-MMD) losses to align intermediate optical and S2O features. Extensive experiments on SAR2Opt and SEN1-2 datasets demonstrate that DOGAN achieves state-of-the-art performance in both translation fidelity and structural realism. To the best of our knowledge, this is the first work to leverage DINO-based optical priors for the S2OIT task. Jingfei He, Liang Chen 0004, Hao Shi 0006, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Fine-Grained Ship Recognition With Spatial-Aligned Feature Pyramid Network and Adaptive Prototypical Contrastive LearningabstractFine-grained ship recognition endeavors to accurately locate ship targets and recognize their respective fine-grained categories. Current ship recognition methods primarily rely on the feature pyramid network (FPN) for extracting multiscale features. However, FPN exhibits a spatial misalignment issue when fusing features from adjacent-scale feature maps, leading to an inability to extract fine-grained features. Consequently, this limitation constrains the fine-grained recognition capabilities of these recognition methods. Moreover, ship targets possess a high level of intraclass diversity and interclass similarity, yet existing recognition models struggle to extract features with strong category separability, resulting in weakened fine-grained ship recognition performance. In order to solve the spatial misalignment problem that occurs in FPN, a spatial-aligned FPN (SAFPN) is investigated. SAFPN employs a spatial-aware alignment fusion module (SAFM) to effectively extract rich fine-grained features between adjacent-scale feature maps. Moreover, in response to the challenge posed by low category separability in features due to the intraclass diversity and interclass similarity among ship targets, an adaptive prototypical contrastive learning (APCL) method is further proposed. By introducing prototypical contrastive loss, APCL effectively enhances the category separability of ship features, thereby improving the performance of fine-grained ship recognition. Numerous experiments are validated on two fine-grained ship recognition datasets: FGSD and ShipRSImageNet. The experimental results demonstrate that the proposed SAFPN and APCL facilitate the model in extracting fine-grained features with strong category separability, effectively enhancing the performance of fine-grained ship recognition. Our code will be public and available athttps://github.com/liyangfan0/Fine-Grained-Ship-Recognition. Yangfan Li 0002, Liang Chen 0004, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Dual-Domain Representation Modeling With Prototype Contrastive Learning for Cross-Domain Few-Shot Scene ClassificationabstractCross-domain few-shot scene classification (CDF-SSC) aims to establish cross-domain representation between source and target domains, endowing the model few-shot classification ability on the target domain. Recent studies have improved cross-domain representation learning by incorporating unlabeled target data into semi-supervised training with labeled source data. However, these methods struggle to effectively bridge the domain gap between source and target domains, and lack the discriminative feature description ability for unlabeled target domain data, leading to inferior cross-domain representation learning and affecting the few-shot performance. In this article, a dual-domain representation modeling with prototype contrastive learning (DMPC) structure is proposed to improve the robustness of cross-domain representation learning. In DMPC, first, a dual-domain Gaussian representation modeling is designed to model the feature statistics of both source and target data as multivariate Gaussian distributions rather than fixed values and enrich domain representation by random sampling new feature statistics. It helps bridge the domain gap at the feature level, and improves the model’s robustness and generalization to better address unpredictable variations in the target domain. Second, a pseudo-prototype contrastive learning branch is proposed to improve the discriminability of representation for limited unlabeled target data. By leveraging pseudo-prototypes derived from the classifier’s weights as dynamic anchors, it refines feature representation by clustering features of the same pseudo-class and separating those of different pseudo-classes, strengthening the model’s ability to capture distinct and consistent features within the target domain. Finally, the classifier is fine-tuned on few-shot tasks to adapt to specific categories of the target domain. Extensive experimental results exhibit impressive performance of DMPC on 12 RS cross-domain scenarios. Can Li 0005, He Chen 0004, Jianlin Xie, Yin Zhuang, Liang Chen 0004, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | High-Throughput Energy-Efficient Accelerator With Collaborative-Trainable Sparse-Quantization Method for On-Board Remote Sensing ProcessingabstractConvolutional Neural Networks (CNNs) have achieved remarkable breakthroughs on remote sensing tasks in recent years. However, deploying CNNs for real-time remote sensing on-board processing still remains a challenge due to power consumption, real-time and other limitations. Therefore, in this article, a satellite-based real-time remote sensing accelerator is proposed, where algorithm and hardware approaches are proposed to jointly optimize CNNs’ deployment on edge-side aerospace devices. Firstly, a collaborative-trainable sparse-quantization (CTSQ) method is proposed to reduce the model’s storage overhead. In the CTSQ method, analysis of the errors is performed for the sparsity-quantization composition. Besides, the inter-channel correlations among parameters are leveraged, where the structured sparsity and quantization are performed with fine-grained units. Secondly, a modular-system co-optimized (MoSyC) architecture is proposed. A hardware-mapped sparse access (HMSA) strategy is proposed to effectively filter out zero elements in sparse parameters. Moreover, a high-throughput architecture is designed for parallel and pipelined data flow control. Finally, extensive experiments are conducted on both scene classification and object detection tasks with ResNet and YOLOv5 models. The results show that the proposed CTSQ method achieves the compression ratio of more than 13.81 times, and the proposed MoSyC architecture achieves the throughput of more than 1815 GOPS, demonstrating the effectiveness of the proposed accelerator. He Chen 0004, Ning Zhang 0042, Shuo Ni, Xi Zhang 0028, Liang Chen 0004, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Retain and Enhance Modality-Specific Information for Multimodal Remote Sensing Image Land Use/Land Cover ClassificationabstractMultimodal remote sensing (RS) image land use/land cover (LULC) classification using optical and synthetic aperture radar (SAR) images has raised attention for recent studies. Current methods primarily employ multimodal fusion operations to directly explore relationships between multimodal features and obtain fused features, leading to the loss of beneficial modality-specific information problem. To solve this problem, this study introduces a multimodal feature decomposition and fusion (MDF) approach combined with a visual state space (VSS) block, namely MDF-VSS block. The MDF-VSS block emphasizes beneficial modality-specific information and perceives shared land cover information through modality-difference and modality-share features, which are then adaptively integrated to obtain discriminative fused features. Based on the MDF-VSS block, an MDF decoder is designed to retain beneficial multi-scale modality-specific information. Then, a multimodal specific information enhancement (MSIE) decoder is designed to perform modality-difference feature guided auxiliary classification tasks, further enhancing modality-specific information that is expert in classification. Combining the MDF and MSIE decoders, a novel retain-enhance fusion network (REF-Net) is proposed to retain and enhance modality-specific information that benefits classification, thus improving the performance of multimodal RS image LULC classification. Extensive experimental results obtained on three public datasets demonstrate the effectiveness of the proposed REF-Net. The source code will be available at https://github.com/TINYWAI/REF_Net. He Chen 0004, Wenchao Liu 0001, Liang Chen 0004, Jue Wang 0011 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | DATR: Unsupervised Domain Adaptive Detection Transformer With Dataset-Level Adaptation and Prototypical AlignmentabstractWith the success of the DEtection TRansformer (DETR), numerous researchers have explored its effectiveness in addressing unsupervised domain adaptation tasks. Existing methods leverage carefully designed feature alignment techniques to align the backbone or encoder, yielding promising results. However, effectively aligning instance-level features within the unique decoder structure of the detector has largely been neglected. Related techniques primarily align instance-level features in a class-agnostic manner, overlooking distinctions between features from different categories, which results in only limited improvements. Furthermore, the scope of current alignment modules in the decoder is often restricted to a limited batch of images, failing to capture the dataset-level cues, thereby severely constraining the detector's generalization ability to the target domain. To this end, we introduce a strong DETR-based detector named Domain Adaptive detection TRansformer (DATR) for unsupervised domain adaptation of object detection. First, we propose the Class-wise Prototypes Alignment (CPA) module, which effectively aligns cross-domain features in a class-aware manner by bridging the gap between the object detection task and the domain adaptation task. Then, the designed Dataset-level Alignment Scheme (DAS) explicitly guides the detector to achieve global representation and enhance inter-class distinguishability of instance-level features across the entire dataset, which spans both domains, by leveraging contrastive learning. Moreover, DATR incorporates a mean-teacher-based self-training framework, utilizing pseudo-labels generated by the teacher model to further mitigate domain bias. Extensive experimental results demonstrate superior performance and generalization capabilities of our proposed DATR in multiple domain adaptation scenarios. Code is released at https://github.com/h751410234/DATR. Liang Chen 0004, Jianhong Han, Yupei Wang |
IEEE Trans. Image Process. | 1 |
| 2025 | DSENet++: A Coarse-to-Fine Framework for Enhanced Sub-Region Detection in Aerial Images
Xiangjie Wang, Liang Chen 0004, Junjie Zhang 0002, Jian Zhang 0002, Shiming Ge, Dan Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Diffusion-Geo: A Two-Stage Controllable Text-To-Image Generative Model for Remote Sensing ScenariosabstractImage generation is a crucial task to facilitate intelligent interpretation in remote sensing domain. Expanding dataset size through image generation can enhance model performance of downtown task. However, current generative models in remote sensing are mostly unconditional or guided by simple text, resulting in generated images lacking spatial and semantic constraints. This lack of control can negatively optimize downstream task models. To tackle these challenges, a two-stage controllable text-image generative model called Diffusion-Geo is presented. In the first stage, an extensive image-text generation dataset called RS-Control is created through prompt engineering of multimodal large language models (MLLMs) and manual prompts for existing datasets, incorporates diverse conditional controls with rich spatial and semantic information. Then RS-Control dataset is utilized to train a universal controllable image generative model. The second stage involves efficient tuning the universal model for different task datasets, minimizing fine-tuning costs while preserving diversity and high-quality features. Experiments conducted on the RSICD caption dataset and WHU change detection dataset demonstrate the superiority of Diffusion-Geo over other state-of-the-art models in image generation. Miaoxin Cai, Wei Zhang 0389, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Liang Chen 0004, Can Li 0005 |
IGARSS | 6 |
| 2024 | Regression-Guided Positive Sample Refocusing Paradigm for Tiny Object Detection in Aerial ImagesabstractTiny object detection represents a pivotal challenge in remote sensing intelligent interpretation, necessitating detectors to exhibit heightened precision in object localization. However, typical model optimization strategies cannot release the detector’s potential for precisely localizing objects. And the lack of interpretability in detection box filtering based on object classification scores serves as a constraint on further performance improvement. Therefore, this paper proposed a novel model optimization strategy to thoroughly unleash the potential of the detector for precise localization. Then, the utilization of object comprehensive confidence score enhances the interpretability of the post-processing step for detection boxes. Rigorous experiments on the AI-TOD dataset have demonstrated the effectiveness of our method, achieving state-of-the-art performance. Lihui Ge, He Chen 0004, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, Fukun Bi, Liang Chen 0004 |
IGARSS | 7 |
| 2024 | Uncertainty-Injected Cross-Domain Few-Shot Scene Classification From Remote Sensing ImageryabstractCross-domain few-shot scene classification (CDFSSC) is crucial for remote sensing (RS) applications since it aims at transferring knowledge learned from the source domain to the target domain to facilitate the model’s few-shot classification for the target domain. However, existing methods ignored the feature statistic discrepancy caused by domain shifts, leading to an inferior performance on the target domain. In this paper, to facilitate the model’s adaptation of the domain shifts and achieve better cross-domain knowledge transfer, an uncertainty-injected cross-domain framework called UICD is proposed for CDFSSC tasks from RS imagery. First, a semi-supervised teacher-student structure is employed to achieve cross-domain knowledge transfer by conducting supervised learning on labeled source data and establishing consistent predictions on unlabeled target data. Secondly, uncertainty is injected in feature statistic modeling during cross-domain training to obtain more diverse feature statistics for data from both the source and target domains, which could promote the robustness and adaptation of the model to domain shifts, thus enabling the model to better adapt to unforeseen variations in the target domain. Extensive experiment results indicate the efficacy and superiority of the proposed methods. Can Li 0005, He Chen 0004, Yin Zhuang, Liang Chen 0004 |
IGARSS | 5 |
| 2024 | Frequency Spectrum Features Modeling for Real-Time Tiny Object Detection in Remote Sensing ImageabstractRecently, object detection in remote sensing images has achieved rapid advancement. However, due to critical issues, such as low spatial resolution and complex background noises, it is still difficult to achieve satisfactory object detection performance for remote sensing images. For current widely used object detection methods, the feature resolution of the backbone network is decreased gradually with successive pooling operations. In this way, object spatial details are largely lost for the deeper feature layers, resulting in the difficulty of accurate object detection, especially for tiny objects. However, current methods fail to eliminate the adverse effects due to the loss of object details. To this end, considering that high-frequency information is more likely to be overlooked and high-frequency object details may be beneficial for detecting tiny objects, we propose to improve the previous spatial feature modeling pipeline with the learned features in the frequency domain. Specifically, discrete cosine transform (DCT) is first used to transform the original image into the frequency domain, obtaining the corresponding frequency spectrum features. We then utilize a dual-domain feature extraction (DFE) network based on a lightweight attention mechanism to align the features in two different domains. Finally, a domain synergy fusion (DSF) module is further employed to match and fuse the features in the spatial domain and the obtained features in the frequency domain. Extensive experimental results are obtained on the challenging remote sensing datasets, DIOR and DOTA. Experimental results show that our method can increase at least 2.9%${\mathbf {AP}}_{\mathbf {s}}^{\mathbf {50}}$in DIOR and 3.5%${\mathbf {AP}}_{\mathbf {s}}^{\mathbf {50}}$in DOTA compared to the new state-of-the-art methods, which effectively demonstrates the superiority of our proposed method. Zhaoyi Luo, Yupei Wang, Liang Chen 0004, Wenying Yang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Axial-Shift Feature Interaction and Prototype-Guided Penalty Constraint for Remote Sensing Change DetectionabstractAt present, deep learning (DL) methods for remote sensing (RS) change detection (CD) are developing rapidly. However, there are still many challenges in complex RS scenarios. Factors, such as season and illumination, contribute to minimal differences in radiation characteristics between the changed area and the background, making them difficult to distinguish. This letter proposes an axial-shift feature interaction and prototype-guided penalty constraint network (ASPGNet) to address this problem. ASPGNet integrates axial-shift feature interaction (ASFI) module and prototype-guided penalty constraint (PGPC) loss. The ASFI module facilitates interaction among adjacent features through axial-shift operations in the width/height directions, aiming to obtain discriminative feature representations of the changed area. The PGPC loss utilizes prototypes to adaptively identify and weigh confusing pixel features, ensuring distinguishability between change and nonchange features and, thereby, generating accurate CD results. We evaluate the proposed method on the WHU-CD and LEVIR-CD datasets, achieving the$F1$scores of 93.22% and 91.49%, respectively. These results demonstrate the effectiveness of the proposed method. He Chen 0004, Jue Wang 0011, Wenchao Liu 0001, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Encouraging the Mutual Interact Between Dataset-Level and Image-Level Context for Semantic Segmentation of Remote Sensing ImageabstractRecently, semantic segmentation of remote sensing images has witnessed rapid advancement with the adoption of deep neural networks. Contextual cues, referring to the long-range correlation between pixels, are crucial for achieving accurate segmentation results, particularly for objects with less discriminative characteristics in these images. Currently, most studies are centered on incorporating contextual cues by aggregating context information at the dataset-level or image-level. However, current research often treats contextual cue modeling at the dataset-level and image-level as independent procedures, neglecting the intrinsic correlation between these two feature levels. Consequently, the obtained contextual cues are sub-optimal. This issue is particularly critical in the semantic segmentation of remote sensing images. To address this, we propose to encourage mutual interaction between dataset-level and image-level contextual cues. Firstly, we propose an interactive dataset-image context aggregation scheme to obtain complementary and consistent multi-level contextual cues. Additionally, we introduce a parallel feature interaction network that progressively extracts and fuses features across multiple layers, enabling effective integration of multi-level contexts. Furthermore, we introduce an enhanced shifted window-based cross-attention mechanism to augment model efficiency. Extensive experimental results on the widely used Vaihingen, GaoFen-2 and iSAID datasets effectively demonstrate the superiority of our proposed method over state-of-art methods. Yupei Wang, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SWDiff: Stage-Wise Hyperspectral Diffusion Model for Hyperspectral Image ClassificationabstractHyperspectral image classification (HSIC) has been a popular task in recent years. Even benefiting from the rapid development of deep neural networks (DNNs), there are still remaining intrinsic problems, including inadequate utilization of spatial-spectral information and insufficient labeled samples. The recent emergency of diffusion models (DMs) came to the fore because of their impressive refined image generation performance. DMs have been proven to not only can capture the underlying information of data through training the decoder of DMs, but also have more stable training than GANs while retaining even better performance. To better perceive and utilize spectral-spatial information while alleviating insufficient labeled samples simultaneously, we introduce the DM into HSIC from a data generation perspective. Specifically, we propose a stage-wise DM framework (SWDiff), dividing the HSIC task into three stages, including: pretrain the diffusion decoder with the hyperspectral image (HSI); generate new HSI cubes through the well-trained decoder to extra supply the original HSI set; and utilize the supplied dataset to train varied classifiers to obtain a better classification performance. Suitable pretraining could enable the decoder to acquire spatial-spectral information of the HSIs sufficiently via modeling spectral-spatial relationships across samples, leading to better utilization of spectral and spatial information of HSIs. Furthermore, the DM could provide the inference stage with spatial-spectral prior knowledge to ensure the feasibility and plausibility of the dataset complement, which could alleviate the insufficient labeled samples problem. Eventually, the classification stage will benefit from the first two stages. Liang Chen 0004, Jingfei He, Hao Shi 0006, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Regression-Guided Refocusing Learning With Feature Alignment for Remote Sensing Tiny Object DetectionabstractTiny object detection is a formidable challenge in remote sensing intelligent interpretation. Tiny objects are usually fuzzy, densely distributed and highly sensitive to positioning errors, which leads to the mainstream detector usually achieving suboptimal detection performance when facing tiny objects. To address the mismatch of mainstream detector architectures and model optimization strategies in the context of tiny object detection, this paper presents an efficient and interpretable algorithm for tiny object detection, termed the Cross-Attention based Feature Fusion Enhanced tiny object detection Network (CAF2ENet). First, the cross-attention mechanism is introduced to refine the upsampling results of deep features. This refinement improves the precision of multi-scale feature fusion. Second, a training strategy named regression-based refocusing learning is introduced. Deviating from the conventional optimization strategy, our method guides the optimizer to prioritize higher-quality detection boxes by adjusting sample weights. This adjustment significantly amplifies the detector’s potential to achieve superior detection results. Finally, the object composite confidence score is employed for the interpretable filtering of detection boxes. Extensive experiments on Tiny Object Detection in Aerial Images (AI-TOD) and object Detection in Optical Remote sensing images (DIOR) datasets are carried out, and comparison indicate that the proposed CAF2ENet can perform the remarkable performance compared to other state-of-the-art (SOTA) tiny object detection detectors, as it can reach 63.7% Average Precision (AP50) on AI-TOD and 75.4%AP50on DIOR, achieve SOTA performance. Lihui Ge, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Hao Dong 0003, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Remote Sensing Teacher: Cross-Domain Detection Transformer With Learnable Frequency-Enhanced Feature Alignment in Remote Sensing ImageryabstractUnsupervised Domain Adaptation (UDA) is critical for remote sensing object detection in real applications, aiming to address the significant performance degradation issue caused by the domain gap between source and target domain. This method achieves cross-domain alignment by leveraging the unlabeled target domain data, thus avoiding the expensive annotation cost. However, existing works mainly cope with CNN-based object detectors, which are characterized by complex adversarial learning architecture and fail to accurately align the features in remote sensing images with sparsely allocated objects and inevitable background noise. Compared to CNN-based methods, the DEtection TRansformer (DETR) largely simplifies the object detection pipeline and demonstrates the great potential by its intrinsic characteristics of global relation modeling between any pixels. On this basis, we propose the first strong DETR-based baseline, Remote Sensing Teacher, for unsupervised domain adaptation in remote sensing object detection. Specifically, the Remote Sensing Teacher introduces an innovative Learnable Frequency-enhanced feature Alignment (LFA) module. Within this module, we initially transform the features into frequency space to simplify the attention solver and effectively capture domain-specific information. Subsequently, the module significantly enhances the global feature representations of sparsely allocated objects by using a lightweight attention mechanism. Following this, the module incorporates learnable filters with a gated mechanism, enabling selective alignment of features in noisy backgrounds. Additionally, the Remote Sensing Teacher employs a Self-adaptive Pseudo label Assigner (SPA) that can automatically adjust the class-wise confidence threshold according to the model’s learning status, thereby enabling the generation of high-quality pseudo-labels in scenarios with a long-tailed distribution. Leveraging these pseudo-labels further mitigates the domain bias of the detector by establishing alignment at the label level. Extensive experimental results demonstrate superior performance and generalization capabilities of our proposed Remote Sensing Teacher in multiple remote sensing adaptation scenarios. Code is released at https://github.com/h751410234/RemoteSensingTeacher. Jianhong Han, Yupei Wang, Liang Chen 0004, Zhaoyi Luo |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Few-Shot Fine-Grained Classification With Rotation-Invariant Feature Map Complementary Reconstruction NetworkabstractFine-grained classification is of significant importance in the field of remote sensing. However, obtaining valuable and rare target images is often a challenging task, giving rise to the few-shot fine-grained classification problem. In response to this challenge, various meta-learning approaches have been introduced, with the feature map reconstruction network emerging as a prominent method. Targets in remote sensing images exhibit arbitrary orientation, substantial inter-class similarity and intra-class diversity. Nevertheless, the conventional feature map reconstruction network exhibits subpar performance due to its inability to handle rotational variations. Moreover, it only reconstructs features from a single channel dimension of support features, neglecting the interplay between different dimensions and resulting in inaccurate reconstruction errors. To overcome the challenges of imprecise rotational variation features for reconstruction and inaccurate reconstruction errors, we propose a rotation-invariant feature map complementary reconstruction network (RIFCRN). The RIFCRN involves several key innovations. First, we introduce a novel rotation-invariant module (RIM) based on active rotating filters and oriented response pooling, enabling the extraction of rotation-invariant features for reconstruction. This modification enhances the suitability of the feature map reconstruction network for the few-shot fine-grained classification problem. Second, we put forward a novel feature map complementary reconstruction (CPR) method that calculates the complementary reconstruction errors (CRE) which effectively captures relationships among different feature map dimensions and results in more accurate reconstruction errors. Finally, extensive experiments have been conducted to validate the effectiveness of the proposed RIFCRN in addressing the few-shot fine-grained classification problem. The code will be available at https://github.com/liyangfan0/RIFCRN. Yangfan Li 0002, Liang Chen 0004, Wei Li 0032, Nan Wang 0038 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | DECOR: Dynamic Decoupling and Multiobjective Optimization for Long-Tailed Remote Sensing Image ClassificationabstractIn the realm of remote sensing, targets of interest span a range of categories. However, their distribution is not always uniform. Certain categories substantially outnumber others, resulting in what’s termed a ‘long-tailed distribution’ in remote sensing imagery. This imbalanced distribution often biases a classifier’s focus toward the more abundant (head) classes, at the detriment of the less-represented (tail) classes. Such biases undermine the classifier’s generalization performance, particularly in the context of remote sensing image classification (RSIC). While existing mitigation approaches such as resampling, reweighting, and transfer learning offer some respite, they often miss out on in-depth knowledge refinement, rendering them less effective for severe long-tailed RSIC scenarios. To counter these challenges, we introduce DECOR, a dynamic decoupling and multi-objective optimization framework. Within DECOR, the feature extractor and classifier are dynamically decoupled, promoting superior feature representation and classifier training. Then, a multi-objective optimization approach is proposed to delve deeper, refining feature representation at the knowledge level using learnable feature centroids coupled with masked world knowledge learning. Moreover, to combat the pronounced effects of sample imbalance on classifier training, we employ a class-balanced re-sampling technique paired with a parameter-efficient adapter, which sharpens the classifier’s decision boundary and bridges the gap between representation and classification. DECOR’s efficacy is validated through comprehensive experiments on several datasets, including the NWPU-RESISC45-LT (NWPU-LT), AID-LT, and our self-built BIT-AFGR50-LT. Experimental results demonstrate DECOR’s marked enhancement in performance on long-tailed datasets. Our source code is available at: https://github.com/ChloeeGrace/DECOR. Jianlin Xie, Guanqun Wang, Yin Zhuang, Can Li 0005, Tong Zhang 0028, He Chen 0004, Liang Chen 0004, Shanghang Zhang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Weakly Supervised Semantic Segmentation With Consistency-Constrained Multiclass Attention for Remote Sensing ScenesabstractObtaining image-level class labels for Remote Sensing (RS) images is a relatively straightforward process, sparking significant interest in Weakly Supervised Semantic Segmentation (WSSS). However, RS images present challenges beyond those encountered in generic WSSS, including complex backgrounds, densely distributed small objects, and considerable scale variations. To address above issues, we introduce a COnsistency-COnstrained Multi-Class Attention model, noted asCocoaNet. Specifically, CocoaNet endeavors to capture both semantic correlation and class distinctiveness using a Global-Local Adaptive Attention mechanism, which integrates the self-attention to model global correlation, complemented by a Local Perception branch that intensifies focus on local regions. The resulting class-specific attention weights and patch-level pairwise affinity weights are employed to optimize the initial CAMs. This mechanism proves highly effective in mitigating inter-class interference and managing the distribution of densely clustered small objects. Moreover, we invoke a Consistency Constraint to rectify activation inaccuracy. By utilizing a Siamese structure for the mutual supervision of features extracted from images at different scales, we address substantial scale variations in RS scenes. Simultaneously, a Class Contrast Loss is adopted to enhance the discriminativeness of class-specific features. Departing from the conventional CAM optimization, which is rather complex and time-consuming, we harness the prior knowledge from generic Segment Anything model to design a joint optimization strategy that refines target boundaries and further promotes discriminative visual features. We validate the effectiveness of our proposed approach on three benchmark datasets in multi-class RS scenarios, experimental results demonstrate that our model yield promising advancements compared to state-of-the-art methods. Junjie Zhang 0002, Yongshun Gong, Jian Zhang 0002, Liang Chen 0004, Dan Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Heterogeneous Prototype Distillation With Support-Query Correlative Guidance for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification (FSRSSC) aims to identify unseen classes only relying on very limited training samples. However, scarce training samples are insufficient to support a robust classwise representation, which is easily influenced by agnostic biases from diverse testing scenarios. Fortunately, there are abundant spatial contextual clues that exist in very limited training samples to have an enormous potential to establish discriminative and transferable concepts. Thus, in this article, a hybrid architecture called ProtoConViT is proposed to learn a powerful classwise representation based on spatial contextual clues for FSRSSC promotion. First, support-query correlative guidance is designed to generate more stable spatial connections among support and query data based on intermediate convolution neural network (CNN) feature maps, which not only can be embedded into each episodic training task to reduce redundant spatial contextual representation learning space of vision transformer (ViT) but also can assist it in rapidly capturing critical spatial contextual clues to classify query data into one of classes from support set. Second, followed by the designed support-query correlative guidance, a novel heterogeneous prototype distillation is proposed to integrate the advantages of CNN and ViT for heterogeneous prototype construction, which can rapidly set up discriminative and transferable concepts for FSRSSC. Third, corresponding to the proposed ProtoConViT, a joint loss is designed to make the model rapid convergence based on meta-learning. Finally, extensive experiments are carried out on three FSRSSC benchmarks, and comparative results indicate that the proposed ProtoConViT can achieve a superior FSRSSC performance. Yin Zhuang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Hybrid Transformer Network for Change Detection Under Self-Supervised PretrainingabstractThis paper presents a Siamese network architecture based on a multi-scale hybrid convolution-Transformer (CTUNet) for Change Detection (CD) in a pair of co-registered optical remote sensing images. Different form CD frameworks based on convolution neural networks (CNNs) and pure Transformer networks, this method combines a convolution-Transformer hybrid encoder with a multi-scale change information extraction decoder in a Siamese network architecture. It overcomes the inherent limitations of CNN and Transformer and effectively integrates the multi-scale information required for accurate CD. To learn better discriminative representations from various scales, we propose a masked auto-encoder scheme (CTMAE) to adapt to building targets with varying morphological scales, further unleashing the potential of CTUNet. Experiments on two CD datasets show that the proposed self-supervised pre-trained hybrid convolution-Transformer CTUNet architecture achieves better CD performance than previous methods. Yongjing Cui, Yin Zhuang, Shan Dong, Peng Gao 0007, He Chen 0004, Liang Chen 0004 |
IGARSS | 7 |
| 2023 | All Adder Neural Networks for On-Board Remote Sensing Scene Classification
Ning Zhang 0042, Jue Wang 0011, He Chen 0004, Wenchao Liu 0001, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Posterior Instance Injection Detector for Arbitrary-Oriented Object Detection From Optical Remote-Sensing ImageryabstractArbitrary-oriented object detection (AOOD) from optical remote sensing imagery has to correctly generate delicate oriented boundary boxes (OBBs) and meanwhile identify their specific categories. However, how to make detectors learn delicate parameters of OBBs, especially for the crucial orientation information, and identify object category from complex background becomes a challenge task. Therefore, in this article, for exploring a better way to guide the detector to learn specific category and parametric information of OBBs, a novel one-stage anchor-free detector called Posterior Instance Injection Detector (PIIDet) is proposed for AOOD. First, as the anchor-free manner lacks prior information, an object-aware posterior guidance (OAPG) structure is proposed to generate specific-category instances used for conditioning on OBB prediction. This structure can assist the proposed PIIDet in better learning the relative parametric information of OBBs corresponding to their specific categories. Besides, to guarantee a high quality injection of specific-category instances, a new hierarchical feature fusion module is developed to establish a suitable multi-scale feature mapping space. Second, considering the negative optimization of angle regression, which is caused by the boundary discontinuity of angular periods and sudden shifts of the relation between width and height in training phase, a novel binary classification embedded angle regression space (BCE-RegSpace) is devised for providing continuous angle regression space and stable relation between width and height. Finally, extensive experiments are executed on three AOOD benchmarks (e.g., DOTA, DIOR-R and HRSC2016), and results proved that the proposed concise one-stage anchor-free PIIDet can reach the state-of-the-art (SOTA) performance and meanwhile have an impressive inference speed. Tong Zhang 0028, Yin Zhuang, He Chen 0004, Guanqun Wang, Lihui Ge, Liang Chen 0004, Hao Dong 0003, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Geometric Boundary Guided Feature Fusion and Spatial-Semantic Context Aggregation for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images aims to achieve pixel-level semantic category assignment for input images. This task has achieved significant advances with the rapid development of deep neural network. Most current methods mainly focus on effectively fusing the low-level spatial details and high-level semantic cues. Other methods also propose to incorporate the boundary guidance to obtain boundary preserving segmentation. However, current methods treat the multi-level feature fusion and the boundary guidance as two separate tasks, resulting in sub-optimal solutions. Moreover, due to the large inter-class difference and small intra-class consistency within remote sensing images, current methods often fail to accurately aggregate the long-range contextual cues. These critical issues make current methods fail to achieve satisfactory segmentation predictions, which severely hinder downstream applications. To this end, we first propose a novel boundary guided multi-level feature fusion module to seamlessly incorporate the boundary guidance into the multi-level feature fusion operations. Meanwhile, in order to further enforce the boundary guidance effectively, we employ a geometric-similarity-based boundary loss function. In this way, under the explicit guidance of boundary constraint, the multi-level features are effectively combined. In addition, a channel-wise correlation guided spatial-semantic context aggregation module is presented to effectively aggregate the contextual cues. In this way, subtle but meaningful contextual cues about pixel-wise spatial context and channel-wise semantic correlation are effectively aggregated, leading to spatial-semantic context aggregation. Extensive qualitative and quantitative experimental results on ISPRS Vaihingen and GaoFen-2 datasets demonstrate the effectiveness of the proposed method. Yupei Wang, Yongkang Hu, Xiaoxing Hu, Liang Chen 0004, Shanqing Hu |
IEEE Trans. Image Process. | 5 |
| 2022 | Adaptive Local Context Embedding for Small Vehicle Detection from Aerial Optical Remote Sensing ImagesabstractSmall vehicle detection is one of the remaining challenging task because the ambiguous appearance is against complex background interference. Consequently, in order to improve the performance of small vehicle detection from aerial optical remote sensing images, a novel adaptive local context (ALC) embedding way is designed and further introduced into an anchor free detection manner which is called ALC-Net, and in ALC-Net, it can adaptively set up the effective local context feature to improve keypoint description of small vehicles and boost the detection performance without adding extra prior information. Finally, several experiments are carried out on two widely used datasets (e.g., UCAS-AOD [1] and VEDAI [2]) and the results indicate that the proposed ALC-Net can exhibit the competitive small vehicle detection performance than other detectors. Shanjunyu Liu, Yin Zhuang, Hao Dong 0003, Peng Gao 0007, Guanqun Wang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li |
IGARSS | 7 |
| 2022 | Bilateral Semantic Fusion Siamese Network for Change Detection From Multitemporal Optical Remote Sensing ImageryabstractChange detection (CD) is an essential task in optical remote sensing, and it can be used to extract the valid information from sequential multitemporal images. However, since the character of long-term revisiting and very high resolution (VHR) development, the great differences of illumination, season, and interior textures between bitemporal images bring considerable challenges for pixel-wise CD. In this letter, focusing on accurate pixel-wise CD, a bilateral semantic fusion Siamese network (BSFNet) is proposed. First, to better map bitemporal images into semantic feature domain for comparison, a novel BSFNet is designed to effectively integrate shallow and deep semantic features, which can provide pixel-wise CD results with complete regions and clear boundary locations. Then, in order to facilitate the reasonable convergence of the proposed BSFNet, a scale-invariant sample balance (SISB) loss is designed for metric learning to avoid the problems of sample imbalance and scale variance. Finally, extensive experiments are carried out on two published CDD and LEVIR CD datasets, and results indicate that the proposed BSFNet can provide superior performance than the other state-of-the-art methods. Our work is available athttps://github.com/ClarissaDHL/BSFNet. Hailin Du, Yin Zhuang, Shan Dong, Can Li 0005, He Chen 0004, Boya Zhao, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | SAR-to-Optical Image Translating Through Generate-Validate Adversarial NetworksabstractSynthetic aperture radar (SAR) has the advantages of high resolution in all-weather and all-day. However, SAR images are hard to be understood, due to their unique imaging mechanism. The SAR to optical image translation can assist in interpreting and has become a topic of growing interest in the field of remote sensing. In this letter, a SAR to optical image translation network is proposed, called generate-validate adversarial networks (GVANs). More specifically, there are two Pix2Pix networks form the cyclic structure. The validate module is employed to increase the training process and improve the edge retention ability. In order to improve multidomain images adaptability, the embedded layer is proposed. Additionally, the dilation convolution layer is employed in the generator, which is more suitable for the characteristics of SAR images. The proposed method has experimented on the SEN1-2 dataset. The result demonstrates the superiority of the proposed method over state-of-the-art methods. Hao Shi 0006, Bocheng Zhang, Yupei Wang, Zihan Cui, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Dual-Path Sparse Hierarchical Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images aims to label every pixel with the correct semantic category. The core challenge of the current deep convolutional network (ConvNet)-based methods lies in the difficulty of effectively aggregating high-level categorical semantics and low-level local details along the hierarchy of backbone. Most current approaches consider only fusing adjacent feature layers gradually with short-range feature connections, which lack the diversity of feature interactions, such as long-range cross-scale connections. To this end, we propose a novel dual-path sparse hierarchical network that is characterized by rich cross-scale feature interactions. Multiscale features are first sparsely grouped with a predefined interval, which is then aggregated via both long-range and short-range cross-scale connections in a hierarchical manner. Moreover, in order to further enrich the diversity of feature interactions, we also introduce another fusion path in parallel but with different sparsity for feature grouping, forming a dual-path network. In this way, our model is able to effectively aggregate multilevel features by incorporating both long-range and short-range feature interactions in both parallel and hierarchical manner. Meanwhile, the semantic and resolution gap between multilevel features can also be bridged. Yupei Wang, Hao Shi 0006, Shan Dong, Yin Zhuang, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Sphere Loss: Learning Discriminative Features for Scene Classification in a Hyperspherical Feature SpaceabstractThe power of features considerably influences the classification performance of remote sensing scene classification (RSSC). Recently, deep convolutional neural networks (DCNNs) have been used to extract powerful scene features. Nevertheless, confusion and overlap still occur in the feature space, leading to inaccurate RSSC. To alleviate this problem, we propose a novel deep metric learning loss function incorporated into a sphere loss to enhance the discrimination of feature representations. Inspired by two representative loss functions (i.e., angular loss and center loss), the proposed sphere loss learns a unique cluster center for each class in a remote sensing scene. Because the cluster centers and features are restricted by an introduced geometrical constraint, the intraclass distance of features decreases, while the interclass distance increases. Moreover, we introduce a spatial constraint, i.e., a uniformity coefficient on different cluster centers, which causes the centers to form a uniform distribution that maximizes the interclass distances between features. Extensive analysis and experiments on three commonly used RSSC data sets consistently show that, compared with state-of-the-art methods, the proposed sphere loss can effectively learn discriminative feature representations and significantly improve RSSC. Jue Wang 0011, He Chen 0004, Long Ma 0003, Liang Chen 0004, Xiaodong Gong, Wenchao Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Feature Enhanced Centernet for Object Detection in Remote Sensing ImagesabstractMulti-scale object detection in optical remote sensing imagery is a challenging task due to the varied object scales. Existed state-of-art object detection methods have achieved significant growth. However, most of the methods are based on default anchors, which need to be predefined. The multi-scale object detection accuracy still needs to be improved, especially for small and dense objects. To improve the robustness of the detection algorithm and the performance of multi-scale object detection, a novel anchor-free multi-scale object detection method Feature Enhanced CenterNet is proposed in this paper. First, we use the “encoder-decoder” structure and introduce horizontal connections to enhance feature representation capabilities. Second, an context-aware up-sampling method is proposed to obtain feature maps with suitable scale. To demonstrate the performance of the proposed method, we perform abundant experiments on the public remote sensing datasets. The experimental results demonstrate the robustness and effectiveness of the proposed method. Tong Zhang 0028, Guanqun Wang, Yin Zhuang, He Chen 0004, Hao Shi 0006, Liang Chen 0004 |
IGARSS | 6 |
| 2019 | Spatial Enhanced-SSD For Multiclass Object Detection in Remote Sensing ImagesabstractAccurate multiclass object detection in remote sensing images is a challenging task, especially for small objects. Since the scales of objects in remote sensing images have a great variance, almost all of the advanced detection methods have shortcomings. Consequently, improving the accuracy of multiclass objects detection has always been the direction of researchers' efforts. In this paper, a spatial enhanced-Single Shot MultiBox Detector (SE-SSD) is proposed. First, to enhance the spatial information, we enlarge the input image channels with embedding oriented-gradients feature maps. Second, the multiple output layers in the backbone network are changed to reduce one pooling operation. Finally, we design a context module to enhance the receptive field for feature layer description in SE-SSD framework. Experimental results on DOTA dataset demonstrate that Spatial Enhanced-SSD method reaches a much higher mean average precision (mAP) than Faster R-CNN, SSD and other classic detection network. Guanqun Wang, Yin Zhuang, Zhiru Wang, He Chen 0004, Hao Shi 0006, Liang Chen 0004 |
IGARSS | 6 |
| 2018 | A Novel Harbor Detection Method Based on Pattern Coding AlgorithmabstractHarbor automatic detection is a scene interpretation in remote sensing image processing. Fast and accurate harbor detection can significantly improve the performance of inshore ship detection. In order to achieve harbor detection in complex remote sensing images, in this paper, a novel pattern coding algorithm is proposed. The proposed harbor detection method has three steps: First, the harbor water area is extracted by using the definition circle (DC) model. Secondly, pre-generated eight KEY patterns are multiplied with the local scenes, and the probability density function (PDF) of the multiplied local scene is recorded. Finally, the Euclidean distance between the eight patterns' PDFs and the original local scene's PDF is calculated, then compared with the threshold and coded, so as to realize the harbor area detection. Experimental results demonstrate that the novel method has outstanding performance on harbor area detection in complex broad width remote sensing images. Guanqun Wang, Yin Zhuang, He Chen 0004, Liang Chen 0004 |
IGARSS | 4 |
| 2018 | Comprehensive Structure Voting Docked Ship Detection from High-Resolution Optical Satellite Images Based on Combined Multi-Orientation Sparse RepresentationabstractInshore ship detection from high-resolution (HR) optical satellite images is a hot research field. However, HR ships multi-scale and multi-orientation characters and harbor scene various interferences affect docked ship detection performance. Therefore, we proposed a multi-orientations sparse dictionaries (MOSDs) algorithm combining with comprehensive structure voting (CSV) to address existed problem and achieve refined docked ship contour region proposal (RP). Moreover, the comparing experiments use a lot of Google Earth harbour images to demonstrate proposed method effectiveness and robustness of HR ships multi-scale and -orientation changing and various harbour background interferences of docked ship detection. Yin Zhuang, He Chen 0004, Liang Chen 0004, Fukun Bi |
IGARSS | 4 |
| 2018 | IORN: An Effective Remote Sensing Image Scene Classification FrameworkabstractIn recent times, many efforts have been made to improve remote sensing image scene classification, especially using popular deep convolutional neural networks. However, most of these methods do not consider the specific scene orientation of the remote sensing images. In this letter, we propose the improved oriented response network (IORN), which is based on the ORN, to handle the orientation problem in remote sensing image scene classification. We propose average active rotating filters (A-ARFs) in the IORN. While IORNs are being trained, A-ARFs are updated by a method that is different from the ARFs of the ORN, without additional computations. This change helps IORN improve its ability to encode orientation information and speeds up optimization during training. We also propose Squeeze-ORAlign (S-ORAlign) by adding a squeeze layer to ORAlign of ORN. With the squeeze layer, S-ORAlign can address large-scale images, unlike ORAlign. An ablation study and comparison experiments are designed on a public remote sensing image scene classification data set. The experimental results demonstrate the effectiveness and better performance of the proposed model over that of other state-of-the-art models. Jue Wang 0011, Wenchao Liu 0001, Long Ma 0003, He Chen 0004, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | A Modified Fixed-Point Chirp Scaling Algorithm Based on Updating Phase Factors Regionally for Spaceborne SAR Real-Time ImagingabstractTo realize real-time imaging for the spaceborne synthetic aperture radar (SAR) using the chirp scaling (CS) algorithm, high generation rates of phase factors and heavy computation loads are unavoidable. For example, updating phase factors continuously in the 2-D time or frequency space is difficult for generating phase factors for real-time imaging. To solve this problem, a modified CS algorithm based on updating phase factors regionally is proposed. Moreover, conventional floating-point arithmetic is unaffordable for real-time imaging onboard. Therefore, fixed-point processing is adopted. In the proposed imaging algorithm, the phase factors are updated only at some specified times and frequencies, and fixed-point operation is adopted. Furthermore, based on the paired echo theory and the finite word length error mathematical model, the effect of the proposed imaging algorithm on the SAR image quality is studied. Moreover, based on the above analysis and the experiments of Chinese HJ-1C and GF-3 spaceborne SAR practical measured data, a hardware system including digital signal processor and field-programmable gate array boards is constructed. Zegang Ding, Yizhuang Xie, Wenyue Yu, Liang Chen 0004, Teng Long 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | A unified reconfigurable floating-point arithmetic architecture based on CORDIC algorithmabstractThis paper presents the design methodology and implementation of reconfigurable coordinate rotation digital computer (CORDIC) architecture that can be configured to operate in different modes and rotations to achieve singleprecision floating point division, multiplication and square-root operations. Through introducing pre- and post-processing, the float-point operations can be integrated into a unified CORDIC iteration procedure. According to the characteristics of different operations, we propose a pipeline-parallel mixed architecture to optimize the area-delay-efficiency. Finally, the prototype based on Xilinx XC7VX690T has been established to test the performance of the proposed design. The result shows the related error with arithmetic computation is less than 10−6, and the resource-consumption of the proposed design is less than the sum of existing IP cores. Linlin Fang, Yizhuang Xie, He Chen 0004, Liang Chen 0004 |
FPT | 5 |
| 2017 | A novel method of speckle reduction and enhancement for SAR imageabstractA new speckle reduction algorithm is presented in this paper to improve the visualization of synthetic aperture radar (SAR) images. The algorithm contains two steps. First, we propose a speckle reduction filter. This filter use different smoothing effect according to scene of sub-window. So it can maintain the details while noise suppression. Second, we propose a method to enhance the details. We select a neighborhood of pixel and divide it into homogeneous and heterogeneous. Then we use different reconstruction strategy to enhance the texture. After these steps, the speckle noise of SAR images has been decreased and the texture details have been enhanced at the same time. Experiments validate the algorithm with good performance. Hao Shi 0006, Liang Chen 0004, Yin Zhuang, Jian Yang 0011 |
IGARSS | 2 |
| 2017 | Pyramid integral image reconstruction algorithm for infrared remote sensing sea-land segmentationabstractThe middle wave infrared remote (MWIR) images has complex scene information, low contrast ratios, and bipolar problems. To solve these problems, we propose a method that uses pyramid integral image reconstruction algorithm achieving sea-land automation segmentation. First, we calculate a gradient feature map (GFM), which extracts the structural information from an MWIR scene. Then, the GFM uses for sum are table (SAT) generation. The pyramid integral image reconstruction technology uses different scale factor reconstruct MWIR images by using SAT. Then the adaptive threshold method is employed for the sea-land segmentation on the multi-scale integral reconstruction images. Finally, we get sea-land refine segmentation result of MWIR images by synthesis analysis the multi-scale reconstruction images. By using GFM and pyramid integral image reconstruction operation, are avoid with the complex gray scene information. The integral image reconstruction can enhance the structure information and improve the reconstruction image contrast. This paper proposed method is from structure and texture information view point for sea-land segmentation, so the bipolar problem is solve in our method for MWIR images sea-land segmentation. Penglin Wang, Yin Zhuang, He Chen 0004, Liang Chen 0004, Hao Shi 0006, Fukun Bi |
IGARSS | 4 |
| 2017 | An Intensity-Space Domain CFAR Method for Ship Detection in HR SAR ImagesabstractSynthetic aperture radar (SAR) is an indispensable and extensively used sensor in ship detection. As high-resolution SAR introduces more spatial details into images, this letter proposes an intensity-space (IS) domain constant false alarm rate (CFAR) ship detector to make good use of this information. The method fuses intensity of each pixel and correlations between pixels into one characteristic, i.e., IS index. All the detection procedures center on the calculation and analysis of IS index. First, a new transform maps an image into a new IS domain. Structures like ships and wakes are enhanced in IS domain. Second, a CFAR detector picks up high IS index pixels. Third, a chain of target features is checked to screen out false candidate target pixels. Also, enhanced wakes are taken to improve detection results. Experiments on real SAR images validate that the proposed transform does enhance these structures and the whole algorithm is of good performance, especially in the case of low-contrast targets. Chonglei Wang, Fukun Bi, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Joint Amplitude-Phase Compensation for Ionospheric Scintillation in GEO SAR ImagingabstractThe ionospheric scintillation induced by local ionospheric plasma anomalies could lead to significant degradation for geosynchronous earth orbit synthetic aperture radar (SAR) imaging. As radar signals pass through the ionosphere with locally variational plasma density, the signal amplitude and phase fluctuations are induced, which principally affect the azimuthal pulse response function. In this paper, the compensation of signal amplitude and phase fluctuations is studied. First, space-variance problem of scintillation is addressed by image segmentation. Then, SPECAN imaging algorithm is adopted for each image segment, because it is computationally efficient for small imaging scene. Furthermore, an iterative algorithm based on entropy minimum is derived to jointly compensate the signal amplitude and phase fluctuations. Finally, a real SAR scene simulation is used to validate our proposed method, where both the simulated scintillation using phase screen technique and the real GPS-derived scintillation data are adopted to degrade the imaging quality. Rui Wang 0018, Cheng Hu 0001, Yuanhao Li 0001, Stephen E. Hobbs, Weiming Tian, Xichao Dong, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2016 | Feature-Area Optimization: A Novel SAR Image Registration MethodabstractThis letter proposes a synthetic aperture radar (SAR) image registration method named feature-area optimization (FAO). First, the traditional area-based optimization model is reconstructed and decomposed into three key but uncertain factors: initialization, slice set, and regularization. Next, structural features are extracted by scale-invariant feature transform (SIFT) in dual-resolution space (SIFT-DRS), a novel SIFT-like method dedicated to FAO. Then, the three key factors are determined based on these features. Finally, solving the factor-determined optimization model can get the registration result. A series of experiments demonstrate that the proposed method can register multitemporal SAR images accurately and efficiently. Fukun Bi, Liang Chen 0004, Hao Shi 0006 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Capturing and tracking of building area based on structure saliency in airborne remote sensing video
Fukun Bi, Liang Chen 0004, He Chen 0004 |
Sci. China Inf. Sci. | 3 |
| 2015 | A waterborne salient ship detection method on SAR imagery
Long Ma 0003, Liang Chen 0004, He Chen 0004, Nouman Qadeer Soomro |
Sci. China Inf. Sci. | 2 |
| 2015 | Accurate Urban Area Detection in Remote Sensing ImagesabstractAutomatic urban area detection in remote sensing images is an important application in the field of earth observation. Most of the existing methods employ feature classifiers and thereby contain a data training process. Moreover, some methods cannot detect urban areas in complex scenes accurately. This letter proposes an automatic urban area detection method that uses multiple features that have different resolutions. First, a downsampled low-resolution image is used to segment the candidate area. After the corner points of the urban area are extracted, a weighted Gaussian voting matrix technique is employed to integrate the corner points into the candidate area. Then, the edge features and homogeneous region are extracted by using the original high-resolution image. Using these results as the input, the processes of guided filtering and contrast enhancement can finally detect accurately the urban areas. This method combines multiple features, such as corner, edge, and regional characteristics, to detect the urban areas. The experimental results show that the proposed method has better detection accuracy for urban areas than the existing algorithms. Hao Shi 0006, Liang Chen 0004, Fukun Bi, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | A novel autofocusing technique based on PGA for the polarimetric SAR applicationabstractThe phase gradient autofocus (PGA) algorithm is widely used for synthetic aperture radar (SAR) autofocusing. In practice, polarizations of electromagnetic waves result in phase error estimates with different precisions obtained by PGA, yielding distinctly focused images for different polarization channels after compensating the phase errors. In this paper, a new autofocusing technique for the polarimetric application is proposed. The algorithm employs the redundancy of phase error information among different polarization channels. The phase error estimate obtained by different polarization channels which optimizes the image quality with the highest image contrast is treated as the final estimation result. In comparison with the standard PGA, the performance of the proposed approach is demonstrated using real data from airborne SAR. Zegang Ding, Teng Long 0001, Liang Chen 0004 |
IGARSS | 4 |
| 2012 | Synthesizing high resolution profile based on correlation coefficient for stepped-frequency radarabstractThis paper mainly focuses on synthesizing high resolution profile for stepped-frequency radar signal. A novel method based on correlation coefficient is presented. Compared to traditional methods, the limitation of the frequency step size is relaxed. It means that the same range resolution can be achieved with less pulse number, thus reducing the complexity of the system design. Finally, the simulation data is used to demonstrate the performance of this proposed method. Rui Wang 0018, Liang Chen 0004, Tao Zeng 0001 |
IGARSS | 2 |