VLDB 2026 Research / reviewers in the wild / expert
He Chen 0004
dblp:54/2335-4
· DBLP profile ↗
68ranked-venue papers
2as first author
39since 2021 · last 2026
0000-0003-4182-6493ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 63 · 2 first-author · 37 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Connected Subgraph-Based Heuristic Conflict-Free Association for Multi-DroneabstractMulti-drone multi-target association suffers from cognitive conflicts in multi-perspective scenarios due to chain rule aggregation. This letter defines the Conflict-Free Association (CFA) task and proposes Connected Subgraph-Based Heuristic Association (CSHA) with a plug-and-play Conflict Resolution Module (CRM). By mapping associations to graphs and searching for conflict-free subgraphs via a greedy algorithm, CSHA resolves conflicts effectively. Experiments on Rflysim and real-world data validate its superiority. Xuqi Yang, Ning Zhang 0042, He Chen 0004 |
IEEE Signal Process. Lett. | 4 |
| 2025 | High-Throughput and Energy-Efficient FPGA-Based Accelerator for All Adder Neural NetworksabstractNeural networks have been extensively applied across various Internet of Things (IoT) applications, such as drone- and satellite-based remote sensing and autonomous driving. With the increasing resolution and amount of data captured by sensors, the demand for real-time response in IoT applications is markedly increasing. However, it is difficult for existing convolutional neural network (CNN) accelerators for IoT applications on field-programmable gate array (FPGA) platforms to achieve high throughput because of the inherent dense multiplication operations of CNNs, memory bandwidth limitations and inefficient mapping mechanisms. In this article, a high-throughput and energy-efficient all adder neural network (A2NN) accelerator for IoT applications on FPGA platform is proposed to solve this problem. First, a series of hardware-oriented algorithm optimization methods are proposed to simplify the processing flow of A2NN and further minimize its deployment overhead. Second, a novel hardware architecture based on the idea of near-memory computation (NMC) is proposed to eliminate off-chip memory access completely and accelerate the reconstructed A2NN in the pipeline. Third, a set of quantitative analysis methods for the proposed accelerator is presented to balance throughput and energy consumption, allowing the accelerator to adapt to the varying demands of different IoT application scenarios. Extensive experimental results on the AMD-Xilinx VC709 board demonstrate that the proposed accelerator achieves state-of-the-art performance in terms of throughput, energy efficiency, and throughput efficiency. Moreover, experiments on the AMD-Xilinx KV260 board highlight the architecture’s exceptional scalability and energy efficiency, enabling a balance between speed and power consumption tailored to the specific requirements of IoT application scenarios. Ning Zhang 0042, Shuo Ni, Liang Chen 0004, He Chen 0004 |
IEEE Internet Things J. | 5 |
| 2025 | Shape Activated CAM Learning for Weakly Supervised Remote Sensing Semantic SegmentationabstractClass activation map (CAM) based weakly-supervised semantic segmentation (WSSS) of remote sensing (RS) images has attracted extensive research interests for its potential in reducing annotation cost. However, challenged by unconstrained activation issue, existing methods struggle to delineate object boundaries clearly, making them particularly difficult to separate multiple densely packed objects, which are common in RS images. By conducting an in-depth analysis of RS image characteristics, we observed a strong correlation between object shapes and their semantics. Inspired by this finding, we propose an Intrinsic Shape Activation Network (ISANet) to learn the category-relevant shape priors as geometry constraints for target-focused region activation in WSSS of RS images. The key idea is to distill the intrinsic shape priors from the hybrid features that are deterministic in classification. Specifically, we adopt a dual-branch architecture to decouple the learning of shape and texture features and leverage a shape awareness alignment module to generate boundary-clear CAMs for computing pseudo labels. In this way, CAMs are generated with perception of target shapes, which increases the completeness of activation regions and alleviates the ultrarange responses. Extensive experiments demonstrates the superiority of our method in delineating densely-packed objects with clear contours, which is especially beneficial for separating multiple targets in RS images. Our method improves the mIoU of the state-of-the-art method by 7.9% and 3.3% on the NWPU VHR-10 and iSAID dataset respectively. He Chen 0004, Mingyue Dong, Linwei Yue, Xianwei Zheng, Jun Li 0009, Jianya Gong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | HPN-CR: Heterogeneous Parallel Network for SAR-Optical Data Fusion Cloud RemovalabstractSynthetic aperture radar (SAR)-optical data fusion cloud removal is a highly promising cloud removal technology that has attracted considerable attention. In this field, primary researches are based on deep learning, which can be divided into two categories: convolutional neural network (CNN)-based and Transformer-based. In the cases of extensive cloud coverage, CNN-based methods, with their local spatial awareness, effectively capture local structural information in SAR data, preserving clear contours of cloud-removed land covers. However, these methods struggle to capture global land cover information in optical images, often resulting in notable color discrepancies between recovered and cloud-free regions. Conversely, Transformer-based methods, with their global modeling capability and inherent low-pass filtering properties, excel at capturing long-range spatial correlations in optical images, thereby maintaining color consistency across cloud-removal outputs. However, these methods are less effective at capturing the fine structural details in SAR data, which can lead to blurred local contours in the final cloud-removed images. In this context, a novel framework called heterogeneous parallel network for cloud removal (HPN-CR) is proposed to achieve high-quality cloud removal under extensive cloud coverage conditions. HPN-CR employs the proposed heterogeneous encoder with its SAR-optical input approach to effectively extract and fuse the local structural information in cloudy areas from SAR images, with the spectral information about land covers in cloud-free areas from the whole optical images. In particular, it uses a ResNet network with local spatial awareness to extract SAR features. It also uses the proposed Decloudformer, which globally models multiscale spatial correlations, to extract optical features. The output features are fused by the heterogeneous encoder and then reconstructed to cloud-removal images through a pixelshuffle-based decoder. Comprehensive experiments were conducted and the experimental results demonstrated the effectiveness and superiority of the proposed method. The code is available athttps://github.com/G-pz/HPN-CR. Panzhe Gu, Wenchao Liu 0001, Shuyi Feng, Jue Wang 0011, He Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Dual-Domain Representation Modeling With Prototype Contrastive Learning for Cross-Domain Few-Shot Scene ClassificationabstractCross-domain few-shot scene classification (CDF-SSC) aims to establish cross-domain representation between source and target domains, endowing the model few-shot classification ability on the target domain. Recent studies have improved cross-domain representation learning by incorporating unlabeled target data into semi-supervised training with labeled source data. However, these methods struggle to effectively bridge the domain gap between source and target domains, and lack the discriminative feature description ability for unlabeled target domain data, leading to inferior cross-domain representation learning and affecting the few-shot performance. In this article, a dual-domain representation modeling with prototype contrastive learning (DMPC) structure is proposed to improve the robustness of cross-domain representation learning. In DMPC, first, a dual-domain Gaussian representation modeling is designed to model the feature statistics of both source and target data as multivariate Gaussian distributions rather than fixed values and enrich domain representation by random sampling new feature statistics. It helps bridge the domain gap at the feature level, and improves the model’s robustness and generalization to better address unpredictable variations in the target domain. Second, a pseudo-prototype contrastive learning branch is proposed to improve the discriminability of representation for limited unlabeled target data. By leveraging pseudo-prototypes derived from the classifier’s weights as dynamic anchors, it refines feature representation by clustering features of the same pseudo-class and separating those of different pseudo-classes, strengthening the model’s ability to capture distinct and consistent features within the target domain. Finally, the classifier is fine-tuned on few-shot tasks to adapt to specific categories of the target domain. Extensive experimental results exhibit impressive performance of DMPC on 12 RS cross-domain scenarios. Can Li 0005, He Chen 0004, Jianlin Xie, Yin Zhuang, Liang Chen 0004, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Resolution-Difference Embedded Network for Cross-Resolution Remote Sensing Image Change DetectionabstractAt present, most remote sensing image change detection methods are applicable to equal-resolution scenarios, that is, the bi-temporal images are assumed to have the same spatial resolution. Real-world tasks such as disaster emergency response have put forward the need for change detection of bi-temporal images with different spatial resolutions, that is, cross-resolution change detection. However, due to the significant differences in spatial details between high-resolution images and low-resolution images, it is difficult to extract the spatial features of changed landcovers in complex scenarios and distinguish between changed landcovers and unchanged landcovers. To overcome the above issues, we propose a Resolution-Difference Embedding Network (RDENet). RDENet combines two innovative methods: pseudo-continuous resolution sequence representation (PRSR) and bi-temporal resolution-difference modulation (BRDM). The PRSR method effectively extracts features of landcovers in complex scenarios by constructing pseudo image sequence with smooth transition of resolution and developing a resolution-guided feature fusion module. The BRDM method significantly enhances the model’s ability to distinguish between changed and unchanged landcovers in complex scenarios through the design of a resolution-aware self-modulation module and resolution-aware mutual-modulation module, which utilizes the resolution-difference factor as prior knowledge to dynamically enhance key features of changed landcovers in cross-resolution image pairs. Extensive experiments conducted on three publicly available change detection datasets demonstrate that the proposed RDENet achieves superior detection performance in cross-resolution scenarios. He Chen 0004, Tingting Qiao, Jue Wang 0011, Wenchao Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | High-Throughput Energy-Efficient Accelerator With Collaborative-Trainable Sparse-Quantization Method for On-Board Remote Sensing ProcessingabstractConvolutional Neural Networks (CNNs) have achieved remarkable breakthroughs on remote sensing tasks in recent years. However, deploying CNNs for real-time remote sensing on-board processing still remains a challenge due to power consumption, real-time and other limitations. Therefore, in this article, a satellite-based real-time remote sensing accelerator is proposed, where algorithm and hardware approaches are proposed to jointly optimize CNNs’ deployment on edge-side aerospace devices. Firstly, a collaborative-trainable sparse-quantization (CTSQ) method is proposed to reduce the model’s storage overhead. In the CTSQ method, analysis of the errors is performed for the sparsity-quantization composition. Besides, the inter-channel correlations among parameters are leveraged, where the structured sparsity and quantization are performed with fine-grained units. Secondly, a modular-system co-optimized (MoSyC) architecture is proposed. A hardware-mapped sparse access (HMSA) strategy is proposed to effectively filter out zero elements in sparse parameters. Moreover, a high-throughput architecture is designed for parallel and pipelined data flow control. Finally, extensive experiments are conducted on both scene classification and object detection tasks with ResNet and YOLOv5 models. The results show that the proposed CTSQ method achieves the compression ratio of more than 13.81 times, and the proposed MoSyC architecture achieves the throughput of more than 1815 GOPS, demonstrating the effectiveness of the proposed accelerator. He Chen 0004, Ning Zhang 0042, Shuo Ni, Xi Zhang 0028, Liang Chen 0004, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Retain and Enhance Modality-Specific Information for Multimodal Remote Sensing Image Land Use/Land Cover ClassificationabstractMultimodal remote sensing (RS) image land use/land cover (LULC) classification using optical and synthetic aperture radar (SAR) images has raised attention for recent studies. Current methods primarily employ multimodal fusion operations to directly explore relationships between multimodal features and obtain fused features, leading to the loss of beneficial modality-specific information problem. To solve this problem, this study introduces a multimodal feature decomposition and fusion (MDF) approach combined with a visual state space (VSS) block, namely MDF-VSS block. The MDF-VSS block emphasizes beneficial modality-specific information and perceives shared land cover information through modality-difference and modality-share features, which are then adaptively integrated to obtain discriminative fused features. Based on the MDF-VSS block, an MDF decoder is designed to retain beneficial multi-scale modality-specific information. Then, a multimodal specific information enhancement (MSIE) decoder is designed to perform modality-difference feature guided auxiliary classification tasks, further enhancing modality-specific information that is expert in classification. Combining the MDF and MSIE decoders, a novel retain-enhance fusion network (REF-Net) is proposed to retain and enhance modality-specific information that benefits classification, thus improving the performance of multimodal RS image LULC classification. Extensive experimental results obtained on three public datasets demonstrate the effectiveness of the proposed REF-Net. The source code will be available at https://github.com/TINYWAI/REF_Net. He Chen 0004, Wenchao Liu 0001, Liang Chen 0004, Jue Wang 0011 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | LLaMA-Unidetector: An LLaMA-Based Universal Framework for Open-Vocabulary Object Detection in Remote Sensing ImageryabstractObject detection is a crucial task in computer vision for remote sensing applications. However, the reliance of traditional methods on predefined and trained object categories limits their applicability in open-world scenarios. A key challenge in open-vocabulary object detection lies in accurately identifying unseen objects. Existing approaches often focus solely on detecting object locations, struggling to recognize the categories of previously unseen targets. To address this issue, we propose a novel benchmark where models are trained on known base classes and evaluated on their performance in detecting and recognizing unseen or novel classes. To this end, we introduce llama-Unidetector, a universal framework that incorporates textual information into a closed-set detector, enabling the generalization to open-set scenarios. Our llama-Unidetector leverages a decoupled learning strategy that separates localization and recognition. In the first stage, a class-agnostic detector identifies objects, distinguishing only between foreground and background. In the second stage, the detected foreground objects are passed through TerraOV-LLM, a multimodal large language model, for recognition, utilizing the strong generalization capabilities of large language models to infer the correct categories. We propose a self-built Vision Question Answering (VQA) remote sensing dataset, TerraVQA, and conduct extensive experiments on the NWPU-VHR10, DOTA1.0, and DIOR datasets. The llama-Unidetector achieves impressive results, with a performance of 75.46% AP, 50.22% AP and 51.38% AP on the zero-shot detection benchmarks for the NWPU-VHR10, DOTA1.0 and DIOR datasets, respectively. Our source code is available at: https://github.com/ChloeeGrace/LLaMA-Unidetector. Jianlin Xie, Guanqun Wang, Tong Zhang 0028, Yikang Sun, He Chen 0004, Yin Zhuang, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | EarthGPT-X: A Spatial MLLM for Multilevel Multisource Remote Sensing Imagery Understanding With Visual PromptingabstractRecent advances in natural-domain multi-modal large language models (MLLMs) have demonstrated effective spatial reasoning through visual and textual prompting. However, their direct transfer to remote sensing (RS) is hindered by heterogeneous sensing physics, diverse modalities, and unique spatial scales. Existing RS MLLMs are mainly limited to optical imagery and plain language interaction, preventing flexible and scalable real-world applications. In this article, EarthGPT-X is proposed, the first flexible spatial MLLM that unifies multi-source RS imagery comprehension and accomplishes both coarse-grained and fine-grained visual tasks under diverse visual prompts in a single framework. Distinct from prior models, EarthGPT-X introduces: 1) a dual-prompt mechanism combining text instructions with various visual prompts (i.e., point, box, and free-form) to mimic the versatility of referring in human life; 2) a comprehensive multi-source multi-level prompting dataset, the model advances beyond holistic image understanding to support hierarchical spatial reasoning, including scene-level understanding and fine-grained object attributes and relational analysis; 3) a cross-domain one-stage fusion training strategy, enabling efficient and consistent alignment across modalities and tasks. Extensive experiments demonstrate that EarthGPT-X substantially outperforms prior natural and RS MLLMs, establishing the first framework capable of multi-source, multi-task, and multi-level interpretation using visual prompting in RS scenarios. The code and dataset are available athttps://github.com/wivizhang/EarthGPT-X. Wei Zhang 0389, Miaoxin Cai, Yaqian Ning, Tong Zhang 0028, Yin Zhuang, Shijian Lu, He Chen 0004, Jun Li 0009, Xuerui Mao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | A Unified Remote Sensing Object Detector Based on Fourier Contour Parametric LearningabstractA unified object detector needs to integrate various abilities for adapting to different remote sensing object detection tasks. However, there is a lack of a feasible way to integrate multigrained object detection requirements i.e., horizontal bounding box (HBB), oriented bounding box (OBB), and instance segmentation (InSeg) into a unified detection way. Then, it often has to design specific parametric learning ways and their corresponding architectures, which cannot be finely adaptive to various kinds of object detection tasks. Therefore, in this article, a new benchmark is set up to integrate multigrained object detection requirements of HBB, OBB, and InSeg into one challenging task of arbitrary-shaped object contour detection. At the same time, a unified object contour detector (UniconDet) is proposed for achieving multigrained object detection from complicated remote sensing scenes. First, a Fourier contour parametric modeling (FCPM) is defined to project arbitrary-shaped object contours from the spatial domain into the frequency domain. Then, it can unify spatial parametric representations of HBB, OBB, and InSeg as frequency coefficient representations, which can be used for realizing a more generic and robust parametric regression. Second, a multiview cross-attention (MVCA) feature extraction way is designed at each scale of the regression layer, which can assist UniconDet in perceiving Fourier contour parameters by exploring the coupled relations between different discrete contour sampling periods of each object. Third, a center-contour enhancing regression layer (C2-ERL) is designed to generate regional guidance and cascade contour propagation, which can ensure a more accurate center point prediction and Fourier contour parameter regression. Finally, extensive experiments are carried out on benchmarks of HBB, OBB, InSeg, and new multigrained object detection, and the results indicate that our proposed UniconDet can obtain superior performance. The source code is available athttps://github.com/ZhAnGToNG1/UniconDet. Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, LianLin Li, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Controllable Generative Knowledge-Driven Few-Shot Object Detection From Optical Remote Sensing ImageryabstractFew-shot object detection (FSOD) has to learn classification and localization information for unseen object detection under very low-data resource regimes. However, when deficient samples are adopted for model training, it is hard to build powerful location-aware and identification abilities for well coping with agnostic bias from diverse testing scenarios; at the same time, the overfitting phenomenon is easily occurring. Therefore, in this article, a controllable generative knowledge-driven FSOD called CGK-FSOD is proposed for unseen object detection from optical remote sensing imagery. Specifically, to enrich the learnable data space of scarce samples for preventing incomplete agnostic-bias learning, while avoiding the overfitting phenomenon, a visual-textual prompt-based controllable data generation is designed to generate high-quality object detection data based on pretrained foundational models [i.e., the stable diffusion (SD) and contrastive language-image pre-training (CLIP)], which not only can introduce the generalized domain-level knowledge into the remote sensing domain but also sets up an all-round data space to support complete learning of potential agnostic bias. Furthermore, with respect to the denoising generative process of SD, a series of cross-modality generative features in latent representation space are reused for few-shot fine-tuning by the designed cross-modality feature embedding (CMFE), which not only can bring diverse generative abilities into the feature fusion step of the detector but also gracefully sets up feature representation scalability to make the detector better adapt to agnostic bias from diverse testing scenarios of FSOD. Finally, extensive experiments are executed on two public remote sensing datasets (e.g., DIOR and NWPUVHR-10), and the results indicate that the proposed CGK-FSOD is very effective and flexible for FSOD. Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, LianLin Li, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Diffusion-Geo: A Two-Stage Controllable Text-To-Image Generative Model for Remote Sensing ScenariosabstractImage generation is a crucial task to facilitate intelligent interpretation in remote sensing domain. Expanding dataset size through image generation can enhance model performance of downtown task. However, current generative models in remote sensing are mostly unconditional or guided by simple text, resulting in generated images lacking spatial and semantic constraints. This lack of control can negatively optimize downstream task models. To tackle these challenges, a two-stage controllable text-image generative model called Diffusion-Geo is presented. In the first stage, an extensive image-text generation dataset called RS-Control is created through prompt engineering of multimodal large language models (MLLMs) and manual prompts for existing datasets, incorporates diverse conditional controls with rich spatial and semantic information. Then RS-Control dataset is utilized to train a universal controllable image generative model. The second stage involves efficient tuning the universal model for different task datasets, minimizing fine-tuning costs while preserving diversity and high-quality features. Experiments conducted on the RSICD caption dataset and WHU change detection dataset demonstrate the superiority of Diffusion-Geo over other state-of-the-art models in image generation. Miaoxin Cai, Wei Zhang 0389, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Liang Chen 0004, Can Li 0005 |
IGARSS | 5 |
| 2024 | A Triplet Multi-Task Learning Network for Semantic Change DetectionabstractSemantic change detection (SCD) aims to provide the change locations and extend the detailed semantic change categories before and after the observation intervals, being a pivotal task in remote sensing community. Recent studies indicate that multi-task learning paradigm is an efficient solution for modeling the SCD. However, the temporal dependency information and semantic feature separability are not efficiently explored. To overcome the above limitations, this paper proposes a new Triplet Multi-task learning Network (TMNet) for SCD, in which a differential image branch is introduced to extract the temporal information reflecting the difference in original image pairs and two temporal branch to extract the semantic representation of each image. Then, a class-wise contrastive loss is employed to deal with the change types imbalance and improve the category discrimination. Finally, experiments are carried on a public SCD dataset to demonstrated the effectiveness of the proposed method. Shan Dong, Huazhe Guo, He Chen 0004 |
IGARSS | 3 |
| 2024 | Regression-Guided Positive Sample Refocusing Paradigm for Tiny Object Detection in Aerial ImagesabstractTiny object detection represents a pivotal challenge in remote sensing intelligent interpretation, necessitating detectors to exhibit heightened precision in object localization. However, typical model optimization strategies cannot release the detector’s potential for precisely localizing objects. And the lack of interpretability in detection box filtering based on object classification scores serves as a constraint on further performance improvement. Therefore, this paper proposed a novel model optimization strategy to thoroughly unleash the potential of the detector for precise localization. Then, the utilization of object comprehensive confidence score enhances the interpretability of the post-processing step for detection boxes. Rigorous experiments on the AI-TOD dataset have demonstrated the effectiveness of our method, achieving state-of-the-art performance. Lihui Ge, He Chen 0004, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, Fukun Bi, Liang Chen 0004 |
IGARSS | 2 |
| 2024 | Uncertainty-Injected Cross-Domain Few-Shot Scene Classification From Remote Sensing ImageryabstractCross-domain few-shot scene classification (CDFSSC) is crucial for remote sensing (RS) applications since it aims at transferring knowledge learned from the source domain to the target domain to facilitate the model’s few-shot classification for the target domain. However, existing methods ignored the feature statistic discrepancy caused by domain shifts, leading to an inferior performance on the target domain. In this paper, to facilitate the model’s adaptation of the domain shifts and achieve better cross-domain knowledge transfer, an uncertainty-injected cross-domain framework called UICD is proposed for CDFSSC tasks from RS imagery. First, a semi-supervised teacher-student structure is employed to achieve cross-domain knowledge transfer by conducting supervised learning on labeled source data and establishing consistent predictions on unlabeled target data. Secondly, uncertainty is injected in feature statistic modeling during cross-domain training to obtain more diverse feature statistics for data from both the source and target domains, which could promote the robustness and adaptation of the model to domain shifts, thus enabling the model to better adapt to unforeseen variations in the target domain. Extensive experiment results indicate the efficacy and superiority of the proposed methods. Can Li 0005, He Chen 0004, Yin Zhuang, Liang Chen 0004 |
IGARSS | 2 |
| 2024 | Multi-Temporal Images Generation for Building Change Detection Performance PromotionabstractThe changes in building are important basis for urban monitoring. However, due to the rarity and sparsity of the occurrence of changes in buildings, collecting effective bitemporal image pairs is challenging, as it requires long-term observation over several months or even years. Additionally, annotating large-scale change detection datasets is time-consuming and labor-intensive. Consequently, data scarcity issues lead to insufficient training of building change detection models. To address this, we propose a data generation method Building Generation GAN (BG-GAN). Different from other GANs, the BG-GAN is trained based on adversarial consistency loss, enabling the model to generate new bi-temporal image pairs with various types of building changes. To verify the effectiveness of the proposed methods on change detection task, BG-GAN is utilized to perform building change samples generation on two building change detection datasets (LEVIR-CD and WHU-CD). The experimental results demonstrate that the proposed method can improve the robustness and generalization of change detection model to detect pseudo changes. Yute Li, Wei Li 0032, Nan Wang 0038, Chenzhong Gao, Yin Zhuang, He Chen 0004 |
IGARSS | 6 |
| 2024 | Advancing Controllable Diffusion Model for Few-Shot Object Detection in Optical Remote Sensing ImageryabstractFew-shot object detection (FSOD) from optical remote sensing imagery has to detect rare objects given only a few annotated bounding boxes. The limited training data is hard to represent the data distribution of realistic remote sensing scenes, restricting the performance of FSOD. Recently, learning conditional controls for text-to-image diffusion model has achieved great progress, which is capable of precisely generating the controllable yet imaginational images by text prompt and spatially localized input conditions. Accordingly, in this work, we aim to explore the potential of diffusion model and propose a solution for few-shot object detection by controllable data generation. Firstly, draw upon a few annotated objects, their bounding boxes and categories are respectively used as the spatial conditions and text prompts, then employ them into large text-to-image diffusion models for controlled image generation. Secondly, based the generated images, in order to adapt to the scale and orientation variances of remote sensing objects, a data transformation is devised for boosting the robustness of model training. Finally, some experiments were conducted on public remote sensing dataset DIOR, and the results proved its effectiveness. Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, Fukun Bi |
IGARSS | 5 |
| 2024 | Axial-Shift Feature Interaction and Prototype-Guided Penalty Constraint for Remote Sensing Change DetectionabstractAt present, deep learning (DL) methods for remote sensing (RS) change detection (CD) are developing rapidly. However, there are still many challenges in complex RS scenarios. Factors, such as season and illumination, contribute to minimal differences in radiation characteristics between the changed area and the background, making them difficult to distinguish. This letter proposes an axial-shift feature interaction and prototype-guided penalty constraint network (ASPGNet) to address this problem. ASPGNet integrates axial-shift feature interaction (ASFI) module and prototype-guided penalty constraint (PGPC) loss. The ASFI module facilitates interaction among adjacent features through axial-shift operations in the width/height directions, aiming to obtain discriminative feature representations of the changed area. The PGPC loss utilizes prototypes to adaptively identify and weigh confusing pixel features, ensuring distinguishability between change and nonchange features and, thereby, generating accurate CD results. We evaluate the proposed method on the WHU-CD and LEVIR-CD datasets, achieving the$F1$scores of 93.22% and 91.49%, respectively. These results demonstrate the effectiveness of the proposed method. He Chen 0004, Jue Wang 0011, Wenchao Liu 0001, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Regression-Guided Refocusing Learning With Feature Alignment for Remote Sensing Tiny Object DetectionabstractTiny object detection is a formidable challenge in remote sensing intelligent interpretation. Tiny objects are usually fuzzy, densely distributed and highly sensitive to positioning errors, which leads to the mainstream detector usually achieving suboptimal detection performance when facing tiny objects. To address the mismatch of mainstream detector architectures and model optimization strategies in the context of tiny object detection, this paper presents an efficient and interpretable algorithm for tiny object detection, termed the Cross-Attention based Feature Fusion Enhanced tiny object detection Network (CAF2ENet). First, the cross-attention mechanism is introduced to refine the upsampling results of deep features. This refinement improves the precision of multi-scale feature fusion. Second, a training strategy named regression-based refocusing learning is introduced. Deviating from the conventional optimization strategy, our method guides the optimizer to prioritize higher-quality detection boxes by adjusting sample weights. This adjustment significantly amplifies the detector’s potential to achieve superior detection results. Finally, the object composite confidence score is employed for the interpretable filtering of detection boxes. Extensive experiments on Tiny Object Detection in Aerial Images (AI-TOD) and object Detection in Optical Remote sensing images (DIOR) datasets are carried out, and comparison indicate that the proposed CAF2ENet can perform the remarkable performance compared to other state-of-the-art (SOTA) tiny object detection detectors, as it can reach 63.7% Average Precision (AP50) on AI-TOD and 75.4%AP50on DIOR, achieve SOTA performance. Lihui Ge, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Hao Dong 0003, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | DECOR: Dynamic Decoupling and Multiobjective Optimization for Long-Tailed Remote Sensing Image ClassificationabstractIn the realm of remote sensing, targets of interest span a range of categories. However, their distribution is not always uniform. Certain categories substantially outnumber others, resulting in what’s termed a ‘long-tailed distribution’ in remote sensing imagery. This imbalanced distribution often biases a classifier’s focus toward the more abundant (head) classes, at the detriment of the less-represented (tail) classes. Such biases undermine the classifier’s generalization performance, particularly in the context of remote sensing image classification (RSIC). While existing mitigation approaches such as resampling, reweighting, and transfer learning offer some respite, they often miss out on in-depth knowledge refinement, rendering them less effective for severe long-tailed RSIC scenarios. To counter these challenges, we introduce DECOR, a dynamic decoupling and multi-objective optimization framework. Within DECOR, the feature extractor and classifier are dynamically decoupled, promoting superior feature representation and classifier training. Then, a multi-objective optimization approach is proposed to delve deeper, refining feature representation at the knowledge level using learnable feature centroids coupled with masked world knowledge learning. Moreover, to combat the pronounced effects of sample imbalance on classifier training, we employ a class-balanced re-sampling technique paired with a parameter-efficient adapter, which sharpens the classifier’s decision boundary and bridges the gap between representation and classification. DECOR’s efficacy is validated through comprehensive experiments on several datasets, including the NWPU-RESISC45-LT (NWPU-LT), AID-LT, and our self-built BIT-AFGR50-LT. Experimental results demonstrate DECOR’s marked enhancement in performance on long-tailed datasets. Our source code is available at: https://github.com/ChloeeGrace/DECOR. Jianlin Xie, Guanqun Wang, Yin Zhuang, Can Li 0005, Tong Zhang 0028, He Chen 0004, Liang Chen 0004, Shanghang Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Heterogeneous Prototype Distillation With Support-Query Correlative Guidance for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification (FSRSSC) aims to identify unseen classes only relying on very limited training samples. However, scarce training samples are insufficient to support a robust classwise representation, which is easily influenced by agnostic biases from diverse testing scenarios. Fortunately, there are abundant spatial contextual clues that exist in very limited training samples to have an enormous potential to establish discriminative and transferable concepts. Thus, in this article, a hybrid architecture called ProtoConViT is proposed to learn a powerful classwise representation based on spatial contextual clues for FSRSSC promotion. First, support-query correlative guidance is designed to generate more stable spatial connections among support and query data based on intermediate convolution neural network (CNN) feature maps, which not only can be embedded into each episodic training task to reduce redundant spatial contextual representation learning space of vision transformer (ViT) but also can assist it in rapidly capturing critical spatial contextual clues to classify query data into one of classes from support set. Second, followed by the designed support-query correlative guidance, a novel heterogeneous prototype distillation is proposed to integrate the advantages of CNN and ViT for heterogeneous prototype construction, which can rapidly set up discriminative and transferable concepts for FSRSSC. Third, corresponding to the proposed ProtoConViT, a joint loss is designed to make the model rapid convergence based on meta-learning. Finally, extensive experiments are carried out on three FSRSSC benchmarks, and comparative results indicate that the proposed ProtoConViT can achieve a superior FSRSSC performance. Yin Zhuang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Hybrid Transformer Network for Change Detection Under Self-Supervised PretrainingabstractThis paper presents a Siamese network architecture based on a multi-scale hybrid convolution-Transformer (CTUNet) for Change Detection (CD) in a pair of co-registered optical remote sensing images. Different form CD frameworks based on convolution neural networks (CNNs) and pure Transformer networks, this method combines a convolution-Transformer hybrid encoder with a multi-scale change information extraction decoder in a Siamese network architecture. It overcomes the inherent limitations of CNN and Transformer and effectively integrates the multi-scale information required for accurate CD. To learn better discriminative representations from various scales, we propose a masked auto-encoder scheme (CTMAE) to adapt to building targets with varying morphological scales, further unleashing the potential of CTUNet. Experiments on two CD datasets show that the proposed self-supervised pre-trained hybrid convolution-Transformer CTUNet architecture achieves better CD performance than previous methods. Yongjing Cui, Yin Zhuang, Shan Dong, Peng Gao 0007, He Chen 0004, Liang Chen 0004 |
IGARSS | 6 |
| 2023 | Uncertainty-Aware Dynamic Learning for Cross-Domain Few-Shot Scene Classification from Remote Sensing ImageryabstractCross-domain few-shot scene classification (CDFSSC) is devoted to transferring knowledge from the source domain to the target domain and facilitating few-shot classification for the target domain. However, due to the domain shifts between source and target domains, high uncertainty would be generated in the knowledge transfer process, leading to unreliable cross-domain learning, which degenerates classification performance on the target domain severely. Thus, in this paper, aiming to reduce the interference of high uncertainty and improve the reliability of cross-domain knowledge transfer, a novel uncertainty-aware dynamic learning (UDL) framework is proposed for CDFSSC from remote sensing imagery. First, a mean-teacher architecture combining pseudo-labeling and consistency regularization is utilized to achieve cross-domain learning. Second, a UDL strategy is proposed to divide data into positive and negative samples based on a well-designed uncertainty-aware dynamic threshold, conducting positive and negative learning respectively, to advance a more reliable knowledge transfer. Third, to further improve cross-domain capability, a self-entropy loss is designed to reduce the epistemic uncertainty of the model. Extensive experiment results indicate the superiority of our proposed methods. Can Li 0005, He Chen 0004, Yin Zhuang, Shanghang Zhang |
IGARSS | 2 |
| 2023 | Multi-Grained Global-Local Semantic Feature Fusion for Few Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification aims to classify unseen scenes by using only a few labeled samples. Hence, how to set up a more effective feature description according to a few labeled samples, becomes an important issue. In this paper, in view of more complicated remote sensing scenes containing several hierarchical and coupled spatial relations (e.g., internal and external spatial contexts), which severely hinder the feature extraction under few-shot learning scenarios, a multi-grained global-local semantic feature fusion (MGGL-SFF) method is proposed for few-shot remote sensing scene classification, which can better combine the global discriminative spatial semantic features with local transferable fragment features to set a powerful prototype representation up for few shot learning. Finally, experiments are carried out on defined few-shot remote sensing scene classification benchmark, and results proved the proposed MGGL-SFF can achieve a new state-of-the-art performance. Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004 |
IGARSS | 5 |
| 2023 | Contour Modeling Arbitrary-Oriented Ship Detection From Very High-Resolution Optical Remote Sensing ImageryabstractUnder the multiscale distribution, due to dramatic aspect ratio variance leading to prominent arbitrary-oriented character of ships in very high-resolution (VHR) optical remote sensing imagery, how to generate accurate oriented bounding box (OBB) becomes a hot research topic for arbitrary-oriented ship detection. Consequently, in this letter, a concise and effective one-stage anchor-free contour modeling detector called CMDet is proposed for accurate arbitrary-oriented ship detection. Different from currently existed methods via carefully decoupling several independently characteristic parameters for OBB modeling and regression, we resolve the OBB modeling by jointly regressing the contour information. Specifically, the contour information is expressed as a series of Fourier transform coefficients, which are generated by setting up the mapping relation of 1-D Fourier contour coefficients and spatial OBB contour. In addition, a new inherent geometry loss is designed to make detector better learn the geometry information in training phase. After that, the proposed CMDet only needs to predict the correct center point of ships and regress the corresponding entire 1-D Fourier contour coefficients to generate accurate OBB for ship detection. Finally, extensive experiments are carried out on two public OBB ship detection datasets (e.g., HRSC2016 and DIOR-ship), and comparison results demonstrate that the proposed CMDet can obtain the competitive result than the state-of-the-art (SOTA) detectors. Yin Zhuang, Yuqun Liu, Tong Zhang 0028, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Full Semantic Constructed Network for Urban Use Classification From Very High-Resolution Optical Remote Sensing ImageryabstractRecently, semantic segmentation technology has been a research hotspot in optical remote sensing urban use classification. However, because of coupled semantic relations in very high-resolution and complex urban scenes, a more effective semantic description for pixelwise urban use interpretation has become a challenge. Then, aiming to set up a more effective semantic description, the effective receptive field (ERF) is analyzed in general convolutional neural networks. The unreasonable ERF distribution in the stacked convolutional layers of the encoder would lead to a large amound of small ERFs and fewer not large enough ERFs that form a naive semantic description in decoder. Therefore, in this article, a novel full semantic constructed network (FSCNet) is proposed to improve the naive semantic description and set up an effective semantic description. First, to avoid noise from shallow feature layers, a residual refinement convolution is designed to optimize the full-scale skip connections based on the U-shaped encoder–decoder. Second, an interscale fusion module is newly designed for multiscale feature fusion, which can generate three initial semantic modalities that are prepared for redefining the full semantic description. Third, a multiscale local context spatial attention module and boundary supervision are designed for an initial shallow semantic modality to capture the pure boundary information, and then, pyramid spatial pooling is employed for an initial deep semantic modality to further enlarge the ERF and obtain more abstract global information. Next, a self-calibration convolution combined with the atrous spatial pyramid pooling is designed to rectify and enrich an initial middle semantic modality, which can improve the naive semantic description and bridge the semantic gap between the redefined shallow and deep semantic modalities to advance the full semantic feature fusion. Finally, extensive experiments are carried out on three benchmarks (e.g., ISPRS Vaihingen, Potsdam, and DLRSD), and comparative results show that the proposed FSCNet can get remarkable performance compared to state-of-the-art (SOTA) methods. Besides, the code is available athttps://github.com/DorisCV/FSCNet. Shan Dong, Yin Zhuang, He Chen 0004, Tong Zhang 0028, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | All Adder Neural Networks for On-Board Remote Sensing Scene Classification
Ning Zhang 0042, Jue Wang 0011, He Chen 0004, Wenchao Liu 0001, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Posterior Instance Injection Detector for Arbitrary-Oriented Object Detection From Optical Remote-Sensing ImageryabstractArbitrary-oriented object detection (AOOD) from optical remote sensing imagery has to correctly generate delicate oriented boundary boxes (OBBs) and meanwhile identify their specific categories. However, how to make detectors learn delicate parameters of OBBs, especially for the crucial orientation information, and identify object category from complex background becomes a challenge task. Therefore, in this article, for exploring a better way to guide the detector to learn specific category and parametric information of OBBs, a novel one-stage anchor-free detector called Posterior Instance Injection Detector (PIIDet) is proposed for AOOD. First, as the anchor-free manner lacks prior information, an object-aware posterior guidance (OAPG) structure is proposed to generate specific-category instances used for conditioning on OBB prediction. This structure can assist the proposed PIIDet in better learning the relative parametric information of OBBs corresponding to their specific categories. Besides, to guarantee a high quality injection of specific-category instances, a new hierarchical feature fusion module is developed to establish a suitable multi-scale feature mapping space. Second, considering the negative optimization of angle regression, which is caused by the boundary discontinuity of angular periods and sudden shifts of the relation between width and height in training phase, a novel binary classification embedded angle regression space (BCE-RegSpace) is devised for providing continuous angle regression space and stable relation between width and height. Finally, extensive experiments are executed on three AOOD benchmarks (e.g., DOTA, DIOR-R and HRSC2016), and results proved that the proposed concise one-stage anchor-free PIIDet can reach the state-of-the-art (SOTA) performance and meanwhile have an impressive inference speed. Tong Zhang 0028, Yin Zhuang, He Chen 0004, Guanqun Wang, Lihui Ge, Liang Chen 0004, Hao Dong 0003, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Adaptive Local Context Embedding for Small Vehicle Detection from Aerial Optical Remote Sensing ImagesabstractSmall vehicle detection is one of the remaining challenging task because the ambiguous appearance is against complex background interference. Consequently, in order to improve the performance of small vehicle detection from aerial optical remote sensing images, a novel adaptive local context (ALC) embedding way is designed and further introduced into an anchor free detection manner which is called ALC-Net, and in ALC-Net, it can adaptively set up the effective local context feature to improve keypoint description of small vehicles and boost the detection performance without adding extra prior information. Finally, several experiments are carried out on two widely used datasets (e.g., UCAS-AOD [1] and VEDAI [2]) and the results indicate that the proposed ALC-Net can exhibit the competitive small vehicle detection performance than other detectors. Shanjunyu Liu, Yin Zhuang, Hao Dong 0003, Peng Gao 0007, Guanqun Wang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li |
IGARSS | 8 |
| 2022 | Bilateral Semantic Fusion Siamese Network for Change Detection From Multitemporal Optical Remote Sensing ImageryabstractChange detection (CD) is an essential task in optical remote sensing, and it can be used to extract the valid information from sequential multitemporal images. However, since the character of long-term revisiting and very high resolution (VHR) development, the great differences of illumination, season, and interior textures between bitemporal images bring considerable challenges for pixel-wise CD. In this letter, focusing on accurate pixel-wise CD, a bilateral semantic fusion Siamese network (BSFNet) is proposed. First, to better map bitemporal images into semantic feature domain for comparison, a novel BSFNet is designed to effectively integrate shallow and deep semantic features, which can provide pixel-wise CD results with complete regions and clear boundary locations. Then, in order to facilitate the reasonable convergence of the proposed BSFNet, a scale-invariant sample balance (SISB) loss is designed for metric learning to avoid the problems of sample imbalance and scale variance. Finally, extensive experiments are carried out on two published CDD and LEVIR CD datasets, and results indicate that the proposed BSFNet can provide superior performance than the other state-of-the-art methods. Our work is available athttps://github.com/ClarissaDHL/BSFNet. Hailin Du, Yin Zhuang, Shan Dong, Can Li 0005, He Chen 0004, Boya Zhao, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Effective Multiscale Residual Network With High-Order Feature Representation for Optical Remote Sensing Scene ClassificationabstractScene classification of optical remote sensing is a basic but important task because of its broad application in a range of fields. Due to the powerful feature extraction capabilities, convolutional neural networks (CNNs) have been widely used in optical remote sensing scene classification tasks. Despite the remarkable efforts have been achieved, there are still several problems existed, including the effective multiscale feature description for complex scene, rotation, and low interclass diversity problems. In this letter, to address these mentioned problems and construct a powerful CNN for optical remote sensing scene classification, an effective multiscale residual network with a high-order feature representation (MRHNet) is proposed. First, data preprocessing is utilized to adapt the rotation invariance problem. Second, related to the original residual module, a pyramid convolution is introduced to realize the multiscale feature extraction, and then, its feature description ability is further improved by an effective channel attention module. Third, inspired by the tensor decomposition and its completion, a high-order feature representation structure is designed for recovering discriminative fine-scale details into deep layers to solve the low interclass diversity problem. Finally, extensive experiments are carried on two widely used scene classification datasets (e.g., AID and NWPU-RESISC45), and comparing results show that the proposed MRHNet can achieve superior performances. Can Li 0005, Yin Zhuang, Wenchao Liu 0001, Shan Dong, Hailin Du, He Chen 0004, Boya Zhao |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | MFST: A Multi-Level Fusion Network for Remote Sensing Scene ClassificationabstractScene classification has become an active research area in remote sensing (RS) image interpretation. Recently, Transformer-based methods have shown great potential in modeling global semantic information and have been exploited in RS scene classification. In this letter, we propose a multi-level fusion Swin Transformer (MFST), which integrates a multi-level feature merging (MFM) module and an adaptive feature compression (AFC) module to further boost the performance for RS scene classification. The MFM module narrows the semantic gaps in multi-level features via patch merging in lower-level feature maps and lateral connections in the top-down pathway. The AFC module makes multi-level features have smaller dimensions and more coherent semantic information by adaptive channel reduction. We evaluate the proposed network on the aerial image dataset (AID) and NWPU-RESISC45 (NWPU) datasets, and the classification results reveal that the proposed network outperforms several state-of-the-art (SOTA) methods. Ning Zhang 0042, Wenchao Liu 0001, He Chen 0004, Yizhuang Xie |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | NAS-Based CNN Channel Pruning for Remote Sensing Scene ClassificationabstractRecently, convolutional neural network (CNN)-based remote sensing scene classification has achieved great success. However, the prohibitively expensive computation and storage requirements of state-of-the-art models have hindered the deployment of CNNs on on- board platforms. In this letter, we propose a differentiable neural architecture search (NAS)-based channel pruning method to automatically prune the CNN models. In the proposed method, the importance of each output channel is measured by a trainable score. The scores are optimized by an NAS method to search a good-performance pruned structure. After the search process, a global score threshold is adopted to derive the pruned model. A cost-awareness loss is proposed for the search process to encourage the floating-point operation (FLOP) compression ratio of the pruned model coverage to a desired value. We apply the proposed method to ResNet-34 and VGG-16 to verify the performance. The NWPU-RESISC-45 and UC Merced Land-Use (UCM) datasets are used for the performance evaluation. A comparison with state-of-the-art pruning methods demonstrates that the proposed method can achieve competitive performance with a similar reduction in FLOP. Ning Zhang 0042, Wenchao Liu 0001, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Sphere Loss: Learning Discriminative Features for Scene Classification in a Hyperspherical Feature SpaceabstractThe power of features considerably influences the classification performance of remote sensing scene classification (RSSC). Recently, deep convolutional neural networks (DCNNs) have been used to extract powerful scene features. Nevertheless, confusion and overlap still occur in the feature space, leading to inaccurate RSSC. To alleviate this problem, we propose a novel deep metric learning loss function incorporated into a sphere loss to enhance the discrimination of feature representations. Inspired by two representative loss functions (i.e., angular loss and center loss), the proposed sphere loss learns a unique cluster center for each class in a remote sensing scene. Because the cluster centers and features are restricted by an introduced geometrical constraint, the intraclass distance of features decreases, while the interclass distance increases. Moreover, we introduce a spatial constraint, i.e., a uniformity coefficient on different cluster centers, which causes the centers to form a uniform distribution that maximizes the interclass distances between features. Extensive analysis and experiments on three commonly used RSSC data sets consistently show that, compared with state-of-the-art methods, the proposed sphere loss can effectively learn discriminative feature representations and significantly improve RSSC. Jue Wang 0011, He Chen 0004, Long Ma 0003, Liang Chen 0004, Xiaodong Gong, Wenchao Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | FSoD-Net: Full-Scale Object Detection From Optical Remote Sensing ImageryabstractObject detection is an essential task in computer vision. Recently, several convolution neural network (CNN)-based detectors have achieved a great success in natural scenes. However, for optical remote sensing images with a large scale of view, lower proportion of foreground target pixels and drastic differences in object scale present considerable challenges. To address these problems, we propose a novel one-stage detector called the full-scale object detection network (FSoD-Net) which consists of proposed multiscale enhancement network (MSE-Net) backbone cascaded with scale-invariant regression layers (SIRLs). First, MSE-Net provides the multiscale description enhancement by integrated the Laplace kernel with fewer parallel multiscale convolution layers. Second, SIRLs contain three different isolated regression branch layers (i.e., corresponding to small, medium, and large scales), which make default discrete scale bounding boxes (bboxes) cover full-scale object information in regression procedure. A novel specific scale joint loss is also designed that uses the softmax function combined with a strong$L_{1}$-norm constraint in each regression branch layer. It can further speed up the convergence and improve the classification scores of predicted bboxes. Finally, extensive experiments are carried on challenge data sets of large-scale dataset for object detection in aerial images (DOTA) and object detection in optical remote sensing images (DIOR) which contain multiple instances from different imaging platforms, and these results demonstrate that FSoD-Net can achieve better performance than other state-of-the-art one-stage detectors, and it can reach a mean average precision (mAP) of 75.33% on DOTA and 71.80% mAP on DIOR, respectively. Especially, the average precision (AP) of tiny object detection can improve 10%–20% approximately. Guanqun Wang, Yin Zhuang, He Chen 0004, Tong Zhang 0028, LianLin Li, Shan Dong, Qianbo Sang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiscale Semantic Fusion-Guided Fractal Convolutional Object Detection Network for Optical Remote Sensing ImageryabstractOptical remote sensing object detection is a challenging task, because of the complex background interference, ambiguous appearances of tiny objects, densely arranged circumstances, and multiclass object with vaster scale variances and irregular aspect ratios. The performance of object detection is seriously restricted. Thus, in this article, inspired by the anchor-free object detection framework, and aiming to solve these difficulties to improve the optical remote sensing object detection performance, a powerful one-stage detector of multiscale semantic fusion-guided fractal convolution network (MSFC-Net) is proposed. First, facing these strong-coupled semantic relations in each complex scene, a compound semantic feature fusion (CSFF) way is designed for generating an effective semantic description, which is a benefit to pixel-wise object center point interpretation. In addition, it can be easily extended into a semantic segmentation task. Second, in view of accurate multiclass pixel-wise center point predictions based on an effective compound semantic description, a novel fractal convolution (FC) regression layer is designed, which adaptively achieves the regression of multiscale bounding boxes (bboxes) with irregular aspect ratio under no priori information. Third, related to the set up FC regression layer, a specific hybrid loss is designed to make the proposed MSFC-Net converge better. Finally, the extensive experiments on challenge data sets of large-scale dataset for object detection in aerial images (DOTA) and object detection in optical remote sensing images (DIOR) datasets are carried out, and comparisons indicate that the proposed MSFC-Net can perform the remarkable performance than other state-of-the-art one-stage detectors, as it can reach 80.26% mean average precision (mAP) and 79.33% mF1 on DOTA and 70.08% mAP and 73.45% mF1 on DIOR. Then, our work is available athttps://github.com/ZhAnGToNG1/MSFC-Net. Tong Zhang 0028, Yin Zhuang, Guanqun Wang, Shan Dong, He Chen 0004, LianLin Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Task-Driven Regional Saliency Analysis Based on a Global-Local Feature Assembly Network in Complex Optical Remote Sensing ScenesabstractSaliency analysis is an essential task in computer vision and aims to generate distinguishing foreground features from background features. However, due to complex structure distributions in large-scale optical remote sensing scenes, generating effective feature descriptions for regional saliency analysis is challenging. Therefore, in this study, we proposed a novel global-local feature assembly method called GLFA-Net based on a convolution neural network (CNN) that can adaptively learn an effective feature representation for regional saliency analysis to achieve the region-of-interest (ROI) (e.g., aircraft carrier, airport, and urban area) extraction from complex optical remote sensing images. In addition, we also collected these complex optical remote sensing scene images from Google Earth and DOTA data sets to demonstrate the effectiveness of the proposed method. Finally, experimentation shows that the proposed regional saliency analysis method can produce better ROI extraction performance than other methods, reaching a 0.057 mean absolute error (MAE), a 0.703 Kappa coefficient, and a 0.735 F1-score. Yin Zhuang, He Chen 0004, LianLin Li |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Mixed-Precision Quantization for CNN-Based Remote Sensing Scene ClassificationabstractExtensive convolutional neural network (CNN)-based methods have been widely used in remote sensing scene classification. However, the dense operation and huge memory storage of the state-of-the-art models hinder their deployment on low-power embedded devices. In this letter, we propose a mixed-precision quantization method to compress the model size without accuracy degradation. In this method, we propose a symmetric nonlinear quantization scheme to reduce the quantization error. A corresponding three-step training strategy is proposed to improve the performance of the quantized network. Finally, based on the proposed scheme and training strategy, we propose a neural architecture search (NAS)-based quantization bit-width search (NQBS) method. This method can automatically select a bit width for each quantized layer to obtain a mixed-precision network with an optimal model size. We apply the proposed method to the ResNet-34 and SqueezeNet networks and evaluate the quantized networks on the NWPU-RESISC45 data set. The experimental results show that the mixed-precision quantized networks under the proposed method strike a satisfying tradeoff between classification accuracy and model size. He Chen 0004, Wenchao Liu 0001, Yizhuang Xie |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Feature Enhanced Centernet for Object Detection in Remote Sensing ImagesabstractMulti-scale object detection in optical remote sensing imagery is a challenging task due to the varied object scales. Existed state-of-art object detection methods have achieved significant growth. However, most of the methods are based on default anchors, which need to be predefined. The multi-scale object detection accuracy still needs to be improved, especially for small and dense objects. To improve the robustness of the detection algorithm and the performance of multi-scale object detection, a novel anchor-free multi-scale object detection method Feature Enhanced CenterNet is proposed in this paper. First, we use the “encoder-decoder” structure and introduce horizontal connections to enhance feature representation capabilities. Second, an context-aware up-sampling method is proposed to obtain feature maps with suitable scale. To demonstrate the performance of the proposed method, we perform abundant experiments on the public remote sensing datasets. The experimental results demonstrate the robustness and effectiveness of the proposed method. Tong Zhang 0028, Guanqun Wang, Yin Zhuang, He Chen 0004, Hao Shi 0006, Liang Chen 0004 |
IGARSS | 4 |
| 2020 | Land Cover Classification From VHR Optical Remote Sensing Images by Feature Ensemble Deep Learning NetworkabstractLand cover classification is a popular research field in remote sensing applications, which have to both consider the pixel-level classification and boundary mapping comprehensively. Although multi-scale features in deep learning (DL) network have a powerful classification ability, how to use multi-scale feature description to produce an accurate land cover classification from very high resolution (VHR) optical remote sensing image is still a challenging task because of large intraclass or small interclass difference of land covers. Therefore, aiming at achieving more accurate pixel-level land cover classification, we proposed a novel feature ensemble network (FE-Net), which includes the multi-scale feature encapsulation and enhancement two phases. First, there are encapsulated shallow, middle, and deep scale feature layers from Resnet-101 backbone. Second, related to multi-scale feature description enhancement, these 2-D dilation convolutions with different sample rates are employed on each scale feature layer. After that, optimal channel selection works on each intrascale and interscale feature layers sequentially. Finally, extensive experiments proved that the proposed FE-Net combined with a special joint loss function outperforms state-of-the-art DL based methods. It can achieve the 68.08% and 65.16% of the mean of class-wise intersection over union (mIoU) on ISPRS and GID data sets, respectively. Shan Dong, Yin Zhuang, Zhanxin Yang, Long Pang, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | FRF-Net: Land Cover Classification From Large-Scale VHR Optical Remote Sensing ImagesabstractDeep learning (DL) technique is widely applied in remote sensing (RS) applications because of its outstanding nonlinear feature extraction ability. However, with regard to the issues of large-scale and very high-resolution (VHR) land cover classification, multi-object distributions and clear appearance with large intraclass difference become challenges for refined pixelwise land cover mapping. Focusing on these problems, the letter proposed a novel encoding-to-decoding method called the full receptive field (RF) network (FRF-Net) based on two types of attention mechanism. In the FRF-Net, ResNet-101 is used as the basic backbone. Then, the ensemble feature is generated by encoding the high-level features based on the self-attention mechanism which could achieve full RF to capture long-range semantic. Next, the encoding result is decoded by the fusion attention mechanism combined with the low-level feature to produce a fusion feature which contains a refined semantic description for accurate land cover mapping. Extensive experiments based on the GID and ISPRS data sets proved that the proposed network outperforms the state-of-the-art methods. The FRF-Net achieved 66.71% and 64.17% of the mean of classwise Intersection over Union (mIOU) with smaller computation cost on ISPRS and GID, respectively. Qianbo Sang, Yin Zhuang, Shan Dong, Guanqun Wang, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Marginal Center Loss for Deep Remote Sensing Image Scene ClassificationabstractRecently, remote sensing image scene classification technology has been widely applied in many applicable industries. As a result, several remote sensing image scene classification frameworks have been proposed; in particular, those based on deep convolutional neural networks have received considerable attention. However, most of these methods have performance limitations when analyzing images with large intraclass variations. To overcome this limitation, this letter presents the marginal center loss with an adaptive margin. The marginal center loss separates hard samples and enhances the contributions of hard samples to minimize the variations in features of the same class. Experimental results on public remote sensing image scene data sets demonstrate the effectiveness of our method. After the model is trained using the marginal center loss, the variations in the features of the same class are reduced. Furthermore, a comparison with state-of-the-art methods proves that our model has competitive performance in the field of remote sensing image scene classification. Jue Wang 0011, Wenchao Liu 0001, He Chen 0004, Hao Shi 0006 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Inshore Ship Change Detection Based on Spatial-Temporal SaliencyabstractAutomatical inshore ship change detection in optical remote sensing images is meaningful for a wide range of applications, but still a challenging task. In this paper, a visual search inspired model is proposed for inshore ship change detection. The proposed model achieves inshore ship change information extraction based on spatial-temporal saliency detection which mimics visual search mechanism. Experimental results performed on a Google Earth optical remote sensing image data set demonstrate the effectiveness of the proposed model in terms of visual and objective evaluations. Long Ma 0003, Wenchao Liu 0001, Zhong Han, Jue Wang 0011, He Chen 0004 |
IGARSS | 5 |
| 2019 | Spatial Enhanced-SSD For Multiclass Object Detection in Remote Sensing ImagesabstractAccurate multiclass object detection in remote sensing images is a challenging task, especially for small objects. Since the scales of objects in remote sensing images have a great variance, almost all of the advanced detection methods have shortcomings. Consequently, improving the accuracy of multiclass objects detection has always been the direction of researchers' efforts. In this paper, a spatial enhanced-Single Shot MultiBox Detector (SE-SSD) is proposed. First, to enhance the spatial information, we enlarge the input image channels with embedding oriented-gradients feature maps. Second, the multiple output layers in the backbone network are changed to reduce one pooling operation. Finally, we design a context module to enhance the receptive field for feature layer description in SE-SSD framework. Experimental results on DOTA dataset demonstrate that Spatial Enhanced-SSD method reaches a much higher mean average precision (mAP) than Faster R-CNN, SSD and other classic detection network. Guanqun Wang, Yin Zhuang, Zhiru Wang, He Chen 0004, Hao Shi 0006, Liang Chen 0004 |
IGARSS | 4 |
| 2019 | Detection of Multiclass Objects in Optical Remote Sensing ImagesabstractObject detection in complex optical remote sensing images is a challenging problem due to the wide variety of scales, densities, and shapes of object instances on the earth surface. In this letter, we focus on the wide-scale variation problem of multiclass object detection and propose an effective object detection framework in remote sensing images based on YOLOv2. To make the model adaptable to multiscale object detection, we design a network that concatenates feature maps from layers of different depths and adopt a feature introducing strategy based on oriented response dilated convolution. Through this strategy, the performance for small-scale object detection is improved without losing the performance for large-scale object detection. Compared to YOLOv2, the performance of the proposed framework tested in the DOTA (a large-scale data set for object detection in aerial images) data set improves by 4.4% mean average precision without adding extra parameters. The proposed framework achieves real-time detection for$1024\times 1024$image using Titan Xp GPU acceleration.11https://github.com/WenchaoliuMUC/Detection-of-Multiclass-Objects-in-Optical-Remote-Sensing-Images Wenchao Liu 0001, Long Ma 0003, Jue Wang 0011, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2018 | An Automated FPGA-Based Fault Injection Platform for Granularly-Pipelined Fault Tolerant CORDICabstractAugment of integration and complexity makes VLSI circuits more sensitive to errors. Also, soft errors caused by Single Event Upset (SEU) have become a significant threat to modern electronic systems. Therefore, the demand of high reliability on modern electronic systems keeps increasing. Aiming at reliability evaluation of fault tolerant very large scale integrated circuits implemented on SRAM-based FPGA, an automated fault injection platform via Internal Configuration Access Port (ICAP) for rapid fault injection is presented in this paper. We adopt a granularly-pipelined fault tolerant CORDIC processor as the Design Under Test (DUT), and a C++ script is deployed for the external fault injection control environment and automating the fault injection procedure. The proposed method can achieve quantities of repeating fault injection tests and is suitable for any fault tolerant design implemented in SRAM-Based FPGA. Yu Xie 0021, He Chen 0004, Yizhuang Xie, Chuang-An Mao |
FPT | 2 |
| 2018 | Change Detection in Optical Remote Sensing Images with a Fully Object-Level ApproachabstractIn this paper, an efficient and accurate multilevel change detection method is presented, achieved by the combination of an unsupervised object-based correlation analysis and a supervised post-classification comparison. Before the change detection procedure, fast multitemporal segmentation is applied to provide object-level information on two registered images. Then the proposed object-based correlation analysis method is used to extract the potential changed areas efficiently and stably, which improves the accuracy of overall performance. Notably, all the procedures are highly automatic except for the necessary selection of training examples within all supervised algorithms. The experimental results demonstrate the superior performance of our method compared with the four typical state-of-the-art change detection methods. Long Ma 0003, Zhihong Mai, He Chen 0004, Wenchao Liu 0001, Guichi Liu, Nouman Qadeer Soomro |
IGARSS | 3 |
| 2018 | A Novel Harbor Detection Method Based on Pattern Coding AlgorithmabstractHarbor automatic detection is a scene interpretation in remote sensing image processing. Fast and accurate harbor detection can significantly improve the performance of inshore ship detection. In order to achieve harbor detection in complex remote sensing images, in this paper, a novel pattern coding algorithm is proposed. The proposed harbor detection method has three steps: First, the harbor water area is extracted by using the definition circle (DC) model. Secondly, pre-generated eight KEY patterns are multiplied with the local scenes, and the probability density function (PDF) of the multiplied local scene is recorded. Finally, the Euclidean distance between the eight patterns' PDFs and the original local scene's PDF is calculated, then compared with the threshold and coded, so as to realize the harbor area detection. Experimental results demonstrate that the novel method has outstanding performance on harbor area detection in complex broad width remote sensing images. Guanqun Wang, Yin Zhuang, He Chen 0004, Liang Chen 0004 |
IGARSS | 3 |
| 2018 | Comprehensive Structure Voting Docked Ship Detection from High-Resolution Optical Satellite Images Based on Combined Multi-Orientation Sparse RepresentationabstractInshore ship detection from high-resolution (HR) optical satellite images is a hot research field. However, HR ships multi-scale and multi-orientation characters and harbor scene various interferences affect docked ship detection performance. Therefore, we proposed a multi-orientations sparse dictionaries (MOSDs) algorithm combining with comprehensive structure voting (CSV) to address existed problem and achieve refined docked ship contour region proposal (RP). Moreover, the comparing experiments use a lot of Google Earth harbour images to demonstrate proposed method effectiveness and robustness of HR ships multi-scale and -orientation changing and various harbour background interferences of docked ship detection. Yin Zhuang, He Chen 0004, Liang Chen 0004, Fukun Bi |
IGARSS | 2 |
| 2018 | A novel word length optimization method for radix-2 k fixed-point FFT
Chen Yang 0003, Yizhuang Xie, He Chen 0004 |
Sci. China Inf. Sci. | 3 |
| 2018 | Arbitrary-Oriented Ship Detection Framework in Optical Remote-Sensing ImagesabstractShip detection is a challenging problem in complex optical remote-sensing images. In this letter, an effective ship detection framework in remote-sensing images based on the convolutional neural network is proposed. The framework is designed to predict bounding box of ship with orientation angle information. Note that the angle information which is added to bounding box regression makes bounding box accurately fit into the ship region. In order to make the model adaptable to the detection of multiscale ship targets, especially small-sized ships, we design the network with feature maps from the layers of different depths. The whole detection pipeline is a single network and achieves real-time detection for a $704 \times 704$ image with the use of Titan X GPU acceleration. Through experiments, we validate the effectiveness, robustness, and accuracy of the proposed ship detection framework in complex remote-sensing scenes. Wenchao Liu 0001, Long Ma 0003, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | IORN: An Effective Remote Sensing Image Scene Classification FrameworkabstractIn recent times, many efforts have been made to improve remote sensing image scene classification, especially using popular deep convolutional neural networks. However, most of these methods do not consider the specific scene orientation of the remote sensing images. In this letter, we propose the improved oriented response network (IORN), which is based on the ORN, to handle the orientation problem in remote sensing image scene classification. We propose average active rotating filters (A-ARFs) in the IORN. While IORNs are being trained, A-ARFs are updated by a method that is different from the ARFs of the ORN, without additional computations. This change helps IORN improve its ability to encode orientation information and speeds up optimization during training. We also propose Squeeze-ORAlign (S-ORAlign) by adding a squeeze layer to ORAlign of ORN. With the squeeze layer, S-ORAlign can address large-scale images, unlike ORAlign. An ablation study and comparison experiments are designed on a public remote sensing image scene classification data set. The experimental results demonstrate the effectiveness and better performance of the proposed model over that of other state-of-the-art models. Jue Wang 0011, Wenchao Liu 0001, Long Ma 0003, He Chen 0004, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | A unified reconfigurable floating-point arithmetic architecture based on CORDIC algorithmabstractThis paper presents the design methodology and implementation of reconfigurable coordinate rotation digital computer (CORDIC) architecture that can be configured to operate in different modes and rotations to achieve singleprecision floating point division, multiplication and square-root operations. Through introducing pre- and post-processing, the float-point operations can be integrated into a unified CORDIC iteration procedure. According to the characteristics of different operations, we propose a pipeline-parallel mixed architecture to optimize the area-delay-efficiency. Finally, the prototype based on Xilinx XC7VX690T has been established to test the performance of the proposed design. The result shows the related error with arithmetic computation is less than 10−6, and the resource-consumption of the proposed design is less than the sum of existing IP cores. Linlin Fang, Yizhuang Xie, He Chen 0004, Liang Chen 0004 |
FPT | 4 |
| 2017 | Moving target detection via hierarchical spatiotemporal saliency analysisabstractAutomatic detection of moving targets is one of important research area in the remote sensing field. In this paper, we propose a method that accurately detects moving targets in aerial videos using hierarchical spatiotemporal saliency analysis. First, coarse motion regions are extracted by utilizing global temporal saliency analysis. Based on these local candidate regions, spatial saliency methods are used to obtain accurate description of targets. After fusing spatial and temporal saliency values, we can get refined results of the detection. Considering about the inter-frame consistency of motion, trajectory level analysis is added in the proposed method to eliminate false alarms. Experiments conducted on the VIVID dataset validate the effectiveness and efficiency of the proposed method. Long Ma 0003, Yin Zhuang, He Chen 0004, Nouman Qadeer Soomro |
IGARSS | 4 |
| 2017 | Pyramid integral image reconstruction algorithm for infrared remote sensing sea-land segmentationabstractThe middle wave infrared remote (MWIR) images has complex scene information, low contrast ratios, and bipolar problems. To solve these problems, we propose a method that uses pyramid integral image reconstruction algorithm achieving sea-land automation segmentation. First, we calculate a gradient feature map (GFM), which extracts the structural information from an MWIR scene. Then, the GFM uses for sum are table (SAT) generation. The pyramid integral image reconstruction technology uses different scale factor reconstruct MWIR images by using SAT. Then the adaptive threshold method is employed for the sea-land segmentation on the multi-scale integral reconstruction images. Finally, we get sea-land refine segmentation result of MWIR images by synthesis analysis the multi-scale reconstruction images. By using GFM and pyramid integral image reconstruction operation, are avoid with the complex gray scene information. The integral image reconstruction can enhance the structure information and improve the reconstruction image contrast. This paper proposed method is from structure and texture information view point for sea-land segmentation, so the bipolar problem is solve in our method for MWIR images sea-land segmentation. Penglin Wang, Yin Zhuang, He Chen 0004, Liang Chen 0004, Hao Shi 0006, Fukun Bi |
IGARSS | 3 |
| 2017 | A novel sea-land segmentation based on integral image reconstruction in MWIR images
Yin Zhuang, Dechun Guo, He Chen 0004, Fukun Bi, Long Ma 0003, Nouman Qadeer Soomro |
Sci. China Inf. Sci. | 3 |
| 2017 | Sea - Land Segmentation for Panchromatic Remote Sensing Imagery via Integrating Improved MNcut and Chan - Vese ModelabstractSea-land segmentation is a key step for some important applications of panchromatic remote sensing image processing. However, robust and effective sea-land segmentation for high-resolution panchromatic remote sensing images is still a challenging problem. This letter presents an accurate and robust approach by integrating the improved multiscale normalized cut (IMNcut) method and improved Chan-Vese model for sea-land segmentation. At first, the image is downsampled and segmented into multiple regions by the IMNcut method. Next, the homogeneous regions are merged to obtain a coarse segmentation result. Finally, gray intensity and local entropy features are integrated as discriminants of the improved Chan-Vese model, which is used to obtain the final segmentation result through a low- to high-resolution segmentation scheme. Experimental results performed on several real data sets demonstrate the effectiveness of the proposed model in terms of visual and objective evaluations. Wenchao Liu 0001, Long Ma 0003, He Chen 0004, Zhong Han, Nouman Qadeer Soomro |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Harbor Water Area Extraction From Pan-Sharpened Remotely Sensed Images Based on the Definition Circle ModelabstractHarbor water area extraction is a key step in nearshore environment pollution surveillance using remote sensing image processing techniques. This letter proposes the definition circle (DC) model of color gradient to describe color fluctuations in harbor water surface areas based on pan-sharpened remote sensing images. The DC model includes two steps: center setting and radius tuning. In the center setting process, labeled training set pixels are selected in the red, green, and blue color space. Then, center setting is completed in the hue, saturation, and intensity color space using the perceptron model. In the radius tuning process, positive and negative sample pixels are used to tune the radius value. After these two steps, the DC model can describe the color gradient of a water surface area and provide accurate harbor water area extraction. A series of experiments shows that the proposed DC model is robust and performs better than other extraction methods based on pan-sharpened remote sensing images. Yin Zhuang, Penglin Wang, Yiding Yang, Hao Shi 0006, He Chen 0004, Fukun Bi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | Region-of-Interest Detection via Superpixel-to-Pixel Saliency Analysis for Remote Sensing ImageabstractTraditional region-of-interest (ROI) detection methods for remote sensing images are generally formulated at pixel level and are less efficient when applied on large high-resolution images. This letter presents an accurate and efficient approach via superpixel-to-pixel saliency analysis for ROI detection. At first, the image is downsampled and segmented into superpixels by simple linear iterative clustering. Next, structure tensor and background contrast are used to yield superpixel feature maps for texture and color. After fusing the feature maps, the overall superpixel saliency map is obtained and then used to achieve the final pixel-level saliency map by superpixel-to-pixel mapping. Through experimentations, we validate the effectiveness and computational efficiency of the proposed model in comparison with state-of-the-art techniques. Long Ma 0003, He Chen 0004, Nouman Qadeer Soomro |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Capturing and tracking of building area based on structure saliency in airborne remote sensing video
Fukun Bi, Liang Chen 0004, He Chen 0004 |
Sci. China Inf. Sci. | 5 |
| 2015 | A waterborne salient ship detection method on SAR imagery
Long Ma 0003, Liang Chen 0004, He Chen 0004, Nouman Qadeer Soomro |
Sci. China Inf. Sci. | 4 |
| 2015 | Accurate Urban Area Detection in Remote Sensing ImagesabstractAutomatic urban area detection in remote sensing images is an important application in the field of earth observation. Most of the existing methods employ feature classifiers and thereby contain a data training process. Moreover, some methods cannot detect urban areas in complex scenes accurately. This letter proposes an automatic urban area detection method that uses multiple features that have different resolutions. First, a downsampled low-resolution image is used to segment the candidate area. After the corner points of the urban area are extracted, a weighted Gaussian voting matrix technique is employed to integrate the corner points into the candidate area. Then, the edge features and homogeneous region are extracted by using the original high-resolution image. Using these results as the input, the processes of guided filtering and contrast enhancement can finally detect accurately the urban areas. This method combines multiple features, such as corner, edge, and regional characteristics, to detect the urban areas. The experimental results show that the proposed method has better detection accuracy for urban areas than the existing algorithms. Hao Shi 0006, Liang Chen 0004, Fukun Bi, He Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2014 | Simplified addressing scheme for mixed radix FFT algorithmsabstractA mixed radix algorithm for the in-place fast Fourier transform (FFT), which is broadly used in most embedded signal processing fields, can be explicitly expressed by an iterative equation based on the Cooley-Tukey algorithm. The expression can be applied to either decimation-in-time (DIT) or decimation-in-frequency (DIF) FFTs with ordered inputs. For many newly emerging low power portable computing applications, such as mobile high definition video compressing, mobile fast and accurate satellite location, etc., the existing methods perform either resource consuming or non-flexible. In this paper, we propose a new addressing scheme for efficiently implementing mixed radix FFTs. In this scheme, we elaborately design an accumulator that can generate accessing addresses for the operands, as well as the twiddle factors. The analytical results show that the proposed scheme reduces the algorithm complexity meanwhile helps the designer to efficiently choose an arbitrary FFT to design the in-place architecture. Cuimei Ma, Yizhuang Xie, He Chen 0004, Yi Deng 0005, Wen Yan 0007 |
ICASSP | 3 |
| 2014 | A coarse-to-fine image registration method based on visual attention model
Long Ma 0003, Fukun Bi, He Chen 0004 |
Sci. China Inf. Sci. | 5 |
| 2013 | An improved constant coefficient multiplication algorithm based on cascaded adder graph
He Chen 0004, Xiujie Qu, Long Pang, Jiyang Yu |
Sci. China Inf. Sci. | 1 |
| 2013 | A novel conflict-free parallel memory access scheme for FFT constant geometry architectures
Cuimei Ma, He Chen 0004, Jiyang Yu |
Sci. China Inf. Sci. | 2 |
| 2013 | Improved Goldschmidt division method using mapping of divisors
Wen Yan 0007, Xiujie Qu, He Chen 0004, Jiyang Yu |
Sci. China Inf. Sci. | 3 |