DeRen Li

dblp:28/9934 · also De-Ren Li, Deren Li · DBLP profile ↗
← Back
88ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0001-5977-3081ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 70 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorComputer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 GMLNet: a lightweight frequency-gradient framework for gravel-mulched land segmentation in high-resolution optical imagery
Jindou Zhang, Yuyan Yan, Zhizheng Zhang 0009, Boshen Chang, Yunong Chen, Qingwei Zhuang, DeRen Li
Expert Syst. Appl.9
2025 SparseFormer: A Credible Dual-CNN Expert-Guided Transformer for Remote Sensing Image Segmentation With Sparse Point Annotation
abstract
Although significant advances have been made in the semantic segmentation of high-resolution remote sensing (RS) images, obtaining accurate pixelwise annotations remains resource-intensive. We propose SparseFormer, a credible dual-convolutional neural network (CNN) expert-guided Transformer model designed for semantic segmentation using point-level annotations to reduce this annotation burden. SparseFormer comprises three branches, where two CNN branches employ different attention mechanisms to encourage diverse outputs. To enhance the local consistency of pseudolabels, we introduce a pixel-adaptive refinement (PAR) module that dynamically refines CNN output probabilities by incorporating image information during training. A credible assessment is then performed to combine the CNN outputs, producing high-quality pseudolabels that supervise the CNN-Transformer hybrid branch. This hybrid branch integrates global representations with local features, achieving precise segmentation. To further strengthen the CNN branches, we introduce a knowledge distillation strategy that steadily feeds back information from the hybrid branch to CNN branches, mitigating overfitting risks caused by sparse supervision. SparseFormer employs credible assessment to reduce pseudolabel uncertainty, followed by continuous interaction and dynamic information enhancement among the three branches in an end-to-end training process. Extensive experiments on two benchmark datasets demonstrate that SparseFormer significantly outperforms state-of-the-art methods. Our code is available at:https://github.com/Yujia73/SparseFormer.
Hao Cui 0002, Guo Zhang 0001, Zhigang Xie, Haifeng Li 0007, DeRen Li
IEEE Trans. Geosci. Remote. Sens.7
2025 RTO-LLI: Robust Real-Time Image Orientation Method With Rapid Multilevel Matching and Third-Times Optimizations for Low-Overlap Large-Format UAV Images
abstract
UAV real-time photogrammetry is important to promote the rapid generation of photogrammetry 4D product, intelligent information extraction and rapid remote sensing mapping, and efficient large-scale 3D modeling. However, for real-time processing of low-overlap large-format image sequence, there remains two challenges: (1) Large-format images result in greater data volume and computational load, posing challenges for real-time online processing on regular-performance computing units, requiring more efficient algorithms; (2) Low-overlap images make it difficult for matching correspondences to cover the entire overlapping area at real-time, leading to significant challenges for real-time and robust relative orientation. Therefore, this paper proposes a robust Real-Time Orientation method for Low-overlap Large-format UAV Images (RTO-LLI), which can robustly handle these kind of data in real-time. Firstly, robust initialization method for real-time processing of low-overlap large-format images was designed to ensure a high-success-rate of SLAM initialization. Secondly, constant velocity hypothesis tracking enables fast orientation during constant-speed flight. Thirdly, when the second step false, using real-time pose estimation method based on multilevel matching and coarse-to-fine optimization to robustly solve the precision image pose. Fourthly, final (third-level) pose optimization method based on the IRLS algorithm with suitable search area, which can compute higher-precision image pose in real-time. Finally, real-time mapping based on parallel processing for low-overlap images can generate high-precision 3D point maps and complete feature extraction for the next frame in real-time. Experiments conducted on several different types of scenes show that: (1) the processing speed of RTO-LLI significantly surpasses traditional offline methods: PhotoScan, OpenMVG, Colmap. RTO-LLI can handle large-format UAV image sequence (single-imagery has 20-million-pixels) at a speed of 1.5 frames-per-second, meeting the demands of real-time UAV photogrammetry tasks; (2) RTO-LLI is the only method that has successfully completed real-time tasks in all 50-times repeated experiments for four different types of scenes, demonstrating robustness far superior to other classical SLAM solutions; (3) the-displacement-error of the estimated Pose by RTO-LLI is less than 1/2000 of the-trajectory-length, and the average-reprojection-error is less than 1.5 pixels, almost as well as traditional offline methods. RTO-LLI method meets the efficiency, robustness and accuracy requirements of real-time photogrammetry for low-overlap large-format UAV images.
Xiongwu Xiao, Gui-Song Xia, Jianya Gong, DeRen Li
IEEE Trans. Geosci. Remote. Sens.6
2025 UMIS-YOLO: Underwater Multimodal Images Instance Segmentation With YOLO
abstract
Underwater instance segmentation plays a pivotal role in various applications. Among them, coral instance segmentation is of great significance in the fields of marine biology and environmental monitoring, and is crucial for comprehensive understanding of coral reef ecosystems. Traditional methods for underwater instance segmentation predominantly rely on RGB images. However, the complex morphology of corals and strong background interference often result in poor segmentation outcomes. To tackle these problems, this study presents a novel multimodal instance segmentation method, termed UMIS-YOLO, which is grounded in the YOLO architecture. UMIS-YOLO incorporates a dual backbone network design that substantially enhances the feature extraction capabilities for both RGB images and depth images, thereby improving the effectiveness of instance segmentation. At the same time, we propose two innovative plug-and-play modules: the Frequency Domain Feature Enhancement Fusion (FDFEF) module and the Residual Feature Fusion (RFF) module. The FDFEF module leverages Fourier transform to enhance the features of both modalities in the frequency domain, employing learnable weights to enable the complementary integration of amplitude and phase information. While the RFF module utilizes a residual learning strategy to efficiently merge low-level and high-level features prior to the segmentation head, thereby improving pixel-level segmentation accuracy. Additionally, we introduce a challenging high-resolution dataset, UMIS-Coral, which comprises RGB images and depth images captured in complex coral environments. Meanwhile, we expand the depth images for the UIIS dataset to further verify the effectiveness of UMIS-YOLO. The experimental results indicate that the UMIS-YOLO model achieved mAP50 and mAP75 improvements of 2.3 and 3.0 on the UMIS-Coral dataset, as well as 3.9 and 2.8 on the UIIS dataset, respectively. Furthermore, the model is characterized by its lightweight architecture and rapid segmentation capabilities. The source code and the dataset are publicly accessible at https://github.com/zhangsanhulk/UMIS-YOLO.
Yue Yang 0051, Xiaoyi Feng, Ming Li 0037, Xiangyun Hu, Jiangying Qin, Armin Gruen, DeRen Li, Jianya Gong
IEEE Trans. Geosci. Remote. Sens.7
2024 Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
abstract
Recently, there has been a surge in interest in Large Language Models (LLMs), with ChatGPT standing out for its exceptional capabilities in language comprehension, reasoning, and interactive communication. These models have garnered attention from a diverse array of users and researchers across various disciplines. While LLMs have demonstrated remarkable proficiency in mimicking human task execution through natural language, their application in remote sensing interpretation remains largely uncharted. Furthermore, the current lack of automation in remote sensing task planning limits the accessibility of these sophisticated interpretation techniques, especially for non-specialists in the field. To bridge this gap, we introduce Remote Sensing ChatGPT, an innovative LLM-driven agent that integrates ChatGPT with a suite of AI-powered remote sensing models to tackle complex interpretation challenges. This system is designed to interpret user requests, delineate task planning based on the functionalities required, execute each subtask sequentially, and compile the final output by synthesizing the results from each stage. Given that LLMs, trained predominantly on natural language, do not inherently comprehend visual elements present in remote sensing imagery, we have devised a method to incorporate visual cues, effectively embedding the visual context of remote sensing images into the ChatGPT framework. With Remote Sensing ChatGPT, users can effortlessly submit a remote sensing image alongside their query and promptly receive detailed interpretation outcomes along with comprehensive linguistic feedback. Experiments and case studies demonstrate that our method is adept at handling a diverse range of remote sensing tasks and has the potential to be expanded to encompass an even wider array of applications with the integration of more advanced models, such as remote sensing foundation model. The code and demo of Remote Sensing ChatGPT is publicly available at https://github.com/HaonanGuo/Remote-Sensing-ChatGPT.
Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001, DeRen Li
IGARSS6
2024 Assessment of Sun Glint Correction Methods in Unmanned Aerial Vehicle-Based Ocean Optical Remote Sensing
abstract
The sun glint poses a significant challenge in optical remote sensing of the ocean using unmanned aerial vehicles (UAVs). It contaminates the oceanographic information within the images, not only obscuring underwater seafloor features but also increasing the radiance reflectance within the images, which is detrimental for the study and monitoring of the marine environment. Currently, there are two main approaches for sun glint correction in optical imagers (in this context, referring to visible light images captured in RGB channels): methods based on sea surface statistical model and methods based on deep learning. To assess the potential application of sun glint correction methods on UAV RGB imagery, this paper reviews and summarizes both types of methods. Through experimental evaluation, our examines the usability of these methods in UAV-based ocean remote sensing, providing a scientific foundation for obtaining high-quality marine monitoring images. The experimental results indicate that traditional sea surface statistical model-based methods developed for satellite imagery face challenges when transitioning to high-resolution UAV imagery. Conversely, deep learning-based methods show promise for sun glint correction in UAV RGB imagery. Although these methods are still in the early exploration stage, they hold the potential to offer innovative and efficient solutions for mitigating sun glint effects in high-resolution UAV RGB imagery.
Jiangying Qin, Ming Li 0037, Armin Gruen, DeRen Li, Jianya Gong, Xuan Liao
IGARSS4
2024 ANNet: Asymmetric Nested Network for Real-Time Cloud Detection in Remote Sensing
abstract
Cloud detection is one of the crucial tasks in the field of remote sensing, which is regarded as binary classification at the pixel level. Although a large number of recent deep learning-based methods have made great progress, most of their appealing performances come at the expense of a large amount of computation, which reduces the real-time performance accordingly. In order to bridge the gap between segmentation performance and inference speed, we propose a novel asymmetric nested network (ANNet) architecture termed ANNet, which is designed for real-time cloud detection with excellent performance. In the encoder branch of ANNet, we introduce an effective tiny U-shape block (TUB) to enrich detailed spatial contexts in each stage, which also allows ANNet to embed spatial recovery ability earlier, and a lightweight, simple feature fusing module (SFFM) is designed to refine the semantic features map at a lower level of TUB for better performance. Following the encoder, which is diverse from the most symmetric U-shape approaches, an asymmetric and lightweight decoder (ALD) with only convolution and bilinear up-sample operations is employed for spatial recovery. We also, moreover, demonstrated that using a constant channel size instead of a larger channel volume as the network goes deeper is an efficient and effective design for cloud detection tasks. Substantial experiments are performed on GF1_WHU and 95-Cloud datasets, which show that ANNet has achieved excellent performance with low computation cost compared to most of the existing state-of-the-art methods. On GF1_WHU dataset, ANNet-l achieves 93.79% on mIoU at 125 FPS, while ANNet-s with only 29.5 K parameters yields 90.74% mIoU at 251 FPS on the Nvidia RTX 2080Ti.
Niming Fan, DeRen Li, Jun Pan 0001, Shangren Huang
IEEE Trans. Geosci. Remote. Sens.2
2024 A Novel LOD Rendering Method With Multilevel Structure-Keeping Mesh Simplification and Fast Texture Alignment for Realistic 3-D Models
abstract
Fast, high-precision texture maps and high-frame-rate level of detail (LOD) generation for realistic 3-D models are foundational data infrastructures for smart cities. However, LOD generation faces three main issues: reduced model accuracy from mesh simplification, inefficient texture memory utilization, and browsing lag with detail loss. This article proposed a novel LOD rendering method with multilevel structure-keeping mesh simplification and fast texture alignment for realistic 3-D models. First, a multilevel structure-keeping mesh simplification method with mesh segmentation and vertex classification was used to generate a simplified mesh with high-precision structure preservation. Second, a fast texture alignment method was proposed that uses segmentation information and least-squares conformal map (LSCM) parameterization to acquire texture blocks. The method integrated integral images and a precise multitemplate strategy to align texture blocks, to obtain texture maps with high completeness and high occupancy. Finally, by integrating these methods, an LOD generation method with fast multilevel pyramid construction and adaptive tree organization is proposed. This method achieved high-precision multilevel structure keeping, along with a high-occupancy rate of texture maps, facilitating a high frame rate for LOD model construction. Compared with quadratic error function (QEF), quadratic error metrics (QEMs), low-poly, computational geometry algorithms library (CGAL), and Nvdiffrec, the proposed mesh simplification algorithm demonstrated average accuracy improvements of 12.1%, 24.2%, 57.1%, 17.9%, and 3.2%, respectively. Compared with open multi-view environment (OpenMVE), ContextCapture, and Xatlas, the proposed texture alignment algorithm achieved average occupancy improvements of 27.74%, 11.89%, and 4.80%, respectively. Compared with the state-of-the-art ContextCapture and Smart3D, the proposed method for browsing large-scale realistic 3-D models increased the frame rate by 17.3% and 14.4%, respectively.
Yingwei Ge, Xiongwu Xiao, Bingxuan Guo, Jianya Gong, DeRen Li
IEEE Trans. Geosci. Remote. Sens.6
2024 CGGLNet: Semantic Segmentation Network for Remote Sensing Images Based on Category-Guided Global-Local Feature Interaction
abstract
As spatial resolution increases, the information conveyed by remote sensing images becomes more and more complex. Large-scale variation and highly discrete distribution of objects greatly increase the challenge of the semantic segmentation task for remote sensing images. Mainstream approaches usually use implicit attention mechanisms or Transformer modules to achieve global context for good results. However, these approaches fail to explicitly extract intra-object consistency and inter-object saliency features leading to unclear boundaries and incomplete structures. In this paper, we propose a Category-Guided Global-Local Feature Interaction Network (CGGLNet), which utilizes category information to guide the modeling of global contextual information. To better acquire global information, we proposed a Category-Guided Supervised Transformer module (CGSTM). This module guides the modeling of global contextual information by estimating the potential class information of pixels so that features of the same class are more aggregated and those of different classes are more easily distinguished. To enhance the representation of local detailed features of multi-scale objects, we designed the Adaptive Local Feature Extraction Module (ALFEM). By parallel connection of the CGSTM and the ALFEM, our network can extract rich global and local context information contained in the image. Meanwhile, the designed Feature Refinement Segmentation Head (FRSH) helps to reduce the semantic difference between deep and shallow features and realizes the full integration of different levels of information. Extensive ablation and comparison experiments on two public remote sensing datasets (ISPRS Vaihingen dataset and ISPRS Potsdam dataset) indicate that our proposed CGGLNet achieves superior performance compared to the state-of-the-art methods.
Yue Ni, Weijian Chi, DeRen Li
IEEE Trans. Geosci. Remote. Sens.5
2024 A Hyperparameter-Free Attention Module Based on Feature Map Mathematical Calculation for Remote-Sensing Image Scene Classification
abstract
Remote-sensing scene classification (RSSC) is crucial for remote-sensing image interpretation and has become a research hotspot in recent years. However, the high complexity of remote-sensing scenes causes most RSSC models to fail to accurately capture key objects, resulting in low classification accuracy. Meanwhile, it is intractable to effectively distinguish similar scenes, such as forest and meadow, whose semantic labels are mainly determined by wide-scale features. In addition, existing remote-sensing attention mechanisms are heuristic settings, which require expert knowledge and extensive experiments. To solve the above problems, a novel plug-and-play hyperparameter-free attention module (HFAM) based on feature map mathematical calculation is proposed in this work. HFAM uses statistical indicators to quantitatively characterize the fluctuations of feature maps that can accurately locate key features and distinguish different scenes, alleviating the problems of intraclass diversity and interclass similarity. Moreover, HFAM adaptively acquires attention weights by performing simple mathematical calculations on the feature maps, which solves the problem of difficult adjustment of hyperparameters. Our proposed HFAM can be expediently inserted into the existing ConvNet models without increasing the number of model’s parameters. Extensive contrast experiments with several famous plug-and-play attention modules on three mainstream datasets reveal the superiority of our HFAM in accuracy, number of parameters, and calculation amount. Moreover, compared with state-of-the-art methods, it also demonstrated considerable competitiveness.
Qiao Wan, Zhifeng Xiao, Zhenqi Liu, Kai Wang 0080, DeRen Li
IEEE Trans. Geosci. Remote. Sens.6
2024 Global Focal Learning for Semi-Supervised Oriented Object Detection
abstract
Oriented object detectors have achieved great success in aerial detection tasks with the help of ample labeled data. Unlabeled images are easier and less expensive to obtain than labeled aerial images. Therefore, semi-supervised oriented object detection (SSOOD) is becoming a hot task, which can leverage both labeled and unlabeled data to train oriented detectors. Most SSOOD approaches focus on well-designed approaches to generate high-quality pseudo labels (PLs) or positive learning regions, which are limited to complex and variable aerial scenes. This study first analyzes key factors influencing the performance of SSOOD and proposes a global focal learning method (termed as focal teacher) without artificial priori design. It relies on global region and soft regression approaches to blur the boundaries between positive and negative samples, mainly through localization focal loss to achieve. It leverages the localization consistency between the teacher and student model to focus more on hard regions. Moreover, we organize a large remote sensing unlabeled (RSUL) dataset to exploit the performance potential of oriented detectors on mainstream aerial detection datasets (DOTA and DIOR). Adequate experiments reveal that the proposed method achieves the best performance compared with other mainstream SSOOD methods, including partly, fully, and additional data settings on DOTA and DIOR datasets. Semi-supervised mechanisms without preset learning regions can be better applied in dense and complex aerial scenes.
Kai Wang 0080, Zhifeng Xiao, Qiao Wan, Fanfan Xia, Pin Chen, DeRen Li
IEEE Trans. Geosci. Remote. Sens.6
2023 AFDE-Net: Building Change Detection Using Attention-Based Feature Differential Enhancement for Satellite Imagery
abstract
Building Change Detection (BCD) from satellite imagery is critical for monitoring urbanization, managing agricultural land, and updating geospatial databases. However, complex variations in building roofs that resemble the background of their surroundings pose challenges for deep learning-based change detection methods due to their focus on color and texture. Additionally, downsampling can result in the loss of spatial information, leading to incomplete buildings and irregular output boundaries. To address these challenges, a novel Siamese network called AFDE-Net is proposed, which combines differential image features and attention modules using a learnable parameter. The AFDE-Net employs an ensemble spatial-channel attention fusion (ESCAF) module, along with a deep supervision module, to mitigate the loss of spatial information and refine deep features in high-dimensional inputs. Besides, we have created a new dataset (EGY-BCD) comprising high-resolution and multi-temporal satellite images captured in four urban and coastal areas in Egypt to detect building changes. The EGY-BCD dataset includes images with complex types of change, such as tall and dense buildings with roofs that resemble the background of their surroundings, which is a challenge for deep learning algorithms. The proposed method outperforms other methods on the EGY-BCD dataset with an overall accuracy of 94.3%, an F1-score of 88.8%, and an mIoU of 86.6%. The datasets and codes will be released at https://github.com/oshholail/EGY-BCD.
Shimaa Holail, Tamer Saleh, Xiongwu Xiao, DeRen Li
IEEE Geosci. Remote. Sens. Lett.4
2023 Learnable Loss Balancing in Anchor-Free Oriented Detectors for Aerial Object
abstract
Oriented object detection plays an important role in aerial image interpretation. Image processing speed is also essential due to massive amounts of aerial images. Anchor-free oriented detectors with fast processing speed are generally accepted despite the absence of pre-set anchors, contributing to their performance gap with anchor-based detectors. Most anchor-free oriented detectors are carefully designed by defining samples according to target characteristics, which require substantial prior knowledge, to realize improved performance. This study proposes an anchor-free oriented detector (termed as rfpoint) that requires minimal prior knowledge. Moreover, this study mainly aims to introduce a dynamic sample definition strategy. This strategy is modeled as a dynamic regulating process, wherein the classification and box regression interact until the model converges. A rotating quality-driven loss (RQDL) and adaptive-weight box loss (AWBL) are also proposed to realize the aforementioned process. RQDL redefines positive and negative attributes of samples according to the distribution of rotating Intersection-over-Unit (IoU) between predictions and ground truth. AWBL adjusts the importance degree of candidate samples in the box regression through classification scores. The proposed method is then tested on three mainstream aerial image datasets (DOTA, DIOR, and HRSC2016). Results reveal that the proposed method achieves the best performance compared with other oriented detectors, whose mAP are 79.92%, 70.88%, and 90.67%. Moreover, the method maintains the inference speed advantage of anchor-free detectors. An effective sample definition method can bridge the performance gap of anchor-free oriented detectors without minimizing inference speed.
Kai Wang 0080, Zhifeng Xiao, Qiao Wan, Xiaowei Tan, DeRen Li
IEEE Trans. Geosci. Remote. Sens.5
2023 AERNet: An Attention-Guided Edge Refinement Network and a Dataset for Remote Sensing Building Change Detection
abstract
Advancements in Earth observation technology enable the detection of surface changes in intricate urban environments. Building change detection (BCD) plays a crucial role in urban planning and environmental monitoring. However, existing deep learning-based BCD algorithms exhibit limited capability in feature extraction, feature relationship comprehension, sample imbalance mitigation, and accurate boundary identification for changed objects. To address these challenges, we introduce an attention-guided edge refinement network (AERNet) that employs a global context feature aggregation module (GCFAM) to aggregate information from extracted multi-layer context features. Our approach incorporates an attention decoding block (ADB) guided by enhanced coordinate attention (ECA) to capture channel and location associations between features. Furthermore, we utilize an edge refinement module (ERM) to enhance the network’s capacity to sense and refine the edges of changed areas. To tackle the issue of class imbalance and augment the algorithm’s feature learning ability, we devise a novel self-adaptive weighted binary cross-entropy (SWBCE) loss function, combined with a deep supervision (DS) strategy. Experiments are conducted on two publicly available datasets, GDSCD and LEVIR-CD, as well as our newly developed high-resolution complex urban scene BCD dataset, i.e., HRCUS-CD. The latter dataset comprises 11,388 pairs of images at 0.5-meter resolution and over 12,000 labeled change buildings. Comparative experiments indicate that AERNet surpasses advanced competitive methods, while ablation experiments demonstrate the effectiveness of AERNet’s model components and the SWBCE loss function. Efficiency comparison confirms that AERNet achieves comprehensive detection performance with superior effectiveness and robustness.
Jindou Zhang, Xiao Huang 0003, Yu Wang 0140, Xuechao Zhou, DeRen Li
IEEE Trans. Geosci. Remote. Sens.7
2022 Semantic Change Detection Based on a New Chinese Satellite Dataset and a Deep Conditional Random Field Framework
abstract
In this paper, a new semantic change detection (CD) dataset based on Chinese Gaofen-2 (GF-2) satellite images with high spatial resolution (HSR) namely Wuhan Urban Semantic Understanding (WUSU) dataset is built up and a CD framework combining binary and semantic CD tasks based on deep learning and a conditional random field model (SDCRF) is proposed. Existing CD datasets mostly focus on “change/no change”. Traditional CD methods pay attention only on either of the binary CD task or the semantic CD task. Although there are methods to handle both tasks simultaneously but they ignore the inconsistency between the two tasks. In the SDCRF framework, any state-of-the-art feature extraction model can be used to extract the class and change probabilities as the unary potential of a fully connected conditional random field (FC-CRF) model which is adopted as a post-processing to enhance the location information of deep networks and reduce outlier noise.
Sunan Shi, Yanfei Zhong, Yinhe Liu, Jue Wang 0011, DeRen Li
IGARSS5
2022 SAR-DRDNet: A SAR image despeckling network with detail recovery
Wenfu Wu, Xiao Huang 0003, Jiahua Teng, DeRen Li
Neurocomputing5
2022 Adaptive dense pyramid network for object detection in UAV imagery
Ruiqian Zhang, Xiao Huang 0003, Jiaming Wang 0001, Yufeng Wang 0004, DeRen Li
Neurocomputing6
2022 MDANet: Unsupervised, Mixed-Domain Adaptation for Semantic Segmentation of Remote Sensing Images
abstract
The imaging process of optical remote sensing images are easily affected by external conditions. Therefore, remote sensing images under different imaging conditions often show color differences, resulting in feature distribution differences between the source and target domain, hindering the migration of semantic segmentation models between domains. Currently, most domain adaptation methods are for single-source and single-target domains. Here, we proposed a novel and concise method, coined MDANet, for the adaptation of patch images of multi-source and multi-target domains and for reducing the distribution differences of different patch images by projecting them onto the virtual center of a mixed-domain. MDANet is a lightweight and self-supervised network that can be grafted with any semantic segmentation model. Our method significantly improved the segmentation accuracy of semantic segmentation models and showed higher stability and competitiveness than existing methods.
Hao Cui 0002, Guo Zhang 0001, Ji Qi 0001, Haifeng Li 0007, Chao Tao 0001, Shasha Hou, DeRen Li
IEEE Geosci. Remote. Sens. Lett.8
2022 Population, GDP, and Carbon Emissions as Revealed by SNPP-VIIRS Nighttime Light Data in China With Different Scales
abstract
Satellite-based artificial nighttime brightness observations are typically considered proxy measures of socioeconomic indicators at large scales, such as population, gross domestic product (GDP), and carbon emissions. However, few studies have explored and compared the correlations between SNPP-VIIRS nighttime light data and socioeconomic indicators from administrative scale to grid scale, and further analyzed the potential mechanisms for the dissimilar correlations at different grid scales. Using regression model, dissimilarity index, and relief amplitude, the quantitative relationship and potential influence mechanism across different scales was investigated in this letter. Results show that the finer the scale is, the lower the correlations between total nighttime lights (NTL) and socioeconomic indicators when comparing 1 km, town, and county scales. The R2values of the NTL-socioeconomic indicator correlations increase sharply with the increase of grid scale at 1–10 km scale. The R2values increase volatilely between 10–30 km but are relatively stable above 30 km. The differences in R2values may be attributed to the diversity and distribution balance of industrial types and relief amplitude at different scales. This letter provides new insights into estimating and predicting population, GDP, and carbon emissions by using SNPP-VIIRS data.
Kaifang Shi, Yizhen Wu, DeRen Li, Xi Li 0016
IEEE Geosci. Remote. Sens. Lett.3
2022 A Spectral-Spatial-Dependent Global Learning Framework for Insufficient and Imbalanced Hyperspectral Image Classification
abstract
Deep learning techniques have been widely applied to hyperspectral image (HSI) classification and have achieved great success. However, the deep neural network model has a large parameter space and requires a large number of labeled data. Deep learning methods for HSI classification usually follow a patchwise learning framework. Recently, a fast patch-free global learning (FPGA) architecture was proposed for HSI classification according to global spatial context information. However, FPGA has difficulty in extracting the most discriminative features when the sample data are imbalanced. In this article, a spectral-spatial-dependent global learning (SSDGL) framework based on the global convolutional long short-term memory (GCL) and global joint attention mechanism (GJAM) is proposed for insufficient and imbalanced HSI classification. In SSDGL, the hierarchically balanced (H-B) sampling strategy and the weighted softmax loss are proposed to address the imbalanced sample problem. To effectively distinguish similar spectral characteristics of land cover types, the GCL module is introduced to extract the long short-term dependency of spectral features. To learn the most discriminative feature representations, the GJAM module is proposed to extract attention areas. The experimental results obtained with three public HSI datasets show that the SSDGL has powerful performance in insufficient and imbalanced sample problems and is superior to other state-of-the-art methods.
Qiqi Zhu, Weihuan Deng, Zhuo Zheng, Yanfei Zhong, Qingfeng Guan 0001, Weihua Lin, Liangpei Zhang 0001, DeRen Li
IEEE Trans. Cybern.8
2022 Oil Spill Contextual and Boundary-Supervised Detection Network Based on Marine SAR Images
abstract
Oil spills have caused serious harm to the marine environment. Remote sensing technology is one of the important tools for marine environment monitoring. Synthetic aperture radar (SAR) has become an important technology for detecting marine pollution. Identifying dark spots is essential for oil spill detection based on SAR images. Dark spots’ detection can be achieved using image segmentation techniques. However, natural phenomena, such as waves and currents, can also cause dark spots, resulting in consistently uneven intensity, high noise, and blurred boundaries in oil spill images. In addition, existing oil spill detection models often perform well for large targets but have poor detection accuracy for small targets. To solve the above problems, the oil spill contextual and boundary-supervised detection network (CBD-Net) is proposed to extract refined oil spill regions by fusing multiscale features. To improve the internal consistency of oil spill regions, the spatial and channel squeeze excitation (scSE) block is introduced. In CBD-Net, boundary details are enhanced with optimized edge supervision. In addition, a manually labeled dataset is proposed, Deep-SAR Oil Spill (SOS) dataset, aiming to solve the problem of insufficient existing oil spill detection dataset. Experimental results demonstrate that CBD-Net outperforms other comparative models and is able to extract robust and accurate oil spill regions from complex SAR images. The highest mIoU of 83.42% and the highest F1 score of 87.87% were achieved on the SOS dataset. The CBD-Net model proposed in this article can play a guiding role in the marine oil spill decision support system.
Qiqi Zhu, Xiaorui Yan, Qingfeng Guan 0001, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li
IEEE Trans. Geosci. Remote. Sens.8
2022 Real-Time and Accurate UAV Pedestrian Detection for Social Distancing Monitoring in COVID-19 Pandemic
abstract
Coronavirus Disease 2019 (COVID-19) is a highly infectious virus that has created a health crisis for people all over the world. Social distancing has proved to be an effective non-pharmaceutical measure to slow down the spread of COVID-19. As unmanned aerial vehicle (UAV) is a flexible mobile platform, it is a promising option to use UAV for social distance monitoring. Therefore, we propose a lightweight pedestrian detection network to accurately detect pedestrians by human head detection in real-time and then calculate the social distancing between pedestrians on UAV images. In particular, our network follows the PeleeNet as backbone and further incorporates the multi-scale features and spatial attention to enhance the features of small objects, like human heads. The experimental results on Merge-Head dataset show that our method achieves 92.22% AP (average precision) and 76 FPS (frames per second), outperforming YOLOv3 models and SSD models and enabling real-time detection in actual applications. The ablation experiments also indicate that multi-scale feature and spatial attention significantly contribute the performance of pedestrian detection. The test results on UAV-Head dataset show that our method can also achieve high precision pedestrian detection on UAV images with 88.5% AP and 75 FPS. In addition, we have conducted a precision calibration test to obtain the transformation matrix from images (vertical images and tilted images) to real-world coordinate. Based on the accurate pedestrian detection and the transformation matrix, the social distancing monitoring between individuals is reliably achieved.
Gui Cheng, Jiayi Ma 0001, Zhongyuan Wang 0001, Jiaming Wang 0001, DeRen Li
IEEE Trans. Multim.6
2021 Toward Dataset Construction for Remote Sensing Image Interpretation
abstract
With the rapid advancement of remote sensing (RS) technology, RS image interpretation has made great progress and been widely used in broad applications, in which the constructed benchmark datasets for developing and testing intelligent interpretation algorithms have been playing an increasingly critical role. Motivated by the essential prerequisites of dataset in the development of RS image interpretation algorithms, this manuscript provides a discussion on dataset construction for RS image interpretation. Specifically, we first analyze the current challenges of developing algorithms for RS image interpretation and a review on the widespread RS image datasets is conducted through the bibliometric analysis. We then propose some principles and discuss the methodology on constructing benchmark datasets. An implementation on creating the RS scene classification dataset demonstrates the practicability of our proposed framework and the experimental results show that our constructed dataset can serve as a promising benchmark for RS image scene interpretation.
Yang Long 0002, Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001, DeRen Li
IGARSS5
2021 Bias Compensation Model for Sensor Orientation Under Weak Conditions
abstract
The high-precision geometric positioning of satellite images is the basis for the geometric processing of remote-sensing images and acquisition of various geospatial information. It is an important premise for the wide application of high-resolution remote-sensing satellite images. To correct systematic biases in rational function models (RFMs), many compensation methods for system errors inherent of RFMs have been proposed. Thus far, the bias compensation model (BCM) is the most widely accepted method under rigorous conditions, namely narrow camera field, small off-nadir angle, and small attitude error. However, research studies on the compensation effect of the BCM under weak conditions (wide camera field, large off-nadir angle, or large attitude error) are still lacking, and some researchers commented that the BCM is inapplicable under weak conditions in the absence of experiments. This letter analyzes the effect of position and attitude errors on orientation accuracy in the image space. An experiment was conducted using data from the Gaofen-1 (GF-1) wide-field-view-4 (WFV-4) sensor to compare the proposed analysis with the traditional analysis, and the results obtained using the BCM under weak conditions were found to be consistent with our analysis rather than the traditional analysis. This confirms that the BCM can also be used under weak conditions.
Kai Xu 0008, Guo Zhang 0001, Peng Jia 0006, Xiaoyun Hao, DeRen Li
IEEE Geosci. Remote. Sens. Lett.5
2020 Urban Scenes Change Detection Based on Multi-Scale Irregular Bag of Visual Features for High Spatial Resolution Imagery
abstract
Remote sensing scene change detection (SCD) is to detect whether and what changes have occurred in the semantic category of corresponding scenes for a long time at the semantic level. This can provide detailed land use/land cover change information for Urban planning and environmental monitoring. Previous studies take regular patches divided by uniform grid sampling as scene units. This may lead to mosaic phenomenon, and use fixed Window to extract features, ignoring the multi-scale features of ground objects, while extracting scene features. To solve the problems, the multi-scale irregular bag of visual features (MIBVF) framework is proposed for high spatial resolution (HSR) imagery SCD. In this paper, we integrate image classification of the physical characteristics from remote sensing data with the socio-economic attributes from open source geographic data. Road network data is used to preserve the geological significance and semantic integrity of urban scenes, and multi-scale window sampling is used to solve the problem of different object sizes. To confirm the feasibility of the proposed method, experiments with multi-temporal images of the Pudong area in Shanghai indicate that the proposed method achieves a clearly higher change detection accuracy than current state-of-the-art methods.
Jiale Chen 0002, Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li
IGARSS6
2020 Semi-Automatic Fully Sparse Semantic Modeling Framework for Hyperspectral Unmixing
abstract
In order to improve the accuracy of surface classification and meet the needs of sub-pixel-level target detection, spectral unmixing has been one of the hot spots in hyperspectral remote sensing research. The employment of the probabilistic topic model to acquire latent topics of hyperspectral image has been an effective way for spectral unmixing. However, this approach fails to consider the sparsity of the semantic representation and high computational complexity. In addition, the number of endmembers cannot be determined automatically. To solve the problem, in this paper, the novel spectral unmixing method based on semi-automatic fully sparse semantic modeling framework (SFSSM) is proposed. In SFSSM, modestly few arithmetic operations are required to identify the pure spectral signatures (endmembers) and the fractional abundances of the endmembers. Meanwhile, the sparsity and representativeness of the topics generated by SFSSM guarantee that the endmembers can be obtained automatically in low time consumption. The experimental results obtained with two real image confirm that the proposed method significantly improves the performance when compared with the other methods.
Qiqi Zhu, Wen Zeng 0003, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li
IGARSS7
2020 A Modified D-Linknet with Transfer Learning for Road Extraction from High-Resolution Remote Sensing
abstract
Road extraction, which aims to label remote images with a specific road detection, is a fundamental task for understanding remote sensing imagery. Deep learning has strong characteristic learning ability, for example, the state-of-the-art D-Linknet is an effective way to capture the road information. However, existing regularization methods either do not match the performance for large batches, or still exhibit degradation in performance for smaller batches. Besides, the roads in different areas have various characteristics and lack a good transfer. To remedy these issues, we proposed a novel road extraction network which integrated the filter response normalization (FRN) layer with D-Linknet (FND-Linknet). The FRN layer is effective and robust for road extraction task, and can eliminate the dependency on other batch samples. In addition, the multisource road dataset is collected and annotated to improve features transfer. Experimental results on three datasets verify that the proposed FND-Linknet framework outperforms the state-of-the-art methods both in accuracy and connectivity.
Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li
IGARSS6
2020 Super Resolution Generative Adversarial Network Based Image Augmentation for Scene Classification of Remote Sensing Images
abstract
High spatial resolution remote sensing image (RSI) scene classification, aimed at automatically labelling images with the given semantic categories, has been a hot issue. As it's difficult for RSI to quickly obtain a large number of training samples from a specific area. Traditional scene classification researches were mainly using deep learning models to transfer natural images to RSI. Considering the differences between natural images and RSI, we trained several Super Resolution GAN models by using different resolution RSI data from Google earth image. This paper proposed a novel SRGAN-CNN framework. Through transferring the data with scene classification dataset to obtain high resolution fake RSI. The experimental results demonstrate that the proposed framework can enhance transfer effect and help improve the accuracy of scene classification using low resolution RSI.
Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li
IGARSS6
2020 Topic Model for Remote Sensing Data: A Comprehensive Review
abstract
From text analysis to image interpretation, the topic model (TM) always plays an important role. With its powerful semantic mining capabilities, it is able to capture the latent spectral and spatial information from remote sensing (RS) images. Recent years have witnessed widespread use of TM to solve the problems in RS image interpretation, i.e., semantic segmentation, target detection, and scene classification. However, there has not yet been a study expatiating and summarizing the current situation of RS applications with TM. This paper intends to systematically summarize the application of TM in RS images and to conduct several typical experiments for comparison. Specifically, the architecture of our work can be explained as follows: 1) the theory of TM; 2) the applications of RS based on TM; 3) experimental analysis of typical TM methods to provide reference for further understanding, and 4) summary and prospects for guiding further research into TM for RS data.
Qiqi Zhu, Jiangqin Wan, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li
IGARSS6
2020 Video anomaly detection and localization via Gaussian Mixture Fully Convolutional Variational Autoencoder
Yaxiang Fan, GongJian Wen, DeRen Li, Shaohua Qiu, Martin D. Levine
Comput. Vis. Image Underst.3
2020 Using Radar Signatures to Classify Bird Flight Modes Between Flapping and Gliding
abstract
In this work, we find that radar signatures registered by wingbeats can work as a chronophotograph in radar version to record bird flight modes. When a bird flies, its wings and body compose a corner reflector, and the incident radiation upon either face of this wingbeat corner reflector impinges onto the other face and is reflected toward the illuminator, which enhances the bird signal intensity by 0-10 dB. When a bird flaps its wings, the corner works at a right angle, and the wingbeat corner reflector strongly modulates the radar signal; therefore, the contribution is strong and stably approaches 10 dB. During gliding, the wingbeat corner reflector fades away, the modulation effect significantly decreases, and the contribution approaches 0 dB. This difference can assist in classifying bird flight modes between flapping and gliding over a long sampling period. Both the Ku-band data from a duck in an anechoic chamber and the Ku-band radar data from a pigeon in an outside environment support the ability of this signature to classify gliding and flapping modes. We present flight mode transitions between flapping and gliding extracted from the radar echoes of small and large oncoming and outgoing birds.
Jiangkun Gong, Jun Yan 0006, DeRen Li, Ruizhi Chen
IEEE Geosci. Remote. Sens. Lett.3
2020 Context-Aware Convolutional Neural Network for Object Detection in VHR Remote Sensing Imagery
abstract
Object detection in very-high-resolution (VHR) remote sensing imagery remains a challenge. Environmental factors, such as illumination intensity and weather, reduce image quality, resulting in poor feature representation and limited detection accuracy. To enrich the feature representation and mine the underlying context information among objects, this article proposes a context-aware convolutional neural network (CA-CNN) model for object detection that includes proposal generation, context feature extraction, feature fusion, and classification. During feature extraction, we propose integrating a context-regions-of-interests (Context-RoIs) mining layer into the CNN model and extracting context features by mapping Context-RoIs mined from the foreground proposals to multilevel feature maps. Finally, the context features extracted from multilevel layers are fused into a single layer, and the proposals represented by the fused features are classified by a softmax classifier. In this article, through numerous experiments, we thoroughly explore the influence of key factors, such as Context-RoIs, different feature scales, and different spatial context window sizes. Because of the end-to-end network design approach, our proposed model simultaneously maintains high efficiency and effectiveness. We conducted all model testing on the public NWPU VHR-10 data set. The experimental results demonstrate that our proposed CA-CNN model achieves significantly improved model performance and better detection results compared with the state-of-the-art methods.
Yiping Gong, Zhifeng Xiao, Xiaowei Tan, Haigang Sui, Haiwang Duan, DeRen Li
IEEE Trans. Geosci. Remote. Sens.7
2018 Early event detection based on dynamic images of surveillance videos
Yaxiang Fan, GongJian Wen, DeRen Li, Shaohua Qiu, Martin D. Levine
J. Vis. Commun. Image Represent.3
2018 Optimal Segmentation of High-Resolution Remote Sensing Image by Combining Superpixels With the Minimum Spanning Tree
abstract
Image segmentation is the foundation of object-based image analysis, and many researchers have sought optimal segmentation results. The initial image oversegmentation and the optimal segmentation scale are two vital factors in high spatial resolution remote sensing image segmentation. With respect to these two issues, a novel image segmentation method combining superpixels with a minimum spanning tree is proposed in this paper. First, the image is oversegmented using a simple linear iterative clustering algorithm to obtain superpixels. Then, the superpixels are clustered by regionalization with a dynamically constrained agglomerative clustering and partitioning (REDCAP) algorithm using the initial number of segments, and the local variance (LV) and the rate of LV change (ROC-LV) indicator diagrams corresponding to the number of segments are obtained. The suitable number of image segments is determined according to the LV and ROC-LV indicator diagrams corresponding to the number of segments. Finally, the superpixels are reclustered using the REDCAP algorithm based on the suitable number of image segments to obtain the image segmentation result. Through two sets of experiments, the proposed method is compared with two other segmentation algorithms. The experimental results show that the proposed method outperforms the others and obtains good image segmentation results.
Mi Wang, Yufeng Cheng, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2018 Scene Classification Based on the Sparse Homogeneous-Heterogeneous Topic Feature Model
abstract
High spatial resolution (HSR) imagery scene classification has been the subject of increased interest in recent years, and has great potential for many applications, such as urban functional analysis. Rooted in natural information processing, the use of the probabilistic topic model (PTM) to capture latent topics to represent HSR images has been an effective way to bridge the semantic gap. However, how to effectively discover discriminative information to recognize the HSR scenes is a challenging task. In this paper, the sparse homogeneous-heterogeneous topic feature model (SHHTFM) is proposed for HSR image scene classification. Differing from the conventional PTM-based scene classification methods, which utilize only heterogeneous features, SHHTFM explores the effect of the homogeneous information. Based on the union of uniform grid sampling and simple linear iterative clustering superpixel sampling, SHHTFM exploits both the heterogeneous and homogeneous information. After separately mining different types of low-level features and latent topics, the sparse topic inference procedure of SHHTFM further improves the fusion of the sparse heterogeneous and homogeneous topics. In addition, multisource geographical data are effectively integrated, where the water and vegetation boundaries define a more accurate way to restrict the boundaries of different scenes, and are then combined with the road network data to further improve the scene annotation performance. This provides more reliable and applicable results for us to better understand the complex scenes. The experimental results obtained with two HSR image classification data sets and an HSR image annotation data set demonstrate that the proposed SHHTFM framework can solve the scene classification problem, with a high classification accuracy as well as a high time efficiency.
Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li
IEEE Trans. Geosci. Remote. Sens.5
2018 Adaptive Deep Sparse Semantic Modeling Framework for High Spatial Resolution Image Scene Classification
abstract
High spatial resolution (HSR) imagery scene classification, which involves labeling an HSR image with a specific semantic class according to the geographical properties, has received increased attention, and many algorithms have been proposed for this task. The employment of the probabilistic topic model to acquire latent topics and the convolutional neural networks (CNNs) to capture deep features for representing HSR images has been an effective ways to bridge the semantic gap. However, the midlevel topic features are usually local and significant, whereas the high-level deep features convey more global and detailed information. In this paper, to discover more discriminative semantics for HSR images, the adaptive deep sparse semantic modeling (ADSSM) framework combining sparse topics and deep features is proposed for HSR image scene classification. In ADSSM, the fully sparse topic model and a CNN are integrated. To exploit the multilevel semantics for HSR scenes, the sparse topic features and deep features are effectively fused at the semantic level. Based on the difference between the sparse topic features and the deep features, an adaptive feature normalization strategy is proposed to improve the fusion of the different features. The experimental results obtained with four HSR image classification data sets confirm that the proposed method significantly improves the performance when compared with the other state-of-the-art methods.
Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2017 Space-based information service in Internet Plus Era
DeRen Li, Xin Shen 0001, Nengcheng Chen, Zhifeng Xiao
Sci. China Inf. Sci.1
2017 Airport Detection Based on a Multiscale Fusion Feature for Optical Remote Sensing Images
abstract
Automatically detecting airports from remote sensing images has attracted significant attention due to its importance in both military and civilian fields. However, the diversity of illumination intensities and contextual information makes this task difficult. Moreover, auxiliary features both within and surrounding the regions of interest are usually ignored. To address these problems, we propose a novel method that uses a multiscale fusion feature to represent the complementary information of each region proposal, which is extracted by constructing a GoogleNet with a light feature module model that has an additional light fully connected layer. Then, the fusion feature is input to a support vector machine whose performance is enhanced using a hard negative mining method. Finally, a simplified localization method is applied to tackle the problem of box redundancy and to optimize the locations of airports. An experiment demonstrates that the fusion feature outperforms other features on airport detection tasks from remote sensing images containing complicated contextual information.
Zhifeng Xiao, Yiping Gong, Yang Long 0002, DeRen Li, Xiaoying Wang 0002
IEEE Geosci. Remote. Sens. Lett.4
2017 Scene Classification Based on the Fully Sparse Semantic Topic Model
abstract
In high spatial resolution (HSR) imagery scene classification, it is a challenging task to recognize the high-level semantics from a large volume of complex HSR images. The probabilistic topic model (PTM), which focuses on modeling topics, has been proposed to bridge the so-called semantic gap. Conventional PTMs usually model the images with a dense semantic representation and, in general, one topic space is generated for all the different features. However, this approach fails to consider the sparsity of the semantic representation, the classification quality, as well as the time consumption. In this paper, to solve the above problems, a fully sparse semantic topic model (FSSTM) framework is proposed for HSR imagery scene classification. FSSTM, with an elaborately designed modeling procedure, is able to represent the image with sparse but representative semantics. Based on this framework, the topic weights of multiple features are exploited by solving a concave maximization problem, which improves the fusion of the discriminative semantic information at the topic level. Meanwhile, the sparsity and representativeness of the topics generated by FSSTM guarantee that the image is adaptive to the change of a topic number. FSSTM can consistently achieve a good performance with a limited number of training samples, and is robust for HSR image scene classification. The experimental results obtained with three different types of HSR image data sets confirm that the proposed algorithm is effective in improving the performance of scene classification, and is highly efficient in discovering the semantics of HSR images when compared with the state-of-the-art PTM methods.
Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2016 Terrain measurements in CHINA using multi-sensor SAR data
abstract
Terrain measurement and surface motion estimation are key applications for SAR missions. These applications drive several SAR satellite missions in the Dragon partner countries in Europe and China. In this context, we present our work in the Dragon program on terrain measurements with results from several test sites in China.
Mingsheng Liao, Lu Zhang 0034, Timo Balz, DeRen Li
IGARSS4
2016 Robust Approach for Recovery of Rigorous Sensor Model Using Rational Function Model
abstract
The replacement of the rational function model (RFM) by the rigorous sensor model (RSM) has been studied extensively and verified with many types of sensors and remote-sensing applications. However, relatively less research has been conducted on recovering RSM from RFM, and the few relating techniques can only be applied in specific circumstances. This paper proposes a novel linear method to obtain the position, attitude, and interior orientation (IO) elements of satellites based on the orientation information of the rays implied by the RFM. Instead of resection, forward intersection is used to solve for position, and an equivalent body coordinate system is introduced to overcome the strong correlation between the attitude and IO. The orientation information of the rays implied by the RFM is used to calculate the IO pixel by pixel. Experiments using the Ziyuan 3 panchromatic nadir sensor show that this method can recover the exterior orientation and IO elements effectively.
Wen-chao Huang, Guo Zhang 0001, DeRen Li
IEEE Trans. Geosci. Remote. Sens.3
2015 Big data in smart cities
DeRen Li, Jianjun Cao
Sci. China Inf. Sci.1
2015 MOC-Based Parallel Preprocessing of ZY-3 Satellite Images
abstract
The launch of the ZY-3 surveying and mapping satellite (ZSMS) by China has resulted in a significant increase in the volume of image data collected for subsequent processing. In this letter, we present our research on the message passing interface (MPI), open multiprocessing (OpenMP), and compute unified device architecture (CUDA)-based (MOC-based) preprocessing of ZSMS images in a system that consists of multiple central processing units (CPUs) and graphics processing units (GPUs). First, CPUs and GPUs in the system are organized into atomic computing resources (ACRs) by means of an MPI. Then, three cooperative methods are proposed for potential performance improvement of the processors with OpenMP and CUDA. The input/output (I/O) overhead is also addressed in this letter. The experimental results show that the total execution time of the 12 ZSMS nadir images with four ACRs is reduced to 86.10 s, which could provide near-real-time response for the time-critical applications that follow.
Liuyang Fang, Mi Wang, DeRen Li, Jun Pan 0001
IEEE Geosci. Remote. Sens. Lett.3
2015 Measuring the Effectiveness of Various Features for Thematic Information Extraction From Very High Resolution Remote Sensing Imagery
abstract
Generally, some object-based features are more relevant to a thematic class than other features. These strongly relevant features, termed as class-specific features, would significantly contribute to thematic information extraction for very high resolution (VHR) images. However, many existing feature selection methods have been designed to select a good feature subset for all classes, rather than an independent feature subset for the thematic class. The latter might better meet the requirement of thematic information extraction than the former. In addition, the lack of quantitative evaluation of the contribution of the selected features to thematic classes also weakens our understandability of these features. To address the problems, class-specific feature selection methods are developed to measure the effectiveness of features for extracting thematic information from VHR images. First, the one-versus-all scheme is combined with traditional feature selection methods, such as ReliefF and LeastC. Also, one-versus-one scheme is utilized for alleviating the negative impact of a class imbalance problem arising from the one-versus-all scheme. Then, the relative contributions of features to thematic classes are obtained by the class-specific feature selection methods to describe the effectiveness of features for thematic information extraction. Finally, the class-specific feature selection methods are compared with the original methods on three different VHR image data sets by the nearest neighbor and support vector machine. Experimental results show that the class-specific feature selection methods outperform the corresponding conventional methods, and the one-versus-one scheme surpasses one-versus-all scheme. Additionally, many features are evaluated by the class-specific feature selection methods, to provide end users advice on effectiveness of the features.
Xi Chen 0004, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2015 Intensity Correction of Terrestrial Laser Scanning Data by Estimating Laser Transmission Function
abstract
Intensity information, recorded by laser scanning, shows great potential for research and applications (e.g., object recognition and registration of images to 3-D models). Multiple studies not only show the significance and possibility of correcting the intensity value but also highlight the existing problems, i.e., that the near-distance and angle-of-incidence effects prevent the practical correction of a terrestrial laser scanner (TLS). In this paper, we explore the near-distance corrections for Z+F Imager5006i, a commercially available coaxial TLS, and propose a corresponding method to correct the range or distance effects on its output intensity data. This method estimates the parameters of customized range-intensity equations using a novel sample data collection design. Using the angle-of-incidence correction method, practical intensity corrections were conducted on TLS point cloud data from white walls and Mogao Grottoes, Dunhuang, China. The results were visualized and show the effectiveness of the proposed method. In addition, we analyzed the new problems that emerge after correction and outline further studies.
Fan Zhang 0006, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2015 Systematic Error Compensation Based on a Rational Function Model for Ziyuan1-02C
abstract
A rational function model (RFM) can be used directly to convert the relationships between image coordinates and object space coordinates without using any physical imaging parameters (such as satellite position and attitude). Thus, RFMs facilitate versatility and high security during geometric processing of optical satellite imagery. Increasingly, RFMs are offered to users as the basic geolocation model for further geometric processing by imagery vendors. However, imagery vendors might perform inadequate in-orbit geometric calibrations, or the calibrated geometric parameters might not be updated in a timely manner. Thus, the RFMs may suffer from high distortion due mainly to interior errors (such as lens distortion). Using the radiometric correction products of Ziyuan1-02C panchromatic and multispectral sensor as examples, the present study addresses the compensation of systematic errors in RFMs. An undistorted RFM can be generated after calibrating the interior error compensation model once, before high-accuracy registration between the panchromatic imagery and multispectral imagery can be achieved using the undistorted RFM. Experimental evaluations based on the positioning accuracy using a few ground control points (GCPs) with an undistorted RFM matched the accuracy of the GCPs. In addition, our approach greatly improves the accuracy of registration (which surpasses 0.7 panchromatic pixels) between panchromatic and multispectral imagery.
Yonghua Jiang 0001, Guo Zhang 0001, DeRen Li, Xinming Tang, Wen-chao Huang
IEEE Trans. Geosci. Remote. Sens.4
2015 Adaptive-Window Polarimetric SAR Image Speckle Filtering Based on a Homogeneity Measurement
abstract
This paper proposes a polarimetric homogeneity measurement and applies it to the speckle filtering of polarimetric synthetic aperture radar (PolSAR) data. First, a line-and-edge (LAE) detector that can detect both the lines and edges in one scan is developed based on the traditional edge detector. A polarimetric homogeneity measurement is then derived by combining the equivalent number of looks and the LAE maps and is used to distinguish the homogeneous and heterogeneous regions. Finally, a new adaptive-window PolSAR filtering algorithm based on the LAE detector and the polarimetric homogeneity measurement is proposed. The proposed speckle filter adjusts the filtering windows in both shape and size, based on the homogeneity and gradient information. Consequently, it uses small and nonsquare windows in heterogeneous regions to preserve the detail information and uses large and square windows in homogeneous regions to maximize the suppression of speckle noise. EMISAR and ESAR L-band PolSAR data were used to demonstrate the effectiveness of the proposed filter in speckle suppression, detail preservation, and polarimetric information preservation.
Fengkai Lang, Jie Yang 0040, DeRen Li
IEEE Trans. Geosci. Remote. Sens.3
2015 Block Adjustment for Satellite Imagery Based on the Strip Constraint
abstract
Given that long strip satellite images have the same error distribution characteristics, we propose a block adjustment method for satellite images based on the strip constraint. First, the image point coordinates are calculated in the strip image coordinate system based on the offset value of the adjacent image. Second, the rational function model (RFM) of the strip image is regenerated using the RFM of single images, and the compensation grid is also generated. Third, block adjustment of the strip image is implemented based on the RFM with an affine transformation parameter. Finally, the affine transformation parameters of single images are recalculated using the affine transformation parameters of the strip image. Experiments using ZY-3 satellite images showed that block adjustment of satellite images based on a strip constraint (strip adjustment) can produce better results than block adjustment of satellite images based on a single image in sparse control conditions. The test results demonstrated the effectiveness and feasibility of the proposed method.
Guo Zhang 0001, Taoyang Wang, DeRen Li, Xinming Tang, Yonghua Jiang 0001, Wen-chao Huang
IEEE Trans. Geosci. Remote. Sens.3
2014 Polarimetric SAR Image Segmentation Using Statistical Region Merging
abstract
The statistical region merging (SRM) algorithm exhibits efficient performance in solving significant noise corruption and does not depend on the data distribution. These advantages make SRM suitable for the segmentation of synthetic aperture radar (SAR) images, which are characterized by speckle noise and different distributions of various data types and spatial resolutions. However, the original SRM algorithm is designed for RGB and gray images characterized by additive noise and having a range of [0, 255]. In this letter, the SRM algorithm is generalized so that it can be applied to images with larger range and multiplicative noise. The original 4-neighborhood models are also generalized into 8-neighborhood models. The effectiveness of the generalized SRM (GSRM) algorithm is demonstrated by AirSAR and ESAR L-band Polarimetric SAR (PolSAR) data. Given that the input data of the GSRM algorithm can be single- or multi-dimensional, the proposed GSRM algorithm can be used for single- and multi-polarized as well as for fully polarimetric SAR data.
Fengkai Lang, Jie Yang 0040, DeRen Li, Lingli Zhao, Lei Shi 0005
IEEE Geosci. Remote. Sens. Lett.3
2014 Stream Model-Based Orthorectification in a GPU Cluster Environment
abstract
One of the most important tasks in remote sensing data processing is the production of orthorectified images. Such tasks are computationally intensive and can become a bottleneck for remote sensing image processing, particularly in high-throughput environments, such as large satellite imagery processing centers. This letter explores the use of massive parallel processing graphical processing unit (GPU) in a clustered network environment to speed up image processing tasks, such as orthorectification. Our parallelization method is based on inverse sensor model and the stream model for image processing, which allow the flexibility of placing computational units on proper computation units, such as GPU, CPU cores, or nodes in a cluster. In our experiments on images of two satellites, more than 198 times and 50.3 times speedup over one and multiple thread CPU versions have been achieved, respectively.
Zhen Lei 0004, Mi Wang, DeRen Li, Ting L. Lei
IEEE Geosci. Remote. Sens. Lett.3
2014 Geometric Accuracy Validation for ZY-3 Satellite Imagery
abstract
The ZiYuan-3 surveying satellite (ZY-3) is a high-precision civilian satellite imaging sensor. Since its launch on January 9, 2012, it has been in operation for one and a half years. Although the initial postlaunch ZY-3 geometric accuracy was verified during an in-orbit operation period, on-orbit calibration was still necessary from time to time. This on-orbit calibration has vastly improved the location accuracy in planimetry for ZY-3 panchromatic images. This letter briefly describes the principle of on-orbit calibration and production processes of sensor-corrected products. Furthermore, block adjustment based on a rational function model test showed planimetric and vertical accuracy values of 10 m and 5 m, respectively, without ground control points (GCPs). The accuracy values improved to 3 m and 2 m, respectively, with a few GCPs. The statistics results are from ten different regions with independent checkpoints (ICPs). All accuracy values are the root-mean-square error of ICPs. Therefore, ZY-3 can be used for the generation of cartographic maps at the 1 : 50 000 scale and for revision and updates of 1 : 25 000 scale maps. Compared with other mainstream high-resolution satellite images of the same ground resolution, ZY-3's geometric accuracy is almost the same and sometimes even better.
Taoyang Wang, Guo Zhang 0001, DeRen Li, Xinming Tang, Yonghua Jiang 0001, Xiaoyong Zhu
IEEE Geosci. Remote. Sens. Lett.3
2014 Detection and Correction of Relative Attitude Errors for ZY1-02C
abstract
Ziyuan1-02C (ZY1-02C) was launched on December 22, 2011, and it is the first civilian high-resolution remote sensing satellite in China. However, the limited precision of the onboard attitude measurement system causes many errors during attitude transfer by ZY1-02C. Thus, there are complex distortions in the images obtained by ZY1-02C, which restricts its application greatly. In this paper, we consider the feasibility of attitude error correction based on parallel observations with high-resolution cameras, and the method is described in detail. To validate the efficiency of the proposed method, several images and corresponding control data were collected from the Henan, Taihang Mountain, Neimeng, and Taiyuan areas in China. The experimental results indicate that seamless mosaic images without distortion can be obtained using our method. Furthermore, the positioning accuracy with a few ground control points (GCPs) was shown to be better than 1.5 pixels and equivalent to the accuracy of the GCPs.
Yonghua Jiang 0001, Guo Zhang 0001, Xinming Tang, DeRen Li, Wen-chao Huang
IEEE Trans. Geosci. Remote. Sens.4
2014 Geometric Calibration and Accuracy Assessment of ZiYuan-3 Multispectral Images
abstract
The ZiYuan-3 (ZY-3) remote sensing satellite is China's first civilian high-resolution stereo mapping satellite. Because the interior orientation parameters measured before launch are biased, the multispectral (four-band) images collected by ZY-3 exhibit low-accuracy band-to-band registration, which affects their subsequent applications. This paper presents a valid method for interior orientation determination of the ZY-3 multispectral sensor by determining the look angles of the charge-coupled device arrays for all bands. One band is chosen as the benchmark band, and its interior orientation is determined using the relevant ZY-3 image collected over the calibration field and the corresponding digital orthoimage map and digital elevation model. The remaining bands are then calibrated using the benchmark band as control data. The quality of the calibration is further enhanced by shortening the calibration period and by combining images collected over different calibration fields, which decreases the negative effects of errors in the satellite's attitude and position data. The interior orientation of the multispectral sensor in ZY-3 was determined using data sets taken over two calibration fields, namely, Dengfeng (Henan Province) and Tianjin. Evaluation experiments were performed using ZY-3 multispectral images and ground control points (GCPs) collected over several different periods and areas. The positioning accuracy of the ZY-3 multispectral images with a limited number of GCPs after calibration of the interior orientation was better than 0.3 pixels, and the band-to-band registration accuracy was up to 0.15 pixels.
Yonghua Jiang 0001, Guo Zhang 0001, Xinming Tang, DeRen Li, Wen-chao Huang
IEEE Trans. Geosci. Remote. Sens.4
2014 Mean-Shift-Based Speckle Filtering of Polarimetric SAR Data
abstract
The mean shift algorithm, which uses a moving window and utilizes both spatial and range information contained in an image, is widely employed in digital image filtering and segmentation. However, because of the large dynamic range of synthetic aperture radar (SAR) images, applying the conventional mean shift algorithm directly to SAR image filtering will not produce meaningful results. This paper proposes an adaptive variable asymmetric bandwidth selection approach to be used in a newly derived generalized mean shift algorithm. The proposed mean shift algorithm is very versatile and can be used for SAR and polarimetric SAR (PolSAR) image filtering directly without any preprocessing steps. Monte Carlo-simulated PolSAR data are used to demonstrate the effectiveness of the proposed algorithm in speckle filtering by comparing it with other filters. Experimental Synthetic Aperture Radar (ESAR) L-band and Radarsat-2 C-band PolSAR data are used to evaluate its ability to preserve the polarimetric information of PolSAR data. The effects of initial value estimating and multilook processing on the filtered results are discussed at the end of this paper.
Fengkai Lang, Jie Yang 0040, DeRen Li, Lei Shi 0005, Jujie Wei
IEEE Trans. Geosci. Remote. Sens.3
2014 Bi-Temporal Texton Forest for Land Cover Transition Detection on Remotely Sensed Imagery
abstract
With the advancement of machine learning, classification methods have been increasingly used in change (or transition) detection. The texton forest (TF)-based method has received increasing research attention because of its speed, good generalization characteristics, stability, and especially its ability to capture spatial contextual information. In this paper, we propose a TF-based method for transition detection in remotely sensed imagery. We investigate a maximal joint-information gain criterion for random forests to better capture combined information in the bi-temporal images in transition detection, which is implemented by a natural extension of binary-trees in traditional methods into a quad-decision tree structure. We also utilize color-invariant gradient as a feature to help alleviate the impact of difference in imaging conditions on bi-temporal transition detection. The experimental results for transition detection show that our bi-temporal TF classifier achieves better performance than a post-classification comparison method and several other alternative methods.
Zhen Lei 0004, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2013 Image-based self-position and orientation method for moving platform
DeRen Li, Xiuxiao Yuan
Sci. China Inf. Sci.1
2013 Non-rigid registration of mural images and laser scanning data based on the optimization of the edges of interest
Fan Zhang 0006, DeRen Li
Sci. China Inf. Sci.4
2012 Spatial data quality and beyond
abstract
Issues of accuracy, uncertainty, and spatial data quality have been on the top of most GIScience research agendas around the world from the late 1980s. Ever since then, growing research efforts have been directed toward uncertainty characterization in spatial information, analysis, and applications, aiming for better understanding of spatial uncertainty and thus improved methods and techniques for assessing and managing data quality. Impressive progress has been made in various issues concerning data quality. In addition, growing research on extensions to the conventional norms of data quality, such as the quality aspects of geospatial information services, has been observed. Chinese researchers have contributed to this great cause by keeping abreast with the developments abroad and striving for their own innovative work. This paper reviews the past research on data quality-related issues and provides a perspective on future developments. These will be seen not only in continued research on theoretical and technical issues concerning data quality, but also in developments of tools for quality assessment and decision-making under uncertainty through geospatial information processing and applications.
DeRen Li, Jingxiong Zhang, Huayi Wu
Int. J. Geogr. Inf. Sci.1
2012 Rotation-Invariant Object Detection of Remotely Sensed Images Based on Texton Forest and Hough Voting
abstract
The Hough forest method is an effective method for object detection in ground-shot images that has received increasing research attention. However, this method lacks the ability to detect objects with arbitrary orientations. This largely constrains the method from being used in detecting geospatial objects from remotely sensed (RS) images since geospatial objects can have many different orientations. In order to achieve rotation invariance and compensate the associated loss of discriminative power, this paper presents a novel color-enhanced rotation-invariant Hough forest (CRIHF) method for detecting geospatial objects in RS images. In our method, we propose to train a Pose-Estimation-based Rotation-invariant Texton Forest (PE-RTF) which first uses dominant gradient orientations to align local image patches. The orientations are then jointly used with coordinates in Hough voting to detect object position. In order to increase discriminative power, Texton Forest is used in codebook generation. Moreover, theoretically sound color-invariant gradients are employed. By rotating split functions rather than image patches in the RTF and sparsely accumulating Hough votes on grid points, computational times can be reduced by two orders of magnitude. The evaluation of the CRIHF method on a data set containing 525 airplanes and a second data set containing 68 residential buildings shows that our method is rotation invariant and robust. The detector achieves around 90% recall rate on both data sets. Experiments also show that our method is noise resistant and can achieve a decent detection performance at a high level (30%) of “salt and pepper” impulsive noise.
Zhen Lei 0004, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2011 Image City sharing platform and its typical applications
DeRen Li
Sci. China Inf. Sci.2
2011 A multi-scale and multi-orientation image retrieval method based on rotation-invariant texture features
DeRen Li, Xianqiang Zhu
Sci. China Inf. Sci.2
2011 Fault tolerant spatio-temporal fusion for moving vehicle classification in wireless sensor networks
abstract
Wireless sensor networks (WSNs) can implement complicated tasks through collaboration among multiple sensor nodes. The low-cost sensors in WSNs often generate noisy and even faulty measurements, which will degrade the network performance. Therefore developing collaborative signal processing (CSP) algorithms that has high fault tolerance ability is necessary for the increasingly deployed WSNs. In this study, the authors propose a novel fault tolerant fusion scheme to implement reliable vehicle classification by integrating fault detection and correction with a spatio-temporal fusion structure. Sensor faults are detected at the fusion centre and then the fault detection results are fed back to the local sensors to update subsequent classification results. A Dempster–Shafer theory based fault correction strategy is devised to utilise the fusion centre feedback. Simulation results demonstrate that the proposed scheme ensures more than 95% of the classification results to be correct when no larger than 30% of the sensors are faulty and the scheme achieves improved fault detection rate and false alarm rate than the optimum threshold Bayesian fault detection scheme.
Chunting Liu, DeRen Li
IET Commun.4
2011 Land Cover Classification for Remote Sensing Imagery Using Conditional Texton Forest With Historical Land Cover Map
abstract
In this letter, we propose a “conditional texton forest” (CTF) method to utilize widely available historical land cover (HLC) maps in land use/cover classification on high-resolution images. The CTF is based on texton forest (TF), which is a popular and powerful method in image semantic segmentation due to its effective use of spatial contextual information, its high accuracy, and its fast speed in multiclass classification. The proposed CTF method nonparametrically aggregates a bank of TFs according to HLC information and uses the fact that different types of HLC follow different transition rules. The performance of CTF is compared to support vector machine (SVM), Markov random field (MRF), and a naive TF method which uses historical data directly as a feature channel. On average, CTF results in a 2%-5% higher classification accuracy than other classifiers in our experiment. The classifying speed of CTF is similar with TF, five times faster than MRF, and hundreds of times faster than SVM. Given the abundance of HLC data, the proposed method can be expected to be useful in a wide range of socioeconomic and environmental studies.
Zhen Lei 0004, DeRen Li
IEEE Geosci. Remote. Sens. Lett.3
2011 Graph-Based Feature Selection for Object-Oriented Classification in VHR Airborne Imagery
abstract
Linearly nonseparability and class imbalance of very high resolution (VHR) imagery make feature selection for object-oriented classification quite challenging, while such characteristics, especially class imbalance, have usually been ignored in open literature. To cope with the challenges, this paper proposes a new graph-based feature selection method named locally weighted discriminating projection (LWDP). First, the popular graph-based criteria of feature selection are reformulated to present linear or nonlinear mapping in feature space. Second, weight matrices of graphs characterize dissimilarity rather than similarity between pairwise neighbors, to well-preserved local structure when the difference of distance between a sample and its neighbors is large. Finally, LWDP provides a new perspective to alleviate class imbalance at both global and local levels, by restricting the pairwise relationships in the weight matrices. Specifically, neighborhood unions are introduced to employ the local class distribution and class size to constrain pairwise relationships in the weight matrices when classifying unbalanced sample sets. To evaluate the performances of LWDP in low dimensions, a holistic scoring scheme is proposed to stress the performances under low dimensions. In addition, overall accuracy curves and Kappa Index of Agreement (KIA) curves, which exhibit KIA in dimensions, are also used. The experimental results show that LWDP and its kernel extension outperform the other classic or latest methods in processing unbalanced sample set of VHR airborne imagery.
Xi Chen 0004, DeRen Li
IEEE Trans. Geosci. Remote. Sens.4
2011 Shadow Detection in Remotely Sensed Images Based on Self-Adaptive Feature Selection
abstract
Shadows in remotely sensed images create difficulties in many applications; thus, they should be effectively detected prior to further processing. This paper presents a novel semiautomatic shadow detection method that meets the requirements of both high accuracy and wide practicability in remote sensing applications. The proposed method uses only the properties derived from the shadow samples to dynamically generate a feature space and calculate decision parameters; then, it employs a series of transformations to separate shadow and nonshadow regions. The proposed method can detect shadows from both color and gray images. If the chromatic properties of color images do not agree with the defined rules through the shadow samples, then the shadow detection process will automatically reduce to the process for gray images. As the shadow samples are manually selected from the input image by the user, the derived parameters conform well to the characteristics of the input image. Experiments and comparisons indicate that the proposed self-adaptive feature selection algorithm is accurate, effective, and widely applicable to shadow detection in practical applications.
DeRen Li
IEEE Trans. Geosci. Remote. Sens.3
2010 Semisupervised Feature Selection for Unbalanced Sample Sets of VHR Images
abstract
A semisupervised feature selection method, named asymmetrically local discriminant selection (ALDS), is proposed to evaluate the class separability of unbalanced sample sets from very high resolution (VHR) imagery in an object-oriented classification. In order to cope with class imbalance, ALDS incorporates asymmetric misclassification costs of classes into weight matrices. Furthermore, this method locally exploits multiple kinds of relationships between sample pairs to more accurately assess the ability of features in preserving the geometrical and discriminant structures. The experimental results on VHR satellite and airborne imagery attest to the effectiveness and practicability of ALDS.
Xi Chen 0004, DeRen Li
IEEE Geosci. Remote. Sens. Lett.4
2010 A Network-Based Radiometric Equalization Approach for Digital Aerial Orthoimages
abstract
Digital aerial orthoimages have been widely used in surveying, mapping, geographic information systems, visualization, and other applications. However, when producing digital aerial orthoimages, radiometric equalization over large areas is often a most time-consuming and costly process and has become a bottleneck. This letter presents a network-based radiometric equalization approach to eliminate the radiometric differences between images. The network is constructed using the area Voronoi diagrams with overlap and is based on the topological relationship of the constructed network; transferring paths between images are determined, and a global-to-local strategy is used to improve the algorithm, both in its global and local performance. Digital aerial orthoimages from both film-based and digital cameras are used to evaluate the performance of the presented algorithm.
Jun Pan 0001, Mi Wang, DeRen Li, Junli Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2010 Object Classification of Aerial Images With Bag-of-Visual Words
abstract
This letter presents a Bag-of-Visual Words (BOV) representation for object-based classification in land-use/cover mapping of high spatial resolution aerial photograph. The method is introduced to handle the special characteristics of aerial images, i.e., variability of spectral and spatial content. Specifically, patch detection and description are used to divide and represent various subregions of objects comprising multiple homogeneous components. Moreover, the BOV representation is constructed with the statistics of the occurrence of visual words, which are learned from the training data set. A combination of spectral and texture features is verified to be a satisfactory choice through the evaluations of various patch descriptors. Furthermore, a threshold-based method is employed to reduce the impact of outliers on classification in test data. Experiments based on aerial-image data set show that the proposed BOV representation yields better classification performance than the low-level features, such as the spectral and texture features.
DeRen Li
IEEE Geosci. Remote. Sens. Lett.3
2010 An Edge Embedded Marker-Based Watershed Algorithm for High Spatial Resolution Remote Sensing Image Segmentation
abstract
This correspondence proposes an edge embedded marker-based watershed algorithm for high spatial resolution remote sensing image segmentation. Two improvement techniques are proposed for the two key steps of maker extraction and pixel labeling, respectively, to make it more effective and efficient for high spatial resolution image segmentation. Moreover, the edge information, detected by the edge detector embedded with confidence, is used to direct the two key steps for detecting objects with weak boundary and improving the positional accuracy of the objects boundary. Experiments on different images show that the proposed method has a good generality in producing good segmentation results. It performs well both in retaining the weak boundary and reducing the undesired over-segmentation.
DeRen Li, Guifeng Zhang, Zhaocong Wu, Lina Yi
IEEE Trans. Image Process.1
2009 Spatio-temporal fusion for reliable moving vehicle classification in wireless sensor networks
abstract
One of the important tasks in sensor networks is classifying moving vehicles. Fusion of large amount of sensor measurements can improve network performance and reduce the consumption of sensor network resource. We study using continuous measurements of multiple sensor nodes to improve the classification performance by spatio-temporal fusion and fault detection. Time series decisions of single sensor node are aggregated to make a reliable classification estimation. A fusion center combines local classification decisions and evaluates the correctness of these decisions. A correctness status is sent back to each sensor node. Based on the status, sensor nodes can adjust their temporal fusion result. Simulation results demonstrate the validity of our method.
Chunting Liu, DeRen Li
SMC4
2009 The new era for geo-information
DeRen Li
Sci. China Ser. F Inf. Sci.1
2009 Repair approach for DMC images based on hierarchical location using edge curve
Jun Pan 0001, Mi Wang, DeRen Li, TianTian Feng
Sci. China Ser. F Inf. Sci.3
2009 Automatic Generation of Seamline Network Using Area Voronoi Diagrams With Overlap
abstract
The mosaicking of orthoimages has been used to cover a large geographic region for various applications ranging from environmental monitoring to disaster management. However, existing mosaicking methods mainly focus on the generation of seamlines between two adjacent orthoimages. In this paper, we present a novel approach based on the use of a seamline network formed by a novel area Voronoi diagrams with overlap and the use of effective mosaic polygons (EMPs) to define the pixels of each orthoimage for the final mosaic. The generated seamline network is global based and is also optimized after refinement. It gives an effective partitioning for the regions of all orthoimages to form EMPs. The partitioning is unique, seamless, and has no redundancy. The algorithm is parallel, and the EMP of each orthoimage only has relation to orthoimages which have overlaps with it. It can ensure the flexibility and efficiency of mosaicking, without an intermediate process and independent of the sequence of the image composite. The experimental results obtained from the mosaicking of 40 color orthoimages demonstrate considerable potential for generating a seamline network automatically and effectively. This is extremely useful when a seamless mosaic is required to cover a large geographic region.
Jun Pan 0001, Mi Wang, DeRen Li, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2008 Laser Intensity Used in Classification of Lidar Point Cloud Data
abstract
LIDAR (LIght Detection And Ranging) is a powerful remote sensing technology for the acquisition of terrain surface. The LIDAR system not only generates the 3D points cloud with irregular spacing, but also detects the laser impulse reflection data. The algorithms used for the LIDAR data are mostly used to deal with the 3D points cloud, produce the digital terrain model (DTM) and detect the objects, such as buildings. Usually, these objects must be classified as part of the extraction. In order to classify, other information besides the height information of points cloud is required, such as laser intensity information. However, few classification algorithms using intensity data have been deeply investigated. The laser intensity is different from material to material. The intensity of reflection on the same material is similar, while pulsed on different material is differ. Based on this theory, this paper provides a classification algorithm of airborne laser scanning altimetry data combined with the intensity of laser in detail. This paper proposes a classification algorithm for LiDAR points data by fusing height data and intensity data. Initially, the height data is used to generate the original terrain data after filtering. As a result, the terrain data and objects data are separated. Second, a filter algorithm is applied to the raw laser intensity of object points. Third, a histogram of filtered intensity is obtained. Fourth, the statistic result combined with the height information according to the pulse feature of laser intensity, is analyzed.
Liping Di, DeRen Li
IGARSS (2)4
2007 Fast detecting and locating groups of targets in high-resolution SAR images
Gui Gao, Gangyao Kuang, DeRen Li
Pattern Recognit.4
2007 An automatic method for generating affine moment invariants
DeRen Li, Wenbing Tao
Pattern Recognit. Lett.2
2007 Reconstruction of DEMs From ERS-1/2 Tandem Data in Mountainous Area Facilitated by SRTM Data
abstract
A new approach is presented in this paper to produce Digital Elevation Model (DEM) in mountainous areas with steep slope using ERS-1/2 tandem data. In order to reduce the impact of phase errors on the Interferometric Synthetic Aperture Radar (InSAR)-generated DEM, an external DEM such as that from Shuttle Radar Topography Mission (SRTM) is utilized in this approach. The proposed algorithm includes two steps: The first step is to model and remove phase trends with a linear regression analysis before converting phase to height; the second step is to filter unreliable height points before interpolating the DEM from the InSAR height map. The critical points are the following: 1) determining the one-to-one correspondence between the interferogram and the SRTM DEM before knowing the InSAR-derived elevation values and 2) estimating the elevation range of every pixel from SRTM DEM. To solve the first problem, an iteratively geocoding algorithm is performed. A DEM interpolation error model solves the second one. For InSAR data processing, the SRTM DEM is not only usable for modeling systematic phase errors but also for filtering gross height errors. The experiments in Zhangbei and the Three Gorges areas in China show that our approach has improved the accuracy of the resulting DEMs significantly without any ground control points.
Mingsheng Liao, Teng Wang 0001, Lijun Lu, Wenjun Zhouzhou, DeRen Li
IEEE Trans. Geosci. Remote. Sens.5
2005 Partially Supervised Classification - Based on Weighted Unlabeled Samples Support Vector Machine
Zhigang Liu 0012, Wenzhong Shi, DeRen Li, Qianqing Qin
ADMA3
2005 Spatial Information Multi-grid for Data Mining
DeRen Li
ADMA2
2005 A comparison of support vector machine with maximum likelihood classification algorithms on texture features
Shuying Jin, DeRen Li
IGARSS2
2005 Large volume spatial data management based on grid computing
abstract
As the emerging technology, grid computing is applied to many projects. The management of spatial data also faces new challenges. This paper introduces the application of grid computing to the management of spatial data, and emphasizes the deployment and design of a spatial data grid.
DeRen Li, Xinyan Zhu
IGARSS2
2005 A comparative analysis of image fusion methods
abstract
There are many image fusion methods that can be used to produce high-resolution multispectral images from a high-resolution panchromatic image and low-resolution multispectral images. Starting from the physical principle of image formation, this paper presents a comprehensive framework, the general image fusion (GIF) method, which makes it possible to categorize, compare, and evaluate the existing image fusion methods. Using the GIF method, it is shown that the pixel values of the high-resolution multispectral images are determined by the corresponding pixel values of the low-resolution panchromatic image, the approximation of the high-resolution panchromatic image at the low-resolution level. Many of the existing image fusion methods, including, but not limited to, intensity-hue-saturation, Brovey transform, principal component analysis, high-pass filtering, high-pass modulation, the a/spl grave/ trous algorithm-based wavelet transform, and multiresolution analysis-based intensity modulation (MRAIM), are evaluated and found to be particular cases of the GIF method. The performance of each image fusion method is theoretically analyzed based on how the corresponding low-resolution panchromatic image is computed and how the modulation coefficients are set. An experiment based on IKONOS images shows that there is consistency between the theoretical analysis and the experimental results and that the MRAIM method synthesizes the images closest to those the corresponding multisensors would observe at the high-resolution level.
Zhijun Wang 0003, Djemel Ziou, Costas Armenakis, DeRen Li, Qingquan Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2004 From digital map to spatial information multi-grid
abstract
Stalling from the thinking about the representation of geo-spatial data in computer network, the challenge of grid computing environment to geo-spatial information science and technology is pointed in This work. After review and summation of the achieved progress and existing problems of geo-spatial information systems in the past 20 years, the authors propose a new representation method for spatial data and spatial information - the spatial information multi-grid (SIMG), which can not only easily run under grid computing environment, but also properly consider the difference of natural and social characteristics in Earth space as well as the different level of economical development in different area. The system structure, data representation, data storage and data access in SIMG are described with emphasis on key techniques. The data conversion and transferring between SIMG and conventional spatial databases are also discussed. The applicability of SIMG in global, national, provincial and local decision-making is briefly indicated.
DeRen Li, Xinyan Zhu, Yixuan Zhu
IGARSS1
2004 A Try for Handling Uncertainties in Spatial Data Mining
Shuliang Wang 0001, Deyi Li, DeRen Li, Hanning Yuan
KES4
2004 Framework design on video coding system for error-prone heterogeneous network
abstract
This paper describes the complete procedure of framework design on a video coding system for error-prone heterogeneous network environment. First of all, system requirements are analyzed and the targets of system design are emphasized on three categories: 1. improve video performance under the constraints of network bandwidth and computational complexity; 2. provide error robust mechanism when packet loss occur; 3. add capability of layered video coding for heterogenous network transmission. Then, ITU-T H.263+ recommendation is introduced, especially about its 16 optional enhanced coding modes: feature, effect and benefit. System design on mode selection is somewhat a trade-off, one hand is improvement on subjective quality of video or other system targets, the other hand is influence on time delay and complexity (computational load, data dependency, ease of implementation, etc.). Five enhanced coding modes are preferred in strong reason for our purpose: Advanced INTRA Coding mode, Deblocking Filter mode, Modified Quantization mode, Slice Structured mode and Temporal, SNR and Spatial Scalability mode. At last, some useful and important system elements beyond H.263+ recommendation are discussed: RTP packetization, error concealment and error tracking. Application shows that all the targets are easily reached by the implementation based on this design.
Weiming Shen 0002, Yanwen Chong, Jinshen Xiao, DeRen Li
VCIP4
2003 A topological 3D reconstruction of complicated buildings and crossroads
abstract
The generation of 3D model for buildings and crossroads presents a challenge due to the complexity of manmade objects and lack of image understanding algorithms. In this paper, a topological strategy for semi-automatic 3D reconstruction of complicated buildings and crossroads from aerial image pair is presented based on a topology-based 3D data model in which data collection and data modeling are merged into an organic whole. Finally, a prototype software system is developed to prove the validity of the presented approach.
DeRen Li, Qimin Cheng
IGARSS2
2000 An algebraic algorithm for point inclusion query
Huayi Wu, Jianya Gong, DeRen Li, Wenzhong Shi
Comput. Graph.3
1998 Mining Association Rules with Linguistic Cloud Models
Deyi Li, Kaichang Di, DeRen Li, Xuemei Shi
PAKDD3