VLDB 2026 Research / reviewers in the wild / expert
Chunping Qiu
dblp:136/3179
· DBLP profile ↗
33ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0002-7109-5559ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-Aware Stereo Matching via Monocular Disparity Distribution Prior and Gradient EnhancementabstractStereo matching recovers 3D scene information based on the correlation between corresponding pixels. Despite impressive progress, existing methods lack sufficient correlation priors in ill-posed regions such as occlusions, detailed and reflective regions. In this paper, we propose Geometry Aware Stereo Matching Network (GEAStereo) to enhance geometric structure perception and address this issue. We adaptively incorporate the Monocular Disparity Distribution Prior into the stereo cost volume, building Mono-Stereo Fusion Volume (MSFV), which effectively captures global geometric structures and rectifies the correlation information in ill-posed regions. Furthermore, we introduce rich detail information from gradient features and construct a Detail-Aware Volume (DAV) by aggregating the group-wise cost volume under the guidance of gradient spatial attention, thus enhancing the correlation modeling in detailed structures. Jointly, MSFV and DAV provide rich correlation priors for disparity iterative optimization. Experimental results show that our method achieves competitive results on the ETH3D and KITTI2015 benchmarks. Compared with the state-of-the-art methods, our method demonstrates stronger performance in zero-shot generalization. Junze Zhang, Luoxi Jing, Yuanyuan Wang 0002, Guoli Yang, Songchang Jin, Chunping Qiu |
AAAI | 7 |
| 2026 | D3HRL: A distributed hierarchical reinforcement learning approach based on causal discovery and spurious correlation detection
Chenran Zhao, Dian-xi Shi, Mengzhu Wang, Jianqiang Xia, Huanhuan Yang, Songchang Jin, Shaowu Yang, Chunping Qiu |
Neural Networks | 8 |
| 2026 | Survey of automated 3D reconstruction from optical satellite imagery
Anzhu Yu, Danyang Hong, Chunping Qiu, Junyi Fan, Song Ji |
Pattern Recognit. | 5 |
| 2025 | Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative ScenariosabstractLearning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a novel model-based MARL method named Contrastive Latent World for Policy Optimization (CLWPO). In CLWPO, we first design a state representation model to facilitate learning in the latent state space. With the support of this model, we construct the latent world and introduce a contrastive variational bound (CVB) to optimize it. Subsequently, we develop a heuristic policy optimization (HPO) scheme, incorporating model-free learning with model-based planning to obtain robust policies that predict future behaviors. In particular, in the planning, we maintain a queue of teammate models and calculate an adaptive rollout length for each agent to support their self-imagination and reduce the model-based return discrepancy. Finally, we conducted extensive experiments in the PettingZoo benchmark, and results show that CLWPO significantly enhances learning efficiency and improves agent performance compared to state-of-the-art MARL methods. Huanhuan Yang, Dian-xi Shi, Songchang Jin, Guojun Xie, Chunping Qiu, Shaowu Yang |
AAAI | 6 |
| 2025 | UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA BlockabstractDepth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event and image data provides significant advantages, yet effective integration remains challenging. Existing CNN-based fusion methods struggle with occlusions and depth disparities due to limited receptive fields, while Transformer-based fusion methods often lack deep modality interaction. To address these issues, we propose UniCT Depth, an event-image fusion method that unifies CNNs and Transformers to model local and global features. We propose the Convolution-compensated ViT Dual SA (CcViT-DA) Block, designed for the encoder, which integrates Context Modeling Self-Attention (CMSA) to capture spatial dependencies and Modal Fusion Self-Attention (MFSA) for effective cross-modal fusion. Furthermore, we design the tailored Detail Compensation Convolution (DCC) Block to improve texture details and enhances edge representations. Extensive experiments show that UniCT Depth outperforms existing image, event, and fusion-based monocular depth estimation methods across key metrics. Luoxi Jing, Dian-xi Shi, Zhe Liu 0029, Songchang Jin, Chunping Qiu, Ziteng Qiao, Jianqiang Xia |
IJCAI | 5 |
| 2025 | Spatial-Aware Remote Sensing Image Generation From Spatial Relationship DescriptionsabstractRecent advances in stable diffusion models have revolutionized text-to-image generation. However, these models struggle with spatial relationship comprehension in remote sensing (RS) scenarios, limiting their ability to generate spatially accurate imagery. We present a novel framework for generating RS images from spatial relationship descriptions with precise spatial control. Our approach introduces a two-stage pipeline: first, a spatial relationship semantic structuring model converts formalized spatial relationship descriptions into controlled layouts, and second, an enhanced diffusion model incorporates positional prompts and a layout attention mechanism to generate the final image. The positional prompts explicitly encode spatial information, while the layout attention mechanism enables focused region learning. Comprehensive experiments demonstrate that our method achieves superior performance compared with state-of-the-art approaches in both spatial accuracy and image quality. Yaxian Lei, Xiaochong Tong, Chunping Qiu, Haoshuai Song, Congzhou Guo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Rethinking Semantic Segmentation With Multi-Grained Logical PrototypeabstractThe last decade has witnessed significant advances in semantic segmentation brought about by deep learning. However, existing methods only fit the data-label correspondence in a data-driven manner and do not fully conform to the abstraction and structuralization characteristics of the human visual cognition process, which limits the upper bounds of their performance. To this end, a multi-grained logical prototype (MGLP) method is proposed to rethink semantic segmentation based on these two key characteristics. Its novel design can be summarized as follows. 1) For abstraction, prototypes of the same class at different grain levels are established: a label generation method is proposed to automatically generate a multi-grained label space, which can guide the learning of the multi-grained prototypes for each class. 2) For structuralization, the intrinsic logical structure across different semantic levels is explicitly modeled: the horizontal metric relationships are established via metric relation operations on prototypes at the same grain level, to improve the discriminability between classes while taking the vertical semantic hierarchy into account. Moveover, the vertical logical relationships are established as the sub-to-super positive and super-to-sub negative constraints, to strengthen the semantic dependencies among prototypes at different grain levels. 3)MGLP is plug-and-play and can be directly combined with existing segmentation methods. Extensive experimental results indicate that MGLP can significantly improve the segmentation performance of existing methods, which opens up a new avenue for future research. Anzhu Yu, Kuiliang Gao, Xiong You, Yanfei Zhong, Bing Liu 0018, Chunping Qiu |
IEEE Trans. Image Process. | 7 |
| 2024 | ICF-Loc: An Infrared-Based Coarse-to-Fine Approach for UAV Visual Geolocation under GPS-Denied EnvironmentsabstractVisual geolocation plays a crucial role when GPS is unavailable in the Unmanned Aerial Vehicles (UAVs). Many methods rely on visible light cameras, which may not perform well in low-light or foggy conditions. To address this issue, we propose an advanced UAV visual geolocation method called ICF-Loc, which utilizes infrared images in a coarse-to-fine approach. ICF-Loc consists of two stages: a retrieval-based coarse localization stage and a matching-based fine localization stage. The goal of the coarse localization stage is to identify the satellite image that is most similar to the UAV’s infrared image. To bridge the distribution gap between the visible and infrared domains, we propose a feature transfer module. In the fine localization stage, the UAV’s infrared image is matched with the satellite image obtained during coarse localization to estimate the UAV’s position and orientation accurately. We have designed a cross-modal image matching method based on the Fourier transform for precise estimation. The experimental results demonstrate the effectiveness of our proposed approach on both synthetic and real-world datasets. Zhen Wang 0052, Dian-xi Shi, Chunping Qiu, Songchang Jin, Tongyue Li |
ICME | 3 |
| 2024 | JFDI: Joint Feature Differentiation and Interaction for domain adaptive object detection
Ziteng Qiao, Dian-xi Shi, Songchang Jin, Zhen Wang 0052, Chunping Qiu |
Neural Networks | 6 |
| 2024 | Discrete diffusion models with Refined Language-Image Pre-trained representations for remote sensing image captioning
Guannan Leng, Yujie Xiong, Chunping Qiu, Congzhou Guo |
Pattern Recognit. Lett. | 3 |
| 2024 | Sequence Matching for Image-Based UAV-to-Satellite GeolocalizationabstractUAV-to-satellite geolocalization offers accurate drift-free navigation in the absence of external positioning signals. Increased deep-learning-based approaches have demonstrated their potential for high accuracy by framing the problem as a one-to-all retrieval task. However, in real-world scenario, the problem is not just a one-to-all retrieval task, which leads to a gap between research and applications. Based on this observation, We attempt to look closer to the problem instead of designing sophisticated network architectures or objective functions. In this study, we proposed a flexible and simple coarse-to-fine sequence-matching solution with targeted joint use of deep learning and classical machine learning approaches. Our goal is to improve geolocalization accuracy by matching UAV images with a few relevant reference image patches instead of all images. To this end, we first coarsely constructed a sequence of reference satellite image patches corresponding to the UAV trajectory, for which we proposed a deep feature- and manifold learning-based image-sorting method. Once the reference satellite patches are sorted and aligned with the UAV trajectory, the reference sequence is determined. Given a query UAV frame, the search area can be decreased from two to one dimensions. In particular, both deep-learning-based and classical image-matching algorithms can provide competitive accuracy when integrating sequence constraints. We demonstrate that classical manifold learning-based and image matching methods perform exceptionally well for UAV-to-satellite geolocalization when utilized jointly with suitable deep learning techniques. We validated the approach’s unique outperformance on two challenging and realistic UAV-to-satellite geolocalization datasets. Dataset, code and models are available for research purposes at https://seqmatch.geovisuallocalization.com/. Zhen Wang 0052, Dian-xi Shi, Chunping Qiu, Songchang Jin, Tongyue Li, Zhe Liu 0029, Ziteng Qiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MVP: Meta Visual Prompt Tuning for Few-Shot Remote Sensing Image Scene ClassificationabstractVision Transformer (ViT) models have recently emerged as powerful and versatile tools for various visual tasks. In this article, we investigate ViT in a more challenging scenario within the context of few-shot conditions. Recent work has achieved promising results in few-shot image classification by utilizing pre-trained vision transformer models. However, this work employs full fine-tuning for the downstream tasks, leading to significant overfitting and storage issues, especially in the remote sensing domain. In order to tackle these issues, we turn to the recently proposed Parameter-Efficient Tuning (PETuning) methods, which update only the newly added parameters while keeping the pre-trained backbone frozen. Inspired by these methods, we propose the Meta Visual Prompt Tuning (MVP) method. Specifically, we integrate the prompt-tuning-based PETuning method into the meta-learning framework and tailor it for remote sensing datasets, resulting in an efficient framework for Few-Shot Remote Sensing Scene Classification (FS-RSSC). Moreover, we introduce a novel data augmentation scheme that exploits patch embedding recombination to enhance the data diversity and quantity. This scheme is generalizable to any network that employs the ViT architecture as its backbone. Experimental results on the FS-RSSC benchmark demonstrate the superior performance of the proposed MVP over existing methods in various settings, including various-way-various-shot, various-way-one-shot, and cross-domain adaptation. Yiying Li, Naiyang Guan, Zunlin Fan, Chunping Qiu, Xiaodong Yi 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Hybrid Contrastive Prototypical Network for Few-Shot Scene ClassificationabstractFew-shot learning has received widespread attention in remote sensing image scene classification. Many existing methods address this challenge by utilizing meta-learning and metric learning, which focus on developing feature extractors that can quickly adapt to novel few-shot scene classification (FSSC) tasks. However, these methods are often insufficient for real-world datasets with class confusion, where there is high inter-class compactness and intra-class diversity. To overcome this issue, we investigate efficient strategies, i.e., meta-learning-based transferable feature representation and contrastive-based prototypical regularization for learning task-adaptive class boundaries for FSSC. Specifically, we designed a combination of Query-vs-Prototype contrastive loss and Prototype-vs-Prototype contrastive loss to normalize the prototypical representation to be more discriminative in a novel FSSC task. Our proposed model is named the Hybrid Contrastive Prototypical Network (HCP-Net). Experiment results on three popular datasets under two standard benchmarks, i.e., general few-shot classification and few-shot domain generalization, indicate the effectiveness of the proposed method. Chunping Qiu, Mengyuan Dai, Naiyang Guan, Xiaodong Yi 0002 |
ICIP | 3 |
| 2023 | Open Self-Supervised Features for Remote-Sensing Image Scene Classification Using Very Few SamplesabstractBig models, large datasets, and self-supervised learning (SSL) have recently gained substantial research interest due to their potential to alleviate our reliance on annotations. Considering the current high generalization ability of self-supervised models in literature, we explore in the letter how helpful SSL can be for a crucial task in remote sensing (RS), image scene classification, when forced to rely on only a few labeled samples. We proposed a simple prototype-based classification procedure without training and fine-tuning, which uses open self-supervised features from the contrastive language-image pre-training (CLIP). We test our method by exploiting ready-to-use open features on four diversified benchmark datasets, including red-green-blue (RGB) and multispectral (MS) images. Highly competitive accuracy has been obtained compared to work with similar settings, i.e., based on an exceedingly small number of labels. To the best of our knowledge, our model is the first to achieve such high accuracy in austere label conditions. We further analyze our approach from different perspectives, including its advantages and limitations, reasons for its astonishing performance, potential applications, and future improvements. Chunping Qiu, Anzhu Yu, Xiaodong Yi 0002, Naiyang Guan, Dian-xi Shi, Xiaochong Tong |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Learning Visual Representation Clusters for Cross-View Geo-LocationabstractCross-view geo-location is a crucial research field that determines the geographic location from images taken from different viewpoints. It is often studied as a retrieval task, where the query images are with unknown locations, and the database includes images with geo-tags from a different platform. Learning image representations by neural networks is an important step, and one typical training method is using a classification loss, where cross-view images of the same locations are considered the same category. However, existing methods only focus on pushing the representation distances of different categories while ignoring the intra-category representation distances of samples from different platforms. Considering that controlling the intra-category distance can help to guide the model to extract compact category-sharing representations from cross-view images, we propose a categorized cluster loss to learn separate and compact representation clusters. Categorized cluster loss can supervise the network to learn invariant information from samples of different platforms by constraining both the inter-category and intra-category feature distances. Meanwhile, we design a category-view-stratified sampling strategy, which samples balanced inputs in terms of both category and view in each batch during the learning process. We implemented our approach with a lightweight OSNet-based network and achieved higher accuracy with fewer parameters on a typical and challenging cross-view geo-location dataset than most state-of-the-art (SOTA) methods. Haoshuai Song, Zhen Wang 0052, Dian-xi Shi, Xiaochong Tong, Yaxian Lei, Chunping Qiu |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2023 | Learning From Self-Supervised Features for Hashing-Based Remote Sensing Image RetrievalabstractImage retrieval (IR) for practical remote sensing (RS) should have high accuracy, storage, and calculation efficiency, while not relying on big annotations. However, current supervised and unsupervised RSIR methods do not yet fully meet these requirements. To this end, we propose a novel hashing-based IR approach via learning hash codes from open and representative self-supervised features. Specifically, we constructed a model out of a self-supervised pretrained backbone and a small multilayer perceptron (MLP)-based hashing learning neural network. Features from the frozen backbones were used to reconstruct a similarity matrix to guide the hash network learning. This way, the semantic structure can be preserved. To enhance the proposed approach, we propose the exploitation of global high-level semantic information within the similarity reconstruction process by introducing a small set of labeled datasets. Extensive comparative experiments on two commonly used RS image datasets demonstrate the outperformance of our proposed approach and its good balance between the retrieval accuracy and utilized annotations. In these two datasets, the labeled data required by our method accounts for less than 3% of that required by traditional methods, but our obtained mean average precision (mAP) can reach over 90%, which is close to that of current advanced supervised methods. In addition, we analyzed the specific effect of our design and the associated hyperparameters. Dali Wang, Xiaochong Tong, Chunping Qiu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Multi-granularity knowledge distillation and prototype consistency regularization for class-incremental learning
Dian-xi Shi, Ziteng Qiao, Zhen Wang 0052, Shaowu Yang, Chunping Qiu |
Neural Networks | 7 |
| 2023 | Prototype and Context-Enhanced Learning for Unsupervised Domain Adaptation Semantic Segmentation of Remote Sensing ImagesabstractIn unsupervised domain adaptation (UDA) of remote sensing images (RSIs), the huge inter-domain discrepancies and intra-domain variances lead to complicated class-level relations. Specifically, the instances of the same class differ greatly while instances of different classes are similar, whether across different RSIs domains or within the same RSIs domain. However, existing methods cannot fully consider these problems, limiting the performance of UDA semantic segmentation of RSIs. To this end, this paper proposes a novel cross-domain multi-prototypes learning method, the core idea of which is to abstract the cross-and intra-domain class-level relations into multiple prototypes. Specifically, the multiple prototypes belonging to different classes can detailedly describe complex inter-class relations, and the multiple prototypes within the same class can better model rich intra-class relations. Further, the source and target samples are jointly used for prototypes calculation, to fully fuse the feature information of different RSIs. In a nutshell, utilizing the samples from different RSIs domains to learn multiple prototypes for each class can achieve better domain alignment at the class level. In addition, considering that RSIs simultaneously contain large targets with wide coverage and important small targets, two masked consistency learning strategies are designed to better explore the contextual structure of target RSIs and improve the quality of pseudo labels for prototype updating. The global consistency strategy can strengthen the utilization of global context relations, while the local consistency strategy can further improve the learning of local context details. Therefore, the proposed method is actually a prototype and context enhanced learning method for UDA semantic segmentation of RSIs. Extensive experiments demonstrate that the proposed method can achieve better performance than existing state-of-the-art UDA methods. Kuiliang Gao, Anzhu Yu, Xiong You, Chunping Qiu, Bing Liu 0018 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Onboard Data Management Approach Based on a Discrete Grid System for Multi-UAV Cooperative Image LocalizationabstractOnboard image data management and sharing are the foundations for achieving cooperative data processing and analysis in multiple unmanned aerial vehicles (multi-UAVs). However, various challenges, such as the lack of efficient onboard data indices, restrict the development of multi-UAV cooperative applications. Here, we propose a novel and versatile cooperative data management framework based on a discrete grid system for multi-UAV onboard image data. First, we study the image coding methodology employed within the proposed framework. This method transforms original spatiotemporal and attribute information in images into well standardized and structured coded information. Secondly, we introduce a grid-based onboard image data management approach (Grid-OIM) to facilitate cooperative data management among multi-UAVs using code-based index and query methods. Finally, we applied Grid-OIM to cooperative image localization tasks. Experiments were conducted using an edge-computing platform and an embedded database. The image coding method could process > 12,000 images/second while maintaining excellent real-time performance. Moreover, the efficiency of creating and updating the image data index and querying the image data increased by averages of 17.6, 9.3, and 66.1 times, respectively, compared to the image data management method based on R*-Tree, highlighting the substantial advantages of this proposed method. These improvements address the demands of indexing and querying highly dynamic onboard image data effectively. Furthermore, the horizontal accuracy of image localization calculated by the cooperative localization method was improved by 43.4-81.6% compared to that of a single UAV, enhancing reliability. Overall, Grid-OIM presents a feasible and practical solution for multi-UAV cooperative applications. Xiaochong Tong, Chunping Qiu, Yuekun Sun, Congzhou Guo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiscale Feature Learning by Transformer for Building Extraction From Satellite ImagesabstractExtracting buildings from very high-resolution satellite images is a challenging yet important task for applications such as urban monitoring. Multiscale feature learning proves to be a potential solution toward accurate extraction of buildings. This study exploits a powerful multiscale feature learning module, a hierarchical vision transformer by shifted windows (swin), as a backbone within a building extraction network. To this end, we first designed a general structure for building extraction, consisting of a backbone to extract multiscale features and a head network to fuse and refine features. Then, we integrated swin into the structure as a backbone and utilized channel-wise and spatial-wise enhancement in a head network. Experimental results show that our method achieves improvements regarding both F1-score and intersection over union (IoU) compared to the multiple attending path neural network (MAP-Net), which is the current state-of-the-art (SOTA) algorithm for building extraction from remote sensing images. Our study thus confirms the potential of swin transformers as backbones for semantic segmentation tasks based on satellite images. Xin Chen 0088, Chunping Qiu, Wenyue Guo, Anzhu Yu, Xiaochong Tong, Michael Schmitt 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Deep Relearning in the Geospatial Domain for Semantic Remote Sensing Image SegmentationabstractWe present a classification postprocessing (CPP) technique based on fully convolutional neural networks (CNNs) for semantic remote sensing image segmentation. Conventional CPP techniques aim to enhance the classification accuracy by imposing smoothness priors in the image domain. Contrary to that, here, a relearning strategy is proposed where the initial classification outcome of a CNN model is provided to a subsequent CNN model via an extended input space to guide the learning of discriminative feature representations in an end-to-end fashion. This deep relearning CNN (DRCNN) explicitly accounts for the geospatial domain by taking the spatial alignment of preliminary class labels into account. Hereby, we evaluate to learn the DRCNN in a cumulative and noncumulative way, i.e., extending the input space based on all previous or solely preceding model outputs, respectively, during an iterative procedure. Besides, the DRCNN can also be conveniently coupled with alternative CPP techniques such as object-based voting (OBV). The experimental results obtained from two test sites of WorldView-II imagery underline the beneficial performance properties of the DRCNN models. They can increase the accuracies of the initial CNN models on average from 72.64% to 76.01% and from 92.43% to 94.52% in terms of$\kappa $statistic. An additional increase of 1.65 and 2.84 percentage points can be achieved when combining the DRCNN models with an OBV strategy. From an epistemological point of view, our results underline that CNNs can benefit from the consideration of preliminary model outcomes and that conventional CPP techniques can profit from an upstream relearning strategy. Christian Geiß, Yue Zhu 0004, Chunping Qiu, Lichao Mou, Xiao Xiang Zhu 0001, Hannes Taubenböck |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | On-Board Thermal Motion Compensation Method for Pointing Errors of the Remote Sensor Aboard a Three-Axis Stabilized Geostationary SatelliteabstractFor earth observation from a three-axis stabilized geostationary (GEO) remote-sensing satellite, a highly accurate onboard thermal motion compensation (TMC) method for real-time correction of the satellite sensor’s line-of-sight (LOS) pointing error due to thermally induced structural distortion internal to the sensor during the diurnal cycle remains a global concern to date. In this letter, we propose a novel TMC method for GEO sensors. Compared with the traditional TMC methods, this method has the following two advantages. First, we define the LOS misalignment angle to model the comprehensive impact of the sensor’s internal thermal misalignment angles on the sensor’s LOS pointing behavior, which makes the optical path modeling simpler and more generic. Second, we fit the repeatable diurnal variations of the LOS misalignment angle based on long time-series star observations; therefore, the calculation and the use of the TMC amount are no longer limited to the assumption that the sensor’s internal thermal misalignment angles remain constant within a certain period of time, as in the traditional method. The proposed TMC method is theoretically applicable to all types of GEO sensors, including optical and microwave sensors, and is verified by Fengyun-4A advanced geosynchronous radiation imager (FY-4A/AGRI) TMC experiments. Hualong Hu, Xiaochong Tong, Chunping Qiu, Yanfa Shang, Jingtao Shi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Multitask Learning for Human Settlement Extent Regression and Local Climate Zone ClassificationabstractHuman settlement extent (HSE) and local climate zone (LCZ) maps are both essential sources, e.g., for sustainable urban development and Urban Heat Island (UHI) studies. Remote sensing (RS)- and deep learning (DL)-based classification approaches play a significant role by providing the potential for global mapping. However, most of the efforts only focus on one of the two schemes, usually on a specific scale. This leads to unnecessary redundancies since the learned features could be leveraged for both of these related tasks. In this letter, the concept of multitask learning (MTL) is introduced to HSE regression and LCZ classification for the first time. We propose an MTL framework and develop an end-to-end convolutional neural network (CNN), which consists of a backbone network for shared feature learning, attention modules for task-specific feature learning, and a weighting strategy for balancing the two tasks. We additionally propose to exploit HSE predictions as a prior for LCZ classification to enhance the accuracy. The MTL approach was extensively tested with Sentinel-2 data of 13 cities across the world. The results demonstrate that the framework is able to provide a competitive solution for both tasks. Chunping Qiu, Lukas Liebel, Lloyd H. Hughes, Michael Schmitt 0003, Marco Körner 0001, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Pixel-Level Self-Supervised Learning for Semi-Supervised Building Extraction From Remote Sensing ImagesabstractThe building extraction from remote sensed images ash been a challenging yet vital task for applicable purposes such as urban monitoring and cartography. Most of the existing learning based approaches focus on the supervised building extraction methods, of which the models should be trained with images and the corresponding labels. This research exploits a self-supervised approach for building extraction, which could train the backbone within a building extraction network without annotations. Specifically, the backbone is initially trained with a pixel-level self-supervised module instead of commonly used supervised approaches or instance-level self-supervised modules. Next, the pretrained backbone is embedded into a task-specific network followed by tuning with limited annotations. The experiments were conducted on three popular datasets and the results show that our method achieves improvements regarding both intersection over union (IoU) and F1-score compared to supervised approach and instance-level self-supervised methods. Our study thus confirms the potential of pixel-level self-supervised approach for semantic segmentation for remote sensing images. Anzhu Yu, Bing Liu 0018, Xuefeng Cao, Chunping Qiu, Wenyue Guo, Yujun Quan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | SDFL-FC: Semisupervised Deep Feature Learning With Feature Consistency for Hyperspectral Image ClassificationabstractSemisupervised deep learning methods (DLMs) can mitigate the dependence on large amounts of labeled samples using a small number of labeled samples. However, for semisupervised deep feature learning (SDFL), the quality of extracted features cannot be well ensured without a certain amount of labeled samples. To address this issue, we develop the SDFL method with feature consistency (SDFL-FC) for the hyperspectral image (HSI) classification. The SDFL-FC first adopts the convolutional neural network (CNN) to extract spectral–spatial features of HSI and then uses the fully connected layers (FCLs) to model the feature consistency. Moreover, two constraints that enforce both the feature consistency of single pixel (FCS) and feature consistency of group pixels (FCG) are introduced to obtain the representative and discriminative features. The FCS is achieved by the generative adversarial network (GAN) regularization, which can reconstruct the original data from extracted features. The FCG is based on the assumption that the features of group pixels should have similar characteristics within a superpixel, which is embedded in each FCL. The final FCL outputs the class labels, and the cross-entropy (CE) loss is calculated with the labeled samples, while the two losses of FCS and FCG are calculated with all the training samples (both labeled and unlabeled). SDFL-FC integrates the FCS, FCG, and CE loss into a unified objective function and uses a customized iterative optimization algorithm to optimize it. Experiments demonstrate that the SDFL-FC can outperform the related state-of-the-art HSI classification methods. Yuebin Wang, Junhuan Peng, Chunping Qiu, Lei Ding 0008, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | A Novel Approach to Unsupervised Segmentation of Multitemporal VHR Images based on Deep LearningabstractVery-high-resolution (VHR) multi-temporal images are important in remote sensing to monitor the dynamics of the Earth surface. Image semantic segmentation classifies pixels and assigns them label from meaningful object groups. It has been extensively studied in context of single image analysis, however not explored for multi-temporal one. In this paper we propose to extend supervised semantic segmentation to the unsupervised joint segmentation of multi-temporal images. The proposed method processes multi-temporal images by separately feeding them to a deep network comprising of trainable convolutional layers. The training process does not involve any external label. Segmentation labels are obtained from argmax classification of the final layer. Multi-temporal segmentation labels and weights of the trainable layers are jointly optimized in iterations. We tested the method on a VHR dataset from Trento, Italy. Both quantitative and qualitative results demonstrated the effectiveness of the proposed approach. Sudipan Saha, Lichao Mou, Chunping Qiu, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone |
IGARSS | 3 |
| 2020 | Fusing Multiseasonal Sentinel-2 Imagery for Urban Land Cover Classification With Multibranch Residual Convolutional Neural NetworksabstractExploiting multitemporal Sentinel-2 images for urban land cover classification has become an important research topic, since these images have become globally available at relatively fine temporal resolution, thus offering great potential for large-scale land cover mapping. However, appropriate exploitation of the images needs to address problems such as cloud cover inherent to optical satellite imagery. To this end, we propose a simple yet effective decision-level fusion approach for urban land cover prediction from multiseasonal Sentinel-2 images, using the state-of-the-art residual convolutional neural networks (ResNet). We extensively tested the approach in a cross-validation manner over a seven-city study area in central Europe. Both quantitative and qualitative results demonstrated the superior performance of the proposed fusion approach over several baseline approaches, including observation- and feature-level fusion. Chunping Qiu, Lichao Mou, Michael Schmitt 0003, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Unsupervised Deep Joint Segmentation of Multitemporal High-Resolution ImagesabstractHigh/very-high-resolution (HR/VHR) multitemporal images are important in remote sensing to monitor the dynamics of the Earth's surface. Unsupervised object-based image analysis provides an effective solution to analyze such images. Image semantic segmentation assigns pixel labels from meaningful object groups and has been extensively studied in the context of single-image analysis, however not explored for multitemporal one. In this article, we propose to extend supervised semantic segmentation to the unsupervised joint semantic segmentation of multitemporal images. We propose a novel method that processes multitemporal images by separately feeding to a deep network comprising of trainable convolutional layers. The training process does not involve any external label, and segmentation labels are obtained from the argmax classification of the final layer. A novel loss function is used to detect object segments from individual images as well as establish a correspondence between distinct multitemporal segments. Multitemporal semantic labels and weights of the trainable layers are jointly optimized in iterations. We tested the method on three different HR/VHR data sets from Munich, Paris, and Trento, which shows the method to be effective. We further extended the proposed joint segmentation method for change detection (CD) and tested on a VHR multisensor data set from Trento. Sudipan Saha, Lichao Mou, Chunping Qiu, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Fusing Multi-Seasonal Sentinel-2 Images with Residual Convolutional Neural Networks for Local Climate Zone-Derived Urban Land Cover ClassificationabstractThis paper proposes a framework to fuse multi-seasonal Sentinel-2 images, with application on LCZ-derived urban land cover classification. Cross-validation over a seven-city study area in central Europe demonstrates its consistently better performance over several previous approaches, with the same experimental setup. Based on our previous work, we can conclude that decision-level fusion is better than feature-level fusion for similar tasks at similar scale with multi-seasonal Sentinel-2 images. With the framework, urban land cover maps of several cities are produced. The visualization of two exemplary areas shows urban structures that are consistent with existing datasets. This framework can be also generally beneficial for other types of urban mapping. Chunping Qiu, Michael Schmitt 0003, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2019 | Normalized Projection Models for Geostationary Remote Sensing Satellite: A Comprehensive Comparative Analysis (January 2019)abstractNominal grid data of geostationary remote sensing satellites are fundamental for generating the subsequent products. It can be obtained by normalized projection models, mainly based on the imaging mode. However, there are only definitions and primary descriptive equations for the normalized geostationary projection (NGP) model in the existing literature, while the corresponding imaging mode and the physical interpretation are missing, thus hindering the understanding of the produced nominal grid dataset as well as the subsequent products based on the grid. This paper first derived the imaging mode for NGP based on the limited literature. In addition, another new imaging mode was introduced and analyzed based on NGP. The corresponding projection model [nonstandard normalized geostationary projection (NNGP)] was proposed, which is entirely consistent with the situation of America's Geostationary Operational Environmental Satellite-R Series (GOES-R) and Chinese Fengyun-4A (FY-4A). Furthermore, this paper proposed a novel nominal projection model for frame imaging, which is consistent with the imaging mode of China's Gaofen-4. Finally, extensive experiments were designed to comparatively analyze the three nominal grids and demonstrate a detailed difference. By providing a theoretical basis for nominal grid selection, this research is highly significant for the efficient near-real-time production and further applications of geostationary images, as well as the conversion between different datasets resulting from different nominal grid data. In addition, our models are sufficiently tested during the on-orbit running of FY-4A, the first satellite of China's second-generation three-axis stabilized geostationary meteorological satellite series. The algorithms provide the technical support for the high-precision image navigation and registration and play a significant role in robustly producing the meteorological data with similar quality to those from GOES-R. Xiaochong Tong, Lei Yang 0035, Jing Wang 0139, Guangling Lai, Jian Shang, Chunping Qiu, Chengbao Liu, Shengxiong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2018 | Feature Importance Analysis of Sentinel-2 Imagery for Large-Scale Urban Local Climate Zone ClassificationabstractThis paper evaluates different spectral-spatial features that can be extracted from Sentinel-2 imagery regarding their relevance for discriminating different Local Climate Zone (LCZ) classes. The features include spectral reflectance, spectral indices, Morphological Profiles (MPs), as well as Global Urban Footprint (GUF), the Open Street Map layers buildings and land use, and their combinations. Using a residual convolutional neural network (ResNet), a systematic analysis of feature importance is performed with a manually generated dataset distributed in Europe. The results of this evaluation are meant to provide guidance about the choice of both spectral and spatial features for the task of LCZ classification on a global scale. The results show that GUF and OSM can contribute to the classification performance, and ResNet relies less on additional features with the highest accuracy provided by the reflectance only. Chunping Qiu, Michael Schmitt 0003, Pedram Ghamisi, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2017 | Comparative evaluation of signal-based and descriptor-based similarity measures for SAR-optical image matchingabstractThis paper compares different similarity measures for the matching of very-high-resolution SAR and optical images over urban areas. It is meant to provide guidance about the performance of both signal-based and descriptor-based similarity measures in the context of this non-trivial case of multi-sensor correspondence matching. Using an automatically generated training dataset, thresholds for the distinction between correct matches and wrong matches are determined. It is shown that descriptor-based similarity measures outperform signal-based similarity measures significantly. Chunping Qiu, Michael Schmitt 0003, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2004 | DEM generation from stereo SAR images based on polynomial rectification and height displacementabstractThis work introduces the algorithm on DEM generation from stereo SAR images based on polynomial rectification and height displacement, and experiments according to the algorithm on RADARSAT and airborne SAR images in mountain area. To generate DEM from stereo SAR image, there are three kinds of models mostly to be used: 1) the model of range and Doppler equations, 2) the equivalent line central projection model based on the photogrammetry theory, 3) the parallax and elevation relation model which uses the relation between parallax and elevation to calculate elevation difference and then getting the plane coordinates. The model used in This work can be one of the third ones. In our model, image distortion caused by factors other than elevation is corrected by polynomial rectification, and elevation is decided by the difference of height displacement in the pair of the stereo SAR images. As the first step, a certain elevation, for instance the mean elevation of the image pair, is given to the point in process. Secondly, the height displacement in the left image can be corrected, and then the plane coordinates can be gotten by polynomial rectification functions of the left image. Thirdly, by polynomial rectification functions of the right image, the image coordinates of the "same name point" in the right image are available. Finally, from the difference between the coordinates and the actual ones, a new elevation can be gotten by the model of the height displacement of the right image. These steps can be repeated until the new elevation is very close to the old one. According to the above algorithm, programme has been designed. And then, experiment has been down on RADARSAT image in a mountain area in China (Dali, Yunnan). The accuracy is about 3 pixels. For airborne SAR images, some experiments have been down in another mountain area of China (Zhengzhou, Henan) with 1 meter resolution SAR images, and the results are similar to that of the first experiment. So, the new algorithm on DEM created from SAR image pairs introduced by This work is efficient and practicable. After modification, this method can also be used for mixed pair of SAR and optical images. Guoman Huang, Jiankun Guo, Zhou Xiao, Chunping Qiu |
IGARSS | 5 |