EDBT 2026 Demo / reviewers in the wild / expert
Kaiqiang Chen
dblp:211/1966
· DBLP profile ↗
24ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-8314-2375ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SemStereo: Semantic-Constrained Stereo Matching Network for Remote SensingabstractSemantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two heterogeneous tasks are not explicitly modeled, since the pioneering studies either utilize a loosely coupled parallel structure or engage in only implicit interactions, failing to capture the inherent connections. In this work, we explore the connections between the two tasks and propose a new network that imposes semantic constraints on the stereo matching task, both implicitly and explicitly. Implicitly, we transform the traditional parallel structure to a new cascade structure termed Semantic-Guided Cascade structure, where the deep features enriched with semantic information are utilized for the computation of initial disparity maps, enhancing semantic guidance. Explicitly, we propose a Semantic Selective Refinement (SSR) module and a Left-Right Semantic Consistency (LRSC) module. The SSR refines the initial disparity map under the guidance of the semantic map. The LRSC ensures semantic consistency between two views via reducing the semantic divergence after transforming the semantic map from one view to the other using the disparity map. Experiments on the US3D and WHU datasets demonstrate that our method achieves state-of-the-art performance for both semantic segmentation and stereo matching. Chen Chen 0036, Liangjin Zhao, Yuanchun He, Yingxuan Long, Kaiqiang Chen, Zhirui Wang 0003, Yanfeng Hu, Xian Sun 0001 |
AAAI | 5 |
| 2025 | SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldabstractExisting vision-based 3D occupancy prediction methods are inherently limited in accuracy due to their exclusive reliance on street-view imagery, neglecting the potential benefits of incorporating satellite views. We propose SA-Occ, the first Satellite-Assisted 3D occupancy prediction model, which leverages GPS & IMU to integrate historical yet readily available satellite imagery into real-time applications, effectively mitigating limitations of ego-vehicle perceptions, involving occlusions and degraded performance in distant regions. To address the core challenges of cross-view perception, we propose: 1) Dynamic-Decoupling Fusion, which resolves inconsistencies in dynamic regions caused by the temporal asynchrony between satellite and street views; 2) 3D-Proj Guidance, a module that enhances 3D feature extraction from inherently 2D satellite imagery; and 3) Uniform Sampling Alignment, which aligns the sampling density between street and satellite views. Evaluated on Occ3D-nuScenes, SA-Occ achieves state-of-the-art performance, especially among single-frame methods, with a 39.05% mIoU (a 6.97% improvement), while incurring only 6.93 ms of additional latency per frame. Our code and newly curated dataset are available at https://github.com/chenchen235/SA-Occ. Chen Chen 0036, Zhirui Wang 0003, Taowei Sheng, Yundu Li, Peirui Cheng, Luning Zhang, Kaiqiang Chen, Yanfeng Hu, Xue Yang 0005, Xian Sun 0001 |
ICCV | 8 |
| 2025 | EAT: epipolar-aware Transformer for low-light light field enhancement
Xingzheng Wang, Kaiqiang Chen, Zixuan Wang 0006, Yuanlong Deng |
Multim. Tools Appl. | 3 |
| 2025 | SiamTHN: Siamese Target Highlight Network for Visual TrackingabstractSiamese network based trackers develop rapidly in the field of visual object tracking in recent years. The majority of Siamese network based trackers now in use treat each channel in the feature maps generated by the backbone network equally, making the similarity response map sensitive to background influence and hence challenging to focus on the target region. Additionally, there are no structural links between the classification and regression branches in these trackers, and the two branches are optimized separately during training. Therefore, there is a misalignment between the classification and regression branches, which results in less accurate tracking results. In this paper, a Target Highlight Module is proposed to help the generated similarity response maps to be more focused on the target region. To reduce the misalignment and produce more precise tracking results, we propose a corrective loss to train the model. The two branches of the model are jointly tuned with the use of corrective loss to produce more reliable prediction results. Experiments on 5 challenging benchmark datasets reveal that the method outperforms current models in terms of performance, and runs at 38 fps, proving its effectiveness and efficiency. Jiahao Bao, Kaiqiang Chen, Xian Sun 0001, Liangjin Zhao, Wenhui Diao, Menglong Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | UBC-CN: Fine-Grained Building Extraction Dataset For Chinese RegionsabstractIn this paper, we propose a new large scale dataset dedicated to fine-grained building extraction in Chinese regions. To ensure the representativeness of the dataset samples, various factors are taken into consideration, encompassing geographical distribution, appearance variety, rural areas and spatial layout. Consequently, the dataset is composed of 131 K building instances, covering an expansive area of 784.3 square kilometers across 34 cities. Each building is represented with a polygon and a corresponding rooftop type. The rooftops are categorized into 12 fine-grained types. Moreover, we conduct experiments using four classical methods on this dataset to establish baselines for future studies. This dataset will promote the refined development of the city management and planning by providing essential resources for developing the state-of-the-art methods to accurately identify building instances on a city or national scale. Kaiqiang Chen, Xingliang Huang, Taowei Sheng, Xian Sun 0001, Hai Huang 0006 |
IGARSS | 1 |
| 2024 | AI-Powered Flood MapathonabstractFloods represent a pervasive natural hazard with global ramifications, impacting a vast population and resulting in substantial property damage and severe mortality. Particularly worrisome is their disproportionate effect on the least developed countries, which exacerbates developmental imbalances, posing a significant obstacle to the attainment of the United Nations Sustainable Development Goals (UN SDGs). This paper introduces the AI-powered Flood Mapathon activity, co-organized by the Aerospace Information Research Institute under the Chinese Academy of Sciences, in partnership with GEOVIS Technology Co., Ltd., GEOVIS Earth Technology Co., Ltd., and IEEE GRSS IADF. The activity seeks to mobilize individuals worldwide to address the most prevalent natural hazard-floods by collaboratively mapping inundated regions through the analysis of satellite imagery. Gaining widespread attention, the activity has garnered 30,755 submissions from 310 participants across 34 countries. Through collective efforts, participants have curated a semantic segmentation dataset focusing on floods, incorporating annotations of pertinent features related to both floods and human activities. Additionally, the paper elucidates the custom crowdsourcing mapping system, which seamlessly integrates cutting-edge AI technologies to alleviate mapping complexities. The activity contributes to sustainability by drawing extensive public attention, creating a public flood dataset for academic research, and establishing an efficient and intelligent mapping system. Kaiqiang Chen, Xue Lu, Taowei Sheng, Zhirui Wang 0003, Xian Sun 0001, Ronny Hänsch |
IGARSS | 1 |
| 2024 | FS-DCL: Distributed Collaborative Learning for Few-Shot Remote Sensing Image ClassificationabstractWith the development of on-orbit hardware and distributed multiplatform observation systems in satellite remote sensing (RS) scenario, on-orbit collaborative model updating has become a promising trend. Due to restrictions of imaging conditions and storage resources, on-orbit updating is usually carried out with limited samples. However, existing collaborative learning methods rarely consider the few-shot problem. To address this issue, this letter innovatively proposes a distributed collaborative learning method for few-shot RS image classification (FS-DCL), which encourages the collaboration between satellites with similar data distribution to supplement useful information for each satellite, and design on-orbit models to extract more discriminative features. Specifically, a personalized parameter aggregation strategy (PPAS) is proposed to generate personalized parameters for each satellite based on information from satellites with similar data distributions, providing information gain to alleviate problems of insufficient samples. Besides, a feature enhancement method (FEM) is applied to on-orbit models to enhance the feature representation and produce a more discriminative feature space, thus improving the accuracy of few-shot metric classification. Extensive experiments on two RS datasets demonstrate the superiority of FS-DCL. Peirui Cheng, Yuelei Wang, Zhirui Wang 0003, Kaiqiang Chen, Xian Sun 0001, Daobing Zhang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | MDCNet: A Multiplatform Distributed Collaborative Network for Object Detection in Remote Sensing ImageryabstractWith the recent development of remote sensing (RS) technology, the amount of RS platforms has witnessed a substantial increase, and the capacity of Earth observation has been greatly enhanced. The interpretation of RS images has also gradually evolved from traditional centralized ground processing to on-orbit processing. However, the traditional single-platform on-orbit processing is limited to a single source of information, which results in the underutilization of the advantages of multiplatform observation in the current RS field, and restricts the accuracy of inference tasks. To tackle the aforementioned problem, we propose a multiplatform distributed collaborative inference network, which can combine the intermediate features from multiple platforms to improve the accuracy of inference tasks. First, we proposed the collaboration map generator, which generates the collaboration map for optimal collaborator selection autonomously. Second, a spatial feature compression (SFC) module is designed to compress the interplatform transmission features, adapting spatially sparse distribution characteristics of RS objects. Finally, a feature fusion module containing spatial priors is proposed to fuse the features collected from multiple platforms to obtain more precise inference results. We conducted extensive experiments on three public datasets and verified the effectiveness of the proposed framework. On the NWPU VHR-10 dataset, for example, the proposed method improves the detection accuracy by 13.7% and 10.3% under two experimental settings compared with a single platform and compresses the intermediate data transmission between platforms by more than 80%. Shujing Duan, Peirui Cheng, Zhechao Wang, Zhirui Wang 0003, Kaiqiang Chen, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multiview Stereo Reconstruction in Remote SensingabstractResearch on multiview stereo (MVS) based on remote sensing images has promoted the development of large-scale urban 3-D reconstruction. However, remote sensing multiview image data suffer from the problems of occlusion and uneven brightness between views during acquisition, which leads to the problem of blurred details in depth estimation. To solve the above problem, we reexamine the deformable learning method in the MVS task and propose a novel paradigm based on view space and depth deformable learning (SDL-MVS), aiming to learn deformable interactions of features in different view spaces and deformably model the depth ranges and intervals to enable high accurate depth estimation. Specifically, to solve the problem of view noise caused by occlusion and uneven brightness, we propose a progressive space deformable sampling (PSS) mechanism, which performs deformable learning of sampling points in the 3-D frustum space and the 2-D image space in a progressive manner to embed source features to the reference feature adaptively. To further optimize the depth, we introduce depth hypothesis deformable discretization (DHD), which achieves precise positioning of the depth prior by adaptively adjusting the depth range hypothesis and performing deformable discretization of the depth interval hypothesis. Finally, our SDL-MVS achieves explicit modeling of occlusion and uneven brightness faced in MVS through the deformable learning paradigm of view space and depth, achieving accurate multiview depth estimation. Extensive experiments on LuoJia-MVS and WHU datasets show that our SDL-MVS reaches state-of-the-art performance. It is worth noting that our SDL-MVS achieves a mean absolute error (MAE) error of 0.086 and an accuracy of 98.9% for Acc$_{\lt 0.6\,\text {m}}$and 98.9% for Acc$_{\lt 3-\text {interval}}$on the LuoJia-MVS dataset under the premise of three views as input. Yongqiang Mao, Hanbo Bi, Liangyu Xu, Kaiqiang Chen, Zhirui Wang 0003, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid FrameworkabstractIn recent years, remote sensing (RS) vision foundation models such as RingMo have emerged and achieved excellent performance in various downstream tasks. However, the high demand for computing resources limits the application of these models on edge devices. It is necessary to design a more lightweight foundation model to support on-orbit RS image interpretation. Existing methods face challenges in achieving lightweight solutions while retaining generalization in RS image interpretation. This is due to the complex high and low-frequency spectral components in RS images, which make traditional single CNN or Vision Transformer methods unsuitable for the task. Therefore, this paper proposes RingMo-lite, a RS lightweight network with a CNN-Transformer hybrid framework, which effectively exploits the frequency-domain properties of RS to optimize the interpretation process on several tasks like classification, object detection, semantic segmentation, and change detection. It is combined by the Transformer module as a low-pass filter to extract global features of RS images through a dual-branch structure, and the CNN module as a stacked high-pass filter to extract fine-grained details effectively. Furthermore, a novelty-designed frequency-domain masked image modeling (FD-MIM) is employed during the pretraining stage for self-supervised learning, which combines the high-frequency and low-frequency characteristics of each image patch. This approach effectively captures the latent feature representation in RS data. As shown in Fig. 1, compared with RingMo, the proposed RingMo-lite reduces the parameters over 60% in various RS image interpretation tasks, the average accuracy drops by less than 2% in most of the scenes and achieves SOTA performance compared to models of the similar size. In addition, our work will be integrated into the MindSpore computing platform in the near future. Yuelei Wang, Liangjin Zhao, Zhechao Wang, Ziqing Niu, Peirui Cheng, Kaiqiang Chen, Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | PMSNet: Parallel Multi-Scale Network for Accurate Low-Light Light-Field Image EnhancementabstractCurrent low-light light-field (LF) image enhancement algorithms tend to produce blurry results, for (1) loss of spatial details during enhancement and (2) inefficient exploitation of angular correlations, which helps to recover spatial details. Therefore, in this article, we propose a parallel multi-scale network (PMSNet), which attempts to (1) process features of different scales in parallel to aggregate the different contributions of multi-scale features at each layer, thus fully preserve spatial details, and (2) integrate multi-resolution 3D convolution streams to efficiently utilize angular correlations. Specifically, PMSNet consists of three stages: Stage-I employs multi-scale modules (MSMs) to generate local understanding with the aid of adjacent views. Notably, MSM retains high-resolution feature extraction to minimize loss of spatial details. Stage-II processes all views to encode global information. Based on the above extracted local and global information, Stage-III utilizes 3D multi-scale modules (3D-MSMs) to efficiently exploit angular correlations. To validate our idea, we comprehensively evaluate the performance of PMSNet on three publicly available datasets. Experimental results show that our method is superior to the current state-of-the-art methods, achieving an average PSNR of 24.76 dB. Xingzheng Wang, Kaiqiang Chen, Zixuan Wang 0006 |
IEEE Trans. Multim. | 2 |
| 2023 | Light: Joint Individual Building Extraction and Height Estimation from Satellite Images Through a Unified Multitask Learning NetworkabstractBuilding extraction and height estimation are two important basic tasks in remote sensing image interpretation, which are widely used in urban planning, real-world 3D construction, and other fields. Most of the existing research regards the two tasks as independent studies. Therefore the height information cannot be fully used to improve the accuracy of building extraction and vice versa. In this work, we combine the individuaL buIlding extraction and heiGHt estimation through a unified multiTask learning network (LIGHT) for the first time, which simultaneously outputs a height map, bounding boxes, and a segmentation mask map of buildings. Specifically, LIGHT consists of an instance segmentation branch and a height estimation branch. In particular, so as to effectively unify multi-scale feature branches and alleviate feature spans between branches, we propose a Gated Cross Task Interaction (GCTI) module that can efficiently perform feature interaction between branches. Experiments on the DFC2023 dataset show that our LIGHT can achieve superior performance, and our GCTI module with ResNet 101 as the backbone can significantly improve the performance of multitask learning by 2.8% AP50 and 6.5% δ1, respectively. Yongqiang Mao, Xian Sun 0001, Xingliang Huang, Kaiqiang Chen |
IGARSS | 4 |
| 2023 | Urban Building Classification (UBC) V2 - A Benchmark for Global Building Detection and Fine-Grained Classification From Satellite ImageryabstractDatasets play a key role in developing superior building detection approaches. However, most of the previous work focuses on accurate building masks and scale expansion, while the categories are always missing, which hinders the further analysis of urban development and cultures. Therefore, we propose a benchmark for building detection and fine-grained classification from very high-resolution (VHR) satellite imagery. An extensive annotation is performed for about 0.5 million building instances with 12 fine-grained roof types and individual polygons. The annotation of building functions of two cities in the previous version (UBCv1) [1] is also integrated. To ensure the building variety, it consists of VHR optical images of 20 unique cities worldwide with various landforms and styles of architecture. Its variety and fine-grained categories pose great challenges and meanwhile provide a foundation for the building extraction and fine-grained classification on a global scale. Besides, 17 cities are provided with finely aligned Synthetic Aperture Radar (SAR) images, which can be employed for the development and evaluation of approaches optionally based on optical, SAR, or multi-modal images. Significantly, the proposed benchmark is used as the base of the 2023 IEEE GRSS Data Fusion Contest [2]. The dataset and codes of the baseline methods are available at: https://github.com/AICyberTeam/UBC-dataset/tree/UBCv2. Xingliang Huang, Kaiqiang Chen, Deke Tang, Libo Ren, Ronny Hänsch, Michael Schmitt 0003, Xian Sun 0001, Hai Huang 0006, Helmut Mayer 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Elevation Estimation-Driven Building 3-D Reconstruction From Single-View Remote Sensing ImageryabstractBuilding 3D reconstruction from remote sensing images has a wide range of applications in smart cities, photogrammetry and other fields. Methods for automatic 3D urban building modeling typically employ multi-view images as input to algorithms to recover point clouds and 3D models of buildings. However, such models rely heavily on multi-view images of buildings, which are time-intensive and limit the applicability and practicality of the models. To solve these issues, we focus on designing an efficient DSM estimation-driven reconstruction framework (Building3D), which aims to reconstruct 3D building models from the input single-view remote sensing image. Existing DSM estimation networks suffer from the imbalance between local features and global features, which leads to over-smooth DSM estimates at instance boundaries. To address this issue, we propose a Semantic Flow Field-guided DSM Estimation (SFFDE) network, which utilizes the proposed concept of elevation semantic flow to achieve the registration of local and global features. First, in order to make the network semantics globally aware, we propose an Elevation Semantic Globalization (ESG) module to realize the semantic globalization of instances. Further, in order to alleviate the semantic span of global features and original local features, we propose a Local-to-Global Elevation Semantic Registration (L2G-ESR) module based on elevation semantic flow. Our Building3D is rooted in the SFFDE network for building elevation prediction, synchronized with a building extraction network for building masks, and then sequentially performs point cloud reconstruction and surface reconstruction (or CityGML model reconstruction). On this basis, our Building3D can optionally generate CityGML models or surface mesh models of the buildings. Extensive experiments on ISPRS Vaihingen and DFC2019 datasets on the DSM estimation task show that our SFFDE significantly improves upon state-of-the-art and δ1, δ2and δ3metrics of our SFFDE are improved to 0.595, 0.897 and 0.970. Furthermore, our Building3D achieves impressive results in the 3D point cloud and 3D model reconstruction process. Yongqiang Mao, Kaiqiang Chen, Liangjin Zhao, Deke Tang, Wenjie Liu 0016, Zhirui Wang 0003, Wenhui Diao, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CODet: Component Object Detector Extracting Structural Features Based on Target CharacteristicsabstractDeep learning technology has promoted the object detection task in the remote sensing (RS) field to move toward better performance and more demanding requirements. Except for rigid body objects, component objects (COs) with more complex characteristics remain a detection challenge. Its “partial rules and overall disorder” characteristic limits the model learning ability to the structural features. And the internal noise and relatively sparse arrangement are not conducive to optimizing the model by the existing sample assignment strategies. We propose CODet to detect COs in RS scenes. It consists of a cross-hierarchy feature fusion module (CFM) and a noise-sparse sample assignment (NSA) strategy. CFM learns the potential representation and relative position relationship of components by fusing different level features. NSA redefines the optimization process of sample assignment. It aims to alleviate the problems of classification–localization misalignment (CLM) and the positive–negative sample imbalance (PNI) caused by the object’s internal noise and sparse arrangement. The method is verified on the proposed COD dataset of six categories of COs, reaching an average mAP/mAP50of 54.3/86.0. To be closer to the task requirements of the practical RS scene, we also propose a RS large-scale images inference framework. It includes a dataset (APRoI, labeled with COs and rigid body objects), a large-scale image inference strategy, and a set of evaluation metrics. With CODet as the core, the framework can effectively reduce the inference time by three to four times on images with an average of more than 100 million pixels. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Qibin He 0001, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hybrid Multiple Attention Network for Semantic Segmentation in Aerial ImagesabstractSemantic segmentation in very-high-resolution (VHR) aerial images is one of the most challenging tasks in remote sensing image understanding. Most of the current approaches are based on deep convolutional neural networks (DCNNs). However, standard convolution with local receptive fields fails in modeling global dependencies. Prior research works have indicated that attention-based methods can capture long-range dependencies and further reconstruct the feature maps for better representation. Nevertheless, limited by the mere perspective of spatial and channel attention and huge computation complexity of self-attention (SA) mechanism, it is unlikely to model the effective semantic interdependencies between each pixel pair of remote sensing data with complex spectra. In this work, we propose a novel attention-based framework named hybrid multiple attention network (HMANet) to adaptively capture global correlations from the perspective of space, channel, and category in a more effective and efficient manner. Concretely, a class augmented attention (CAA) module embedded with a class channel attention (CCA) module can be used to compute category-based correlation and recalibrate the class-level information. In addition, we introduce a simple yet effective region shuffle attention (RSA) module to reduce feature redundant and improve the efficiency of SA mechanism via regionwise representations. Extensive experimental results on the ISPRS Vaihingen, Potsdam benchmark, and iSAID data set demonstrate the effectiveness and efficiency of our HMANet over other state-of-the-art methods. Ruigang Niu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | AOPDet: Automatic Organized Points Detector for Precisely Localizing Objects in Aerial ImageryabstractWith the development of deep convolutional neural networks, detecting rotating objects in remote-sensing images is of great significance in various fields. Existing rotating object detectors most suffer the problem of ambiguous supervision caused by inappropriate rotating object representations. This problem may result in fuzzy object localization and further lead to misclassification. In this article, we propose an Automatic Organized Points Detector (AOPDet), which derives precise localization results by applying a novel rotating object representation called nonsequential corners representation. To achieve the proposed representation, an Automatic Organization Mechanism (AOM) technique is designed to guide the model to organize points to object corners automatically. An Automatic-Organized-Points-specific (AOP-specific) head structure is also designed and equipped in the model to better focus on the rotating object detection task. On public aerial datasets, experiments show that the AOPDet achieves 17.0 mAP higher than the compared baseline model, reaching the state-of-the-art (SOTA) level. Detailed ablation experiments and error analysis strongly reveal the effectiveness of the proposed model. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Invariant Structure Representation for Remote Sensing Object Detection Based on Graph ModelingabstractDue to the characteristics of vertical orthophoto imaging, the apparent structural features of the object in the remote sensing image are relatively stable, such as the cross-shaped structure of the aircraft, the rectangular structure of the vehicle, etc. Compared with the traditional visual features, using these features is conducive to improving the accuracy of object detection. However, there are few studies on such characteristics. In this paper, we systematically study the invariant structural features of remote sensing objects and propose a Graph Focusing Aggregation Network (GFA-Net) to represent the structural features of remote sensing objects. Among them, in view of the problem that traditional convolutional neural networks (CNNs) are sensitive to the changes in rotation, scale, and other factors, which makes it difficult to extract structural features, we propose the Graph Focusing Process (GFP) based on the idea of graph convolution. Analysis and experiments show that graph structure has significant advantages over Euclidean feature space under CNN in expressing such structural features. In order to realize the end-to-end efficient training of the above model, we design Graph Aggregation Network (GAN) to update the weight of nodes. We verify the effectiveness of our method on the proposed multi-task datasets ACSD and large-scale fine-grained remote sensing dataset FAIR1M. Experiments conducted on the object detection data sets of DOTA and HRSC2016 prove that the proposed method is superior to the current state-of-the-art method. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Effective Fusion of Multi-Modal Data with Group Convolutions for Semantic Segmentation of Aerial ImageryabstractIn this paper, we achieve a semantic segmentation of aerial imagery based on the fusion of multi-modal data in an effective way. The multi-modal data contains a true orthophoto and the corresponding normalized Digital Surface Model (nDSM), which are stacked together before they are fed into a Convolutional Neural Network (CNN). Though the two modalities are fused at the early stage, their features are learned independently with group convolutions firstly and then the learned features of different modalities are fused at multiple scales with standard convolutions. Therefore, the multi-scale fusion of multi-modal features is completed in a single-branch convolutional network. In this way, the computational cost is reduced while the experimental results reveal that we can still get promising results. Kaiqiang Chen, Kun Fu 0001, Menglong Yan, Wenkai Zhang 0002, Yue Zhang 0016, Xian Sun 0001 |
IGARSS | 1 |
| 2019 | Effective Classification of Local Climate Zones Based on Multi-Source Remote Sensing DataabstractThe local climate zone (LCZ) classification divides the urban areas into 17 categories, which are composed of 10 manmade structures and 7 natural landscapes. Though originally designed for temperature study, LCZ classification can be used for studies on economy and population. In this paper, we achieve a LCZ classification with convolutional neural networks based on the multi-source remote sensing data, including the polarimetric synthetic aperture radar (PolSAR) data and the corresponding multi-spectral imagery (MSI). Through experiments we attempt to reveal the contributions of the SAR data and the MSI to the classification performance. Furthermore, we emphasize the crucial importance of the preprocessing on the training data to derive a balanced dataset. We are ranked second in the Tianchi competition rankings when we submit our results. Yingchao Feng, Wenkai Zhang 0002, Yue Zhang 0016, Siyue Wang, Kun Fu 0001, Kaiqiang Chen |
IGARSS | 7 |
| 2018 | Deep Semantic Segmentation of Aerial Imagery Based on Multi-Modal DataabstractIn this paper, we focus on the use of multi-modal data to achieve a semantic segmentation of aerial imagery. Thereby, the multi-modal data is composed of a true orthophoto, the Digital Surface Model (DSM) and further representations derived from these. Taking data of different modalities separately and in combination as input to a Residual Shuffling Convolutional Neural Network (RSCNN), we analyze their value for the classification task given with a benchmark dataset. The derived results reveal an improvement if different types of geometric features extracted from the DSM are used in addition to the true orthophoto. Kaiqiang Chen, Kun Fu 0001, Xian Sun 0001, Michael Weinmann, Stefan Hinz, Boris Jutzi, Martin Weinmann |
IGARSS | 1 |
| 2018 | Semantic Segmentation of Aerial Images With Shuffling Convolutional Neural NetworksabstractSemantic segmentation of aerial images refers to assigning one land cover category to each pixel. This is a challenging task due to the great differences in the appearances of ground objects. Many attempts have been made during the past decades. In recent years, convolutional neural networks (CNNs) have been introduced in the remote sensing field, and various solutions have been proposed to realize dense semantic labeling with CNNs. In this letter, we propose shuffling CNNs to realize semantic segmentation of aerial images in a periodic shuffling manner. This approach is a supplement to current methods for semantic segmentation of aerial images. We propose a naive version and a deeper version of this method, and both are adept at detecting small objects. Additionally, we propose a method called field-of-view (FoV) enhancement that can enhance the predictions. This method can be applied to various networks, and our experiments verify its effectiveness. The final results are further improved through an ensemble method that averages the score maps generated by the models at different checkpoints of the same network. We evaluate our models using the ISPRS Vaihingen and Potsdam data sets, and we acquire promising results using these two data sets. Kaiqiang Chen, Kun Fu 0001, Menglong Yan, Xian Sun 0001, Xin Wei 0004 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Flat-roofed building reconstruction based on layover modelling and MCMC methodabstractIn this paper, we propose a top-down building reconstruction technique based on layover modelling and MCMC method. Through representing the layover with parameterized geometrical models, the problem is converted into an optimization problem under the Bayesian scheme. The energy function consists of two parts: region part and edge part. In order to obtain global optima, simulated annealing algorithm with MCMC is used in the optimization stage. Two groups of transmission kernels which are responsible for model updates are designed according to the model. This method is tested both on simulated SAR image and HR TanDEM-X data. At this moment, only qualitative analysis for this method is provided. It proves the effectiveness of the presented method. Detailed quantitative evaluation will be added when we submit the final version of this paper. Yue Zhang 0016, Xian Sun 0001, Kun Fu 0001, Kaiqiang Chen |
IGARSS | 4 |
| 2017 | Building extraction from remote sensing images with deep learning in a supervised mannerabstractBuilding extraction from remote sensing images is a longstanding topic in land use analysis and applications of remote sensing. Variations in shape and appearance of buildings, occlusions and other unpredictable factors increase the hardness of automatic building extraction. Numerous methods have been proposed during the last several decays, but most of these works are task oriented and lack of generalization. This paper applys deep learning to building extraction in a supervised manner. A deep deconvolution neural network with 27 Convolution/Deconvolution weight layers is designed to realize building extraction in pixel level. As such a deep network is prone to overfitting, a data augment method that suits pixel-wise prediction tasks in remote sensing is suggested. Moreover, an overall training and inferencing architecture is proposed. Our methods are finally applied to building extraction tasks and get competitive results with other methods published. Kaiqiang Chen, Kun Fu 0001, Menglong Yan, Xian Sun 0001 |
IGARSS | 1 |