Xiaochong Tong

dblp:16/8503 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Spatial-Aware Remote Sensing Image Generation From Spatial Relationship Descriptions
abstract
Recent advances in stable diffusion models have revolutionized text-to-image generation. However, these models struggle with spatial relationship comprehension in remote sensing (RS) scenarios, limiting their ability to generate spatially accurate imagery. We present a novel framework for generating RS images from spatial relationship descriptions with precise spatial control. Our approach introduces a two-stage pipeline: first, a spatial relationship semantic structuring model converts formalized spatial relationship descriptions into controlled layouts, and second, an enhanced diffusion model incorporates positional prompts and a layout attention mechanism to generate the final image. The positional prompts explicitly encode spatial information, while the layout attention mechanism enables focused region learning. Comprehensive experiments demonstrate that our method achieves superior performance compared with state-of-the-art approaches in both spatial accuracy and image quality.
Yaxian Lei, Xiaochong Tong, Chunping Qiu, Haoshuai Song, Congzhou Guo
IEEE Geosci. Remote. Sens. Lett.2
2025 Attention-Driven Object Encoding and Multiscale Contextual Perception for Improved Cross-View Object Geo-Localization
abstract
Cross-view object geo-localization is essential for applications like navigation and intelligent city management. By identifying objects in street-view/drone-view and precisely locating them in satellite imagery, more accurate geo-localization can be achieved compared to retrieval-based methods. However, existing approaches fail to account for query object shape/size and significant scale variations in remote sensing images. To address these limitations, we propose an Attention-Driven Multi-Scale Perception Network (AMPNet) for cross-view geo-localization. AMPNet employs an attention-driven object encoding (ADOE) based on segmentation, which provides prior information to enable learning more discriminative representations of the query object. Additionally, AMPNet introduces a cross-view multi-scale perception (CVMSP) module that captures multiscale contextual information using varying convolution kernels, and applies an MLP to enhance channel-wise feature interactions. Experimental results demonstrate that AMPNet outperforms state-of-the-art methods in both ground-to-satellite and drone-to-satellite object localization tasks on a challenging dataset.
Haoshuai Song, Xiaochong Tong, Xiaoyu Zhang 0008, Yaxian Lei, He Li 0035, Congzhou Guo
IEEE Geosci. Remote. Sens. Lett.2
2023 Open Self-Supervised Features for Remote-Sensing Image Scene Classification Using Very Few Samples
abstract
Big models, large datasets, and self-supervised learning (SSL) have recently gained substantial research interest due to their potential to alleviate our reliance on annotations. Considering the current high generalization ability of self-supervised models in literature, we explore in the letter how helpful SSL can be for a crucial task in remote sensing (RS), image scene classification, when forced to rely on only a few labeled samples. We proposed a simple prototype-based classification procedure without training and fine-tuning, which uses open self-supervised features from the contrastive language-image pre-training (CLIP). We test our method by exploiting ready-to-use open features on four diversified benchmark datasets, including red-green-blue (RGB) and multispectral (MS) images. Highly competitive accuracy has been obtained compared to work with similar settings, i.e., based on an exceedingly small number of labels. To the best of our knowledge, our model is the first to achieve such high accuracy in austere label conditions. We further analyze our approach from different perspectives, including its advantages and limitations, reasons for its astonishing performance, potential applications, and future improvements.
Chunping Qiu, Anzhu Yu, Xiaodong Yi 0002, Naiyang Guan, Dian-xi Shi, Xiaochong Tong
IEEE Geosci. Remote. Sens. Lett.6
2023 Learning Visual Representation Clusters for Cross-View Geo-Location
abstract
Cross-view geo-location is a crucial research field that determines the geographic location from images taken from different viewpoints. It is often studied as a retrieval task, where the query images are with unknown locations, and the database includes images with geo-tags from a different platform. Learning image representations by neural networks is an important step, and one typical training method is using a classification loss, where cross-view images of the same locations are considered the same category. However, existing methods only focus on pushing the representation distances of different categories while ignoring the intra-category representation distances of samples from different platforms. Considering that controlling the intra-category distance can help to guide the model to extract compact category-sharing representations from cross-view images, we propose a categorized cluster loss to learn separate and compact representation clusters. Categorized cluster loss can supervise the network to learn invariant information from samples of different platforms by constraining both the inter-category and intra-category feature distances. Meanwhile, we design a category-view-stratified sampling strategy, which samples balanced inputs in terms of both category and view in each batch during the learning process. We implemented our approach with a lightweight OSNet-based network and achieved higher accuracy with fewer parameters on a typical and challenging cross-view geo-location dataset than most state-of-the-art (SOTA) methods.
Haoshuai Song, Zhen Wang 0052, Dian-xi Shi, Xiaochong Tong, Yaxian Lei, Chunping Qiu
IEEE Geosci. Remote. Sens. Lett.5
2023 Learning From Self-Supervised Features for Hashing-Based Remote Sensing Image Retrieval
abstract
Image retrieval (IR) for practical remote sensing (RS) should have high accuracy, storage, and calculation efficiency, while not relying on big annotations. However, current supervised and unsupervised RSIR methods do not yet fully meet these requirements. To this end, we propose a novel hashing-based IR approach via learning hash codes from open and representative self-supervised features. Specifically, we constructed a model out of a self-supervised pretrained backbone and a small multilayer perceptron (MLP)-based hashing learning neural network. Features from the frozen backbones were used to reconstruct a similarity matrix to guide the hash network learning. This way, the semantic structure can be preserved. To enhance the proposed approach, we propose the exploitation of global high-level semantic information within the similarity reconstruction process by introducing a small set of labeled datasets. Extensive comparative experiments on two commonly used RS image datasets demonstrate the outperformance of our proposed approach and its good balance between the retrieval accuracy and utilized annotations. In these two datasets, the labeled data required by our method accounts for less than 3% of that required by traditional methods, but our obtained mean average precision (mAP) can reach over 90%, which is close to that of current advanced supervised methods. In addition, we analyzed the specific effect of our design and the associated hyperparameters.
Dali Wang, Xiaochong Tong, Chunping Qiu
IEEE Geosci. Remote. Sens. Lett.3
2023 Onboard Data Management Approach Based on a Discrete Grid System for Multi-UAV Cooperative Image Localization
abstract
Onboard image data management and sharing are the foundations for achieving cooperative data processing and analysis in multiple unmanned aerial vehicles (multi-UAVs). However, various challenges, such as the lack of efficient onboard data indices, restrict the development of multi-UAV cooperative applications. Here, we propose a novel and versatile cooperative data management framework based on a discrete grid system for multi-UAV onboard image data. First, we study the image coding methodology employed within the proposed framework. This method transforms original spatiotemporal and attribute information in images into well standardized and structured coded information. Secondly, we introduce a grid-based onboard image data management approach (Grid-OIM) to facilitate cooperative data management among multi-UAVs using code-based index and query methods. Finally, we applied Grid-OIM to cooperative image localization tasks. Experiments were conducted using an edge-computing platform and an embedded database. The image coding method could process > 12,000 images/second while maintaining excellent real-time performance. Moreover, the efficiency of creating and updating the image data index and querying the image data increased by averages of 17.6, 9.3, and 66.1 times, respectively, compared to the image data management method based on R*-Tree, highlighting the substantial advantages of this proposed method. These improvements address the demands of indexing and querying highly dynamic onboard image data effectively. Furthermore, the horizontal accuracy of image localization calculated by the cooperative localization method was improved by 43.4-81.6% compared to that of a single UAV, enhancing reliability. Overall, Grid-OIM presents a feasible and practical solution for multi-UAV cooperative applications.
Xiaochong Tong, Chunping Qiu, Yuekun Sun, Congzhou Guo
IEEE Trans. Geosci. Remote. Sens.2
2022 Corrigendum to "An efficient integer coding index algorithm for multi-scale time information management" [Data Knowl. Eng. 119 (2019) 123-138]
Xiaochong Tong, Chengqi Cheng, Guangling Lai, Bo Chen 0015
Data Knowl. Eng.1
2022 Multiscale Feature Learning by Transformer for Building Extraction From Satellite Images
abstract
Extracting buildings from very high-resolution satellite images is a challenging yet important task for applications such as urban monitoring. Multiscale feature learning proves to be a potential solution toward accurate extraction of buildings. This study exploits a powerful multiscale feature learning module, a hierarchical vision transformer by shifted windows (swin), as a backbone within a building extraction network. To this end, we first designed a general structure for building extraction, consisting of a backbone to extract multiscale features and a head network to fuse and refine features. Then, we integrated swin into the structure as a backbone and utilized channel-wise and spatial-wise enhancement in a head network. Experimental results show that our method achieves improvements regarding both F1-score and intersection over union (IoU) compared to the multiple attending path neural network (MAP-Net), which is the current state-of-the-art (SOTA) algorithm for building extraction from remote sensing images. Our study thus confirms the potential of swin transformers as backbones for semantic segmentation tasks based on satellite images.
Xin Chen 0088, Chunping Qiu, Wenyue Guo, Anzhu Yu, Xiaochong Tong, Michael Schmitt 0003
IEEE Geosci. Remote. Sens. Lett.5
2022 On-Board Thermal Motion Compensation Method for Pointing Errors of the Remote Sensor Aboard a Three-Axis Stabilized Geostationary Satellite
abstract
For earth observation from a three-axis stabilized geostationary (GEO) remote-sensing satellite, a highly accurate onboard thermal motion compensation (TMC) method for real-time correction of the satellite sensor’s line-of-sight (LOS) pointing error due to thermally induced structural distortion internal to the sensor during the diurnal cycle remains a global concern to date. In this letter, we propose a novel TMC method for GEO sensors. Compared with the traditional TMC methods, this method has the following two advantages. First, we define the LOS misalignment angle to model the comprehensive impact of the sensor’s internal thermal misalignment angles on the sensor’s LOS pointing behavior, which makes the optical path modeling simpler and more generic. Second, we fit the repeatable diurnal variations of the LOS misalignment angle based on long time-series star observations; therefore, the calculation and the use of the TMC amount are no longer limited to the assumption that the sensor’s internal thermal misalignment angles remain constant within a certain period of time, as in the traditional method. The proposed TMC method is theoretically applicable to all types of GEO sensors, including optical and microwave sensors, and is verified by Fengyun-4A advanced geosynchronous radiation imager (FY-4A/AGRI) TMC experiments.
Hualong Hu, Xiaochong Tong, Chunping Qiu, Yanfa Shang, Jingtao Shi
IEEE Geosci. Remote. Sens. Lett.2
2019 An efficient integer coding index algorithm for multi-scale time information management
Xiaochong Tong, Chengqi Cheng, Guangling Lai, Bo Chen 0015
Data Knowl. Eng.1
2019 Normalized Projection Models for Geostationary Remote Sensing Satellite: A Comprehensive Comparative Analysis (January 2019)
abstract
Nominal grid data of geostationary remote sensing satellites are fundamental for generating the subsequent products. It can be obtained by normalized projection models, mainly based on the imaging mode. However, there are only definitions and primary descriptive equations for the normalized geostationary projection (NGP) model in the existing literature, while the corresponding imaging mode and the physical interpretation are missing, thus hindering the understanding of the produced nominal grid dataset as well as the subsequent products based on the grid. This paper first derived the imaging mode for NGP based on the limited literature. In addition, another new imaging mode was introduced and analyzed based on NGP. The corresponding projection model [nonstandard normalized geostationary projection (NNGP)] was proposed, which is entirely consistent with the situation of America's Geostationary Operational Environmental Satellite-R Series (GOES-R) and Chinese Fengyun-4A (FY-4A). Furthermore, this paper proposed a novel nominal projection model for frame imaging, which is consistent with the imaging mode of China's Gaofen-4. Finally, extensive experiments were designed to comparatively analyze the three nominal grids and demonstrate a detailed difference. By providing a theoretical basis for nominal grid selection, this research is highly significant for the efficient near-real-time production and further applications of geostationary images, as well as the conversion between different datasets resulting from different nominal grid data. In addition, our models are sufficiently tested during the on-orbit running of FY-4A, the first satellite of China's second-generation three-axis stabilized geostationary meteorological satellite series. The algorithms provide the technical support for the high-precision image navigation and registration and play a significant role in robustly producing the meteorological data with similar quality to those from GOES-R.
Xiaochong Tong, Lei Yang 0035, Jing Wang 0139, Guangling Lai, Jian Shang, Chunping Qiu, Chengbao Liu, Shengxiong Zhou
IEEE Trans. Geosci. Remote. Sens.1
2013 Efficient encoding and spatial operation scheme for aperture 4 hexagonal discrete global grid system
abstract
Discrete global grid systems (DGGSs) are considered to be promising structures for global geospatial information representation. Square and triangular DGGSs have had the advantage over hexagonal ones in geospatial data processing over the past few decades. Despite a significant body of research supporting hexagonal grids as the superior alternative, the application thereof has been hindered partly owing to the lack of a hierarchy. This study presents an original perspective to combine two types of aperture 4 hexagonal discrete grid systems into a hierarchy. Each cell of the hierarchy is assigned a unique code using a linear quadtree that constructs the hexagonal quaternary balanced structure (HQBS). The mathematical system described by HQBS addressing and the vector operations, including addition, subtraction, multiplication, and division, are defined. Essential spatial operations for HQBS cell retrieval, transformation between HQBS codes and other coordinate systems, and arrangement of HQBS cells on spherical surfaces were studied and implemented. The accuracy and efficiency of algorithms were validated through experiments. The results indicate that the average efficiency of cell retrieval using the HQBS is higher than that using other schemes, thus proving it to be more efficient.
Xiaochong Tong, Jin Ben, Tao Pei
Int. J. Geogr. Inf. Sci.1