VLDB 2026 Research / reviewers in the wild / expert
Xiuyuan Zhang
dblp:194/7160
· DBLP profile ↗
14ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-9350-9126ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap DataabstractIntegrating ground-level geospatial data with rich geographic context, like OpenStreetMap (OSM), into remote sensing (RS) foundation models (FMs) is essential for advancing geospatial intelligence and supporting a broad spectrum of tasks. However, modality gap between RS and OSM data, including differences in data structure, content, and spatial granularity, makes effective synergy highly challenging, and most existing RS FMs focus on imagery alone. To this end, this study presents GeoLink, a multimodal framework that leverages OSM data to enhance RS FM during both the pretraining and downstream task stages.
Specifically, GeoLink enhances RS self-supervised pretraining using multi-granularity learning signals derived from OSM data, guided by cross-modal spatial correlations for information interaction and collaboration. It also introduces image mask-reconstruction to enable sparse input for efficient pretraining. For downstream tasks, GeoLink generates both unimodal and multimodal fine-grained encodings to support a wide range of applications, from common RS interpretation tasks like land cover classification to more comprehensive geographic tasks like urban function zone mapping. Extensive experiments show that incorporating OSM data during pretraining enhances the performance of the RS image encoder, while fusing RS and OSM data in downstream tasks improves the FM’s adaptability to complex geographic scenarios. These results underscore the potential of multimodal synergy in advancing high-level geospatial artificial intelligence. Moreover, we find that spatial correlation plays a crucial role in enabling effective multimodal geospatial data integration. Code, checkpoints, and using examples are released at [GitHub](https://github.com/bailubin/GeoLink_NeurIPS2025). Lubin Bai, Xiuyuan Zhang, Shihong Du |
NeurIPS | 2 |
| 2025 | Estimating Individual Building Heights by Integrating Spaceborne LiDAR and Multisource Remote Sensing Data: A CNN-Transformer Model and a Semi-Supervised Sample Augmentation ApproachabstractBuilding height is a key metric in transitioning urban analysis from two-dimensional to three-dimensional perspectives. It serves as a fundamental indicator for assessing urban development, population density, and energy consumption. Regardless of whether machine learning or deep learning methods are used, accurate height estimation inevitably relies on large quantities of labeled building height samples for model training. However, existing reference datasets often suffer from ambiguity due to floor-to-height conversions and temporal obsolescence. The advent of spaceborne light detection and ranging (LiDAR) provides a promising avenue to overcome these limitations, enabling the acquisition of large-scale, high-precision, and time-sensitive reference height data. Moreover, while buildings are inherently individual and discrete units, most existing height estimation studies operate at gridded scales, thereby obscuring inter-building height variability. In this study, we propose a novel approach to estimate individual building heights by integrating spaceborne LiDAR with multisource remote sensing data. First, leveraging building footprint data, we derive reference height samples by jointly retrieving building heights from the Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) and the Global Ecosystem Dynamics Investigation (GEDI). Then, we construct a multimodal feature set by extracting seasonal features from a combination of multisource datasets, such as Sentinel-1, Sentinel-2, and SDGSAT-1. To model the spatially complex height distribution, we propose the Neighborhood-aware CNN-Transformer Network (NeCT-Net), which estimates building heights at the pixel level. To address the scarcity of tall building samples, we incorporate a semi-supervised learning paradigm. Specifically, pseudo-labels for high-rise buildings are generated using predictions from a teacher model and iteratively refined through self-training, thereby enhancing the student model’s learning capacity in tall-structure scenarios. Finally, we propose a post-processing approach that constrains height predictions using building morphological information, enabling the conversion from pixel-level predictions to vector-based building-level estimations. We apply the proposed method to two iconic urban areas—Beijing and Seattle. The results demonstrate strong predictive performance, with RMSE, MAE, and R values of 9.01 m, 6.08 m, and 0.93 for Beijing, and 8.72 m, 5.31 m, and 0.79 for Seattle. Compared to baseline methods, our approach reduces RMSE by 29.66% in Beijing and 7.43% in Seattle, while improving R by 47.62% and 46.06%. When benchmarked against existing individual building height products, our model achieves RMSE reductions of 31.43% and 21.16%, and R improvements of 57.63% and 23.44% in the two cities. These results highlight the methodological advancement and robustness of our approach. This work presents a new paradigm for large-scale, individual building height estimation through the integration of spaceborne LiDAR and multisource remote sensing data. The source code will be made available at https://github.com/Lunar-Elf/NeCT-Net. Shouhang Du, Hao Liu 0105, Jianghe Xing, Xiuyuan Zhang, Xunyu Guan, Shihong Du |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Object-Based Urban Land-Use Change Detection With Siamese Network and Hierarchical ClusteringabstractUrban land use (ULU) undergoes continuous and dynamic changes, profoundly reshaping the spatial configuration of urban areas. Accurately quantifying these changes is critical for the precise planning and efficient management of limited urban land resources during urbanization. However, existing change detection methods struggle to address the unique challenges of ULU change detection, leading to a lack of reliable ULU change maps. Two major challenges hinder the study of ULU changes: (1) high internal heterogeneity, which obscures the main features of ULU, and (2) inaccurate boundary delineation. To overcome these challenges, we propose an object-based method for detecting ULU changes. Our approach employs a Siamese neural network equipped with a self-attention mechanism to extract global features, a Fusion module to integrate global and local features, and a multi-scale Amplifier to filter and enhance the main features. Subsequently, a hierarchical clustering framework aggregates the segmented objects with the main features into coherent ULU units. Experimental results demonstrate that the proposed method effectively addresses the two major challenges of ULU change detection. The proposed method achieves the best performance compared to representative methods in Beijing and Shanghai, two of China’s largest and most typical cities. Furthermore, the analysis of ULU change results reveals that both cities experienced rapid outward urban expansion between 2015 and 2020. Song Ouyang, Shihong Du, Xiuyuan Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Retrieving Historical Images at 10-m Resolution From 1985 to 2015 Through STARS-Net: A Spatiotemporal Attention Referenced Super-Resolution NetworkabstractHigh-resolution (HR) satellite imagery over long time series is crucial for monitoring fine-scale land use and land cover changes. However, existing freely available Earth observation data are limited to medium to coarse resolutions or short time spans. Our goal is to reconstruct 10-m resolution Sentinel-2-like imagery from 1985 to 2015 using Landsat imagery. Unlike previous studies that heavily rely on HR reference images from the same period, we utilize the overlapping Landsat and Sentinel-2 observations to reconstruct 10-m Sentinel-2-like imagery from 1985 to 2015. Specifically, we propose the STARS-Net, which can effectively transfer multiple similar textures from Sentinel-2 to Landsat images, ensuring accurate reconstruction even with significant land changes. The innovation of our method lies in its ability to integrate any number of HR images, without limiting these images to similar times or the same locations, leading to more flexible and accurate reconstructions. As a result, a 10-m resolution dataset of five major China cities from 1985 to 2015 is available athttps://figshare.com/s/186e9b25426067e0dcf4with a five-year interval. Extensive comparative experiments demonstrate that STARS-Net outperforms state-of-the-art super-resolution (SR) methods and the results highlight that STARS-Net exhibits strong robustness in various scenarios, including cloud cover, image missing, and land cover changes. More importantly, our method produces the first historical image retrospect at 10-m resolution, which has a finer resolution compared to Landsat imagery and an important time extend compared with the Sentinel-2 data. Therefore, our dataset can help to facilitate a deeper understanding of terrestrial transformations over the past four decades. Shuping Xiong, Xiuyuan Zhang, Yichen Lei, Ge Tan, Shihong Du |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | AP-Semi: Improving the Semi-Supervised Semantic Segmentation for VHR Images Through Adaptive Data Augmentation and Prototypical Sample GuidanceabstractAs a method that can incorporate unlabeled data into model training, semi-supervised semantic segmentation (SSS) can mitigate the burden of manual annotation in geographic mapping tasks with very-high-resolution (VHR) images. Various SSS approaches have been proposed for VHR images, but most of them either rely on complex models or additional training procedures, increasing the computation cost. In this study, we propose a simple-yet-effective SSS approach named AP-semi for VHR images, which neither increases the number of learnable parameters nor introduces additional training steps. First, since data augmentation can influence the training data and process significantly, we refine the traditional CutMix transformation, an important data augmentation strategy, by generating adaptively sized cut boxes, optimizing the intensity of perturbation and making it better suited for the training phase. Second, in order to make all the unlabeled samples contribute to the model training, we design a prototype-based feature-level consistency loss, which can enforce the consistency of model predictions in the feature space. By combining the adaptive data augmentation strategy and prototype-based loss, AP-semi can achieve impressive results with a limited number of labeled samples, surpassing several baseline models in extensive experiments. Through experiments, we find that AP-semi can adapt to both convolutional neural networks (CNNs)- and transformer-based segmentation models. In the simulation experiment mimicking the routine operations in geographic mapping tasks, AP-semi demonstrates significant time-saving benefits. Lubin Bai, Xiuyuan Zhang, Bo Liu 0080, Shihong Du |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Children's developing understanding of learning as improvement over time
Xiuyuan Zhang, Brandon Carrillo, Julia A. Leonard |
CogSci | 1 |
| 2023 | Synergistic Classification of Multilevel Land Patches (SC-MLPs): Reducing Conflicts and Improving Mapping Results for Land Uses and Functional Spaces With Very-High-Resolution Satellite ImageryabstractLand uses (e.g., commercial, residential, and industrial lands) and functional spaces (e.g., living, productive, and ecological spaces) are two-level landscape patches and totally work as basic units for urban planning. The two-level patches are interrelated and mutually binding, but existing mapping methods extracted them separately, leading to substantial conflicts and errors in their mapping results. Accordingly, this study proposes a synergistic classification of multilevel land patches (SC-MLPs). It considers a multitask learning strategy and proposes a novel correlation loss function to measure the correlations between land uses and functional spaces, which is expected to resolve conflicts and improve the accuracy of the two-level land patch mapping results. Consequently, land-use and functional-space maps of three major Chinese cities are generated, which generally have a high resolution of 2 m and high overall accuracies of 90.1% for land uses and 93.8% for functional spaces. Compared to state-of-the-art land-use and functional-space mapping methods, our results have not only higher accuracies but also a better consistency which is improved by 36%. Accordingly, the proposed SC-MLP can generate not only accurate but also consistent maps of land uses and functional spaces, which plays a fundamental role in land system research and urban planning. Xiuyuan Zhang, Shuping Xiong, Xiaoyan Dong, Shihong Du |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Willingness to Interact Increases When Opponents Offer Specific Evidence
Benjamin F. Scheve, Xiuyuan Zhang, Frank C. Keil |
CogSci | 2 |
| 2022 | People Have Systematically Different Intuitions about Ownership even in Seemingly Simple Cases
Xiuyuan Zhang, Paul Bloom, Julian Jara-Ettinger |
CogSci | 1 |
| 2022 | Thinking about doing: Representations of skill learning
Xiuyuan Zhang, Samuel D. McDougle, Julia A. Leonard |
CogSci | 1 |
| 2022 | Domain Adaptation for Remote Sensing Image Semantic Segmentation: An Integrated Approach of Contrastive Learning and Adversarial LearningabstractAlthough semantic segmentation models based on deep neural networks (DNNs) have achieved excellent results, generalizing well from one remote sensing dataset (source domain) to another dataset with different acquisition conditions (target domain) remains a major challenge. Many domain adaptation (DA) approaches have been proposed to address this problem. DA aims to help DNNs learn a generalizable representation space in which source and target domains have similar feature distributions, but most of the existing DA approaches have difficulty in aligning the high-dimensional image representations of two domains directly. In this study, we proposed a model integrating contrastive learning and adversarial learning in a unified framework for aligning two domains in both representation space and spatial layout. Specifically, the model consists of a semantic segmentation network for feature extraction and two branches for DA. The first branch is used for adaptation in representation space directly by a proposed pixelwise contrastive loss, while the second branch is used for adaptation in predicted results to help two domains have similar spatial layouts through a novel but simple entropy-based similarity discriminator. Additionally, a training strategy called category similarity matching sampling was proposed to provide source and target image pairs with similar category composition for each training iteration, which can help the two branches work better. Extensive experiments indicated that the two branches can benefit each other to gain a superior performance and DA pretraining by our methods can achieve impressive results with only a small number of target labeled samples. Lubin Bai, Shihong Du, Xiuyuan Zhang, Bo Liu 0080, Song Ouyang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Interpretation of Generic Language is Dependent on Listener's Background Knowledge
Xiuyuan Zhang, Daniel Yurovsky |
CogSci | 1 |
| 2019 | Learning Self-Adaptive Scales for Extracting Urban Functional Zones From Very-High-Resolution Satellite ImagesabstractUrban functional zones (e.g. commercial, residential, and industrial) are basic units for city planning and management, and play an important role in city studies. However, functional zones are difficult to extract from very-high-resolution (VHR) remote sensing images, as they are various in components, sizes, and heterogeneities, leading to different segmentation scales. To resolve this issue, this study uses selfhood scale, a local optimum scale, to extract functional zones. Firstly, geoscene segmentation is used to delineate functional zones at multiple scales. Then, selfhood scales are calculated to measure the local optimum scales of segmenting functional zones, based on which multiscale segmentation results can be finally assembled into one layer to generate functional-zone boundaries. The experimental results indicate this method that is effective to delineate functional zones in Beijing, adapting to local built environments. Xiuyuan Zhang, Shihong Du |
IGARSS | 1 |
| 2017 | Classifying natural-language spatial relation terms with random forest algorithmabstractThe exponential growth of natural language text data in social media has contributed a rich data source for geographic information. However, incorporating such data source for GIS analysis faces tremendous challenges as existing GIS data tend to be geometry based while natural language text data tend to rely on natural language spatial relation (NLSR) terms. To alleviate this problem, one critical step is to translate geometric configurations into NLSR terms, but existing methods to date (e.g. mean value or decision tree algorithm) are insufficient to obtain a precise translation. This study addresses this issue by adopting the random forest (RF) algorithm to automatically learn a robust mapping model from a large number of samples and to evaluate the importance of each variable for each NLSR term. Because the semantic similarity of the collected terms reduces the classification accuracy, different grouping schemes of NLSR terms are used, with their influences on classification results being evaluated. The experiment results demonstrate that the learned model can accurately transform geometric configurations into NLSR terms, and that recognizing different groups of terms require different sets of variables. More importantly, the results of variable importance evaluation indicate that the importance of topology types determined by the 9-intersection model is weaker than metric variables in defining NLSR terms, which contrasts to the assertion of ‘topology matters, metric refines’ in existing studies. Shihong Du, Chen-Chieh Feng, Xiuyuan Zhang |
Int. J. Geogr. Inf. Sci. | 4 |