Yongyang Xu

dblp:154/4706 · DBLP profile ↗
← Back
9ranked-venue papers in the field
5as first author
7since 2021 · last 2026
0000-0001-7421-4915ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 9 (5 first)
YearPublicationVenuePosition
2026 Integrating geographic knowledge into self-supervised contrastive learning of street view imagery for representing urban space
abstract
In this study, a novel self-supervised contrastive learning framework that integrates geographic knowledge into the analysis of street view imagery for representing urban space is proposed. Traditional methods that rely solely on deep learning often struggle to capture the complex spatial characteristics of urban environments. To address this issue, in the proposed framework, we first extracted visual knowledge (VK) from street view imagery using semantic segmentation and then constructed contrastive samples through VK–imagery pairs. Finally, we introduced distance-weighted temperatures into the contrastive loss function to adjust the similarities of urban features encoded in the representations on the basis of geographical proximity, thereby integrating semantic knowledge among geographical locations (GSK). The method was validated through two case studies: urban village classification in Guangzhou and Foshan and urban-mobility pattern prediction in Shenzhen. The results showed significant improvements in classification accuracy (OA: 0.967) and reduced prediction error (MAE: 28.069) compared with conventional approaches. This research demonstrates the effectiveness of integrating geographic knowledge (including VK and GSK) into contrastive learning. The approach enhances model interpretability and generalizability for urban studies and provides a tool for analyzing urban development patterns and mobility needs that will be useful for urban planning and policymaking.
Sheng Hu 0001, Hanfa Xing, Zhonglin Yang, Jiaju Li, Yongyang Xu, Liang Wu 0005
Int. J. Geogr. Inf. Sci.8
2025 Building pattern recognition by using an edge-attention multi-head graph convolutional network
abstract
Effective building pattern recognition, a complex task that requires the simultaneous consideration of individual building features and spatial relations, is essential for successfully generalizing maps. However, existing deep learning approaches must still be adequately comprehensive in jointly quantifying the individual features and spatial relationships of buildings, suggesting further improvement in the quantitative representation of building spaces. This study presents a novel edge-attention multi-head graph convolutional network (GCN) that concurrently considers the quantitative modeling and representation of individual features and spatial relations, enhancing building pattern recognition. The proposed method captures individual building features and spatial relations, including proximity and arrangement similarity, by using spatial relationship descriptors and attention mechanisms to generate spatial relevance coefficients. These coefficients are then integrated into a weighted multi-head GCN to participate in the quantitative expression of individual features, facilitating the quantitative analysis and modeling of building features, and thus, improving recognition performance. Our experimental analysis confirms the method’s superior capability in recognizing complex spatial features. The method also demonstrates strong generalization across different scales and areas, underscoring its efficacy and potential for enhancing geospatial analyses.
Yongyang Xu, Anna Hu, Xuejing Xie, Siqiong Chen, Zhong Xie
Int. J. Geogr. Inf. Sci.2
2024 Matching the building footprints of different vector spatial datasets at a similar scale based on one-class support vector machines
abstract
Automatic matching of multisource data is an important technique for achieving change detection, fusion and updating spatial data. However, most current learning methods for building footprint matching require a large number of samples, and labeling these samples is costly in terms of labor and time. Moreover, multisource building footprint data are complex and diverse leading to recognizing the different matching relationships is a hard task. Thus, this study proposes a learning-based method for recognizing multisource building footprints matching relationships by using a one-class support vector machine (OCSVM). The OCSVM was trained using only positive samples. First, a set of geometric indicators was designed to train a model and realize initial matching recognition. Then, a contextual metric was calculated based on the rough matching results, and geometric and contextual metrics were combined to train the model and realize relaxed matching recognition. Relaxed matching is an optimization process implemented after initial matching to recognize more relaxed matching relationships. In relaxed matching, a convex hull is used to recognize matching relationships besides 1:1, such as 1:n, m:1 and m:n. The experimental results showed that the proposed method outperformed indicator-weighted (weighted average) and learning-based matching methods, such as traditional SVMs and decision trees (DTs). The precision scores of the proposed model were 97.1%, 95% and 97.2% for the Wuhan (China), Beijing (China) and Richmond Hill (Canada) datasets, respectively. Furthermore, the proposed model identified the matching relationships of buildings with complex geometric features and high-density spatial distributions.
Yongyang Xu, Xuejing Xie, Zhong Xie
Int. J. Geogr. Inf. Sci.1
2023 Uncovering the association between traffic crashes and street-level built-environment features using street view images
abstract
Investigating the relationship between built environment factors and roadway safety is crucial for preventing road traffic accidents. Although studies have analyzed traffic-related built environment factors based on pre-determined zonal units, conclusive evidence regarding the relationship between streetscape features and traffic accidents at a fine-grained road segment level is still lacking. With the widespread availability of large-scale street view images, automatically analyzing urban built environments on a large scale is possible. Therefore, the aim of this study was to investigate the relationship between streetscape features and traffic accidents at a fine-grained road segment level using street view images. Specifically, we employed semantic image segmentation to extract streetscape elements from urban street view images, and then created traffic crash-related variables, including the street-level built environment variables, traffic variables, land-use indices, and proximity characteristics, at the road-segment level. Finally, we adopted a classification-then-regression strategy to model the number of traffic crashes while considering the zero-inflated and spatial heterogeneity issues. Our findings suggest that streetscape features can effectively reflect built-environment characteristics at the road-segment level. Moreover, a comparison of our proposed modeling method with existing models demonstrates its superior performance. The results provide insight into the development of effective planning strategies to improve traffic safety.
Sheng Hu 0001, Hanfa Xing, Wei Luo 0010, Liang Wu 0005, Yongyang Xu, Weiming Huang 0001
Int. J. Geogr. Inf. Sci.5
2023 A geometry-aware attention network for semantic segmentation of MLS point clouds
abstract
Semantic segmentation of mobile laser scanning (MLS) point clouds can provide meaningful 3 D semantic information of urban facilities for various applications. However, it still remains a challenge to extract accurate 3 D semantic information from MLS point cloud data due to its irregular 3 D geometric structure in a large-scale outdoor scene. To this end, this study develops a geometry-aware attention point network (GAANet) with geometric properties of the point cloud as a reference. Specifically, the proposed method first builds a graph-like region for each input point to establish the geometric correlation toward its neighbors for robustly descripting local geometry-aware features. Thereafter, the method introduces a novel multi-head attention mechanism to efficiently learn local discriminative features on the constructed graphs and a feature combination operation to capture both local and global geometric dependencies inside fused point features for significantly facilitating the segmentation of small or incomplete 3 D objects at point-level. Finally, an adaptive loss function is appended to handle class imbalance for the overall performance improvement. The validation experiments on two challenging benchmarks demonstrate the effectiveness and powerful generation ability of the proposed method, which achieves state-of-the-art performance with mean IoU of 65.09% and 95.20% in the Toronto-3D and Oakland 3-D MLS dataset, respectively.
Yongyang Xu, Qinjun Qiu, Zhong Xie
Int. J. Geogr. Inf. Sci.2
2022 Application of a graph convolutional network with visual and semantic features to classify urban scenes
abstract
Urban scenes consist of visual and semantic features and exhibit spatial relationships among land-use types (e.g. industrial areas are far away from the residential zones). This study applied a graph convolutional network with neighborhood information (henceforth, named the neighbour supporting graph convolutional neural network), to learn spatial relationships for urban scene classification. Furthermore, a co-occurrence analysis with visual and semantic features proceeded to improve the accuracy of urban scene classification. We tested the proposed method with the fifth ring road of Beijing with an overall classification accuracy of 0.827 and a Kappa coefficient of 0.769. In comparison with other methods, such as support vector machine, random forest, and general graph convolutional network, the case study showed that the proposed method improved about 10% in urban scene classification.
Yongyang Xu, Shuai Jin, Zhanlong Chen, Xuejing Xie, Sheng Hu 0001, Zhong Xie
Int. J. Geogr. Inf. Sci.1
2021 Measuring the similarity between multipolygons using convex hulls and position graphs
abstract
Polygon similarity can play an important role in geographic information retrieval, map matching and updating, and spatial data mining applications. Geographic information science (GIS) represents various spatial objects as polygons, including simple polygons and polygons with holes, as well as multipolygons. Spatial objects of multipolygons possess complex structure which makes it difficult to assess their similarity. This study develops a method based on convex hulls and position graphs to measure the similarity between multipolygons. The proposed method first finds correspondences between subpolygons in the two multipolygons based on a control polygon. Thereafter, the method constructs a position graph to denote the distribution of these subpolygons and applies a turning function to compute the similarity between various graphs. Fourier transformation and moment invariants were combined to characterize the different matching relationships among subpolygons. The experiments involve three different kinds multipolygons to verify the effectiveness and robustness of proposed method. The experiments show that this approach effectively measures similarity between multipolygons. Moreover, the proposed method accounts for the relationships across the entire complex geometrical shape and components of multipolygon during measuring similarity.
Yongyang Xu, Zhong Xie, Zhanlong Chen, Mingyu Xie
Int. J. Geogr. Inf. Sci.1
2017 Quality assessment of building footprint data using a deep autoencoder network
abstract
Volunteered geographic information (VGI), OpenStreetMap (OSM), has been used in many applications, especially when official spatial data are unavailable or outdated. However, the quality of VGI remains a valid concern. In this paper, we use the matched results between OSM building footprints and official data as the samples for training an autoencoder network, which encodes and reconstructs the sample populations according to unknown complex multivariate probability distributions. Then, the OSM data are assessed based on the theory that small probability samples contribute little to the autoencoder network and that they can be recognized by the higher reconstructed errors during training. In the method described here, the selected measures, including data completeness, positional accuracy, shape accuracy, semantic accuracy and orientation consistency between OSM and official data, are used as the inputs for a deep autoencoder network. Finally, building footprint data from Toronto, Canada, are evaluated, and experiments show that the proposed method can assess the OSM data comprehensively, objectively and accurately.
Yongyang Xu, Zhanlong Chen, Zhong Xie, Liang Wu 0005
Int. J. Geogr. Inf. Sci.1
2017 Shape similarity measurement model for holed polygons based on position graphs and Fourier descriptors
abstract
In geographic information retrieval and spatial data mining, similarity is used to resolve shape matching and clustering. Many approaches have been developed to calculate similarity between simple geometric shapes. However, complex spatial objects are common in spatial database systems, spatial query languages and Geographic Information Science (GIS) applications. With holed polygons, many similarity measurement approaches are restricted to address the relationships between holes or between the holes and the entire complex geometric shape. A successful method should remove the restrictions due to these complex relations and retain invariant during geometric translation (rotation, moving and scaling). To overcome these deficiencies, we utilize position graphs to describe the distribution of holes in complex geometric shapes by storing invariants, such as angles and distances. In addition, Fourier descriptors and the position graph-based method are used to measure the similarity between holed polygons. Experiments show that the proposed method takes into account the relationships in an entire complex geometric shape. It can effectively calculate the similarity of holed polygons, even if they contain different numbers of holes.
Yongyang Xu, Zhong Xie, Zhanlong Chen, Liang Wu 0005
Int. J. Geogr. Inf. Sci.1