EDBT 2026 Demo / reviewers in the wild / expert
A-Xing Zhu
dblp:125/8552 · also A.-Xing Zhu, Axing Zhu
· DBLP profile ↗
31ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0000-0002-5725-0460ORCID · reported
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 29 (2 first)Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatial prediction of environmental suitability for dengue fever transmission based on geographical environmental similarity with spatial connectivityabstractSpatial prediction of environmental suitability for vector-borne disease transmission is crucial for public health, with spatial connectivity being a critical factor. However, existing models often face a trade-off between model interpretability and the accurate representation of spatial connectivity. Simpler, interpretable models tend to oversimplify connectivity, while more advanced models can lack interpretability and often rely on difficult-to-obtain fine-grained mobility data. To address this challenge, this paper proposes an integrated framework (GES-SC) for the spatial prediction of environmental suitability for vector-borne disease transmission, which combines Geographical Environmental Similarity (GES) with quantified Spatial Connectivity (SC). The method operates on the principle that environmentally similar and highly spatially connected locations have similar transmission potential. It quantifies connectivity using available road network data and distance decay, providing an effective alternative to direct mobility data. A case study in Guangzhou, China, demonstrates that integrating this quantified spatial connectivity enhances prediction accuracy. Compared to Geographically Weighted Regression, Bayesian Conditional Autoregressive, XGBoost, and a Simple Neural Network, the GES-SC method performed better in cross-validation, achieving lower RMSE and MAE, and a higher R2. The proposed framework effectively addresses the accuracy-interpretability trade-off, improving spatial epidemiological prediction accuracy and providing an uncertainty measure for public health decisions. Haiwen Du, A-Xing Zhu, Tianwu Ma |
Int. J. Geogr. Inf. Sci. | 2 |
| 2025 | A multivariate spatial structure indicator based on geographic similarityabstractQuantification of spatial structure reveals the distribution patterns of geographic features, which is essential for geographic analysis. For quantitative measurements of multivariate spatial structure, existing methods often neglect either the geographic meaning or the spatial combination of the multivariate dataset. In this paper, a new indicator for multivariate spatial structure (MuSS) is proposed. Since multivariate datasets characterize geographic conditions, the correlation between multivariate attributes at different locations can be measured as the similarity of geographic conditions. The MuSS indicator evaluates whether location pairs with closer distances have higher geographic similarities. Experimental results show that MuSS outperforms existing methods in differentiate multivariate datasets with varied spatial distribution patterns. A MuSS value deviating from 1 suggests that geographic similarity between location pairs is relevant to their distance, and the statistical significance of the captured distribution patterns is evaluated using a p-value from a permutation test. MuSS is also applied to real geographic data at different spatial resolutions in two study areas with diverse distribution patterns. Case studies show that MuSS provides consistent comparison results for spatial structure levels between study areas, while existing methods cannot. Fang-He Zhao, Cheng-Zhi Qin 0001, A-Xing Zhu, Tao Pei |
Int. J. Geogr. Inf. Sci. | 3 |
| 2023 | An adaptive uncertainty-guided sampling method for geospatial prediction and its application in digital soil mappingabstractSampling design can significantly reduce the uncertainty in geospatial predictions. In this paper, we developed an adaptive uncertainty-guided stepwise sampling (AUGSS) method to select sampling locations to supplement existing legacy sample points whose representation should be improved. The proposed method selects supplemental samples in a stepwise manner as guided by an objective function with two weighted sub-objectives. One reduces the area with high prediction uncertainty, and the other minimizes the overall prediction uncertainty for the entire area. The method takes an adaptive approach to adjust weights for the two sub-objectives and to tune an uncertainty threshold controlling whether a location can be reliably predicted during the sampling procedure. A case study on soil property prediction shows that AUGSS outperforms the stratified random sampling (SRS) and the non-adaptive uncertainty guided sampling method (UGSS) in terms of RMSE and Lin’s concordance correlation coefficient with different sample sizes. This study shows that the AUGSS method offers a potential for effectively adding supplemental samples to existing samples which are insufficient for spatial prediction. The adaptive strategy guided by predicted uncertainty provides an efficient support to improve the spatial pattern of samples, which plays a key role in the result accuracy of geospatial predictive mapping. Lei Zhang 0151, A-Xing Zhu, Junzhi Liu, Tianwu Ma, Lin Yang 0018, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 2 |
| 2023 | Spatial prediction of groundwater level change based on the Third Law of GeographyabstractSpatial prediction methods are an important means of predicting the spatial variation of groundwater level change. Existing methods extract spatial or statistical relationships from samples to represent the study area for inference and require a representative sample set that is usually in large quantity and is distributed across geographic or covariate space. However, samples for groundwater are usually sparsely and unevenly distributed. In this paper, an approach based on the Third Law of Geography is proposed to make predictions by comparing the similarity between each individual sample and unmeasured site. The approach requires no specific number or distribution of samples and provides individual uncertainty measures at each location. Experiments in three different watersheds across the U.S. show that the proposed methods outperform machine learning methods when available samples do not well represent the area. The provided uncertainty measures are indicative of prediction accuracy by location. The results of this study also show that the spatial prediction based on the Third Law of Geography can also be successfully applied to dynamic variables such as groundwater level change. Fang-He Zhao, A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 3 |
| 2022 | A load-balancing strategy for data domain decomposition in parallel programming libraries of raster-based geocomputationabstractParallel programming libraries have been proposed to simplify programming for parallel raster-based geocomputation through hiding parallel programming details for users. However, the strategy of data domain decomposition used in existing libraries often leads to load imbalance owing to inherent characteristics of geocomputation including not only irregular spatial data distribution, but also spatial variation in the amount of computation, thereby impeding their parallel performances. This paper thus proposes a load-balancing strategy of data domain decomposition in parallel programming libraries for raster-based geocomputation based on the concept of spatial computational domain, which characterizes the distribution of computational intensity based on geocomputation characteristics. By implementing the proposed strategy with the message passing interface (MPI), a set of parallel raster-based geocomputation operators across different parallel computing platforms (known as PaRGO V2) was upgraded to improve load-balancing parallelization. The proposed strategy was evaluated by parallelizing two typical geocomputation algorithms (i.e. inverse distance weight interpolation and fuzzy c-means clustering) using PaRGO V2 with uneven distributed computational intensity. The results show that the proposed strategy with PaRGO V2, compared with the previously adopted data domain decomposition strategy, yielded significant improvements to the load balance (i.e. better parallel performance). Yu-Jing Wang, Bei-Bei Ai, Cheng-Zhi Qin 0001, A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 4 |
| 2022 | Incorporation of spatial anisotropy in urban expansion modelling with cellular automataabstractCellular Automata (CA) models have become the most commonly used tool for simulating urban expansion. To improve the accuracy of CA models, various driving factors like spatial proximity and neighbourhood effects have been explored in previous studies, but the inclusion of these factors does not address the directional differences in urban expansion. To address this issue, this study develops a method to measure urban spatial anisotropy (SA) with respect to 18 variables at both the global and local scales, and integrates all these SA variables into a logistic regression-based CA model. The revised CA model is evaluated with a case study for Huizhou, China. The case study shows that the simulation results for the CA model with SA exhibit 89% overall accuracy; compared to CA models that do not consider SA, the revised CA model can improve precision by 5% on newly developed cells. The consideration of SA in CA models proves promising in improving the accuracy of urban expansion simulations. Jinqu Zhang, Yu Ling, A-Xing Zhu, Hongyun Zeng, Jia Song 0001, Yunqiang Zhu, Lang Qian |
Int. J. Geogr. Inf. Sci. | 3 |
| 2019 | Reflections and speculations on the progress in Geographic Information Systems (GIS): a geographic perspectiveabstractGreat strides have been made in Geographic Information Systems (GIS) research over the past half-century. However, this progress has created both opportunities and challenges. From a geographic perspective, certain challenges remain, including the modelling of geographic-featured environments with GIS data model, the enhancement of GIS’s analysis functions for comprehensive geographic analysis and achieving human-oriented geographic information presentation. Several basic theoretical and technical ideas that follow the workflow and processes of geographic information induction, geographic scenario modelling, geographic process analysis and geographic environment representation are proposed to fill the gaps between GIS and geography. We also call for designing methods for big geographic data-oriented analysis, making best use of videos and developing virtual geographic scenario-based GIS for further evolution. Guonian Lv, Michael Batty, Josef Strobl, Hui Lin 0002, A-Xing Zhu, Min Chen 0008 |
Int. J. Geogr. Inf. Sci. | 5 |
| 2019 | A representativeness-directed approach to mitigate spatial bias in VGI for the predictive mapping of geographic phenomenaabstractVolunteered geographic information (VGI) contains valuable field observations that represent the spatial distribution of geographic phenomena. As such, it has the potential to provide regularly updated low-cost field samples for predictively mapping the spatial variations of geographic phenomena. The predictive mapping of geographic phenomena often requires representative samples for high mapping accuracy, but samples consisting of VGI observations are often not representative as they concentrate on specific geographic areas (i.e. spatial bias) due to the opportunistic nature of voluntary observation efforts. In this article, we propose a representativeness-directed approach to mitigate spatial bias in VGI for predictive mapping. The proposed approach defines and quantifies sample representativeness by comparing the probability distributions of sample locations and the mapping area in the environmental covariate space. Spatial bias is mitigated by weighting the sample locations to maximize their representativeness. The approach is evaluated using species habit suitability mapping as a case study. The results show that the accuracy of predictive mapping using weighted sample locations is higher than using unweighted sample locations. A positive relationship between sample representativeness and mapping accuracy is also observed, suggesting that sample representativeness is a valid indicator of predictive mapping accuracy. This approach mitigates spatial bias in VGI to improve predictive mapping accuracy. Guiming Zhang, A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 2 |
| 2018 | A discrete global grid system for earth system modelingabstractTo support Earth system modeling, we propose a discrete global grid system that expresses multi-resolution spatial data. Specifically, a unified coding model that expresses a grid of nodes, edges, and cells is constructed for a triangular discrete global grid system. To fulfill the requirements of practical applications, we design a code-based topological query method for this grid system and an algorithm to transform between grid codes and geographic coordinates. We evaluate the Global Finite Volume Community Ocean Model (Global-FVCOM) on the triangular discrete global grid system in the proposed uniform coding model. The ocean tidal waves simulated by the Global-FVCOM running on the coded grid are then compared with results obtained using a traditional irregular spherical grid system, and the results display comparable accuracy. The uniform coding model proposed in this paper provides a triangular discrete global grid system that can represent multi-resolution spatial data and can be used in Earth system models. This unified coding model can also be applied to the geographic coordinate system made up of latitudes and longitudes, as well as diamond and hexagonal grids. Bingxian Lin, Liangchen Zhou, Depeng Xu 0003, A-Xing Zhu, Guonian Lv |
Int. J. Geogr. Inf. Sci. | 4 |
| 2017 | Template-based GIS computation: a geometric algebra approachabstractThe tight coupling between geospatial data and spatial analysis results in high costs in terms of efficiency when developing algorithms to accommodate different types of data, even when the analysis tasks are the same. Universal GIS (Geographic information system) algorithms, as alternatives to tightly coupled approaches, can reduce development costs. However, a unified representation of spatial data is necessary to support the development of universal GIS algorithms. To this end, this research proposes and implements a template-based approach using geometric algebra to create a unified representation of multidimensional data. The template is composed of parameters and operators for GIS representation and computation. The template approach can support general GIS analyses with parameter unfolding and operator integration methods. A case study of intersection analysis shows that developing programming scripts based on computation templates is much simpler than traditional methods. The results suggest that the template-based method is more efficient than traditional methods and more convenient for high-dimensional applications. Wen Luo 0004, Zhaoyuan Yu, Linwang Yuan, A-Xing Zhu, Guonian Lv |
Int. J. Geogr. Inf. Sci. | 5 |
| 2017 | A peak-cluster assessment method for the identification of upland planation surfacesabstractResidual upland planation surfaces serve as strong evidence of peneplains during long intervals of base-level stability in the peneplanation process. Multi-stage planation surfaces could aid the calculation of uplift rates and the reconstruction of upland plateau evolution. However, most planation surfaces have been damaged by crustal uplift, tectonic deformation, and surface erosion, thus increasing the difficulty in automatically identifying residual planation surfaces. This study proposes a peak-cluster assessment method for the automatic identification of potential upland planation surfaces. It consists of two steps: peak extraction and peak-cluster characterization. Three critical parameters, namely, landform planation index (LPI), peak elevation standard deviation, and peak density, are employed to assess peak clusters. The proposed method is applied and validated in five case areas in the Tibetan Plateau using a Shuttle Radar Topography Mission digital elevation model (SRTM DEM) with 3 arc-second resolution. Results show that the proposed method can effectively extract potential planation surfaces, which are found to be stable with different resolutions of DEM data. A significant planation characteristic can be obtained in the relatively flat areas of the Gangdise–Nyainqentanglha Mountains and Qaidam Basin. Several vestiges of potential former planation areas are also extracted in the hilly-gully areas of the western part of the Himalaya Mountains, the northern part of the Tangula–Hengduan Mountains, and the northeastern part of the Kunlun–Qinling Mountains despite the absence of significant topographical features characterized by low slope angles or low terrain reliefs. Vestiges of planation surfaces are also identified in these hilly-gully upland areas. Hence, the proposed method can be effectively used to extract potential upland planation surfaces not only in flat areas but also in hilly-gully areas. Liyang Xiong, Guoan Tang, A-Xing Zhu, Yeqing Qian |
Int. J. Geogr. Inf. Sci. | 3 |
| 2017 | A GPU-accelerated adaptive kernel density estimation approach for efficient point pattern analysis on spatial big dataabstractKernel density estimation (KDE) is a classic approach for spatial point pattern analysis. In many applications, KDE with spatially adaptive bandwidths (adaptive KDE) is preferred over KDE with an invariant bandwidth (fixed KDE). However, bandwidths determination for adaptive KDE is extremely computationally intensive, particularly for point pattern analysis tasks of large problem sizes. This computational challenge impedes the application of adaptive KDE to analyze large point data sets, which are common in this big data era. This article presents a graphics processing units (GPUs)-accelerated adaptive KDE algorithm for efficient spatial point pattern analysis on spatial big data. First, optimizations were designed to reduce the algorithmic complexity of the bandwidth determination algorithm for adaptive KDE. The massively parallel computing resources on GPU were then exploited to further speed up the optimized algorithm. Experimental results demonstrated that the proposed optimizations effectively improved the performance by a factor of tens. Compared to the sequential algorithm and an Open Multiprocessing (OpenMP)-based algorithm leveraging multiple central processing unit cores for adaptive KDE, the GPU-enabled algorithm accelerated point pattern analysis tasks by a factor of hundreds and tens, respectively. Additionally, the GPU-accelerated adaptive KDE algorithm scales reasonably well while increasing the size of data sets. Given the significant acceleration brought by the GPU-enabled adaptive KDE algorithm, point pattern analysis with the adaptive KDE approach on large point data sets can be performed efficiently. Point pattern analysis on spatial big data, computationally prohibitive with the sequential algorithm, can be conducted routinely with the GPU-accelerated algorithm. The GPU-accelerated adaptive KDE approach contributes to the geospatial computational toolbox that facilitates geographic knowledge discovery from spatial big data. Guiming Zhang, A-Xing Zhu, Qunying Huang |
Int. J. Geogr. Inf. Sci. | 2 |
| 2017 | A similarity-based automatic data recommendation approach for geographic modelsabstractThe complexity of geographic modelling is increasing; hence, preparing data to drive geographic models is becoming a time-consuming and difficult task that may significantly hinder the application of such models. Meanwhile, a huge number of data sets have been shared and have become publicly accessible through the Internet. This study presents a data similarity-based approach to automatically recommend available data sets to fulfil the data requirements of geographic models. Unified description factors are adopted to provide a consistent description of public data sets and input data requirements of geographic models. Five elementary data similarities between them, specifically content, spatial coverage, temporal coverage, spatial precision, and temporal granularity similarities, are calculated. An overall similarity is estimated from aggregating the elementary data similarities. Thereafter, the candidate data for running the models are recommended in the order of overall data similarity. As a case study, the approach has been applied to recommend data from the China National Data Sharing Platform of Earth System Science to drive the population spatialization model (PSM). The approach has successfully recommended the most related data sets to run PSM. The result also suggests that the data recommendation approach can facilitate the intelligent identification of geographic data and the building of links between the open data sets. Yunqiang Zhu, A-Xing Zhu, Jia Song 0001, Jie Yang 0002, Qiuyi Zhang 0003, Kai Sun 0009, Jinqu Zhang, Ling Yao |
Int. J. Geogr. Inf. Sci. | 2 |
| 2016 | A function-based linear map symbol building and rendering method using shader languageabstractMaps are widely used to visualize geo-information so that map users can develop related understandings about the real world. Such a process for communicating information is largely dependent on the rendering of map elements using different symbols (points and linear and area symbols). To meet the demand of more dynamic and comprehensive visualization in map rendering, it is essential to improve the rendering efficiency. This paper focuses on these research topics, especially the difficulty in constructing and drawing linear map symbols. By employing shader language, a function-based linear symbol building and rendering method is presented in this paper. The basic idea of this function-based method is to build a map-rendering solution that employs graphic processing unit (GPU) acceleration technology to improve the rendering efficiency. A ‘function’ is used to represent the algorithm that draws certain simple or complex linear map symbols. This function reflects the structure of a linear map symbol (describing the symbol construction information) and also the rendering process of the symbolized linear map elements (handled on a per-pixel basis by the shader program). Based on the Open Geospatial Consortium (OGC), Styled Layer Descriptor (SLD) specifications, four basic line types (i.e., solid lines, dashed lines, gradient color lines, and transition lines) are implemented in the proposed method, and the implementation of line markers, line joins and line caps is also discussed. Three experiments are conducted to demonstrate improvements in map rendering. The results show that a variety of linear map symbols can be constructed in a uniform way, which suggests that the proposed method addresses the difficulty in drawing linear map symbols. With this method, the efficiency of rendering linear map elements is substantially improved compared to using the graphics device interface plus (GDI+) and anti-grain geometry (AGG) methods; it also provides an applicable approach for developing map rendering systems. Using this function-based concept, the complexity of building linear map symbols and drawing linear map elements can be decreased. Songshan Yue, Jianshun Yang, Min Chen 0008, Guonian Lv, A-Xing Zhu, Yongning Wen |
Int. J. Geogr. Inf. Sci. | 5 |
| 2016 | Enabling point pattern analysis on spatial big data using cloud computing: optimizing and accelerating Ripley's K functionabstractPerforming point pattern analysis using Ripley’s K function on point events of large size is computationally intensive as it involves massive point-wise comparisons, time-consuming edge effect correction weights calculation, and a large number of simulations. This article presented two strategies to optimize the algorithm for point pattern analysis using Ripley’s K function and utilized cloud computing to further accelerate the optimized algorithm. The first optimization sorted the points on their x and y coordinates and thus narrowed the scope of searching for neighboring points down to a rectangular area around each point in estimating K function. Using the actual study area in computing edge effect correction weights is essential to estimate an unbiased K function, but is very computationally intensive if the study area is of complex shape. The second optimization reused the previously computed weights to avoid repeating expensive weights calculation. The optimized algorithm was then parallelized using Open Multi-Processing (OpenMP) and hybrid Message Passing Interface (MPI)/OpenMP on the cloud computing platform. Performance testing showed that the optimizations effectively accelerated point pattern analysis using K function by a factor of 8 using both the sequential version and the OpenMP-parallel version of the optimized algorithm. While the OpenMP-based parallelization achieved good scalability with respect to the number of CPU cores utilized and the problem size, the hybrid MPI/OpenMP-based parallelization significantly shortened the time for estimating K function and performing simulations by utilizing computing resources on multiple computing nodes. Computational challenge imposed by point pattern analysis tasks on point events of large size involving a large number of simulations can be addressed by utilizing elastic, distributed cloud resources. Guiming Zhang, Qunying Huang, A-Xing Zhu, John H. Keel |
Int. J. Geogr. Inf. Sci. | 3 |
| 2015 | Change detection for 3D vector data: a CGA-based Delaunay-TIN intersection approachabstractIn this paper, conformal geometric algebra (CGA) is introduced to construct a Delaunay–Triangulated Irregular Network (DTIN) intersection for change detection with 3D vector data. A multivector-based representation model is first constructed to unify the representation and organization of the multidimensional objects of DTIN. The intersection relations between DTINs are obtained using the meet operator with a sphere-tree index. The change of area/volume between objects at different times can then be extracted by topological reconstruction. This method has been tested with the Antarctica ice change simulation data. The characteristics and efficiency of our method are compared with those of the Möller method as well as those from the Guigue–Devillers method. The comparison shows that this new method produces five times less redundant segments for DTIN intersection. The computational complexity of the new method is comparable to Möller’s and that of Guigue–Devillers methods. In addition, our method can be easily implemented in a parallel computation environment as shown in our case study. The new method not only realizes the unified expression of multidimensional objects with DTIN but also achieves the unification of geometry and topology in change detection. Our method can also serve as an effective candidate method for universal vector data change detection. Zhaoyuan Yu, Wen Luo 0004, Linwang Yuan, A-Xing Zhu, Guonian Lv |
Int. J. Geogr. Inf. Sci. | 5 |
| 2015 | A citizen data-based approach to predictive mapping of spatial variation of natural phenomenaabstractThe vast accumulation of environmental data and the rapid development of geospatial visualization and analytical techniques make it possible for scientists to solicit information from local citizens to map spatial variation of geographic phenomena. However, data provided by citizens (referred to as citizen data in this article) suffer two limitations for mapping: bias in spatial coverage and imprecision in spatial location. This article presents an approach to minimizing the impacts of these two limitations of citizen data using geospatial analysis techniques. The approach reduces location imprecision by adopting a frequency-sampling strategy to identify representative presence locations from areas over which citizens observed the geographic phenomenon. The approach compensates for the spatial bias by weighting presence locations with cumulative visibility (the frequency at which a given location can be seen by local citizens). As a case study to demonstrate the principle, this approach was applied to map the habitat suitability of the black-and-white snub-nosed monkey (Rhinopithecus bieti) in Yunnan, China. Sightings of R. bieti were elicited from local citizens using a geovisualization platform and then processed with the proposed approach to predict a habitat suitability map. Presence locations of R. bieti recorded by biologists through intensive field tracking were used to validate the predicted habitat suitability map. Validation showed that the continuous Boyce index (Bcont(0.1)) calculated on the suitability map was 0.873 (95% CI: [0.810, 0.917]), indicating that the map was highly consistent with the field-observed distribution of R. bieti. Bcont(0.1) was much lower (0.173) for the suitability map predicted based on citizen data when location imprecision was not reduced and even lower (−0.048) when there was no compensation for spatial bias. This indicates that the proposed approach effectively minimized the impacts of location imprecision and spatial bias in citizen data and therefore effectively improved the quality of mapped spatial variation using citizen data. It further implies that, with the application of geospatial analysis techniques to properly account for limitations in citizen data, valuable information embedded in such data can be extracted and used for scientific mapping. A-Xing Zhu, Guiming Zhang, Zhi-Pang Huang, Ge-Sang Dunzhu, Guopeng Ren, Cheng-Zhi Qin 0001, Lin Yang 0018, Tao Pei, Shengtian Yang |
Int. J. Geogr. Inf. Sci. | 1 |
| 2015 | A Hierarchical Tensor-Based Approach to Compressing, Updating and Querying Geospatial DataabstractWith the rapid development of data observation and model simulation in geoscience, spatial-temporal data have become increasingly multidimensional, massive and are consistently being updated. As a result, the integrated maintenance of these data is becoming a challenge. This paper presents a blocked hierarchical tensor representation within the split-and-merge paradigm for the compressed storage, continuously updating and data querying of multidimensional geospatial field data. The original multidimensional geospatial field data are split into small blocks according to their spatial-temporal references. These blocks are represented and compressed hierarchically, and then combined into a single hierarchical tree as the representation of original data. With a buffered binary tree data structure and corresponding optimized operation algorithms, the original multidimensional geospatial field data can be continuously compressed, appended, and queried. Data from the 20th Century Reanalysis Monthly Mean Composites are used to evaluate the performance of this approach. Compared to traditional methods, the new approach is shown to retain the quality of the original data with much lower storage costs and faster computational performance. The result suggests that the blocked hierarchical tensor representation provides an effective structure for integrated storage, presentation and computation of multidimensional geospatial field data. Linwang Yuan, Zhaoyuan Yu, Wen Luo 0004, Linyao Feng, A-Xing Zhu |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2014 | A strategy for raster-based geocomputation under different parallel computing platformsabstractThe demand for parallel geocomputation based on raster data is constantly increasing with the increase of the volume of raster data for applications and the complexity of geocomputation processing. The difficulty of parallel programming and the poor portability of parallel programs between different parallel computing platforms greatly limit the development and application of parallel raster-based geocomputation algorithms. A strategy that hides the parallel details from the developer of raster-based geocomputation algorithms provides a promising way towards solving this problem. However, existing parallel raster-based libraries cannot solve the problem of the poor portability of parallel programs. This paper presents such a strategy to overcome the poor portability, along with a set of parallel raster-based geocomputation operators (PaRGO) designed and implemented under this strategy. The developed operators are compatible with three popular types of parallel computing platforms: graphics processing unit supported by compute unified device architecture, Beowulf cluster supported by message passing interface (MPI), and symmetrical multiprocessing cluster supported by MPI and open multiprocessing, which make the details of the parallel programming and the parallel hardware architecture transparent to users. By using PaRGO in a style similar to sequential program coding, geocomputation developers can quickly develop parallel raster-based geocomputation algorithms compatible with three popular parallel computing platforms. Practical applications in implementing two algorithms for digital terrain analysis show the effectiveness of PaRGO. Cheng-Zhi Qin 0001, Li-Jun Zhan, A-Xing Zhu, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 3 |
| 2013 | Artificial surfaces simulating complex terrain types for evaluating grid-based flow direction algorithmsabstractThis article presents a set of artificial surfaces simulating complex terrain types for evaluating the performances of grid-based flow direction algorithms. The proposed artificial surfaces were developed based on sine and cosine functions and thus can simulate four complex terrain types: a convex-centred slope, concave-centred slope, saddle-centred slope and straight-ridge slope; such features are typical and widespread in real landscapes. We analytically solved the theoretical values of specific catchment area (SCA) for the proposed artificial surfaces. Compared with existing artificial surfaces for evaluating flow direction algorithms, the proposed artificial surfaces provide a better representation of common terrain types in the real world. To analyse the feasibility of the proposed artificial surfaces, two sets of artificial digital elevation models (DEMs) were created by sampling the proposed artificial surfaces with different reliefs at a series of resolutions (i.e. 1, 5, 10 and 20 m). Four representative flow direction algorithms were applied to these artificial DEMs: D8, D-inf, FD8 and MFD-md. The root mean square error, mean error and standard deviation in the computed SCA from flow direction algorithm show that MFD-md generally yielded lower error under simulated terrain conditions than D8, D-inf and FD8. The cumulative frequency distributions of errors from the tested flow direction algorithms with the proposed artificial DEMs can effectively reflect the inherent characteristics of each algorithm. MFD-md performed more similarly to FD8 in low-relief terrains and more similarly to D-inf in high-relief terrains. The map of errors from each tested algorithm is available for a spatially explicit evaluation of the occurrence of errors. Over most of the area, the D-inf algorithm underestimated the SCA when FD8 overestimated the SCA. Cheng-Zhi Qin 0001, Li-Li Bao, A-Xing Zhu, Biao Qin |
Int. J. Geogr. Inf. Sci. | 3 |
| 2013 | Uncertainty due to DEM error in landslide susceptibility mappingabstractTerrain attributes such as slope gradient and slope shape, computed from a gridded digital elevation model (DEM), are important input data for landslide susceptibility mapping. Errors in DEM can cause uncertainty in terrain attributes and thus influence landslide susceptibility mapping. Monte Carlo simulations have been used in this article to compare uncertainties due to DEM error in two representative landslide susceptibility mapping approaches: a recently developed expert knowledge and fuzzy logic-based approach to landslide susceptibility mapping (efLandslides), and a logistic regression approach that is representative of multivariate statistical approaches to landslide susceptibility mapping. The study area is located in the middle and upper reaches of the Yangtze River, China, and includes two adjacent areas with similar environmental conditions – one for efLandslides model development (approximately 250 km2) and the other for model extrapolation (approximately 4600 km2). Sequential Gaussian simulation was used to simulate DEM error fields at 25-m resolution with different magnitudes and spatial autocorrelation levels. Nine sets of simulations were generated. Each set included 100 realizations derived from a DEM error field specified by possible combinations of three standard deviation values (1, 7.5, and 15 m) for error magnitude and three range values (0, 60, and 120 m) for spatial autocorrelation. The overall uncertainties of both efLandslides and the logistic regression approach attributable to each model-simulated DEM error were evaluated based on a map of standard deviations of landslide susceptibility realizations. The uncertainty assessment showed that the overall uncertainty in efLandslides was less sensitive to DEM error than that in the logistic regression approach and that the overall uncertainties in both efLandslides and the logistic regression approach for the model-extrapolation area were generally lower than in the model-development area used in this study. Boxplots were produced by associating an independent validation set of 205 observed landslides in the model-extrapolation area with the resulting landslide susceptibility realizations. These boxplots showed that for all simulations, efLandslides produced more reasonable results than logistic regression. Cheng-Zhi Qin 0001, Li-Li Bao, A-Xing Zhu, Rong-Xun Wang |
Int. J. Geogr. Inf. Sci. | 3 |
| 2013 | An integrative hierarchical stepwise sampling strategy for spatial sampling and its application in digital soil mappingabstractSampling design plays an important role in spatial modeling. Existing methods often require a large amount of samples to achieve desired mapping accuracy, but imply considerable cost. When there are not enough resources for collecting a large set of samples at once, stepwise sampling approach is often the only option for collecting the needed large sample set, especially in the case of field surveying over large areas. This article proposes an integrative hierarchical stepwise sampling strategy which makes the samples collected at different stages an integrative one. The strategy is based on samples' representativeness of the geographic feature at different scales. The basic idea is to sample at locations that are representative of large-scale spatial patterns first and then add samples that represent more local patterns in a stepwise fashion. Based on the relationships between a geographic feature and its environmental covariates, the proposed sampling method approximates a hierarchy of spatial variations of the geographic feature under concern by delineating natural aggregates (clusters) of its relevant environmental covariates at different scales. The natural occurrence of such aggregates is modeled using a fuzzy c-means clustering method. We iterate through different numbers of clusters from only a few to many more to be able to reveal clusters at different spatial scales. At a particular iteration, locations that bear high similarity to the cluster prototypes are identified. If a location is consistently identified at multiple iterations, it is then considered to be more representative of the general or large-scale spatial patterns. Locations that are identified less during the iterations are representative of local patterns. The integrative stepwise sampling design then gives higher sampling priority to the locations that are more representative of the large-scale patterns than local ones. We applied this sampling design in a digital soil mapping case study. Different representative samples were obtained and used for soil inference. We started with samples that are the most representative of the large-scale patterns and then gradually included the samples representative of local patterns. Field evaluation indicated that the additions of more samples with lower representativeness lead to improvements of accuracy with a decreasing marginal gain. When cost-effectiveness is considered, the representative grade could provide essential information on the number and order of samples to be sampled for an effective sampling design. Lin Yang 0018, A-Xing Zhu, Cheng-Zhi Qin 0001, T. Pei |
Int. J. Geogr. Inf. Sci. | 2 |
| 2013 | The recent advancement in digital terrain analysis and modelingabstractTerrain analysis (or geomorphometry) is ‘the science of quantitative land-surface analysis’ (cited in Pike et al. (2009, p. 3)). It collects, analyzes, evaluates, and interprets geographical inform... Qiming Zhou, A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 2 |
| 2012 | Neighborhood size and spatial scale in raster-based slope calculationsabstractRaster-based slope estimation is routine in GIS. Like many other terrain attributes, the slope at a location is determined from elevations of surrounding cells. This spatial extent – ‘neighborhood size’ – is often treated as the ‘spatial scale’ of the calculation. In fact, neighborhood size and spatial scale are two connected yet different concepts, but few studies have investigated the relationship between them. The distinction is important because neighborhood size is under user control whereas spatial scale is merely implicit in the computational method. This article attempts to clarify and provide a more precise meaning of the two terms by considering slope operators from the standpoint of the frequency (or wavenumber) domain. This article derives analytical expressions for the amplitude response functions of four popular slope estimators. These are used to characterize the individual methods and also to show that the neighborhood size and spatial scale of a slope calculation are not numerically the same. In fact, because there is no single spatial scale that can be unambiguously associated with a given neighborhood size, neighborhood size cannot be an adequate indicator of spatial scale. Furthermore, this article shows that different indices of ‘scale’ yield different impressions about the action of a slope estimator and its response to changing neighborhood size. Therefore, it is necessary to examine the amplitude response function when investigating the spatial scale. The article also provides guidance for GIS practitioners when selecting a slope estimation method. James E. Burt, A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 3 |
| 2010 | Windowed nearest neighbour method for mining spatio-temporal clusters in the presence of noiseabstractIn a spatio-temporal data set, identifying spatio-temporal clusters is difficult because of the coupling of time and space and the interference of noise. Previous methods employ either the window scanning technique or the spatio-temporal distance technique to identify spatio-temporal clusters. Although easily implemented, they suffer from the subjectivity in the choice of parameters for classification. In this article, we use the windowed kth nearest (WKN) distance (the geographic distance between an event and its kth geographical nearest neighbour among those events from which to the event the temporal distances are no larger than the half of a specified time window width [TWW]) to differentiate clusters from noise in spatio-temporal data. The windowed nearest neighbour (WNN) method is composed of four steps. The first is to construct a sequence of TWW factors, with which the WKN distances of events can be computed at different temporal scales. Second, the appropriate values of TWW (i.e. the appropriate temporal scales, at which the number of false positives may reach the lowest value when classifying the events) are indicated by the local maximum values of densities of identified clustered events, which are calculated over varying TWW by using the expectation-maximization algorithm. Third, the thresholds of the WKN distance for classification are then derived with the determined TWW. In the fourth step, clustered events identified at the determined TWW are connected into clusters according to their density connectivity in geographic–temporal space. Results of simulated data and a seismic case study showed that the WNN method is efficient in identifying spatio-temporal clusters. The novelty of WNN is that it can not only identify spatio-temporal clusters with arbitrary shapes and different spatio-temporal densities but also significantly reduce the subjectivity in the classification process. Tao Pei, Chenghu Zhou, A-Xing Zhu, Cheng-Zhi Qin 0001 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2009 | DECODE: a new method for discovering clusters of different densities in spatial data
Tao Pei, Ajay Jasra, David J. Hand, A-Xing Zhu, Chenghu Zhou |
Data Min. Knowl. Discov. | 4 |
| 2009 | Mapping with words: A new approach to automated digital soil surveyabstractSoil survey (soil mapping) is based on soil-landscape knowledge of soil scientists. Current automated approaches to soil survey cannot take such knowledge as direct input, because the knowledge is descriptive in nature. This paper presents a “mapping with words” solution by using fuzzy logic. Environmental variables used to describe landscape conditions are treated as linguistic variables. Each descriptive term used to characterize an environmental variable is treated as a fuzzy granule and is represented with a fuzzy membership function. Fuzzy membership functions are defined through gathering samples of expert perception on the landscape. Using the granule—fuzzy membership functions pairs as a dictionary, an inference can decode input descriptive knowledge accordingly and conduct soil inference. The proposed approach has been tested in a case study in Dane County, Wisconsin, USA via a soil inference approach (soil–land inference model, SoLIM). The mapping result shows that the mapping with word version of SoLIM has an 85% accuracy based on collected field points, better than a comparable earlier version (about 78%). Traditional soil survey maps usually have a mapping accuracy about 60%. The proposed methodology can be adapted to other knowledge-based natural resource mapping with slight modifications. © 2009 Wiley Periodicals, Inc. A-Xing Zhu |
Int. J. Intell. Syst. | 2 |
| 2007 | An adaptive approach to selecting a flow-partition exponent for a multiple-flow-direction algorithmabstractMost multiple‐flow‐direction algorithms (MFDs) use a flow‐partition coefficient (exponent) to determine the fractions draining to all downslope neighbours. The commonly used MFD often employs a fixed exponent over an entire watershed. The fixed coefficient strategy cannot effectively model the impact of local terrain conditions on the dispersion of local flow. This paper addresses this problem based on the idea that dispersion of local flow varies over space due to the spatial variation of local terrain conditions. Thus, the flow‐partition exponent of an MFD should also vary over space. We present an adaptive approach for determining the flow‐partition exponent based on local topographic attribute which controls local flow partitioning. In our approach, the influence of local terrain on flow partition is modelled by a flow‐partition function which is based on local maximum downslope gradient (we refer to this approach as MFD based on maximum downslope gradient, MFD‐md for short). With this new approach, a steep terrain which induces a convergent flow condition can be modelled using a large value for the flow‐partition exponent. Similarly, a gentle terrain can be modelled using a small value for the flow‐partition exponent. MFD‐md is quantitatively evaluated using four types of mathematical surfaces and their theoretical ‘true’ value of Specific Catchment Area (SCA). The Root Mean Square Error (RMSE) shows that the error of SCA computed by MFD‐md is lower than that of SCA computed by the widely used SFD and MFD algorithms. Application of the new approach using a real DEM of a watershed in Northeast China shows that the flow accumulation computed by MFD‐md is better adapted to terrain conditions based on visual judgement. Cheng-Zhi Qin 0001, A-Xing Zhu, Tao Pei, Baoluo Li, Chenghu Zhou, Lin Yang 0018 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2006 | A new approach to the nearest-neighbour method to discover cluster features in overlaid spatial point processesabstractWhen two spatial point processes are overlaid, the one with the higher rate is shown as clustered points, and the other one with the lower rate is often perceived to be background. Usually, we consider the clustered points as feature and the background as noise. Revealing these point clusters allows us to further examine and understand the spatial point process. Two important aspects in discerning spatial cluster features from a set of points are the removal of noise and the determination of the number of spatial clusters. Until now, few methods were able to deal with these two aspects at the same time in an automated way. In this study, we combine the nearest‐neighbour (NN) method and the concept of density‐connected to address these two aspects. First, the removal of noise can be achieved using the NN method; then, the number of clusters can be determined by finding the density‐connected clusters. The complexity for finding density‐connected clusters is reduced in our algorithm. Since the number of clusters depends on the value of k (the kth nearest neighbour), we introduce the concept of lifetime for the number of clusters in order to measure how stable the segmentation results (or number of clusters) are. The number of clusters with the longest lifetime is considered to be the final number of clusters. Finally, a seismic example of the west part of China is used as a case study to examine the validity of our method. In this seismic case study, we discovered three seismic clusters: one as the foreshocks of the Songpan quake (M = 7.2), and the other two as aftershocks related to the Kangding‐Jiulong (M = 6.2) quake and Daguan quake (M = 7.1), respectively. Through this case study, we conclude that the approach we proposed is effective in removing noise and determining the number of feature clusters. Tao Pei, A-Xing Zhu, Chenghu Zhou, Cheng-Zhi Qin 0001 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2003 | Knowledge discovery from soil maps using inductive learningabstractThis paper develops a knowledge discovery procedure for extracting knowledge of soil-landscape models from a soil map. It has broad relevance to knowledge discovery from other natural resource maps. The procedure consists of four major steps: data preparation, data preprocessing, pattern extraction, and knowledge consolidation. In order to recover true expert knowledge from the error-prone soil maps, our study pays specific attention to the reduction of representation noise in soil maps. The data preprocessing step has exhibited an important role in obtaining greater accuracy. A specific method for sampling pixels based on modes of environmental histograms has proven to be effective in terms of reducing noise and constructing representative sample sets. Three inductive learning algorithms, the See5 decision tree algorithm, Naïve Bayes, and artificial neural network, are investigated for a comparison concerning learning accuracy and result comprehensibility. See5 proves to be an accurate method and produces the most comprehensible results, which are consistent with the rules (expert knowledge) used in producing the soil map. The incorporation of spatial information into the knowledge discovery process is found not only to improve the accuracy of the extracted knowledge, but also to add to the explicitness and extensiveness of the extracted soil-landscape model. A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 2 |
| 1999 | A personal construct-based knowledge acquisition process for natural resource mappingabstractThis paper presents an iterative, structured knowledge-acquisition process for extracting human understanding of relationships between a natural resource and its environment. This understanding can then be used to map natural resources as spatial continua. The knowledge acquisition process is based on personal construct theory and consists of several iterations. Each iteration has five structured interview sessions: preparation, key development, description, comparison, and quantification. The knowledge derived from each iteration is represented as a set of membership functions that describes the degree to which a given environmental condition impacts the status of the given resource. The final set of membership functions, which is the final version of knowledge, is derived through the comparison and 'fusion' of the membership functions from each iteration. The comparison of the membership functions among different iterations is also used to measure the consistency (integrity) of an expert's understanding of the relationships. In a soil mapping case study, knowledge on soil-environment relationships was acquired from a local soil scientist using the knowledge acquisition process. The case study shows that knowledge sets extracted a year apart were consistent with each other. The study also shows that the soil expert was more familiar with the relationships between soils and some environmental variables than with other environmental variables. The expert's understanding about soil-environmental relationships also differed among soil series. Although it was designed to extract expert knowledge for mapping natural resources as spatial continua under a GIS environment, this knowledge elicitation process can be easily adapted to extract expert knowledge for other knowledge-based applications. A-Xing Zhu |
Int. J. Geogr. Inf. Sci. | 1 |