EDBT 2026 Demo / reviewers in the wild / expert
Tao Pei
dblp:40/6706
· DBLP profile ↗
30ranked-venue papers in the field
8as first author
11since 2021 · last 2026
0000-0002-5311-8761ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 26 (6 first)Data Mining & Knowledge Discovery · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identify different types of urban renewal implementations at the streetscape scaleabstractUnderstanding the different types of urban renewal implementation processes can inform ways to improve residents’ quality of life and optimize sustainable development strategies. Existing research has primarily focused on detecting pixel-level or object-level changes in urban physical space, but it frequently overlooks the semantic complexity inherent in urban renewal. This complexity involves an integration of what changed, where the change occurred, and how it occurred, and is important to distinguish the different types of renewal. To address this gap, this study provides a multi-type urban renewal identification (MTURI) framework using street view images (SVIs) to combine multi-source information at the street level. We constructed an SVI benchmark dataset and developed a comprehensive indicator system comprising built features, semantic attributes and street-level environmental factors. Using a two-stage recognition algorithm, the MTURI model identifies various types of urban renewal activities. We also applied the model to four renewal community cases to evaluate the potential applications of the method. The findings demonstrate that our model can effectively identify various types of renewal and provide evidence-based assessments of the effectiveness and benefits of urban policy interventions. Our paradigm introduces new tools and insights into research and practice in SVI identification. Tao Pei, Daojing Zhou, Jie Chen 0077, Dayu Cheng |
Int. J. Geogr. Inf. Sci. | 3 |
| 2025 | A novel approach for cluster detection in trajectory data with low cluster-to-noise density ratioabstractA spatial cluster of trajectories refers to objects that follow similar paths, revealing shared movement trends and aiding in anomaly detection. However, detecting clusters in trajectory data becomes challenging when the cluster-to-noise density ratio (CNDR) is low. For example, clusters in free-range sheep movements are easily seen due to their group behaviour, whereas the diversity of human movement introduces significant noise, making clustering difficult. The L-function, widely used for clustering detection in various data types (e.g. point or OD flow data), captures aggregation changes across scales without relying on predefined thresholds, offering potential for low CNDR trajectory data. Thus, we define a trajectory space to derive the Trajectory L (TL)-function for multipoint trajectories. Then we use the second derivative of the TL-function and the local TL-function to identify cluster sizes and extract clusters. Inflection points in the second derivatives enable the detection of subtle changes in aggregation, allowing for precise and sensitive cluster identification. Simulation experiments show that our method outperforms four state-of-the-art approaches in detecting clusters under low CNDR conditions while avoiding parameter dependency. We validated the generality and robustness of our method using both taxi GPS trajectories and mobile phone signalling trajectories. Furthermore, our work lays a rigorous and extensible foundation for the future formulation of spatiotemporal statistical frameworks tailored to trajectory data. Zidong Fang, Tao Pei, Xiaorui Yan, Linfeng Jiang, Hua Shu 0001, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2025 | Identification of indoor states of individuals based on mobile phone dataabstractIn urban regions, individuals predominantly spend their time indoors. Accurately identifying these indoor states is essential for many fields, such as public health and urban planning. Existing methods generally rely on sensors placed at specific locations or on volunteers’ mobile phones, limiting their applicability to those locations and a fraction of the population. To overcome these limitations, we propose a novel framework that leverages cellular signaling data—providing extensive spatio-temporal and population-wide coverage—to identify individuals’ indoor states comprehensively. We extract three types of features: interaction between individuals’ mobile phones and cells, individuals’ moving and stationary, and environmental context. Using these features, we apply three machine learning models—CatBoost, Random Forest (RF) and Support Vector Machine (SVM)—along with an interpretable machine learning model, Associative Tree (AT), to identify the indoor states. Evaluation with a ground truth dataset shows that CatBoost outperforms the other models, with an F1 score of 97.21% in quantifying the time individuals spend indoors. To our knowledge, this is the first study to identify the indoor states of individuals using cellular signaling data. We argue that this study can contribute to advancements in areas such as public health and urban planning. Linfeng Jiang, Tao Pei, Mingbo Wu, Zidong Fang, Meng Gao 0001, Xiaorui Yan, Dasheng Ge |
Int. J. Geogr. Inf. Sci. | 2 |
| 2025 | Enhanced scan statistic with tightened window for detecting irregularly shaped hotspotsabstractIn spatial point data, a hotspot is defined as a group of points with a significantly higher density within an arbitrarily shaped area. Among existing hotspot identification methods, spatial scan statistic, known for its simple mechanism in locating hotspots, has been extensively studied and applied in various fields. However, existing methods rely on pre-defined scanning window shapes, e.g. generic geometries like circles or pre-divided regions, like administrative divisions, and thereby may not accurately capture the irregular shapes of hotspots. This study enhances the spatial scan statistic by introducing a tightened window, which is defined as the window tightened to align with the hotspot’s shape. In our method, without the necessity of outlining the exact geometry, the area of the tightened window, estimated using the nearest distance statistics, is used for calculating the objective function. Experiments with simulated data demonstrate that our method outperforms existing methods in terms of testing hotspots’ significance, identifying arbitrarily shaped hotspots, estimating hotspots’ spatial extent, and reducing subjectivity in parameter selection. An empirical study using taxi pick-up point data shows our method can identify regions with high taxi demand and potential traffic congestion, including subway exits and commercial streets. Xiaorui Yan, Zhuoting Fu, Tao Pei, Zidong Fang, Meng Gao 0001, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2025 | A multivariate spatial structure indicator based on geographic similarityabstractQuantification of spatial structure reveals the distribution patterns of geographic features, which is essential for geographic analysis. For quantitative measurements of multivariate spatial structure, existing methods often neglect either the geographic meaning or the spatial combination of the multivariate dataset. In this paper, a new indicator for multivariate spatial structure (MuSS) is proposed. Since multivariate datasets characterize geographic conditions, the correlation between multivariate attributes at different locations can be measured as the similarity of geographic conditions. The MuSS indicator evaluates whether location pairs with closer distances have higher geographic similarities. Experimental results show that MuSS outperforms existing methods in differentiate multivariate datasets with varied spatial distribution patterns. A MuSS value deviating from 1 suggests that geographic similarity between location pairs is relevant to their distance, and the statistical significance of the captured distribution patterns is evaluated using a p-value from a permutation test. MuSS is also applied to real geographic data at different spatial resolutions in two study areas with diverse distribution patterns. Case studies show that MuSS provides consistent comparison results for spatial structure levels between study areas, while existing methods cannot. Fang-He Zhao, Cheng-Zhi Qin 0001, A-Xing Zhu, Tao Pei |
Int. J. Geogr. Inf. Sci. | 4 |
| 2024 | Spatiotemporal mobility network of global scientists, 1970-2020abstractThe mobility of scientists, manifested by movements to new academic institutions, grows with globalization and plays a crucial role in individual careers, institutional productivity, and knowledge dissemination. Current research on scientists’ mobility focuses on aggregated levels such as inter-country mobility, with little attention paid to fine-grained institutional level, leading to a simplified spatial portrayal of the mobility. To fill the gap, we take scientists in geography as examples, and reconstructed their dynamic mobility network among institutions from 1970 to 2020 based on massive literature metadata. Our findings reveal the spatial mobility pattern that is now dominated by North America, Western and Northern Europe, East Asia, and Oceania, with the trend of intensification, multipolarity, and inequality over time. Specifically, the mobility network exhibits clear community structure largely constrained by spatial proximity and national borders. We also uncovered a universal downward mobility pattern embedded in the hierarchical structure. Our quantitative analysis further suggest that mobility is facilitated by multiple realities, including spatial, cultural, and scientific proximity, institutional rankings and national economic levels, cooperation, and visa-free policies, with varying dynamics. These results contribute to spatiotemporal insights into the mechanisms of scientific development in theory, and the basis for talent policymaking in practice. Tao Pei, Zidong Fang, Mingbo Wu, Xiaorui Yan, Jingyu Jiang, Linfeng Jiang, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2023 | A kriging interpolation model for geographical flowsabstractThe kriging model can accommodate various spatial supports and has been extensively applied in hydrology, meteorology, soil science, and other domains. With the expansion of applications, it is essential to extend the kriging model for new spatial support of high-dimensional data. Geographical flows can depict the movements of geographical objects and imply the underlying mobility patterns in geographical phenomena. However, due to the bias, sparsity, and uneven quality of flow data in the real world, research about flows remains hindered by the lack of complete flow data and effective flow interpolation methods. In this study, we design a kriging interpolation model for flows based on several flow-related concepts and the autocorrelation of flows. We also analyze the second-order stationarity and anisotropy in the flow spatial random field. To illustrate the effectiveness and applicability of our method, we conduct two case studies. The former case study compares several experiments of flow density interpolation using Beijing mobile signaling data and illustrates the conditions of applicable areas. The latter case study extends our model to other flow attributes, such as travel time uncertainty, using Beijing taxi origin-destination flow data. The results of these cases demonstrate the effectiveness and high accuracy of our model. Ya Fang, Tao Pei, Jie Chen 0077, Yaxi Liu 0002 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2023 | Trend surface analysis of geographic flowsabstractAn origin-destination (OD) flow is the movement of objects from an origin to a destination. Determining how the flows vary across geographic locations helps understand the mechanism of flow distributions; however, it has rarely been studied. Here, we propose a trend surface model with polynomial functions to quantify the flow distribution with coordinates in the flow space. This model assumes that an observed data-record is composed of the trend value and the residual, and is represented by the orthogonal polynomial with O and D coordinates as independent variables and flow properties as dependent variables. The simulation experiments based on the linear and quadratic models indicated that the trend surface function could reflect the increasing/decreasing variation of flows with OD locations (i.e. flow trends) in different patterns. Applying this model to a case study of taxi OD flows in the broad Central Business District of Beijing, we found that the flows exhibited a rising trend toward the southwest. The trend surface characteristics are associated with the distributions of urban functional patches, where the workplaces and residences increased toward the southwest in the study area. Notably, the spatial deviations of trend surface model can help in identifying site pairs that attract flows at a high density (e.g. commerce centers and big communities), facilitating the planning of public transportation to mitigate the congestion. Beiyang Guo, Tao Pei, Hua Shu 0001, Mingbo Wu, Sihui Guo, Jingyu Jiang, Peijun Du |
Int. J. Geogr. Inf. Sci. | 2 |
| 2023 | Spatiotemporal Flow L-function: a new method for identifying spatiotemporal clusters in geographical flow dataabstractA geographical flow (hereafter flow) is defined as a movement between locations at two different times. A group of spatiotemporal flows can be viewed as a cluster if their origins and destinations are both spatiotemporally concentrated. Identifying spatiotemporal flow clusters may help reveal underlying spatiotemporal mobility trends or intensive relationships between regions. Despite recent advances in flow clustering methods, most only consider spatial attributes and ignore temporal information, and may fail to differentiate space-close but time-separated clusters. To this end, we derive global and local versions of the Spatiotemporal Flow L-function, extended from the classical L-function for points, and thereby construct a clustering method. First, the global version is utilized to check whether flow data contain clusters and estimate the spatial and temporal scales of the clusters. The local version is then employed to extract the clusters with the estimated scales. Experiments of simulated data demonstrate that our method outperforms three state-of-the-art methods in identifying spatiotemporal flow clusters with arbitrary shapes and different densities and reducing subjectivity in the parameter selection process. A case study with taxi data shows that our method reveals residents’ spatiotemporal moving patterns, including rush-hour commuting and whole-daytime transferring among railway stations. Xiaorui Yan, Tao Pei, Hua Shu 0001, Mingbo Wu, Zidong Fang, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2022 | Density-based clustering for bivariate-flow dataabstractGeographical flows reflect the movements, spatial interactions or connections among locations and are generally abstracted as origin-destination (OD) flows. In this context, clustering is a spatial pattern describing a group of flows with adjacent O and D points. For data composed of two types of flows (bivariate-flow data), a bivariate-flow cluster is a cluster comprising two types of flows, at least one of which exhibits a clustering pattern. In a bivariate-flow cluster, varying flow density combinations imply different meanings. For instance, a cluster with high-density travel flows on both weekdays (type A) and weekends (type B) may be associated with entertainment, whereas high-density flows on weekdays and sparse flows on weekends may reveal work-related travel. However, identifying bivariate-flow clusters with different flow density combinations is still an unsolved problem. To this end, we extend a bivariate-point clustering method and propose a density-based clustering method for bivariate flows. The simulation experiments verify model robustness. In a case study, we apply this method to extract clusters of bivariate-flow data comprising Beijing taxi OD flows of different periods, and identify clusters of work-related, entertainment, tourism, or egress and return travels. These results demonstrate the capability of our method in detecting bivariate-flow clusters. Hua Shu 0001, Tao Pei, Jie Chen 0077, Sihui Guo, Yaxi Liu 0002, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 2 |
| 2021 | L-function of geographical flowsabstractGeographical flow (hereafter flow) can be modeled as an orderly connected point pair composed of an origin (O) and a destination (D). Aggregation is the most common form of spatial heterogeneity of flows, which we define as their deviation from complete spatial randomness (CSR), and the aggregation scale is an important indicator for its perception. Nevertheless, quantifying the aggregation scale of flows is still an unsolved problem. In this paper, we propose the L-function for flows as a solution, derive theoretical null models of the K-function and L-function in a flow space. We conduct simulation experiments to validate the L-function and its capability to detect aggregation scales. Finally, we apply the solution in a case study with taxi data in Beijing and identify nine aggregation scales of taxi OD flows, ranging from 170 m to 22.1 km. These scales correspond to three classes: less than 300 m, from 600 m to 700 m and more than 1500 m. The classes are related to the sizes of the urban facilities where the dominant flow clusters occur, indicating that the L-function in flow space can detect the aggregation scale of flows at the building scale, the block scale and the district scale. Hua Shu 0001, Tao Pei, Sihui Guo, Yaxi Liu 0002, Jie Chen 0077, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 2 |
| 2019 | A proportional odds model of human mobility and migration patternsabstractThe modelling of human mobility and migration patterns has received much attention due to its substantial importance. Despite long-term efforts, we still lack a modelling framework that captures mobility patterns and further obtains a prospective view of movement trends with regards to diverse impacting factors. Here, we propose a proportional odds model of human mobility and migration (POM-HM) that takes a probabilistic approach to model human movements. Our model is based on the migration probability with a log-logistic distribution under the proportional odds assumption. Explanatory variables are introduced into the model by re-parameterizing the probability distribution function. The two resultant functions, namely, the migration strength and cumulative hazard, are used to estimate regional differences among travel fluxes and their tendencies. The performance of the POM-HM in terms of its validity and accuracy is examined and compared with the gravity model and the radiation model. The probability-based modelling framework enables us to investigate regional variations in migrant fluxes consequently further predict potential future patterns. In short, our modelling approach captures the probabilistic nature of human mobility and migration and furthers our understanding of both the spatiotemporal patterns of population movements and the impacts of various driving forces. Ting Ma 0002, Jianghao Wang, Tao Pei, Yunyan Du, Chenghu Zhou, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 5 |
| 2019 | Quantifying the spatial heterogeneity of pointsabstractVariation in the spatial heterogeneity of points reflects the evolutionary process or mechanism of geographical events. The key to depicting this variation is quantifying spatial heterogeneity. In this paper, the spatial heterogeneity of a point pattern is defined as the degree of aggregation-type deviation from complete spatial randomness. In such a case, a goodness-of-fit-type statistic based on the distribution of nearest-neighbor distances called the level of heterogeneity (LH*) is regarded as a standard measurement, and a normalized version called the normalized level of heterogeneity (NLH*) is proposed for datasets with different point numbers and study region areas. Considering the complex integration calculation of LH* and NLH*, simulation experiments are implemented to test the capability of some classic nearest-neighbor statistics in quantifying spatial heterogeneity. The results showed that except for the standard LH* statistic, only Clark and Evans’ statistic (A-w) and Byth and Ripley’s statistic (H-xw) are robust. Statistics NLH*, (A-w) and (H-xw) are validated by quantifying the spatial heterogeneity of two-dimensional crime events, three-dimensional earthquake events and four-dimensional origin-destination (OD) events. The results indicate that these statistics all have a reasonable explanation in quantifying spatial heterogeneity for real-world geographical events of different types and with different dimensions. Compared with NLH*, Clark and Evans’ (A-w) statistic and Byth and Ripley’s (H-xw) statistic are recommended from the perspective of accessibility. Hua Shu 0001, Tao Pei, Ting Ma 0002, Yunyan Du, Zide Fan, Sihui Guo |
Int. J. Geogr. Inf. Sci. | 2 |
| 2019 | Detecting arbitrarily shaped clusters in origin-destination flows using ant colony optimizationabstractAn origin-destination (OD) flow can be defined as the movement of objects between two locations. These movements must be determined for a range of purposes, and strong interactions can be visually represented via clustering of OD flows. Identification of such clusters may be useful in urban planning, traffic planning and logistics management research. However, few methods can identify arbitrarily shaped flow clusters. Here, we present a spatial scan statistical approach based on ant colony optimization (ACO) for detecting arbitrarily shaped clusters of OD flows (AntScan_flow). In this study, an OD flow cluster is defined as a regional pair with significant log likelihood ratio (LLR), and the ACO is employed to detect the clusters with maximum LLRs in the search space. Simulation experiments based on AntScan_flow and SaTScan_flow show that AntScan_flow yields better performance based on accuracy but requires a large computational demand. Finally, a case study of the morning commuting flows of Beijing residents was conducted. The AntScan_flow results show that the regions associated with moderate- and long-distance commuting OD flow clusters are highly consistent with subway lines and highways in the city. Additionally, the regions of short-distance commuting OD flow clusters are more likely to exhibit ‘residential-area to work-area’ patterns. Tao Pei, Ting Ma 0002, Yunyan Du, Hua Shu 0001, Sihui Guo, Zide Fan |
Int. J. Geogr. Inf. Sci. | 2 |
| 2018 | Integrating spatial and temporal contexts into a factorization model for POI recommendationabstractMatrix factorization is one of the most popular methods in recommendation systems. However, it faces two challenges related to the check-in data in point of interest (POI) recommendation: data scarcity and implicit feedback. To solve these problems, we propose a Feature-Space Separated Factorization Model (FSS-FM) in this paper. The model represents the POI feature spaces as separate slices, each of which represents a type of feature. Thus, spatial and temporal information and other contexts can be easily added to compensate for scarce data. Moreover, two commonly used objective functions for the factorization model, the weighted least squares and pairwise ranking functions, are combined to construct a hybrid optimization function. Extensive experiments are conducted on two real-life data sets: Gowalla and Foursquare, and the results are compared with those of baseline methods to evaluate the model. The results suggest that the FSS-FM performs better than state-of-the-art methods in terms of precision and recall on both data sets. The model with separate feature spaces can improve the performance of recommendation. The inclusion of spatial and temporal contexts further leverages the performance, and the spatial context is more influential than the temporal context. In addition, the capacity of hybrid optimization in improving POI recommendation is demonstrated. Jun Xu 0020, Tao Pei |
Int. J. Geogr. Inf. Sci. | 4 |
| 2018 | Fine-grained prediction of urban population using mobile phone location dataabstractFine-grained prediction of urban population is of great practical significance in many domains that require temporally and spatially detailed population information. However, fine-grained population modeling has been challenging because the urban population is highly dynamic and its mobility pattern is complex in space and time. In this study, we propose a method to predict the population at a large spatiotemporal scale in a city. This method models the temporal dependency of population by estimating the future inflow population with the current inflow pattern and models the spatial correlation of population using an artificial neural network. With a large dataset of mobile phone locations, the model’s prediction error is low and only increases gradually as the temporal prediction granularity increases, and this model is adaptive to sudden changes in population caused by special events. Jie Chen 0077, Tao Pei, Shih-Lung Shaw, Feng Lu 0004, Mingxiao Li 0001, Shifen Cheng, Xiliang Liu, Hengcai Zhang |
Int. J. Geogr. Inf. Sci. | 2 |
| 2018 | Local multi-feature hashing based fast matching for aerial images
Suting Chen, Yunjiao Shi, Jiaojiao Zhao, Tao Pei |
Inf. Sci. | 6 |
| 2016 | A new assessment model for evacuation vulnerability in urban areasabstractIn the high-speed urbanization process of China, the urban population has been increasing significantly, leading to a high-density aggregation of population. However, the sharp increase in population density has not produced commensurate improvements in the road networks. On the contrary, the population increase induced a serious evacuation vulnerability, which cities experience during various hazards and catastrophic events. Therefore, research on evacuation vulnerability is important to urban planning. To assess the evacuation vulnerability, the optimal and worst scenarios should be considered because all possible evacuation plans occur between these extremes. However, most previous evacuation vulnerability studies are based on the worst-case scenario, only providing an upper bound of a potential evacuation assessment. To provide a more comprehensive theoretical basis for decision-makers to understand the consequences caused by all possible evacuations, this paper proposes an optimal evacuation vulnerability assessment model that provides the lower bound on potential evacuation difficulties. The model is solved by a stepwise spreading algorithm based on Graph Theory. Subsequently, to evaluate the effectiveness of the model, the study adopts the model to assess the evacuation capability of different road network topologies. A comparison with previous research was performed. The model was demonstrated in an application to the South Luogu Alley of Beijing, China. The significance of this paper is that the combination of our model with previous research may provide a more complete theoretical basis for an evacuation vulnerability assessment. Xiaoyi Ma, Tao Pei, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 2 |
| 2015 | Density-based clustering for data containing two types of pointsabstractWhen only one type of point is distributed in a region, clustered points can be seen as an anomaly. When two different types of points coexist in a region, they overlap at different places with various densities. In such cases, the meaning of a cluster of one type of point may be altered if points of the other type show different densities within the same cluster. If we consider the origins and destinations (OD) of taxicab trips, the clustering of both in the morning may indicate a transportation hub, whereas clustered origins and sparse destinations (a hot spot where taxis are in short supply) could suggest a densely populated residential area. This cannot be identified by previous clustering methods, so it is worthwhile studying a clustering method for two types of points. The concept of two-component clustering is first defined in this paper as a group containing two types of points, at least one of which exhibits clustering. We then propose a density-based method for identifying two-component clusters. The method is divided into four steps. The first estimates the clustering scale of the point data. The second transforms the point data into the 2D density domain, where the x and y axes represent the local density of each type of point around each point, respectively. The third determines the thresholds for extracting the clusters, and the fourth generates two-component clusters using a density-connectivity mechanism. The method is applied to taxicab trip data in Beijing. Three types of two-component clusters are identified: high-density origins and destinations, high-density origins and low-density destinations, and low-density origins and high-density destinations. The clustering results are verified by the spatial relationship between the cluster locations and their land-use types over different periods of the day. Tao Pei, Hengcai Zhang, Ting Ma 0002, Yunyan Du, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 1 |
| 2015 | A citizen data-based approach to predictive mapping of spatial variation of natural phenomenaabstractThe vast accumulation of environmental data and the rapid development of geospatial visualization and analytical techniques make it possible for scientists to solicit information from local citizens to map spatial variation of geographic phenomena. However, data provided by citizens (referred to as citizen data in this article) suffer two limitations for mapping: bias in spatial coverage and imprecision in spatial location. This article presents an approach to minimizing the impacts of these two limitations of citizen data using geospatial analysis techniques. The approach reduces location imprecision by adopting a frequency-sampling strategy to identify representative presence locations from areas over which citizens observed the geographic phenomenon. The approach compensates for the spatial bias by weighting presence locations with cumulative visibility (the frequency at which a given location can be seen by local citizens). As a case study to demonstrate the principle, this approach was applied to map the habitat suitability of the black-and-white snub-nosed monkey (Rhinopithecus bieti) in Yunnan, China. Sightings of R. bieti were elicited from local citizens using a geovisualization platform and then processed with the proposed approach to predict a habitat suitability map. Presence locations of R. bieti recorded by biologists through intensive field tracking were used to validate the predicted habitat suitability map. Validation showed that the continuous Boyce index (Bcont(0.1)) calculated on the suitability map was 0.873 (95% CI: [0.810, 0.917]), indicating that the map was highly consistent with the field-observed distribution of R. bieti. Bcont(0.1) was much lower (0.173) for the suitability map predicted based on citizen data when location imprecision was not reduced and even lower (−0.048) when there was no compensation for spatial bias. This indicates that the proposed approach effectively minimized the impacts of location imprecision and spatial bias in citizen data and therefore effectively improved the quality of mapped spatial variation using citizen data. It further implies that, with the application of geospatial analysis techniques to properly account for limitations in citizen data, valuable information embedded in such data can be extracted and used for scientific mapping. A-Xing Zhu, Guiming Zhang, Zhi-Pang Huang, Ge-Sang Dunzhu, Guopeng Ren, Cheng-Zhi Qin 0001, Lin Yang 0018, Tao Pei, Shengtian Yang |
Int. J. Geogr. Inf. Sci. | 10 |
| 2014 | A new insight into land use classification based on aggregated mobile phone dataabstractLand-use classification is essential for urban planning. Urban land-use types can be differentiated either by their physical characteristics (such as reflectivity and texture) or social functions. Remote sensing techniques have been recognized as a vital method for urban land-use classification because of their ability to capture the physical characteristics of land use. Although significant progress has been achieved in remote sensing methods designed for urban land-use classification, most techniques focus on physical characteristics, whereas knowledge of social functions is not adequately used. Owing to the wide usage of mobile phones, the activities of residents, which can be retrieved from the mobile phone data, can be determined in order to indicate the social function of land use. This could bring about the opportunity to derive land-use information from mobile phone data. To verify the application of this new data source to urban land-use classification, we first construct a vector of aggregated mobile phone data to characterize land-use types. This vector is composed of two aspects: the normalized hourly call volume and the total call volume. A semi-supervised fuzzy c-means clustering approach is then applied to infer the land-use types. The method is validated using mobile phone data collected in Singapore. Land use is determined with a detection rate of 58.03%. An analysis of the land-use classification results shows that the detection rate decreases as the heterogeneity of land use increases, and increases as the density of cell phone towers increases. Tao Pei, Stanislav Sobolevsky, Carlo Ratti, Shih-Lung Shaw, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 1 |
| 2013 | Clustering of temporal event processesabstractA temporal point process is a sequence of points, each representing the occurrence time of an event. Each temporal point process is related to the behavior of an entity. As a result, clustering of temporal point processes can help differentiate between entities, thereby revealing patterns of behaviors. This study proposes a hierarchical cluster method for clustering temporal point processes based on the discrete Fréchet (DF) distance. The DF cluster method is divided into four steps: (1) constructing a DF similarity matrix between temporal point processes; (2) constructing a complete linkage hierarchical tree based on the DF similarity matrix; (3) clustering the point processes with a threshold determined by locating the local maxima on the curve of the pseudo-F statistic (an index which measures the separability between clusters and the compactness in clusters); and (4) identifying inner patterns for each cluster formed by a series of dense intervals, each of which contains at least one event of all processes of the cluster. The contributions of the article are: (1) the proposed DF cluster method can cluster temporal point processes into different groups and (2) more importantly, it can identify the inner pattern of each cluster. Two synthetic data sets were created to illustrate the DF distance between temporal point process clusters (the first data set) and validate the proposed DF cluster method (the second data set), respectively. An experiment and a comparison with a method based on dynamic time warping show that DF cluster successfully identifies the preconfigured patterns in the second synthetic data set. The cluster method was then applied to a population migration history data set for the Northern Plains of the United States, revealing some interesting population migration patterns. Tao Pei, Xi Gong, Shih-Lung Shaw, Ting Ma 0002, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 1 |
| 2013 | Efficient encoding and spatial operation scheme for aperture 4 hexagonal discrete global grid systemabstractDiscrete global grid systems (DGGSs) are considered to be promising structures for global geospatial information representation. Square and triangular DGGSs have had the advantage over hexagonal ones in geospatial data processing over the past few decades. Despite a significant body of research supporting hexagonal grids as the superior alternative, the application thereof has been hindered partly owing to the lack of a hierarchy. This study presents an original perspective to combine two types of aperture 4 hexagonal discrete grid systems into a hierarchy. Each cell of the hierarchy is assigned a unique code using a linear quadtree that constructs the hexagonal quaternary balanced structure (HQBS). The mathematical system described by HQBS addressing and the vector operations, including addition, subtraction, multiplication, and division, are defined. Essential spatial operations for HQBS cell retrieval, transformation between HQBS codes and other coordinate systems, and arrangement of HQBS cells on spherical surfaces were studied and implemented. The accuracy and efficiency of algorithms were validated through experiments. The results indicate that the average efficiency of cell retrieval using the HQBS is higher than that using other schemes, thus proving it to be more efficient. Xiaochong Tong, Jin Ben, Tao Pei |
Int. J. Geogr. Inf. Sci. | 5 |
| 2012 | Multi-scale decomposition of point process data
Tao Pei, Jianhuan Gao, Ting Ma 0002, Chenghu Zhou |
GeoInformatica | 1 |
| 2011 | Detecting arbitrarily shaped clusters using ant colony optimizationabstractIn the map of geo-referenced population and cases, the detection of the most likely cluster (MLC), which is made up of many connected polygons (e.g., the boundaries of census tracts), may face two difficulties. One is the irregularity of the shape of the cluster and the other is the heterogeneity of the cluster. A heterogeneous cluster is referred to as the cluster containing depression links (a polygon is a depression link if it satisfies two conditions: (1) the ratio between the case number and the population in the polygon is below the average ratio of the whole map; (2) the removal of the polygon will disconnect the cluster). Previous studies have successfully solved the problem of detecting arbitrarily shaped clusters not containing depression links. However, for a heterogeneous cluster, existing methods may generate mistakes, for example, missing some parts of the cluster. In this article, a spatial scanning method based on the ant colony optimization (AntScan) is proposed to improve the detection power. If a polygon can be simplified as a node, the research area consisting of many polygons then can be seen as a graph. So the detection of the MLC can be seen as the search of the best subgraph (with the largest likelihood value) in the graph. The comparison between AntScan, GAScan (the spatial scan method based on the genetic optimization), and SAScan (the spatial scan method based on the simulated annealing optimization) indicates that (1) the performance of GAScan and SAScan is significantly influenced by the parameter of the fraction value (the maximum allowed size of the detected cluster), which can only be estimated by multiple trials, while no such parameter is needed in AntScan; (2) AntScan shows superior power over GAScan and SAScan in detecting heterogeneous clusters. The case study on esophageal cancer in North China demonstrates that the cluster identified by AntScan has the larger likelihood value than that detected by SAScan and covers all high-risk regions of esophageal cancer whereas SAScan misses some high-risk regions (the region in the southwest of Shandong province, eastern China) due to the existence of a depression link. Tao Pei, You Wan, Yong Jiang 0002, Chenxu Qu, Chenghu Zhou, Youlin Qiao |
Int. J. Geogr. Inf. Sci. | 1 |
| 2010 | Windowed nearest neighbour method for mining spatio-temporal clusters in the presence of noiseabstractIn a spatio-temporal data set, identifying spatio-temporal clusters is difficult because of the coupling of time and space and the interference of noise. Previous methods employ either the window scanning technique or the spatio-temporal distance technique to identify spatio-temporal clusters. Although easily implemented, they suffer from the subjectivity in the choice of parameters for classification. In this article, we use the windowed kth nearest (WKN) distance (the geographic distance between an event and its kth geographical nearest neighbour among those events from which to the event the temporal distances are no larger than the half of a specified time window width [TWW]) to differentiate clusters from noise in spatio-temporal data. The windowed nearest neighbour (WNN) method is composed of four steps. The first is to construct a sequence of TWW factors, with which the WKN distances of events can be computed at different temporal scales. Second, the appropriate values of TWW (i.e. the appropriate temporal scales, at which the number of false positives may reach the lowest value when classifying the events) are indicated by the local maximum values of densities of identified clustered events, which are calculated over varying TWW by using the expectation-maximization algorithm. Third, the thresholds of the WKN distance for classification are then derived with the determined TWW. In the fourth step, clustered events identified at the determined TWW are connected into clusters according to their density connectivity in geographic–temporal space. Results of simulated data and a seismic case study showed that the WNN method is efficient in identifying spatio-temporal clusters. The novelty of WNN is that it can not only identify spatio-temporal clusters with arbitrary shapes and different spatio-temporal densities but also significantly reduce the subjectivity in the classification process. Tao Pei, Chenghu Zhou, A-Xing Zhu, Cheng-Zhi Qin 0001 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2009 | DECODE: a new method for discovering clusters of different densities in spatial data
Tao Pei, Ajay Jasra, David J. Hand, A-Xing Zhu, Chenghu Zhou |
Data Min. Knowl. Discov. | 1 |
| 2007 | An adaptive approach to selecting a flow-partition exponent for a multiple-flow-direction algorithmabstractMost multiple‐flow‐direction algorithms (MFDs) use a flow‐partition coefficient (exponent) to determine the fractions draining to all downslope neighbours. The commonly used MFD often employs a fixed exponent over an entire watershed. The fixed coefficient strategy cannot effectively model the impact of local terrain conditions on the dispersion of local flow. This paper addresses this problem based on the idea that dispersion of local flow varies over space due to the spatial variation of local terrain conditions. Thus, the flow‐partition exponent of an MFD should also vary over space. We present an adaptive approach for determining the flow‐partition exponent based on local topographic attribute which controls local flow partitioning. In our approach, the influence of local terrain on flow partition is modelled by a flow‐partition function which is based on local maximum downslope gradient (we refer to this approach as MFD based on maximum downslope gradient, MFD‐md for short). With this new approach, a steep terrain which induces a convergent flow condition can be modelled using a large value for the flow‐partition exponent. Similarly, a gentle terrain can be modelled using a small value for the flow‐partition exponent. MFD‐md is quantitatively evaluated using four types of mathematical surfaces and their theoretical ‘true’ value of Specific Catchment Area (SCA). The Root Mean Square Error (RMSE) shows that the error of SCA computed by MFD‐md is lower than that of SCA computed by the widely used SFD and MFD algorithms. Application of the new approach using a real DEM of a watershed in Northeast China shows that the flow accumulation computed by MFD‐md is better adapted to terrain conditions based on visual judgement. Cheng-Zhi Qin 0001, A-Xing Zhu, Tao Pei, Baoluo Li, Chenghu Zhou, Lin Yang 0018 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2006 | A Mathematical Morphology Based Scale Space Method for the Mining of Linear Features in Geographic Data
Yee Leung, Chenghu Zhou, Tao Pei, Jiancheng Luo |
Data Min. Knowl. Discov. | 4 |
| 2006 | A new approach to the nearest-neighbour method to discover cluster features in overlaid spatial point processesabstractWhen two spatial point processes are overlaid, the one with the higher rate is shown as clustered points, and the other one with the lower rate is often perceived to be background. Usually, we consider the clustered points as feature and the background as noise. Revealing these point clusters allows us to further examine and understand the spatial point process. Two important aspects in discerning spatial cluster features from a set of points are the removal of noise and the determination of the number of spatial clusters. Until now, few methods were able to deal with these two aspects at the same time in an automated way. In this study, we combine the nearest‐neighbour (NN) method and the concept of density‐connected to address these two aspects. First, the removal of noise can be achieved using the NN method; then, the number of clusters can be determined by finding the density‐connected clusters. The complexity for finding density‐connected clusters is reduced in our algorithm. Since the number of clusters depends on the value of k (the kth nearest neighbour), we introduce the concept of lifetime for the number of clusters in order to measure how stable the segmentation results (or number of clusters) are. The number of clusters with the longest lifetime is considered to be the final number of clusters. Finally, a seismic example of the west part of China is used as a case study to examine the validity of our method. In this seismic case study, we discovered three seismic clusters: one as the foreshocks of the Songpan quake (M = 7.2), and the other two as aftershocks related to the Kangding‐Jiulong (M = 6.2) quake and Daguan quake (M = 7.1), respectively. Through this case study, we conclude that the approach we proposed is effective in removing noise and determining the number of feature clusters. Tao Pei, A-Xing Zhu, Chenghu Zhou, Cheng-Zhi Qin 0001 |
Int. J. Geogr. Inf. Sci. | 1 |