EDBT 2026 Demo / reviewers in the wild / expert
Hua Shu 0001
dblp:06/162-1
· DBLP profile ↗
7ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0001-7694-564XORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 7 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A novel approach for cluster detection in trajectory data with low cluster-to-noise density ratioabstractA spatial cluster of trajectories refers to objects that follow similar paths, revealing shared movement trends and aiding in anomaly detection. However, detecting clusters in trajectory data becomes challenging when the cluster-to-noise density ratio (CNDR) is low. For example, clusters in free-range sheep movements are easily seen due to their group behaviour, whereas the diversity of human movement introduces significant noise, making clustering difficult. The L-function, widely used for clustering detection in various data types (e.g. point or OD flow data), captures aggregation changes across scales without relying on predefined thresholds, offering potential for low CNDR trajectory data. Thus, we define a trajectory space to derive the Trajectory L (TL)-function for multipoint trajectories. Then we use the second derivative of the TL-function and the local TL-function to identify cluster sizes and extract clusters. Inflection points in the second derivatives enable the detection of subtle changes in aggregation, allowing for precise and sensitive cluster identification. Simulation experiments show that our method outperforms four state-of-the-art approaches in detecting clusters under low CNDR conditions while avoiding parameter dependency. We validated the generality and robustness of our method using both taxi GPS trajectories and mobile phone signalling trajectories. Furthermore, our work lays a rigorous and extensible foundation for the future formulation of spatiotemporal statistical frameworks tailored to trajectory data. Zidong Fang, Tao Pei, Xiaorui Yan, Linfeng Jiang, Hua Shu 0001, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 8 |
| 2023 | Trend surface analysis of geographic flowsabstractAn origin-destination (OD) flow is the movement of objects from an origin to a destination. Determining how the flows vary across geographic locations helps understand the mechanism of flow distributions; however, it has rarely been studied. Here, we propose a trend surface model with polynomial functions to quantify the flow distribution with coordinates in the flow space. This model assumes that an observed data-record is composed of the trend value and the residual, and is represented by the orthogonal polynomial with O and D coordinates as independent variables and flow properties as dependent variables. The simulation experiments based on the linear and quadratic models indicated that the trend surface function could reflect the increasing/decreasing variation of flows with OD locations (i.e. flow trends) in different patterns. Applying this model to a case study of taxi OD flows in the broad Central Business District of Beijing, we found that the flows exhibited a rising trend toward the southwest. The trend surface characteristics are associated with the distributions of urban functional patches, where the workplaces and residences increased toward the southwest in the study area. Notably, the spatial deviations of trend surface model can help in identifying site pairs that attract flows at a high density (e.g. commerce centers and big communities), facilitating the planning of public transportation to mitigate the congestion. Beiyang Guo, Tao Pei, Hua Shu 0001, Mingbo Wu, Sihui Guo, Jingyu Jiang, Peijun Du |
Int. J. Geogr. Inf. Sci. | 4 |
| 2023 | Spatiotemporal Flow L-function: a new method for identifying spatiotemporal clusters in geographical flow dataabstractA geographical flow (hereafter flow) is defined as a movement between locations at two different times. A group of spatiotemporal flows can be viewed as a cluster if their origins and destinations are both spatiotemporally concentrated. Identifying spatiotemporal flow clusters may help reveal underlying spatiotemporal mobility trends or intensive relationships between regions. Despite recent advances in flow clustering methods, most only consider spatial attributes and ignore temporal information, and may fail to differentiate space-close but time-separated clusters. To this end, we derive global and local versions of the Spatiotemporal Flow L-function, extended from the classical L-function for points, and thereby construct a clustering method. First, the global version is utilized to check whether flow data contain clusters and estimate the spatial and temporal scales of the clusters. The local version is then employed to extract the clusters with the estimated scales. Experiments of simulated data demonstrate that our method outperforms three state-of-the-art methods in identifying spatiotemporal flow clusters with arbitrary shapes and different densities and reducing subjectivity in the parameter selection process. A case study with taxi data shows that our method reveals residents’ spatiotemporal moving patterns, including rush-hour commuting and whole-daytime transferring among railway stations. Xiaorui Yan, Tao Pei, Hua Shu 0001, Mingbo Wu, Zidong Fang, Jie Chen 0077 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2022 | Density-based clustering for bivariate-flow dataabstractGeographical flows reflect the movements, spatial interactions or connections among locations and are generally abstracted as origin-destination (OD) flows. In this context, clustering is a spatial pattern describing a group of flows with adjacent O and D points. For data composed of two types of flows (bivariate-flow data), a bivariate-flow cluster is a cluster comprising two types of flows, at least one of which exhibits a clustering pattern. In a bivariate-flow cluster, varying flow density combinations imply different meanings. For instance, a cluster with high-density travel flows on both weekdays (type A) and weekends (type B) may be associated with entertainment, whereas high-density flows on weekdays and sparse flows on weekends may reveal work-related travel. However, identifying bivariate-flow clusters with different flow density combinations is still an unsolved problem. To this end, we extend a bivariate-point clustering method and propose a density-based clustering method for bivariate flows. The simulation experiments verify model robustness. In a case study, we apply this method to extract clusters of bivariate-flow data comprising Beijing taxi OD flows of different periods, and identify clusters of work-related, entertainment, tourism, or egress and return travels. These results demonstrate the capability of our method in detecting bivariate-flow clusters. Hua Shu 0001, Tao Pei, Jie Chen 0077, Sihui Guo, Yaxi Liu 0002, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 1 |
| 2021 | L-function of geographical flowsabstractGeographical flow (hereafter flow) can be modeled as an orderly connected point pair composed of an origin (O) and a destination (D). Aggregation is the most common form of spatial heterogeneity of flows, which we define as their deviation from complete spatial randomness (CSR), and the aggregation scale is an important indicator for its perception. Nevertheless, quantifying the aggregation scale of flows is still an unsolved problem. In this paper, we propose the L-function for flows as a solution, derive theoretical null models of the K-function and L-function in a flow space. We conduct simulation experiments to validate the L-function and its capability to detect aggregation scales. Finally, we apply the solution in a case study with taxi data in Beijing and identify nine aggregation scales of taxi OD flows, ranging from 170 m to 22.1 km. These scales correspond to three classes: less than 300 m, from 600 m to 700 m and more than 1500 m. The classes are related to the sizes of the urban facilities where the dominant flow clusters occur, indicating that the L-function in flow space can detect the aggregation scale of flows at the building scale, the block scale and the district scale. Hua Shu 0001, Tao Pei, Sihui Guo, Yaxi Liu 0002, Jie Chen 0077, Chenghu Zhou |
Int. J. Geogr. Inf. Sci. | 1 |
| 2019 | Quantifying the spatial heterogeneity of pointsabstractVariation in the spatial heterogeneity of points reflects the evolutionary process or mechanism of geographical events. The key to depicting this variation is quantifying spatial heterogeneity. In this paper, the spatial heterogeneity of a point pattern is defined as the degree of aggregation-type deviation from complete spatial randomness. In such a case, a goodness-of-fit-type statistic based on the distribution of nearest-neighbor distances called the level of heterogeneity (LH*) is regarded as a standard measurement, and a normalized version called the normalized level of heterogeneity (NLH*) is proposed for datasets with different point numbers and study region areas. Considering the complex integration calculation of LH* and NLH*, simulation experiments are implemented to test the capability of some classic nearest-neighbor statistics in quantifying spatial heterogeneity. The results showed that except for the standard LH* statistic, only Clark and Evans’ statistic (A-w) and Byth and Ripley’s statistic (H-xw) are robust. Statistics NLH*, (A-w) and (H-xw) are validated by quantifying the spatial heterogeneity of two-dimensional crime events, three-dimensional earthquake events and four-dimensional origin-destination (OD) events. The results indicate that these statistics all have a reasonable explanation in quantifying spatial heterogeneity for real-world geographical events of different types and with different dimensions. Compared with NLH*, Clark and Evans’ (A-w) statistic and Byth and Ripley’s (H-xw) statistic are recommended from the perspective of accessibility. Hua Shu 0001, Tao Pei, Ting Ma 0002, Yunyan Du, Zide Fan, Sihui Guo |
Int. J. Geogr. Inf. Sci. | 1 |
| 2019 | Detecting arbitrarily shaped clusters in origin-destination flows using ant colony optimizationabstractAn origin-destination (OD) flow can be defined as the movement of objects between two locations. These movements must be determined for a range of purposes, and strong interactions can be visually represented via clustering of OD flows. Identification of such clusters may be useful in urban planning, traffic planning and logistics management research. However, few methods can identify arbitrarily shaped flow clusters. Here, we present a spatial scan statistical approach based on ant colony optimization (ACO) for detecting arbitrarily shaped clusters of OD flows (AntScan_flow). In this study, an OD flow cluster is defined as a regional pair with significant log likelihood ratio (LLR), and the ACO is employed to detect the clusters with maximum LLRs in the search space. Simulation experiments based on AntScan_flow and SaTScan_flow show that AntScan_flow yields better performance based on accuracy but requires a large computational demand. Finally, a case study of the morning commuting flows of Beijing residents was conducted. The AntScan_flow results show that the regions associated with moderate- and long-distance commuting OD flow clusters are highly consistent with subway lines and highways in the city. Additionally, the regions of short-distance commuting OD flow clusters are more likely to exhibit ‘residential-area to work-area’ patterns. Tao Pei, Ting Ma 0002, Yunyan Du, Hua Shu 0001, Sihui Guo, Zide Fan |
Int. J. Geogr. Inf. Sci. | 5 |