EDBT 2026 Demo / reviewers in the wild / expert
Song Gao 0001
dblp:92/357-1
· DBLP profile ↗
31ranked-venue papers in the field
3as first author
21since 2021 · last 2026
0000-0003-4359-6302ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 28 (2 first)Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeoXCP: uncertainty quantification of spatial explanations in explainable AIabstractUnderstanding and explaining complex geographic phenomena—ranging from climate change to socioeconomic disparities—is a central focus in both geography and the broader scientific community. Various methods have been developed to elucidate relationships between variables, from coefficient estimates in linear regression models to the increasingly dominant use of feature attribution scores in Explainable AI (XAI) techniques. However, explanations generated by XAI methods often carry uncertainty, stemming from the model itself and the data used to train the model. Despite the critical importance of accounting for such uncertainty, this issue remains largely overlooked in the geospatial domain. In this study, we developed an uncertainty quantification framework for XAI explanations based on conformal prediction, termed Geospatial eXplanation Conformal Prediction (GeoXCP). By incorporating spatial dependence into the modeling process, GeoXCP produced spatially adaptive explanations with calibrated uncertainty estimates. We validated the effectiveness of GeoXCP through extensive simulation experiments and real-world datasets. The results demonstrated that GeoXCP provided reliable explanations while effectively quantifying uncertainty across diverse geospatial scenarios. Our approach represented a significant advancement in explainable geospatial machine learning, enabling decision-makers to better assess the trustworthiness of model-driven insights. The proposed framework was implemented in a python package, named GeoXCP. Xiayin Lou, Peng Luo 0001, Song Gao 0001, Liqiu Meng |
Int. J. Geogr. Inf. Sci. | 4 |
| 2025 | Whose Truth? Pluralistic Geo-Alignment for (Agentic) AIabstractAI alignment describes the challenge of ensuring (future) AI systems behave in accordance with societal norms, values, and goals. Alignment is now central to research on foundation models and AI agents. Most recent work focuses on methods to prevent potentially harmful biases, account for social inequalities, improve AI safety, and enhance explainability. Notably, the debiasing 'corrections' applied to various stages of AI/ML workflows may lead to outcomes that diverge strongly from current statistical realities on the ground. For instance, text-to-image models may depict a balanced gender ratio of company leadership, despite existing imbalances. However, an often overlooked dimension is the geographic variability of alignment. What is considered appropriate, truthful, or legal can vary greatly between regions due to cultural differences, political realities, or legislation. Hence, some model outputs align without further knowledge of the user's geospatial context, while others are highly sensitive to it. Put differently, whether these outputs align varies geographically. E.g., statements about Kashmir cannot be generated without understanding the user's origin and current location. From a common-sense perspective, this problem is hardly new. In fact, Google Maps will render different administrative borders based on the user's location. Interestingly, in both knowledge representation and representation learning, spatiotemporal context, e.g., due to the monotonic nature of reasoning, remains a major challenge. Until very recently, these were largely theoretical problems. What is truly novel is the scale and level of automation at which AI systems now mediate knowledge, express opinions, and represent reality to millions of users across borders, often with little transparency or oversight regarding how context is handled. With agentic AI on the horizon, the urgency for pluralistic, geographically aware alignment, rather than one-size-fits-all solutions, is growing. Here, we motivate and formalize the vision of geo-alignment, outline how it goes beyond pluralistic alignment by offering learnable spatially explicit patterns, and suggest concrete avenues for future research. Krzysztof Janowicz, Zilong Liu 0003, Gengchen Mai, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao 0001 |
SIGSPATIAL/GIS | 8 |
| 2025 | Scalable Inter-County Food Flow Prediction Using Graph Neural NetworksabstractUnderstanding the spatial distribution of food flows is critical for assessing commodity supply chain resilience, infrastructure planning, and environmental impact. However, high-resolution trade data, especially at the county level and across specific commodity groups, are rarely available. In this study, we present a deep learning framework based on a Graph Neural Network (GNN) to predict food flows at the US county level across seven major commodity categories (Standard Classification of Transported Goods(SCTG) 01-07). Our approach trains GNN models solely on the coarser zone-level data of the Freight Analysis Framework (FAF), leveraging node attributes and inter-regional connectivity to enable cross-scale generalization. Each commodity category is modeled independently and the predictions are evaluated against the FAF5 county-level benchmarks. To support spatial generalization, our model integrates detailed node attributes with enriched edge features, including gravity-based interaction terms, distance transformations, and port infrastructure indicators. The GNNs achieve an average classification accuracy of 0.724 and an average Common Part of Flows (CPF) score of 0.22 at the FAF zone level, effectively capturing structural characteristics of US food distribution. Furthermore, the same models are applied to predict county-level flows, providing scalable and trustworthy estimations that outperform traditional models by over 200% improvement in CPF and generalize across spatial scales without requiring county-level training data. By producing spatially explicit commodity-level flow estimates from coarse training data, our method offers a generalizable and data-efficient tool for applications in supply chain vulnerability analysis, environmental footprint modeling, and regional food system planning. Qianheng Zhang, Dev Paul, Michelle Miller 0006, Alfonso Morales, Song Gao 0001 |
SIGSPATIAL/GIS | 5 |
| 2025 | Foundation models for geospatial reasoning: assessing the capabilities of large language models in understanding geometries and topological spatial relationsabstractAI foundation models have demonstrated some capabilities for the understanding of geospatial semantics. However, applying such pre-trained models directly to geospatial datasets remains challenging due to their limited ability to represent and reason with geographical entities, specifically vector-based geometries and natural language descriptions of complex spatial relations. To address these issues, we investigate the extent to which a well-known-text (WKT) representation of geometries and their spatial relations (e.g., topological predicates) are preserved during spatial reasoning when the geospatial vector data are passed to large language models (LLMs) including GPT-3.5-turbo, GPT-4, and DeepSeek-R1-14B. Our workflow employs three distinct approaches to complete the spatial reasoning tasks for comparison, i.e., geometry embedding-based, prompt engineering-based, and everyday language-based evaluation. Our experiment results demonstrate that both the embedding-based and prompt engineering-based approaches to geospatial question-answering tasks with GPT models can achieve an accuracy of over 0.6 on average for the identification of topological spatial relations between two geometries. Among the evaluated models, GPT-4 with few-shot prompting achieved the highest performance with over 0.66 accuracy on topological spatial relation inference. Additionally, GPT-based reasoner is capable of properly comprehending inverse topological spatial relations and including an LLM-generated geometry can enhance the effectiveness for geographic entity retrieval. GPT-4 also exhibits the ability to translate certain vernacular descriptions about places into formal topological relations, and adding the geometry-type or place-type context in prompts may improve inference accuracy, but it varies by instance. The performance of these spatial reasoning tasks unveils the strengths and limitations of the current LLMs in the processing and comprehension of geospatial vector data and offers valuable insights for the refinement of LLMs with geographical knowledge towards the development of geo-foundation models capable of geospatial reasoning. Yuhan Ji, Song Gao 0001, Ivan Majic, Krzysztof Janowicz |
Int. J. Geogr. Inf. Sci. | 2 |
| 2025 | Understanding of the predictability and uncertainty in population distributions empowered by visual analyticsabstractUnderstanding the intricacies of fine-grained population distribution, including both predictability and uncertainty, is crucial for urban planning, social equity, and environmental sustainability. The spatial processes associated with the distribution of populations are complex, and enhancing their predictability involves revealing nonlinear interactions among various explanatory variables. Additionally, population distribution is influenced by various factors that are often challenging to quantify, thereby introducing uncertainty into predictive models. Although the development of explainable artificial intelligence (XAI) helps identify underlying factors, the complex geographical processes and the special nature of spatial data present challenges for purely statistical-based explanation methods, leading to incomplete or incorrect explanations. To address these challenges, we introduce GeoVisX, a geospatial visual analytics framework integrated with XAI. GeoVisX integrates XAI with visual analytics to dissect the spatial processes. Through a case study of Munich, GeoVisX demonstrates its utility in analyzing spatial distribution and identifying key factors impacting population distribution at the 100 m grid level. Our findings highlight the GeoVisX’s capability to enhance understanding of geographical phenomena, contributing to more informed urban policy and planning strategies. This study not only validates the effectiveness of GeoVisX but also emphasizes the importance of incorporating visual analytics and explainable methodologies for addressing complex geographical issues. Peng Luo 0001, Song Gao 0001, Xianfeng Zhang, Deng Majok Chol, Liqiu Meng |
Int. J. Geogr. Inf. Sci. | 3 |
| 2024 | Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
Vardaan Pahuja, Weidi Luo, Yu Gu 0016, Cheng-Hao Tu 0001, Hong-You Chen, Tanya Y. Berger-Wolf, Charles V. Stewart, Song Gao 0001, Wei-Lun Chao, Yu Su 0001 |
CIKM | 8 |
| 2024 | ChatGPT for Intelligent Spatial Analysis Workflow ConstructionabstractThe capability of constructing executable scientific workflows that integrate data, models, and domain-specific knowledge using AI tools remains quite limited. This study utilizes large language model (LLM)-powered chatbots such as ChatGPT to streamline the process of automating spatial analysis workflow construction. By exploring prompt engineering techniques in ChatGPT, we aim to generate scientific workflows that address spatial analysis problems using ArcGIS geoprocessing tools. We identify key challenges and propose strategies involving geo-knowledge and structured prompts to enhance GIS workflow generation. Our preliminary results show that the accuracy and effectiveness of ChatGPT-generated GIS workflow are improved by adopting these strategies. Ying Nie 0002, Song Gao 0001 |
SIGSPATIAL/GIS | 2 |
| 2024 | Automating Geospatial Analysis Workflows Using ChatGPT-4abstractThe field of Geospatial Artificial Intelligence (GeoAI) has significantly impacted domain applications such as urban analytics, environmental monitoring, and disaster management. While powerful geoprocessing tools in geographic information systems (GIS) like ArcGIS Pro are available, automating these workflows with Python scripting using AI chatbots remains a challenge, especially for non-expert users. This study investigates whether ChatGPT-4 can automate GIS workflows by generating ArcPy functions based on structured instructions. We tested prompt engineering's ability on helping large language models (LLMs) understand spatial data and GIS workflows. The overall task success rate reaches 80.5%. It is a valid and easy to implement approach for domain scientists who want to use ArcPy to automate their workflows. Qianheng Zhang, Song Gao 0001 |
SIGSPATIAL/GIS | 2 |
| 2024 | Act2Loc: a synthetic trajectory generation method by combining machine learning and mechanistic modelsabstractHuman mobility data play a crucial role in many fields such as infectious diseases, transportation, and public safety. Although the development of Information and Communication Technologies (ICTs) has made it easy to collect individual-level positioning records, raw individual trajectory data are still limited in availability and usability due to privacy issues. Developing models to generate synthetic trajectories that are statistically close to the real data is a promising solution. This study proposed a novel trajectory generation method called Act2Loc (Activity to Location), which combined machine learning and mechanistic models. First, an activity-sequence generation model was constructed based on machine learning models (i.e. K-medoids and Transformer) to generate individual activity sequences aligning with human activity patterns. Then, a spatial-location selection model was proposed based on mechanistic models (e.g. Universal Opportunity model) to explicitly determine the specific locations of the activities in each generated sequence. Experimental results showed that compared to baselines based on purely machine learning or mechanistic models, Act2Loc can better reproduce the spatio-temporal characteristics of the real data, with additional advantage of low data requirements for training, proving its potential for generating synthetic trajectories in practice. This research offers new insights on knowledge-guided GeoAI models for human mobility. Kang Liu 0010, Shifen Cheng, Song Gao 0001, Ling Yin 0001, Feng Lu 0004 |
Int. J. Geogr. Inf. Sci. | 4 |
| 2023 | Geo-knowledge-informed Deep Learning for Auto-identification of Supraglacial Lakes on the Greenland Ice Sheet from Satellite ImageryabstractMelting glaciers are indicators of global climate change. The formation and dynamics of supraglacial lakes on the Greenland's ice sheet, which appear during the summer melt season, are of particular interest due to their impact on ice dynamics. Detecting these lakes is essential, yet challenging. This paper presents a comprehensive geo-knowledge-informed deep leaning workflow for auto-identification of supraglacial lakes from Sentinel-2 satellite images. The U-Net deep neural network is employed for pixel-level segmentation, distinguishing lakes from other features. Post-processing techniques filter false positives and incorporate geographical knowledge to re-fine results. The experiment results demonstrate the superior performance of our proposed method with F1-score of 0.78. This research contributes to the understanding of supraglacial lake dynamics and the development of robust automated detection methods. Song Gao 0001 |
SIGSPATIAL/GIS | 2 |
| 2023 | Building Privacy-Preserving and Secure Geospatial Artificial Intelligence Foundation Models (Vision Paper)abstractIn recent years we have seen substantial advances in foundation models for artificial intelligence, including language, vision, and multimodal models. Recent studies have highlighted the potential of using foundation models in geospatial artificial intelligence, known as GeoAI Foundation Models, for geographic question answering, remote sensing image understanding, map generation, and location-based services, among others. However, the development and application of GeoAI foundation models can pose serious privacy and security risks, which have not been fully discussed or addressed to date. This paper introduces the potential privacy and security risks throughout the lifecycle of GeoAI foundation models and proposes a comprehensive blueprint for research directions and preventative and control strategies. Through this vision paper, we hope to draw the attention of researchers and policymakers in geospatial domains to these privacy and security risks inherent in GeoAI foundation models and advocate for the development of privacy-preserving and secure GeoAI foundation models. Jinmeng Rao, Song Gao 0001, Gengchen Mai, Krzysztof Janowicz |
SIGSPATIAL/GIS | 2 |
| 2023 | Geo-Foundation Models: Reality, Gaps and OpportunitiesabstractWith the recent rapid advances of revolutionary AI models such as ChatGPT, foundation models have become a main topic for the discussion of future AI. Despite the excitement, the success is still limited to specific types of tasks. Particularly, ChatGPT and similar foundation models have unique characteristics that are difficult to replicate for most geospatial tasks. This paper envisions several major challenges and opportunities in the creation of geospatial foundation (geo-foundation) models, as well as potential future adoption scenarios. We also expect that a major success story is necessary for geo-foundation models to take off in the long term. Yiqun Xie, Zhaonan Wang 0001, Gengchen Mai, Xiaowei Jia, Song Gao 0001, Shaowen Wang 0001 |
SIGSPATIAL/GIS | 6 |
| 2023 | Special issue on geospatial artificial intelligence
Song Gao 0001, Yingjie Hu 0001, Wenwen Li 0002, Lei Zou 0002 |
GeoInformatica | 1 |
| 2023 | CATS: Conditional Adversarial Trajectory Synthesis for privacy-preserving trajectory data publication using deep learning approachesabstractThe prevalence of ubiquitous location-aware devices and mobile Internet enables us to collect massive individual-level trajectory dataset from users. Such trajectory big data bring new opportunities to human mobility research but also raise public concerns with regard to location privacy. In this work, we present the Conditional Adversarial Trajectory Synthesis (CATS), a deep-learning-based GeoAI methodological framework for privacy-preserving trajectory data generation and publication. CATS applies K-anonymity to the underlying spatiotemporal distributions of human movements, which provides a distributional-level strong privacy guarantee. By leveraging conditional adversarial training on K-anonymized human mobility matrices, trajectory global context learning using the attention-based mechanism, and recurrent bipartite graph matching of adjacent trajectory points, CATS is able to reconstruct trajectory topology from conditionally sampled locations and generate high-quality individual-level synthetic trajectory data, which can serve as supplements or alternatives to raw data for privacy-preserving trajectory data publication. The experiment results on over 90k GPS trajectories show that our method has a better performance in privacy preservation, spatiotemporal characteristic preservation, and downstream utility compared with baseline methods, which brings new insights into privacy-preserving human mobility research using generative AI techniques and explores data ethics issues in GIScience. Jinmeng Rao, Song Gao 0001, Sijia Zhu |
Int. J. Geogr. Inf. Sci. | 2 |
| 2023 | Incorporating multimodal context information into traffic speed forecasting through graph deep learningabstractAccurate traffic speed forecasting is a prerequisite for anticipating future traffic status and increasing the resilience of intelligent transportation systems. However, most studies ignore the involvement of context information ubiquitously distributed over the urban environment to boost speed prediction. The diversity and complexity of context information also hinder incorporating it into traffic forecasting. Therefore, this study proposes a multimodal context-based graph convolutional neural network (MCGCN) model to fuse context data into traffic speed prediction, including spatial and temporal contexts. The proposed model comprises three modules, ie (a) hierarchical spatial embedding to learn spatial representations by organizing spatial contexts from different dimensions, (b) multivariate temporal modeling to learn temporal representations by capturing dependencies of multivariate temporal contexts and (c) attention-based multimodal fusion to integrate traffic speed with the spatial and temporal context representations for multi-step speed prediction. We conduct extensive experiments in Singapore. Compared to the baseline model (spatial-temporal graph convolutional network, STGCN), our results demonstrate the importance of multimodal contexts with the mean-absolute-error improvement of 0.29 km/h, 0.45 km/h and 0.89 km/h in 30-min, 60-min and 120-min speed prediction, respectively. We also explore how different contexts affect traffic speed forecasting, providing references for stakeholders to understand the relationship between context information and transportation systems. Yatao Zhang, Tianhong Zhao, Song Gao 0001, Martin Raubal |
Int. J. Geogr. Inf. Sci. | 3 |
| 2022 | Exploring multilevel regularity in human mobility patterns using a feature engineering approach: a case study in chicagoabstractFine-grained trajectory data offers the opportunity to advance our understanding of regularity in individual mobility choices. Existing research on regular location-based mobility patterns cannot fully capture the complexity of day-to-day individual trajectories that are critical for downstream predictive models. To bridge the gap, we construct a comprehensive mobility profile using interpretable features to quantify the regularity of human mobility from three levels, e.g. locations, daily itineraries, and routes. An empirical study uses over 93k trips from 776 car drivers in the Chicago metropolitan area validates the routinely regular patterns in users' mobility choices. A feature engineering approach is then designed for user segmentation. Six user clusters are discovered from the users' mobility profiles. The clusters exhibit heterogeneous commuting behavior and preference for motif choices. The improved multilevel understanding of repeated travel behavior can further assist transportation modeling and planning. Yuhan Ji, Song Gao 0001, Jacob Kruse, Tam Huynh, James Triveri, Chris Scheele, Collin Bennett, Yicheng Wen |
SIGSPATIAL/GIS | 2 |
| 2022 | Region2Vec: community detection on spatial networks using graph embedding with node attributes and spatial interactionsabstractCommunity Detection algorithms are used to detect densely connected components in complex networks and reveal underlying relationships among components. As a special type of networks, spatial networks are usually generated by the connections among geographic regions. Identifying the spatial network communities can help reveal the spatial interaction patterns, understand the hidden regional structures and support regional development decision-making. Given the recent development of Graph Convolutional Networks (GCN) and its powerful performance in identifying multi-scale spatial interactions, we proposed an unsupervised GCN-based community detection method region2vec on spatial networks. Our method first generates node embeddings for regions that share common attributes and have intense spatial interactions, and then applies clustering algorithms to detect communities based on their embedding similarity and spatial adjacency. Experimental results show that while existing methods trade off either attribute similarities or spatial interactions for one another, region2vec maintains a great balance between both and performs the best when one wants to maximize both attribute similarities and spatial interactions within communities. Yunlei Liang, Wen Ye 0001, Song Gao 0001 |
SIGSPATIAL/GIS | 4 |
| 2022 | Spatiotemporal heterogeneities of the associations between human mobility and close contacts with COVID-19 infections in the United StatesabstractIt has been well-established that human mobility has an inseparable relationship with COVID-19 infections. As the COVID-19 pandemic progresses, our knowledge on how human behaviors including mobility and close contact associates with the pandemic also need to stay updated. In this paper, we examine the relationship of the effective reproduction number (Rt) of COVID-19 daily cases with the two indices that provide mobility insights: Mobility Index (CMI) and Contact Index (CCI). Both relationships are evaluated through Maximal Information Coefficient (MIC). Using the Bayesian Change Point Detection and the KShape clustering algorithms, we found significant temporal and spatial heterogeneities among the relationship between two indices and the daily confirmed COVID-19 cases. Although CMI has demonstrated high correlation with COVID-19 cases in 2020, CCI became much more correlated with COVID-19 cases than CMI in 2021. During the first wave in 2020, it is also shown that mobility has a high impact on states outside of Farwest and Southeast than those states within that region. Wen Ye 0001, Song Gao 0001 |
SIGSPATIAL/GIS | 2 |
| 2022 | STICC: a multivariate spatial clustering method for repeated geographic pattern discovery with consideration of spatial contiguityabstractSpatial clustering has been widely used for spatial data mining and knowledge discovery. An ideal multivariate spatial clustering should consider both spatial contiguity and aspatial attributes. Existing spatial clustering approaches may face challenges for discovering repeated geographic patterns with spatial contiguity maintained. In this paper, we propose a Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method that considers both attributes and spatial relationships of geographic objects for multivariate spatial clustering. A subregion is created for each geographic object serving as the basic unit when performing clustering. A Markov random field is then constructed to characterize the attribute dependencies of subregions. Using a spatial consistency strategy, nearby objects are encouraged to belong to the same cluster. To test the performance of the proposed STICC algorithm, we apply it in two use cases. The comparison results with several baseline methods show that the STICC outperforms others significantly in terms of adjusted rand index and macro-F1 score. Join count statistics is also calculated and shows that the spatial contiguity is well preserved by STICC. Such a spatial clustering method may benefit various applications in the fields of geography, remote sensing, transportation, and urban planning, etc. Yuhao Kang, Kunlin Wu, Song Gao 0001, Ignavier Ng, Jinmeng Rao, Shan Ye, Fan Zhang 0011, Teng Fei 0001 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2022 | A review of location encoding for GeoAI: methods and applicationsabstractA common need for artificial intelligence models in the broader geoscience is to encode various types of spatial data, such as points, polylines, polygons, graphs, or rasters, in a hidden embedding space so that they can be readily incorporated into deep learning models. One fundamental step is to encode a single point location into an embedding space, such that this embedding is learning-friendly for downstream machine learning models. We call this process location encoding. However, there lacks a systematic review on location encoding, its potential applications, and key challenges that need to be addressed. This paper aims to fill this gap. We first provide a formal definition of location encoding, and discuss the necessity of it for GeoAI research. Next, we provide a comprehensive survey about the current landscape of location encoding research. We classify location encoding models into different categories based on their inputs and encoding methods, and compare them based on whether they are parametric, multi-scale, distance preserving, and direction aware. We demonstrate that existing location encoders can be unified under one formulation framework. We also discuss the application of location encoding. Finally, we point out several challenges that need to be solved in the future. Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao |
Int. J. Geogr. Inf. Sci. | 4 |
| 2021 | Prediction of human activity intensity using the interactions in physical and social spaces through graph convolutional networksabstractDynamic human activity intensity information is of great importance in many location-based applications. However, two limitations remain in the prediction of human activity intensity. First, it is hard to learn the spatial interaction patterns across scales for predicting human activities. Second, social interaction can help model the activity intensity variation but is rarely considered in the existing literature. To mitigate these limitations, we proposed a novel dynamic activity intensity prediction method with deep learning on graphs using the interactions in both physical and social spaces. In this method, the physical interactions and social interactions between spatial units were integrated into a fused graph convolutional network to model multi-type spatial interaction patterns. The future activity intensity variation was predicted by combining the spatial interaction pattern and the temporal pattern of activity intensity series. The method was verified with a country-scale anonymized mobile phone dataset. The results demonstrated that our proposed deep learning method with combining graph convolutional networks and recurrent neural networks outperformed other baseline approaches. This method enables dynamic human activity intensity prediction from a more spatially and socially integrated perspective, which helps improve the performance of modeling human dynamics. Mingxiao Li 0001, Song Gao 0001, Feng Lu 0004, Kang Liu 0010, Hengcai Zhang, Wei Tu 0001 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2020 | Progress in computational movement analysis - towards movement data scienceabstractThere has not been a time in the history of GIScience when movement analytics and mobility insights have played such an important role in policymaking as in today’s global responses to the COVID-19... Somayeh Dodge, Song Gao 0001, Martin Tomko 0001, Robert Weibel |
Int. J. Geogr. Inf. Sci. | 2 |
| 2020 | GeoAI: spatially explicit artificial intelligence techniques for geographic knowledge discovery and beyondabstractRecent progress in Artificial Intelligence (AI) techniques, the large-scale availability of high-quality data, as well as advances in both hardware and software to efficiently process these data, a... Krzysztof Janowicz, Song Gao 0001, Grant McKenzie, Yingjie Hu 0001, Budhendra L. Bhaduri |
Int. J. Geogr. Inf. Sci. | 2 |
| 2019 | A Data-Driven Approach to Understanding and Predicting the Spatiotemporal Availability of Street ParkingabstractSearching for a parking spot in metropolitan areas is a great challenge comparable to the Hunger Games, especially in highly populated areas such as downtown districts and job centers. On-street parking is often a cost-effective choice compared to parking facilities such as garages and parking lots. However, limited space and complex parking regulation rules make the search process of on-street parking very difficult. To this end, we propose a data-driven framework for understanding and predicting the spatiotemporal availability of on-street parking using the NYC parking tickets open data, points of interest (POI) data and human mobility data. Four popular types of spatial analysis units (i.e., point, street, census tract, and grid) are used to examine the effects of spatial scale in machine learning predictive models. The results show that random forest works the best with the highest accuracy scores for the spatiotemporal availability classification across all four spatial analysis scales. Mingxiao Li 0001, Song Gao 0001, Yunlei Liang, Joseph Marks, Yuhao Kang, Moyin Li |
SIGSPATIAL/GIS | 2 |
| 2019 | Exploring the uncertainty of activity zone detection using digital footprints with multi-scaled DBSCANabstractThe density-based spatial clustering of applications with noise (DBSCAN) method is often used to identify individual activity clusters (i.e., zones) using digital footprints captured from social networks. However, DBSCAN is sensitive to the two parameters, eps and minpts. This paper introduces an improved density-based clustering algorithm, Multi-Scaled DBSCAN (M-DBSCAN), to mitigate the detection uncertainty of clusters produced by DBSCAN at different scales of density and cluster size. M-DBSCAN iteratively calibrates suitable local eps and minpts values instead of using one global parameter setting as DBSCAN for detecting clusters of varying densities, and proves to be effective for detecting potential activity zones. Besides, M-DBSCAN can significantly reduce the noise ratio by identifying all points capturing the activities performed in each zone. Using the historic geo-tagged tweets of users in Washington, D.C. and in Madison, Wisconsin, the results reveal that: 1) M-DBSCAN can capture dispersed clusters with low density of points, and therefore detecting more activity zones for each user; 2) A value of 40 m or higher should be used for eps to reduce the possibility of collapsing distinctive activity zones; and 3) A value between 200 and 300 m is recommended for eps while using DBSCAN for detecting activity zones. Qunying Huang, Song Gao 0001 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2018 | A context-based geoprocessing framework for optimizing meetup location of multiple moving objects along road networksabstractGiven different types of constraints on human life, people must make decisions that satisfy social activity needs. Minimizing costs (i.e. distance, time, or money) associated with travel plays an important role in perceived and realized social quality of life. Identifying optimal interaction locations on road networks when there are multiple moving objects (MMO) with space–time constraints remains a challenge. In this research, we formalize the problem of finding dynamic ideal interaction locations for MMO as a spatial optimization model and introduce a context-based geoprocessing heuristic framework to address this problem. As a proof of concept, a case study involving identification of a meetup location for multiple people under traffic conditions is used to validate the proposed geoprocessing framework. Five heuristic methods with regard to efficient shortest-path search space have been tested. We find that the R* tree-based algorithm performs the best with high quality solutions and low computation time. This framework is implemented in a geographic information systems environment to facilitate integration with external geographic contextual information, e.g. temporary road barriers, points of interest, and real-time traffic information, when dynamically searching for ideal meetup sites. The proposed method can be applied in trip planning, carpooling services, collaborative interaction, and logistics management. Shaohua Wang 0001, Song Gao 0001, Xin Feng 0003, Alan T. Murray |
Int. J. Geogr. Inf. Sci. | 2 |
| 2017 | From ITDL to Place2Vec: Reasoning About Place Type Similarity and Relatedness by Learning Embeddings From Augmented Spatial ContextsabstractUnderstanding, representing, and reasoning about Points Of Interest (POI) types such as Auto Repair, Body Shop, Gas Stations, or Planetarium, is a key aspect of geographic information retrieval, recommender systems, geographic knowledge graphs, as well as studying urban spaces in general, e.g., for extracting functional or vague cognitive regions from user-generated content. One prerequisite to these tasks is the ability to capture the similarity and relatedness between POI types. Intuitively, a spatial search that returns body shops or even gas stations in the absence of auto repair places is still likely to satisfy some user needs while returning planetariums will not. Place hierarchies are frequently used for query expansion, but most of the existing hierarchies are relatively shallow and structured from a single perspective, thereby putting POI types that may be closely related regarding some characteristics far apart from another. This leads to the question of how to learn POI type representations from data. Models such as Word2Vec that produces word embeddings from linguistic contexts are a novel and promising approach as they come with an intuitive notion of similarity. However, the structure of geographic space, e.g., the interactions between POI types, differs substantially from linguistics. In this work, we present a novel method to augment the spatial contexts of POI types using a distance-binned, information-theoretic approach to generate embeddings. We demonstrate that our work outperforms Word2Vec and other models using three different evaluation tasks and strongly correlates with human assessments of POI type similarity. We published the resulting embeddings for 570 place types as well as a collection of human similarity assessments online for others to use. Bo Yan 0003, Krzysztof Janowicz, Gengchen Mai, Song Gao 0001 |
SIGSPATIAL/GIS | 4 |
| 2017 | A data-synthesis-driven method for detecting and extracting vague cognitive regionsabstractCognitive regions and places are notoriously difficult to represent in geographic information science and systems. The exact delineation of cognitive regions is challenging insofar as borders are vague, membership within the regions varies non-monotonically, and raters cannot be assumed to assess membership consistently and homogeneously. In a study published in this journal in 2014, researchers devised a novel grid-based task in which participants rated the membership of individual cells in a given region and contrasted this approach to a standard boundary-drawing task. Specifically, the authors assessed the vague cognitive regions of Northern California and Southern California. The boundary between these cognitive regions was found to have variable width, and region membership peaked not at the most northern or southern cells but at substantially less extreme latitudes. The authors thus concluded that region membership is about attitude, not just latitude. In the present work, we reproduce this study by approaching it from a computational fourth-paradigm perspective, i.e., by the synthesis of high volumes of heterogeneous data from various sources. We compare the regions which we identify to those from the human-participants study of 2014, identifying differences and commonalities. Our results show a significant positive correlation to those in the original study. Beyond the extracted regions themselves, we compare and contrast the empirical and analytical approaches of these two methods, one a conventional human-participants study and the other an application of increasingly popular data-synthesis-driven research methods in GIScience. Song Gao 0001, Krzysztof Janowicz, Daniel R. Montello, Yingjie Hu 0001, Jiue-An Yang, Grant McKenzie, Yiting Ju, Benjamin Adams, Bo Yan 0003 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2016 | VOLT: A Provenance-Producing, Transparent SPARQL Proxy for the On-Demand Computation of Linked Data and its Application to Spatiotemporally Dependent Data
Blake Regalia, Krzysztof Janowicz, Song Gao 0001 |
ESWC | 3 |
| 2016 | ADCN: an anisotropic density-based clustering algorithmabstractIn this work we introduce an anisotropic density-based clustering algorithm. It outperforms DBSCAN and OPTICS for the detection of anisotropic spatial point patterns and performs equally well in cases that do not explicitly benefit from an anisotropic perspective. ADCN has the same time complexity as DBSCAN and OPTICS, namely O(n log n) when using a spatial index, O(n2) otherwise. Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001 |
SIGSPATIAL/GIS | 4 |
| 2013 | A spatiotemporal scientometrics framework for exploring the citation impact of publications and scientistsabstractThe research field of scientometrics is concerned with measuring and analyzing science. In practice, this is often done by restricting the impact of publications, journals, and researchers to a mere frequency. However, scientific activities (co-publication, citation, labor mobility) display clear spatiotemporal patterns, and such patterns have rarely been considered in traditional scientometrics. In this work we focus on the study of citations and present a spatiotemporal scientometrics framework to measure the citation impact of research output by taking physical space, place, and time into account. Specifically, we use the statistics of categorical places (institutions, cities, and countries), spatiotemporal kernel density estimations, cartograms, distance distribution curves, and point-pattern analysis to identify spatiotemporal citation patterns. Moreover, we propose a series of s-indices, such as S_institution-index, S_city-index, and S_country-index to evaluate a scientist's impact as a complement to non-spatial citation indicators, e.g., h-index and g-index. In addition, we have developed an interactive web application which allows users to visually explore research topics, authors, publications, as well as the spread of citations through space and time. Our work offers insights on the role of location in scientific knowledge diffusion. Song Gao 0001, Yingjie Hu 0001, Krzysztof Janowicz, Grant McKenzie |
SIGSPATIAL/GIS | 1 |