Gengchen Mai

dblp:151/5583 · DBLP profile ↗
← Back
31ranked-venue papers in the field
7as first author
25since 2021 · last 2026
0000-0002-7818-7309ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 18 (4 first)Knowledge Engineering, Semantic Web & Information Systems · 8 (2 first)Information Retrieval & Web Search · 2Other / Interdisciplinary · 2 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Street semantic tree: a knowledge-driven GeoAI framework for urban e-scooter ridership classification
abstract
Recently, geospatial artificial intelligence (GeoAI) has risen as a set of essential technologies for urban mobility pattern mining and understanding. However, traditional deep learning models are constrained by their high data dependency and limited interpretability. This study introduces the Knowledge-Driven Semantic Tree (KD-ST) model, a novel GeoAI framework that integrates structured semantic descriptions with graph-based learning to enhance geospatial modeling on e-scooter ridership classification. By incorporating a street knowledge structure into the GeoAI model architecture, KD-ST bridges the gap between purely data-driven methods and knowledge-informed urban analytics, improving classification performance and model transparency. We conducted case studies in four major U.S. cities, including Austin, Phoenix, Denver, and Washington, D.C., to evaluate the proposed KD-ST model’s performance. The proposed model outperformed baseline models by 12.1% to 156.5% as for the F1 score. Moreover, to enhance transparency and reliability, key internal parameters were extracted to visualize and analyze the learned hierarchical knowledge structure. Results indicate that domain knowledge provides useful information for the design of deep learning models and improves model performance. Furthermore, the model achieved higher transferability among cities with more similar urban contexts, which provides valuable insights for e-scooter planners on model choice.
Huihai Wang, William Davis, Justin Yu, Gengchen Mai, Junfeng Jiao
Int. J. Geogr. Inf. Sci.5
2026 SpatialCausal : a spatially-aware causal inference deep learning model for out-of-hospital cardiac arrest survival prediction
abstract
Recently, numerous machine learning methods have been effectively applied to uncover spatial relationships between health risk factors and health outcomes. However, traditional machine learning methods often fail to address confounding bias, which arises when a common factor simultaneously influences both the treatment and the outcome – a challenge frequently encountered in observational studies. Deep learning-based causal inference models seek to mitigate confounding bias by learning balanced representations of covariates between treated and control groups, thereby reducing the dependence of treatment assignment on covariates. This enables accurate estimation of causal effects on health outcomes. Moreover, distinct geospatial patterns of risk exposure and health outcomes are common in many chronic diseases. Therefore, developing a spatially-aware causal inference model is essential for guiding geospatial health interventions. Here, we propose SpatialCausal, a spatially-aware deep learning-based causal inference model that explicitly integrates spatial, non-spatial, and unmeasured confounders, enabling accurate estimation of spatially-aware causal effects. We demonstrate the effectiveness of our approach through an application to Out-of-Hospital Cardiac Arrest survival outcome prediction. Our method surpasses state-of-the-art approaches and exhibits robust adaptability to various geospatial disease scenarios, making it a valuable tool for spatially-aware causal effect estimation in health geography.
Jielu Zhang, Lan Mu, Gengchen Mai, Andrew Grundstein, Zhongliang Zhou, Donglan Zhang
Int. J. Geogr. Inf. Sci.3
2026 SpaCE: a spatial counterfactual explainable deep learning model for predicting out-of-hospital cardiac arrest survival outcome
abstract
Understanding the relationship between risk factors, geospatial patterns, and disease outcomes is essential in health geography research. These relationships can inform the implementation of healthcare and public health strategies to improve health outcomes. To accurately uncover such complex relationships, it is necessary to have a predictive model capable of integrating both health variables and spatial information to forecast health outcomes, along with a tool to interpret and reveal the patterns identified by this model. We developed a Spatial Counterfactual Explainable Deep Learning model (SpaCE), comprising a spatially explicit health outcome predictor and a prototype-guided counterfactual explanation. The SpaCE model unifies geospatial and health variables to improve predictions and generates hypothetical examples with minimal changes but opposite outcomes. Using these counterfactuals, SpaCE assesses the impact of each variable in different spatial contexts. We evaluated the model for predicting cardiac arrest survival outcomes. With a 0.682 AUCROC score, the SpaCE exceeds baseline models by 10.2%. Further analysis also reveals that the geospatial context significantly affects how various risk factors affect the survival outcomes of patients. Overall, the SpaCE model significantly improves predictive accuracy and explainability. It provides targeted interventions at both individual and geographic levels, and the cardiac arrest case study shows its high adaptability to various disease scenarios.
Jielu Zhang, Lan Mu, Donglan Zhang, Zhuo Chen 0012, Janani Rajbhandari-Thapa, José A. Pagán, Yan Li 0017, Gengchen Mai, Zhongliang Zhou
Int. J. Geogr. Inf. Sci.8
2025 Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
abstract
AI alignment describes the challenge of ensuring (future) AI systems behave in accordance with societal norms, values, and goals. Alignment is now central to research on foundation models and AI agents. Most recent work focuses on methods to prevent potentially harmful biases, account for social inequalities, improve AI safety, and enhance explainability. Notably, the debiasing 'corrections' applied to various stages of AI/ML workflows may lead to outcomes that diverge strongly from current statistical realities on the ground. For instance, text-to-image models may depict a balanced gender ratio of company leadership, despite existing imbalances. However, an often overlooked dimension is the geographic variability of alignment. What is considered appropriate, truthful, or legal can vary greatly between regions due to cultural differences, political realities, or legislation. Hence, some model outputs align without further knowledge of the user's geospatial context, while others are highly sensitive to it. Put differently, whether these outputs align varies geographically. E.g., statements about Kashmir cannot be generated without understanding the user's origin and current location. From a common-sense perspective, this problem is hardly new. In fact, Google Maps will render different administrative borders based on the user's location. Interestingly, in both knowledge representation and representation learning, spatiotemporal context, e.g., due to the monotonic nature of reasoning, remains a major challenge. Until very recently, these were largely theoretical problems. What is truly novel is the scale and level of automation at which AI systems now mediate knowledge, express opinions, and represent reality to millions of users across borders, often with little transparency or oversight regarding how context is handled. With agentic AI on the horizon, the urgency for pluralistic, geographically aware alignment, rather than one-size-fits-all solutions, is growing. Here, we motivate and formalize the vision of geo-alignment, outline how it goes beyond pluralistic alignment by offering learnable spatially explicit patterns, and suggest concrete avenues for future research.
Krzysztof Janowicz, Zilong Liu 0003, Gengchen Mai, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao 0001
SIGSPATIAL/GIS3
2025 Scenario-Based Evaluation of Probabilistic Time Series Forecasting for Solar Energy
abstract
Probabilistic time-series forecasting plays a vital role in decisionmaking under uncertainty, especially in applications like solar energy, where forecast reliability directly impacts energy planning and grid stability. While recent models have improved in generating predictive distributions rather than single-point estimates, existing evaluations often focus on average performance and overlook how model quality varies across different real-world scenarios. In solar energy monitoring, for example, the difficulty of forecasting can change significantly due to atmospheric variability, sensor types, and climate conditions. This work addresses the need for scenario-aware evaluation of probabilistic models by benchmarking state-of-the-art forecasting methods using SolarCube-a large-scale solar radiation dataset spanning diverse regions, cloud regimes, and environmental conditions. We define structured "easy" and "hard" cases across four scenarios and examine how different probabilistic model families (e.g., diffusion, VAE, flow-based) capture uncertainty under these conditions. Our goal is to move beyond overall metrics and reveal how model reliability changes across scenarios that are critical for downstream applications.
Yiqun Xie, Xiaowei Jia, Gengchen Mai, Sophia Hou, Zhili Li
SIGSPATIAL/GIS4
2025 ZooplanktonBench: A Geo-Aware Zooplankton Recognition and Classification Dataset from Marine Observations
abstract
Plankton are small drifting organisms found throughout the world's oceans and can be indicators of ocean health. One component of this plankton community is the zooplankton, which includes gelatinous animals and crustaceans (e.g. shrimp), as well as the early life stages (i.e., eggs and larvae) of many commercially important fishes. Being able to monitor zooplankton abundances accurately and understand how populations change in relation to ocean conditions is invaluable to marine science research, with important implications for future marine seafood productivity. While new imaging technologies generate massive amounts of video data of zooplankton, analyzing them using general-purpose computer vision tools turns out to be highly challenging due to the high similarity in appearance between the zooplankton and its background (e.g., marine snow). In this work, we present the ZooplanktonBench, a benchmark dataset containing images and videos of zooplankton associated with rich geospatial metadata (e.g., geographic coordinates, depth, etc.) in various water ecosystems. ZooplanktonBench defines a collection of tasks to detect, classify, and track zooplankton in challenging settings, including highly cluttered environments, living vs non-living classification, objects with similar shapes, and relatively small objects. Our dataset presents unique challenges and opportunities for state-of-the-art computer vision systems to evolve and improve visual understanding in dynamic environments characterized by significant variation and the need for geo-awareness. The code and settings described in this paper can be found on our website: https://lfk118.github.io/ZooplanktonBench_Webpage.
Fukun Liu, Adam T. Greer, Gengchen Mai, Jin Sun 0011
KDD (2)3
2025 GeoFM: how will geo-foundation models reshape spatial data science and GeoAI?
abstract
The emerging field of geo-foundation models (GeoFM) has the potential to reshape GeoAI and spatial data science research, education, and practice. In this work, we motivate and define the term and put it into its historic context within GeoAI and spatial data science more broadly. Next, we review core datasets, models, and benchmarks. Based on this overview of the state-of-the-art, we introduce key research challenges for future GeoFM research, such as GeoAI scaling laws, geo-alignment of AI, truly multimodal GeoFM, and so on. Finally, we discuss potential risks of GeoFM research and outline the road ahead with a specific focus on the increasing role of international large-scale collaborations and the future of GeoAI and spatial data science education.
Krzysztof Janowicz, Gengchen Mai, Weiming Huang 0001, Rui Zhu 0008, Ni Lao, Ling Cai 0002
Int. J. Geogr. Inf. Sci.2
2025 Machine-learning-enabled spatial pattern mining: evaluating the impact of imperfect inputs
abstract
Spatial pattern mining (SPM) aims to detect geographic locations or areas that present interesting, nontrivial, and potentially useful patterns. Traditional formulations of point-based SPM tasks are mainly based on true observations, which tend to have limited spatial coverage, availability, and timeliness. While machine learning (ML) has the potential to extend the range of usable data, the uncertainty of model-predicted labels presents new challenges for their usability in the SPM context. This paper formulates the task of ML-enabled SPM using predicted labels by ML models. Given the ever-expanding family of spatial patterns, we consider four widely-adopted patterns – hotspots, co-locations, mixture patterns, and spatial outliers – to scope our study to make the discussion concrete. We develop soft-label versions of SPM algorithms that can directly execute on uncertain predictions generated by ML models. Additionally, we evaluate the ML-enabled SPM results for both categorical and real-valued datasets across a spectrum of prediction quality. The results show that certain spatial patterns such as multinomial scan statistic-based mixture patterns and normal-model-based hotspots can more robustly maintain the detection quality at different error levels, while others such as spatial outliers are more sensitive to incorrect predictions. This provides helpful guidance on using learning-based predictions for SPM.
Zhili Li, Yiqun Xie, Xiaowei Jia, Gengchen Mai, Weiye Chen
Int. J. Geogr. Inf. Sci.4
2025 The KnowWhereGraph ontology
abstract
KnowWhereGraph is one of the largest fully publicly available geospatial knowledge graphs. It includes data from 30 layers on natural hazards (e.g., hurricanes, wildfires), climate variables (e.g., air temperature, precipitation), soil properties, crop and land-cover types, demographics, and human health, various place and region identifiers, among other themes. These have been leveraged through the graph by a variety of applications to address challenges in food security and agricultural supply chains; sustainability related to soil conservation practices and farm labor; and delivery of emergency humanitarian aid following a disaster. In this paper, we introduce the ontology that acts as the schema for KnowWhereGraph. This broad overview provides insight into the requirements and design specifications for the graph and its schema, including the development methodology (modular ontology modeling) and the resources utilized to implement, materialize, and deploy KnowWhereGraph with its end-user interfaces and public query SPARQL endpoint.
Cogan Shimizu, Shirly Stephen, Adrita Barua, Ling Cai 0002, Antrea Christou, Kitty Currier, Abhilekha Dalal, Colby K. Fisher, Pascal Hitzler, Krzysztof Janowicz, Wenwen Li 0002, Zilong Liu 0003, Mohammad Saeid Mahdavinejad, Gengchen Mai, Dean Rehberger, Mark Schildhauer, Meilin Shi, Sanaz Saki Norouzi, Yuanyuan Tian 0002, Joseph Zalewski, Lu Zhou 0005, Rui Zhu 0008
J. Web Semant.14
2024 SRL: Towards a General-Purpose Framework for Spatial Representation Learning
abstract
Representation learning (RL) techniques are widely adopted in areas such as natural language processing and computer vision, with prominent examples such as attention and ConvNet architectures. In comparison, many GeoAI works still rely on feature engineering or data conversion to represent spatial data (e.g., points, polylines, polygons, 3D building models, etc.) as features in formats that are easier for neural networks to handle. The neural network architectures remain unchanged, and the need for feature engineering has become a bottleneck for applying deep learning to new tasks in the age of big data. In this paper, we advocate the idea of developing learnable spatial representation modules, which not only enable spatial reasoning but also enable neural nets to directly consume (i.e., encoding) or generate (i.e., decoding) spatial data. We propose Spatial Representation Learning (SRL), a new general-purpose representation learning framework for spatial reasoning. We discuss the key challenges of spatial representation learning including multi-scale RL, continuous RL, shape-centric RL, noise-robust RL, heterogeneity-aware RL, and fairness-aware RL. We also discuss the critical role and potential of SRL in various geospatial subdomains and how this technique can lead to a new generation of GeoAI.
Gengchen Mai, Xiaobai Angela Yao, Yiqun Xie, Jinmeng Rao, Hao Li 0019, Qing Zhu 0011, Ni Lao
SIGSPATIAL/GIS1
2024 Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
abstract
Geolocating precise locations from images presents a challenging problem in computer vision and information retrieval. Traditional methods typically employ either classification-dividing the Earth's surface into grid cells and classifying images accordingly, or retrieval-identifying locations by matching images with a database of image-location pairs. However, classification-based approaches are limited by the cell size and cannot yield precise predictions, while retrieval-based systems usually suffer from poor search quality and inadequate coverage of the global landscape at varied scale and aggregation levels. To overcome these drawbacks, we present Img2Loc, a novel system that redefines image geolocalization as a text generation task. This is achieved using cutting-edge large multi-modality models (LMMs) like GPT-4V or LLaVA with retrieval augmented generation. Img2Loc first employs CLIP-based representations to generate an image-based coordinate query database. It then uniquely combines query results with images itself, forming elaborate prompts customized for LMMs. When tested on benchmark datasets such as Im2GPS3k and YFCC4k, Img2Loc not only surpasses the performance of previous state-of-the-art models but does so without any model training. A video demonstration of the system can be accessed via this link https://drive.google.com/file/d/16A6A-mc7AyUoKHRH3_WBRToRC13sn7tU/view?usp=sharing
Zhongliang Zhou, Jielu Zhang, Zihan Guan 0001, Mengxuan Hu, Ni Lao, Lan Mu, Sheng Li 0001, Gengchen Mai
SIGIR8
2024 BB-GeoGPT: A framework for learning a large language model for geographic information science
Yifan Zhang 0009, Zhiyun Wang, Zhengting He, Gengchen Mai, Jianfeng Lin 0004, Wenhao Yu 0001
Inf. Process. Manag.5
2023 Rethink Geographical Generalizability with Unsupervised Self-Attention Model Ensemble: A Case Study of OpenStreetMap Missing Building Detection in Africa
abstract
The recent advance of adapting pre-trained task-agnostic artificial intelligence (AI) models leads to great successes in downstream tasks via fine-tuning, or low-resource (i.e., few-shot and zero-shot) learning. However, when adapting such pre-trained AI models to geographical applications, it is still challenging to find the "sweet spot" of the model's generalizability and specializability (e.g., geographic generalizability v.s. spatial heterogeneity). For instance, a building detection task may require vision models with different parameters across different geographic areas of the world. In this paper, we rethink this interesting topic, namely Geographical Generalizability of GeoAI models, with a case study of detecting OpenStreetMap (OSM) missing buildings across different countries in sub-Saharan Africa. We consider a real-world scenario, in which we first train a Single-Shot Multibox Detection (SSD) base model for OSM missing building detection in Kakola, Tanzania, where a previous humanitarian mapping project of OSM was organized to map all possible buildings. Then we extrapolate this base model using Few-Shot Transfer Learning (FSTL) to a set of areas in the proximity of the test area in Cameroon. Here, we develop a Geographical Weighted Model Ensemble (GWME) method to improve Geographical Generalizability of GeoAI models. Moreover, we compare four unsupervised model ensemble weighting strategies: 1) Average weighting, 2) Image similarity weighting, 3) Geographical distance weighting, and 4) Self-attention-based weighting. Experiments show promising results of the proposed GWME method, which implicitly generates model weights from their location embedding and image feature embedding in an unsupervised manner. More specifically, the self-attention-based model ensemble achieves the highest performance. The results shed inspiring light on improving the generalizability and replicability of GeoAI models across geographic areas. Data and code are available at https://github.com/tum-bgd/GWME.
Hao Li 0019, Jiapan Wang, Johann Maximilian Zollner, Gengchen Mai, Ni Lao, Martin Werner 0001
SIGSPATIAL/GIS4
2023 Building Privacy-Preserving and Secure Geospatial Artificial Intelligence Foundation Models (Vision Paper)
abstract
In recent years we have seen substantial advances in foundation models for artificial intelligence, including language, vision, and multimodal models. Recent studies have highlighted the potential of using foundation models in geospatial artificial intelligence, known as GeoAI Foundation Models, for geographic question answering, remote sensing image understanding, map generation, and location-based services, among others. However, the development and application of GeoAI foundation models can pose serious privacy and security risks, which have not been fully discussed or addressed to date. This paper introduces the potential privacy and security risks throughout the lifecycle of GeoAI foundation models and proposes a comprehensive blueprint for research directions and preventative and control strategies. Through this vision paper, we hope to draw the attention of researchers and policymakers in geospatial domains to these privacy and security risks inherent in GeoAI foundation models and advocate for the development of privacy-preserving and secure GeoAI foundation models.
Jinmeng Rao, Song Gao 0001, Gengchen Mai, Krzysztof Janowicz
SIGSPATIAL/GIS3
2023 Geo-Foundation Models: Reality, Gaps and Opportunities
abstract
With the recent rapid advances of revolutionary AI models such as ChatGPT, foundation models have become a main topic for the discussion of future AI. Despite the excitement, the success is still limited to specific types of tasks. Particularly, ChatGPT and similar foundation models have unique characteristics that are difficult to replicate for most geospatial tasks. This paper envisions several major challenges and opportunities in the creation of geospatial foundation (geo-foundation) models, as well as potential future adoption scenarios. We also expect that a major success story is necessary for geo-foundation models to take off in the long term.
Yiqun Xie, Zhaonan Wang 0001, Gengchen Mai, Xiaowei Jia, Song Gao 0001, Shaowen Wang 0001
SIGSPATIAL/GIS3
2023 HyperQuaternionE: A hyperbolic embedding model for qualitative spatial and temporal reasoning
abstract
Qualitative spatial/temporal reasoning (QSR/QTR) plays a key role in research on human cognition, e.g., as it relates to navigation, as well as in work on robotics and artificial intelligence. Although previous work has mainly focused on various spatial and temporal calculi, more recently representation learning techniques such as embedding have been applied to reasoning and inference tasks such as query answering and knowledge base completion. These subsymbolic and learnable representations are well suited for handling noise and efficiency problems that plagued prior work. However, applying embedding techniques to spatial and temporal reasoning has received little attention to date. In this paper, we explore two research questions: (1) How do embedding-based methods perform empirically compared to traditional reasoning methods on QSR/QTR problems? (2) If the embedding-based methods are better, what causes this superiority? In order to answer these questions, we first propose a hyperbolic embedding model, called HyperQuaternionE, to capture varying properties of relations (such as symmetry and anti-symmetry), to learn inversion relations and relation compositions (i.e., composition tables), and to model hierarchical structures over entities induced by transitive relations. We conduct various experiments on two synthetic datasets to demonstrate the advantages of our proposed embedding-based method against existing embedding models as well as traditional reasoners with respect to entity inference and relation inference. Additionally, our qualitative analysis reveals that our method is able to learn conceptual neighborhoods implicitly. We conclude that the success of our method is attributed to its ability to model composition tables and learn conceptual neighbors, which are among the core building blocks of QSR/QTR.
Ling Cai 0002, Krzysztof Janowicz, Rui Zhu 0008, Gengchen Mai, Bo Yan 0003
GeoInformatica4
2023 Towards general-purpose representation learning of polygonal geometries
Gengchen Mai, Chiyu Max Jiang, Rui Zhu 0008, Yao Xuan, Ling Cai 0002, Krzysztof Janowicz, Stefano Ermon, Ni Lao
GeoInformatica1
2023 Geo-knowledge-guided GPT models improve the extraction of location descriptions from disaster-related social media messages
abstract
Social media messages posted by people during natural disasters often contain important location descriptions, such as the locations of victims. Recent research has shown that many of these location descriptions go beyond simple place names, such as city names and street names, and are difficult to extract using typical named entity recognition (NER) tools. While advanced machine learning models could be trained, they require large labeled training datasets that can be time-consuming and labor-intensive to create. In this work, we propose a method that fuses geo-knowledge of location descriptions and a Generative Pre-trained Transformer (GPT) model, such as ChatGPT and GPT-4. The result is a geo-knowledge-guided GPT model that can accurately extract location descriptions from disaster-related social media messages. Also, only 22 training examples encoding geo-knowledge are used in our method. We conduct experiments to compare this method with nine alternative approaches on a dataset of tweets from Hurricane Harvey. Our method demonstrates an over 40% improvement over typically used NER approaches. The experiment results also show that geo-knowledge is indispensable for guiding the behavior of GPT models. The extracted location descriptions can help disaster responders reach victims more quickly and may even save lives.
Yingjie Hu 0001, Gengchen Mai, Chris Cundy, Kristy Choi, Ni Lao, Gaurish Lakhanpal, Ryan Zhenqi Zhou, Kenneth Joseph
Int. J. Geogr. Inf. Sci.2
2022 LD Connect: A Linked Data Portal for IOS Press Scientometrics
Zilong Liu 0003, Meilin Shi, Krzysztof Janowicz, Blake Regalia, Stephanie Delbecque, Gengchen Mai, Rui Zhu 0008, Pascal Hitzler
ESWC6
2022 Understanding economic development in rural Africa using satellite imagery, building footprints and deep models
abstract
Recent advancements in machine learning enable cost effective methods for understanding societal and economic activities in developing countries using publicly available satellite imagery. However, this progress remains stagnant in rural areas where the largest population under poverty line resides. In this work, we explore deep models' performance in rural areas in Africa and investigate methods that improve the performance. We argue that the geographic displacement noise present in ground surveys for anonymization purposes causes misalignments between input imagery and labels and therefore hampers accuracy, which exacerbates in rural areas. We then propose to incorporate building footprints data and a novel self-attention mechanism to provide more robust and accurate predictions of socioeconomic development. We test our framework against three socioeconomic measures in 21 African countries. Our best models outperform previous baselines in most of these tasks.
Amna Elmustafa, Erik Rozi, Gengchen Mai, Stefano Ermon, Marshall Burke, David B. Lobell
SIGSPATIAL/GIS4
2022 Towards a foundation model for geospatial artificial intelligence (vision paper)
abstract
Large pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine tuning, few-shot, or even zero-shot learning. Despite their successes in language and vision tasks, we have yet to see an attempt to develop foundation models for geospatial artificial intelligence (GeoAI). In this work, we explore the promises and challenges for developing multimodal foundation models for GeoAI. We first show the advantages of this idea by testing the performance of existing Large pre-trained Language Models (LLMs) (e.g. GPT-2 and GPT-3) on two geospatial semantics tasks. Results indicate that these task-agnostic LLMs can outperform task-specific fully-supervised models on both tasks with 2--9% improvement in a few-shot learning setting. However, we also show the limitations of these existing foundation models given the multimodality nature of GeoAI, especially when dealing with geometries in conjunction with other modalities. So we discuss the possibility of a multimodal foundation model which can reason over various types of geospatial data through geospatial alignments. We conclude this paper by discussing the unique risks and challenges to develop such model for GeoAI.
Gengchen Mai, Chris Cundy, Kristy Choi, Yingjie Hu 0001, Ni Lao, Stefano Ermon
SIGSPATIAL/GIS1
2022 A review of location encoding for GeoAI: methods and applications
abstract
A common need for artificial intelligence models in the broader geoscience is to encode various types of spatial data, such as points, polylines, polygons, graphs, or rasters, in a hidden embedding space so that they can be readily incorporated into deep learning models. One fundamental step is to encode a single point location into an embedding space, such that this embedding is learning-friendly for downstream machine learning models. We call this process location encoding. However, there lacks a systematic review on location encoding, its potential applications, and key challenges that need to be addressed. This paper aims to fill this gap. We first provide a formal definition of location encoding, and discuss the necessity of it for GeoAI research. Next, we provide a comprehensive survey about the current landscape of location encoding research. We classify location encoding models into different categories based on their inputs and encoding methods, and compare them based on whether they are parametric, multi-scale, distance preserving, and direction aware. We demonstrate that existing location encoders can be unified under one formulation framework. We also discuss the application of location encoding. Finally, we point out several challenges that need to be solved in the future.
Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao
Int. J. Geogr. Inf. Sci.1
2022 Reasoning over higher-order qualitative spatial relations via spatially explicit neural networks
abstract
Qualitative spatial reasoning has been a core research topic in GIScience and AI for decades. It has been adopted in a wide range of applications such as wayfinding, question answering, and robotics. Most developed spatial inference engines use symbolic representation and reasoning, which focuses on small and densely connected data sets, and struggles to deal with noise and vagueness. However, with more sensors becoming available, reasoning over spatial relations on large-scale and noisy geospatial data sets requires more robust alternatives. This paper, therefore, proposes a subsymbolic approach using neural networks to facilitate qualitative spatial reasoning. More specifically, we focus on higher-order spatial relations as those have been largely ignored due to the binary nature of the underlying representations, e.g. knowledge graphs. We specifically explore the use of neural networks to reason over ternary projective relations such as between. We consider multiple types of spatial constraint, including higher-order relatedness and the conceptual neighborhood of ternary projective relations to make the proposed model spatially explicit. We introduce evaluating results demonstrating that the proposed spatially explicit method substantially outperforms the existing baseline by about 20%.
Rui Zhu 0008, Krzysztof Janowicz, Ling Cai 0002, Gengchen Mai
Int. J. Geogr. Inf. Sci.4
2021 Providing Humanitarian Relief Support through Knowledge Graphs
abstract
Disasters are often unpredictable and complex events, requiring humanitarian organizations to understand and respond to many different issues simultaneously and immediately. Often the biggest challenge to improving the effectiveness of the response is quickly finding the right expert, with the right expertise concerning a specific disaster type/disaster and geographic region. To assist in achieving such a goal, this paper demonstrates a knowledge graph-based search engine developed on top of an expert knowledge graph. It accommodates three modes of information retrieval, including a follow-your-nose search, an expert similarity search, and a SPARQL query interface. We will demonstrate utilizing the system to rapidly navigate from a hazard event to a specific expert who may be helpful, for example. More importantly, as the data is fully integrated including links between hazards and their abstract topics, we can find experts who have relevant expertise while navigating the graph.
Rui Zhu 0008, Ling Cai 0002, Gengchen Mai, Cogan Shimizu, Colby K. Fisher, Krzysztof Janowicz, Anna Lopez-Carr, Andrew Schroeder, Mark Schildhauer, Yuanyuan Tian 0002, Shirly Stephen, Zilong Liu 0003
K-CAP3
2021 Time in a Box: Advancing Knowledge Graph Completion with Temporal Scopes
abstract
Almost all statements in knowledge bases have a temporal scope during which they are valid. Hence, knowledge base completion (KBC) on temporal knowledge bases (TKB), where each statementmay be associated with a temporal scope, has attracted growing attention. Prior works assume that each statement in a TKBmust be associated with a temporal scope. This ignores the fact that the scoping information is commonly missing in a KB. Thus prior work is typically incapable of handling generic use cases where a TKB is composed of temporal statements with/without a known temporal scope. In order to address this issue, we establish a new knowledge base embedding framework, called TIME2BOX, that can deal with atemporal and temporal statements of different types simultaneously. Our main insight is that answers to a temporal query always belong to a subset of answers to a time-agnostic counterpart. Put differently, time is a filter that helps pick out answers to be correct during certain periods. We introduce boxes to represent a set of answer entities to a time-agnostic query. The filtering functionality of time is modeled by intersections over these boxes. In addition, we generalize current evaluation protocols on time interval prediction. We describe experiments on two datasets and show that the proposed method outperforms state-of-the-art (SOTA) methods on both link prediction and time prediction.
Ling Cai 0002, Krzysztof Janowicz, Bo Yan 0003, Rui Zhu 0008, Gengchen Mai
K-CAP5
2019 TransGCN: Coupling Transformation Assumptions with Graph Convolutional Networks for Link Prediction
abstract
Link prediction is an important and frequently studied task that contributes to an understanding of the structure of knowledge graphs (KGs) in statistical relational learning. Inspired by the success of graph convolutional networks (GCN) in modeling graph data, we propose a unified GCN framework, named TransGCN, to address this task, in which relation and entity embeddings are learned simultaneously. To handle heterogeneous relations in KGs, we introduce a novel way of representing heterogeneous neighborhood by introducing transformation assumptions on the relationship between the subject, the relation, and the object of a triple. Specifically, a relation is treated as a transformation operator transforming a head entity to a tail entity. Both translation assumption in TransE and rotation assumption in RotatE are explored in our framework. Additionally, instead of only learning entity embeddings in the convolution-based encoder while learning relation embeddings in the decoder as done by the state-of-art models, e.g., R-GCN, the TransGCN framework trains relation embeddings and entity embeddings simultaneously during the graph convolution operation, thus having fewer parameters compared with R-GCN. Experiments show that our models outperform the-state-of-arts methods on both FB15K-237 and WN18RR.
Ling Cai 0002, Bo Yan 0003, Gengchen Mai, Krzysztof Janowicz, Rui Zhu 0008
K-CAP3
2019 Contextual Graph Attention for Answering Logical Queries over Incomplete Knowledge Graphs
abstract
Recently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attention mechanism to handle the unequal contribution of different query paths. However, commonly used graph attention assumes that the center node embedding is provided, which is unavailable in this task since the center node is to be predicted. To solve this problem we propose a multi-head attention-based end-to-end logical query answering model, called Contextual Graph Attention model (CGA), which uses an initial neighborhood aggregation layer to generate the center embedding, and the whole model is trained jointly on the original KG structure as well as the sampled query-answer pairs. We also introduce two new datasets, DB18 and WikiGeo19, which are rather large in size compared to the existing datasets and contain many more relation types, and use them to evaluate the performance of the proposed model. Our result shows that the proposed CGA with fewer learnable parameters consistently outperforms the baseline models on both datasets as well as Bio dataset.
Gengchen Mai, Krzysztof Janowicz, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao
K-CAP1
2018 Support and Centrality: Learning Weights for Knowledge Graph Embedding Models
Gengchen Mai, Krzysztof Janowicz, Bo Yan 0003
EKAW1
2018 GNIS-LD: Serving and Visualizing the Geographic Names Information System Gazetteer as Linked Data
Blake Regalia, Krzysztof Janowicz, Gengchen Mai, Dalia Varanka, E. Lynn Usery
ESWC3
2017 From ITDL to Place2Vec: Reasoning About Place Type Similarity and Relatedness by Learning Embeddings From Augmented Spatial Contexts
abstract
Understanding, representing, and reasoning about Points Of Interest (POI) types such as Auto Repair, Body Shop, Gas Stations, or Planetarium, is a key aspect of geographic information retrieval, recommender systems, geographic knowledge graphs, as well as studying urban spaces in general, e.g., for extracting functional or vague cognitive regions from user-generated content. One prerequisite to these tasks is the ability to capture the similarity and relatedness between POI types. Intuitively, a spatial search that returns body shops or even gas stations in the absence of auto repair places is still likely to satisfy some user needs while returning planetariums will not. Place hierarchies are frequently used for query expansion, but most of the existing hierarchies are relatively shallow and structured from a single perspective, thereby putting POI types that may be closely related regarding some characteristics far apart from another. This leads to the question of how to learn POI type representations from data. Models such as Word2Vec that produces word embeddings from linguistic contexts are a novel and promising approach as they come with an intuitive notion of similarity. However, the structure of geographic space, e.g., the interactions between POI types, differs substantially from linguistics. In this work, we present a novel method to augment the spatial contexts of POI types using a distance-binned, information-theoretic approach to generate embeddings. We demonstrate that our work outperforms Word2Vec and other models using three different evaluation tasks and strongly correlates with human assessments of POI type similarity. We published the resulting embeddings for 570 place types as well as a collection of human similarity assessments online for others to use.
Bo Yan 0003, Krzysztof Janowicz, Gengchen Mai, Song Gao 0001
SIGSPATIAL/GIS3
2016 ADCN: an anisotropic density-based clustering algorithm
abstract
In this work we introduce an anisotropic density-based clustering algorithm. It outperforms DBSCAN and OPTICS for the detection of anisotropic spatial point patterns and performs equally well in cases that do not explicitly benefit from an anisotropic perspective. ADCN has the same time complexity as DBSCAN and OPTICS, namely O(n log n) when using a spatial index, O(n2) otherwise.
Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001
SIGSPATIAL/GIS1