Ling Cai 0002

dblp:51/1001-2 · DBLP profile ↗
← Back
10ranked-venue papers in the field
3as first author
8since 2021 · last 2025
0000-0001-7106-4907ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 5 (2 first)Database Systems & Data Management · 3Other / Interdisciplinary · 2 (1 first)
YearPublicationVenuePosition
2025 GeoFM: how will geo-foundation models reshape spatial data science and GeoAI?
abstract
The emerging field of geo-foundation models (GeoFM) has the potential to reshape GeoAI and spatial data science research, education, and practice. In this work, we motivate and define the term and put it into its historic context within GeoAI and spatial data science more broadly. Next, we review core datasets, models, and benchmarks. Based on this overview of the state-of-the-art, we introduce key research challenges for future GeoFM research, such as GeoAI scaling laws, geo-alignment of AI, truly multimodal GeoFM, and so on. Finally, we discuss potential risks of GeoFM research and outline the road ahead with a specific focus on the increasing role of international large-scale collaborations and the future of GeoAI and spatial data science education.
Krzysztof Janowicz, Gengchen Mai, Weiming Huang 0001, Rui Zhu 0008, Ni Lao, Ling Cai 0002
Int. J. Geogr. Inf. Sci.6
2025 The KnowWhereGraph ontology
abstract
KnowWhereGraph is one of the largest fully publicly available geospatial knowledge graphs. It includes data from 30 layers on natural hazards (e.g., hurricanes, wildfires), climate variables (e.g., air temperature, precipitation), soil properties, crop and land-cover types, demographics, and human health, various place and region identifiers, among other themes. These have been leveraged through the graph by a variety of applications to address challenges in food security and agricultural supply chains; sustainability related to soil conservation practices and farm labor; and delivery of emergency humanitarian aid following a disaster. In this paper, we introduce the ontology that acts as the schema for KnowWhereGraph. This broad overview provides insight into the requirements and design specifications for the graph and its schema, including the development methodology (modular ontology modeling) and the resources utilized to implement, materialize, and deploy KnowWhereGraph with its end-user interfaces and public query SPARQL endpoint.
Cogan Shimizu, Shirly Stephen, Adrita Barua, Ling Cai 0002, Antrea Christou, Kitty Currier, Abhilekha Dalal, Colby K. Fisher, Pascal Hitzler, Krzysztof Janowicz, Wenwen Li 0002, Zilong Liu 0003, Mohammad Saeid Mahdavinejad, Gengchen Mai, Dean Rehberger, Mark Schildhauer, Meilin Shi, Sanaz Saki Norouzi, Yuanyuan Tian 0002, Joseph Zalewski, Lu Zhou 0005, Rui Zhu 0008
J. Web Semant.4
2023 HyperQuaternionE: A hyperbolic embedding model for qualitative spatial and temporal reasoning
abstract
Qualitative spatial/temporal reasoning (QSR/QTR) plays a key role in research on human cognition, e.g., as it relates to navigation, as well as in work on robotics and artificial intelligence. Although previous work has mainly focused on various spatial and temporal calculi, more recently representation learning techniques such as embedding have been applied to reasoning and inference tasks such as query answering and knowledge base completion. These subsymbolic and learnable representations are well suited for handling noise and efficiency problems that plagued prior work. However, applying embedding techniques to spatial and temporal reasoning has received little attention to date. In this paper, we explore two research questions: (1) How do embedding-based methods perform empirically compared to traditional reasoning methods on QSR/QTR problems? (2) If the embedding-based methods are better, what causes this superiority? In order to answer these questions, we first propose a hyperbolic embedding model, called HyperQuaternionE, to capture varying properties of relations (such as symmetry and anti-symmetry), to learn inversion relations and relation compositions (i.e., composition tables), and to model hierarchical structures over entities induced by transitive relations. We conduct various experiments on two synthetic datasets to demonstrate the advantages of our proposed embedding-based method against existing embedding models as well as traditional reasoners with respect to entity inference and relation inference. Additionally, our qualitative analysis reveals that our method is able to learn conceptual neighborhoods implicitly. We conclude that the success of our method is attributed to its ability to model composition tables and learn conceptual neighbors, which are among the core building blocks of QSR/QTR.
Ling Cai 0002, Krzysztof Janowicz, Rui Zhu 0008, Gengchen Mai, Bo Yan 0003
GeoInformatica1
2023 Towards general-purpose representation learning of polygonal geometries
Gengchen Mai, Chiyu Max Jiang, Rui Zhu 0008, Yao Xuan, Ling Cai 0002, Krzysztof Janowicz, Stefano Ermon, Ni Lao
GeoInformatica6
2022 A review of location encoding for GeoAI: methods and applications
abstract
A common need for artificial intelligence models in the broader geoscience is to encode various types of spatial data, such as points, polylines, polygons, graphs, or rasters, in a hidden embedding space so that they can be readily incorporated into deep learning models. One fundamental step is to encode a single point location into an embedding space, such that this embedding is learning-friendly for downstream machine learning models. We call this process location encoding. However, there lacks a systematic review on location encoding, its potential applications, and key challenges that need to be addressed. This paper aims to fill this gap. We first provide a formal definition of location encoding, and discuss the necessity of it for GeoAI research. Next, we provide a comprehensive survey about the current landscape of location encoding research. We classify location encoding models into different categories based on their inputs and encoding methods, and compare them based on whether they are parametric, multi-scale, distance preserving, and direction aware. We demonstrate that existing location encoders can be unified under one formulation framework. We also discuss the application of location encoding. Finally, we point out several challenges that need to be solved in the future.
Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao
Int. J. Geogr. Inf. Sci.7
2022 Reasoning over higher-order qualitative spatial relations via spatially explicit neural networks
abstract
Qualitative spatial reasoning has been a core research topic in GIScience and AI for decades. It has been adopted in a wide range of applications such as wayfinding, question answering, and robotics. Most developed spatial inference engines use symbolic representation and reasoning, which focuses on small and densely connected data sets, and struggles to deal with noise and vagueness. However, with more sensors becoming available, reasoning over spatial relations on large-scale and noisy geospatial data sets requires more robust alternatives. This paper, therefore, proposes a subsymbolic approach using neural networks to facilitate qualitative spatial reasoning. More specifically, we focus on higher-order spatial relations as those have been largely ignored due to the binary nature of the underlying representations, e.g. knowledge graphs. We specifically explore the use of neural networks to reason over ternary projective relations such as between. We consider multiple types of spatial constraint, including higher-order relatedness and the conceptual neighborhood of ternary projective relations to make the proposed model spatially explicit. We introduce evaluating results demonstrating that the proposed spatially explicit method substantially outperforms the existing baseline by about 20%.
Rui Zhu 0008, Krzysztof Janowicz, Ling Cai 0002, Gengchen Mai
Int. J. Geogr. Inf. Sci.3
2021 Providing Humanitarian Relief Support through Knowledge Graphs
abstract
Disasters are often unpredictable and complex events, requiring humanitarian organizations to understand and respond to many different issues simultaneously and immediately. Often the biggest challenge to improving the effectiveness of the response is quickly finding the right expert, with the right expertise concerning a specific disaster type/disaster and geographic region. To assist in achieving such a goal, this paper demonstrates a knowledge graph-based search engine developed on top of an expert knowledge graph. It accommodates three modes of information retrieval, including a follow-your-nose search, an expert similarity search, and a SPARQL query interface. We will demonstrate utilizing the system to rapidly navigate from a hazard event to a specific expert who may be helpful, for example. More importantly, as the data is fully integrated including links between hazards and their abstract topics, we can find experts who have relevant expertise while navigating the graph.
Rui Zhu 0008, Ling Cai 0002, Gengchen Mai, Cogan Shimizu, Colby K. Fisher, Krzysztof Janowicz, Anna Lopez-Carr, Andrew Schroeder, Mark Schildhauer, Yuanyuan Tian 0002, Shirly Stephen, Zilong Liu 0003
K-CAP2
2021 Time in a Box: Advancing Knowledge Graph Completion with Temporal Scopes
abstract
Almost all statements in knowledge bases have a temporal scope during which they are valid. Hence, knowledge base completion (KBC) on temporal knowledge bases (TKB), where each statementmay be associated with a temporal scope, has attracted growing attention. Prior works assume that each statement in a TKBmust be associated with a temporal scope. This ignores the fact that the scoping information is commonly missing in a KB. Thus prior work is typically incapable of handling generic use cases where a TKB is composed of temporal statements with/without a known temporal scope. In order to address this issue, we establish a new knowledge base embedding framework, called TIME2BOX, that can deal with atemporal and temporal statements of different types simultaneously. Our main insight is that answers to a temporal query always belong to a subset of answers to a time-agnostic counterpart. Put differently, time is a filter that helps pick out answers to be correct during certain periods. We introduce boxes to represent a set of answer entities to a time-agnostic query. The filtering functionality of time is modeled by intersections over these boxes. In addition, we generalize current evaluation protocols on time interval prediction. We describe experiments on two datasets and show that the proposed method outperforms state-of-the-art (SOTA) methods on both link prediction and time prediction.
Ling Cai 0002, Krzysztof Janowicz, Bo Yan 0003, Rui Zhu 0008, Gengchen Mai
K-CAP1
2019 TransGCN: Coupling Transformation Assumptions with Graph Convolutional Networks for Link Prediction
abstract
Link prediction is an important and frequently studied task that contributes to an understanding of the structure of knowledge graphs (KGs) in statistical relational learning. Inspired by the success of graph convolutional networks (GCN) in modeling graph data, we propose a unified GCN framework, named TransGCN, to address this task, in which relation and entity embeddings are learned simultaneously. To handle heterogeneous relations in KGs, we introduce a novel way of representing heterogeneous neighborhood by introducing transformation assumptions on the relationship between the subject, the relation, and the object of a triple. Specifically, a relation is treated as a transformation operator transforming a head entity to a tail entity. Both translation assumption in TransE and rotation assumption in RotatE are explored in our framework. Additionally, instead of only learning entity embeddings in the convolution-based encoder while learning relation embeddings in the decoder as done by the state-of-art models, e.g., R-GCN, the TransGCN framework trains relation embeddings and entity embeddings simultaneously during the graph convolution operation, thus having fewer parameters compared with R-GCN. Experiments show that our models outperform the-state-of-arts methods on both FB15K-237 and WN18RR.
Ling Cai 0002, Bo Yan 0003, Gengchen Mai, Krzysztof Janowicz, Rui Zhu 0008
K-CAP1
2019 Contextual Graph Attention for Answering Logical Queries over Incomplete Knowledge Graphs
abstract
Recently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attention mechanism to handle the unequal contribution of different query paths. However, commonly used graph attention assumes that the center node embedding is provided, which is unavailable in this task since the center node is to be predicted. To solve this problem we propose a multi-head attention-based end-to-end logical query answering model, called Contextual Graph Attention model (CGA), which uses an initial neighborhood aggregation layer to generate the center embedding, and the whole model is trained jointly on the original KG structure as well as the sampled query-answer pairs. We also introduce two new datasets, DB18 and WikiGeo19, which are rather large in size compared to the existing datasets and contain many more relation types, and use them to evaluate the performance of the proposed model. Our result shows that the proposed CGA with fewer learnable parameters consistently outperforms the baseline models on both datasets as well as Bio dataset.
Gengchen Mai, Krzysztof Janowicz, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao
K-CAP5