EDBT 2026 Demo / reviewers in the wild / expert
Krzysztof Janowicz
dblp:95/5567 · also Krzysztof W. Janowicz
· DBLP profile ↗
45ranked-venue papers in the field
7as first author
16since 2021 · last 2025
0009-0003-1968-887XORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 23 (5 first)Knowledge Engineering, Semantic Web & Information Systems · 16 (2 first)Other / Interdisciplinary · 3Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Whose Truth? Pluralistic Geo-Alignment for (Agentic) AIabstractAI alignment describes the challenge of ensuring (future) AI systems behave in accordance with societal norms, values, and goals. Alignment is now central to research on foundation models and AI agents. Most recent work focuses on methods to prevent potentially harmful biases, account for social inequalities, improve AI safety, and enhance explainability. Notably, the debiasing 'corrections' applied to various stages of AI/ML workflows may lead to outcomes that diverge strongly from current statistical realities on the ground. For instance, text-to-image models may depict a balanced gender ratio of company leadership, despite existing imbalances. However, an often overlooked dimension is the geographic variability of alignment. What is considered appropriate, truthful, or legal can vary greatly between regions due to cultural differences, political realities, or legislation. Hence, some model outputs align without further knowledge of the user's geospatial context, while others are highly sensitive to it. Put differently, whether these outputs align varies geographically. E.g., statements about Kashmir cannot be generated without understanding the user's origin and current location. From a common-sense perspective, this problem is hardly new. In fact, Google Maps will render different administrative borders based on the user's location. Interestingly, in both knowledge representation and representation learning, spatiotemporal context, e.g., due to the monotonic nature of reasoning, remains a major challenge. Until very recently, these were largely theoretical problems. What is truly novel is the scale and level of automation at which AI systems now mediate knowledge, express opinions, and represent reality to millions of users across borders, often with little transparency or oversight regarding how context is handled. With agentic AI on the horizon, the urgency for pluralistic, geographically aware alignment, rather than one-size-fits-all solutions, is growing. Here, we motivate and formalize the vision of geo-alignment, outline how it goes beyond pluralistic alignment by offering learnable spatially explicit patterns, and suggest concrete avenues for future research. Krzysztof Janowicz, Zilong Liu 0003, Gengchen Mai, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao 0001 |
SIGSPATIAL/GIS | 1 |
| 2025 | MobilityDL: a review of deep learning from trajectory dataabstractAbstract Trajectory data combines the complexities of time series, spatial data, and (sometimes irrational) movement behavior. As data availability and computing power have increased, so has the popularity of deep learning from trajectory data. This review paper provides the first comprehensive overview of deep learning approaches for trajectory data. We have identified eight specific mobility use cases which we analyze with regards to the deep learning models and the training data used. Besides a comprehensive quantitative review of the literature since 2018, the main contribution of our work is the data-centric analysis of recent work in this field, placing it along the mobility data continuum which ranges from detailed dense trajectories of individual movers (quasi-continuous tracking data), to sparse trajectories (such as check-in data), and aggregated trajectories (crowd information). Anita Graser, Anahid N. Jalali, Jasmin Lampert, Axel Weissenfeld, Krzysztof Janowicz |
GeoInformatica | 5 |
| 2025 | GeoFM: how will geo-foundation models reshape spatial data science and GeoAI?abstractThe emerging field of geo-foundation models (GeoFM) has the potential to reshape GeoAI and spatial data science research, education, and practice. In this work, we motivate and define the term and put it into its historic context within GeoAI and spatial data science more broadly. Next, we review core datasets, models, and benchmarks. Based on this overview of the state-of-the-art, we introduce key research challenges for future GeoFM research, such as GeoAI scaling laws, geo-alignment of AI, truly multimodal GeoFM, and so on. Finally, we discuss potential risks of GeoFM research and outline the road ahead with a specific focus on the increasing role of international large-scale collaborations and the future of GeoAI and spatial data science education. Krzysztof Janowicz, Gengchen Mai, Weiming Huang 0001, Rui Zhu 0008, Ni Lao, Ling Cai 0002 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2025 | Foundation models for geospatial reasoning: assessing the capabilities of large language models in understanding geometries and topological spatial relationsabstractAI foundation models have demonstrated some capabilities for the understanding of geospatial semantics. However, applying such pre-trained models directly to geospatial datasets remains challenging due to their limited ability to represent and reason with geographical entities, specifically vector-based geometries and natural language descriptions of complex spatial relations. To address these issues, we investigate the extent to which a well-known-text (WKT) representation of geometries and their spatial relations (e.g., topological predicates) are preserved during spatial reasoning when the geospatial vector data are passed to large language models (LLMs) including GPT-3.5-turbo, GPT-4, and DeepSeek-R1-14B. Our workflow employs three distinct approaches to complete the spatial reasoning tasks for comparison, i.e., geometry embedding-based, prompt engineering-based, and everyday language-based evaluation. Our experiment results demonstrate that both the embedding-based and prompt engineering-based approaches to geospatial question-answering tasks with GPT models can achieve an accuracy of over 0.6 on average for the identification of topological spatial relations between two geometries. Among the evaluated models, GPT-4 with few-shot prompting achieved the highest performance with over 0.66 accuracy on topological spatial relation inference. Additionally, GPT-based reasoner is capable of properly comprehending inverse topological spatial relations and including an LLM-generated geometry can enhance the effectiveness for geographic entity retrieval. GPT-4 also exhibits the ability to translate certain vernacular descriptions about places into formal topological relations, and adding the geometry-type or place-type context in prompts may improve inference accuracy, but it varies by instance. The performance of these spatial reasoning tasks unveils the strengths and limitations of the current LLMs in the processing and comprehension of geospatial vector data and offers valuable insights for the refinement of LLMs with geographical knowledge towards the development of geo-foundation models capable of geospatial reasoning. Yuhan Ji, Song Gao 0001, Ivan Majic, Krzysztof Janowicz |
Int. J. Geogr. Inf. Sci. | 5 |
| 2025 | The KnowWhereGraph ontologyabstractKnowWhereGraph is one of the largest fully publicly available geospatial knowledge graphs. It includes data from 30 layers on natural hazards (e.g., hurricanes, wildfires), climate variables (e.g., air temperature, precipitation), soil properties, crop and land-cover types, demographics, and human health, various place and region identifiers, among other themes. These have been leveraged through the graph by a variety of applications to address challenges in food security and agricultural supply chains; sustainability related to soil conservation practices and farm labor; and delivery of emergency humanitarian aid following a disaster. In this paper, we introduce the ontology that acts as the schema for KnowWhereGraph. This broad overview provides insight into the requirements and design specifications for the graph and its schema, including the development methodology (modular ontology modeling) and the resources utilized to implement, materialize, and deploy KnowWhereGraph with its end-user interfaces and public query SPARQL endpoint. Cogan Shimizu, Shirly Stephen, Adrita Barua, Ling Cai 0002, Antrea Christou, Kitty Currier, Abhilekha Dalal, Colby K. Fisher, Pascal Hitzler, Krzysztof Janowicz, Wenwen Li 0002, Zilong Liu 0003, Mohammad Saeid Mahdavinejad, Gengchen Mai, Dean Rehberger, Mark Schildhauer, Meilin Shi, Sanaz Saki Norouzi, Yuanyuan Tian 0002, Joseph Zalewski, Lu Zhou 0005, Rui Zhu 0008 |
J. Web Semant. | 10 |
| 2023 | Building Privacy-Preserving and Secure Geospatial Artificial Intelligence Foundation Models (Vision Paper)abstractIn recent years we have seen substantial advances in foundation models for artificial intelligence, including language, vision, and multimodal models. Recent studies have highlighted the potential of using foundation models in geospatial artificial intelligence, known as GeoAI Foundation Models, for geographic question answering, remote sensing image understanding, map generation, and location-based services, among others. However, the development and application of GeoAI foundation models can pose serious privacy and security risks, which have not been fully discussed or addressed to date. This paper introduces the potential privacy and security risks throughout the lifecycle of GeoAI foundation models and proposes a comprehensive blueprint for research directions and preventative and control strategies. Through this vision paper, we hope to draw the attention of researchers and policymakers in geospatial domains to these privacy and security risks inherent in GeoAI foundation models and advocate for the development of privacy-preserving and secure GeoAI foundation models. Jinmeng Rao, Song Gao 0001, Gengchen Mai, Krzysztof Janowicz |
SIGSPATIAL/GIS | 4 |
| 2023 | HyperQuaternionE: A hyperbolic embedding model for qualitative spatial and temporal reasoningabstractQualitative spatial/temporal reasoning (QSR/QTR) plays a key role in research on human cognition, e.g., as it relates to navigation, as well as in work on robotics and artificial intelligence. Although previous work has mainly focused on various spatial and temporal calculi, more recently representation learning techniques such as embedding have been applied to reasoning and inference tasks such as query answering and knowledge base completion. These subsymbolic and learnable representations are well suited for handling noise and efficiency problems that plagued prior work. However, applying embedding techniques to spatial and temporal reasoning has received little attention to date. In this paper, we explore two research questions: (1) How do embedding-based methods perform empirically compared to traditional reasoning methods on QSR/QTR problems? (2) If the embedding-based methods are better, what causes this superiority? In order to answer these questions, we first propose a hyperbolic embedding model, called HyperQuaternionE, to capture varying properties of relations (such as symmetry and anti-symmetry), to learn inversion relations and relation compositions (i.e., composition tables), and to model hierarchical structures over entities induced by transitive relations. We conduct various experiments on two synthetic datasets to demonstrate the advantages of our proposed embedding-based method against existing embedding models as well as traditional reasoners with respect to entity inference and relation inference. Additionally, our qualitative analysis reveals that our method is able to learn conceptual neighborhoods implicitly. We conclude that the success of our method is attributed to its ability to model composition tables and learn conceptual neighbors, which are among the core building blocks of QSR/QTR. Ling Cai 0002, Krzysztof Janowicz, Rui Zhu 0008, Gengchen Mai, Bo Yan 0003 |
GeoInformatica | 2 |
| 2023 | Towards general-purpose representation learning of polygonal geometries
Gengchen Mai, Chiyu Max Jiang, Rui Zhu 0008, Yao Xuan, Ling Cai 0002, Krzysztof Janowicz, Stefano Ermon, Ni Lao |
GeoInformatica | 7 |
| 2022 | LD Connect: A Linked Data Portal for IOS Press Scientometrics
Zilong Liu 0003, Meilin Shi, Krzysztof Janowicz, Blake Regalia, Stephanie Delbecque, Gengchen Mai, Rui Zhu 0008, Pascal Hitzler |
ESWC | 3 |
| 2022 | Knowledge explorer: exploring the 12-billion-statement KnowWhereGraph using faceted search (demo paper)abstractKnowledge graphs are a rapidly growing paradigm and technology stack for integrating large-scale, heterogeneous data in an AI-ready form, i.e., combining data with the formal semantics required to understand it. However, toolchains that support data synthesis and knowledge discovery through information organization, search, filtering, and visualization have been developed at a pace lagging knowledge graph technology. In this paper, we present Knowledge Explorer, an open-source faceted search interface that provides environmentally intelligent services for interactively browsing and navigating KnowWhereGraph. Currently one of the largest open knowledge graphs, KnowWhereGraph contains over 12 billion statements with rich spatial and temporal information from more than 30 data layers. With an extensive collection of facets, Knowledge Explorer enables spatial, temporal, full-text, and expert search with dereferencing functionality to support "follow-your-nose"exploration, and it allows users to narrow their search by selecting facets. Given the size of the underlying graph and dependency on GeoSPARQL, we have improved query performance by implementing Elasticsearch indexing, spatial query generation, and caching. Knowledge Explorer is capable of retrieving information within seconds, answering a wide variety of competency questions posed by researchers, humanitarian relief organizations, and the broader public, thus helping better perform tasks such as cross-gazetteer place retrieval and disaster assessment from global to local geographic scales. Zilong Liu 0003, Zhining Gu, Thomas Thelen, Seila Gonzalez Estrecha, Rui Zhu 0008, Colby K. Fisher, Anthony D'Onofrio, Cogan Shimizu, Krzysztof Janowicz, Mark Schildhauer, Shirly Stephen, Dean Rehberger, Wenwen Li 0002, Pascal Hitzler |
SIGSPATIAL/GIS | 9 |
| 2022 | International Workshop on Knowledge Graphs: Open Knowledge NetworkabstractKnowledge networks/graphs provide a powerful approach for data discovery, integration, and reuse. The NSF's new Convergence Accelerator program, which focuses on transitioning research to practice and translational research, announced Track A on the Open Knowledge Network (OKN). The program calls for multidisciplinary and multi-sector teams to work together to build a cooperative and shared open knowledge network infrastructure to drive innovation across science, engineering, and humanities. This workshop aims to invite researchers, practitioners, and the general public to brainstorm the ideas related to OKN, collaboratively build KGs for different domains or applications, develop AI algorithms to provide intelligent services based on OKN, and discuss the social and economic implications related to OKN. Ying Ding 0001, Amit P. Sheth, Krzysztof Janowicz, Sergio Baranzini, Sharat Israni, Ilkay Altintas, Lilit Yeghiazarian, Ellie Young, Sam Klein |
KDD | 3 |
| 2022 | A review of location encoding for GeoAI: methods and applicationsabstractA common need for artificial intelligence models in the broader geoscience is to encode various types of spatial data, such as points, polylines, polygons, graphs, or rasters, in a hidden embedding space so that they can be readily incorporated into deep learning models. One fundamental step is to encode a single point location into an embedding space, such that this embedding is learning-friendly for downstream machine learning models. We call this process location encoding. However, there lacks a systematic review on location encoding, its potential applications, and key challenges that need to be addressed. This paper aims to fill this gap. We first provide a formal definition of location encoding, and discuss the necessity of it for GeoAI research. Next, we provide a comprehensive survey about the current landscape of location encoding research. We classify location encoding models into different categories based on their inputs and encoding methods, and compare them based on whether they are parametric, multi-scale, distance preserving, and direction aware. We demonstrate that existing location encoders can be unified under one formulation framework. We also discuss the application of location encoding. Finally, we point out several challenges that need to be solved in the future. Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao |
Int. J. Geogr. Inf. Sci. | 2 |
| 2022 | Reasoning over higher-order qualitative spatial relations via spatially explicit neural networksabstractQualitative spatial reasoning has been a core research topic in GIScience and AI for decades. It has been adopted in a wide range of applications such as wayfinding, question answering, and robotics. Most developed spatial inference engines use symbolic representation and reasoning, which focuses on small and densely connected data sets, and struggles to deal with noise and vagueness. However, with more sensors becoming available, reasoning over spatial relations on large-scale and noisy geospatial data sets requires more robust alternatives. This paper, therefore, proposes a subsymbolic approach using neural networks to facilitate qualitative spatial reasoning. More specifically, we focus on higher-order spatial relations as those have been largely ignored due to the binary nature of the underlying representations, e.g. knowledge graphs. We specifically explore the use of neural networks to reason over ternary projective relations such as between. We consider multiple types of spatial constraint, including higher-order relatedness and the conceptual neighborhood of ternary projective relations to make the proposed model spatially explicit. We introduce evaluating results demonstrating that the proposed spatially explicit method substantially outperforms the existing baseline by about 20%. Rui Zhu 0008, Krzysztof Janowicz, Ling Cai 0002, Gengchen Mai |
Int. J. Geogr. Inf. Sci. | 2 |
| 2021 | Semantic Compression with Region Calculi in Nested Hierarchical GridsabstractWe propose the combining of region connection calculi with nested hierarchical grids for representing spatial region data in the context of knowledge graphs, thereby avoiding reliance on vector representations. We present a resulting region calculus, and provide qualitative and formal evidence that this representation can be favorable with large data volumes in the context of knowledge graphs; in particular we study means of efficiently choosing which triples to store to minimize space requirements when data is represented this way, and we provide an algorithm for finding the smallest possible set of triples for this purpose including an asymptotic measure of the size of this set for a special case. We prove that a known constraint calculus is adequate for the reconstruction of all triples describing a region from such a pruned representation, but problematic for reasoning with hierarchical grids in general. Joseph Zalewski, Pascal Hitzler, Krzysztof Janowicz |
SIGSPATIAL/GIS | 3 |
| 2021 | Providing Humanitarian Relief Support through Knowledge GraphsabstractDisasters are often unpredictable and complex events, requiring humanitarian organizations to understand and respond to many different issues simultaneously and immediately. Often the biggest challenge to improving the effectiveness of the response is quickly finding the right expert, with the right expertise concerning a specific disaster type/disaster and geographic region. To assist in achieving such a goal, this paper demonstrates a knowledge graph-based search engine developed on top of an expert knowledge graph. It accommodates three modes of information retrieval, including a follow-your-nose search, an expert similarity search, and a SPARQL query interface. We will demonstrate utilizing the system to rapidly navigate from a hazard event to a specific expert who may be helpful, for example. More importantly, as the data is fully integrated including links between hazards and their abstract topics, we can find experts who have relevant expertise while navigating the graph. Rui Zhu 0008, Ling Cai 0002, Gengchen Mai, Cogan Shimizu, Colby K. Fisher, Krzysztof Janowicz, Anna Lopez-Carr, Andrew Schroeder, Mark Schildhauer, Yuanyuan Tian 0002, Shirly Stephen, Zilong Liu 0003 |
K-CAP | 6 |
| 2021 | Time in a Box: Advancing Knowledge Graph Completion with Temporal ScopesabstractAlmost all statements in knowledge bases have a temporal scope during which they are valid. Hence, knowledge base completion (KBC) on temporal knowledge bases (TKB), where each statementmay be associated with a temporal scope, has attracted growing attention. Prior works assume that each statement in a TKBmust be associated with a temporal scope. This ignores the fact that the scoping information is commonly missing in a KB. Thus prior work is typically incapable of handling generic use cases where a TKB is composed of temporal statements with/without a known temporal scope. In order to address this issue, we establish a new knowledge base embedding framework, called TIME2BOX, that can deal with atemporal and temporal statements of different types simultaneously. Our main insight is that answers to a temporal query always belong to a subset of answers to a time-agnostic counterpart. Put differently, time is a filter that helps pick out answers to be correct during certain periods. We introduce boxes to represent a set of answer entities to a time-agnostic query. The filtering functionality of time is modeled by intersections over these boxes. In addition, we generalize current evaluation protocols on time interval prediction. We describe experiments on two datasets and show that the proposed method outperforms state-of-the-art (SOTA) methods on both link prediction and time prediction. Ling Cai 0002, Krzysztof Janowicz, Bo Yan 0003, Rui Zhu 0008, Gengchen Mai |
K-CAP | 2 |
| 2020 | GeoAI: spatially explicit artificial intelligence techniques for geographic knowledge discovery and beyondabstractRecent progress in Artificial Intelligence (AI) techniques, the large-scale availability of high-quality data, as well as advances in both hardware and software to efficiently process these data, a... Krzysztof Janowicz, Song Gao 0001, Grant McKenzie, Yingjie Hu 0001, Budhendra L. Bhaduri |
Int. J. Geogr. Inf. Sci. | 1 |
| 2019 | TransGCN: Coupling Transformation Assumptions with Graph Convolutional Networks for Link PredictionabstractLink prediction is an important and frequently studied task that contributes to an understanding of the structure of knowledge graphs (KGs) in statistical relational learning. Inspired by the success of graph convolutional networks (GCN) in modeling graph data, we propose a unified GCN framework, named TransGCN, to address this task, in which relation and entity embeddings are learned simultaneously. To handle heterogeneous relations in KGs, we introduce a novel way of representing heterogeneous neighborhood by introducing transformation assumptions on the relationship between the subject, the relation, and the object of a triple. Specifically, a relation is treated as a transformation operator transforming a head entity to a tail entity. Both translation assumption in TransE and rotation assumption in RotatE are explored in our framework. Additionally, instead of only learning entity embeddings in the convolution-based encoder while learning relation embeddings in the decoder as done by the state-of-art models, e.g., R-GCN, the TransGCN framework trains relation embeddings and entity embeddings simultaneously during the graph convolution operation, thus having fewer parameters compared with R-GCN. Experiments show that our models outperform the-state-of-arts methods on both FB15K-237 and WN18RR. Ling Cai 0002, Bo Yan 0003, Gengchen Mai, Krzysztof Janowicz, Rui Zhu 0008 |
K-CAP | 4 |
| 2019 | Contextual Graph Attention for Answering Logical Queries over Incomplete Knowledge GraphsabstractRecently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attention mechanism to handle the unequal contribution of different query paths. However, commonly used graph attention assumes that the center node embedding is provided, which is unavailable in this task since the center node is to be predicted. To solve this problem we propose a multi-head attention-based end-to-end logical query answering model, called Contextual Graph Attention model (CGA), which uses an initial neighborhood aggregation layer to generate the center embedding, and the whole model is trained jointly on the original KG structure as well as the sampled query-answer pairs. We also introduce two new datasets, DB18 and WikiGeo19, which are rather large in size compared to the existing datasets and contain many more relation types, and use them to evaluate the performance of the proposed model. Our result shows that the proposed CGA with fewer learnable parameters consistently outperforms the baseline models on both datasets as well as Bio dataset. Gengchen Mai, Krzysztof Janowicz, Bo Yan 0003, Rui Zhu 0008, Ling Cai 0002, Ni Lao |
K-CAP | 2 |
| 2019 | SOSA: A lightweight ontology for sensors, observations, samples, and actuators
Krzysztof Janowicz, Armin Haller, Simon J. D. Cox, Danh Le Phuoc, Maxime Lefrançois |
J. Web Semant. | 1 |
| 2018 | Support and Centrality: Learning Weights for Knowledge Graph Embedding Models
Gengchen Mai, Krzysztof Janowicz, Bo Yan 0003 |
EKAW | 2 |
| 2018 | GNIS-LD: Serving and Visualizing the Geographic Names Information System Gazetteer as Linked Data
Blake Regalia, Krzysztof Janowicz, Gengchen Mai, Dalia Varanka, E. Lynn Usery |
ESWC | 2 |
| 2017 | How "Alternative" are Alternative Facts?: Towards Measuring Statement Coherence via Spatial AnalysisabstractFollowing the AAA principle by which anybody can say anything about any topic, the Web is no stranger to alternative facts. Nonetheless, with the increasing volume and velocity at which content is being published and difficulties to assess the credibility of information and the trustworthiness of sources, alternative facts are becoming a major challenge and an instrument for spreading disinformation. Interestingly, the diversity of today's data sources can also help us to counter alternative facts by measuring their coherence, i.e., the degree to which data from one source confirms or contradict data from another source. While a single dataset can be biased towards supporting or discrediting a statement, the diverse sources of data across media types that are publicly accessible today offer unique perspectives on which to assess a given statement. To give an intuitive example, a statement about the comparison of crowd sizes should align with photos of said crowds. However, these photos could be taken at different times, from different viewpoints, and could lead to different, sample-based estimations. Adding further data from heterogeneous sources, such as metro ridership, can either further support a statement or contradict it. We use three thought experiments to discuss the role of geographic data, knowledge graphs, and spatial analysis in approaching alternative facts from a novel angle, namely by studying their coherence, i.e., whether they align with other statements, instead of trying to falsify them. In doing so, we aim at increasing the costs for maintaining alternative facts. Krzysztof Janowicz, Grant McKenzie |
SIGSPATIAL/GIS | 1 |
| 2017 | From ITDL to Place2Vec: Reasoning About Place Type Similarity and Relatedness by Learning Embeddings From Augmented Spatial ContextsabstractUnderstanding, representing, and reasoning about Points Of Interest (POI) types such as Auto Repair, Body Shop, Gas Stations, or Planetarium, is a key aspect of geographic information retrieval, recommender systems, geographic knowledge graphs, as well as studying urban spaces in general, e.g., for extracting functional or vague cognitive regions from user-generated content. One prerequisite to these tasks is the ability to capture the similarity and relatedness between POI types. Intuitively, a spatial search that returns body shops or even gas stations in the absence of auto repair places is still likely to satisfy some user needs while returning planetariums will not. Place hierarchies are frequently used for query expansion, but most of the existing hierarchies are relatively shallow and structured from a single perspective, thereby putting POI types that may be closely related regarding some characteristics far apart from another. This leads to the question of how to learn POI type representations from data. Models such as Word2Vec that produces word embeddings from linguistic contexts are a novel and promising approach as they come with an intuitive notion of similarity. However, the structure of geographic space, e.g., the interactions between POI types, differs substantially from linguistics. In this work, we present a novel method to augment the spatial contexts of POI types using a distance-binned, information-theoretic approach to generate embeddings. We demonstrate that our work outperforms Word2Vec and other models using three different evaluation tasks and strongly correlates with human assessments of POI type similarity. We published the resulting embeddings for 570 place types as well as a collection of human similarity assessments online for others to use. Bo Yan 0003, Krzysztof Janowicz, Gengchen Mai, Song Gao 0001 |
SIGSPATIAL/GIS | 2 |
| 2017 | A data-synthesis-driven method for detecting and extracting vague cognitive regionsabstractCognitive regions and places are notoriously difficult to represent in geographic information science and systems. The exact delineation of cognitive regions is challenging insofar as borders are vague, membership within the regions varies non-monotonically, and raters cannot be assumed to assess membership consistently and homogeneously. In a study published in this journal in 2014, researchers devised a novel grid-based task in which participants rated the membership of individual cells in a given region and contrasted this approach to a standard boundary-drawing task. Specifically, the authors assessed the vague cognitive regions of Northern California and Southern California. The boundary between these cognitive regions was found to have variable width, and region membership peaked not at the most northern or southern cells but at substantially less extreme latitudes. The authors thus concluded that region membership is about attitude, not just latitude. In the present work, we reproduce this study by approaching it from a computational fourth-paradigm perspective, i.e., by the synthesis of high volumes of heterogeneous data from various sources. We compare the regions which we identify to those from the human-participants study of 2014, identifying differences and commonalities. Our results show a significant positive correlation to those in the original study. Beyond the extracted regions themselves, we compare and contrast the empirical and analytical approaches of these two methods, one a conventional human-participants study and the other an application of increasingly popular data-synthesis-driven research methods in GIScience. Song Gao 0001, Krzysztof Janowicz, Daniel R. Montello, Yingjie Hu 0001, Jiue-An Yang, Grant McKenzie, Yiting Ju, Benjamin Adams, Bo Yan 0003 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2016 | Things and Strings: Improving Place Name Disambiguation from Short Texts by Combining Entity Co-Occurrence with Topic Modeling
Yiting Ju, Benjamin Adams, Krzysztof Janowicz, Yingjie Hu 0001, Bo Yan 0003, Grant McKenzie |
EKAW | 3 |
| 2016 | VOLT: A Provenance-Producing, Transparent SPARQL Proxy for the On-Demand Computation of Linked Data and its Application to Spatiotemporally Dependent Data
Blake Regalia, Krzysztof Janowicz, Song Gao 0001 |
ESWC | 2 |
| 2016 | ADCN: an anisotropic density-based clustering algorithmabstractIn this work we introduce an anisotropic density-based clustering algorithm. It outperforms DBSCAN and OPTICS for the detection of anisotropic spatial point patterns and performs equally well in cases that do not explicitly benefit from an anisotropic perspective. ADCN has the same time complexity as DBSCAN and OPTICS, namely O(n log n) when using a spatial index, O(n2) otherwise. Gengchen Mai, Krzysztof Janowicz, Yingjie Hu 0001, Song Gao 0001 |
SIGSPATIAL/GIS | 2 |
| 2016 | Task-oriented information value measurement based on space-time prismsabstractRecent years have witnessed a large increase in the amount of information available from the Web and many other sources. Such an information deluge presents a challenge for individuals who have to identify useful information items to complete particular tasks in hand. Information value theory (IVT) from economics and artificial intelligence has provided some guidance on this issue. However, existing IVT studies often focus on monetary values, while ignoring the spatiotemporal properties which can play important roles in everyday tasks. In this paper, we propose a theoretical framework for task-oriented information value measurement. This framework integrates IVT with the space-time prism from time geography and measures the value of information based on its impact on an individual’s space-time prisms and its capability of improving task planning. We develop and formalize this framework by extending the utility function from space-time accessibility studies and elaborate it using a simplified example from time geography. We conduct a simulation on a real-world transportation network using the proposed framework. Our research could be applied to improving information display on small-screen mobile devices (e.g., smartwatches) by assigning priorities to different information items. Yingjie Hu 0001, Krzysztof Janowicz, Yuqi Chen 0003 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2015 | The GeoLink Modular Oceanography Ontology
Adila Krisnadhi, Yingjie Hu 0001, Krzysztof Janowicz, Pascal Hitzler, Robert A. Arko, Suzanne Carbotte, Cynthia Chandler, Michelle Cheatham, Douglas Fils, Tim Finin, Matthew B. Jones, Nazifa Karima, Kerstin A. Lehnert, Audrey Mickle, Thomas W. Narock, Margaret O'Brien, Lisa Raymond, Adam Shepherd, Mark Schildhauer, Peter H. Wiebe |
ISWC (2) | 3 |
| 2015 | Thematic signatures for cleansing and enriching place-related linked dataabstractThere has been significant progress transforming semi-structured data about places into knowledge graphs that can be used in a wide variety of geographic information systems such as digital gazetteers or geographic information retrieval systems. For instance, in addition to information about events, actors, and objects, DBpedia contains data about hundreds of thousands of places from Wikipedia and publishes it as Linked Data. Repositories that store data about places are among the most interlinked hubs on the Linked Data cloud. However, most content about places resides in unstructured natural language text, and therefore it is not captured in these knowledge graphs. Instead, place representations are limited to facts such as their population counts, geographic locations, and relations to other entities, for example, headquarters of companies or historical figures. In this paper, we present a novel method to enrich the information stored about places in knowledge graphs using thematic signatures that are derived from unstructured text through the process of topic modeling. As proof of concept, we demonstrate that this enables the automatic categorization of articles into place types defined in the DBpedia ontology (e.g., mountain) and also provides a mechanism to infer relationships between place types that are not captured in existing ontologies. This method can also be used to uncover miscategorized places, which is a common problem arising from the automatic lifting of unstructured and semi-structured data. Benjamin Adams, Krzysztof Janowicz |
Int. J. Geogr. Inf. Sci. | 2 |
| 2013 | An Ontology Design Pattern for Cartographic Map Scaling
David Carral, Simon Scheider, Krzysztof Janowicz, Charles Vardeman, Adila Krisnadhi, Pascal Hitzler |
ESWC | 3 |
| 2013 | A spatiotemporal scientometrics framework for exploring the citation impact of publications and scientistsabstractThe research field of scientometrics is concerned with measuring and analyzing science. In practice, this is often done by restricting the impact of publications, journals, and researchers to a mere frequency. However, scientific activities (co-publication, citation, labor mobility) display clear spatiotemporal patterns, and such patterns have rarely been considered in traditional scientometrics. In this work we focus on the study of citations and present a spatiotemporal scientometrics framework to measure the citation impact of research output by taking physical space, place, and time into account. Specifically, we use the statistics of categorical places (institutions, cities, and countries), spatiotemporal kernel density estimations, cartograms, distance distribution curves, and point-pattern analysis to identify spatiotemporal citation patterns. Moreover, we propose a series of s-indices, such as S_institution-index, S_city-index, and S_country-index to evaluate a scientist's impact as a complement to non-spatial citation indicators, e.g., h-index and g-index. In addition, we have developed an interactive web application which allows users to visually explore research topics, authors, publications, as well as the spread of citations through space and time. Our work offers insights on the role of location in scientific knowledge diffusion. Song Gao 0001, Yingjie Hu 0001, Krzysztof Janowicz, Grant McKenzie |
SIGSPATIAL/GIS | 3 |
| 2013 | Weighted multi-attribute matching of user-generated points of interestabstractTo a large degree, the attraction of Big Data lies in the variety of its heterogeneous multi-thematic and multi-dimensional data sources and not merely its volume. To fully exploit this variety, however, requires conflation. This is a two step process. First, one has to establish identity relations between information entities across the different data sources; and second, attribute values have to be merged according to certain procedures which avoid logical contradictions. The first step, also called matching, can be thought of as a weighted combination of common attributes according to some similarity measures. In this work, we propose such a matching based on multiple attributes of Points of Interests (POI) from the Location-based Social Network Foursquare and the Yelp local directory service. While both contain overlapping attributes that can be use for matching, they have specific strengths and weaknesses which makes their conflation desirable. We present a weighted multi-attribute matching strategy and evaluate its performance. Our strategy can automatically match 97% of randomly selected Yelp POI to their corresponding Foursquare entities. Grant McKenzie, Krzysztof Janowicz, Benjamin Adams |
SIGSPATIAL/GIS | 2 |
| 2013 | A Linked-Data-Driven and Semantically-Enabled Journal Portal for ScientometricsabstractThe Semantic Web journal by IOS Press follows a unique open and transparent process during which each submitted manuscript is available online together with the full history of its successive decision statuses, assigned editors, solicited and voluntary reviewers, their full text reviews, and in many cases also the authors’ response letters. Combined with a highly-customized, Drupal-based journal management system, this provides the journal with semantically rich manuscript time lines and networked data about authors, reviewers, and editors. These data are now exposed using a SPARQL endpoint, an extended Bibo ontology, and a modular Linked Data portal that provides interactive scientometrics based on established and new analysis methods. The portal can be customized for other journals as well. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yingjie Hu 0001, Krzysztof Janowicz, Grant McKenzie, Kunal Sengupta, Pascal Hitzler |
ISWC (2) | 2 |
| 2012 | Improving personal information management by integrating activities in the physical world with the semantic desktopabstractSemantic desktops are a novel approach to improve user interfaces by recording, semantically annotating, and learning from the user's activities to create a personalized user experience and improve search. Such activities, however, are restricted to the information universe, i.e., they only cover events on the local desktop. A next step towards smart mobile devices is the integration of those desktop events with the user's activities in the physical world. Establishing such mappings enables the device to draw conclusions from the recorded desktop events to those that the user is likely performing in the physical world. A Personal Information Management (PIM) system can then better assist the user in task planning and routing. In this work, we propose activity ontologies as blueprints to model the user's activities in the physical world, and use these ontologies to link the Semantic Desktop and the information available on the Web of Linked Data. We discuss the principles of designing the activity ontologies and how to employ them to associate local files and applications with complementary information from the Web. We design a specific activity ontology for a conference use case and present a user interface that extends the Zeitgeist Semantic Desktop to evaluate our approach. Yingjie Hu 0001, Krzysztof Janowicz |
SIGSPATIAL/GIS | 2 |
| 2012 | A logical geo-ontology design pattern for quantifying over typesabstractOntology design patterns ease the engineering of ontologies, improve their quality, foster reusability, and support the alignment of ontologies by acting as common building blocks or strategies for reoccurring modeling problems. This makes ontology design patterns key enablers of semantic interoperability and, hence, a crucial technology for representing the body of knowledge of such heterogeneous domains as the geosciences. While different types of patterns can be distinguished, existing work on geo-ontology design patterns has solely focused on content patterns, i.e., design solutions for domain classes and relationships. In this work, we propose a logical pattern that addresses a frequent modeling problem that has hampered the development of sophisticated geo-ontologies in the past, namely how to model the quantification over types. We argue for the need for such a pattern, explain why it is difficult to model, demonstrate how to implement it using the Web Ontology Language OWL, and finally show how it can be applied to modeling concepts such as biodiversity. David Carral, Krzysztof Janowicz, Pascal Hitzler |
SIGSPATIAL/GIS | 2 |
| 2012 | On the Geo-Indicativeness of Non-Georeferenced Text
Benjamin Adams, Krzysztof Janowicz |
ICWSM | 2 |
| 2012 | The SSN ontology of the W3C semantic sensor network incubator groupabstractThe W3C Semantic Sensor Network Incubator group (the SSN-XG) produced an OWL 2 ontology to describe sensors and observations — the SSN ontology, available at http://purl.oclc.org/NET/ssnx/ssn. The SSN ontology can describe sensors in terms of capabilities, measurement processes, observations and deployments. This article describes the SSN ontology. It further gives an example and describes the use of the ontology in recent research projects. Michael Compton, Payam M. Barnaghi, Luis Bermudez, Raúl García-Castro, Óscar Corcho, Simon J. D. Cox, John B. Graybeal, Manfred Hauswirth, Cory A. Henson, Arthur Herzog, Vincent Huang 0002, Krzysztof Janowicz, W. David Kelsey, Danh Le Phuoc, Laurent Lefort, Myriam Leggieri, Holger Neuhaus, Andriy Nikolov, Kevin R. Page, Alexandre Passant, Amit P. Sheth, Kerry L. Taylor |
J. Web Semant. | 12 |
| 2011 | Constructing geo-ontologies by reification of observation dataabstractThe semantic integration of heterogeneous, spatiotemporal information is a major challenge for achieving the vision of a multi-thematic and multi-perspective Digital Earth. The Semantic Web technology stack has been proposed to address the integration problem by knowledge representation languages and reasoning. However approaches such as the Web Ontology Languages (OWL) were developed with decidability in mind. They do not integrate well with established modeling paradigms in the geosciences that are dominated by numerical and geometric methods. Additionally, work on the Semantic Web is mostly feature-centric and a field-based view is difficult to integrate. A layer specifying the transition from observation data to classes and relations is missing. In this work we combine OWL with geometric and topological language constructs based on similarity spaces. Our approach provides three main benefits. First, class constructors can be built from a larger palette of mathematical operations based on vector algebra. Second, it affords the representation of prototype-based classes. Third, it facilitates the representation of classes derived from machine learning classifiers that utilize a multi-dimensional feature space. Instead of following a one-size-fits-all approach, our work allows one to derive contextualized OWL ontologies by reification of observation data. Benjamin Adams, Krzysztof Janowicz |
GIS | 2 |
| 2011 | What you are is when you are: the temporal dimension of feature types in location-based social networksabstractFeature types play a crucial role in understanding and analyzing geographic information. Usually, these types are defined, standardized, and controlled by domain experts and cover geographic features on the mesoscale level, e.g., populated places, forests, or lakes. While feature types also underlie most Location-Based Services (LBS), assigning a consistent typing schema for Points Of Interest (POI) across different data sets is challenging. In case of Volunteered Geographic Information (VGI), types are assigned as tags by a heterogeneous community with different backgrounds and applications in mind. Consequently, VGI research is shifting away from data completeness and positional accuracy as quality measures towards attribute accuracy. As tags can be assigned by everybody and have no formal or stable definition, we propose to study category tags via indirect observations. We extract user check-ins from massive real-world data crawled from Location-based Social Networks to understand the temporal dimension of Points Of Interest. While users may assign different category tags to places, we argue that their temporal characteristics, e.g., opening times, will show distinguishable patterns. Mao Ye 0002, Krzysztof Janowicz, Christoph Mülligann, Wang-Chien Lee |
GIS | 2 |
| 2011 | On the semantic annotation of places in location-based social networksabstractIn this paper, we develop a semantic annotation technique for location-based social networks to automatically annotate all places with category tags which are a crucial prerequisite for location search, recommendation services, or data cleaning. Our annotation algorithm learns a binary support vector machine (SVM) classifier for each tag in the tag space to support multi-label classification. Based on the check-in behavior of users, we extract features of places from i) explicit patterns (EP) of individual places and ii) implicit relatedness (IR) among similar places. The features extracted from EP are summarized from all check-ins at a specific place. The features from IR are derived by building a novel network of related places (NRP) where similar places are linked by virtual edges. Upon NRP, we determine the probability of a category tag for each place by exploring the relatedness of places. Finally, we conduct a comprehensive experimental study based on a real dataset collected from a location-based social network, Whrrl. The results demonstrate the suitability of our approach and show the strength of taking both EP and IR into account in feature extraction. Mao Ye 0002, Dong Shou, Wang-Chien Lee, Peifeng Yin, Krzysztof Janowicz |
KDD | 5 |
| 2009 | SIM-DLA: A Novel Semantic Similarity Measure for Description Logics Reducing Inter-concept to Inter-instance Similarity
Krzysztof Janowicz, Marc Wilkes |
ESWC | 1 |
| 2009 | An agenda for the next generation gazetteer: geographic information contribution and retrievalabstractGazetteers are key components of georeferenced information systems, including applications such as Web-based mapping services. Existing gazetteers lack the capabilities to fully integrate user-contributed and vernacular geographic information, as well as to support complex queries. To address these issues, a next generation gazetteer should leverage formal semantics, harvesting of implicit geographic information - such as geotagged photos - as well as models of trust for contributors. In this paper, we discuss these requirements in detail. We elucidate how existing standards can be integrated to realize a gazetteer infrastructure allowing for bottom-up contribution as well as information exchange between different gazetteers. We show how to ensure the quality of user-contributed information and demonstrate how to improve querying and navigation using semantics-based information retrieval. Carsten Keßler, Krzysztof Janowicz, Mohamed Bishr |
GIS | 2 |
| 2008 | The role of ontology in improving gazetteer interactionabstractGazetteers are more than basic place name directories containing names and locations for named geographic places. Most of them contain additional information, including a categorization of gazetteer entries using a typing scheme. This paper focuses on the nature of these categorization schemes. We argue that gazetteers can benefit from an ontological approach to typing schemes, providing a formalization that will better support gazetteer applications, maintenance, interoperability, and semi‐automatic feature annotation. We discuss the process of developing such an ontology as a modification of an existing feature type thesaurus; the difficulties in mapping from thesauri to ontologies are described in detail. To demonstrate the benefits of a categorization based on ontologies, a new gazetteer Web (and programming) interface is introduced and the impact on gazetteer interoperability is discussed. Krzysztof Janowicz, Carsten Keßler |
Int. J. Geogr. Inf. Sci. | 1 |