EDBT 2026 Demo / reviewers in the wild / expert
Grant McKenzie
dblp:32/11518 · also Grant D. McKenzie
· DBLP profile ↗
18ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-3247-2777ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WeatherArchive: A Benchmark for Retrieval-Augmented Reasoning over Historical Weather ArchivesabstractHistorical news segments on weather events are collections of enduring primary source records that offer rich, untapped narratives of how societies have experienced and responded to extreme weather events. These qualitative accounts provide insights into societal vulnerability and resilience that are largely absent from meteorological records, making them valuable for climate scientists to understand societal responses. However, their large scale, noise in optical character recognition (OCR), and archaic language make it difficult to transform them into structured knowledge for climate research. To address this challenge, we introduce øurmethod, the first large-scale benchmark for evaluating end-to-end retrieval-augmented generation (RAG) systems on historical weather archives. WeatherArchive-Bench comprises two tasks: WeatherArchive-Retrieval, which measures a system's ability to locate historically relevant news segments from over one million archival news segments, and WeatherArchive-Assessment, which evaluates whether Large Language Models (LLMs) can classify societal vulnerability and resilience indicators from extreme weather narratives and answer queries using the segments retrieved. Extensive experiments across sparse, dense, and re-ranking retrievers, as well as a diverse set of LLMs, reveal that dense retrievers often fail on historical terminology, while LLMs frequently misinterpret vulnerability and resilience concepts. These findings highlight key limitations in reasoning about complex societal indicators and provide insights for designing more robust climate-focused RAG systems from archival contexts. The constructed dataset and evaluation framework are available at: https://github.com/Weather-Archival-Rescue/WeatherArchive-Bench. Yongan Yu, Xianda Du, Qingchen Hu, Jingwei Ni, Dan Qiang, Grant McKenzie, Renée Sieber, Fengran Mo |
SIGIR | 8 |
| 2025 | Whose Truth? Pluralistic Geo-Alignment for (Agentic) AIabstractAI alignment describes the challenge of ensuring (future) AI systems behave in accordance with societal norms, values, and goals. Alignment is now central to research on foundation models and AI agents. Most recent work focuses on methods to prevent potentially harmful biases, account for social inequalities, improve AI safety, and enhance explainability. Notably, the debiasing 'corrections' applied to various stages of AI/ML workflows may lead to outcomes that diverge strongly from current statistical realities on the ground. For instance, text-to-image models may depict a balanced gender ratio of company leadership, despite existing imbalances. However, an often overlooked dimension is the geographic variability of alignment. What is considered appropriate, truthful, or legal can vary greatly between regions due to cultural differences, political realities, or legislation. Hence, some model outputs align without further knowledge of the user's geospatial context, while others are highly sensitive to it. Put differently, whether these outputs align varies geographically. E.g., statements about Kashmir cannot be generated without understanding the user's origin and current location. From a common-sense perspective, this problem is hardly new. In fact, Google Maps will render different administrative borders based on the user's location. Interestingly, in both knowledge representation and representation learning, spatiotemporal context, e.g., due to the monotonic nature of reasoning, remains a major challenge. Until very recently, these were largely theoretical problems. What is truly novel is the scale and level of automation at which AI systems now mediate knowledge, express opinions, and represent reality to millions of users across borders, often with little transparency or oversight regarding how context is handled. With agentic AI on the horizon, the urgency for pluralistic, geographically aware alignment, rather than one-size-fits-all solutions, is growing. Here, we motivate and formalize the vision of geo-alignment, outline how it goes beyond pluralistic alignment by offering learnable spatially explicit patterns, and suggest concrete avenues for future research. Krzysztof Janowicz, Zilong Liu 0003, Gengchen Mai, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao 0001 |
SIGSPATIAL/GIS | 7 |
| 2025 | A research agenda for GIScience in a time of disruptionsabstractSocial issues, AI, and climate change are just a few of the disruptive focuses impacting science. The field of GIScience is well positioned to respond to accelerating disruptions due to the interdisciplinary nature of the field and the ability of GIScience approaches to be used in support of decision-making. This manuscript aims to start a conversation that will establish a research agenda for GIScience in an age of disruptions. We outline three guiding principles: (1) focusing on the relevance and real-world impact of research, (2) adopting systems-based thinking and contextual approaches and (3) emphasizing inclusive practices. We then outline prioritized research areas organized by what topics are important focal areas (Data and Infrastructure, Artificial Intelligence, and Causality and Generalizability), and what approaches to science we should be attentive to (Impactful Open Science, Collaborative and Convergent Science, and through Diverse Participation and Partnerships). We conclude with a call to increase impact by balancing slow science with practical and policy-oriented research. We also recognize that while broad adoption of spatial approaches is a signal of GIScience's success, we should continue to work together to advance core knowledge centered on spatial thinking and approaches. Trisalyn A. Nelson, Amy E. Frazier, Peter Kedron, Somayeh Dodge, Bo Zhao 0036, Michael F. Goodchild, Alan T. Murray, Sarah E. Battersby, Lauren Bennett, Justine I. Blanford, Carmen Cabrera Arnau, Christophe Claramunt, Rachel S. Franklin, Joseph Holler, Caglar Koylu, Steven M. Manson, Grant McKenzie, Harvey J. Miller, Taylor Oshan, Sergio J. Rey, Francisco Rowe, Seda Salap-Ayça, Eric Shook, Seth Spielman, Wenfei Xu, John P. Wilson |
Int. J. Geogr. Inf. Sci. | 18 |
| 2023 | Toward a Critical Toponymy Framework for Named Entity Recognition: A Case Study of Airbnb in New York CityabstractCritical toponymy examines the dynamics of power, capital, and resistance through place names and the sites to which they refer.Studies here have traditionally focused on the semantic content of toponyms and the top-down institutional processes that produce them.However, they have generally ignored the ways in which toponyms are used by ordinary people in everyday discourse, as well as the other strategies of geospatial description that accompany and contextualize toponymic reference.Here, we develop computational methods to measure how cultural and economic capital shape the ways in which people refer to places, through a novel annotated dataset of 47,440 New York City Airbnb listings from the 2010s.Building on this dataset, we introduce a new named entity recognition (NER) model able to identify important discourse categories integral to the characterization of place.Our findings point toward new directions for critical toponymy and to a range of previously understudied linguistic signals relevant to research on neighborhood status, housing and tourism markets, and gentrification. Mikael Brunila, Jack LaViolette, Sky CH-Wang, Clara Féré, Grant McKenzie |
EMNLP | 6 |
| 2022 | Towards place-based privacy: Challenges and opportunities in the "smart" worldabstractThe emergence of “smart” technologies has given rise to new interaction models merging our physical realities with our digital environments. As a result, new privacy threats have emerged, substantially impacting both individuals and groups. In this short paper, we summarize many of the privacy challenges we face in the smart and connected world, and identify opportunities for further research. Drawing from the recent literature on geoprivacy, user-tailored privacy, and group privacy, we explore this topic through the lens of contextually aware, place-based, or platial, information analysis. Grant McKenzie |
ISTAS | 2 |
| 2020 | GeoAI: spatially explicit artificial intelligence techniques for geographic knowledge discovery and beyondabstractRecent progress in Artificial Intelligence (AI) techniques, the large-scale availability of high-quality data, as well as advances in both hardware and software to efficiently process these data, a... Krzysztof Janowicz, Song Gao 0001, Grant McKenzie, Yingjie Hu 0001, Budhendra L. Bhaduri |
Int. J. Geogr. Inf. Sci. | 3 |
| 2019 | A natural language processing and geospatial clustering framework for harvesting local place names from geotagged housing advertisementsabstractLocal place names are frequently used by residents living in a geographic region. Such place names may not be recorded in existing gazetteers, due to their vernacular nature, relative insignificance to a gazetteer covering a large area (e.g. the entire world), recent establishment (e.g. the name of a newly-opened shopping center) or other reasons. While not always recorded, local place names play important roles in many applications, from supporting public participation in urban planning to locating victims in disaster response. In this paper, we propose a computational framework for harvesting local place names from geotagged housing advertisements. We make use of those advertisements posted on local-oriented websites, such as Craigslist, where local place names are often mentioned. The proposed framework consists of two stages: natural language processing (NLP) and geospatial clustering. The NLP stage examines the textual content of housing advertisements and extracts place name candidates. The geospatial stage focuses on the coordinates associated with the extracted place name candidates and performs multiscale geospatial clustering to filter out the non-place names. We evaluate our framework by comparing its performance with those of six baselines. We also compare our result with four existing gazetteers to demonstrate the not-yet-recorded local place names discovered by our framework. Yingjie Hu 0001, Huina Mao, Grant McKenzie |
Int. J. Geogr. Inf. Sci. | 3 |
| 2017 | Juxtaposing Thematic Regions Derived from Spatial and Platial User-Generated ContentabstractTypical approaches to defining regions, districts or neighborhoods within a city often focus on place instances of a similar type that are grouped together. For example, most cities have at least one bar district defined as such by the clustering of bars within a few city blocks. In reality, it is not the presence of spatial locations labeled as bars that contribute to a bar region, but rather the popularity of the bars themselves. Following the principle that places, and by extension, place-type regions exist via the people that have given space meaning, we explore user-contributed content as a way of extracting this meaning. Kernel density estimation models of place-based social check-ins are compared to spatially tagged social posts with the goal of identifying thematic regions within the city of Los Angeles, CA. Dynamic human activity patterns, represented as temporal signatures, are included in this analysis to demonstrate how regions change over time. Grant McKenzie, Benjamin Adams |
COSIT | 1 |
| 2017 | How "Alternative" are Alternative Facts?: Towards Measuring Statement Coherence via Spatial AnalysisabstractFollowing the AAA principle by which anybody can say anything about any topic, the Web is no stranger to alternative facts. Nonetheless, with the increasing volume and velocity at which content is being published and difficulties to assess the credibility of information and the trustworthiness of sources, alternative facts are becoming a major challenge and an instrument for spreading disinformation. Interestingly, the diversity of today's data sources can also help us to counter alternative facts by measuring their coherence, i.e., the degree to which data from one source confirms or contradict data from another source. While a single dataset can be biased towards supporting or discrediting a statement, the diverse sources of data across media types that are publicly accessible today offer unique perspectives on which to assess a given statement. To give an intuitive example, a statement about the comparison of crowd sizes should align with photos of said crowds. However, these photos could be taken at different times, from different viewpoints, and could lead to different, sample-based estimations. Adding further data from heterogeneous sources, such as metro ridership, can either further support a statement or contradict it. We use three thought experiments to discuss the role of geographic data, knowledge graphs, and spatial analysis in approaching alternative facts from a novel angle, namely by studying their coherence, i.e., whether they align with other statements, instead of trying to falsify them. In doing so, we aim at increasing the costs for maintaining alternative facts. Krzysztof Janowicz, Grant McKenzie |
SIGSPATIAL/GIS | 2 |
| 2017 | A data-synthesis-driven method for detecting and extracting vague cognitive regionsabstractCognitive regions and places are notoriously difficult to represent in geographic information science and systems. The exact delineation of cognitive regions is challenging insofar as borders are vague, membership within the regions varies non-monotonically, and raters cannot be assumed to assess membership consistently and homogeneously. In a study published in this journal in 2014, researchers devised a novel grid-based task in which participants rated the membership of individual cells in a given region and contrasted this approach to a standard boundary-drawing task. Specifically, the authors assessed the vague cognitive regions of Northern California and Southern California. The boundary between these cognitive regions was found to have variable width, and region membership peaked not at the most northern or southern cells but at substantially less extreme latitudes. The authors thus concluded that region membership is about attitude, not just latitude. In the present work, we reproduce this study by approaching it from a computational fourth-paradigm perspective, i.e., by the synthesis of high volumes of heterogeneous data from various sources. We compare the regions which we identify to those from the human-participants study of 2014, identifying differences and commonalities. Our results show a significant positive correlation to those in the original study. Beyond the extracted regions themselves, we compare and contrast the empirical and analytical approaches of these two methods, one a conventional human-participants study and the other an application of increasingly popular data-synthesis-driven research methods in GIScience. Song Gao 0001, Krzysztof Janowicz, Daniel R. Montello, Yingjie Hu 0001, Jiue-An Yang, Grant McKenzie, Yiting Ju, Benjamin Adams, Bo Yan 0003 |
Int. J. Geogr. Inf. Sci. | 6 |
| 2016 | Things and Strings: Improving Place Name Disambiguation from Short Texts by Combining Entity Co-Occurrence with Topic Modeling
Yiting Ju, Benjamin Adams, Krzysztof Janowicz, Yingjie Hu 0001, Bo Yan 0003, Grant McKenzie |
EKAW | 6 |
| 2016 | Assessing the effectiveness of different visualizations for judgments of positional uncertaintyabstractMany techniques have been proposed for visualizing uncertainty in geospatial data. Previous empirical research on the effectiveness of visualizations of geospatial uncertainty has focused primarily on user intuitions rather than objective measures of performance when reasoning under uncertainty. Framed in the context of Google’s blue dot, we examined the effectiveness of four alternative visualizations for representing positional uncertainty when reasoning about self-location data. Our task presents a mobile mapping scenario in which GPS satellite location readings produce location estimates with varying levels of uncertainty. Given a known location and two smartphone estimates of that known location, participants were asked to judge which smartphone produces the better location reading, taking uncertainty into account. We produced visualizations that vary by glyph type (uniform blue circle with border vs. Gaussian fade) and visibility of a centroid dot (visible vs. not visible) to produce the four visualization formats. Participants viewing the uniform blue circle are most likely to respond in accordance with the actual probability density of points sampled from bivariate normal distributions and additionally respond most rapidly. Participants reported a number of simple heuristics on which they based their judgments, and consistency with these heuristics was highly predictive of their judgments. Grant McKenzie, Mary Hegarty, Trevor J. Barrett, Michael F. Goodchild |
Int. J. Geogr. Inf. Sci. | 1 |
| 2015 | Interpreting Visualizations of Uncertainty on Smartphone Displays
Trevor J. Barrett, Mary Hegarty, Grant McKenzie, Michael F. Goodchild |
CogSci | 3 |
| 2015 | Frankenplace: Interactive Thematic Mapping for Ad Hoc Exploratory SearchabstractAd hoc keyword search engines built using modern information retrieval methods do a good job of handling fine-grained queries. However, they perform poorly at facilitating spatial and spatially-embedded thematic exploration of the results, despite the fact that many queries, e.g. "civil war," refer to different documents and topics in different places. This is not for lack of data: geographic information, such as place names, events, and coordinates are common in unstructured document collections on the web. The associations between geographic and thematic contents in these documents can provide a rich groundwork to organize information for exploratory research. In this paper we describe the architecture of an interactive thematic map search engine, Frankenplace, designed to facilitate document exploration at the intersection of theme and place. The map interface enables a user to zoom the geographic context of their query in and out, and quickly explore through thousands of search results in a meaningful way. And by combining topic models with geographically contextualized search results, users can discover related topics based on geographic context. Frankenplace utilizes a novel indexing method called geoboost for boosting terms associated with cells on a discrete global grid. The resulting index factors in the geographic scale of the place or feature mentioned in related text, the relative textual scope of the place reference, and the overall importance of the containing document in the document network. The system is currently indexed with over 5 million documents from the web, including the English Wikipedia and online travel blog entries. We demonstrate that Frankenplace can support four distinct types of exploratory search tasks while being adaptive to scale and location of interest. Benjamin Adams, Grant McKenzie, Mark Gahegan |
WWW | 2 |
| 2013 | A spatiotemporal scientometrics framework for exploring the citation impact of publications and scientistsabstractThe research field of scientometrics is concerned with measuring and analyzing science. In practice, this is often done by restricting the impact of publications, journals, and researchers to a mere frequency. However, scientific activities (co-publication, citation, labor mobility) display clear spatiotemporal patterns, and such patterns have rarely been considered in traditional scientometrics. In this work we focus on the study of citations and present a spatiotemporal scientometrics framework to measure the citation impact of research output by taking physical space, place, and time into account. Specifically, we use the statistics of categorical places (institutions, cities, and countries), spatiotemporal kernel density estimations, cartograms, distance distribution curves, and point-pattern analysis to identify spatiotemporal citation patterns. Moreover, we propose a series of s-indices, such as S_institution-index, S_city-index, and S_country-index to evaluate a scientist's impact as a complement to non-spatial citation indicators, e.g., h-index and g-index. In addition, we have developed an interactive web application which allows users to visually explore research topics, authors, publications, as well as the spread of citations through space and time. Our work offers insights on the role of location in scientific knowledge diffusion. Song Gao 0001, Yingjie Hu 0001, Krzysztof Janowicz, Grant McKenzie |
SIGSPATIAL/GIS | 4 |
| 2013 | Weighted multi-attribute matching of user-generated points of interestabstractTo a large degree, the attraction of Big Data lies in the variety of its heterogeneous multi-thematic and multi-dimensional data sources and not merely its volume. To fully exploit this variety, however, requires conflation. This is a two step process. First, one has to establish identity relations between information entities across the different data sources; and second, attribute values have to be merged according to certain procedures which avoid logical contradictions. The first step, also called matching, can be thought of as a weighted combination of common attributes according to some similarity measures. In this work, we propose such a matching based on multiple attributes of Points of Interests (POI) from the Location-based Social Network Foursquare and the Yelp local directory service. While both contain overlapping attributes that can be use for matching, they have specific strengths and weaknesses which makes their conflation desirable. We present a weighted multi-attribute matching strategy and evaluate its performance. Our strategy can automatically match 97% of randomly selected Yelp POI to their corresponding Foursquare entities. Grant McKenzie, Krzysztof Janowicz, Benjamin Adams |
SIGSPATIAL/GIS | 1 |
| 2013 | A Linked-Data-Driven and Semantically-Enabled Journal Portal for ScientometricsabstractThe Semantic Web journal by IOS Press follows a unique open and transparent process during which each submitted manuscript is available online together with the full history of its successive decision statuses, assigned editors, solicited and voluntary reviewers, their full text reviews, and in many cases also the authors’ response letters. Combined with a highly-customized, Drupal-based journal management system, this provides the journal with semantically rich manuscript time lines and networked data about authors, reviewers, and editors. These data are now exposed using a SPARQL endpoint, an extended Bibo ontology, and a modular Linked Data portal that provides interactive scientometrics based on established and new analysis methods. The portal can be customized for other journals as well. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yingjie Hu 0001, Krzysztof Janowicz, Grant McKenzie, Kunal Sengupta, Pascal Hitzler |
ISWC (2) | 3 |
| 2012 | Frankenplace: An Application for Similarity-Based Place Search
Benjamin Adams, Grant McKenzie |
ICWSM | 2 |