VLDB 2026 Research / reviewers in the wild / expert
Benjamin Adams
dblp:03/740
· DBLP profile ↗
19ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-1657-9809ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 5 first-authorArtificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Dialectal, Longitudinal Corpus of Human-AI Hybrid Language ProductionabstractThis paper presents a multi-dialectal, longitudinal corpus of human-AI hybrid language production, comprising purely human-written texts, purely LLM-generated texts, and hybrid texts produced under different LLM-assistance modes (e.g., stylistic suggestions, short continuations, partial essay generation). The corpus includes 693 participants from five national English dialects, with natural and hybrid samples paired within individuals over a four-week period. This design enables investigation of both short- and longer-term effects of LLM assistance on language use across geographic and social contexts. To illustrate the corpus’s utility, we analyze linguistic features across three dimensions: lexical diversity, syntactic complexity, and stylistic variation. The results show that LLM assistance enhances lexical diversity without a corresponding increase in syntactic complexity, revealing distinct effects across linguistic dimensions. Overall, this corpus offers a valuable resource for studying human-AI interaction, dialectal variation, and the influence of AI assistance on written language. Qiao Gan, Jonathan Dunn, Andrea Nini, Benjamin Adams |
LREC | 4 |
| 2025 | Growing manipulators through feeding material outside-in: Inversion robotsabstractSoft eversion robots have demonstrated significant advantages in navigating within confined spaces with minimal friction, making them promising candidates for various intraluminal applications in medical, industrial, and exploratory domains. While these type of growing robots enable frictionless movement within hollow structures, no existing soft robotic actuation mechanism can grow along the outer surface of an environment without generating friction. A robot with these capabilities could open new possibilities, such as endoscopic vein harvesting for coronary artery bypass graft surgery. This paper introduces a novel growing robotic manipulator based on an outside-in material feeding mechanism - the inversion robot. Unlike conventional eversion robots, which expand by feeding material from the inside out, the inversion robot draws material from the outside to the inside, encapsulating its external environment within an inner sleeve to achieve frictionless movement. We present the design, implementation, and experimental validation of this inversion robot, investigating its growth behavior under varying pressure values and with different diameters, its ability to navigate along defined trajectories, and with a tools mounted to its tip. This inversion robot could enable vein dissection while preserving the surrounding fat layer, making it a promising innovation for minimally invasive vascular surgery and beyond. Xinyi Pi, Junke Yao, Benjamin Adams, Antonia Gerontati, Helge A. Wurdemann |
IROS | 3 |
| 2025 | KEA Explain: Explanations of Hallucinations using Graph Kernel AnalysisabstractLarge Language Models (LLMs) frequently generate hallucinations: statements that are syntactically plausible but lack factual grounding. This research presents KEA (Kernel-Enriched AI) Explain: a neurosymbolic framework that detects and explains such hallucinations by comparing knowledge graphs constructed from LLM outputs with ground truth data from Wikidata or contextual documents. Using graph kernels and semantic clustering, the method provides explanations for detected hallucinations, ensuring both robustness and interpretability. Our framework achieves competitive accuracy in detecting hallucinations across both open- and closed-domain tasks, and is able to generate contrastive explanations, enhancing transparency. This research advances the reliability of LLMs in high-stakes domains and provides a foundation for future work on precision improvements and multi-source knowledge integration. Reilly Haskins, Benjamin Adams |
NeSy | 2 |
| 2024 | Pre-Trained Language Models Represent Some Geographic Populations Better than OthersabstractThis paper measures the skew in how well two families of LLMs represent diverse geographic populations. A spatial probing task is used with geo-referenced corpora to measure the degree to which pre-trained language models from the OPT and BLOOM series represent diverse populations around the world. Results show that these models perform much better for some populations than others. In particular, populations across the US and the UK are represented quite well while those in South and Southeast Asia are poorly represented. Analysis shows that both families of models largely share the same skew across populations. At the same time, this skew cannot be fully explained by sociolinguistic factors, economic factors, or geographic factors. The basic conclusion from this analysis is that pre-trained models do not equally represent the world’s population: there is a strong skew towards specific geographic populations. This finding challenges the idea that a single model can be used for all populations. Jonathan Dunn, Benjamin Adams, Harish Tayyar Madabushi |
LREC/COLING | 2 |
| 2023 | LivePublication: The Science Workflow Creates and Updates the PublicationabstractThe uptake of computational methods to support research has led to some remarkable new tools and methods to improve outcomes. But one unintended consequence is that the scientific record ends up being fragmented and distributed amongst several distinct systems. The research we report aims to gather together all of the components of an experiment into a single container—including the publication itself. We describe the architecture of such a system that marries together distributed workflows (Globus) with research object containers (RO-Crate) and adds new methods to describe, update and 'publish: the details of the workflow and its outcomes. Finally, we demonstrate the system with a natural language processing research use case. Augustus Ellerm, Mark Gahegan, Benjamin Adams |
e-Science | 3 |
| 2022 | Enabling LivePublicationabstractThis paper presents a prototype implementation of LivePublication, a method of publishing which integrates live eScience infrastructure and scientific methods with dynamic natural language research articles. With the maturation of eScience infrastructure, bridging the gap between live computational processes and research articles becomes more tractable and brings long term value to the publication process. The prototype demonstrates the value of integrating computational processes with research articles and provides a first step into LivePublication research and development. Augustus Ellerm, Benjamin Adams, Mark Gahegan, Lukas Trombach |
e-Science | 2 |
| 2020 | Geographically-Balanced Gigaword Corpora for 50 Language VarietiesabstractWhile text corpora have been steadily increasing in overall size, even very large corpora are not designed to represent global population demographics. For example, recent work has shown that existing English gigaword corpora over-represent inner-circle varieties from the US and the UK. To correct implicit geographic and demographic biases, this paper uses country-level population demographics to guide the construction of gigaword web corpora. The resulting corpora explicitly match the ground-truth geographic distribution of each language, thus equally representing language users from around the world. This is important because it ensures that speakers of under-resourced language varieties (i.e., Indian English or Algerian French) are represented, both in the corpora themselves but also in derivative resources like word embeddings. Jonathan Dunn, Benjamin Adams |
LREC | 2 |
| 2019 | Gaze-Guided Narratives: Adapting Audio Guide Content to Gaze in Virtual and Real EnvironmentsabstractExploring a city panorama from a vantage point is a popular tourist activity. Typical audio guides that support this activity are limited by their lack of responsiveness to user behavior and by the difficulty of matching audio descriptions to the panorama. These limitations can inhibit the acquisition of information and negatively affect user experience. This paper proposes Gaze-Guided Narratives as a novel interaction concept that helps tourists find specific features in the panorama (gaze guidance) while adapting the audio content to what has been previously looked at (content adaptation). Results from a controlled study in a virtual environment (n=60) revealed that a system featuring both gaze guidance and content adaptation obtained better user experience, lower cognitive load, and led to better performance in a mapping task compared to a classic audio guide. A second study with tourists situated at a vantage point (n=16) further demonstrated the feasibility of this approach in the real world. Tiffany C. K. Kwok, Peter Kiefer, Victor R. Schinazi, Benjamin Adams, Martin Raubal |
CHI | 4 |
| 2017 | Juxtaposing Thematic Regions Derived from Spatial and Platial User-Generated ContentabstractTypical approaches to defining regions, districts or neighborhoods within a city often focus on place instances of a similar type that are grouped together. For example, most cities have at least one bar district defined as such by the clustering of bars within a few city blocks. In reality, it is not the presence of spatial locations labeled as bars that contribute to a bar region, but rather the popularity of the bars themselves. Following the principle that places, and by extension, place-type regions exist via the people that have given space meaning, we explore user-contributed content as a way of extracting this meaning. Kernel density estimation models of place-based social check-ins are compared to spatially tagged social posts with the goal of identifying thematic regions within the city of Los Angeles, CA. Dynamic human activity patterns, represented as temporal signatures, are included in this analysis to demonstrate how regions change over time. Grant McKenzie, Benjamin Adams |
COSIT | 2 |
| 2017 | Why good data analysts need to be critical synthesists. Determining the role of semantics in data analysis
Simon Scheider, Frank O. Ostermann, Benjamin Adams |
Future Gener. Comput. Syst. | 3 |
| 2017 | A data-synthesis-driven method for detecting and extracting vague cognitive regionsabstractCognitive regions and places are notoriously difficult to represent in geographic information science and systems. The exact delineation of cognitive regions is challenging insofar as borders are vague, membership within the regions varies non-monotonically, and raters cannot be assumed to assess membership consistently and homogeneously. In a study published in this journal in 2014, researchers devised a novel grid-based task in which participants rated the membership of individual cells in a given region and contrasted this approach to a standard boundary-drawing task. Specifically, the authors assessed the vague cognitive regions of Northern California and Southern California. The boundary between these cognitive regions was found to have variable width, and region membership peaked not at the most northern or southern cells but at substantially less extreme latitudes. The authors thus concluded that region membership is about attitude, not just latitude. In the present work, we reproduce this study by approaching it from a computational fourth-paradigm perspective, i.e., by the synthesis of high volumes of heterogeneous data from various sources. We compare the regions which we identify to those from the human-participants study of 2014, identifying differences and commonalities. Our results show a significant positive correlation to those in the original study. Beyond the extracted regions themselves, we compare and contrast the empirical and analytical approaches of these two methods, one a conventional human-participants study and the other an application of increasingly popular data-synthesis-driven research methods in GIScience. Song Gao 0001, Krzysztof Janowicz, Daniel R. Montello, Yingjie Hu 0001, Jiue-An Yang, Grant McKenzie, Yiting Ju, Benjamin Adams, Bo Yan 0003 |
Int. J. Geogr. Inf. Sci. | 9 |
| 2016 | Things and Strings: Improving Place Name Disambiguation from Short Texts by Combining Entity Co-Occurrence with Topic Modeling
Yiting Ju, Benjamin Adams, Krzysztof Janowicz, Yingjie Hu 0001, Bo Yan 0003, Grant McKenzie |
EKAW | 2 |
| 2015 | Frankenplace: Interactive Thematic Mapping for Ad Hoc Exploratory SearchabstractAd hoc keyword search engines built using modern information retrieval methods do a good job of handling fine-grained queries. However, they perform poorly at facilitating spatial and spatially-embedded thematic exploration of the results, despite the fact that many queries, e.g. "civil war," refer to different documents and topics in different places. This is not for lack of data: geographic information, such as place names, events, and coordinates are common in unstructured document collections on the web. The associations between geographic and thematic contents in these documents can provide a rich groundwork to organize information for exploratory research. In this paper we describe the architecture of an interactive thematic map search engine, Frankenplace, designed to facilitate document exploration at the intersection of theme and place. The map interface enables a user to zoom the geographic context of their query in and out, and quickly explore through thousands of search results in a meaningful way. And by combining topic models with geographically contextualized search results, users can discover related topics based on geographic context. Frankenplace utilizes a novel indexing method called geoboost for boosting terms associated with cells on a discrete global grid. The resulting index factors in the geographic scale of the place or feature mentioned in related text, the relative textual scope of the place reference, and the overall importance of the containing document in the document network. The system is currently indexed with over 5 million documents from the web, including the English Wikipedia and online travel blog entries. We demonstrate that Frankenplace can support four distinct types of exploratory search tasks while being adaptive to scale and location of interest. Benjamin Adams, Grant McKenzie, Mark Gahegan |
WWW | 1 |
| 2015 | Thematic signatures for cleansing and enriching place-related linked dataabstractThere has been significant progress transforming semi-structured data about places into knowledge graphs that can be used in a wide variety of geographic information systems such as digital gazetteers or geographic information retrieval systems. For instance, in addition to information about events, actors, and objects, DBpedia contains data about hundreds of thousands of places from Wikipedia and publishes it as Linked Data. Repositories that store data about places are among the most interlinked hubs on the Linked Data cloud. However, most content about places resides in unstructured natural language text, and therefore it is not captured in these knowledge graphs. Instead, place representations are limited to facts such as their population counts, geographic locations, and relations to other entities, for example, headquarters of companies or historical figures. In this paper, we present a novel method to enrich the information stored about places in knowledge graphs using thematic signatures that are derived from unstructured text through the process of topic modeling. As proof of concept, we demonstrate that this enables the automatic categorization of articles into place types defined in the DBpedia ontology (e.g., mountain) and also provides a mechanism to infer relationships between place types that are not captured in existing ontologies. This method can also be used to uncover miscategorized places, which is a common problem arising from the automatic lifting of unstructured and semi-structured data. Benjamin Adams, Krzysztof Janowicz |
Int. J. Geogr. Inf. Sci. | 1 |
| 2013 | Weighted multi-attribute matching of user-generated points of interestabstractTo a large degree, the attraction of Big Data lies in the variety of its heterogeneous multi-thematic and multi-dimensional data sources and not merely its volume. To fully exploit this variety, however, requires conflation. This is a two step process. First, one has to establish identity relations between information entities across the different data sources; and second, attribute values have to be merged according to certain procedures which avoid logical contradictions. The first step, also called matching, can be thought of as a weighted combination of common attributes according to some similarity measures. In this work, we propose such a matching based on multiple attributes of Points of Interests (POI) from the Location-based Social Network Foursquare and the Yelp local directory service. While both contain overlapping attributes that can be use for matching, they have specific strengths and weaknesses which makes their conflation desirable. We present a weighted multi-attribute matching strategy and evaluate its performance. Our strategy can automatically match 97% of randomly selected Yelp POI to their corresponding Foursquare entities. Grant McKenzie, Krzysztof Janowicz, Benjamin Adams |
SIGSPATIAL/GIS | 3 |
| 2012 | On the Geo-Indicativeness of Non-Georeferenced Text
Benjamin Adams, Krzysztof Janowicz |
ICWSM | 1 |
| 2012 | Frankenplace: An Application for Similarity-Based Place Search
Benjamin Adams, Grant McKenzie |
ICWSM | 1 |
| 2011 | Constructing geo-ontologies by reification of observation dataabstractThe semantic integration of heterogeneous, spatiotemporal information is a major challenge for achieving the vision of a multi-thematic and multi-perspective Digital Earth. The Semantic Web technology stack has been proposed to address the integration problem by knowledge representation languages and reasoning. However approaches such as the Web Ontology Languages (OWL) were developed with decidability in mind. They do not integrate well with established modeling paradigms in the geosciences that are dominated by numerical and geometric methods. Additionally, work on the Semantic Web is mostly feature-centric and a field-based view is difficult to integrate. A layer specifying the transition from observation data to classes and relations is missing. In this work we combine OWL with geometric and topological language constructs based on similarity spaces. Our approach provides three main benefits. First, class constructors can be built from a larger palette of mathematical operations based on vector algebra. Second, it affords the representation of prototype-based classes. Third, it facilitates the representation of classes derived from machine learning classifiers that utilize a multi-dimensional feature space. Instead of following a one-size-fits-all approach, our work allows one to derive contextualized OWL ontologies by reification of observation data. Benjamin Adams, Krzysztof Janowicz |
GIS | 1 |
| 2009 | A Metric Conceptual Space Algebra
Benjamin Adams, Martin Raubal |
COSIT | 1 |