VLDB 2026 Research / reviewers in the wild / expert
Shirly Stephen
dblp:168/7650
· DBLP profile ↗
10ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-3547-8058ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ContaminOSO: Ontological Foundations and Design Choices for an Ontology for Environmental Contamination DataabstractContamination by heavy metals, per- and polyfluoroalkyl substances (PFAS), and other emerging pollutants poses serious risks to environmental and human health. Effective monitoring and tracing require integrating data from diverse sources. A knowledge graph approach enables semantic integration, but relies on an ontology that supports intuitive and robust querying and reasoning. To address this, we present the Contaminant Observations and Samples Ontology (ContaminOSO), a framework for semantically enriching environmental contaminant data. Built on SOSA and QUDT ontologies, ContaminOSO introduces key extensions to meet contamination-specific needs and real-world data challenges. This paper highlights four of its core design solutions: (1) extending SOSA to model multiple features of interest; (2) using QUDT to standardize the representation of contaminants and observed properties; (3) developing a detailed and nuanced pattern for measurement result representation using QUDT and STAD; and (4) adopting a pragmatic approach for connecting to existing taxonomies from the OBO Foundry, such as the NCBI organismal classification and relevant subsets of the Food Ontology (FoodOn), for classifying samples. Torsten Hahmann, Katrina Schweikert, Shirly Stephen, David K. Kedrowski |
FOIS | 3 |
| 2025 | The KnowWhereGraph ontologyabstractKnowWhereGraph is one of the largest fully publicly available geospatial knowledge graphs. It includes data from 30 layers on natural hazards (e.g., hurricanes, wildfires), climate variables (e.g., air temperature, precipitation), soil properties, crop and land-cover types, demographics, and human health, various place and region identifiers, among other themes. These have been leveraged through the graph by a variety of applications to address challenges in food security and agricultural supply chains; sustainability related to soil conservation practices and farm labor; and delivery of emergency humanitarian aid following a disaster. In this paper, we introduce the ontology that acts as the schema for KnowWhereGraph. This broad overview provides insight into the requirements and design specifications for the graph and its schema, including the development methodology (modular ontology modeling) and the resources utilized to implement, materialize, and deploy KnowWhereGraph with its end-user interfaces and public query SPARQL endpoint. Cogan Shimizu, Shirly Stephen, Adrita Barua, Ling Cai 0002, Antrea Christou, Kitty Currier, Abhilekha Dalal, Colby K. Fisher, Pascal Hitzler, Krzysztof Janowicz, Wenwen Li 0002, Zilong Liu 0003, Mohammad Saeid Mahdavinejad, Gengchen Mai, Dean Rehberger, Mark Schildhauer, Meilin Shi, Sanaz Saki Norouzi, Yuanyuan Tian 0002, Joseph Zalewski, Lu Zhou 0005, Rui Zhu 0008 |
J. Web Semant. | 2 |
| 2022 | Knowledge explorer: exploring the 12-billion-statement KnowWhereGraph using faceted search (demo paper)abstractKnowledge graphs are a rapidly growing paradigm and technology stack for integrating large-scale, heterogeneous data in an AI-ready form, i.e., combining data with the formal semantics required to understand it. However, toolchains that support data synthesis and knowledge discovery through information organization, search, filtering, and visualization have been developed at a pace lagging knowledge graph technology. In this paper, we present Knowledge Explorer, an open-source faceted search interface that provides environmentally intelligent services for interactively browsing and navigating KnowWhereGraph. Currently one of the largest open knowledge graphs, KnowWhereGraph contains over 12 billion statements with rich spatial and temporal information from more than 30 data layers. With an extensive collection of facets, Knowledge Explorer enables spatial, temporal, full-text, and expert search with dereferencing functionality to support "follow-your-nose"exploration, and it allows users to narrow their search by selecting facets. Given the size of the underlying graph and dependency on GeoSPARQL, we have improved query performance by implementing Elasticsearch indexing, spatial query generation, and caching. Knowledge Explorer is capable of retrieving information within seconds, answering a wide variety of competency questions posed by researchers, humanitarian relief organizations, and the broader public, thus helping better perform tasks such as cross-gazetteer place retrieval and disaster assessment from global to local geographic scales. Zilong Liu 0003, Zhining Gu, Thomas Thelen, Seila Gonzalez Estrecha, Rui Zhu 0008, Colby K. Fisher, Anthony D'Onofrio, Cogan Shimizu, Krzysztof Janowicz, Mark Schildhauer, Shirly Stephen, Dean Rehberger, Wenwen Li 0002, Pascal Hitzler |
SIGSPATIAL/GIS | 11 |
| 2021 | Providing Humanitarian Relief Support through Knowledge GraphsabstractDisasters are often unpredictable and complex events, requiring humanitarian organizations to understand and respond to many different issues simultaneously and immediately. Often the biggest challenge to improving the effectiveness of the response is quickly finding the right expert, with the right expertise concerning a specific disaster type/disaster and geographic region. To assist in achieving such a goal, this paper demonstrates a knowledge graph-based search engine developed on top of an expert knowledge graph. It accommodates three modes of information retrieval, including a follow-your-nose search, an expert similarity search, and a SPARQL query interface. We will demonstrate utilizing the system to rapidly navigate from a hazard event to a specific expert who may be helpful, for example. More importantly, as the data is fully integrated including links between hazards and their abstract topics, we can find experts who have relevant expertise while navigating the graph. Rui Zhu 0008, Ling Cai 0002, Gengchen Mai, Cogan Shimizu, Colby K. Fisher, Krzysztof Janowicz, Anna Lopez-Carr, Andrew Schroeder, Mark Schildhauer, Yuanyuan Tian 0002, Shirly Stephen, Zilong Liu 0003 |
K-CAP | 11 |
| 2020 | Model-Finding for Externally Verifying FOL Ontologies: A Study of Spatial OntologiesabstractUse and reuse of an ontology requires prior ontology verification which encompasses, at least, proving that the ontology is internally consistent and consistent with representative datasets. First-order logic (FOL) model finders are among the only available tools to aid us in this undertaking, but proving consistency of FOL ontologies is theoretically intractable while also rarely succeeding in practice, with FOL model finders scaling even worse than FOL theorem provers. This issue is further exacerbated when verifying FOL ontologies against datasets, which requires constructing models with larger domain sizes. This paper presents a first systematic study of the general feasibility of SAT-based model finding with FOL ontologies. We use select spatial ontologies and carefully controlled synthetic datasets to identify key measures that determine the size and difficulty of the resulting SAT problems. We experimentally show that these measures are closely correlated with the runtimes of Vampire and Paradox, two state-of-the-art model finders. We propose a definition elimination technique and demonstrate that it can be a highly effective measure for reducing the problem size and improving the runtime and scalability of model finding. Shirly Stephen, Torsten Hahmann |
FOIS | 1 |
| 2019 | Identifying Bottlenecks in Practical SAT-Based Model Finding for First-Order Logic Ontologies with DatasetsabstractSatisfiability of first-order logic (FOL) ontologies is typically verified by translation to propositional satisfiability (SAT) problems, which is then tackled by a SAT solver. Unfortunately, SAT solvers often experience scalability issues when reasoning with FOL ontologies and even moderately sized datasets. While SAT solvers have been found to capably handle complex axiomatizations, finding models of datasets gets considerably more complex and time-intensive as the number of clause exponentially increases with increase in individuals and axiomatic complexity. We identify FOL definitions as a specific bottleneck and demonstrate via experiments that the presence of many defined terms of the highest arity significantly slows down model finding. We also show that removing optional definitions and substituting these terms by their definiens leads to a reduction in the number of clauses, which makes SAT-based model finding practical for over 100 individuals in a FOL theory. Shirly Stephen, Torsten Hahmann |
AAAI | 1 |
| 2019 | Formal Qualitative Spatial Augmentation of the Simple Feature Access ModelabstractThe need to share and integrate heterogeneous geospatial data has resulted in the development of geospatial data standards such as the OGC/ISO standard Simple Feature Access (SFA), that standardize operations and simple topological and mereotopological relations over various geometric features such as points, line segments, polylines, polygons, and polyhedral surfaces. While SFA’s supplied relations enable qualitative querying over the geometric features, the relations' semantics are not formalized. This lack of formalization prevents further automated reasoning - apart from simple querying - with the geometric data, either in isolation or in conjunction with external purely qualitative information as one might extract from textual sources, such as social media. To enable joint qualitative reasoning over geometric and qualitative spatial information, this work formalizes the semantics of SFA’s geometric features and mereotopological relations by defining or restricting them in terms of the spatial entity types and relations provided by CODIB, a first-order logical theory from an existing logical formalization of multidimensional qualitative space. Shirly Stephen, Torsten Hahmann |
COSIT | 1 |
| 2018 | Using a hydro-reference ontology to provide improved computer-interpretable semantics for the groundwater markup language (GWML2)abstractComprehensive water data management requires semantically integrating various data models and ontologies that represent hydrologic knowledge. But integration is hampered by nuances in the use of water-related vocabulary (e.g. terms such as water body, aquifer, reservoir, well, etc.) across water representations and by the reliance on a mix of formal and informal specifications of how these terms are interpreted in each representation. Reconciliation of only partially formal encodings of the semantics of water representations requires manual inspection using tools from ontological analysis. This paper investigates as to what extent a domain reference ontology that is fully formalized in first-order logic can guide the ontological analysis.In particular, it is studied as to what extent the Hydro Foundational Ontology (HyFO), which encodes the semantics of a small set of unifying water concepts and associated relations in first-order logic, can serve as a reference ontology for the water domain to steer the ontological analysis of individual water representations, and to formalize their semantics more fully. This is specifically tested on the Groundwater Markup Language (GWML2). The result is GWML2-FOL, a concise logical description of GWML2’s key terms as a logical extension of HyFO. GWML2-FOL is structured into three layers of terms (mostly classes) of increasing specificity. The top layer consists of terms shareable across the earth and physical sciences, an intermediate layer includes HyFO’s hydro terms that span surface and subsurface water storage, and the bottom layer encapsulates groundwater specific GWML2 terms. The analysis and stratification uncover semantic ambiguities in GWML2 and suggest terminological and semantic clarifications and modifications in preparation for integrating GWML2 with other semantic water representations.The analysis also identifies two necessary additions to the HyFO: the concept of a hydro rock body as a hybrid of water and solid matter, which generalizes key groundwater terms such as aquifers or wells, and the concept of dependent hydrologic features such as springs, water tables, or divides. More broadly, differences between domain ontologies and a domain-reference ontology and their respective complementary roles in semantic-enabled geosciences are outlined. Torsten Hahmann, Shirly Stephen |
Int. J. Geogr. Inf. Sci. | 2 |
| 2017 | An Ontological Framework for Characterizing Hydrological Flow ProcessesabstractThe spatio-temporal processes that describe hydrologic flow - the movement of water above and below the surface of the Earth -- are currently underrepresented in formal semantic representations of the water domain. This paper analyses basic flow processes in the hydrology domain and systematically studies the hydrogeological entities, such as different rock and water bodies, the ground surface or subsurface zones, that participate in them. It identifies the source and goal entities and the transported water (the theme) as common participants in hydrologic flow and constructs a taxonomy of different flow patterns based on differences in source and goal participants. The taxonomy and related concepts are axiomatized in first-order logic as refinements of DOLCE's participation relation and reusing hydrogeological concepts from the Hydro Foundational Ontology (HyFO). The formalization further enhances HyFO and contributes to improved knowledge integration in the hydrology domain. Shirly Stephen, Torsten Hahmann |
COSIT | 1 |
| 2015 | Swiss Canton Regions: A Model for Complex Objects in Geographic Partitions
Matthew P. Dube, Max J. Egenhofer, Joshua A. Lewis, Shirly Stephen, Mark A. Plummer |
COSIT | 4 |