EDBT 2026 Demo / reviewers in the wild / expert
Guohui Xiao 0001
dblp:55/3586
· DBLP profile ↗
30ranked-venue papers in the field
3as first author
4since 2021 · last 2025
0000-0002-5115-4769ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 17 (3 first)Information Retrieval & Web Search · 4Big Data, Cloud & Distributed Data Systems · 4Database Systems & Data Management · 3Other / Interdisciplinary · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A rapid cross-validation computing for three-way decisions in imbalanced data
Zhenzhen Gu, Guohui Xiao 0001 |
Inf. Sci. | 4 |
| 2021 | SUMA: A Partial Materialization-Based Scalable Query Answering in OWL 2 DLabstractAbstract Ontology-mediated querying (OMQ) provides a paradigm for query answering according to which users not only query records at the database but also query implicit information inferred from ontology. A key challenge in OMQ is that the implicit information may be infinite, which cannot be stored at the database and queried by off -the -shelf query engine. The commonly adopted technique to deal with infinite entailments is query rewriting, which, however, comes at the cost of query rewriting at runtime. In this work, the partial materialization method is proposed to ensure that the extension is always finite. The partial materialization technology does not rewrite query but instead computes partial consequences entailed by ontology before the online query. Besides, a query analysis algorithm is designed to ensure the completeness of querying rooted and Boolean conjunctive queries over partial materialization. We also soundly and incompletely expand our method to support highly expressive ontology language, OWL 2 DL. Finally, we further optimize the materialization efficiency by role rewriting algorithm and implement our approach as a prototype system SUMA by integrating off-the-shelf efficient SPARQL query engine. The experiments show that SUMA is complete on each test ontology and each test query, which is the same as Pellet and outperforms PAGOdA. Besides, SUMA is highly scalable on large datasets. Xiaowang Zhang, Muhammad Qasim Yasin, Zhiyong Feng 0002, Guohui Xiao 0001 |
Data Sci. Eng. | 6 |
| 2021 | Consistency assessment for open geodata integration: an ontology-based approach
Linfang Ding, Guohui Xiao 0001, Diego Calvanese, Liqiu Meng |
GeoInformatica | 2 |
| 2021 | Towards the next generation of the LinkedGeoData project using virtual knowledge graphsabstractWith the advancement of Semantic Technologies, large geospatial data sources have been increasingly published as Linked data on the Web. The LinkedGeoData project is one of the most prominent such projects to create a large knowledge graph from OpenStreetMap (OSM) with global coverage and interlinking of other data sources. In this paper, we report on the ongoing effort of exposing the relational database in LinkedGeoData as a SPARQL endpoint using Virtual Knowledge Graph (VKG) technology. Specifically, we present two realizations of VKGs, using the two systems Sparqlify and Ontop. In order to improve compliance with the OGC GeoSPARQL standard, we have implemented GeoSPARQL support in Ontop v4. Moreover, we have evaluated the VKG-powered LinkedGeoData in the test areas of Italy and Germany. Our experiments demonstrate that such system supports complex GeoSPARQL queries, which confirms that query answering in the VKG approach is efficient. Linfang Ding, Guohui Xiao 0001, Albulen Pano, Claus Stadler, Diego Calvanese |
J. Web Semant. | 2 |
| 2020 | A Partial Materialization-Based Approach to Scalable Query Answering in OWL 2 DL
Xiaowang Zhang, Muhammad Qasim Yasin, Zhiyong Feng 0002, Guohui Xiao 0001 |
DASFAA (3) | 6 |
| 2020 | Semantic Integration of Bosch Manufacturing Data Using Virtual Knowledge Graphs
Elem Guzel Kalayci, Irlán Grangel-González, Felix Lösch, Guohui Xiao 0001, Anees Mehdi, Evgeny Kharlamov, Diego Calvanese |
ISWC (2) | 4 |
| 2020 | The Virtual Knowledge Graph System Ontop
Guohui Xiao 0001, Davide Lanti, Roman Kontchakov, Sarah Komla-Ebri, Elem Guzel Kalayci, Linfang Ding, Julien Corman, Benjamin Cogrel, Diego Calvanese, Elena Botoeva |
ISWC (2) | 1 |
| 2019 | Optimizing Horn- SHIQ Reasoning for OBDA
Labinot Bajraktari, Magdalena Ortiz 0001, Guohui Xiao 0001 |
ISWC (1) | 3 |
| 2019 | Ontop-spatial: Ontop of geospatial databases
Konstantina Bereta, Guohui Xiao 0001, Manolis Koubarakis |
J. Web Semant. | 2 |
| 2019 | Semantically-enhanced rule-based diagnostics for industrial Internet of Things: The SDRL language and case study for Siemens trains and turbines
Evgeny Kharlamov, Gulnar Mehdi, Ognjen Savkovic, Guohui Xiao 0001, Elem Guzel Kalayci, Mikhail Roshchin |
J. Web Semant. | 4 |
| 2018 | Towards Simplification of Analytical Workflows With Semantics at Siemens (Extended Abstract)abstractAnalytical workflows are heavily used in large and data intensive companies. An important application of such workflows in Siemens is equipment analytics when equipment KPIs and reports are computed by aggregating equipment's operational, master, and analytical data. In Siemens this data satisfies big data dimensions and this dependence poses significant challenges in authoring, reuse, and maintenance of analytical workflows by engineers and data scientists. In this work we propose to address these problems by relying on semantic technologies: we use ontologies to give a high level representation of equipment's operational and master data and offer a high level language to express KPIs over ontologies. We implemented our approach, integrated it with KNIME, and evaluated at Siemens. This is a preliminary work and we are excited about its further extensions. Evgeny Kharlamov, Gulnar Mehdi, Ognjen Savkovic, Guohui Xiao 0001, Steffen Lamparter, Ian Horrocks 0001, Arild Waaler |
IEEE BigData | 4 |
| 2018 | Finding Data Should be Easier than Finding OilabstractThe competitiveness of modern enterprises heavily depends on their ability to make the right business decisions by relying on efficient and timely analysis of the right business critical data. In large and data intensive companies such as Equinor, a Norwegian multinational oil and gas company with more than 20,000 employees, gathering such data is not a trivial task due to the growing size and complexity of corporate information sources. As a result, the data gathering task is often the most time-consuming part of the decision making process, in particular when it comes to the work processes of Equinor’s exploration geologists that should find in a timely manner new exploitable accumulations of oil or gas in given areas by analysing data about these areas. In this work we present our experience in addressing this data challenge tast at Equinor. We have developed and deployed at Equinor a semantic data access system that relies on the Ontology Based Data Access (OBDA) approach. Our system is based on our solid theoretical contributions and has been extensively evaluated at Equinor. Evgeny Kharlamov, Martin G. Skjæveland, Dag Hovland, Theofilos P. Mailis, Ernesto Jiménez-Ruiz, Guohui Xiao 0001, Ahmet Soylu, Ian Horrocks 0001, Arild Waaler |
IEEE BigData | 6 |
| 2018 | BigSR: real-time expressive RDF stream reasoning on modern Big Data platformsabstractShifting from Big Data to Big Knowledge requires systems that are able to cope with the large volume and high-velocity dimensions in a scalable and inference-enabled manner. In this work, we are focusing on stream processing and reasoning using the graph-based RDF data model. We are aiming to explore the ability of modern distributed computing frameworks to process highly expressive knowledge inference queries over Big Data streams. To do so, we consider queries expressed as a positive fragment of a temporal logic framework based on Answer Set Programming and propose solutions to process such queries, based on the two main execution models adopted by major parallel and distributed execution frameworks: Bulk Synchronous Parallel (BSP) and Recordat-A-Time (RAT). We implement our solution named BigSR and conduct a series of experiments with 15 queries from 4 different datasets. Our experiments show that BigSR achieves high throughput beyond million-triples per second using a rather small cluster of machines. Xiangnan Ren, Olivier Curé, Hubert Naacke, Guohui Xiao 0001 |
IEEE BigData | 4 |
| 2018 | Semantic Technologies for Data Access and IntegrationabstractRecently, semantic technologies have been successfully deployed to overcome the typical difficulties in accessing and integrating data stored in different kinds of legacy sources. In particular, knowledge graphs are being used as a mechanism to provide a uniform representation of heterogeneous information. Such graphs represent data in the RDF format, which is complemented by an ontology and can be queried using the standard SPARQL language. The RDF graph is often obtained by materializing source data, following the traditional extract-transform-load workflow. Alternatively, the sources are declaratively mapped to the ontology, and the RDF graph is maintained virtual. In such an approach, usually called ontology-based data access/integration (OBDA/I), query answering is based on sophisticated query transformation techniques. In this tutorial: (i) we provide a general introduction to semantic technologies; (ii) we illustrate the principles underlying OBDA/I, providing insights into its theoretical foundations, and describing well-established algorithms, techniques, and tools; (iii) we discuss relevant use-cases for OBDA/I; (iv) we provide an overview on some recent advancements. Diego Calvanese, Guohui Xiao 0001 |
CIKM | 2 |
| 2018 | Ontop-temporal: A Tool for Ontology-based Query Answering over Temporal DataabstractWe present Ontop-temporal, an extension of the ontology-based data access system Ontop for query answering with temporal data and ontologies. Ontop is a system to answer SPARQL queries over various data stores, using standard R2RML mappings and an OWL2QL domain ontology to produce high-level conceptual views over the raw data. The Ontop-temporal extension is designed to handle timestamped log data, by additionally using (i) mappings supporting validity time specification, and (ii) rules based on metric temporal logic to define temporalised concepts. In this demo we present how Ontop-temporal can be used to facilitate the access to the MIMIC-III critical care unit dataset containing log data on hospital admissions, procedures, and diagnoses. We use the ICD9CM diagnoses ontology and temporal rules formalising the selection of patients for clinical trials taken from the clinicaltrials.gov database. We demonstrate how high-level queries can be answered by Ontop-temporal to identify patients eligible for the trials. Elem Guzel Kalayci, Guohui Xiao 0001, Vladislav Ryzhikov, Tahir Emre Kalayci, Diego Calvanese |
CIKM | 2 |
| 2018 | Efficient Ontology-Based Data Integration with Canonical IRIs
Guohui Xiao 0001, Dag Hovland, Dimitris Bilidas, Martín Rezk, Martin Giese, Diego Calvanese |
ESWC | 1 |
| 2018 | Expressivity and Complexity of MongoDB QueriesabstractA significant number of novel database architectures and data models have been proposed during the last decade. While some of these new systems have gained in popularity, they lack a proper formalization, and a precise understanding of the expressivity and the computational properties of the associated query languages. In this paper, we aim at filling this gap, and we do so by considering MongoDB, a widely adopted document database managing complex (tree structured) values represented in a JSON-based data model, equipped with a powerful query mechanism. We provide a formalization of the MongoDB data model, and of a core fragment, called MQuery, of the MongoDB query language. We study the expressivity of MQuery, showing its equivalence with nested relational algebra. We further investigate the computational complexity of significant fragments of it, obtaining several (tight) bounds in combined complexity, which range from LOGSPACE to alternating exponential-time with a polynomial number of alternations. As a consequence, we obtain also a characterization of the combined complexity of nested relational algebra query evaluation. Elena Botoeva, Diego Calvanese, Benjamin Cogrel, Guohui Xiao 0001 |
ICDT | 4 |
| 2018 | Efficient Handling of SPARQL OPTIONAL for OBDA
Guohui Xiao 0001, Roman Kontchakov, Benjamin Cogrel, Diego Calvanese, Elena Botoeva |
ISWC (1) | 1 |
| 2017 | Semantic Rules for Machine Diagnostics: Execution and ManagementabstractRule-based diagnostics of equipment is an important task in industry. In this paper we present how semantic technologies can enhance diagnostics. In particular, we present our semantic rule language sigRL that is inspired by the real diagnostic languages used in Siemens. SigRL allows to write compact yet powerful diagnostic programs by relying on a high level data independent vocabulary, diagnostic ontologies, and queries over these ontologies. We study computational complexity of SigRL: execution of diagnostic programs, provenance computation, as well as automatic verification of redundancy and inconsistency in diagnostic programs. Evgeny Kharlamov, Ognjen Savkovic, Guohui Xiao 0001, Rafael Peñaloza, Gulnar Mehdi, Mikhail Roshchin, Ian Horrocks 0001 |
CIKM | 3 |
| 2017 | SemDia: Semantic Rule-Based Equipment Diagnostics ToolabstractRule-based diagnostics of power generating equipment is an important task in industry. In this demo we present how semantic technologies can enhance diagnostics. In particular, we present our semantic rule language sigRL that is inspired by the real diagnostic languages in Siemens. SigRL allows to write compact yet powerful diagnostic programs by relying on a high level data independent vocabulary, diagnostic ontologies, and queries over these ontologies. We present our diagnostic system SemDia. The attendees will be able to write diagnostic programs in SemDia using sigRL over 50 Siemens turbines. We also present how such programs can be automatically verified for redundancy and inconsistency. Moreover, the attendees will see the provenance service that SemDia provides to trace the origin of diagnostic results. Gulnar Mehdi, Evgeny Kharlamov, Ognjen Savkovic, Guohui Xiao 0001, Elem Guzel Kalayci, Sebastian Brandt 0001, Ian Horrocks 0001, Mikhail Roshchin, Thomas A. Runkler |
CIKM | 4 |
| 2017 | Cost-Driven Ontology-Based Data Access
Davide Lanti, Guohui Xiao 0001, Diego Calvanese |
ISWC (1) | 2 |
| 2017 | Semantic Rule-Based Equipment Diagnostics
Gulnar Mehdi, Evgeny Kharlamov, Ognjen Savkovic, Guohui Xiao 0001, Elem Guzel Kalayci, Sebastian Brandt 0001, Ian Horrocks 0001, Mikhail Roshchin, Thomas A. Runkler |
ISWC (2) | 4 |
| 2017 | Ontology Based Data Access in Statoil
Evgeny Kharlamov, Dag Hovland, Martin G. Skjæveland, Dimitris Bilidas, Ernesto Jiménez-Ruiz, Guohui Xiao 0001, Ahmet Soylu, Davide Lanti, Martín Rezk, Dmitriy Zheleznyakov, Martin Giese, Hallstein Lie, Yannis E. Ioannidis, Yannis Kotidis, Manolis Koubarakis, Arild Waaler |
J. Web Semant. | 6 |
| 2016 | A semantic approach to polystoresabstractIn the database community Polystores is an emerging and promising approach for data federation that aims at designing a unified querying layer over multiple data models. In the Semantic Web community a similar in spirit approach of Ontology-Based Data Access (OBDA) has been recently proposed, attracted a lot of attention, and proved its success in several industrial scenarios. In this paper we discuss a semantic approach to building polystores using the OBDA paradigm. We also present our system Optique that is utilized in an industrial application of performing turbine diagnostics in Siemens. Evgeny Kharlamov, Theofilos P. Mailis, Konstantina Bereta, Dimitris Bilidas, Sebastian Brandt 0001, Ernesto Jiménez-Ruiz, Steffen Lamparter, Christian Neuenstadt, Özgür L. Özçep, Ahmet Soylu, Christoforos Svingos, Guohui Xiao 0001, Dmitriy Zheleznyakov, Diego Calvanese, Ian Horrocks 0001, Martin Giese, Yannis E. Ioannidis, Yannis Kotidis, Ralf Möller 0001, Arild Waaler |
IEEE BigData | 12 |
| 2016 | Ontology-Based Data Access for Maritime Security
Stefan Brüggemann, Konstantina Bereta, Guohui Xiao 0001, Manolis Koubarakis |
ESWC | 3 |
| 2015 | The NPD Benchmark: Reality Check for OBDA SystemsabstractIn the last decades we moved from a world in which an enterprise had one central database---rather small for todays' standards---to a world in which many different---and big---databases must interact and operate, providing the user an integrated and understandable view of the data. Ontology-Based Data Access (OBDA) is becoming a popular approach to cope with this new scenario. OBDA separates the user from the data sources by means of a conceptual view of the data (ontology) that provides clients with a convenient query vocabulary. The ontology is connected to the data sources through a declarative specification given in terms of mappings. Although prototype OBDA systems providing the ability to answer SPARQL queries over the ontology are available, a significant challenge remains when it comes to use these systems in industrial environments: performance. To properly evaluate OBDA systems, benchmarks tailored towards the requirements in this setting are needed. In this work, we propose a novel benchmark for OBDA systems based on real data coming from the oil industry: the Norwegian Petroleum Directorate (NPD) FactPages. Our benchmark comes with novel techniques to generate, from the NPD data, datasets of increasing size, taking into account the requirements dictated by the OBDA setting. We validate our benchmark on significant OBDA systems, showing that it is more adequate than previous benchmarks not tailored for OBDA. Davide Lanti, Martín Rezk, Guohui Xiao 0001, Diego Calvanese |
EDBT | 3 |
| 2015 | Ontology Based Access to Exploration Data at Statoil
Evgeny Kharlamov, Dag Hovland, Ernesto Jiménez-Ruiz, Davide Lanti, Hallstein Lie, Christoph Pinkel, Martín Rezk, Martin G. Skjæveland, Evgenij Thorstensen, Guohui Xiao 0001, Dmitriy Zheleznyakov, Ian Horrocks 0001 |
ISWC (2) | 10 |
| 2014 | Answering SPARQL Queries over Databases under OWL 2 QL Entailment Regime
Roman Kontchakov, Martín Rezk, Mariano Rodriguez-Muro, Guohui Xiao 0001, Michael Zakharyaschev |
ISWC (1) | 4 |
| 2009 | A Tableau Algorithm for Handling Inconsistency in OWL
Xiaowang Zhang, Guohui Xiao 0001, Zuoquan Lin |
ESWC | 2 |
| 2009 | An Anytime Algorithm for Computing Inconsistency Measurement
Yue Ma 0009, Guilin Qi, Guohui Xiao 0001, Pascal Hitzler, Zuoquan Lin |
KSEM | 3 |