Evgeny Kharlamov

dblp:20/4833 · DBLP profile ↗
← Back
84ranked-venue papers in the field
17as first author
35since 2021 · last 2026
0000-0003-3247-4166ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 36 (9 first)Information Retrieval & Web Search · 21 (2 first)Database Systems & Data Management · 14 (2 first)Data Mining & Knowledge Discovery · 7Big Data, Cloud & Distributed Data Systems · 6 (4 first)
YearPublicationVenuePosition
2026 Integrating Meta-features with Knowledge Graph Embeddings for Meta-learning
Antonis Klironomos, Ioannis Dasoulas, Francesco Periti, Mohamed H. Gad-Elrab, Heiko Paulheim, Anastasia Dimou, Evgeny Kharlamov
ESWC (1)7
2026 DiKGRec: Generative Recommender Model with Diffusion and Knowledge Graph-Based Reasoning
abstract
Generative AI has shown remarkable advancements across various tasks, including recommender systems, where recent research leverages generative approaches to provide personalised recommendations based on user-item historical interaction data. However, the inherent sparsity of the interaction data poses a significant challenge to the advancement of generative recommender models. While some discriminative models have explored incorporating knowledge graphs (KGs) to address this issue, they often struggle with noise sensitivity, lack of explainability, and difficulties in handling cold-start scenarios, where new items with little or no historical user interaction data are involved. In this paper, we propose a novel dual-architecture generative model that intuitively integrates a diffusion model with KG-based reasoning, which reflects the propagation of user preference in a KG towards items. Our approach not only improves recommendation accuracy significantly, but also introduces explainability by leveraging the structured insights from KGs. Furthermore, the KG-based reasoning enables our model to effectively address cold-start scenarios. By utilising the semantic connections in the KG, our model can recommend these new items with confidence, overcoming a common limitation of traditional methods. We evaluate our model on three benchmark datasets, demonstrating superior performance (beat SOTA by over 10% in recall@20 in average).
Zhuoxun Zheng, Baifan Zhou, Ahmet Soylu, Jie Tang 0001, Evgeny Kharlamov
KDD (1)5
2026 Caddie: A prototype of content-based ad hoc RDF dataset retrieval
abstract
The rapid growth of open and structured RDF data on the Web has promoted the development of dataset search as an important research topic. The core function of existing systems is ad hoc dataset retrieval (AHDR) based on the metadata of datasets, which contains limited information and often suffers from quality issues. To overcome the limitations, in this article, we systematically investigate content-based AHDR to exploit the actual RDF data in datasets. We address three main tasks of content-based AHDR with novel methods for handling the large size and complex structure of RDF data to facilitate dataset retrieval, deduplication, and snippet extraction. These methods are integrated into an online and open-source prototype called Caddie . The effectiveness and practicability of its components are evaluated on a public test collection and by a user study.
Xiaxia Wang 0001, Qiaosheng Chen, Weiqing Luo, Jeff Z. Pan, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
J. Web Semant.7
2025 Graph Constraint Language for Industrial Knowledge Graphs and Machine Learning
Zhuoxun Zheng, Ognjen Savkovic, Baifan Zhou, Antonis Klironomos, Evgeny Kharlamov, Ahmet Soylu
DaWaK5
2025 Towards Hybrid Graphs: Unifying Property Graphs and Time Series
Mouna Ammar, Christopher Rost, Riccardo Tommasini 0001, Shubhangi Agarwal 0001, Angela Bonifati, Petra Selmer, Evgeny Kharlamov, Erhard Rahm
EDBT7
2025 ReaLitE: Enrichment of Relation Embeddings in Knowledge Graphs Using Numeric Literals
Antonis Klironomos, Baifan Zhou, Zhuoxun Zheng, Mohamed H. Gad-Elrab, Heiko Paulheim, Evgeny Kharlamov
ESWC (1)6
2025 LLM-Supported Mapping Generation for Semantic Manufacturing Treasure Hunting
Wilma Johanna Schmidt, Irlán Grangel-González, Tobias Huschle, Lena Wagner, Evgeny Kharlamov, Adrian Paschke
ESWC (2)5
2025 GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model
abstract
Anomaly detection on text-rich graphs is widely prevalent in real life, such as detecting incorrectly assigned academic papers to authors and detecting bots in social networks. The remarkable capabilities of large language models (LLMs) pave a new revenue by utilizing rich-text information for effective anomaly detection. However, simply introducing rich texts into LLMs can obscure essential detection cues and introduce high fine-tuning costs. Moreover, LLMs often overlook the intrinsic structural bias of graphs which is vital for distinguishing normal from abnormal node patterns. To this end, this paper introduces GuARD, a text-rich and graph-informed language model that combines key structural features from graph-based methods with fine-grained semantic attributes extracted via small language models for effective anomaly detection on text-rich graphs. GuARD is optimized with the progressive multimodal multi-turn instruction tuning framework in the task-guided instruction tuning regime tailed to incorporate both rich-text and structural modalities. Extensive experiments on four datasets reveal that GuARD outperforms graph-based and LLM-based anomaly detection methods, while offering up to 5× speedup in training and 10× speedup in inference over vanilla long-context LLMs on the large-scale WhoIsWho dataset.
Yunhe Pang 0001, Bo Chen 0026, Fanjin Zhang, Yanghui Rao, Evgeny Kharlamov, Jie Tang 0001
KDD (2)5
2025 ExeKGLib: A Platform for Machine Learning Analytics Based on Knowledge Graphs
Antonis Klironomos, Baifan Zhou, Zhuoxun Zheng, Mohamed H. Gad-Elrab, Heiko Paulheim, Evgeny Kharlamov
ISWC (2)7
2025 DAGE: DAG Query Answering via Relational Combinator with Logical Constraints
abstract
Predicting answers to queries over knowledge graphs is called a complex reasoning task because answering a query requires subdividing it into subqueries. Existing query embedding methods use this decomposition to compute the embedding of a query as the combination of the embedding of the subqueries. This requirement limits the answerable queries to queries having a single free variable and being decomposable, which are called tree-form queries and correspond to the SROI- description logic. In this paper, we define a more general set of queries, called DAG queries and formulated in the ALCOIR description logic, propose a query embedding method for them, called DAGE, and a new benchmark to evaluate query embeddings on them. Given the computational graph of a DAG query, DAGE combines the possibly multiple paths between two nodes into a single path with a trainable operator that represents the intersection of relations and learns DAG-DL concepts from tautologies. We implement DAGE on top of existing query embedding methods, and we empirically measure the improvement of our method over the results of vanilla methods evaluated in tree-form queries that approximate the DAG queries of our proposed benchmark.
Yunjie He, Bo Xiong 0001, Daniel Hernández 0002, Yuqicheng Zhu, Evgeny Kharlamov, Steffen Staab
WWW5
2024 Low-Dimensional Hyperbolic Knowledge Graph Embedding for Better Extrapolation to Under-Represented Data
Zhuoxun Zheng, Baifan Zhou, Arild Waaler, Evgeny Kharlamov, Ahmet Soylu
ESWC (1)6
2024 Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed Graphs
abstract
The text-attributed graph (TAG) is one kind of important real-world graph-structured data with each node associated with raw texts. For TAGs, traditional few-shot node classification methods directly conduct training on the pre-processed node features and do not consider the raw texts. The performance is highly dependent on the choice of the feature pre-processing method. In this paper, we propose P2TAG, a framework designed for few-shot node classification on TAGs with graph pre-training and prompting. P2TAG first pre-trains the language model (LM) and graph neural network (GNN) on TAGs with self-supervised loss. To fully utilize the ability of language models, we adapt the masked language modeling objective for our framework. The pre-trained model is then used for the few-shot node classification with a mixed prompt method, which simultaneously considers both text and graph information. We conduct experiments on six real-world TAGs, including paper citation networks and product co-purchasing networks. Experimental results demonstrate that our proposed framework outperforms existing graph few-shot learning methods on these datasets with +18.98% ~ +32.14% improvements.
Huanjing Zhao, Beining Yang, Yukuo Cen, Junyu Ren, Yuxiao Dong, Evgeny Kharlamov, Shu Zhao 0005, Jie Tang 0001
KDD7
2024 Alleviating Over-Smoothing via Aggregation over Compact Manifolds
Dongzhuoran Zhou, Bo Xiong 0001, Yue Ma 0009, Evgeny Kharlamov
PAKDD (2)5
2024 ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries
abstract
Dataset search, or more specifically, ad hoc dataset retrieval which is a trending specialized IR task, has received increasing attention in both academia and industry. While methods and systems continue evolving, existing test collections for this task exhibit shortcomings, particularly suffering from lexical bias in pooling and limited to keyword-style queries for evaluation. To address these limitations, in this paper, we construct ACORDAR 2.0, a new test collection for this task which is also the largest to date. To reduce lexical bias in pooling, we adapt dense retrieval models to large structured data, using them to find an extended set of semantically relevant datasets to be annotated. To diversify query forms, we employ a large language model to rewrite keyword queries into high-quality question-style queries. We use the test collection to evaluate popular sparse and dense retrieval models to establish a baseline for future studies. The test collection and source code are publicly available.
Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang 0001, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
SIGIR9
2024 Knowledge graph embedding closed under composition
abstract
Abstract Knowledge Graph Embedding (KGE) has attracted increasing attention. Relation patterns, such as symmetry and inversion, have received considerable focus. Among them, composition patterns are particularly important, as they involve nearly all relations in KGs. However, prior KGE approaches often consider relations to be compositional only if they are well-represented in the training data. Consequently, it can lead to performance degradation, especially for under-represented composition patterns. To this end, we propose HolmE, a general form of KGE with its relation embedding space closed under composition, namely that the composition of any two given relation embeddings remains within the embedding space. This property ensures that every relation embedding can compose, or be composed by other relation embeddings. It enhances HolmE’s capability to model under-represented (also called long-tail) composition patterns with limited learning instances. To our best knowledge, our work is pioneering in discussing KGE with this property of being closed under composition. We provide detailed theoretical proof and extensive experiments to demonstrate the notable advantages of HolmE in modelling composition patterns, particularly for long-tail patterns. Our results also highlight HolmE’s effectiveness in extrapolating to unseen relations through composition and its state-of-the-art performance on benchmark datasets.
Zhuoxun Zheng, Baifan Zhou, Zequn Sun 0001, Chunnong Li, Arild Waaler, Evgeny Kharlamov, Ahmet Soylu
Data Min. Knowl. Discov.8
2023 Literal-Aware Knowledge Graph Embedding for Welding Quality Monitoring: A Bosch Case
Baifan Zhou, Zhuoxun Zheng, Ognjen Savkovic, Irlán Grangel-González, Ahmet Soylu, Evgeny Kharlamov
ISWC8
2023 Scaling Data Science Solutions with Semantics and Machine Learning: Bosch Case
Baifan Zhou, Nikolay Nikolov, Zhuoxun Zheng, Xianghui Luo, Ognjen Savkovic, Dumitru Roman, Ahmet Soylu, Evgeny Kharlamov
ISWC8
2023 GraphMAE2: A Decoding-Enhanced Masked Self-Supervised Graph Learner
abstract
Graph self-supervised learning (SSL), including contrastive and generative approaches, offers great potential to address the fundamental challenge of label scarcity in real-world graph data. Among both sets of graph SSL techniques, the masked graph autoencoders (e.g., GraphMAE)—one type of generative methods—have recently produced promising results. The idea behind this is to reconstruct the node features (or structures)—that are randomly masked from the input—with the autoencoder architecture. However, the performance of masked feature reconstruction naturally relies on the discriminability of the input features and is usually vulnerable to disturbance in the features. In this paper, we present a masked self-supervised learning framework1 GraphMAE2 with the goal of overcoming this issue. The idea is to impose regularization on feature reconstruction for graph SSL. Specifically, we design the strategies of multi-view random re-mask decoding and latent representation prediction to regularize the feature reconstruction. The multi-view random re-mask decoding is to introduce randomness into reconstruction in the feature space, while the latent representation prediction is to enforce the reconstruction in the embedding space. Extensive experiments show that GraphMAE2 can consistently generate top results on various public datasets, including at least 2.45% improvements over state-of-the-art baselines on ogbn-Papers100M with 111M nodes and 1.6B edges.
Yukuo Cen, Xiao Liu 0036, Yuxiao Dong, Evgeny Kharlamov, Jie Tang 0001
WWW6
2023 ApeGNN: Node-Wise Adaptive Aggregation in GNNs for Recommendation
abstract
In recent years, graph neural networks (GNNs) have made great progress in recommendation. The core mechanism of GNNs-based recommender system is to iteratively aggregate neighboring information on the user-item interaction graph. However, existing GNNs treat users and items equally and cannot distinguish diverse local patterns of each node, which makes them suboptimal in the recommendation scenario. To resolve this challenge, we present a node-wise adaptive graph neural network framework ApeGNN. ApeGNN develops a node-wise adaptive diffusion mechanism for information aggregation, in which each node is enabled to adaptively decide its diffusion weights based on the local structure (e.g., degree). We perform experiments on six widely-used recommendation datasets. The experimental results show that the proposed ApeGNN is superior to the most advanced GNN-based recommender methods (up to 48.94%), demonstrating the effectiveness of node-wise adaptive aggregation.
Yifan Zhu 0001, Yuxiao Dong, Yuandong Wang 0002, Wenzheng Feng, Evgeny Kharlamov, Jie Tang 0001
WWW6
2023 GCCAD: Graph Contrastive Coding for Anomaly Detection
abstract
Graph-based anomaly detection has been widely used for detecting malicious activities in real-world applications. Existing attempts to address this problem have thus far focused on structural feature engineering or learning in the binary classification regime. In this work, we propose to leverage graph contrastive learning and present the supervised GCCAD model for contrasting abnormal nodes with normal ones in terms of their distances to the global context (e.g., the average of all nodes). To handle scenarios with scarce labels, we further enable GCCAD as a self-supervised framework by designing a graph corrupting strategy for generating synthetic node labels. To achieve the contrastive objective, we design a graph neural network encoder that can infer and further remove suspicious links during message passing, as well as learn the global context of the input graph. We conduct extensive experiments on four public datasets, demonstrating that 1) GCCAD significantly and consistently outperforms various advanced baselines and 2) its self-supervised version without fine-tuning can achieve comparable performance with its fully supervised version.
Bo Chen 0026, Jing Zhang 0001, Yuxiao Dong, Jian Song 0016, Peng Zhang 0077, Kaibo Xu, Evgeny Kharlamov, Jie Tang 0001
IEEE Trans. Knowl. Data Eng.8
2023 BANDAR: Benchmarking Snippet Generation Algorithms for (RDF) Dataset Search
abstract
The large volume of open data on the Web is expected to be reused and create value. Finding the right data to reuse is a non-trivial task addressed by the recent dataset search systems, which retrieve datasets relevant to a keyword query. An important component of such systems is snippet generation, extracting data from a retrieved dataset to exemplify its content and explain its relevance to the query. Snippet generation algorithms have emerged but were mainly evaluated by user studies. More efficient and reproducible evaluation methods are needed. To meet this challenge, in this article, we present a set of quality metrics for assessing the usefulness of a snippet from different perspectives, and we select and aggregate them into quality profiles for different stages of a dataset search process. Furthermore, we create a benchmark from thousands of collected real-world data needs and datasets, on which we apply the presented quality metrics and profiles to evaluate snippets generated by two existing algorithms and three adapted algorithms. The results, which are reproducible as they are automatically computed without human interaction, show the pros and cons of the tested algorithms and highlight directions for future research. The benchmark data is publicly available.
Xiaxia Wang 0001, Gong Cheng 0001, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
IEEE Trans. Knowl. Data Eng.4
2023 OAG: Linking Entities Across Large-Scale Heterogeneous Knowledge Graphs
abstract
Different knowledge graphs for the same domain are often uniquely housed on the Web. Effectively linking entities from different graphs is critical for building an open and comprehensive knowledge graph. However, linking entities across different sources has thus far faced various challenges, including the increasingly large-scale volume of the data, the heterogeneity of the graphs, and the ambiguity of real-world entities. To address them, we propose a unified framework LinKG. Specifically, we decouple the problem into different linking tasks based on the unique properties of each type of entity. To link word sequence based entities, we propose an LSTM-based method to capture word dependencies. To link entities of large scale, we utilize the hashing technique and convolutional neural networks for scalable and accurate linking. To link ambiguous entities, we propose heterogeneous graph attention networks to leverage heterogeneous structural information. Finally, to validate the design choices of different LinKG modules, we characterize the relationships between different tasks based on the single-domain and multi-domain transfer models. Extensive experiments demonstrate the effectiveness of LinKG with an overall F1-score of 95.15%, based on which we deploy and release the Open Academic Graph (OAG)—the largest publicly available heterogeneous academic graph to date.
Fanjin Zhang, Xiao Liu 0036, Jie Tang 0001, Yuxiao Dong, Peiran Yao, Jie Zhang 0078, Xiaotao Gu, Yan Wang 0120, Evgeny Kharlamov, Kuansan Wang
IEEE Trans. Knowl. Data Eng.9
2022 ExeKG: Executable Knowledge Graph System for User-friendly Data Analytics
abstract
Data analytics including machine learning (ML) is essential to extract insights from production data in modern industries. However, industrial ML is affected by: the low transparency of ML towards non-ML experts; poor and non-unified descriptions of ML practices for reviewing or comprehension; ad-hoc fashion of ML solutions tailored to specific applications, which affects their re-usability. To address these challenges, we propose the concept and a system of executable knowledge graph (KG), which represent KGs that rely on semantic technologies to formally encode ML knowledge and solutions. These KGs can be translated to executable scripts in a reusable and modularised fashion. The demo attendees will use our system to modify, integrate and create executable KGs via a graphic user interface, which offer a user-friendly way to understand, configure, reuse, and create data analytics pipelines.
Zhuoxun Zheng, Baifan Zhou, Dongzhuoran Zhou, Ahmet Soylu, Evgeny Kharlamov
CIKM5
2022 Executable Knowledge Graph for Transparent Machine Learning in Welding Monitoring at Bosch
abstract
With the development of Industry 4.0 technology, modern industries such as Bosch's welding monitoring witnessed the rapid widespread of machine learning (ML) based data analytical applications, which in the case of welding monitoring has led to more efficient and accurate welding monitoring quality. However, industrial ML is affected by the low transparency of ML towards non-ML experts needs. The lack of understanding by domain experts of ML methods hampers the application of ML methods in industry and the reuse of developed ML pipelines, as ML methods are often developed in an ad hoc manner for specific problems. To address these challenges, we propose the concept and a system of executable Knowledge Graph (KG), which formally encode ML knowledge and solutions in KGs, which serve as common language between ML experts and non-ML experts, thus facilitate their communication and increase the transparency of ML methods. We evaluated our system extensively with an industrial use case at Bosch, showing promising results.
Zhuoxun Zheng, Baifan Zhou, Dongzhuoran Zhou, Ahmet Soylu, Evgeny Kharlamov
CIKM5
2022 ScheRe: Schema Reshaping for Enhancing Knowledge Graph Construction
abstract
Automatic knowledge graph (KG) construction is widely used for e.g. data integration, question answering and semantic search. There are many approaches of automatic KG construction. Among which, an important approach is to map the raw data to a given domain KG schema, e.g., domain ontology or conceptual graph, and construct the entities and properties according to the domain KG schema. However, the existing approaches to construct KGs are not always efficient enough and the resulting KGs are not sufficiently application and user-friendly. The main challenge arises from the trade-off: the domain KG schema should be domain-generic and knowledge-oriented, to reflect the general domain knowledge rather than data particularities; while a KG schema should be data-oriented, to cover all data features. If the former is directly used for KG construction, this can cause issues like a high load of blank nodes, which are technical nodes in the KGs that represent unknown entities. To this end, we propose our ScheRe system in the demo, which relies on a schema reshaping algorithm and other two semantic modules for enhancing KG construction. The demo attendees will use ScheRe to reshape a domain KG schema to data specific KG schema, build KGs with industrial data, and experience more user-friendly querying.
Dongzhuoran Zhou, Baifan Zhou, Zhuoxun Zheng, Ahmet Soylu, Ognjen Savkovic, Egor V. Kostylev, Evgeny Kharlamov
CIKM7
2022 Executable Knowledge Graphs for Machine Learning: A Bosch Case of Welding Monitoring
Zhuoxun Zheng, Baifan Zhou, Dongzhuoran Zhou, Xianda Zheng, Gong Cheng 0001, Ahmet Soylu, Evgeny Kharlamov
ISWC7
2022 Ontology Reshaping for Knowledge Graph Construction: Applied on Bosch Welding Case
Dongzhuoran Zhou, Baifan Zhou, Zhuoxun Zheng, Ahmet Soylu, Gong Cheng 0001, Ernesto Jiménez-Ruiz, Egor V. Kostylev, Evgeny Kharlamov
ISWC8
2022 ACORDAR: A Test Collection for Ad Hoc Content-Based (RDF) Dataset Retrieval
abstract
Ad hoc dataset retrieval is a trending topic in IR research. Methods and systems are evolving from metadata-based to content-based ones which exploit the data itself for improving retrieval accuracy but thus far lack a specialized test collection. In this paper, we build and release the first test collection for ad hoc content-based dataset retrieval, where content-oriented dataset queries and content-based relevance judgments are annotated by human experts who are assisted with a dashboard designed specifically for comprehensively and conveniently browsing both the metadata and data of a dataset. We conduct extensive experiments on the test collection to analyze its difficulty and provide insights into the underlying task.
Tengteng Lin, Qiaosheng Chen, Gong Cheng 0001, Ahmet Soylu, Basil Ell, Ruoqi Zhao, Xiaxia Wang 0001, Yu Gu 0016, Evgeny Kharlamov
SIGIR10
2022 GRAND+: Scalable Graph Random Neural Networks
abstract
Graph neural networks (GNNs) have been widely adopted for semi-supervised learning on graphs. A recent study shows that the graph random neural network (GRAND) model can generate state-of-the-art performance for this problem. However, it is difficult for GRAND to handle large-scale graphs since its effectiveness relies on computationally expensive data augmentation procedures. In this work, we present a scalable and high-performance GNN framework GRAND+ for semi-supervised graph learning. To address the above issue, we develop a generalized forward push (GFPush) algorithm in GRAND+ to pre-compute a general propagation matrix and employ it to perform graph data augmentation in a mini-batch manner. We show that both the low time and space complexities of GFPush enable GRAND+ to efficiently scale to large graphs. Furthermore, we introduce a confidence-aware consistency loss into the model optimization of GRAND+, facilitating GRAND+’s generalization superiority. We conduct extensive experiments on seven public datasets of different sizes. The results demonstrate that GRAND+ 1) is able to scale to large graphs and costs less running time than existing scalable GNNs, and 2) can offer consistent accuracy improvements over both full-batch and scalable GNNs across all datasets.
Wenzheng Feng, Yuxiao Dong, Evgeny Kharlamov, Jie Tang 0001
WWW6
2022 SelfKG: Self-Supervised Entity Alignment in Knowledge Graphs
abstract
Entity alignment, aiming to identify equivalent entities across different knowledge graphs (KGs), is a fundamental problem for constructing Web-scale KGs. Over the course of its development, the label supervision has been considered necessary for accurate alignments. Inspired by the recent progress of self-supervised learning, we explore the extent to which we can get rid of supervision for entity alignment. Commonly, the label information (positive entity pairs) is used to supervise the process of pulling the aligned entities in each positive pair closer. However, our theoretical analysis suggests that the learning of entity alignment can actually benefit more from pushing unlabeled negative pairs far away from each other than pulling labeled positive pairs close. By leveraging this discovery, we develop the self-supervised learning objective for entity alignment. We present SelfKG with efficient strategies to optimize this objective for aligning entities without label supervision. Extensive experiments on benchmark datasets demonstrate that SelfKG without supervision can match or achieve comparable results with state-of-the-art supervised baselines. The performance of SelfKG suggests that self-supervised learning offers great potential for entity alignment in KGs. The code and data are available at https://github.com/THUDM/SelfKG.
Xiao Liu 0036, Haoyun Hong, Zeyi Chen, Evgeny Kharlamov, Yuxiao Dong, Jie Tang 0001
WWW5
2021 TDGIA: Effective Injection Attacks on Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have achieved promising performance in various real-world applications. However, recent studies have shown that GNNs are vulnerable to adversarial attacks. In this paper, we study a recently-introduced realistic attack scenario on graphs---graph injection attack (GIA). In the GIA scenario, the adversary is not able to modify the existing link structure and node attributes of the input graph, instead the attack is performed by injecting adversarial nodes into it. We present an analysis on the topological vulnerability of GNNs under GIA setting, based on which we propose the Topological Defective Graph Injection Attack (TDGIA) for effective injection attacks. TDGIA first introduces the topological defective edge selection strategy to choose the original nodes for connecting with the injected ones. It then designs the smooth feature optimization objective to generate the features for the injected nodes. Extensive experiments on large-scale datasets show that TDGIA can consistently and significantly outperform various attack baselines in attacking dozens of defense GNN models. Notably, the performance drop on target GNNs resultant from TDGIA is more than double the damage brought by the best attack solution among hundreds of submissions on KDD-CUP 2020.
Xu Zou 0001, Qinkai Zheng, Yuxiao Dong, Evgeny Kharlamov, Jie Tang 0001
KDD5
2021 PCSG: Pattern-Coverage Snippet Generation for RDF Datasets
Xiaxia Wang 0001, Gong Cheng 0001, Tengteng Lin, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
ISWC6
2021 Efficient Computation of Semantically Cohesive Subgraphs for Keyword-Based Knowledge Graph Exploration
abstract
A knowledge graph (KG) represents a set of entities and their relations. To explore the content of a large and complex KG, a convenient way is keyword-based querying. Traditional methods assign small weights to salient entities or relations, and answer an exploratory keyword query by computing a group Steiner tree (GST), which is a minimum-weight subgraph that connects all the keywords in the query. Recent studies have suggested improving the semantic cohesiveness of a query answer by minimizing the pairwise semantic distances between the entities in a subgraph, but it remains unclear how to efficiently compute such a semantically cohesive subgraph. In this paper, we formulate it as a quadratic group Steiner tree problem (QGSTP) by extending the classical minimum-weight GST problem which is NP-hard. We design two approximation algorithms for QGSTP and prove their approximation ratios. Furthermore, to improve their practical performance, we present heuristics including pruning and ranking strategies.
Gong Cheng 0001, Trung Kien Tran, Evgeny Kharlamov
WWW4
2021 View Selection over Knowledge Graphs in Triple Stores
abstract
Knowledge Graphs (KGs) are collections of interconnected and annotated entities that have become powerful assets for data integration, search enhancement, and other industrial applications. Knowledge Graphs such as DBPEDIA may contain billion of triple relations and are intensively queried with millions of queries per day. A prominent approach to enhance query answering on Knowledge Graph databases is View Materialization, ie., the materialization of an appropriate set of computations that will improve query performance. We study the problem of view materialization and propose a view selection methodology for processing query workloads with more than a million queries. Our approach heavily relies on subgraph pattern mining techniques that allow to create efficient summarizations of massive query workloads while also identifying the candidate views for materialization. In the core of our work is the correspondence between the view selection problem to that of Maximizing a Nondecreasing Submodular Set Function Subject to a Knapsack Constraint . The latter leads to a tractable view-selection process for native triple stores that allows a (1 - e ---1 )-approximation of the optimal selection of views. Our experimental evaluation shows that all the steps of the view-selection process are completed in a few minutes, while the corresponding rewritings accelerate 67.68% of the queries in the DBPEDIA query workload. Those queries are executed in 2.19% of their initial time on average.
Theofilos P. Mailis, Yannis Kotidis, Stamatis Christoforidis, Evgeny Kharlamov, Yannis E. Ioannidis
Proc. VLDB Endow.4
2021 SemML: Facilitating development of ML models for condition monitoring with semantics
abstract
Monitoring of the state, performance, quality of operations and other parameters of equipment and production processes, which is typically referred to as condition monitoring, is an important common practice in many industries including manufacturing, oil and gas, chemical and process industry. In the age of Industry 4.0, where the aim is a deep degree of production automation, unprecedented amounts of data are generated by equipment and processes, and this enables adoption of Machine Learning (ML) approaches for condition monitoring. Development of such ML models is challenging. On the one hand, it requires collaborative work of experts from different areas, including data scientists, engineers, process experts, and managers with asymmetric backgrounds. On the other hand, there is high variety and diversity of data relevant for condition monitoring. Both factors hampers ML modelling for condition monitoring. In this work, we address these challenges by empowering ML-based condition monitoring with semantic technologies. To this end we propose a software system SemML that allows to reuse and generalise ML pipelines for conditions monitoring by relying on semantics. In particular, SemML has several novel components and relies on ontologies and ontology templates for ML task negotiation and for data and ML feature annotation. SemML also allows to instantiate parametrised ML pipelines by semantic annotation of industrial data. With SemML, users do not need to dive into data and ML scripts when new datasets of a studied application scenario arrive. They only need to annotate data and then ML models will be constructed through the combination of semantic reasoning and ML modules. We demonstrate the benefits of SemML on a Bosch use-case of electric resistance welding with very promising results.
Baifan Zhou, Yulia Svetashova, Andre Gusmao, Ahmet Soylu, Gong Cheng 0001, Ralf Mikut, Arild Waaler, Evgeny Kharlamov
J. Web Semant.8
2020 Predicting Quality of Automated Welding with Machine Learning and Semantics: A Bosch Case Study
abstract
Manufacturing of car bodies heavily relies on demanding welding processes of joining body parts together that introduce thousands of joining welding spots in each car. Quality monitoring for these spots impacts production efficiency and cost. In this paper we develop an ML pipeline to predict the spot quality before the actual welding happens. This pipeline is based on a Feature Engineering~(FE) approach to manually design features using domain knowledge. We evaluated the pipeline with two datasets from industrial plants, achieving very promising results with prediction errors around 2%. Then, we develop an approach to semantically enhance FE pipelines in order to automate the ML process without compromising the prediction accuracy and to facilitate generalisation and transfer of FE-based models to other datasets and processes. Our ML pipeline has been deployed offline on various Bosch manufacturing datasets in a controlled environment since early 2019 and evaluated.
Baifan Zhou, Yulia Svetashova, Seongsu Byeon, Tim Pychynski, Ralf Mikut, Evgeny Kharlamov
CIKM6
2020 SemFE: Facilitating ML Pipeline Development with Semantics
abstract
Machine learning (ML) based data analysis has attracted an increasing attention in the manufacturing industry, however, many challenges hamper their wide spread adoption. The main challenges are the high costs of labour-intensive data preparation from diverse sources and processes, the asymmetrical backgrounds of the experts involved in manufacturing analyses that impede efficient communication between them, and the lack of generalisability of ML models tailored to specific applications. Our semantically enhanced ML pipeline, SemFE, with feature engineering addresses these challenges, serving as a bridge to bring the endeavours of experts together, and making data science accessible to non-ML-experts. SemFE relies on ontologies for discrete manufacturing monitoring that encapsulate domain and ML knowledge; it has five novel semantic modules for automation of ML-pipeline development and user-friendly GUIs. The demo attendees will be able to use our system to build manufacturing monitoring ML pipelines, and to design their own pipelines with minimal prior knowledge of machine learning.
Baifan Zhou, Yulia Svetashova, Tim Pychynski, Ildar Baimuratov, Ahmet Soylu, Evgeny Kharlamov
CIKM6
2020 Entity Summarization with User Feedback
Qingxia Liu, Yue Chen 0039, Gong Cheng 0001, Evgeny Kharlamov, Junyou Li, Yuzhong Qu
ESWC4
2020 On Equivalence and Cores for Incomplete Databases in Open and Closed Worlds
abstract
Data exchange heavily relies on the notion of incomplete database instances. Several semantics for such instances have been proposed and include open (OWA), closed (CWA), and open-closed (OCWA) world. For all these semantics important questions are: whether one incomplete instance semantically implies another; when two are semantically equivalent; and whether a smaller or smallest semantically equivalent instance exists. For OWA and CWA these questions are fully answered. For several variants of OCWA, however, they remain open. In this work we adress these questions for Closed Powerset semantics and the OCWA semantics of Libkin and Sirangelo, 2011. We define a new OCWA semantics, called OCWA*, in terms of homomorphic covers that subsumes both semantics, and characterize semantic implication and equivalence in terms of such covers. This characterization yields a guess-and-check algorithm to decide equivalence, and shows that the problem is NP-complete. For the minimization problem we show that for several common notions of minimality there is in general no unique minimal equivalent instance for Closed Powerset semantics, and consequently not for the more expressive OCWA* either. However, for Closed Powerset semantics we show that one can find, for any incomplete database, a unique finite set of its subinstances which are subinstances (up to renaming of nulls) of all instances semantically equivalent to the original incomplete one. We study properties of this set, and extend the analysis to OCWA*.
Henrik Forssell, Evgeny Kharlamov, Evgenij Thorstensen
ICDT2
2020 Semantic Integration of Bosch Manufacturing Data Using Virtual Knowledge Graphs
Elem Guzel Kalayci, Irlán Grangel-González, Felix Lösch, Guohui Xiao 0001, Anees Mehdi, Evgeny Kharlamov, Diego Calvanese
ISWC (2)6
2020 Ontology-Enhanced Machine Learning: A Bosch Use Case of Welding Quality Monitoring
Yulia Svetashova, Baifan Zhou, Tim Pychynski, York Sure-Vetter, Ralf Mikut, Evgeny Kharlamov
ISWC (2)7
2020 Keyword Search over Knowledge Graphs via Static and Dynamic Hub Labelings
abstract
Keyword search is a prominent approach to querying Web data. For graph-structured data, a widely accepted semantics for keywords is based on group Steiner trees. For this NP-hard problem, existing algorithms with provable quality guarantees have prohibitive run time on large graphs. In this paper, we propose practical approximation algorithms with a guaranteed quality of computed answers and very low run time. Our algorithms rely on Hub Labeling (HL), a structure that labels each vertex in a graph with a list of vertices reachable from it, which we use to compute distances and shortest paths. We devise two HLs: a conventional static HL that uses a new heuristic to improve pruned landmark labeling, and a novel dynamic HL that inverts and aggregates query-relevant static labels to more efficiently process vertex sets. Our approach allows to compute a reasonably good approximation of answers to keyword queries in milliseconds on million-scale knowledge graphs.
Gong Cheng 0001, Evgeny Kharlamov
WWW3
2020 Fast Computation of Explanations for Inconsistency in Large-Scale Knowledge Graphs
abstract
Knowledge graphs (KGs) are essential resources for many applications including Web search and question answering. As KGs are often automatically constructed, they may contain incorrect facts. Detecting them is a crucial, yet extremely expensive task. Prominent solutions detect and explain inconsistency in KGs with respect to accompanying ontologies that describe the KG domain of interest. Compared to machine learning methods they are more reliable and human-interpretable but scale poorly on large KGs. In this paper, we present a novel approach to dramatically speed up the process of detecting and explaining inconsistency in large KGs by exploiting KG abstractions that capture prominent data patterns. Though much smaller, KG abstractions preserve inconsistency and their explanations. Our experiments with large KGs (e.g., DBpedia and Yago) demonstrate the feasibility of our approach and show that it significantly outperforms the popular baseline.
Trung Kien Tran, Mohamed H. Gad-Elrab, Daria Stepanova 0001, Evgeny Kharlamov, Jannik Strötgen
WWW4
2019 Towards More Usable Dataset Search: From Query Characterization to Snippet Generation
abstract
Reusing published datasets on the Web is of great interest to researchers and developers. Their data needs may be met by submitting queries to a dataset search engine to retrieve relevant datasets. In this ongoing work towards developing a more usable dataset search engine, we characterize real data needs by annotating the semantics of 1,947 queries using a novel fine-grained scheme, to provide implications for enhancing dataset search. Based on the findings, we present a query-centered framework for dataset search, and explore the implementation of snippet generation and evaluate it with a preliminary user study.
Jinchi Chen, Xiaxia Wang 0001, Gong Cheng 0001, Evgeny Kharlamov, Yuzhong Qu
CIKM4
2019 MiCRon: Making Sense of News via Relationship Subgraphs
abstract
Knowledge graphs (KGs) have been extensively used to annotate text, e.g., news articles, in order to enhance its comprehension by readers. This requires to map entities occurring in the news to the target entities of the KG and to extract a so-called relationship sub-graph (RSG) that spans these entities. RSG extraction is computationally demanding and cannot scale to large KGs. Existing approximation algorithms that focus on structurally compact RSGs are not satisfactory since they often return no answers. We address this problem and develop an efficient algorithm to find approximations that connect the most salient subset of the target entities. Moreover, we propose a context-aware method to rank RSGs by their relevance to the news and their semantic cohesion. In the demo we will present our approach and the attendees will be able to experience how our system MiCRon helps to make sense of news article by computing and presenting RSGs relevant to these articles.
Zixian Huang, Gong Cheng 0001, Evgeny Kharlamov, Yuzhong Qu
CIKM4
2019 Validation of SHACL Constraints over KGs with OWL 2 QL Ontologies via Rewriting
abstract
Constraints have traditionally been used to ensure data quality. Recently, several constraint languages such as SHACL, as well as mechanisms for constraint validation, have been proposed for Knowledge Graphs (KGs). KGs are often enhanced with ontologies that define relevant background knowledge in a formal language such as OWL 2 QL. However, existing systems for constraint validation either ignore these ontologies, or compile ontologies and constraints into rules that should be executed by some rule engine. In the latter case, one has to rely on different systems when validating constrains over KGs and over ontology-enhanced KGs. In this work, we address this problem by defining rewriting techniques that allow to compile an OWL 2 QL ontology and a set of SHACL constraints into another set of SHACL constraints. We show that in the general case the rewriting may not exists, but it always exists for the positive fragment of SHACL. Our rewriting techniques allow to validate constraints over KGs with and without ontologies using the same SHACL validation engines.
Ognjen Savkovic, Evgeny Kharlamov, Steffen Lamparter
ESWC2
2019 A Framework for Evaluating Snippet Generation for Dataset Search
Xiaxia Wang 0001, Jinchi Chen, Gong Cheng 0001, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
ISWC (1)6
2019 An Efficient Index for RDF Query Containment
abstract
Query containment is a fundamental operation used to expedite query processing in view materialisation and query caching techniques. Since query containment has been shown to be NP-complete for arbitrary conjunctive queries on RDF graphs, we introduce a simpler form of conjunctive queries that we name f-graph queries. We first show that containment checking for f-graph queries can be solved in polynomial time. Based on this observation, we propose a novel indexing structure, named mv-index, that allows for fast containment checking between a single f-graph query and an arbitrary number of stored queries. Search is performed in polynomial time in the combined size of the query and the index. We then show how our algorithms and structures can be extended for arbitrary conjunctive queries on RDF graphs by introducing f-graph witnesses, i.e., f-graph representatives of conjunctive queries. F-graph witnesses have the following interesting property, a conjunctive query for RDF graphs is contained in another query only if its corresponding f-graph witness is also contained in it. The latter allows to use our indexing structure for the general case of conjunctive query containment. This translates in practice to microseconds or less for the containment test against hundreds of thousands of queries that are indexed within our structure.
Theofilos P. Mailis, Yannis Kotidis, Vaggelis Nikolopoulos, Evgeny Kharlamov, Ian Horrocks 0001, Yannis E. Ioannidis
SIGMOD Conference4
2019 An ontology-mediated analytics-aware approach to support monitoring and diagnostics of static and streaming data
Evgeny Kharlamov, Yannis Kotidis, Theofilos P. Mailis, Christian Neuenstadt, Charalampos Nikolaou, Özgür L. Özçep, Christoforos Svingos, Dmitriy Zheleznyakov, Yannis E. Ioannidis, Steffen Lamparter, Ralf Möller 0001, Arild Waaler
J. Web Semant.1
2019 Semantically-enhanced rule-based diagnostics for industrial Internet of Things: The SDRL language and case study for Siemens trains and turbines
Evgeny Kharlamov, Gulnar Mehdi, Ognjen Savkovic, Guohui Xiao 0001, Elem Guzel Kalayci, Mikhail Roshchin
J. Web Semant.1
2019 On expansion and contraction of DL-Lite knowledge bases
Dmitriy Zheleznyakov, Evgeny Kharlamov, Werner Nutt, Diego Calvanese
J. Web Semant.2
2018 Towards Semantically Enhanced Digital Twins
abstract
Digital twins (DTs) are a powerful mechanism for representing complex industrial assets such as oil platforms as digital models. These models can facilitate temporal analyses and computer simulations of assets. In order to enable this, DTs should be able to capture characteristics of an asset as specified by the manufacturer, its state during the run time, as well as how the asset interacts with other assets in a complex system. We argue that semantic technologies and in particular semantic models or ontologies is promising modelling paradigm for DTs. Semantic models allow to capture complex systems in an intuitive fashion, can be written in standardised ontology languages, and come with a wide range of off-the-shelf systems to design, maintain, query, and navigate semantic models. In this work we report our preliminary results on developing a system that would support semantic-based DTs. In particular, we plan to augment the PI System developed by OSIsoft with ontologies and show how the resulting solution can help in simplifying analytical and machine learning routines for DTs.
Evgeny Kharlamov, Francisco Martín-Recuerda, Brandon Perry, David B. Cameron, Roar Fjellheim, Arild Waaler
IEEE BigData1
2018 Towards Simplification of Analytical Workflows With Semantics at Siemens (Extended Abstract)
abstract
Analytical workflows are heavily used in large and data intensive companies. An important application of such workflows in Siemens is equipment analytics when equipment KPIs and reports are computed by aggregating equipment's operational, master, and analytical data. In Siemens this data satisfies big data dimensions and this dependence poses significant challenges in authoring, reuse, and maintenance of analytical workflows by engineers and data scientists. In this work we propose to address these problems by relying on semantic technologies: we use ontologies to give a high level representation of equipment's operational and master data and offer a high level language to express KPIs over ontologies. We implemented our approach, integrated it with KNIME, and evaluated at Siemens. This is a preliminary work and we are excited about its further extensions.
Evgeny Kharlamov, Gulnar Mehdi, Ognjen Savkovic, Guohui Xiao 0001, Steffen Lamparter, Ian Horrocks 0001, Arild Waaler
IEEE BigData1
2018 Finding Data Should be Easier than Finding Oil
abstract
The competitiveness of modern enterprises heavily depends on their ability to make the right business decisions by relying on efficient and timely analysis of the right business critical data. In large and data intensive companies such as Equinor, a Norwegian multinational oil and gas company with more than 20,000 employees, gathering such data is not a trivial task due to the growing size and complexity of corporate information sources. As a result, the data gathering task is often the most time-consuming part of the decision making process, in particular when it comes to the work processes of Equinor’s exploration geologists that should find in a timely manner new exploitable accumulations of oil or gas in given areas by analysing data about these areas. In this work we present our experience in addressing this data challenge tast at Equinor. We have developed and deployed at Equinor a semantic data access system that relies on the Ontology Based Data Access (OBDA) approach. Our system is based on our solid theoretical contributions and has been extensively evaluated at Equinor.
Evgeny Kharlamov, Martin G. Skjæveland, Dag Hovland, Theofilos P. Mailis, Ernesto Jiménez-Ruiz, Guohui Xiao 0001, Ahmet Soylu, Ian Horrocks 0001, Arild Waaler
IEEE BigData1
2018 Event-Enhanced Learning for KG Completion
Martin Ringsquandl, Evgeny Kharlamov, Daria Stepanova 0001, Marcel Hildebrandt, Steffen Lamparter, Raffaello Lepratti, Ian Horrocks 0001, Peer Kröger
ESWC2
2018 Rule Learning from Knowledge Graphs Guided by Embedding Models
Vinh Thinh Ho, Daria Stepanova 0001, Mohamed H. Gad-Elrab, Evgeny Kharlamov, Gerhard Weikum
ISWC (1)4
2017 Towards a semantic keyword search over industrial knowledge graphs (extended abstract)
abstract
Knowlege graphs have become powerful assets for enhancing search and data integration and are now widely used in both academia and industry. Existing semantics of keyword search over knowledge graphs, especially in industrial context, has limitations. In this extended abstract we discuss these limitations and ingredients for alternative semantics.
Gong Cheng 0001, Evgeny Kharlamov
IEEE BigData2
2017 On event-driven knowledge graph completion in digital factories
abstract
Smart factories are equipped with machines that can sense their manufacturing environments, interact with each other, and control production processes. Smooth operation of such factories requires that the machines and engineering personnel that conduct their monitoring and diagnostics share a detailed common industrial knowledge about the factory, e.g., in the form of knowledge graphs. Creation and maintenance of such knowledge is expensive and requires automation. In this work we show how machine learning that is specifically tailored towards industrial applications can help in knowledge graph completion. In particular, we show how knowledge completion can benefit from event logs that are common in smart factories. We evaluate this on the knowledge graph from a real world-inspired smart factory with encouraging results.
Martin Ringsquandl, Evgeny Kharlamov, Daria Stepanova 0001, Steffen Lamparter, Raffaello Lepratti, Ian Horrocks 0001, Peer Kröger
IEEE BigData2
2017 SemFacet: Making Hard Faceted Search Easier
abstract
Faceted search is a prominent search paradigm that became the standard in many Web applications and has also been recently proposed as a suitable paradigm for exploring and querying RDF graphs. One of the main challenges that hampers usability of faceted search systems especially in the RDF context is information overload, that is, when the size of faceted interfaces becomes comparable to the size of the data over which the search is performed. In this demo we present (an extension of) our faceted search system SemFacet and focus on features that address the information overload: ranking, aggregation, and reachability. The demo attendees will be able to try our system on an RDF graph that models online shopping over a catalogs with up to millions of products.
Evgeny Kharlamov, Luca Giacomelli, Evgeny Sherkhonov, Bernardo Cuenca Grau, Egor V. Kostylev, Ian Horrocks 0001
CIKM1
2017 Semantic Rules for Machine Diagnostics: Execution and Management
abstract
Rule-based diagnostics of equipment is an important task in industry. In this paper we present how semantic technologies can enhance diagnostics. In particular, we present our semantic rule language sigRL that is inspired by the real diagnostic languages used in Siemens. SigRL allows to write compact yet powerful diagnostic programs by relying on a high level data independent vocabulary, diagnostic ontologies, and queries over these ontologies. We study computational complexity of SigRL: execution of diagnostic programs, provenance computation, as well as automatic verification of redundancy and inconsistency in diagnostic programs.
Evgeny Kharlamov, Ognjen Savkovic, Guohui Xiao 0001, Rafael Peñaloza, Gulnar Mehdi, Mikhail Roshchin, Ian Horrocks 0001
CIKM1
2017 SemDia: Semantic Rule-Based Equipment Diagnostics Tool
abstract
Rule-based diagnostics of power generating equipment is an important task in industry. In this demo we present how semantic technologies can enhance diagnostics. In particular, we present our semantic rule language sigRL that is inspired by the real diagnostic languages in Siemens. SigRL allows to write compact yet powerful diagnostic programs by relying on a high level data independent vocabulary, diagnostic ontologies, and queries over these ontologies. We present our diagnostic system SemDia. The attendees will be able to write diagnostic programs in SemDia using sigRL over 50 Siemens turbines. We also present how such programs can be automatically verified for redundancy and inconsistency. Moreover, the attendees will see the provenance service that SemDia provides to trace the origin of diagnostic results.
Gulnar Mehdi, Evgeny Kharlamov, Ognjen Savkovic, Guohui Xiao 0001, Elem Guzel Kalayci, Sebastian Brandt 0001, Ian Horrocks 0001, Mikhail Roshchin, Thomas A. Runkler
CIKM2
2017 Semantic Rule-Based Equipment Diagnostics
Gulnar Mehdi, Evgeny Kharlamov, Ognjen Savkovic, Guohui Xiao 0001, Elem Guzel Kalayci, Sebastian Brandt 0001, Ian Horrocks 0001, Mikhail Roshchin, Thomas A. Runkler
ISWC (2)2
2017 Semantic Faceted Search with Aggregation and Recursion
Evgeny Sherkhonov, Bernardo Cuenca Grau, Evgeny Kharlamov, Egor V. Kostylev
ISWC (1)3
2017 Ontology Based Data Access in Statoil
Evgeny Kharlamov, Dag Hovland, Martin G. Skjæveland, Dimitris Bilidas, Ernesto Jiménez-Ruiz, Guohui Xiao 0001, Ahmet Soylu, Davide Lanti, Martín Rezk, Dmitriy Zheleznyakov, Martin Giese, Hallstein Lie, Yannis E. Ioannidis, Yannis Kotidis, Manolis Koubarakis, Arild Waaler
J. Web Semant.1
2017 Semantic access to streaming and static data at Siemens
Evgeny Kharlamov, Theofilos P. Mailis, Gulnar Mehdi, Christian Neuenstadt, Özgür L. Özçep, Mikhail Roshchin, Nina Solomakhina, Ahmet Soylu, Christoforos Svingos, Sebastian Brandt 0001, Martin Giese, Yannis E. Ioannidis, Steffen Lamparter, Ralf Möller 0001, Yannis Kotidis, Arild Waaler
J. Web Semant.1
2016 A semantic approach to polystores
abstract
In the database community Polystores is an emerging and promising approach for data federation that aims at designing a unified querying layer over multiple data models. In the Semantic Web community a similar in spirit approach of Ontology-Based Data Access (OBDA) has been recently proposed, attracted a lot of attention, and proved its success in several industrial scenarios. In this paper we discuss a semantic approach to building polystores using the OBDA paradigm. We also present our system Optique that is utilized in an industrial application of performing turbine diagnostics in Siemens.
Evgeny Kharlamov, Theofilos P. Mailis, Konstantina Bereta, Dimitris Bilidas, Sebastian Brandt 0001, Ernesto Jiménez-Ruiz, Steffen Lamparter, Christian Neuenstadt, Özgür L. Özçep, Ahmet Soylu, Christoforos Svingos, Guohui Xiao 0001, Dmitriy Zheleznyakov, Diego Calvanese, Ian Horrocks 0001, Martin Giese, Yannis E. Ioannidis, Yannis Kotidis, Ralf Möller 0001, Arild Waaler
IEEE BigData1
2016 Capturing Industrial Information Models with Ontologies and Constraints
Evgeny Kharlamov, Bernardo Cuenca Grau, Ernesto Jiménez-Ruiz, Steffen Lamparter, Gulnar Mehdi, Martin Ringsquandl, Yavor Nenov, Stephan Grimm, Mikhail Roshchin, Ian Horrocks 0001
ISWC (2)1
2016 Towards Analytics Aware Ontology Based Access to Static and Streaming Data
Evgeny Kharlamov, Yannis Kotidis, Theofilos P. Mailis, Christian Neuenstadt, Charalampos Nikolaou, Özgür L. Özçep, Christoforos Svingos, Dmitriy Zheleznyakov, Sebastian Brandt 0001, Ian Horrocks 0001, Yannis E. Ioannidis, Steffen Lamparter, Ralf Möller 0001
ISWC (2)1
2016 Ontology-Based Integration of Streaming and Static Relational Data with Optique
abstract
Real-time processing of data coming from multiple heterogeneous data streams and static databases is a typical task in many industrial scenarios such as diagnostics of large machines. A complex diagnostic task may require a collection of up to hundreds of queries over such data. Although many of these queries retrieve data of the same kind, such as temperature measurements, they access structurally different data sources. In this work we show how Semantic Technologies implemented in our system optique can simplify such complex diagnostics by providing an abstraction layer---ontology---that integrates heterogeneous data. In a nutshell, optique allows complex diagnostic tasks to be expressed with just a few high-level semantic queries. The system can then automatically enrich these queries, translate them into a collection with a large number of low-level data queries, and finally optimise and efficiently execute the collection in a heavily distributed environment. We will demo the benefits of optique on a real world scenario from Siemens.
Evgeny Kharlamov, Sebastian Brandt 0001, Ernesto Jiménez-Ruiz, Yannis Kotidis, Steffen Lamparter, Theofilos P. Mailis, Christian Neuenstadt, Özgür L. Özçep, Christoph Pinkel, Christoforos Svingos, Dmitriy Zheleznyakov, Ian Horrocks 0001, Yannis E. Ioannidis, Ralf Möller 0001
SIGMOD Conference1
2016 Faceted search over RDF-based knowledge graphs
Marcelo Arenas, Bernardo Cuenca Grau, Evgeny Kharlamov, Sarunas Marciuska, Dmitriy Zheleznyakov
J. Web Semant.3
2015 RODI: A Benchmark for Automatic Mapping Generation in Relational-to-Ontology Data Integration
Christoph Pinkel, Carsten Binnig, Ernesto Jiménez-Ruiz, Wolfgang May, Dominique Ritze, Martin G. Skjæveland, Alessandro Solimando, Evgeny Kharlamov
ESWC8
2015 BootOX: Practical Mapping of RDBs to OWL 2
Ernesto Jiménez-Ruiz, Evgeny Kharlamov, Dmitriy Zheleznyakov, Ian Horrocks 0001, Christoph Pinkel, Martin G. Skjæveland, Evgenij Thorstensen, Jose Mora
ISWC (2)2
2015 Ontology Based Access to Exploration Data at Statoil
Evgeny Kharlamov, Dag Hovland, Ernesto Jiménez-Ruiz, Davide Lanti, Hallstein Lie, Christoph Pinkel, Martín Rezk, Martin G. Skjæveland, Evgenij Thorstensen, Guohui Xiao 0001, Dmitriy Zheleznyakov, Ian Horrocks 0001
ISWC (2)1
2014 Faceted Search over Ontology-Enhanced RDF Data
abstract
An increasing number of applications rely on RDF, OWL 2, and SPARQL for storing and querying data. SPARQL, however, is not targeted towards end-users, and suitable query interfaces are needed. Faceted search is a prominent approach for end-user data access, and several RDF-based faceted search systems have been developed. There is, however, a lack of rigorous theoretical underpinning for faceted search in the context of RDF and OWL 2. In this paper, we provide such solid foundations. We formalise faceted interfaces for this context, identify a fragment of first-order logic capturing the underlying queries, and study the complexity of answering such queries for RDF and OWL 2 profiles. We then study interface generation and update, and devise efficiently implementable algorithms. Finally, we have implemented and tested our faceted search algorithms for scalability, with encouraging results.
Marcelo Arenas, Bernardo Cuenca Grau, Evgeny Kharlamov, Sarunas Marciuska, Dmitriy Zheleznyakov
CIKM3
2014 How Semantic Technologies Can Enhance Data Access at Siemens Energy
Evgeny Kharlamov, Nina Solomakhina, Özgür L. Özçep, Dmitriy Zheleznyakov, Thomas Hubauer, Steffen Lamparter, Mikhail Roshchin, Ahmet Soylu, Stuart Watson
ISWC (1)1
2013 Controlled Query Evaluation over OWL 2 RL Ontologies
Bernardo Cuenca Grau, Evgeny Kharlamov, Egor V. Kostylev, Dmitriy Zheleznyakov
ISWC (1)2
2012 QUASAR: querying annotation, structure, and reasoning
abstract
An increasing number of systems provide the ability to semantically annotate documents. OpenCalais [4], Evri API [2], Zemanta [6], and Alchemy API [1] are web-hosted systems that return annotated documents, i. e. documents with annotations that are overlayed on the document structure. Many of the annotations can be linked to standard ontologies, such as DBpedia and YAGO. These annotations give insight as to the meaning of documents in a variety of ways, identifying entities and relationships inside them, classifying them according to topic or theme, and giving the attitude or sentiment of a document or document fragment. In order for users (or applications) to make use of these annotations with a means to access and manipulate documents that contain them, we provide a query language for doing this and demonstrate its utility on a demo system built on top of diverse semantic annotators and external ontologies. We explain how integrating semantic annotations and utilizing external knowledge helps in increasing the quality of query answers over annotated documents by both filtering out irrelevant answers and obtaining extra answers that are not explicitly available in the annotated documents.
Luying Chen, Michael Benedikt, Evgeny Kharlamov
EDBT3
2012 Answering Queries using Views over Probabilistic XML: Complexity and Tractability
abstract
We study the complexity of query answering using views in a probabilistic XML setting, identifying large classes of XPath queries -- with child and descendant navigation and predicates -- for which there are efficient (PTime) algorithms. We consider this problem under the two possible semantics for XML query results: with persistent node identifiers and in their absence. Accordingly, we consider rewritings that can exploit a single view, by means of compensation, and rewritings that can use multiple views, by means of intersection. Since in a probabilistic setting queries return answers with probabilities, the problem of rewriting goes beyond the classic one of retrieving XML answers from views. For both semantics of XML queries, we show that, even when XML answers can be retrieved from views, their probabilities may not be computable. For rewritings that use only compensation, we describe a PTime decision procedure, based on easily verifiable criteria that distinguish between the feasible cases -- when probabilistic XML results are computable -- and the unfeasible ones. For rewritings that can use multiple views, with compensation and intersection, we identify the most permissive conditions that make probabilistic rewriting feasible, and we describe an algorithm that is sound in general, and becomes complete under fairly permissive restrictions, running in PTime modulo worst-case exponential time equivalence tests. This is the best we can hope for since intersection makes query equivalence intractable already over deterministic data. Our algorithm runs in PTime whenever deterministic rewritings can be found in PTime.
Bogdan Cautis, Evgeny Kharlamov
Proc. VLDB Endow.2
2011 Capturing Instance Level Ontology Evolution for DL-Lite
Evgeny Kharlamov, Dmitriy Zheleznyakov
ISWC (1)1
2011 Capturing continuous data and answering aggregate queries in probabilistic XML
abstract
Sources of data uncertainty and imprecision are numerous. A way to handle this uncertainty is to associate probabilistic annotations to data. Many such probabilistic database models have been proposed, both in the relational and in the semi-structured setting. The latter is particularly well adapted to the management of uncertain data coming from a variety of automatic processes. An important problem, in the context of probabilistic XML databases, is that of answering aggregate queries (count, sum, avg, etc.), which has received limited attention so far. In a model unifying the various (discrete) semi-structured probabilistic models studied up to now, we present algorithms to compute the distribution of the aggregation values (exploiting some regularity properties of the aggregate functions) and probabilistic moments (especially expectation and variance) of this distribution. We also prove the intractability of some of these problems and investigate approximation techniques. We finally extend the discrete model to a continuous one, in order to take into account continuous data values, such as measurements from sensor networks, and extend our algorithms and complexity results to the continuous case.
Serge Abiteboul, T.-H. Hubert Chan, Evgeny Kharlamov, Werner Nutt, Pierre Senellart
ACM Trans. Database Syst.3
2010 Aggregate queries for discrete and continuous probabilistic XML
abstract
Sources of data uncertainty and imprecision are numerous. A way to handle this uncertainty is to associate probabilistic annotations to data. Many such probabilistic database models have been proposed, both in the relational and in the semi-structured setting. The latter is particularly well adapted to the management of uncertain data coming from a variety of automatic processes. An important problem, in the context of probabilistic XML databases, is that of answering aggregate queries (count, sum, avg, etc.), which has received limited attention so far. In a model unifying the various (discrete) semi-structured probabilistic models studied up to now, we present algorithms to compute the distribution of the aggregation values (exploiting some regularity properties of the aggregate functions) and probabilistic moments (especially, expectation and variance) of this distribution. We also prove the intractability of some of these problems and investigate approximation techniques. We finally extend the discrete model to a continuous one, in order to take into account continuous data values, such as measurements from sensor networks, and present algorithms to compute distribution functions and moments for various classes of continuous distributions of data values.
Serge Abiteboul, T.-H. Hubert Chan, Evgeny Kharlamov, Werner Nutt, Pierre Senellart
ICDT3
2010 Evolution of DL-Lite Knowledge Bases
Diego Calvanese, Evgeny Kharlamov, Werner Nutt, Dmitriy Zheleznyakov
ISWC (1)2
2010 Probabilistic XML via Markov Chains
abstract
We show how Recursive Markov Chains (RMCs) and their restrictions can define probabilistic distributions over XML documents, and study tractability of querying over such models. We show that RMCs subsume several existing probabilistic XML models. In contrast to the latter, RMC models (i) capture probabilistic versions of XML schema languages such as DTDs, (ii) can be exponentially more succinct, and (iii) do not restrict the domain of probability distributions to be finite. We investigate RMC models for which tractability can be achieved, and identify several tractable fragments that subsume known tractable probabilistic XML models. We then look at the space of models between existing probabilistic XML formalisms and RMCs, giving results on the expressiveness and succinctness of RMC subclasses, both with each other and with prior formalisms.
Michael Benedikt, Evgeny Kharlamov, Dan Olteanu, Pierre Senellart
Proc. VLDB Endow.2
2008 Incompleteness in information integration
abstract
Information integration is becoming a critical problem for both businesses and individuals. The data, especially the one that comes from the Web, is naturally incomplete, that is, some data values may be unknown or lost because of communication problems, hidden due to privacy considerations. At the same time research in (virtual) integration in the community focusses on null-free sources and addresses limited forms of incompleteness only. In our work we aim to extend current results on virtual integration by considering various forms of incompleteness at the level of the sources, the integrated database and the queries (we call this Incomplete Information Integration , or III). More specifically, we aim to extend current query answering techniques for local-, and global-as-view integration to integration of tables with SQL nulls, Codd tables, etc. We also aim to consider incomplete answers as a natural extension of the classical approach. Our main research issues are (i) semantics of III, (ii) semantics of query answering in III, (iii) complexity of query answering, and (iv) algorithms (possibly approximate) to compute the answers.
Evgeny Kharlamov, Werner Nutt
Proc. VLDB Endow.1