Michael Cochez

dblp:83/11448 · DBLP profile ↗
← Back
38ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0001-5726-4638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Theory of computation · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Contextualized Interpretation of Machine-Learning Predictions in Rare Disease Omics: Integrating SHAP, Biomedical Knowledge, and Language Models
Daniel Daza, Yorrick Jaspers, Alberto Bernardi, Luca Costabello, Christophe Guéret, Michael Cochez, Marc Engelen, Stephan Kemp, Martijn Schut
AIME (2)6
2026 SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao, Andreas Zimmerer, Olga Ovcharenko, Tianji Cong, Matthew Russo, Gerardo Vitagliano, Michael Cochez, Fatma Özcan 0001, Gautam Gupta, Thibaud Hottelier, H. V. Jagadish, Kris Kissel, Sebastian Schelter, Andreas Kipf, Immanuel Trummer
Proc. VLDB Endow.7
2025 Bias by Design: Diversity Quantification to Mitigate Structural Bias Effects in AIG Logic Optimization
abstract
And-Inverter Graphs (AIGs) are a fundamental data structure in logic optimization, widely used in modern electronic design automation. A persistent challenge in AIG optimization is structural bias, where the initial graph structure strongly influences optimization quality by restricting the search space, often resulting in subpar outcomes. Existing methods address this issue by running multiple optimization workflows in parallel, relying on a trial-and-error approach that lacks a systematic way to measure structural diversity or assess effectiveness, making them computationally expensive and inefficient. This paper introduces a novel framework for systematically evaluating and reducing structural bias by measuring structural diversity, defined as the degree of dissimilarity between AIG graphs. Several traditional graph similarity measures and newly proposed AIG-specific metrics, including the Rewrite, Refactor, and Resub Scores, are explored. Results reveal limitations in traditional graph similarity metrics and highlight the effectiveness of the proposed AIG-specific measures in quantifying structural dissimilarity. Notably, the RRR Score shows a strong correlation (Pearson correlation coefficient,$r$= 0.79) with post-optimization structural differences, demonstrating the reliability of the metric in capturing meaningful variations between AIG structures. This work addresses the challenge of quantifying structural bias and offers a methodology that can potentially improve optimization outcomes, with future extensions applicable to other logic graph types.
Isabella Venancia Gardner, Marcel Walter, Yukio Miyasaka, Robert Wille, Michael Cochez
DATE5
2025 The ART of Link Prediction with KGEs
abstract
Link Prediction (LP) in Knowledge Graphs (KGs) is typically framed as ranking candidate entities for a query of the form $(entity, relation,?)$, with models evaluated on their ability to rank the correct entities for each query. At the same time, Knowledge Graph Embedding (KGE) models used for this task produce unnormalised scores, making it unclear how to interpret their belief in the truthfulness of triples across different queries. Together, these two factors create a blind spot: models can achieve perfect rankings while assigning scores that are not comparable across queries, limiting their utility in downstream tasks or even in identifying the most plausible triples overall. Indeed, this issue becomes clear when test triples are ranked globally and evaluated with IR metrics, revealing that models with unnormalized scores often perform poorly due to inconsistent scoring across queries. To address this problem, we propose a new KGE model, called ART, which exploits probabilistic Auto-Regressive modelling and hence is normalised by design. Despite its conceptual simplicity, we show that ART outperforms prior art for discriminative and generative LP as well as other post-hoc calibration techniques.
Yannick Brunink, Michael Cochez, Jacopo Urbani
NeSy2
2025 Do Graph Neural Network States Contain Graph Properties?
abstract
Deep neural networks (DNNs) achieve state-of-the-art performance on many tasks, but this often requires increasingly larger model sizes, which in turn leads to more complex internal representations. Explainability techniques (XAI) have made remarkable progress in the interpretability of ML models. However, the non-euclidean nature of Graph Neural Networks (GNNs) makes it difficult to reuse already existing XAI methods. While other works have focused on instance-based explanation methods for GNNs, very few have investigated model-based methods and, to our knowledge, none have tried to probe the embedding of the GNNs for structural graph properties. In this paper we present a model agnostic explainability pipeline for Graph Neural Networks (GNNs) employing diagnostic classifiers. We propose to consider graph-theoretic properties as the features of choice for studying the emergence of representations in GNNs. This pipeline aims to probe and interpret the learned representations in GNNs across various architectures and datasets, refining our understanding and trust in these models.
Tom Pelletreau-Duris, Ruud van Bakel, Michael Cochez
NeSy3
2024 QAGCN: Answering Multi-relation Questions via Single-Step Implicit Reasoning over Knowledge Graphs
Ruijie Wang 0003, Luca Rossetto, Michael Cochez, Abraham Bernstein
ESWC (1)3
2024 Workshop on Deep Learning and Large Language Models for Knowledge Graphs (DL4KG)
abstract
The use of Knowledge Graphs (KGs) which constitute large networks of real-world entities and their interrelationships, has grown rapidly. A substantial body of research has emerged, exploring the integration of deep learning (DL) and large language models (LLMs) with KGs. This workshop aims to bring together leading researchers in the field to discuss and foster collaborations on the intersection of KG and DL/LLMs.
Mehwish Alam, Davide Buscaldi, Michael Cochez, Genet Asefa Gesese, Francesco Osborne, Diego Reforgiato Recupero
KDD3
2023 A Machine with Short-Term, Episodic, and Semantic Memory Systems
abstract
Inspired by the cognitive science theory of the explicit human memory systems, we have modeled an agent with short-term, episodic, and semantic memory systems, each of which is modeled with a knowledge graph. To evaluate this system and analyze the behavior of this agent, we designed and released our own reinforcement learning agent environment, “the Room”, where an agent has to learn how to encode, store, and retrieve memories to maximize its return by answering questions. We show that our deep Q-learning based agent successfully learns whether a short-term memory should be forgotten, or rather be stored in the episodic or semantic memory systems. Our experiments indicate that an agent with human-like memory systems can outperform an agent without this memory structure in the environment.
Taewoon Kim 0002, Michael Cochez, Vincent François-Lavet, Mark A. Neerincx, Piek Vossen
AAAI2
2023 Reasoning beyond Triples: Recent Advances in Knowledge Graph Embeddings
abstract
Knowledge Graphs (KGs) are a collection of facts describing entities connected by relationships. KG embeddings map entities and relations into a vector space while preserving their relational semantics. This enables effective inference of missing knowledge from the embedding space. Most KG embedding approaches focused on triple-shaped KGs. A great amount of real-world knowledge, however, cannot simply be represented by triples. In this tutorial, we give a systematic introduction to KG embeddings that go beyond the triple representation. In particular, the tutorial will focus on temporal facts where the triples are enriched with temporal information, hyper-relational facts where the triples are enriched with qualifiers, n-ary facts describing relationships between multiple entities, and also facts that are augmented with literal and text descriptions. During the tutorial, we will introduce both fundamental knowledge and advanced topics for understanding recent embedding approaches for beyond-triple representations.
Bo Xiong 0001, Mojtaba Nayyeri, Daniel Daza, Michael Cochez
CIKM4
2023 The Graph-Massivizer Approach Toward a European Sustainable Data Center Digital Twin
abstract
Modeling and understanding an expensive next-generation data center operating at a sustainable exascale performance remains a challenge yet to solve. The paper presents the approach taken by the Graph-Massivizer project, funded by the European Union, towards a sustainable data center, targeting a massive graph representation and analysis of its digital twin. We introduce five interoperable open-source tools that support this undertaking, creating an automated, sustainable loop of graph creation, analytics, optimization, sustainable resource management, and operation, emphasizing state-of-the-art progress. We plan to employ the tools for designing a massive data center graph, representing a digital twin describing spatial, semantic, and temporal relationships between the monitoring metrics, hardware nodes, cooling equipment, and jobs. The project aims to strengthen Bologna Technopole as a leading European supercomputing and big data hub offering sustainable green computing for improved societally relevant science throughput.
Martin Molan, Junaid Ahmed Khan, Andrea Bartolini, Roberta Turra, Giorgio Pedrazzi, Michael Cochez, Alexandru Iosup, Dumitru Roman, Joze M. Rozanec, Ana Lucia Varbanescu, Radu Prodan
COMPSAC6
2023 Adapting Neural Link Predictors for Data-Efficient Complex Query Answering
abstract
Answering complex queries on incomplete knowledge graphs is a challenging task where a model needs to answer complex logical queries in the presence of missing knowledge. Prior work in the literature has proposed to address this problem by designing architectures trained end-to-end for the complex query answering task with a reasoning process that is hard to interpret while requiring data and resource-intensive training. Other lines of research have proposed re-using simple neural link predictors to answer complex queries, reducing the amount of training data by orders of magnitude while providing interpretable answers. The neural link predictor used in such approaches is not explicitly optimised for the complex query answering task, implying that its scores are not calibrated to interact together. We propose to address these problems via CQD$^{\mathcal{A}}$, a parameter-efficient score \emph{adaptation} model optimised to re-calibrate neural link prediction scores for the complex query answering task. While the neural link predictor is frozen, the adaptation component -- which only increases the number of model parameters by $0.03\%$ -- is trained on the downstream complex query answering task. Furthermore, the calibration component enables us to support reasoning over queries that include atomic negations, which was previously impossible with link predictors. In our experiments, CQD$^{\mathcal{A}}$ produces significantly more accurate results than current state-of-the-art methods, improving from $34.4$ to $35.1$ Mean Reciprocal Rank values averaged across all datasets and query types while using $\leq 30\%$ of the available training query types. We further show that CQD$^{\mathcal{A}}$ is data-efficient, achieving competitive results with only $1\%$ of the complex training queries and robust in out-of-domain evaluations. Source code and datasets are available at https://github.com/EdinburghNLP/adaptive-cqd.
Erik Arakelyan, Pasquale Minervini, Daniel Daza, Michael Cochez, Isabelle Augenstein
NeurIPS4
2023 Explainable AI for Bioinformatics: Methods, Tools and Applications
abstract
Artificial intelligence (AI) systems utilizing deep neural networks and machine learning (ML) algorithms are widely used for solving critical problems in bioinformatics, biomedical informatics and precision medicine. However, complex ML models that are often perceived as opaque and black-box methods make it difficult to understand the reasoning behind their decisions. This lack of transparency can be a challenge for both end-users and decision-makers, as well as AI developers. In sensitive areas such as healthcare, explainability and accountability are not only desirable properties but also legally required for AI systems that can have a significant impact on human lives. Fairness is another growing concern, as algorithmic decisions should not show bias or discrimination towards certain groups or individuals based on sensitive attributes. Explainable AI (XAI) aims to overcome the opaqueness of black-box models and to provide transparency in how AI systems make decisions. Interpretable ML models can explain how they make predictions and identify factors that influence their outcomes. However, the majority of the state-of-the-art interpretable ML methods are domain-agnostic and have evolved from fields such as computer vision, automated reasoning or statistics, making direct application to bioinformatics problems challenging without customization and domain adaptation. In this paper, we discuss the importance of explainability and algorithmic transparency in the context of bioinformatics. We provide an overview of model-specific and model-agnostic interpretable ML methods and tools and outline their potential limitations. We discuss how existing interpretable ML methods can be customized and fit to bioinformatics research problems. Further, through case studies in bioimaging, cancer genomics and text mining, we demonstrate how XAI methods can improve transparency and decision fairness. Our review aims at providing valuable insights and serving as a starting point for researchers wanting to enhance explainability and decision transparency while solving bioinformatics problems. GitHub: https://github.com/rezacsedu/XAI-for-bioinformatics.
Md. Rezaul Karim 0001, Tanhim Islam, Md Shajalal, Oya Beyan, Christoph Lange 0002, Michael Cochez, Dietrich Rebholz-Schuhmann, Stefan Decker
Briefings Bioinform.6
2023 A knowledge graph approach to predict and interpret disease-causing gene interactions
abstract
BACKGROUND: Understanding the impact of gene interactions on disease phenotypes is increasingly recognised as a crucial aspect of genetic disease research. This trend is reflected by the growing amount of clinical research on oligogenic diseases, where disease manifestations are influenced by combinations of variants on a few specific genes. Although statistical machine-learning methods have been developed to identify relevant genetic variant or gene combinations associated with oligogenic diseases, they rely on abstract features and black-box models, posing challenges to interpretability for medical experts and impeding their ability to comprehend and validate predictions. In this work, we present a novel, interpretable predictive approach based on a knowledge graph that not only provides accurate predictions of disease-causing gene interactions but also offers explanations for these results. RESULTS: We introduce BOCK, a knowledge graph constructed to explore disease-causing genetic interactions, integrating curated information on oligogenic diseases from clinical cases with relevant biomedical networks and ontologies. Using this graph, we developed a novel predictive framework based on heterogenous paths connecting gene pairs. This method trains an interpretable decision set model that not only accurately predicts pathogenic gene interactions, but also unveils the patterns associated with these diseases. A unique aspect of our approach is its ability to offer, along with each positive prediction, explanations in the form of subgraphs, revealing the specific entities and relationships that led to each pathogenic prediction. CONCLUSION: Our method, built with interpretability in mind, leverages heterogenous path information in knowledge graphs to predict pathogenic gene interactions and generate meaningful explanations. This not only broadens our understanding of the molecular mechanisms underlying oligogenic diseases, but also presents a novel application of knowledge graphs in creating more transparent and insightful predictors for genetic research.
Alexandre Renaux, Chloé Terwagne, Michael Cochez, Ilaria Tiddi, Ann Nowé, Tom Lenaerts
BMC Bioinform.3
2022 Query Embedding on Hyper-Relational Knowledge Graphs
Dimitrios Alivanistos, Max Berrendorf, Michael Cochez, Michael Galkin
ICLR3
2022 Complex Query Answering with Neural Link Predictors (Extended Abstract)
abstract
Neural link predictors are useful for identifying missing edges in large scale Knowledge Graphs. However, it is still not clear how to use these models for answering more complex queries containing logical conjunctions (∧), disjunctions (∨), and existential quantifiers (∃). We propose a framework for efficiently answering complex queries on in- complete Knowledge Graphs. We translate each query into an end-to-end differentiable objective, where the truth value of each atom is computed by a pre-trained neural link predictor. We then analyse two solutions to the optimisation problem, including gradient-based and combinatorial search. In our experiments, the proposed approach produces more accurate results than state-of-the-art methods — black-box models trained on millions of generated queries — without the need for training on a large and diverse set of complex queries. Using orders of magnitude less training data, we obtain relative improvements ranging from 8% up to 40% in Hits@3 across multiple knowledge graphs. We find that it is possible to explain the outcome of our model in terms of the intermediate solutions identified for each of the complex query atoms. All our source code and datasets are available online (https://github.com/uclnlp/cqd).
Pasquale Minervini, Erik Arakelyan, Daniel Daza, Michael Cochez
IJCAI4
2022 Scientific Item Recommendation Using a Citation Network
Xu Wang 0032, Frank van Harmelen, Michael Cochez, Zhisheng Huang
KSEM (2)3
2022 Hyperbolic Embedding Inference for Structured Multi-Label Prediction
abstract
We consider a structured multi-label prediction problem where the labels are organized under implication and mutual exclusion constraints. A major concern is to produce predictions that are logically consistent with these constraints. To do so, we formulate this problem as an embedding inference problem where the constraints are imposed onto the embeddings of labels by geometric construction. Particularly, we consider a hyperbolic Poincaré ball model in which we encode labels as Poincaré hyperplanes that work as linear decision boundaries. The hyperplanes are interpreted as convex regions such that the logical relationships (implication and exclusion) are geometrically encoded using the insideness and disjointedness of these regions, respectively. We show theoretical groundings of the method for preserving logical relationships in the embedding space. Extensive experiments on 12 datasets show 1) significant improvements in mean average precision; 2) lower number of constraint violations; 3) an order of magnitude fewer dimensions than baselines.
Bo Xiong 0001, Michael Cochez, Mojtaba Nayyeri, Steffen Staab
NeurIPS2
2022 Convolutional Embedded Networks for Population Scale Clustering and Bio-Ancestry Inferencing
abstract
The study of genetic variants (GVs) can help find correlating population groups and to identify cohorts that are predisposed to common diseases and explain differences in disease susceptibility and how patients react to drugs. Machine learning techniques are increasingly being applied to identify interacting GVs to understand their complex phenotypic traits. Since the performance of a learning algorithm not only depends on the size and nature of the data but also on the quality of underlying representation, deep neural networks (DNNs) can learn non-linear mappings that allow transforming GVs data into more clustering and classification friendly representations than manual feature selection. In this paper, we propose convolutional embedded networks (CEN) in which we combine two DNN architectures called convolutional embedded clustering (CEC) and convolutional autoencoder (CAE) classifier for clustering individuals and predicting geographic ethnicity based on GVs, respectively. We employed CAE-based representation learning to 95 million GVs from the '1000 genomes' (covering 2,504 individuals from 26 ethnic origins) and 'Simons genome diversity' (covering 279 individuals from 130 ethnic origins) projects. Quantitative and qualitative analyses with a focus on accuracy and scalability show that our approach outperforms state-of-the-art approaches such as VariantSpark and ADMIXTURE. In particular, CEC can cluster targeted population groups in 22 hours with an adjusted rand index (ARI) of 0.915, the normalized mutual information (NMI) of 0.92, and the clustering accuracy (ACC) of 89 percent. Contrarily, the CAE classifier can predict the geographic ethnicity of unknown samples with an F1 and Mathews correlation coefficient (MCC) score of 0.9004 and 0.8245, respectively. Further, to provide interpretations of the predictions, we identify significant biomarkers using gradient boosted trees (GBT) and SHapley Additive exPlanations (SHAP). Overall, our approach is transparent and faster than the baseline methods, and scalable for 5 to 100 percent of the full human genome.
Md. Rezaul Karim 0001, Michael Cochez, Achille Zappa, Ratnesh Sahay, Dietrich Rebholz-Schuhmann, Oya Beyan, Stefan Decker
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Unsupervised Feature Selection for Efficient Exploration of High Dimensional Data
Arnab Chakrabarti, Abhijeet Das, Michael Cochez, Christoph Quix
ADBIS3
2021 Complex Query Answering with Neural Link Predictors
Erik Arakelyan, Daniel Daza, Pasquale Minervini, Michael Cochez
ICLR4
2021 Inductive Entity Representations from Text via Link Prediction
abstract
Knowledge Graphs (KG) are of vital importance for multiple applications on the web, including information retrieval, recommender systems, and metadata annotation.
Daniel Daza, Michael Cochez, Paul Groth
WWW2
2021 Deep learning-based clustering approaches for bioinformatics
abstract
Clustering is central to many data-driven bioinformatics research and serves a powerful computational method. In particular, clustering helps at analyzing unstructured and high-dimensional data in the form of sequences, expressions, texts and images. Further, clustering is used to gain insights into biological processes in the genomics level, e.g. clustering of gene expressions provides insights on the natural structure inherent in the data, understanding gene functions, cellular processes, subtypes of cells and understanding gene regulations. Subsequently, clustering approaches, including hierarchical, centroid-based, distribution-based, density-based and self-organizing maps, have long been studied and used in classical machine learning settings. In contrast, deep learning (DL)-based representation and feature learning for clustering have not been reviewed and employed extensively. Since the quality of clustering is not only dependent on the distribution of data points but also on the learned representation, deep neural networks can be effective means to transform mappings from a high-dimensional data space into a lower-dimensional feature space, leading to improved clustering results. In this paper, we review state-of-the-art DL-based approaches for cluster analysis that are based on representation learning, which we hope to be useful, particularly for bioinformatics research. Further, we explore in detail the training procedures of DL-based clustering algorithms, point out different clustering quality metrics and evaluate several DL-based approaches on three bioinformatics use cases, including bioimaging, cancer genomics and biomedical text mining. We believe this review and the evaluation results will provide valuable insights and serve a starting point for researchers wanting to apply DL-based unsupervised methods to solve emerging bioinformatics research problems.
Md. Rezaul Karim 0001, Oya Beyan, Achille Zappa, Ivan G. Costa, Dietrich Rebholz-Schuhmann, Michael Cochez, Stefan Decker
Briefings Bioinform.6
2020 DeepCOVIDExplainer: Explainable COVID-19 Diagnosis from Chest X-ray Images
abstract
In this paper1, we proposed an explainable deep neural networks (DNN)-based method for automatic detection of COVID-19 symptoms from chest radiography (CXR) images, which we call ‘DeepCOVIDExplainer’. We used 15,959 CXR images of 15,854 patients, covering normal, pneumonia, and COVID-19 cases. CXR images are first comprehensively preprocessed and augmented before classifying with a neural ensemble method, followed by highlighting class-discriminating regions using gradient-guided class activation maps (Grad-CAM ++) and layer-wise relevance propagation (LRP). Further, we provide human-interpretable explanations for the diagnosis. Evaluation results show that our approach can identify COVID-19 cases with a positive predictive value (PPV) of 91.6%, 92.45%, and 96.12%, respectively for normal, pneumonia, and COVID-19 cases, respectively, outperforming recent approaches.1Read longer version of this paper: https://arxiv.org/pdf/2004.04582.pdf
Md. Rezaul Karim 0001, Till Döhmen, Michael Cochez, Oya Beyan, Dietrich Rebholz-Schuhmann, Stefan Decker
BIBM3
2020 Classification Benchmarks for Under-resourced Bengali Language based on Multichannel Convolutional-LSTM Network
abstract
Exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices, but also enables people to express anti-social behavior like online harassment, cyberbul-lying, and hate speech. Numerous works have been proposed to utilize these data for social and anti-social behavior analysis, document characterization, and sentiment analysis by predicting the contexts mostly for highly resourced languages like English. However, some languages are under-resources, e.g., South Asian languages like Bengali, Tamil, Assamese, Malayalam, that lack of computational resources for natural language processing. In this paper1, we provide several classification benchmarks for Bengali, an under-resourced language. We prepared three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis. We built the largest Bengali word embedding models to date based on 250 million articles, which we call BengFastText. We perform three experiments, covering document classification, sentiment analysis, and hate speech detection. We incorporate word embeddings into a Multichannel Convolutional-LSTM (MC-LSTM) network for predicting different types of hate speech, document classification, and sentiment analysis. Experiments demonstrate that BengFastText can capture the semantics of words from respective contexts correctly. Evaluations against several baseline embedding models, e.g., Word2Vec and GloVe yield up to 92.30%, 82.25%, and 90.45% F1-scores in case of document classification, sentiment analysis, and hate speech detection, respectively during 5-fold cross-validation tests.
Md. Rezaul Karim 0001, Bharathi Raja Chakravarthi, John P. McCrae, Michael Cochez
DSAA4
2020 SchemaTree: Maximum-Likelihood Property Recommendation for Wikidata
Lars Christoph Gleim, Rafael Schimassek, Dominik Hüser, Maximilian Peters, Christoph Krämer, Michael Cochez, Stefan Decker
ESWC6
2020 GEval: A Modular and Extensible Evaluation Framework for Graph Embedding Techniques
Maria Angela Pellegrino, Abdulrahman Altabba, Martina Garofalo, Petar Ristoski, Michael Cochez
ESWC5
2020 Structured query construction via knowledge graph embedding
Ruijie Wang 0003, Meng Wang 0009, Jun Liu 0002, Michael Cochez, Stefan Decker
Knowl. Inf. Syst.4
2019 OncoNetExplainer: Explainable Predictions of Cancer Types Based on Gene Expression Data
abstract
The discovery of important biomarkers is a significant step towards understanding the molecular mechanisms of carcinogenesis; enabling accurate diagnosis for, and prognosis of, a certain cancer type. Before recommending any diagnosis, genomics data such as gene expressions (GE) and clinical outcomes need to be analyzed. However, complex nature, high dimensionality, and heterogeneity in genomics data make the overall analysis challenging. Convolutional neural networks (CNN) have shown tremendous success in solving such problems. However, neural network models are perceived mostly as 'black box' methods because of their not well-understood internal functioning. However, interpretability is important to provide insights on why a given cancer case has a certain type. Besides, finding the most important biomarkers can help in recommending more accurate treatments and drug repositioning. Moreover, the 'right to explanation' of the EU GDPR gives patients the right to know why and how an algorithm made a diagnosis decision. Hence, in this paper, we propose a new approach called OncoNetExplainer to make explainable predictions of cancer types based on GE data. We used genomics data about 9,074 cancer patients covering 33 different cancer types from the Pan-Cancer Atlas on which we trained CNN and VGG16 networks using guided-gradient class activation maps++ (GradCAM++). Further, we generate class-specific heat maps to identify significant biomarkers and computed feature importance in terms of mean absolute impact to rank top genes across all the cancer types. Quantitative and qualitative analyses show that both models exhibit high confidence at predicting the cancer types correctly giving an average precision of 96.25%. To provide comparisons with the baselines, we identified top genes, and cancer-specific driver genes using gradient boosted trees and SHapley Additive exPlanations (SHAP). Finally, our findings were validated with the annotations provided by the TumorPortal.
Md. Rezaul Karim 0001, Michael Cochez, Oya Beyan, Stefan Decker, Christoph Lange 0002
BIBE2
2019 Message Passing for Complex Question Answering over Knowledge Graphs
abstract
Question answering over knowledge graphs (KGQA) has evolved from simple single-fact questions to complex questions that require graph traversal and aggregation. We propose a novel approach for complex KGQA that uses unsupervised message passing, which propagates confidence scores obtained by parsing an input question and matching terms in the knowledge graph to a set of possible answers. First, we identify entity, relationship, and class names mentioned in a natural language question, and map these to their counterparts in the graph. Then, the confidence scores of these mappings propagate through the graph structure to locate the answer entities. Finally, these are aggregated depending on the identified question type. This approach can be efficiently implemented as a series of sparse matrix multiplications mimicking joins over small local subgraphs. Our evaluation results show that the proposed approach outperforms the state of the art on the LC-QuAD benchmark. Moreover, we show that the performance of the approach depends only on the quality of the question interpretation results, i.e., given a correct relevance score distribution, our approach always produces a correct answer ranking. Our error analysis reveals correct answers missing from the benchmark dataset and inconsistencies in the DBpedia knowledge graph. Finally, we provide a comprehensive evaluation of the proposed approach accompanied with an ablation study and an error analysis, which showcase the pitfalls for each of the question answering components in more detail.
Svitlana Vakulenko, Javier D. Fernández, Axel Polleres, Maarten de Rijke, Michael Cochez
CIKM5
2019 Leveraging Knowledge Graph Embeddings for Natural Language Question Answering
Ruijie Wang 0003, Meng Wang 0009, Jun Liu 0002, Weitong Chen 0001, Michael Cochez, Stefan Decker
DASFAA (1)5
2018 Measuring Semantic Coherence of a Conversation
Svitlana Vakulenko, Maarten de Rijke, Michael Cochez, Vadim Savenkov, Axel Polleres
ISWC (1)3
2018 Mining maximal frequent patterns in transactional databases and dynamic data streams: A spark-based approach
Md. Rezaul Karim 0001, Michael Cochez, Oya Beyan, Chowdhury Farhan Ahmed, Stefan Decker
Inf. Sci.2
2017 The Future of the Semantic Web: Prototypes on a Global Distributed Filesystem
abstract
An important part of the Semantic Web vision is the idea that data is shared seamlessly and that world wide distributed, accessible, and interlinked knowledge bases can be created. However, the current incarnation of the Semantic Web falls short of this vision: while some necessary infrastructure (e.g., Linked Data) has been put in place, the current use of Linked Data in the Semantic Web is still happening in data silos, and sharing and reusing of knowledge is cumbersome and not straightforward. Recently the idea of prototypical objects was proposed to remedy this situation. This concept, known as Prototypes originates from early Frame systems and is also adopted in programming languages such as Javascript. In this vision paper we describe how a distributed file system forms a natural habitat for prototype knowledge representation, advancing the Semantic Web. In particular, we describe how we envision the deployment of Linked Data and Prototype Knowledge bases atop of the InterPlanetary File System (IPFS), which has several useful features matching the needs for knowledge representation based on prototypes.
Michael Cochez, Dominik Hüser, Stefan Decker
ICDCS1
2017 Global RDF Vector Space Embeddings
Michael Cochez, Petar Ristoski, Simone Paolo Ponzetto, Heiko Paulheim
ISWC (1)1
2016 Knowledge Representation on the Web Revisited: The Case for Prototypes
Michael Cochez, Stefan Decker, Eric Prud'hommeaux
ISWC (1)1
2015 Twister Tries: Approximate Hierarchical Agglomerative Clustering for Average Distance in Linear Time
abstract
Many commonly used data-mining techniques utilized across research fields perform poorly when used for large data sets. Sequential agglomerative hierarchical non-overlapping clustering is one technique for which the algorithms' scaling properties prohibit clustering of a large amount of items. Besides the unfavorable time complexity of O(n2), these algorithms have a space complexity of O(n2), which can be reduced to O(n) if the time complexity is allowed to rise to O(n2 log2n). In this paper, we propose the use of locality-sensitive hashing combined with a novel data structure called twister tries to provide an approximate clustering for average linkage. Our approach requires only linear space. Furthermore, its time complexity is linear in the number of items to be clustered, making it feasible to apply it on a larger scale. We evaluate the approach both analytically and by applying it to several data sets.
Michael Cochez, Hao Mou
SIGMOD Conference1
2015 Anomaly Detection Algorithms for the Sleeping Cell Detection in LTE Networks
abstract
The Sleeping Cell problem is a particular type of cell degradation in Long-Term Evolution (LTE) networks. In practice such cell outage leads to the lack of network service and sometimes it can be revealed only after multiple user complains by an operator. In this study a cell becomes sleeping because of a Random Access Channel (RACH) failure, which may happen due to software or hardware problems. For the detection of malfunctioning cells, we introduce a data mining based framework. In its core is the analysis of event sequences reported by a User Equipment (UE) to a serving Base Station (BS). The crucial element of the developed framework is an anomaly detection algorithm. We compare performances of distance, centroid distance and probabilistic based methods, using Receiver Operating Characteristic (ROC) and Precision-Recall curves. Moreover, the theoretical comparison of the methods' computational efficiencies is provided. The sleeping cell detection framework is verified by means of a dynamic LTE system simulator, using Minimization of Drive Testing (MDT) functionality. It is shown that the sleeping cell can be pinpointed.
Sergey Chernov 0003, Michael Cochez, Tapani Ristaniemi
VTC Spring2
2013 Issues with a course that emphasizes self-direction
abstract
In this paper, we examine a master's level course that emphasizes self-direction on the part of students. The course is run by weekly group assignments and requires independent work such that only one mandatory classroom session is arranged each week. Our specific research interests are how students responded to the setting of this kind and whether they demonstrated self-direction during the course. We surveyed the students' view of the course, their group work experience, and their study habits, and analyzed the resultant survey data for themes. The results suggest that while the pass rate was considerably high and the course was regarded as well-organized by the students, there were several concerns related to whether we could prompt self-directed study habits.
Ville Isomöttönen, Ville Tirronen, Michael Cochez
ITiCSE3