EDBT 2026 Demo / reviewers in the wild / expert
Vít Novácek
dblp:79/1204
· DBLP profile ↗
29ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-6578-5449ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large language model vs. traditional machine learning: Evaluating predictive models for early detection of tumor relapseabstractIn this study, we evaluate the effectiveness of foundational artificial intelligence (AI) models, particularly large language models (LLMs), in comparison to traditional machine learning methods for predicting tumor relapse in patients with non-small-cell lung cancer (NSCLC). With a high recurrence risk in NSCLC, early and accurate prediction is essential for improving patient outcomes and guiding treatment decisions. Our analysis utilizes a dataset of 1,348 patients, examining the performance of traditional machine learning models such as Random Forest, alongside cutting-edge LLMs like Mistral-7B, LLaMA-7B, Falcon-7B, and GPT-based models. While the Random Forest model slightly outperforms Mistral-7B in precision–recall for relapse prediction, the comparable results suggest that both approaches offer valuable insights for early relapse detection. This study underscores the potential of integrating classical machine learning with foundational AI models to enhance predictive accuracy in cancer prognosis, providing pathways for more personalized medical interventions. • Study contrasts AI and traditional methods for NSCLC relapse prediction. • Early, accurate predictions are crucial for NSCLC patient care. • Analysis includes 1,348 NSCLC patient data points. • Random Forest edges out Mistral-7B in precision–recall. • Both models show potential for early detection of lung cancer recurrence. Mohan Timilsina, Samuele Buosi, Maria Torrente, Mariano Provencio, Manuel Cobo, Delvys Rodriguez Abreu, Rafael Castro, Enric Carcereny, Edward Curry, Vít Novácek |
Expert Syst. Appl. | 10 |
| 2024 | Machine learning estimated probability of relapse in early-stage non-small-cell lung cancer patients with aneuploidy imputation scores and knowledge graph embeddingsabstractLow-stage lung cancer is known to recur unpredictably, and patients receiving various treatment methods like radiation, chemotherapy, and immunotherapies have been seen to respond very differently. Identifying a priori if a patient is going to relapse or not could make a difference in terms of saving lives and personalized care offered. In this work, we provide an answer to the following research question: Is it possible to enhance the machine learning (ML) of the estimated probability of relapse in early-stage non-small-cell lung cancer (NSCLC) patients with aneuploidy imputation scores? To predict recurrence in 1,348 early-stage (I-II) NSCLC patients, we train graph ML models utilizing the Spanish pulmonary cancer group knowledge graph enriched with triples from pathway imputation. ML models trained on Knowledge graph data enriched with triples from pathway score imputation present an 82% Precision and 91% Specificity in predicting relapse over 200 patients from a held-out test set. ML models trained using graphs data could prove useful supplemental tool in the TNM classification systems and improve a lung cancer patient’s prognosis. Samuele Buosi, Mohan Timilsina, Adrianna Janik, Luca Costabello, Maria Torrente, Mariano Provencio, Dirk Fey, Vít Novácek |
Expert Syst. Appl. | 8 |
| 2023 | Unsupervised extraction, classification and visualization of clinical note segments using the MIMIC-III datasetabstractThis paper presents a text-mining approach to extracting and organizing segments from unstructured clinical notes in an unsupervised way. Our work is motivated by the real challenge of poor semantic integration between clinical notes produced by different doctors, departments, or hospitals. This can lead to clinicians overlooking important information, especially for patients with long and varied medical histories. This work extends a previous approach developed for Czech breast cancer patients and validates it on the publicly accessible MIMIC-III English dataset, demonstrating its universal and language-independent applicability. Our work is a stepping stone to a broad array of downstream tasks, such as summarizing or integrating patient records, extracting structured information, or computing patient embeddings. Additionally, the paper presents a clustering analysis of the latent space of note segment types, using hierarchical clustering and an interactive treemap visualization. The presented results demonstrate that this approach generalizes well for MIMIC and English. Petr Zelina, Jana Halámková, Vít Novácek |
BIBM | 3 |
| 2023 | Machine Learning Survival Models for Relapse Prediction in a Early Stage Lung Cancer PatientabstractLung cancer is one of the leading health complications causing high mortality worldwide. The relapsing behavior of medically treated early-stage lung cancer makes this disease even more complicated. Thus predicting such relapse using a data-centric approach provides a complementary perspective for clinicians to understand the disease. In this preliminary work, we explored off-the-shelf survival models to predict the relapse of early-stage lung cancer patients. We analyzed the survival models on a cohort of 1348 early-stage non-small cell lung cancer (NSCLC) patients in different timestamps. Using the prediction explanation model SHAP (SHapley Additive exPlanations), we further explained the best-performing survival model's predictions. Our explainable predictive model is a potential tool for oncologists that address an unmet clinical need for post-treatment patient stratification based on the relapse hazard. Mohan Timilsina, Samuele Buosi, Adrianna Janik, Pasquale Minervini, Luca Costabello, Maria Torrente, Mariano Provencio, Virginia Calvo, Carlos Camps, Ana L. Ortega, Bartomeu Massutí, M. Rosario Garcia Campelo, Edel del Barco, Joaquim Bosch-Barrera, Vít Novácek |
IJCNN | 15 |
| 2023 | Synergy between imputed genetic pathway and clinical information for predicting recurrence in early stage non-small cell lung cancerabstractOBJECTIVE: Lung cancer exhibits unpredictable recurrence in low-stage tumors and variable responses to different therapeutic interventions. Predicting relapse in early-stage lung cancer can facilitate precision medicine and improve patient survivability. While existing machine learning models rely on clinical data, incorporating genomic information could enhance their efficiency. This study aims to impute and integrate specific types of genomic data with clinical data to improve the accuracy of machine learning models for predicting relapse in early-stage, non-small cell lung cancer patients. METHODS: The study utilized a publicly available TCGA lung cancer cohort and imputed genetic pathway scores into the Spanish Lung Cancer Group (SLCG) data, specifically in 1348 early-stage patients. Initially, tumor recurrence was predicted without imputed pathway scores. Subsequently, the SLCG data were augmented with pathway scores imputed from TCGA. The integrative approach aimed to enhance relapse risk prediction performance. RESULTS: The integrative approach achieved improved relapse risk prediction with the following evaluation metrics: an area under the precision-recall curve (PR-AUC) score of 0.75, an area under the ROC (ROC-AUC) score of 0.80, an F1 score of 0.61, and a Precision of 0.80. The prediction explanation model SHAP (SHapley Additive exPlanations) was employed to explain the machine learning model's predictions. CONCLUSION: We conclude that our explainable predictive model is a promising tool for oncologists that addresses an unmet clinical need of post-treatment patient stratification based on the relapse risk while also improving the predictive power by incorporating proxy genomic data not available for specific patients. Mohan Timilsina, Dirk Fey, Samuele Buosi, Adrianna Janik, Luca Costabello, Enric Carcereny, Delvys Rodriguez Abreu, Manuel Cobo, Rafael Castro, Reyes Bernabé, Pasquale Minervini, Maria Torrente, Mariano Provencio, Vít Novácek |
J. Biomed. Informatics | 14 |
| 2022 | Integration of Clinical Information and Imputed Aneuploidy Scores to Enhance Relapse Prediction in Early Stage Lung Cancer Patients
Mohan Timilsina, Samuele Bousi, Dirk Fey, Adrianna Janik, Maria Torrente, Mariano Provencio, Alberto Bermúdez, Enric Carcereny, Luca Costabello, Delvys Rodriguez Abreu, Manuel Cobo, Rafael Castro, Reyes Bernabé, Maria Guirado, Pasquale Minervini, Vít Novácek |
AMIA | 16 |
| 2022 | Unsupervised extraction, labelling and clustering of segments from clinical notesabstractThis work is motivated by the scarcity of tools for accurate, unsupervised information extraction from unstructured clinical notes in computationally underrepresented languages, such as Czech. We introduce a stepping stone to a broad array of downstream tasks such as summarisation or integration of individual patient records, extraction of structured information for national cancer registry reporting or building of semi-structured semantic patient representations for computing patient embeddings. More specifically, we present a method for unsupervised extraction of semantically-labelled textual segments from clinical notes and test it out on a dataset of Czech breast cancer patients, provided by Masaryk Memorial Cancer Institute (the largest Czech hospital specialising in oncology). Our goal was to extract, classify (i.e. label) and cluster segments of the free-text notes that correspond to specific clinical features (e.g., family background, comorbidities or toxicities). The presented results demonstrate the practical relevance of the proposed approach for building more sophisticated extraction and analytical pipelines deployed on Czech clinical notes. Petr Zelina, Jana Halámková, Vít Novácek |
BIBM | 3 |
| 2022 | Boundary heat diffusion classifier for a semi-supervised learning in a multilayer network embeddingabstractThe scarcity of high-quality annotations in many application scenarios has recently led to an increasing interest in devising learning techniques that combine unlabeled data with labeled data in a network. In this work, we focus on the label propagation problem in multilayer networks. Our approach is inspired by the heat diffusion model, which shows usefulness in machine learning problems such as classification and dimensionality reduction. We propose a novel boundary-based heat diffusion algorithm that guarantees a closed-form solution with an efficient implementation. We experimentally validated our method on synthetic networks and five real-world multilayer network datasets representing scientific coauthorship, spreading drug adoption among physicians, two bibliographic networks, and a movie network. The results demonstrate the benefits of the proposed algorithm, where our boundary-based heat diffusion dominates the performance of the state-of-the-art methods. Mohan Timilsina, Vít Novácek, Mathieu d'Aquin, Haixuan Yang |
Neural Networks | 2 |
| 2021 | On Predicting Recurrence in Early Stage Non-small Cell Lung Cancer
Sameh K. Mohamed, Brian Walsh, Mohan Timilsina, Vít Novácek, Maria Torrente, Fabio Franco, Mariano Provencio, Adrianna Janik, Luca Costabello, Pontus Stenetorp, Pasquale Minervini |
AMIA | 4 |
| 2021 | Biological applications of knowledge graph embedding modelsabstractComplex biological systems are traditionally modelled as graphs of interconnected biological entities. These graphs, i.e. biological knowledge graphs, are then processed using graph exploratory approaches to perform different types of analytical and predictive tasks. Despite the high predictive accuracy of these approaches, they have limited scalability due to their dependency on time-consuming path exploratory procedures. In recent years, owing to the rapid advances of computational technologies, new approaches for modelling graphs and mining them with high accuracy and scalability have emerged. These approaches, i.e. knowledge graph embedding (KGE) models, operate by learning low-rank vector representations of graph nodes and edges that preserve the graph's inherent structure. These approaches were used to analyse knowledge graphs from different domains where they showed superior performance and accuracy compared to previous graph exploratory approaches. In this work, we study this class of models in the context of biological knowledge graphs and their different applications. We then show how KGE models can be a natural fit for representing complex biological knowledge modelled as graphs. We also discuss their predictive and analytical capabilities in different biology applications. In this regard, we present two example case studies that demonstrate the capabilities of KGE models: prediction of drug-target interactions and polypharmacy side effects. Finally, we analyse different practical considerations for KGEs, and we discuss possible opportunities and challenges related to adopting them for modelling biological systems. Sameh K. Mohamed, Aayah Nounu, Vít Novácek |
Briefings Bioinform. | 3 |
| 2020 | BioKG: A Knowledge Graph for Relational Learning On Biological DataabstractKnowledge graphs became a popular means for modeling complex biological systems where they model the interactions between biological entities and their effects on the biological system. They also provide support for relational learning models which are known to provide highly scalable and accurate predictions of associations between biological entities. Despite the success of the combination of biological knowledge graph and relation learning models in biological predictive tasks, there is a lack of unified biological knowledge graph resources. This forced all current efforts and studies for applying a relational learning model on biological data to compile and build biological knowledge graphs from open biological databases. This process is often performed inconsistently across such efforts, especially in terms of choosing the original resources, aligning identifiers of the different databases, and assessing the quality of included data. To make relational learning on biomedical data more standardised and reproducible, we propose a new biological knowledge graph which provides a compilation of curated relational data from open biological databases in a unified format with common, interlinked identifiers. We also provide a new module for mapping identifiers and labels from different databases which can be used to align our knowledge graph with biological data from other heterogeneous sources. Finally, to illustrate the practical relevance of our work, we provide a set of benchmarks based on the presented data that can be used to train and assess the relational learning models in various tasks related to pathway and drug discovery. Brian Walsh, Sameh K. Mohamed, Vít Novácek |
CIKM | 3 |
| 2020 | Discovering protein drug targets using knowledge graph embeddingsabstractMOTIVATION: Computational approaches for predicting drug-target interactions (DTIs) can provide valuable insights into the drug mechanism of action. DTI predictions can help to quickly identify new promising (on-target) or unintended (off-target) effects of drugs. However, existing models face several challenges. Many can only process a limited number of drugs and/or have poor proteome coverage. The current approaches also often suffer from high false positive prediction rates. RESULTS: We propose a novel computational approach for predicting drug target proteins. The approach is based on formulating the problem as a link prediction in knowledge graphs (robust, machine-readable representations of networked knowledge). We use biomedical knowledge bases to create a knowledge graph of entities connected to both drugs and their potential targets. We propose a specific knowledge graph embedding model, TriModel, to learn vector representations (i.e. embeddings) for all drugs and targets in the created knowledge graph. These representations are consequently used to infer candidate drug target interactions based on their scores computed by the trained TriModel model. We have experimentally evaluated our method using computer simulations and compared it to five existing models. This has shown that our approach outperforms all previous ones in terms of both area under ROC and precision-recall curves in standard benchmark tests. AVAILABILITY AND IMPLEMENTATION: The data, predictions and models are available at: drugtargets.insight-centre.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sameh K. Mohamed, Vít Novácek, Aayah Nounu |
Bioinform. | 2 |
| 2020 | Accurate prediction of kinase-substrate networks using knowledge graphsabstractPhosphorylation of specific substrates by protein kinases is a key control mechanism for vital cell-fate decisions and other cellular processes. However, discovering specific kinase-substrate relationships is time-consuming and often rather serendipitous. Computational predictions alleviate these challenges, but the current approaches suffer from limitations like restricted kinome coverage and inaccuracy. They also typically utilise only local features without reflecting broader interaction context. To address these limitations, we have developed an alternative predictive model. It uses statistical relational learning on top of phosphorylation networks interpreted as knowledge graphs, a simple yet robust model for representing networked knowledge. Compared to a representative selection of six existing systems, our model has the highest kinome coverage and produces biologically valid high-confidence predictions not possible with the other tools. Specifically, we have experimentally validated predictions of previously unknown phosphorylations by the LATS1, AKT1, PKA and MST2 kinases in human. Thus, our tool is useful for focusing phosphoproteomic experiments, and facilitates the discovery of new phosphorylation reactions. Our model can be accessed publicly via an easy-to-use web interface (LinkPhinder). Vít Novácek, Gavin McGauran, David Matallanas, Adrián Vallejo Blanco, Piero Conca, Emir Muñoz, Luca Costabello, Kamalesh Kanakaraj, Muhammad Zeeshan Nawaz, Brian Walsh, Sameh K. Mohamed, Pierre-Yves Vandenbussche, Colm J. Ryan, Walter Kolch, Dirk Fey |
PLoS Comput. Biol. | 1 |
| 2019 | Link Prediction Using Multi Part EmbeddingsabstractKnowledge graph embeddings models are widely used to provide scalable and efficient link prediction for knowledge graphs. They use different techniques to model embeddings interactions, where their tensor factorisation based versions are known to provide state-of-the-art results. In recent works, developments on factorisation based knowledge graph embedding models were mostly limited to enhancing the ComplEx and the DistMult models, as they can efficiently provide predictions within linear time and space complexity. In this work, we aim to extend the works of the ComplEx and the DistMult models by proposing a new factorisation model, TriModel, which uses three part embeddings to model a combination of symmetric and asymmetric interactions between embeddings. We perform an empirical evaluation for the TriModel model compared to other tensor factorisation models on different training configurations (loss functions and regularisation terms), and we show that the TriModel model provides the state-of-the-art results in all configurations. In our experiments, we use standard benchmarking datasets (WN18, WN18RR, FB15k, FB15k-237, YAGO10) along with a new NELL based benchmarking dataset (NELL239) that we have developed. Sameh K. Mohamed, Vít Novácek |
ESWC | 2 |
| 2019 | Facilitating prediction of adverse drug reactions by using knowledge graphs and multi-label learning modelsabstractTimely identification of adverse drug reactions (ADRs) is highly important in the domains of public health and pharmacology. Early discovery of potential ADRs can limit their effect on patient lives and also make drug development pipelines more robust and efficient. Reliable in silico prediction of ADRs can be helpful in this context, and thus, it has been intensely studied. Recent works achieved promising results using machine learning. The presented work focuses on machine learning methods that use drug profiles for making predictions and use features from multiple data sources. We argue that despite promising results, existing works have limitations, especially regarding flexibility in experimenting with different data sets and/or predictive models. We suggest to address these limitations by generalization of the key principles used by the state of the art. Namely, we explore effects of: (1) using knowledge graphs-machine-readable interlinked representations of biomedical knowledge-as a convenient uniform representation of heterogeneous data; and (2) casting ADR prediction as a multi-label ranking problem. We present a specific way of using knowledge graphs to generate different feature sets and demonstrate favourable performance of selected off-the-shelf multi-label learning models in comparison with existing works. Our experiments suggest better suitability of certain multi-label learning methods for applications where ranking is preferred. The presented approach can be easily extended to other feature sources or machine learning methods, making it flexible for experiments tuned toward specific requirements of end users. Our work also provides a clearly defined and reproducible baseline for any future related experiments. Emir Muñoz, Vít Novácek, Pierre-Yves Vandenbussche |
Briefings Bioinform. | 2 |
| 2017 | Identifying Equivalent Relation Paths in Knowledge Graphs
Sameh K. Mohamed, Emir Muñoz, Vít Novácek, Pierre-Yves Vandenbussche |
LDK | 3 |
| 2017 | Regularizing Knowledge Graph Embeddings via Equivalence and Inversion Axioms
Pasquale Minervini, Luca Costabello, Emir Muñoz, Vít Novácek, Pierre-Yves Vandenbussche |
ECML/PKDD (1) | 4 |
| 2016 | Using Drug Similarities for Discovery of Possible Adverse Reactions
Emir Muñoz, Vít Novácek, Pierre-Yves Vandenbussche |
AMIA | 2 |
| 2014 | A Method for Building Burst-Annotated Co-Occurrence Networks for Analysing Trends in Textual Data
Yutaka Mitsuishi, Vít Novácek, Pierre-Yves Vandenbussche |
LREC | 2 |
| 2013 | Linking the scientific and clinical data with KI2NA-LHC - An outlineabstractWe introduce KI2NA-LHC (Linked Health Care) a system for data and knowledge integration in life sciences. In particular, we focus on linking clinical resources (electronic patient records) with scientific documents and data (research articles, biomedical ontologies and databases). Our motivation is two-fold. Firstly, we aim to instantly provide scientific context of particular patient cases for clinicians in order for them to propose treatments in a more informed way. Secondly, we want to build a technical infrastructure for researchers that will allow them to semi-automatically formulate and evaluate their hypothesis against longitudinal patient data. This paper outlines the proposed system and its services in a broader context of KI2NA, an ongoing collaboration between the DERI research institute and Fujitsu Laboratories. Vít Novácek, Aisha Naseer |
CBMS | 1 |
| 2011 | Getting the Meaning Right: A Complementary Distributional Layer for the Web Semantics
Vít Novácek, Siegfried Handschuh, Stefan Decker |
ISWC (1) | 1 |
| 2010 | CORAAL - Dive into publications, bathe in the knowledge
Vít Novácek, Tudor Groza, Siegfried Handschuh, Stefan Decker |
J. Web Semant. | 1 |
| 2009 | CORAAL - Towards Deep Exploitation of Textual Resources in Life Sciences
Vít Novácek, Tudor Groza, Siegfried Handschuh |
AIME | 1 |
| 2009 | Knowledge-based search for oncological literatureabstractUsing the current state of the art in life science publication search (e.g., PubMed), one can efficiently search for resources containing particular key-words or their combinations. It is impossible to search for abstract concepts and expressive relations between them (e.g., type of, different from or part of), though. Nevertheless, such a more expressive - semantic - search could largely reduce the efforts related to finding appropriate answers in biomedical articles. In this paper we identify challenges related to building a semantic publication search engine. Then we describe the architecture and usage principles of a tool tackling them. Eventually, we report on the tool's deployment on oncological literature data and preliminary tests with domain experts. Vít Novácek, Tudor Groza, Siegfried Handschuh |
CBMS | 1 |
| 2009 | Towards Lightweight and Robust Large Scale Emergent Knowledge Processing
Vít Novácek, Stefan Decker |
ISWC | 1 |
| 2008 | Infrastructure for dynamic knowledge integration - Automated biomedical ontology extension using textual resources
Vít Novácek, Loredana Laera, Siegfried Handschuh, Brian Davis 0001 |
J. Biomed. Informatics | 1 |
| 2006 | Empirical Merging of Ontologies - A Proposal of Universal Uncertainty Representation Framework
Vít Novácek, Pavel Smrz |
ESWC | 1 |
| 2006 | Text Mining for Semantic Relations as a Support Base of a Scientific Portal Generator
Vít Novácek, Pavel Smrz, Jan Pomikálek |
LREC | 1 |
| 2006 | Ontology Acquisition for Automatic Building of Scientific Portals
Pavel Smrz, Vít Novácek |
SOFSEM | 2 |