EDBT 2026 Demo / reviewers in the wild / expert
Nikola Simidjievski
dblp:168/9151
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0003-3948-6370ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Deep learning architectures and training · 21% Trustworthy machine learning · 20% Generative modeling · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 73% Medical and health informatics · 27% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.6 | 2 | 2025 | Measuring Cross-Modal Interactions in Multimodal Models · AAAI 2025 ProtoGate: Prototype-based Neural Networks with Global-to-local Feature Selection for Tabular Biomedical Data · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
1.4 | 2 | 2024 | ProtoGate: Prototype-based Neural Networks with Global-to-local Feature Selection for Tabular Biomedical Data · ICML 2024 Weight Predictor Network with Feature Selection for Small Sample Tabular Biomedical Data · AAAI 2023 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.9 | 1 | 2025 | Measuring Cross-Modal Interactions in Multimodal Models · AAAI 2025 |
Machine learning › Efficient and distributed learning
model merging |
0.9 | 1 | 2025 | Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in Biomedicine · ICLR 2025 |
Medical and health informatics
clinical decision support |
0.9 | 1 | 2025 | Measuring Cross-Modal Interactions in Multimodal Models · AAAI 2025 |
Computer vision › Vision and language
cross-modal attention |
0.8 | 1 | 2024 | HEALNet: Multimodal Fusion for Heterogeneous Biomedical Data · NeurIPS 2024 |
Machine learning › Generative modeling
energy-based model |
0.8 | 1 | 2024 | TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based Models · NeurIPS 2024 |
Machine learning › Generative modeling › synthetic data generation
tabular data augmentation |
0.8 | 1 | 2024 | TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based Models · NeurIPS 2024 |
Machine learning › Learning theory › sample complexity
small sample learning |
0.7 | 1 | 2023 | Weight Predictor Network with Feature Selection for Small Sample Tabular Biomedical Data · AAAI 2023 |
Machine learning › Deep learning architectures and training
tabular data learning |
0.7 | 1 | 2023 | Weight Predictor Network with Feature Selection for Small Sample Tabular Biomedical Data · AAAI 2023 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.6 | 1 | 2022 | Attentional Meta-learners for Few-shot Polythetic Classification · ICML 2022 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.6 | 1 | 2022 | Attentional Meta-learners for Few-shot Polythetic Classification · ICML 2022 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.6 | 1 | 2022 | Attentional Meta-learners for Few-shot Polythetic Classification · ICML 2022 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.6 | 1 | 2022 | Attentional Meta-learners for Few-shot Polythetic Classification · ICML 2022 |
Bioinformatics and computational biology
cancer genomics |
0.6 | 1 | 2022 | Unsupervised construction of computational graphs for gene expression data with explicit structural inductive biases · Bioinform. 2022 |
Bioinformatics and computational biology
gene expression analysis |
0.6 | 1 | 2022 | Unsupervised construction of computational graphs for gene expression data with explicit structural inductive biases · Bioinform. 2022 |
Bioinformatics and computational biology › statistical genetics
phenotype prediction |
0.6 | 1 | 2022 | Unsupervised construction of computational graphs for gene expression data with explicit structural inductive biases · Bioinform. 2022 |
Machine learning › Generative modeling
latent space regularization |
0.4 | 1 | 2020 | Constraining Variational Inference with Geometric Jensen-Shannon Divergence · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations |
0.4 | 1 | 2020 | On Second Order Behaviour in Augmented Neural ODEs · NeurIPS 2020 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2020 | Constraining Variational Inference with Geometric Jensen-Shannon Divergence · NeurIPS 2020 |
Medical and health informatics › biomedical data science › cancer informatics
cancer data analysis |
0.2 | 1 | 2024 | HEALNet: Multimodal Fusion for Heterogeneous Biomedical Data · NeurIPS 2024 |
Machine learning › Graph learning
graph neural network |
0.2 | 1 | 2022 | Unsupervised construction of computational graphs for gene expression data with explicit structural inductive biases · Bioinform. 2022 |
Methods — techniques the papers use, named apart from their topics
shapley interaction index · 1.7multimodal fusion · 1.7frequency-domain representation · 1.7SHAP · 1.7missing modality handling · 1.5early fusion · 1.5attention · 1.5energy-based model · 0.8weight predictor network · 0.7neural network · 0.7topological clustering · 0.6protein-protein interaction network · 0.6inductive bias · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Measuring Cross-Modal Interactions in Multimodal ModelsabstractIntegrating AI in healthcare can greatly improve patient care and system efficiency. However, the lack of explainability in AI systems (XAI) hinders their clinical adoption, especially in multimodal decision-making that combines various data sources. The majority of existing XAI methods focus on unimodal models, which fail to capture cross-modal interactions that are crucial for understanding the combined impact of multiple data sources. Existing methods for quantifying cross-modal interactions are limited to two modalities, rely on labelled data, and depend on model performance, which is problematic in healthcare, where XAI must handle multiple data sources and provide individualised explanations. This paper introduces InterSHAP, a cross-modal interaction score that addresses the limitations of existing approaches. InterSHAP uses the Shapley interaction index to precisely separate and quantify the contributions of the individual modalities and their interactions without approximations. By integrating an open-source implementation with the SHAP package, we enhance reproducibility and ease of use. We show that InterSHAP accurately measures the presence of cross-modal interactions, can handle multiple modalities, and provides detailed explanations at a local level for individual data points. Furthermore, we apply InterSHAP to real medical multimodal datasets, and demonstrate its practical applicability for individualised explanations. Laura Wenderoth, Konstantin Hemker, Nikola Simidjievski, Mateja Jamnik |
AAAI | 3 |
| 2025 | Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in BiomedicineabstractLearning holistic computational representations in physical, chemical or biological systems requires the ability to process information from different distributions and modalities within the same model. Thus, the demand for multimodal machine learning models has sharply risen for modalities that go beyond vision and language, such as sequences, graphs, time series, or tabular data. While there are many available multimodal fusion and alignment approaches, most of them require end-to-end training, scale quadratically with the number of modalities, cannot handle cases of high modality imbalance in the training set, or are highly topology-specific, making them too restrictive for many biomedical learning tasks. This paper presents Multimodal Lego (MM-Lego), a general-purpose fusion framework to turn any set of encoders into a competitive multimodal model with no or minimal fine-tuning. We achieve this by introducing a wrapper for any unimodal encoder that enforces shape consistency between modality representations. It harmonises these representations by learning features in the frequency domain to enable model merging with little signal interference. We show that MM-Lego 1) can be used as a model merging method which achieves competitive performance with end-to-end fusion models without any fine-tuning, 2) can operate on any unimodal encoder, and 3) is a model fusion method that, with minimal fine-tuning, surpasses all benchmarks in five out of seven datasets. Konstantin Hemker, Nikola Simidjievski, Mateja Jamnik |
ICLR | 2 |
| 2024 | ProtoGate: Prototype-based Neural Networks with Global-to-local Feature Selection for Tabular Biomedical DataabstractTabular biomedical data poses challenges in machine learning because it is often high-dimensional and typically low-sample-size (HDLSS). Previous research has attempted to address these challenges via local feature selection, but existing approaches often fail to achieve optimal performance due to their limitation in identifying globally important features and their susceptibility to the co-adaptation problem. In this paper, we propose ProtoGate, a prototype-based neural model for feature selection on HDLSS data. ProtoGate first selects instance-wise features via adaptively balancing global and local feature selection. Furthermore, ProtoGate employs a non-parametric prototype-based prediction mechanism to tackle the co-adaptation problem, ensuring the feature selection results and predictions are consistent with underlying data clusters. We conduct comprehensive experiments to evaluate the performance and interpretability of ProtoGate on synthetic and real-world datasets. The results show that ProtoGate generally outperforms state-of-the-art methods in prediction accuracy by a clear margin while providing high-fidelity feature selection and explainable predictions. Code is available at https://github.com/SilenceX12138/ProtoGate. Xiangjian Jiang, Andrei Margeloiu, Nikola Simidjievski, Mateja Jamnik |
ICML | 3 |
| 2024 | HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataabstractTechnological advances in medical data collection, such as high-throughput genomic sequencing and digital high-resolution histopathology, have contributed to the rising requirement for multimodal biomedical modelling, specifically for image, tabular and graph data. Most multimodal deep learning approaches use modality-specific architectures that are often trained separately and cannot capture the crucial cross-modal information that motivates the integration of different data sources. This paper presents the **H**ybrid **E**arly-fusion **A**ttention **L**earning **Net**work (HEALNet) – a flexible multimodal fusion architecture, which: a) preserves modality-specific structural information, b) captures the cross-modal interactions and structural information in a shared latent space, c) can effectively handle missing modalities during training and inference, and d) enables intuitive model inspection by learning on the raw data input instead of opaque embeddings. We conduct multimodal survival analysis on Whole Slide Images and Multi-omic data on four cancer datasets from The Cancer Genome Atlas (TCGA). HEALNet achieves state-of-the-art performance compared to other end-to-end trained fusion models, substantially improving over unimodal and multimodal baselines whilst being robust in scenarios with missing modalities. The code is available at https://github.com/konst-int-i/healnet. Konstantin Hemker, Nikola Simidjievski, Mateja Jamnik |
NeurIPS | 2 |
| 2024 | TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based ModelsabstractData collection is often difficult in critical fields such as medicine, physics, and chemistry, yielding typically only small tabular datasets. However, classification methods tend to struggle with these small datasets, leading to poor predictive performance. Increasing the training set with additional synthetic data, similar to data augmentation in images, is commonly believed to improve downstream tabular classification performance. However, current tabular generative methods that learn either the joint distribution $ p(\mathbf{x}, y) $ or the class-conditional distribution $ p(\mathbf{x} \mid y) $ often overfit on small datasets, resulting in poor-quality synthetic data, usually worsening classification performance compared to using real data alone. To solve these challenges, we introduce TabEBM, a novel class-conditional generative method using Energy-Based Models (EBMs). Unlike existing tabular methods that use a shared model to approximate all class-conditional densities, our key innovation is to create distinct EBM generative models for each class, each modelling its class-specific data distribution individually. This approach creates robust energy landscapes, even in ambiguous class distributions. Our experiments show that TabEBM generates synthetic data with higher quality and better statistical fidelity than existing methods. When used for data augmentation, our synthetic data consistently leads to improved classification performance across diverse datasets of various sizes, especially small ones. Code is available at https://github.com/andreimargeloiu/TabEBM. Andrei Margeloiu, Xiangjian Jiang, Nikola Simidjievski, Mateja Jamnik |
NeurIPS | 3 |
| 2024 | In-Domain Self-Supervised Learning Improves Remote Sensing Image Scene ClassificationabstractWe investigate the utility of in-domain self-supervised pre-training of vision models in the analysis of remote sensing imagery. Self-supervised learning (SSL) has emerged as a promising approach for remote sensing image classification due to its ability to exploit large amounts of unlabeled data. Unlike traditional supervised learning, SSL aims to learn representations of data without the need for explicit labels. This is achieved by formulating auxiliary tasks that can be used for pre-training models before fine-tuning them on a given downstream task. A common approach in practice to SSL pre-training is utilizing standard pre-training datasets, such as ImageNet. While relevant, such a general approach can have a sub-optimal influence on the downstream performance of models, especially on tasks from challenging domains such as remote sensing. In this paper, we analyze the effectiveness of SSL pre-training by employing the iBOT framework coupled with Vision transformers trained on Million-AID, a large and unlabeled remote sensing dataset. We present a comprehensive study of different self-supervised pre-training strategies and evaluate their effect across 14 downstream datasets with diverse properties. Our results demonstrate that leveraging large in-domain datasets for self-supervised pre-training consistently leads to improved predictive downstream performance, compared to the standard approaches found in practice. Ivica Dimitrovski, Ivan Kitanovski, Nikola Simidjievski, Dragi Kocev |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Weight Predictor Network with Feature Selection for Small Sample Tabular Biomedical DataabstractTabular biomedical data is often high-dimensional but with a very small number of samples. Although recent work showed that well-regularised simple neural networks could outperform more sophisticated architectures on tabular data, they are still prone to overfitting on tiny datasets with many potentially irrelevant features. To combat these issues, we propose Weight Predictor Network with Feature Selection (WPFS) for learning neural networks from high-dimensional and small sample data by reducing the number of learnable parameters and simultaneously performing feature selection. In addition to the classification network, WPFS uses two small auxiliary networks that together output the weights of the first layer of the classification model. We evaluate on nine real-world biomedical datasets and demonstrate that WPFS outperforms other standard as well as more recent methods typically applied to tabular data. Furthermore, we investigate the proposed feature selection mechanism and show that it improves performance while providing useful insights into the learning task. Andrei Margeloiu, Nikola Simidjievski, Pietro Liò, Mateja Jamnik |
AAAI | 2 |
| 2022 | Attentional Meta-learners for Few-shot Polythetic ClassificationabstractPolythetic classifications, based on shared patterns of features that need neither be universal nor constant among members of a class, are common in the natural world and greatly outnumber monothetic classifications over a set of features. We show that threshold meta-learners, such as Prototypical Networks, require an embedding dimension that is exponential in the number of task-relevant features to emulate these functions. In contrast, attentional classifiers, such as Matching Networks, are polythetic by default and able to solve these problems with a linear embedding dimension. However, we find that in the presence of task-irrelevant features, inherent to meta-learning problems, attentional models are susceptible to misclassification. To address this challenge, we propose a self-attention feature-selection mechanism that adaptively dilutes non-discriminative features. We demonstrate the effectiveness of our approach in meta-learning Boolean functions, and synthetic and real-world few-shot learning tasks. Ben Day, Ramón Viñas 0001, Nikola Simidjievski, Pietro Liò |
ICML | 3 |
| 2022 | Unsupervised construction of computational graphs for gene expression data with explicit structural inductive biasesabstractMOTIVATION: Gene expression data are commonly used at the intersection of cancer research and machine learning for better understanding of the molecular status of tumour tissue. Deep learning predictive models have been employed for gene expression data due to their ability to scale and remove the need for manual feature engineering. However, gene expression data are often very high dimensional, noisy and presented with a low number of samples. This poses significant problems for learning algorithms: models often overfit, learn noise and struggle to capture biologically relevant information. In this article, we utilize external biological knowledge embedded within structures of gene interaction graphs such as protein-protein interaction (PPI) networks to guide the construction of predictive models. RESULTS: We present Gene Interaction Network Constrained Construction (GINCCo), an unsupervised method for automated construction of computational graph models for gene expression data that are structurally constrained by prior knowledge of gene interaction networks. We employ this methodology in a case study on incorporating a PPI network in cancer phenotype prediction tasks. Our computational graphs are structurally constructed using topological clustering algorithms on the PPI networks which incorporate inductive biases stemming from network biology research on protein complex discovery. Each of the entities in the GINCCo computational graph represents biological entities such as genes, candidate protein complexes and phenotypes instead of arbitrary hidden nodes of a neural network. This provides a biologically relevant mechanism for model regularization yielding strong predictive performance while drastically reducing the number of model parameters and enabling guided post-hoc enrichment analyses of influential gene sets with respect to target phenotypes. Our experiments analysing a variety of cancer phenotypes show that GINCCo often outperforms support vector machine, Fully Connected Multi-layer Perceptrons (MLP) and Randomly Connected MLPs despite greatly reduced model complexity. AVAILABILITY AND IMPLEMENTATION: https://github.com/paulmorio/gincco contains the source code for our approach. We also release a library with algorithms for protein complex discovery within PPI networks at https://github.com/paulmorio/protclus. This repository contains implementations of the clustering algorithms used in this article. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Paul Scherer, Maja Trebacz, Nikola Simidjievski, Ramón Viñas 0001, Zohreh Shams, Helena Andrés-Terré, Mateja Jamnik, Pietro Liò |
Bioinform. | 3 |
| 2020 | Constraining Variational Inference with Geometric Jensen-Shannon DivergenceabstractWe examine the problem of controlling divergences for latent space regularisation in variational autoencoders. Specifically, when aiming to reconstruct example $x\in\mathbb{R}^{m}$ via latent space $z\in\mathbb{R}^{n}$ ($n\leq m$), while balancing this against the need for generalisable latent representations. We present a regularisation mechanism based on the {\em skew-geometric Jensen-Shannon divergence} $\left(\textrm{JS}^{\textrm{G}_{\alpha}}\right)$. We find a variation in $\textrm{JS}^{\textrm{G}_{\alpha}}$, motivated by limiting cases, which leads to an intuitive interpolation between forward and reverse KL in the space of both distributions and divergences. We motivate its potential benefits for VAEs through low-dimensional examples, before presenting quantitative and qualitative results. Our experiments demonstrate that skewing our variant of $\textrm{JS}^{\textrm{G}_{\alpha}}$, in the context of $\textrm{JS}^{\textrm{G}_{\alpha}}$-VAEs, leads to better reconstruction and generation when compared to several baseline VAEs. Our approach is entirely unsupervised and utilises only one hyperparameter which can be easily interpreted in latent space. Jacob Deasy, Nikola Simidjievski, Pietro Liò |
NeurIPS | 2 |
| 2020 | On Second Order Behaviour in Augmented Neural ODEsabstractNeural Ordinary Differential Equations (NODEs) are a new class of models that transform data continuously through infinite-depth architectures. The continuous nature of NODEs has made them particularly suitable for learning the dynamics of complex physical systems. While previous work has mostly been focused on first order ODEs, the dynamics of many systems, especially in classical physics, are governed by second order laws. In this work, we consider Second Order Neural ODEs (SONODEs). We show how the adjoint sensitivity method can be extended to SONODEs and prove that the optimisation of a first order coupled ODE is equivalent and computationally more efficient. Furthermore, we extend the theoretical understanding of the broader class of Augmented NODEs (ANODEs) by showing they can also learn higher order dynamics with a minimal number of augmented dimensions, but at the cost of interpretability. This indicates that the advantages of ANODEs go beyond the extra space offered by the augmented dimensions, as originally thought. Finally, we compare SONODEs and ANODEs on synthetic and real dynamical systems and demonstrate that the inductive biases of the former generally result in faster training and better performance. Alexander Norcliffe, Cristian Bodnar, Ben Day, Nikola Simidjievski, Pietro Liò |
NeurIPS | 4 |
| 2017 | Process-Based Modeling and Design of Dynamical Systems
Jovan Tanevski, Nikola Simidjievski, Ljupco Todorovski, Saso Dzeroski |
ECML/PKDD (3) | 2 |
| 2016 | Learning Ensembles of Process-Based Models by Bagging of Random Library Samples
Nikola Simidjievski, Ljupco Todorovski, Saso Dzeroski |
DS | 1 |
| 2015 | Predicting long-term population dynamics with bagging and boosting of process-based models
Nikola Simidjievski, Ljupco Todorovski, Saso Dzeroski |
Expert Syst. Appl. | 1 |