VLDB 2026 Research / reviewers in the wild / expert
Dragi Kocev
dblp:18/6014
· DBLP profile ↗
46ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-0687-0878ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fully- and semi-supervised hierarchical multi-label image classification with graph learning
Marjan Stoimchev, Boshko Koloski, Jurica Levatic, Dragi Kocev, Saso Dzeroski |
Inf. Sci. | 4 |
| 2024 | Semi-Supervised Predictive Clustering Trees for (Hierarchical) Multi-Label ClassificationabstractSemi-supervised learning (SSL) is a common approach to learning predictive models using not only labeled, but also unlabeled examples. While SSL for the simple tasks of classification and regression has received much attention from the research community, this is not the case for complex prediction tasks with structurally dependent variables, such as multi-label classification and hierarchical multi-label classification. These tasks may require additional information, possibly coming from the underlying distribution in the descriptive space provided by unlabeled examples, to better face the challenging task of simultaneously predicting multiple class labels. In this paper, we investigate this aspect and propose a (hierarchical) multi-label classification method based on semi-supervised learning of predictive clustering trees, which we also extend towards ensemble learning. Extensive experimental evaluation conducted on 24 datasets shows significant advantages of the proposed method and its extension with respect to their supervised counterparts. Moreover, the method preserves interpretability of classical tree-based models. Jurica Levatic, Michelangelo Ceci, Dragi Kocev, Saso Dzeroski |
Int. J. Intell. Syst. | 3 |
| 2024 | In-Domain Self-Supervised Learning Improves Remote Sensing Image Scene ClassificationabstractWe investigate the utility of in-domain self-supervised pre-training of vision models in the analysis of remote sensing imagery. Self-supervised learning (SSL) has emerged as a promising approach for remote sensing image classification due to its ability to exploit large amounts of unlabeled data. Unlike traditional supervised learning, SSL aims to learn representations of data without the need for explicit labels. This is achieved by formulating auxiliary tasks that can be used for pre-training models before fine-tuning them on a given downstream task. A common approach in practice to SSL pre-training is utilizing standard pre-training datasets, such as ImageNet. While relevant, such a general approach can have a sub-optimal influence on the downstream performance of models, especially on tasks from challenging domains such as remote sensing. In this paper, we analyze the effectiveness of SSL pre-training by employing the iBOT framework coupled with Vision transformers trained on Million-AID, a large and unlabeled remote sensing dataset. We present a comprehensive study of different self-supervised pre-training strategies and evaluate their effect across 14 downstream datasets with diverse properties. Our results demonstrate that leveraging large in-domain datasets for self-supervised pre-training consistently leads to improved predictive downstream performance, compared to the standard approaches found in practice. Ivica Dimitrovski, Ivan Kitanovski, Nikola Simidjievski, Dragi Kocev |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Semi-Supervised Multi-Label Classification of Land Use/Land Cover in Remote Sensing Images With Predictive Clustering Trees and EnsemblesabstractThe task of remote sensing image (RSI) classification has been studied extensively in the geoscience and remote sensing (RS) community. While deep learning methods have shown great success in solving this task, their reliance on large-scale labeled datasets is a serious limitation when dealing with complex labels and multiple semantic categories. The process of annotating such datasets can be time-consuming and tedious, leading to limited availability of labeled data and reduced performance of supervised learning methods. To address this issue, semi-supervised learning (SSL) methods can be applied, as they use both the limited labeled data and the abundant unlabeled data. In this article, we propose an effective SSL framework for RSI classification, which combines two key concepts. First, we employ a deep convolutional feature extractor to learn feature representations that encode the images into a lower dimensional feature space, capturing the rich semantic context present in RSI. Second, we utilize semi-supervised predictive clustering trees (PCTs) and ensembles thereof to learn from both the labeled and unlabeled data. To evaluate the effectiveness of the proposed framework, we compare it against several state-of-the-art self-supervised and semi-supervised methods from the literature. We conduct extensive experiments on ten publicly available land use/land cover RSI classification datasets: five for multiclass classification (MCC) and five for multi-label classification (MLC). The results demonstrate that the proposed framework has superior predictive performance compared to state-of-the-art methods from the literature, highlighting its effectiveness in semi-supervised RSI classification. Marjan Stoimchev, Jurica Levatic, Dragi Kocev, Saso Dzeroski |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Feature ranking for semi-supervised learningabstractAbstract The data used for analysis are becoming increasingly complex along several directions: high dimensionality, number of examples and availability of labels for the examples. This poses a variety of challenges for the existing machine learning methods, related to analyzing datasets with a large number of examples that are described in a high-dimensional space, where not all examples have labels provided. For example, when investigating the toxicity of chemical compounds, there are many compounds available that can be described with information-rich high-dimensional representations, but not all of the compounds have information on their toxicity. To address these challenges, we propose methods for semi-supervised learning (SSL) of feature rankings. The feature rankings are learned in the context of classification and regression, as well as in the context of structured output prediction (multi-label classification, MLC, hierarchical multi-label classification, HMLC and multi-target regression, MTR) tasks. This is the first work that treats the task of feature ranking uniformly across various tasks of semi-supervised structured output prediction. To the best of our knowledge, it is also the first work on SSL of feature rankings for the tasks of HMLC and MTR. More specifically, we propose two approaches—based on predictive clustering tree ensembles and the Relief family of algorithms—and evaluate their performance across 38 benchmark datasets. The extensive evaluation reveals that rankings based on Random Forest ensembles perform the best for classification tasks (incl. MLC and HMLC tasks) and are the fastest for all tasks, while ensembles based on extremely randomized trees work best for the regression tasks. Semi-supervised feature rankings outperform their supervised counterparts across the majority of datasets for all of the different tasks, showing the benefit of using unlabeled in addition to labeled data. Matej Petkovic, Saso Dzeroski, Dragi Kocev |
Mach. Learn. | 3 |
| 2022 | Crop Type Prediction Across Countries and Years: Slovenia, Denmark and the NetherlandsabstractCrop type prediction is a very relevant and a very challenging task. The increasing availability of high-quality satellite imagery and machine learning have enabled the development of automatic crop type classification methods. In this paper, we present a crop type prediction data suite that consists of crop type information from three countries (Denmark, the Netherlands, and Slovenia) across three years (2017, 2018 and 2019). By considering the complex challenges contained by this data suite, we investigate the robustness of 7 deep learning methods used for crop type prediction (TempCNN, MSResNet, InceptionTime, OmniscaleCNN, LSTM, StarRNN, and Transformer networks). The comprehensive experiments reveal that the recurrence-based methods perform the best (with LSTM being the best performing). The methods can achieve very good predictive performance - up to a weighted F1 score of 0.8432. Elena Merdjanovska, Ivan Kitanovski, Ziga Kokalj, Ivica Dimitrovski, Dragi Kocev |
IGARSS | 5 |
| 2022 | Comprehensive comparative study of multi-label classification methodsabstractMulti-label classification (MLC) has recently attracted increasing interest in the machine learning community. Several studies provide surveys of methods and datasets for MLC, and a few provide empirical comparisons of MLC methods. However, they are limited in the number of methods and datasets considered. This paper provides a comprehensive empirical investigation of a wide range of MLC methods on a wealth of datasets from different domains. More specifically, our study evaluates 26 methods on 42 benchmark datasets using 20 evaluation measures. The evaluation methodology used meets the highest literature standards for designing and conducting large-scale, time-limited experimental studies. First, the methods were selected based on their use in the community to ensure a balanced representation of methods across the MLC taxonomy of methods within the study. Second, the datasets cover a wide range of complexity and application domains. The selected evaluation measures assess the predictive performance and efficiency of the methods. The results of the analysis identify RFPCT, RFDTBR, ECCJ48, EBRJ48, and AdaBoost.MH as the best-performing methods across the spectrum of performance measures. Whenever a new method is introduced, it should be compared with different subsets of MLC methods selected according to relevant (and possibly different) evaluation criteria. Jasmin Bogatinovski, Ljupco Todorovski, Saso Dzeroski, Dragi Kocev |
Expert Syst. Appl. | 4 |
| 2022 | Explaining the performance of multilabel classification methods with data set propertiesabstractMeta learning generalizes the empirical experience with different learning tasks and holds promise for providing important empirical insight into the behavior of machine learning algorithms.In this paper, we present a comprehensive meta-learning study of data sets and methods for multilabel classification (MLC).MLC is a practically relevant machine learning task where each example is labeled with multiple labels simultaneously.Here, we analyze 40 MLC data sets by using 50 meta features describing different properties of the data.The main findings of this study are as follows.First, the most prominent meta features that describe the space of MLC data sets are the ones assessing different aspects of the label space.Second, the meta models show that the most important meta features describe the label space, and, the meta features describing the relationships among the labels tend to occur a bit more often than the meta features describing the distributions between and within the individual labels.Third, the optimization of the hyperparameters can improve the predictive performance, however, quite often the extent of the Jasmin Bogatinovski, Ljupco Todorovski, Saso Dzeroski, Dragi Kocev |
Int. J. Intell. Syst. | 4 |
| 2021 | Ensemble- and distance-based feature ranking for unsupervised learningabstractIn this study, we propose two novel (groups of) methods for unsupervised feature ranking and selection. The first group includes feature ranking scores (Genie3 score, RandomForest score) that are computed from ensembles of predictive clustering trees. The second method is URelief, the unsupervised extension of the Relief family of feature ranking algorithms. Using 26 benchmark data sets and 5 baselines, we show that both the Genie3 score (computed from the ensemble of extra trees) and the URelief method outperform the existing methods and that Genie3 performs best overall, in terms of predictive power of the top-ranked features. Additionally, we analyze the influence of the hyper-parameters of the proposed methods on their performance and show that for the Genie3 score the highest quality is achieved by the most efficient parameter configuration. Finally, we propose a way of discovering the location of the features in the ranking, which are the most relevant in reality. Matej Petkovic, Dragi Kocev, Blaz Skrlj, Saso Dzeroski |
Int. J. Intell. Syst. | 2 |
| 2021 | Oblique predictive clustering trees
Tomaz Stepisnik Perdih, Dragi Kocev |
Knowl. Based Syst. | 2 |
| 2020 | Hyperbolic Embeddings for Hierarchical Multi-label Classification
Tomaz Stepisnik Perdih, Dragi Kocev |
ISMIS | 2 |
| 2020 | Multivariate Predictive Clustering Trees for Classification
Tomaz Stepisnik Perdih, Dragi Kocev |
ISMIS | 2 |
| 2020 | Semi-supervised regression trees with application to QSAR modelling
Jurica Levatic, Michelangelo Ceci, Tomaz Stepisnik Perdih, Saso Dzeroski, Dragi Kocev |
Expert Syst. Appl. | 5 |
| 2020 | Ensembles of extremely randomized predictive clustering trees for predicting structured outputs
Dragi Kocev, Michelangelo Ceci, Tomaz Stepisnik Perdih |
Mach. Learn. | 1 |
| 2020 | Multi-label feature ranking with ensemble methods
Matej Petkovic, Saso Dzeroski, Dragi Kocev |
Mach. Learn. | 3 |
| 2020 | Feature ranking for multi-target regression
Matej Petkovic, Dragi Kocev, Saso Dzeroski |
Mach. Learn. | 2 |
| 2019 | Mix and Rank: A Framework for Benchmarking Recommender SystemsabstractRecommender systems use big data methods, and are widely used in various social-network, e-commerce, and content platforms. With their increased relevance, online platforms and developers are in need of better ways to choose the systems that are most suitable for their use-cases. At the same time, the research literature on recommender systems describes a multitude of measures to evaluate the performance of different algorithms. For the end-user however, the large number of available measures do not provide much help in deciding which algorithm to deploy. Some of the measures are correlated, while others deal with different aspects of recommendation performance like accuracy and coverage. To address this problem, we propose a novel benchmarking framework that mixes different evaluation measures in order to rank the recommender systems on each benchmark dataset, separately. Additionally, our approach discovers sets of correlated measures as well as sets of evaluation measures that are least correlated. We investigate the robustness of the proposed methodology using published results from an experimental study involving multiple big datasets and evaluation measures. Our work provides a general framework that can handle an arbitrary number of evaluation measures and help end-users rank the systems available to them. Bibek Paudel, Dragi Kocev, Tome Eftimov |
IEEE BigData | 2 |
| 2019 | Ensemble-Based Feature Ranking for Semi-supervised Classification
Matej Petkovic, Saso Dzeroski, Dragi Kocev |
DS | 3 |
| 2019 | Predicting Thermal Power Consumption of the Mars Express Satellite with Data Stream Mining
Bozhidar Stevanoski, Dragi Kocev, Aljaz Osojnik, Ivica Dimitrovski, Saso Dzeroski |
DS | 2 |
| 2019 | Web genre classification with methods for structured output prediction
Gjorgji Madjarov, Vedrana Vidulin, Ivica Dimitrovski, Dragi Kocev |
Inf. Sci. | 4 |
| 2018 | Feature Ranking with Relief for Multi-label Classification: Does Distance Matter?
Matej Petkovic, Dragi Kocev, Saso Dzeroski |
DS | 2 |
| 2018 | Semi-supervised trees for multi-target regression
Jurica Levatic, Dragi Kocev, Michelangelo Ceci, Saso Dzeroski |
Inf. Sci. | 2 |
| 2018 | Ensembles for multi-target regression with random output selections
Martin Breskvar, Dragi Kocev, Saso Dzeroski |
Mach. Learn. | 2 |
| 2017 | Modelling Time-Series of Glucose Measurements from Diabetes Patients Using Predictive Clustering Trees
Mate Bestek, Dragi Kocev, Saso Dzeroski, Andrej Brodnik, Rade Iljaz |
AIME | 2 |
| 2017 | Multi-label Classification Using Random Label Subset Selections
Martin Breskvar, Dragi Kocev, Saso Dzeroski |
DS | 2 |
| 2017 | Option Predictive Clustering Trees for Hierarchical Multi-label Classification
Tomaz Stepisnik Perdih, Aljaz Osojnik, Saso Dzeroski, Dragi Kocev |
DS | 4 |
| 2017 | Feature Ranking for Multi-target Regression with Tree Ensemble Methods
Matej Petkovic, Saso Dzeroski, Dragi Kocev |
DS | 3 |
| 2017 | Predictive Clustering Trees for Hierarchical Multi-Target Regression
Vanja Mileski, Saso Dzeroski, Dragi Kocev |
IDA | 3 |
| 2017 | Image Representation, Annotation and Retrieval with Predictive Clustering Trees
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski |
ECML/PKDD (3) | 2 |
| 2017 | Introduction to the special issue dedicated to the Journal Track of ECML PKDD 2017
Kurt Driessens, Dragi Kocev, Marko Robnik-Sikonja, Myra Spiliopoulou |
Data Min. Knowl. Discov. | 2 |
| 2017 | Semi-supervised classification trees
Jurica Levatic, Michelangelo Ceci, Dragi Kocev, Saso Dzeroski |
J. Intell. Inf. Syst. | 3 |
| 2017 | Self-training for multi-target regression with tree ensembles
Jurica Levatic, Michelangelo Ceci, Dragi Kocev, Saso Dzeroski |
Knowl. Based Syst. | 3 |
| 2017 | Introduction to the special issue dedicated to the Journal Track of ECML PKDD 2017
Kurt Driessens, Dragi Kocev, Marko Robnik-Sikonja, Myra Spiliopoulou |
Mach. Learn. | 2 |
| 2016 | Option Predictive Clustering Trees for Multi-target Regression
Aljaz Osojnik, Saso Dzeroski, Dragi Kocev |
DS | 3 |
| 2016 | Improving bag-of-visual-words image retrieval with predictive clustering trees
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski |
Inf. Sci. | 2 |
| 2016 | Special issue on discovery science
Saso Dzeroski, Dragi Kocev, Pance Panov |
Mach. Learn. | 2 |
| 2015 | Ensembles of Extremely Randomized Trees for Multi-target Regression
Dragi Kocev, Michelangelo Ceci |
Discovery Science | 1 |
| 2015 | Web Genre Classification via Hierarchical Multi-label Classification
Gjorgji Madjarov, Vedrana Vidulin, Ivica Dimitrovski, Dragi Kocev |
IDEAL | 4 |
| 2015 | The importance of the label hierarchy in hierarchical multi-label classification
Jurica Levatic, Dragi Kocev, Saso Dzeroski |
J. Intell. Inf. Syst. | 2 |
| 2014 | Fast and efficient visual codebook construction for multi-label annotation using predictive clustering trees
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski |
Pattern Recognit. Lett. | 2 |
| 2013 | Fast and Scalable Image Retrieval Using Predictive Clustering Trees
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski |
Discovery Science | 2 |
| 2013 | Tree ensembles for predicting structured outputs
Dragi Kocev, Celine Vens, Jan Struyf, Saso Dzeroski |
Pattern Recognit. | 1 |
| 2012 | An extensive experimental comparison of methods for multi-label learning
Gjorgji Madjarov, Dragi Kocev, Dejan Gjorgjevikj, Saso Dzeroski |
Pattern Recognit. | 2 |
| 2011 | Hierarchical annotation of medical images
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski |
Pattern Recognit. | 2 |
| 2010 | Predicting gene function using hierarchical multi-label decision tree ensemblesabstractBACKGROUND: S. cerevisiae, A. thaliana and M. musculus are well-studied organisms in biology and the sequencing of their genomes was completed many years ago. It is still a challenge, however, to develop methods that assign biological functions to the ORFs in these genomes automatically. Different machine learning methods have been proposed to this end, but it remains unclear which method is to be preferred in terms of predictive performance, efficiency and usability. RESULTS: We study the use of decision tree based models for predicting the multiple functions of ORFs. First, we describe an algorithm for learning hierarchical multi-label decision trees. These can simultaneously predict all the functions of an ORF, while respecting a given hierarchy of gene functions (such as FunCat or GO). We present new results obtained with this algorithm, showing that the trees found by it exhibit clearly better predictive performance than the trees found by previously described methods. Nevertheless, the predictive performance of individual trees is lower than that of some recently proposed statistical learning methods. We show that ensembles of such trees are more accurate than single trees and are competitive with state-of-the-art statistical learning and functional linkage methods. Moreover, the ensemble method is computationally efficient and easy to use. CONCLUSIONS: Our results suggest that decision tree based methods are a state-of-the-art, efficient and easy-to-use approach to ORF function prediction. Leander Schietgat, Celine Vens, Jan Struyf, Hendrik Blockeel, Dragi Kocev, Saso Dzeroski |
BMC Bioinform. | 5 |
| 2007 | Ensembles of Multi-Objective Decision Trees
Dragi Kocev, Celine Vens, Jan Struyf, Saso Dzeroski |
ECML | 1 |