Saso Dzeroski

dblp:d/SasoDzeroski · DBLP profile ↗
← Back
52ranked-venue papers in the field
9as first author
6since 2021 · last 2026
0000-0003-2363-712XORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 35 (7 first)Knowledge Engineering, Semantic Web & Information Systems · 12 (1 first)Other / Interdisciplinary · 4Database Systems & Data Management · 1 (1 first)
YearPublicationVenuePosition
2026 Dynamic instance weighting for online learning in multi-cryptocurrency price and trend forecasting
abstract
Abstract The cryptocurrency market represents a significant innovation in the financial ecosystem, built upon cryptographic principles to ensure secure and transparent transactions. Cryptocurrencies experienced a global adoption, driven by their decentralized nature that enables borderless transactions without third-party intermediaries. The price of cryptocurrencies is characterized by a significant volatility, that introduces both opportunities and challenges. In this context, the development of accurate methods for the forecasting of price variation, able to work in real-time on data streams, has become vital for various stakeholders. In this paper, we propose a novel approach, called LEMON, for the online prediction of the price variation of cryptocurrencies, that leverages possible temporal correlations among them. Our approach stems from the empirical evidence that cryptocurrencies tend to form groups characterized by similar trends, a behavior often attributed to shared market dynamics and common external factors. Through the analysis of temporal correlations, LEMON dynamically identifies these groups, that are then exploited to learn multiple multi-target tree-based models, specifically designed for processing continuous data streams. LEMON also introduces a novel adaptive non-parametric weighting scheme, that automatically adjusts the importance of each instance based on the observed data distribution in real-time, improving the forecasting of the price variation. Our experiments, performed on 16 datasets related to 16 cryptocurrencies, demonstrate that LEMON outperforms state-of-the-art approaches in two distinct prediction tasks: forecasting the closing price variation (regression) and predicting the market trend direction (classification), making it an effective tool to support stakeholders requiring accurate real-time predictions.
Antonio Pellicani, Gianvito Pio, Saso Dzeroski, Michelangelo Ceci
Data Min. Knowl. Discov.3
2026 Fully- and semi-supervised hierarchical multi-label image classification with graph learning
Marjan Stoimchev, Boshko Koloski, Jurica Levatic, Dragi Kocev, Saso Dzeroski
Inf. Sci.5
2024 Semi-Supervised Predictive Clustering Trees for (Hierarchical) Multi-Label Classification
abstract
Semi-supervised learning (SSL) is a common approach to learning predictive models using not only labeled, but also unlabeled examples. While SSL for the simple tasks of classification and regression has received much attention from the research community, this is not the case for complex prediction tasks with structurally dependent variables, such as multi-label classification and hierarchical multi-label classification. These tasks may require additional information, possibly coming from the underlying distribution in the descriptive space provided by unlabeled examples, to better face the challenging task of simultaneously predicting multiple class labels. In this paper, we investigate this aspect and propose a (hierarchical) multi-label classification method based on semi-supervised learning of predictive clustering trees, which we also extend towards ensemble learning. Extensive experimental evaluation conducted on 24 datasets shows significant advantages of the proposed method and its extension with respect to their supervised counterparts. Moreover, the method preserves interpretability of classical tree-based models.
Jurica Levatic, Michelangelo Ceci, Dragi Kocev, Saso Dzeroski
Int. J. Intell. Syst.4
2023 Dimensionally-consistent equation discovery through probabilistic attribute grammars
abstract
Equation discovery, also known as symbolic regression, is a machine learning task of inducing closed-form equations from data and background knowledge. The latter takes various forms. Domain-specific knowledge can constrain the space of candidate equations to those that make sense in the scientific or engineering domain of use. Cross-domain knowledge, on the other hand, imposes general rules for model acceptability, such as parsimony, understandability, or consistency of the equations with the dimensional units of the variables. In this paper, we propose using attribute grammars to ensure the induced equations' dimensional consistency. Attribute grammars are flexible enough to combine cross-domain knowledge on dimensional consistency with domain-specific knowledge expressed as a probabilistic context-free grammar. At the same time, we show that attribute grammars can be efficiently transformed into probabilistic context-free grammars for equation discovery with existing algorithms. Finally, we provide empirical evidence that attribute grammars ensuring dimensional consistency of equations can significantly improve the performance of equation discovery on the standard set of a hundred Feynman benchmarks.
Jure Brence, Saso Dzeroski, Ljupco Todorovski
Inf. Sci.2
2022 Explaining the performance of multilabel classification methods with data set properties
abstract
Meta learning generalizes the empirical experience with different learning tasks and holds promise for providing important empirical insight into the behavior of machine learning algorithms.In this paper, we present a comprehensive meta-learning study of data sets and methods for multilabel classification (MLC).MLC is a practically relevant machine learning task where each example is labeled with multiple labels simultaneously.Here, we analyze 40 MLC data sets by using 50 meta features describing different properties of the data.The main findings of this study are as follows.First, the most prominent meta features that describe the space of MLC data sets are the ones assessing different aspects of the label space.Second, the meta models show that the most important meta features describe the label space, and, the meta features describing the relationships among the labels tend to occur a bit more often than the meta features describing the distributions between and within the individual labels.Third, the optimization of the hyperparameters can improve the predictive performance, however, quite often the extent of the
Jasmin Bogatinovski, Ljupco Todorovski, Saso Dzeroski, Dragi Kocev
Int. J. Intell. Syst.3
2021 Ensemble- and distance-based feature ranking for unsupervised learning
abstract
In this study, we propose two novel (groups of) methods for unsupervised feature ranking and selection. The first group includes feature ranking scores (Genie3 score, RandomForest score) that are computed from ensembles of predictive clustering trees. The second method is URelief, the unsupervised extension of the Relief family of feature ranking algorithms. Using 26 benchmark data sets and 5 baselines, we show that both the Genie3 score (computed from the ensemble of extra trees) and the URelief method outperform the existing methods and that Genie3 performs best overall, in terms of predictive power of the top-ranked features. Additionally, we analyze the influence of the hyper-parameters of the proposed methods on their performance and show that for the Genie3 score the highest quality is achieved by the most efficient parameter configuration. Finally, we propose a way of discovering the location of the features in the ranking, which are the most relevant in reality.
Matej Petkovic, Dragi Kocev, Blaz Skrlj, Saso Dzeroski
Int. J. Intell. Syst.4
2018 MetaBags: Bagged Meta-Decision Trees for Regression
Jihed Khiari, Luís Moreira-Matias, Ammar Shaker, Bernard Zenko, Saso Dzeroski
ECML/PKDD (1)5
2018 Semi-supervised trees for multi-target regression
Jurica Levatic, Dragi Kocev, Michelangelo Ceci, Saso Dzeroski
Inf. Sci.4
2018 Redescription mining augmented with random forest of multi-target predictive clustering trees
Matej Mihelcic, Saso Dzeroski, Nada Lavrac, Tomislav Smuc
J. Intell. Inf. Syst.2
2018 Tree-based methods for online multi-target regression
Aljaz Osojnik, Pance Panov, Saso Dzeroski
J. Intell. Inf. Syst.3
2017 Predictive Clustering Trees for Hierarchical Multi-Target Regression
Vanja Mileski, Saso Dzeroski, Dragi Kocev
IDA2
2017 Image Representation, Annotation and Retrieval with Predictive Clustering Trees
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski
ECML/PKDD (3)4
2017 Process-Based Modeling and Design of Dynamical Systems
Jovan Tanevski, Nikola Simidjievski, Ljupco Todorovski, Saso Dzeroski
ECML/PKDD (3)4
2017 Semi-supervised classification trees
Jurica Levatic, Michelangelo Ceci, Dragi Kocev, Saso Dzeroski
J. Intell. Inf. Syst.4
2017 Erratum to: The use of data-derived label hierarchies in multi-label classification
Gjorgji Madjarov, Dejan Gjorgjevikj, Ivica Dimitrovski, Saso Dzeroski
J. Intell. Inf. Syst.4
2016 Improving bag-of-visual-words image retrieval with predictive clustering trees
Ivica Dimitrovski, Dragi Kocev, Suzana Loskovska, Saso Dzeroski
Inf. Sci.4
2016 Generic ontology of datatypes
abstract
We present OntoDT, a generic ontology for the representation of scientific knowledge about datatypes. OntoDT defines basic entities, such as datatype, properties of datatypes, specifications, characterizing operations, and a datatype taxonomy. We demonstrate the utility of OntoDT on several use cases. OntoDT was used within an Ontology of core data mining entities for constructing taxonomies of datasets, data mining tasks, generalizations and data mining algorithms. Furthermore, we show how OntoDT can be used to annotate and query dataset repositories. We also show how OntoDT can improve the representation of datatypes in the BioXSD exchange format for basic bio-informatics types of data. The generic nature of OntoDT enables it to support a wide range of other applications, especially in combination with other domain specific ontologies: the construction of data mining workflows, annotation of software and algorithms, semantic annotation of scientific articles, etc. OntoDT is open source and is available at http://www.ontodt.com.
Pance Panov, Larisa N. Soldatova, Saso Dzeroski
Inf. Sci.3
2016 The use of data-derived label hierarchies in multi-label classification
Gjorgji Madjarov, Dejan Gjorgjevikj, Ivica Dimitrovski, Saso Dzeroski
J. Intell. Inf. Syst.4
2015 Model-Tree Ensembles for noise-tolerant system identification
Darko Aleksovski, Jus Kocijan, Saso Dzeroski
Adv. Eng. Informatics3
2015 The importance of the label hierarchy in hierarchical multi-label classification
Jurica Levatic, Dragi Kocev, Saso Dzeroski
J. Intell. Inf. Syst.3
2014 Ontology of core data mining entities
Pance Panov, Larisa N. Soldatova, Saso Dzeroski
Data Min. Knowl. Discov.3
2012 Network regression with predictive clustering trees
Daniela Stojanova, Michelangelo Ceci, Annalisa Appice, Saso Dzeroski
Data Min. Knowl. Discov.4
2012 Estimating the risk of fire outbreaks in the natural environment
Daniela Stojanova, Andrej Kobler, Peter Ogrinc, Bernard Zenko, Saso Dzeroski
Data Min. Knowl. Discov.5
2011 Network Regression with Predictive Clustering Trees
Daniela Stojanova, Michelangelo Ceci, Annalisa Appice, Saso Dzeroski
ECML/PKDD (3)4
2011 Learning model trees from evolving data streams
Elena Ikonomovska, João Gama 0001, Saso Dzeroski
Data Min. Knowl. Discov.3
2009 Rule Ensembles for Multi-target Regression
abstract
Methods for learning decision rules are being successfully applied to many problem domains, especially where understanding and interpretation of the learned model is necessary. In many real life problems, we would like to predict multiple related (nominal or numeric) target attributes simultaneously. Methods for learning rules that predict multiple targets at once already exist, but are unfortunately based on the covering algorithm, which is not very well suited for regression problems. A better solution for regression problems may be a rule ensemble approach that transcribes an ensemble of decision trees into a large collection of rules. An optimization procedure is then used for selecting the best (and much smaller) subset of these rules, and to determine their weights. Using the rule ensembles approach we have developed a new system for learning rule ensembles for multi-target regression problems. The newly developed method was extensively evaluated and the results show that the accuracy of multi-target regression rule ensembles is better than the accuracy of multi-target regression trees, but somewhat worse than the accuracy of multi-target random forests. The rules are significantly more concise than random forests, and it is also possible to create very small rule sets that are still comparable in accuracy to single regression trees.
Timo Aho, Bernard Zenko, Saso Dzeroski
ICDM3
2008 A Minimal Description Length Scheme for Polynomial Regression
Aleksandar Peckov, Saso Dzeroski, Ljupco Todorovski
PAKDD2
2008 Learning Classification Rules for Multiple Target Attributes
Bernard Zenko, Saso Dzeroski
PAKDD2
2007 Stepwise Induction of Multi-target Model Trees
Annalisa Appice, Saso Dzeroski
ECML2
2007 Ensembles of Multi-Objective Decision Trees
Dragi Kocev, Celine Vens, Jan Struyf, Saso Dzeroski
ECML4
2007 Clustering Trees with Instance Level Constraints
Jan Struyf, Saso Dzeroski
ECML2
2007 Combining Bagging and Random Subspaces to Create Better Ensembles
Pance Panov, Saso Dzeroski
IDA2
2006 Decision Trees for Hierarchical Multilabel Classification: A Case Study in Functional Genomics
Hendrik Blockeel, Leander Schietgat, Jan Struyf, Saso Dzeroski, Amanda Clare
PKDD4
2004 Inductive Databases of Polynomial Equations
Saso Dzeroski, Ljupco Todorovski, Peter Ljubic
DaWaK1
2004 Inducing Polynomial Equations for Regression
Ljupco Todorovski, Peter Ljubic, Saso Dzeroski
ECML3
2003 Using Domain Specific Knowledge for Automated Modeling
Ljupco Todorovski, Saso Dzeroski
IDA2
2002 Ranking with Predictive Clustering Trees
Ljupco Todorovski, Hendrik Blockeel, Saso Dzeroski
ECML3
2002 Stacking with an Extended Set of Meta-level Attributes and MLR
Bernard Zenko, Saso Dzeroski
ECML2
2001 Using Domain Knowledge on Population Dynamics Modeling for Equation Discovery
Ljupco Todorovski, Saso Dzeroski
ECML2
2001 A Comparison of Stacking with Meta Decision Trees to Bagging, Boosting, and Stacking with other Methods
abstract
Meta decision trees (MDTs) are a method for combining multiple classifiers. We present an integration of the algorithm MLC4.5 for learning MDTs into the Weka data mining suite. We compare classifier ensembles combined with MDTs to bagged and boosted decision trees, and to classifier ensembles combined with other methods: voting and stacking with three different meta-level classifiers (ordinary decision trees, naive Bayes, and multi-response linear regression - MLR). Meta decision trees. Techniques for combining predictions obtained from multiple base-level classifiers can be clustered in three combining frameworks: voting (used in bagging and boosting), stacked generalization or stacking [7] and cascading. Meta decision trees (MDTs) [5] adopt the stacking framework of combining base-level classifiers. The difference between meta and ordinary decision trees (ODTs) is that MDT leaves specify which base-level classifier should be used, instead of predicting the class value directly. Th...
Bernard Zenko, Ljupco Todorovski, Saso Dzeroski
ICDM3
2000 Supporting Discovery in Medicine by Association Rule Mining of Bibliographic Databases
Dimitar Hristovski, Saso Dzeroski, Borut Peterlin, Anamarija Rozic-Hristovski
PKDD2
2000 Combining Multiple Models with Meta Decision Trees
Ljupco Todorovski, Saso Dzeroski
PKDD2
1999 Simultaneous Prediction of Mulriple Chemical Parameters of River Water Quality with TILDE
Hendrik Blockeel, Saso Dzeroski, Jasna Grbovic
PKDD2
1999 Experiments in Meta-level Learning with ILP
Ljupco Todorovski, Saso Dzeroski
PKDD2
1999 Editorial
Saso Dzeroski, Nada Lavrac
Data Min. Knowl. Discov.1
1998 ILP Experiments in Detecting Traffic Problems
Saso Dzeroski, Nico Jacobs, Martin Molina, Carlos Moure
ECML1
1995 Handling Real Numbers in ILP: A Step Towards Better Behavioural Clones (Extended Abstract)
Saso Dzeroski, Ljupco Todorovski, Tanja Urbancic
ECML1
1995 Knowledge Discovery in a Water Quality Database
Saso Dzeroski
KDD1
1995 Discovering Dynamics: From Inductive Logic Programming to Machine Discovery
Saso Dzeroski, Ljupco Todorovski
J. Intell. Inf. Syst.1
1994 Discovering Dynamics with Genetic Programming
Saso Dzeroski, Igor Petrovski
ECML1
1993 Learnability of Constrained Logic Programs
Saso Dzeroski, Stephen H. Muggleton, Stuart Russell 0001
ECML1
1993 Inductive Learning in Deductive Databases
abstract
Most current applications of inductive learning in databases take place in the context of a single extensional relation. The authors place inductive learning in the context of a set of relations defined either extensionally or intentionally in the framework of deductive databases. LINUS, an inductive logic programming system that induces virtual relations from example positive and negative tuples and already defined relations in a deductive database, is presented. Based on the idea of transforming the problem of learning relations to attribute-value form, several attribute-value learning systems are incorporated. As the latter handle noisy data successfully, LINUS is able to learn relations from real-life noisy databases. The use of LINUS for learning virtual relations is illustrated, and a study of its performance on noisy data is presented.>
Saso Dzeroski, Nada Lavrac
IEEE Trans. Knowl. Data Eng.1