Meelis Kull

dblp:20/5835 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-9257-595XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021
YearPublicationVenuePosition
2025 Probability Density from Latent Diffusion Models for Out-of-Distribution Detection
abstract
Despite rapid advances in AI, safety remains the main bottleneck to deploying machine-learning systems. A critical safety component is out-of-distribution (OOD) detection: given an input, decide whether it comes from the same distribution as the training data. In generative models, the most natural OOD score is the data likelihood. Actually, under the assumption of uniformly distributed OOD data, the likelihood is even the optimal OOD detector, as we show in this work. However, earlier work reported that likelihood often fails in practice, raising doubts about its usefulness. We explore if, in practice, the representation space also suffers from the inability to learn good density estimation for OOD detection, or if it is merely a problem of the pixel-space typically used in generative models. To test this, we trained a Variational Diffusion Model (VDM) not on images, but on the representation space of a pre-trained ResNet-18 to assess the performance of our likelihood-based-detector in comparison to the state-of-the-art methods from OpenOOD suite.
Joonas Järve, Karl Kaspar Haavel, Meelis Kull
ECAI3
2025 Aligning the Evaluation of Probabilistic Predictions with Downstream Value
abstract
Every prediction is ultimately used in a downstream task. Consequently, evaluating prediction quality is more meaningful when considered in the context of its downstream use. Metrics based solely on predictive performance often diverge from measures of real-world downstream impact. Existing approaches incorporate the downstream view by relying on multiple task-specific metrics, which can be burdensome to analyze, or by formulating cost-sensitive evaluations that require an explicit cost structure, typically assumed to be known a priori. We frame this mismatch as an evaluation alignment problem and propose a data-driven method to learn a proxy evaluation function aligned with the downstream evaluation. Building on the theory of proper scoring rules, we explore transformations of scoring rules that ensure the preservation of propriety. Our approach leverages weighted scoring rules parametrized by a neural network, where weighting is learned to align with the performance in the downstream task. This enables fast and scalable evaluation cycles across tasks where the weighting is complex or unknown a priori. We showcase our framework through synthetic and real-data experiments for regression tasks, demonstrating its potential to bridge the gap between predictive evaluation and downstream utility in modular prediction systems.
Novin Shahroudi, Viacheslav Komisarenko, Meelis Kull
ECAI3
2025 On the usefulness of the fit-on-test view on evaluating calibration of classifiers
abstract
Abstract Calibrated uncertainty estimates are essential for classifiers used in safety-critical applications. If a classifier is uncalibrated, then there is a unique way to calibrate its uncertainty using the idealistic true calibration map corresponding to this classifier. Although the true calibration map is typically unknown in practice, it can be estimated with many post-hoc calibration methods which fit some family of potential calibration functions on a validation dataset. This paper examines the connection between such post-hoc calibration methods and calibration evaluation. Despite the negative connotations of fitting on test data in machine learning, we claim that fitting calibration maps on test data as part of the calibration evaluation process is a method worth considering, and we refer to this view as fit-on-test. This view enables the usage of any post-hoc calibration method as an evaluation measure, unlocking missed opportunities in development of evaluation methods. We prove that even ECE, which is the most common calibration evaluation method, is actually a fit-on-test measure. This observation leads us to a new method of tuning the number of bins in ECE with cross-validation. Fitting on test data can lead to test-time overfitting, and therefore, we discuss the limitations and concerns with the fit-on-test view. Our contributions also include: (1) enhancement of reliability diagrams with diagonal filling; (2) development of new calibration map families PL and PL3; and (3) an experimental study of which families perform strongly both as post-hoc calibrators and calibration evaluators.
Markus Kängsepp, Kaspar Valk, Meelis Kull
Mach. Learn.3
2025 Cost-sensitive classification with cost uncertainty: do we need surrogate losses?
abstract
Abstract In many binary classification applications, the costs of false positives and negatives are imbalanced. Furthermore, there is often uncertainty about the exact costs of these errors. A natural measure-of-interest to be minimised in such scenarios is the expected misclassification cost. We identify many situations where this measure has analytic gradients, and thus it can be used as a training loss and optimised directly using empirical risk minimisation. In particular, we derive such losses from the Beta, Gamma and Gaussian distributions to model different kinds of cost uncertainty. The Beta family includes commonly used losses such as cross-entropy, squared error and 0–1 loss as special cases. The question then arises as to when it is appropriate to directly optimize the measure-of-interest, versus using a standard surrogate like cross-entropy or focal loss during training. After revisiting the theory of surrogate losses, proper losses and cost-sensitive learning to obtain good candidate surrogates out of derived families, we conduct an empirical comparison of derived training losses that, to our knowledge, were never tried on deep neural networks before, with the aim to minimise cost-sensitive measures-of-interest. The findings show that using Beta losses in training leads to improved performance compared to traditional training objectives like cross-entropy, label smoothing, and focal loss. This improvement is seen not only in terms of misclassification cost metrics, but (perhaps surprisingly) also in conventional metrics such as accuracy, mean squared error, and the area under the ROC curve.
Viacheslav Komisarenko, Meelis Kull
Mach. Learn.2
2024 Cautious Calibration in Binary Classification
abstract
Being cautious is crucial for enhancing the trustworthiness of machine learning systems integrated into decision-making pipelines. Although calibrated probabilities help in optimal decision-making, perfect calibration remains unattainable, leading to estimates that fluctuate between under- and overconfidence. This becomes a critical issue in high-risk scenarios, where even occasional overestimation can lead to extreme expected costs. In these scenarios, it is important for each predicted probability to lean towards underconfidence, rather than just achieving an average balance. In this study, we introduce the novel concept of cautious calibration in binary classification. This approach aims to produce probability estimates that are intentionally underconfident for each predicted probability. We highlight the importance of this approach in a high-risk scenario and propose a theoretically grounded method for learning cautious calibration maps. Through experiments, we explore and compare our method to various approaches, including methods originally not devised for cautious calibration but applicable in this context. We show that our approach is the most consistent in providing cautious estimates. Our work establishes a strong baseline for further developments in this novel framework.
Mari-Liis Allikivi, Joonas Järve, Meelis Kull
ECAI3
2024 Improving Calibration by Relating Focal Loss, Temperature Scaling, and Properness
abstract
Proper losses such as cross-entropy incentivize classifiers to produce class probabilities that are well-calibrated on the training data. Due to the generalization gap, these classifiers tend to become overconfident on the test data, mandating calibration methods such as temperature scaling. The focal loss is not proper, but training with it has been shown to often result in classifiers that are better calibrated on test data. Our first contribution is a simple explanation about why focal loss training often leads to better calibration than cross-entropy training. For this, we prove that focal loss can be decomposed into a confidence-raising transformation and a proper loss. This is why focal loss pushes the model to provide under-confident predictions on the training data, resulting in being better calibrated on the test data, due to the generalization gap. Secondly, we reveal a strong connection between temperature scaling and focal loss through its confidence-raising transformation, which we refer to as the focal calibration map. Thirdly, we propose focal temperature scaling - a new post-hoc calibration method combining focal calibration and temperature scaling. Our experiments on three image classification datasets demonstrate that focal temperature scaling outperforms standard temperature scaling.
Viacheslav Komisarenko, Meelis Kull
ECAI2
2024 Evaluation of Trajectory Distribution Predictions with Energy Score
abstract
Predicting the future trajectory of surrounding objects is inherently uncertain and vital in the safe and reliable planning of autonomous systems such as in self-driving cars. Although trajectory prediction models have become increasingly sophisticated in dealing with the complexities of spatiotemporal data, the evaluation methods used to assess these models have not kept pace. "Minimum of N" is a common family of metrics used to assess the rich outputs of such models. We critically examine the Minimum of N within the proper scoring rules framework to show that it is not strictly proper and demonstrate how that could lead to a misleading assessment of multimodal trajectory predictions. As an alternative, we propose using Energy Score-based evaluation measures, leveraging their proven propriety for a more reliable evaluation of trajectory distribution predictions.
Novin Shahroudi, Mihkel Lepson, Meelis Kull
ICML3
2023 Generality-Training of a Classifier for Improved Calibration in Unseen Contexts
abstract
Abstract Artificial neural networks tend to output class probabilities that are miscalibrated, i.e., their reported uncertainty is not a very good indicator of how much we should trust the model. Consequently, methods have been developed to improve the model’s predictive uncertainty, both during training and post-hoc. Even if the model is calibrated on the domain used in training, it typically becomes over-confident when applied on slightly different target domains, e.g. due to perturbations or shifts in the data. The model can be recalibrated for a fixed list of target domains, but its performance can still be poor on unseen target domains. To address this issue, we propose a generality-training procedure that learns a modified head for the neural network to achieve better calibration generalization to new domains while retaining calibration performance on the given domains. This generality-head is trained on multiple domains using a new objective function with increased emphasis on the calibration loss compared to cross-entropy. Such training results in a more general model in the sense of not only better calibration but also better accuracy on unseen domains, as we demonstrate experimentally on multiple datasets. The code and supplementary for the paper is available ( https://github.com/bsl-traveller/CaliGen.git ).
Bhawani Shankar Leelar, Meelis Kull
ECML/PKDD (5)2
2023 Classifier calibration: a survey on how to assess and improve predicted class probabilities
abstract
Abstract This paper provides both an introduction to and a detailed overview of the principles and practice of classifier calibration. A well-calibrated classifier correctly quantifies the level of uncertainty or confidence associated with its instance-wise predictions. This is essential for critical applications, optimal decision making, cost-sensitive classification, and for some types of context change. Calibration research has a rich history which predates the birth of machine learning as an academic field by decades. However, a recent increase in the interest on calibration has led to new methods and the extension from binary to the multiclass setting. The space of options and issues to consider is large, and navigating it requires the right set of concepts and tools. We provide both introductory material and up-to-date technical details of the main concepts and methods, including proper scoring rules and other evaluation metrics, visualisation approaches, a comprehensive account of post-hoc calibration methods for binary and multiclass classification, and several advanced topics.
Telmo de Menezes e Silva Filho, Hao Song 0007, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Meelis Kull, Peter A. Flach
Mach. Learn.5
2021 Instance-based Label Smoothing For Better Calibrated Classification Networks
abstract
Label smoothing is widely used in deep neural networks for multi-class classification. While it enhances model generalization and reduces overconfidence by aiming to lower the probability for the predicted class, it distorts the predicted probabilities of other classes resulting in poor class-wise calibration. Another method for enhancing model generalization is self-distillation where the predictions of a teacher network trained with one-hot labels are used as the target for training a student network. We take inspiration from both label smoothing and self-distillation and propose two novel instance-based label smoothing approaches, where a teacher network trained with hard one-hot labels is used to determine the amount of per class smoothness applied to each instance. The assigned smoothing factor is non-uniformly distributed along with the classes according to their similarity with the actual class. Our methods show better generalization and calibration over standard label smoothing on various deep neural architectures and image classification datasets.
Mohamed Maher 0001, Meelis Kull
ICMLA2
2021 CRISP-DM Twenty Years Later: From Data Mining Processes to Data Science Trajectories
abstract
CRISP-DM(CRoss-Industry Standard Process for Data Mining) has its origins in the second half of the nineties and is thus about two decades old. According to many surveys and user polls it is still the de facto standard for developing data mining and knowledge discovery projects. However, undoubtedly the field has moved on considerably in twenty years, with data science now the leading term being favoured over data mining. In this paper we investigate whether, and in what contexts, CRISP-DM is still fit for purpose for data science projects. We argue that if the project is goal-directed and process-driven the process model view still largely holds. On the other hand, when data science projects become more exploratory the paths that the project can take become more varied, and a more flexible model is called for. We suggest what the outlines of such a trajectory-based model might look like and how it can be used to categorise data science projects (goal-directed, exploratory or data management). We examine seven real-life exemplars where exploratory activities play an important role and compare them against 51 use cases extracted from the NIST Big Data Public Working Group. We anticipate this categorisation can help project planning in terms of time and cost characteristics.
Fernando Martínez-Plumed, Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Meelis Kull, Nicolas Lachiche, María José Ramírez-Quintana, Peter A. Flach
IEEE Trans. Knowl. Data Eng.5
2019 Distribution calibration for regression
abstract
We are concerned with obtaining well-calibrated output distributions from regression models. Such distributions allow us to quantify the uncertainty that the model has regarding the predicted target value. We introduce the novel concept of distribution calibration, and demonstrate its advantages over the existing definition of quantile calibration. We further propose a post-hoc approach to improving the predictions from previously trained regression models, using multi-output Gaussian Processes with a novel Beta link function. The proposed method is experimentally verified on a set of common regression models and shows improvements for both distribution-level and quantile-level calibration.
Hao Song 0007, Tom Diethe, Meelis Kull, Peter A. Flach
ICML3
2019 Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with Dirichlet calibration
abstract
Class probabilities predicted by most multiclass classifiers are uncalibrated, often tending towards over-confidence. With neural networks, calibration can be improved by temperature scaling, a method to learn a single corrective multiplicative factor for inputs to the last softmax layer. On non-neural models the existing methods apply binary calibration in a pairwise or one-vs-rest fashion. We propose a natively multiclass calibration method applicable to classifiers from any model class, derived from Dirichlet distributions and generalising the beta calibration method from binary classification. It is easily implemented with neural nets since it is equivalent to log-transforming the uncalibrated probabilities, followed by one linear layer and softmax. Experiments demonstrate improved probabilistic predictions according to multiple measures (confidence-ECE, classwise-ECE, log-loss, Brier score) across a wide range of datasets and classifiers. Parameters of the learned Dirichlet calibration map provide insights to the biases in the uncalibrated model.
Meelis Kull, Miquel Perelló-Nieto, Markus Kängsepp, Telmo de Menezes e Silva Filho, Hao Song 0007, Peter A. Flach
NeurIPS1
2019 Non-parametric Bayesian Isotonic Calibration: Fighting Over-Confidence in Binary Classification
Mari-Liis Allikivi, Meelis Kull
ECML/PKDD (2)2
2019 Shift Happens: Adjusting Classifiers
Theodore James Thibault Heiser, Mari-Liis Allikivi, Meelis Kull
ECML/PKDD (2)3
2018 Releasing eHealth Analytics into the Wild: Lessons Learnt from the SPHERE Project
abstract
The SPHERE project is devoted to advancing eHealth in a smart-home context, and supports full-scale sensing and data analysis to enable a generic healthcare service. We describe, from a data-science perspective, our experience of taking the system out of the laboratory into more than thirty homes in Bristol, UK. We describe the infrastructure and processes that had to be developed along the way, describe how we train and deploy Machine Learning systems in this context, and give a realistic appraisal of the state of the deployed systems.
Tom Diethe, Mike Holmes, Meelis Kull, Miquel Perelló-Nieto, Kacper Sokol, Hao Song 0007, Emma Tonkin, Niall Twomey, Peter A. Flach
KDD3
2017 Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers
abstract
For optimal decision making under variable class distributions and misclassification costs a classifier needs to produce well-calibrated estimates of the posterior probability. Isotonic calibration is a powerful non-parametric method that is however prone to overfitting on smaller datasets; hence a parametric method based on the logistic curve is commonly used. While logistic calibration is designed for normally distributed per-class scores, we demonstrate experimentally that many classifiers including Naive Bayes and Adaboost suffer from a particular distortion where these score distributions are heavily skewed. In such cases logistic calibration can easily yield probability estimates that are worse than the original scores. Moreover, the logistic curve family does not include the identity function, and hence logistic calibration can easily uncalibrate a perfectly calibrated classifier. In this paper we solve all these problems with a richer class of calibration maps based on the beta distribution. We derive the method from first principles and show that fitting it is as easy as fitting a logistic curve. Extensive experiments show that beta calibration is superior to logistic calibration for Naive Bayes and Adaboost.
Meelis Kull, Telmo de Menezes e Silva Filho, Peter A. Flach
AISTATS1
2016 Declaratively Capturing Local Label Correlations with Multi-Label Trees
abstract
The goal of multi-label classification is to predict multiple labels per data point simultaneously. Real-world applications tend to have high-dimensional label spaces, employing hundreds or even thousands of labels. While these labels could be predicted separately, by capturing label correlation we might achieve better predictive performance. In contrast with previous attempts in the literature that have modelled label correlations globally, this paper proposes a novel algorithm to model correlations and cluster labels locally. LaCovaC is a multi-label decision tree classifier that clusters labels into several dependent subsets at various points during training. The clusters are obtained locally by identifying the conditionally-dependent labels in localised regions of the feature space using the label correlation matrix. LaCovaC interleaves between two main decisions on the label matrix with training instances in rows and labels in columns: splitting this matrix vertically by partitioning the labels into subsets, or splitting it horizontally using features in the conventional way. Experiments on 13 benchmark datasets demonstrate that our proposal achieves competitive performance over a wide range of evaluation metrics when compared with the state-of-the-art multi-label classifiers.
Reem Alotaibi, Meelis Kull, Peter A. Flach
ECAI2
2016 Background Check: A General Technique to Build More Reliable and Versatile Classifiers
abstract
We introduce a powerful technique to make classifiers more reliable and versatile. Background Check equips classifiers with the ability to assess the difference of unlabelled test data from the training data. In particular, Background Check gives classifiers the capability to (i) perform cautious classification with a reject option, (ii) identify outliers, and (iii) better assess the confidence in their predictions. We derive the method from first principles and consider four particular relationships between background and foreground distributions. One of these assumes an affine relationship with two parameters, and we show how this bivariate parameter space naturally interpolates between the above capabilities. We demonstrate the versatility of the approach by comparing it experimentally with published special-purpose solutions for outlier detection and confident classification on 41 benchmark datasets. Results show that Background Check can match and in many cases surpass the performances of specialised approaches.
Miquel Perelló-Nieto, Telmo de Menezes e Silva Filho, Meelis Kull, Peter A. Flach
ICDM3
2016 Subgroup Discovery with Proper Scoring Rules
Hao Song 0007, Meelis Kull, Peter A. Flach, Georgios Kalogridis
ECML/PKDD (2)2
2016 Cost-sensitive boosting algorithms: Do we really need them?
abstract
We provide a unifying perspective for two decades of work on cost-sensitive Boosting algorithms. When analyzing the literature 1997–2016, we find 15 distinct cost-sensitive variants of the original algorithm; each of these has its own motivation and claims to superiority—so who should we believe? In this work we critique the Boosting literature using four theoretical frameworks: Bayesian decision theory, the functional gradient descent view, margin theory, and probabilistic modelling. Our finding is that only three algorithms are fully supported—and the probabilistic model view suggests that all require their outputs to be calibrated for best performance. Experiments on 18 datasets across 21 degrees of imbalance support the hypothesis—showing that once calibrated, they perform equivalently, and outperform all others. Our final recommendation—based on simplicity, flexibility and performance—is to use the original Adaboost algorithm with a shifted decision threshold and calibrated probability estimates.
Nikolaos Nikolaou, Narayanan Unny Edakunni, Meelis Kull, Peter A. Flach, Gavin Brown 0001
Mach. Learn.3
2015 Reframing in Frequent Pattern Mining
abstract
Mining frequent patterns is a crucial task in data mining. Most of the existing frequent pattern mining methods find the complete set of frequent patterns from a given dataset. However, in real-life scenarios we often need to predict the future frequent patterns for different tasks such as business policy making, web page recommendation, stock-market behavior and road traffic analysis. Predicting future frequent patterns from the currently available set of frequent patterns is challenging due to dataset shift where data distributions may change from one dataset to another. In this paper, we propose a new approach called reframing in frequent pattern mining to solve this task. Moreover, we experimentally show the existence of dataset shift in two real-life transactional datasets and the capability of our approach to handle these unknown shifts.
Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Nicolas Lachiche, Meelis Kull, Peter A. Flach
ICTAI4
2015 Precision-Recall-Gain Curves: PR Analysis Done Right
abstract
Precision-Recall analysis abounds in applications of binary classification where true negatives do not add value and hence should not affect assessment of the classifier's performance. Perhaps inspired by the many advantages of receiver operating characteristic (ROC) curves and the area under such curves for accuracy-based performance assessment, many researchers have taken to report Precision-Recall (PR) curves and associated areas as performance metric. We demonstrate in this paper that this practice is fraught with difficulties, mainly because of incoherent scale assumptions -- e.g., the area under a PR curve takes the arithmetic mean of precision values whereas the $F_{\beta}$ score applies the harmonic mean. We show how to fix this by plotting PR curves in a different coordinate system, and demonstrate that the new Precision-Recall-Gain curves inherit all key advantages of ROC curves. In particular, the area under Precision-Recall-Gain curves conveys an expected $F_1$ score on a harmonic scale, and the convex hull of a Precision-Recall-Gain curve allows us to calibrate the classifier's scores so as to determine, for each operating point on the convex hull, the interval of $\beta$ values for which the point optimises $F_{\beta}$. We demonstrate experimentally that the area under traditional PR curves can easily favour models with lower expected $F_1$ score than others, and so the use of Precision-Recall-Gain curves will result in better model selection.
Peter A. Flach, Meelis Kull
NIPS2
2015 Versatile Decision Trees for Learning Over Multiple Contexts
Reem Alotaibi, Ricardo B. C. Prudêncio, Meelis Kull, Peter A. Flach
ECML/PKDD (1)3
2015 Novel Decompositions of Proper Scoring Rules for Classification: Score Adjustment as Precursor to Calibration
Meelis Kull, Peter A. Flach
ECML/PKDD (1)1
2014 LaCova: A Tree-Based Multi-label Classifier Using Label Covariance as Splitting Criterion
abstract
Dealing with multiple labels is a supervised learning problem of increasing importance. Multi-label classifiers face the challenge of exploiting correlations between labels. While in existing work these correlations are often modelled globally, in this paper we use the divide-and-conquer approach of decision trees which enables taking local decisions about how best to model label dependency. The resulting algorithm establishes a tree-based multi-label classifier called LaCova which dynamically interpolates between two well-known baseline methods: Binary Relevance, which assumes all labels independent, and Label Power set, which learns the joint label distribution. The key idea is a splitting criterion based on the label covariance matrix at that node, which allows us to choose between a horizontal split (branching on a feature) and a vertical split (separating the labels). Empirical results on 12 data sets show strong performance of the proposed method, particularly on data sets with hundreds of labels.
Reem Alotaibi, Meelis Kull, Peter A. Flach
ICMLA2
2014 Reliability Maps: A Tool to Enhance Probability Estimates and Improve Classification Accuracy
Meelis Kull, Peter A. Flach
ECML/PKDD (2)1
2014 Rate-Oriented Point-Wise Confidence Bounds for ROC Curves
Louise A. C. Millard, Meelis Kull, Peter A. Flach
ECML/PKDD (2)2