Juhani Kivimäki

dblp:353/5194 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-9673-9760ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 72% Transfer learning and domain adaptation · 28%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
1.122025
Estimating Model Performance Under Covariate Shift Without Labels · NeurIPS 2025
Confidence-based Estimators for Predictive Performance in Model Monitoring (Abstract Reprint) · IJCAI 2025
Machine learning › Trustworthy machine learning
accuracy estimation
0.912025
Estimating Model Performance Under Covariate Shift Without Labels · NeurIPS 2025
Machine learning › Trustworthy machine learning
model monitoring
0.912025
Confidence-based Estimators for Predictive Performance in Model Monitoring (Abstract Reprint) · IJCAI 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
Confidence-based Estimators for Predictive Performance in Model Monitoring (Abstract Reprint) · IJCAI 2025
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.312025
Estimating Model Performance Under Covariate Shift Without Labels · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

unbiased estimator · 0.9probabilistic adaptive performance estimation · 0.9consistency analysis · 0.9confusion matrix · 0.9confidence interval estimation · 0.9
YearPublicationVenuePosition
2026 Performance Estimation in Binary Classification Using Calibrated Confidence
abstract
Abstract Model monitoring is a critical component of the machine learning lifecycle, safeguarding against undetected drops in the model’s performance after deployment. Traditionally, performance monitoring has required access to ground truth labels, which are not always readily available. This can result in unacceptable latency or render performance monitoring altogether impossible. Recently, methods designed to estimate the accuracy of classifier models without access to labels have shown promising results. However, there are various other metrics that might be more suitable for assessing model performance in many cases. Until now, none of these important metrics has received similar interest from the scientific community. In this work, we address this gap by presenting Confidence-based Performance estimation (CBPE), a novel method that can estimate any binary classification metric defined using the confusion matrix. In particular, we choose four metrics from this large family: accuracy, precision, recall, and $$\hbox {F}_1$$ , to demonstrate our method. CBPE treats the elements of the confusion matrix as random variables and leverages calibrated confidence scores of the model to estimate their distributions. The desired metric is then also treated as a random variable, whose full probability distribution can be derived from the estimated confusion matrix. CBPE is shown to produce estimates that come with strong theoretical guarantees and valid confidence intervals.
Juhani Kivimäki, Jakub Bialek, Wojtek Kuberski, Jukka K. Nurminen
Mach. Learn.1
2025 Confidence-based Estimators for Predictive Performance in Model Monitoring (Abstract Reprint)
abstract
After a machine learning model has been deployed into production, its predictive performance needs to be monitored. Ideally, such monitoring can be carried out by comparing the model’s predictions against ground truth labels. For this to be possible, the ground truth labels must be available relatively soon after inference. However, there are many use cases where ground truth labels are available only after a significant delay, or in the worst case, not at all. In such cases, directly monitoring the model’s predictive performance is impossible. Recently, novel methods for estimating the predictive performance of a model when ground truth is unavailable have been developed. Many of these methods leverage model confidence or other uncertainty estimates and are experimentally compared against a naive baseline method, namely Average Confidence (AC), which estimates model accuracy as the average of confidence scores for a given set of predictions. However, until now the theoretical properties of the AC method have not been properly explored. In this paper, we bridge this gap by reviewing the AC method and show that under certain general assumptions, it is an unbiased and consistent estimator of model accuracy. We also augment the AC method by deriving valid confidence intervals for the estimates it produces. These contributions elevate AC from an ad-hoc estimator to a principled one, encouraging its use in practice. We complement our theoretical results with empirical experiments, comparing AC against more complex estimators in a monitoring setting under covariate shift. We conduct our experiments using synthetic datasets, which allow for full control over the nature of the shift. Our experiments with binary classifiers show that the AC method is able to beat other estimators in many cases. However, the comparative quality of the different estimators is found to be heavily case-dependent.
Juhani Kivimäki, Jakub Bialek, Jukka K. Nurminen, Wojtek Kuberski
IJCAI1
2025 Estimating Model Performance Under Covariate Shift Without Labels
abstract
After deployment, machine learning models often experience performance degradation due to shifts in data distribution. It is challenging to assess post-deployment performance accurately when labels are missing or delayed. Existing proxy methods, such as data drift detection, fail to measure the effects of these shifts adequately. To address this, we introduce a new method for evaluating binary classification models on unlabeled tabular data that accurately estimates model performance under covariate shift and call it Probabilistic Adaptive Performance Estimation (PAPE). It can be applied to any performance metric defined with elements of the confusion matrix. Crucially, PAPE operates independently of the original model, relying only on its predictions and probability estimates, and does not need any assumptions about the nature of covariate shift, learning directly from data instead. We tested PAPE using over 900 dataset-model combinations from US census data, assessing its performance against several benchmarks through various metrics. Our findings show that PAPE outperforms other methodologies, making it a superior choice for estimating the performance of binary classification models.
Jakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos Perrakis
NeurIPS2
2025 Confidence-based Estimators for Predictive Performance in Model Monitoring
abstract
After a machine learning model has been deployed into production, its predictive performance needs to be monitored. Ideally, such monitoring can be carried out by comparing the model’s predictions against ground truth labels. For this to be possible, the ground truth labels must be available relatively soon after inference. However, there are many use cases where ground truth labels are available only after a significant delay, or in the worst case, not at all. In such cases, directly monitoring the model’s predictive performance is impossible. Recently, novel methods for estimating the predictive performance of a model when ground truth is unavailable have been developed. Many of these methods leverage model confidence or other uncertainty estimates and are experimentally compared against a naive baseline method, namely Average Confidence (AC), which estimates model accuracy as the average of confidence scores for a given set of predictions. However, until now the theoretical properties of the AC method have not been properly explored. In this paper, we bridge this gap by reviewing the AC method and show that under certain general assumptions, it is an unbiased and consistent estimator of model accuracy. We also augment the AC method by deriving valid confidence intervals for the estimates it produces. These contributions elevate AC from an ad-hoc estimator to a principled one, encouraging its use in practice. We complement our theoretical results with empirical experiments, comparing AC against more complex estimators in a monitoring setting under covariate shift. We conduct our experiments using synthetic datasets, which allow for full control over the nature of the shift. Our experiments with binary classifiers show that the AC method is able to beat other estimators in many cases. However, the comparative quality of the different estimators is found to be heavily case-dependent.
Juhani Kivimäki, Jukka K. Nurminen, Jakub Bialek, Wojtek Kuberski
J. Artif. Intell. Res.1
2023 Failure Prediction in 2D Document Information Extraction with Calibrated Confidence Scores
abstract
Modern machine learning models can achieve impressive results in many tasks, but often fail to express reliably how confident they are with their predictions. In an industrial setting, the end goal is usually not a prediction of a model, but a decision based on that prediction. It is often not sufficient to generate high-accuracy predictions on average. One also needs to estimate the uncertainty and risks involved when making related decisions. Thus, having reliable and calibrated uncertainty estimates is highly useful for any model used in automated decision-making.In this paper, we present a case study, where we propose a novel method to improve the uncertainty estimates of an in-production machine learning model operating in an industrial setting with real-life data. This model is used by Basware, a Finnish software company, to extract information from invoices in the form of machine-readable PDFs. The solution we propose is shown to produce calibrated confidence estimates, which outperform legacy estimates on several relevant metrics, increasing coverage of automated invoices from 65.6% to 73.2% with no increase in error rate.
Juhani Kivimäki, Aleksey Lebedev, Jukka K. Nurminen
COMPSAC1
2023 Anomaly Localization in Audio via Feature Pyramid Matching
abstract
Sound anomaly detection is a task that aims at identifying unusual or abnormal sounds within audio data. These sounds could be caused by different factors, such as background noise, equipment malfunctions, or unexpected events. Anomaly detection in sound is a well-studied topic, with a lot of research being done in the field. Anomaly localization refers to the process of identifying the specific location or region within a sample where an anomaly (or outlier) occurs. When applied to audio signals, anomaly localization can involve analyzing the spectral content of the sound to detect regions that deviate from the typical or expected pattern.In this study, we present a simple yet effective model based on the Student-Teacher Feature Pyramid Matching Method for locating anomalies in audio data. Utilizing the MIMII dataset by augmenting it with synthetic anomalies, we evaluate the method’s accuracy. Our results demonstrate that the proposed model can accurately locate artificially created anomalies within the spectrograms, both in terms of time and frequency. This approach offers a promising solution for identifying and determining the precise location of anomalies in various audio applications.
Jorma Valjakka, Juha Mylläri, Lalli Myllyaho, Juhani Kivimäki, Jukka K. Nurminen
COMPSAC4