Sebastian Gruber 0001

dblp:72/11348-1 · also Sebastian G. Gruber, Sebastian Gregor Gruber · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-5714-1870ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 82% Generative modeling · 11% Transfer learning and domain adaptation · 7%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
uncertainty estimation
2.332026
Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral Approach · AAAI 2026
A Bias-Variance-Covariance Decomposition of Kernel Scores for Generative Models · ICML 2024
Post-Hoc Uncertainty Calibration for Domain Drift Scenarios · CVPR 2021
Machine learning › Trustworthy machine learning › uncertainty estimation
large language model uncertainty
1.012026
Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral Approach · AAAI 2026
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty decomposition
1.012026
Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral Approach · AAAI 2026
Machine learning › Generative modeling
generative model evaluation
0.812024
A Bias-Variance-Covariance Decomposition of Kernel Scores for Generative Models · ICML 2024
Machine learning › Trustworthy machine learning › calibration
recalibration
0.612022
Better Uncertainty Calibration via Proper Scores for Classification and Beyond · NeurIPS 2022
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration
0.612022
Better Uncertainty Calibration via Proper Scores for Classification and Beyond · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation
domain shift
0.512021
Post-Hoc Uncertainty Calibration for Domain Drift Scenarios · CVPR 2021
Machine learning › Trustworthy machine learning › calibration
post-hoc calibration
0.512021
Post-Hoc Uncertainty Calibration for Domain Drift Scenarios · CVPR 2021

Methods — techniques the papers use, named apart from their topics

von neumann entropy · 1.0spectral methods · 1.0semantic similarity · 1.0proper scores · 0.6post-hoc calibration · 0.5perturbation-based calibration · 0.5
YearPublicationVenuePosition
2026 Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral Approach
abstract
As Large Language Models (LLMs) are increasingly integrated in diverse applications, obtaining reliable measures of their predictive uncertainty has become critically important. A precise distinction between aleatoric uncertainty, arising from inherent ambiguities within input data, and epistemic uncertainty, originating exclusively from model limitations, is essential to effectively address each uncertainty source. In this paper, we introduce Spectral Uncertainty, a novel approach to quantifying and decomposing uncertainties in LLMs. Leveraging the Von Neumann entropy from quantum information theory, Spectral Uncertainty provides a rigorous theoretical foundation for separating total uncertainty into distinct aleatoric and epistemic components. Unlike existing baseline methods, our approach incorporates a fine-grained representation of semantic similarity, enabling nuanced differentiation among various semantic interpretations in model responses. Empirical evaluations demonstrate that Spectral Uncertainty outperforms state-of-the-art methods in estimating both aleatoric and total uncertainty across diverse models and benchmark datasets.
Nassim Walha, Sebastian Gruber 0001, Thomas Decker 0004, Yinchong Yang, Alireza Javanmardi, Eyke Hüllermeier, Florian Buettner 0001
AAAI2
2025 Application-driven validation of posteriors in inverse problems
abstract
Current deep learning-based solutions for image analysis tasks are commonly incapable of handling problems to which multiple different plausible solutions exist. In response, posterior-based methods such as conditional Diffusion Models and Invertible Neural Networks have emerged; however, their translation is hampered by a lack of research on adequate validation. In other words, the way progress is measured often does not reflect the needs of the driving practical application. Closing this gap in the literature, we present the first systematic framework for the application-driven validation of posterior-based methods in inverse problems. As a methodological novelty, it adopts key principles from the field of object detection validation, which has a long history of addressing the question of how to locate and match multiple object instances in an image. Treating modes as instances enables us to perform mode-centric validation, using well-interpretable metrics from the application perspective. We demonstrate the value of our framework through instantiations for a synthetic toy example and two medical vision use cases: pose estimation in surgery and imaging-based quantification of functional tissue parameters for diagnostics. Our framework offers key advantages over common approaches to posterior validation in all three examples and could thus revolutionize performance assessment in inverse problems.
Tim Adler, Jan-Hinrich Nölke, Annika Reinke, Minu Tizabi, Sebastian Gruber 0001, Dasha Trofimova, Lynton Ardizzone, Paul F. Jaeger, Florian Buettner 0001, Ullrich Köthe, Lena Maier-Hein
Medical Image Anal.5
2024 Consistent and Asymptotically Unbiased Estimation of Proper Calibration Errors
abstract
Proper scoring rules evaluate the quality of probabilistic predictions, playing an essential role in the pursuit of accurate and well-calibrated models. Every proper score decomposes into two fundamental components – proper calibration error and refinement – utilizing a Bregman divergence. While uncertainty calibration has gained significant attention, current literature lacks a general estimator for these quantities with known statistical properties. To address this gap, we propose a method that allows consistent, and asymptotically unbiased estimation of all proper calibration errors and refinement terms. In particular, we introduce Kullback-Leibler calibration error, induced by the commonly used cross-entropy loss. As part of our results, we prove the relation between refinement and f-divergences, which implies information monotonicity in neural networks, regardless of which proper scoring rule is optimized. Our experiments validate empirically the claimed properties of the proposed estimator and suggest that the selection of a post-hoc calibration method should be determined by the particular calibration error of interest.
Teodora Popordanoska, Sebastian Gruber 0001, Aleksei Tiulpin, Florian Buettner 0001, Matthew B. Blaschko
AISTATS2
2024 A Bias-Variance-Covariance Decomposition of Kernel Scores for Generative Models
abstract
Generative models, like large language models, are becoming increasingly relevant in our daily lives, yet a theoretical framework to assess their generalization behavior and uncertainty does not exist. Particularly, the problem of uncertainty estimation is commonly solved in an ad-hoc and task-dependent manner. For example, natural language approaches cannot be transferred to image generation. In this paper, we introduce the first bias-variance-covariance decomposition for kernel scores. This decomposition represents a theoretical framework from which we derive a kernel-based variance and entropy for uncertainty estimation. We propose unbiased and consistent estimators for each quantity which only require generated samples but not the underlying model itself. Based on the wide applicability of kernels, we demonstrate our framework via generalization and uncertainty experiments for image, audio, and language generation. Specifically, kernel entropy for uncertainty estimation is more predictive of performance on CoQA and TriviaQA question answering datasets than existing baselines and can also be applied to closed-source models.
Sebastian Gruber 0001, Florian Buettner 0001
ICML1
2023 Uncertainty Estimates of Predictions via a General Bias-Variance Decomposition
abstract
Reliably estimating the uncertainty of a prediction throughout the model lifecycle is crucial in many safety-critical applications. The most common way to measure this uncertainty is via the predicted confidence. While this tends to work well for in-domain samples, these estimates are unreliable under domain drift and restricted to classification. Alternatively, proper scores can be used for most predictive tasks but a bias-variance decomposition for model uncertainty does not exist in the current literature. In this work we introduce a general bias-variance decomposition for proper scores, giving rise to the Bregman Information as the variance term. We discover how exponential families and the classification log-likelihood are special cases and provide novel formulations. Surprisingly, we can express the classification case purely in the logit space. We showcase the practical relevance of this decomposition on several downstream tasks, including model ensembles and confidence regions. Further, we demonstrate how different approximations of the instance-level Bregman Information allow reliable out-of-distribution detection for all degrees of domain drift.
Sebastian Gruber 0001, Florian Buettner 0001
AISTATS1
2022 Better Uncertainty Calibration via Proper Scores for Classification and Beyond
abstract
With model trustworthiness being crucial for sensitive real-world applications, practitioners are putting more and more focus on improving the uncertainty calibration of deep neural networks.Calibration errors are designed to quantify the reliability of probabilistic predictions but their estimators are usually biased and inconsistent.In this work, we introduce the framework of \textit{proper calibration errors}, which relates every calibration error to a proper score and provides a respective upper bound with optimal estimation properties.This relationship can be used to reliably quantify the model calibration improvement.We theoretically and empirically demonstrate the shortcomings of commonly used estimators compared to our approach.Due to the wide applicability of proper scores, this gives a natural extension of recalibration beyond classification.
Sebastian Gruber 0001, Florian Buettner 0001
NeurIPS1
2021 Post-Hoc Uncertainty Calibration for Domain Drift Scenarios
abstract
We address the problem of uncertainty calibration. While standard deep neural networks typically yield uncalibrated predictions, calibrated confidence scores that are representative of the true likelihood of a prediction can be achieved using post-hoc calibration methods. However, to date, the focus of these approaches has been on in-domain calibration. Our contribution is two-fold. First, we show that existing post-hoc calibration methods yield highly over-confident predictions under domain shift. Second, we introduce a simple strategy where perturbations are applied to samples in the validation set before performing the post-hoc calibration step. In extensive experiments, we demonstrate that this perturbation step results in substantially better calibration under domain shift on a wide range of architectures and modelling tasks.
Christian Tomani, Sebastian Gruber 0001, Muhammed Ebrar Erdem, Daniel Cremers, Florian Buettner 0001
CVPR2