EDBT 2026 Demo / reviewers in the wild / expert
Alvaro Henrique Chaim Correia
dblp:222/2873 · also Alvaro H. C. Correia
· DBLP profile ↗
9ranked-venue papers
8as first author
5since 2021 · last 2025
0000-0001-5291-0653ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 54% Probabilistic and Bayesian machine learning · 27% Transfer learning and domain adaptation · 8% | |
| Theoretical computer science
1 paper |
Information theory · 50% Mathematical optimization · 50% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
2.5 | 3 | 2025 | Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled Data · NeurIPS 2025 Approximating Full Conformal Prediction for Neural Network Regression with Gauss-Newton Influence · ICLR 2025 An Information Theoretic Perspective on Conformal Prediction · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
2.5 | 3 | 2025 | Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled Data · NeurIPS 2025 Approximating Full Conformal Prediction for Neural Network Regression with Gauss-Newton Influence · ICLR 2025 An Information Theoretic Perspective on Conformal Prediction · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › tractable probabilistic model
probabilistic circuit |
1.1 | 2 | 2023 | Continuous Mixtures of Tractable Probabilistic Models · AAAI 2023 Joints in Random Forests · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift |
0.9 | 1 | 2025 | Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled Data · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.9 | 1 | 2025 | Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled Data · NeurIPS 2025 |
Information theory › information measures › entropy
conditional entropy |
0.8 | 1 | 2024 | An Information Theoretic Perspective on Conformal Prediction · NeurIPS 2024 |
Mathematical optimization › uncertainty quantification
uncertainty estimation |
0.8 | 1 | 2024 | An Information Theoretic Perspective on Conformal Prediction · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
continuous latent-variable models |
0.7 | 1 | 2023 | Continuous Mixtures of Tractable Probabilistic Models · AAAI 2023 |
Machine learning › Probabilistic and Bayesian machine learning
tractable probabilistic model |
0.7 | 1 | 2023 | Continuous Mixtures of Tractable Probabilistic Models · AAAI 2023 |
Machine learning › Kernel, tree and ensemble methods
decision tree |
0.4 | 1 | 2020 | Joints in Random Forests · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning
generative and discriminative models |
0.4 | 1 | 2020 | Joints in Random Forests · NeurIPS 2020 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest |
0.4 | 1 | 2020 | Joints in Random Forests · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.4 | 1 | 2019 | Human-in-the-Loop Feature Selection · AAAI 2019 |
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals |
0.3 | 1 | 2025 | Approximating Full Conformal Prediction for Neural Network Regression with Gauss-Newton Influence · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
marginalization |
0.1 | 1 | 2020 | Joints in Random Forests · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning
missing data |
0.1 | 1 | 2020 | Joints in Random Forests · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
information-theoretic inequalities · 1.5federated learning · 1.5probabilistic circuits · 1.1optimal transport · 0.9laplace approximation · 0.9gauss-newton influence · 0.9reinforcement learning · 0.8numerical integration · 0.7bayes consistency · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Approximating Full Conformal Prediction for Neural Network Regression with Gauss-Newton InfluenceabstractUncertainty quantification is an important prerequisite for the deployment of deep learning models in safety-critical areas. Yet, this hinges on the uncertainty estimates being useful to the extent the prediction intervals are well-calibrated and sharp. In the absence of inherent uncertainty estimates (e.g. pretrained models predicting only point estimates), popular approaches that operate post-hoc include Laplace’s method and split conformal prediction (split-CP). However, Laplace’s method can be miscalibrated when the model is misspecified and split-CP requires sample splitting, and thus comes at the expense of statistical efficiency. In this work, we construct prediction intervals for neural network regressors post-hoc without held-out data. This is achieved by approximating the full conformal prediction method (full-CP). Whilst full-CP nominally requires retraining the model for every test point and candidate label, we propose to train just once and locally perturb model parameters using Gauss-Newton influence to approximate the effect of retraining. Coupled with linearization of the network, we express the absolute residual nonconformity score as a piecewise linear function of the candidate label allowing for an efficient procedure that avoids the exhaustive search over the output space. On standard regression benchmarks and bounding box localization, we show the resulting prediction intervals are locally-adaptive and often tighter than those of split-CP. Dharmesh Tailor, Alvaro Henrique Chaim Correia, Eric T. Nalisnick, Christos Louizos |
ICLR | 2 |
| 2025 | Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled DataabstractConformal prediction is a distribution-free uncertainty quantification method that has gained popularity in the machine learning community due to its finite-sample guarantees and ease of use. Its most common variant, dubbed split conformal prediction, is also computationally efficient as it boils down to collecting statistics of the model predictions on some calibration data not yet seen by the model. Nonetheless, these guarantees only hold if the calibration and test data are exchangeable, a condition that is difficult to verify and often violated in practice due to so-called distribution shifts. The literature is rife with methods to mitigate the loss in coverage in this non-exchangeable setting, but these methods require some prior information on the type of distribution shift to be expected at test time. In this work, we study this problem via a new perspective, through the lens of optimal transport, and show that it is possible to estimate the loss in coverage and mitigate arbitrary distribution shifts, offering a principled and broadly applicable solution. Alvaro Henrique Chaim Correia, Christos Louizos |
NeurIPS | 1 |
| 2024 | An Information Theoretic Perspective on Conformal PredictionabstractConformal Prediction (CP) is a distribution-free uncertainty estimation framework that constructs prediction sets guaranteed to contain the true answer with a user-specified probability. Intuitively, the size of the prediction set encodes a general notion of uncertainty, with larger sets associated with higher degrees of uncertainty. In this work, we leverage information theory to connect conformal prediction to other notions of uncertainty. More precisely, we prove three different ways to upper bound the intrinsic uncertainty, as described by the conditional entropy of the target variable given the inputs, by combining CP with information theoretical inequalities. Moreover, we demonstrate two direct and useful applications of such connection between conformal prediction and information theory: (i) more principled and effective conformal training objectives that generalize previous approaches and enable end-to-end training of machine learning models from scratch, and (ii) a natural mechanism to incorporate side information into conformal prediction. We empirically validate both applications in centralized and federated learning settings, showing our theoretical results translate to lower inefficiency (average prediction set size) for popular CP methods. Alvaro Henrique Chaim Correia, Fabio Valerio Massoli, Christos Louizos, Arash Behboodi |
NeurIPS | 1 |
| 2023 | Continuous Mixtures of Tractable Probabilistic ModelsabstractProbabilistic models based on continuous latent spaces, such as variational autoencoders, can be understood as uncountable mixture models where components depend continuously on the latent code. They have proven to be expressive tools for generative and probabilistic modelling, but are at odds with tractable probabilistic inference, that is, computing marginals and conditionals of the represented probability distribution. Meanwhile, tractable probabilistic models such as probabilistic circuits (PCs) can be understood as hierarchical discrete mixture models, and thus are capable of performing exact inference efficiently but often show subpar performance in comparison to continuous latent-space models. In this paper, we investigate a hybrid approach, namely continuous mixtures of tractable models with a small latent dimension. While these models are analytically intractable, they are well amenable to numerical integration schemes based on a finite set of integration points. With a large enough number of integration points the approximation becomes de-facto exact. Moreover, for a finite set of integration points, the integration method effectively compiles the continuous mixture into a standard PC. In experiments, we show that this simple scheme proves remarkably effective, as PCs learnt this way set new state of the art for tractable models on many standard density estimation benchmarks. Alvaro Henrique Chaim Correia, Gennaro Gala, Erik Quaeghebeur, Cassio P. de Campos, Robert Peharz |
AAAI | 1 |
| 2023 | Neural Simulated AnnealingabstractSimulated annealing (SA) is a stochastic global optimisation metaheuristic applicable to a wide range of discrete and continuous variable problems. Despite its simplicity, SA hinges on carefully handpicked components, viz. proposal distribution and annealing schedule, that often have to be fine tuned to individual problem instances. In this work, we seek to make SA more effective and easier to use by framing its proposal distribution as a reinforcement learning policy that can be optimised for higher solution quality given a computational budget. The result is Neural SA, a competitive and general machine learning method for combinatorial optimisation that is efficient, and easy to design and train. We show Neural SA with such a learnt proposal distribution, parametrised by small equivariant neural networks, outperforms SA baselines on several problems: Rosenbrock’s function and the Knapsack, Bin Packing and Travelling Salesperson problems. We also show Neural SA scales well to large problems (generalising to much larger instances than those seen during training) while getting comparable performance to popular off-the-shelf solvers and machine learning methods in terms of solution quality and wall-clock time. Alvaro Henrique Chaim Correia, Daniel E. Worrall, Roberto Bondesan |
AISTATS | 1 |
| 2020 | On Pruning for Score-Based Bayesian Network Structure LearningabstractMany algorithms for score-based Bayesian network structure learning (BNSL), in particular exact ones, take as input a collection of potentially optimal parent sets for each variable in the data. Constructing such collections naively is computationally intensive since the number of parent sets grows exponentially with the number of variables. Thus, pruning techniques are not only desirable but essential. While good pruning rules exist for the Bayesian Information Criterion (BIC), current results for the Bayesian Dirichlet equivalent uniform (BDeu) score reduce the search space very modestly, hampering the use of the (often preferred) BDeu. We derive new non-trivial theoretical upper bounds for the BDeu score that considerably improve on the state-of-the-art. Since the new bounds are mathematically proven to be tighter than previous ones and at little extra computational cost, they are a promising addition to BNSL methods. Alvaro Henrique Chaim Correia, James Cussens, Cassio P. de Campos |
AISTATS | 1 |
| 2020 | Joints in Random ForestsabstractDecision Trees (DTs) and Random Forests (RFs) are powerful discriminative learners and tools of central importance to the everyday machine learning practitioner and data scientist. Due to their discriminative nature, however, they lack principled methods to process inputs with missing features or to detect outliers, which requires pairing them with imputation techniques or a separate generative model. In this paper, we demonstrate that DTs and RFs can naturally be interpreted as generative models, by drawing a connection to Probabilistic Circuits, a prominent class of tractable probabilistic models. This reinterpretation equips them with a full joint distribution over the feature space and leads to Generative Decision Trees (GeDTs) and Generative Forests (GeFs), a family of novel hybrid generative-discriminative models. This family of models retains the overall characteristics of DTs and RFs while additionally being able to handle missing features by means of marginalisation. Under certain assumptions, frequently made for Bayes consistency results, we show that consistency in GeDTs and GeFs extend to any pattern of missing input features, if missing at random. Empirically, we show that our models often outperform common routines to treat missing data, such as K-nearest neighbour imputation, and moreover, that our models can naturally detect outliers by monitoring the marginal probability of input features. Alvaro Henrique Chaim Correia, Robert Peharz, Cassio P. de Campos |
NeurIPS | 1 |
| 2019 | Human-in-the-Loop Feature SelectionabstractFeature selection is a crucial step in the conception of Machine Learning models, which is often performed via datadriven approaches that overlook the possibility of tapping into the human decision-making of the model’s designers and users. We present a human-in-the-loop framework that interacts with domain experts by collecting their feedback regarding the variables (of few samples) they evaluate as the most relevant for the task at hand. Such information can be modeled via Reinforcement Learning to derive a per-example feature selection method that tries to minimize the model’s loss function by focusing on the most pertinent variables from a human perspective. We report results on a proof-of-concept image classification dataset and on a real-world risk classification task in which the model successfully incorporated feedback from experts to improve its accuracy. Alvaro Henrique Chaim Correia, Freddy Lécué |
AAAI | 1 |
| 2018 | A Fully Attention-Based Information RetrieverabstractRecurrent neural networks are now the state-of-the-art in natural language processing because they can build rich contextual representations and process texts of arbitrary length. However, recent developments on attention mechanisms have equipped feedforward networks with similar capabilities, hence enabling faster computations due to the increase in the number of operations that can be parallelized. We explore this new type of architecture in the domain of question-answering and propose a novel approach that we call Fully Attention Based Information Retriever (FABIR). We show that FABIR achieves competitive results in the Stanford Question Answering Dataset (SQuAD) while having fewer parameters and being faster at both learning and inference than rival methods. Alvaro Henrique Chaim Correia, Jorge Luiz Moreira Silva, Thiago de Castro Martins, Fábio G. Cozman |
IJCNN | 1 |