VLDB 2026 Research / reviewers in the wild / expert
Mohamed Hebiri
dblp:78/8006
· DBLP profile ↗
12ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 since 2021Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Trustworthy machine learning · 74% Learning theory · 14% Learning paradigms · 4% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 50% Algorithms and data structures · 50% |
Topics — the 21 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.8 | 5 | 2024 | Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024 Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024 Fair regression via plug-in estimator and recalibration with statistical guarantees · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › fairness › fairness criteria
demographic parity |
2.0 | 3 | 2024 | Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024 Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024 Fair regression via plug-in estimator and recalibration with statistical guarantees · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › fairness › fair prediction
fair regression |
1.6 | 3 | 2024 | Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024 Fair regression via plug-in estimator and recalibration with statistical guarantees · NeurIPS 2020 Fair regression with Wasserstein barycenters · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
post-processing fairness |
0.8 | 1 | 2024 | Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024 |
Machine learning › Learning theory
classification |
0.5 | 2 | 2024 | Confidence Sets with Expected Sizes for Multiclass Classification · J. Mach. Learn. Res. 2017 Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024 |
Machine learning › Learning paradigms
semi-supervised learning |
0.4 | 1 | 2020 | Regression with reject option and application to kNN · NeurIPS 2020 |
Machine learning › Learning theory
statistical learning theory |
0.4 | 1 | 2020 | Regression with reject option and application to kNN · NeurIPS 2020 |
Algorithms and data structures › similarity search › nearest neighbor search
k-nearest neighbors |
0.4 | 1 | 2020 | Regression with reject option and application to kNN · NeurIPS 2020 |
Algorithms and data structures › similarity search
nearest neighbor search |
0.4 | 1 | 2020 | Regression with reject option and application to kNN · NeurIPS 2020 |
Mathematical optimization
optimal transport |
0.4 | 1 | 2020 | Fair regression with Wasserstein barycenters · NeurIPS 2020 |
Mathematical optimization › optimal transport
wasserstein barycenter |
0.4 | 1 | 2020 | Fair regression with Wasserstein barycenters · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › fairness › fairness criteria
equal opportunity |
0.4 | 1 | 2019 | Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › fairness › fair classification
fair binary classification |
0.4 | 1 | 2019 | Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification · NeurIPS 2019 |
Machine learning › Learning theory › classification
multiclass classification |
0.3 | 2 | 2024 | Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024 Confidence Sets with Expected Sizes for Multiclass Classification · J. Mach. Learn. Res. 2017 |
Machine learning › Reinforcement learning › exploration
confidence sets |
0.3 | 1 | 2017 | Confidence Sets with Expected Sizes for Multiclass Classification · J. Mach. Learn. Res. 2017 |
Machine learning › Optimization for machine learning
convex optimization |
0.2 | 1 | 2013 | Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
heteroscedastic modeling |
0.2 | 1 | 2013 | Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013 |
Machine learning › Learning theory
high-dimensional statistics |
0.2 | 1 | 2013 | How Correlations Influence Lasso Prediction · IEEE Trans. Inf. Theory 2013 |
Machine learning › Optimization for machine learning › regularized risk minimization
regularized regression |
0.2 | 1 | 2013 | How Correlations Influence Lasso Prediction · IEEE Trans. Inf. Theory 2013 |
Machine learning › Learning theory
high-dimensional regression |
0.0 | 1 | 2013 | Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013 |
Machine learning › Learning theory › high-dimensional statistics
sparse estimation |
0.0 | 1 | 2013 | Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013 |
Methods — techniques the papers use, named apart from their topics
plug-in estimator · 1.6wasserstein barycenter · 0.9post-processing · 0.9convergence rates · 0.9conditional variance thresholding · 0.9recalibration · 0.8stochastic convex optimization · 0.8discretization · 0.8smooth optimization · 0.4group-dependent threshold · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Class conditional conformal prediction for multiple inputs by p-value aggregationabstractConformal prediction methods are statistical tools designed to quantify uncertainty and generate predictive sets with guaranteed coverage probabilities. This work introduces an innovative refinement to these methods for classification tasks, specifically tailored for scenarios where multiple observations (multi-inputs) of a single instance are available at prediction time. Our approach is particularly motivated by applications in citizen science, where multiple images of the same plant or animal are captured by individuals.
Our method integrates the information from each observation into conformal prediction, enabling a reduction in the size of the predicted label set while preserving the required class-conditional coverage guarantee. The approach is based on the aggregation of conformal p-values computed from each observation of a multi-input.
By exploiting the exact distribution of these p-values, we propose a general aggregation framework using an abstract scoring function, encompassing many classical statistical tools. Knowledge of this distribution also enables refined versions of standard strategies, such as majority voting.
We evaluate our method on simulated and real data, with a particular focus on Pl@ntNet, a prominent citizen science platform that facilitates the collection and identification of plant species through user-submitted images. Jean-Baptiste Fermanian, Mohamed Hebiri, Joseph Salmon |
NeurIPS | 2 |
| 2025 | EERO: Early Exit with Reject Option for Efficient Classification with limited budgetabstractThe increasing complexity of advanced machine learning models requires innovative approaches to manage computational resources effectively. One such method is the Early Exit strategy, which allows for adaptive computation by providing a mechanism to shorten the processing path for simpler data instances. In this paper, we propose EERO, a new methodology to translate the problem of early exiting to a problem of using multiple classifiers with reject option in order to better select the exiting head for each instance. We calibrate the probabilities of exiting at the different heads using aggregation with exponential weights to guarantee a fixed budget. We consider factors such as Bayesian risk, budget constraints, and head-specific budget consumption. Experimental results demonstrate that our method achieves competitive compromise between budget allocation and accuracy. Florian Valade, Mohamed Hebiri, Paul Gay |
UAI | 2 |
| 2024 | Regression under demographic parity constraints via unlabeled post-processingabstractWe address the problem of performing regression while ensuring demographic parity, even without access to sensitive attributes during inference. We present a general-purpose post-processing algorithm that, using accurate estimates of the regression function and a sensitive attribute predictor, generates predictions that meet the demographic parity constraint. Our method involves discretization and stochastic minimization of a smooth convex function. It is suitable for online post-processing and multi-class classification tasks only involving unlabeled data for the post-processing. Unlike prior methods, our approach is fully theory-driven. We require precise control over the gradient norm of the convex function, and thus, we rely on more advanced techniques than standard stochastic gradient descent. Our algorithm is backed by finite-sample analysis and post-processing bounds, with experimental results validating our theoretical findings. Gayane Taturyan, Evgenii Chzhen, Mohamed Hebiri |
NeurIPS | 3 |
| 2024 | Fairness guarantees in multi-class classification with demographic parityabstractAlgorithmic Fairness is an established area of machine learning, willing to reduce the influence of hidden bias in the data. Yet, despite its wide range of applications, very few works consider the multi-class classification setting from the fairness perspective. We focus on this question and extend the definition of approximate fairness in the case of Demographic Parity to multi-class classification. We specify the corresponding expressions of the optimal fair classifiers in the attribute-aware case and both for binary and multi-categorical sensitive attributes. This suggests a plug-in data-driven procedure, for which we establish theoretical guarantees. The enhanced estimator is proved to mimic the behavior of the optimal rule both in terms of fairness and risk. Notably, fairness guarantees are distribution-free. The approach is evaluated on both synthetic and real datasets and reveals very effective in decision making with a preset level of unfairness. In addition, our method is competitive (if not better) with the state-of-the-art in binary and multi-class tasks. Christophe Denis, Romuald Elie, Mohamed Hebiri, François Hu |
J. Mach. Learn. Res. | 3 |
| 2024 | Active learning algorithm through the lens of rejection argumentsabstractAbstract Active learning is a paradigm of machine learning which aims at reducing the amount of labeled data needed to train a classifier. Its overall principle is to sequentially select the most informative data points, which amounts to determining the uncertainty of regions of the input space. The main challenge lies in building a procedure that is computationally efficient and that offers appealing theoretical properties; most of the current methods satisfy only one or the other. In this paper, we use the classification with rejection in a novel way to estimate the uncertain regions. We provide an active learning algorithm and prove its theoretical benefits under classical assumptions. In addition to the theoretical results, numerical experiments are carried out on synthetic and non-synthetic datasets. These experiments provide empirical evidence that the use of rejection arguments in our active learning algorithm is beneficial and allows good performance in various statistical situations. Christophe Denis, Mohamed Hebiri, Boris Ndjia Njike, Xavier Siebert |
Mach. Learn. | 2 |
| 2020 | Fair regression with Wasserstein barycentersabstractWe study the problem of learning a real-valued function that satisfies the Demographic Parity constraint. It demands the distribution of the predicted output to be independent of the sensitive attribute. We consider the case that the sensitive attribute is available for prediction. We establish a connection between fair regression and optimal transport theory, based on which we derive a close form expression for the optimal fair predictor. Specifically, we show that the distribution of this optimum is the Wasserstein barycenter of the distributions induced by the standard regression function on the sensitive groups. This result offers an intuitive interpretation of the optimal fair prediction and suggests a simple post-processing algorithm to achieve fairness. We establish risk and distribution-free fairness guarantees for this procedure. Numerical experiments indicate that our method is very effective in learning fair models, with a relative increase in error rate that is inferior to the relative gain in fairness. Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil |
NeurIPS | 3 |
| 2020 | Fair regression via plug-in estimator and recalibration with statistical guaranteesabstractWe study the problem of learning an optimal regression function subject to a fairness constraint. It requires that, conditionally on the sensitive feature, the distribution of the function output remains the same. This constraint naturally extends the notion of demographic parity, often used in classification, to the regression setting. We tackle this problem by leveraging on a proxy-discretized version, for which we derive an explicit expression of the optimal fair predictor. This result naturally suggests a two stage approach, in which we first estimate the (unconstrained) regression function from a set of labeled data and then we recalibrate it with another set of unlabeled data. The recalibration step can be efficiently performed via a smooth optimization. We derive rates of convergence of the proposed estimator to the optimal fair predictor both in terms of the risk and fairness constraint. Finally, we present numerical experiments illustrating that the proposed method is often superior or competitive with state-of-the-art methods. Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil |
NeurIPS | 3 |
| 2020 | Regression with reject option and application to kNNabstractWe investigate the problem of regression where one is allowed to abstain from predicting. We refer to this framework as regression with reject option as an extension of classification with reject option. In this context, we focus on the case where the rejection rate is fixed and derive the optimal rule which relies on thresholding the conditional variance function. We provide a semi-supervised estimation procedure of the optimal rule involving two datasets: a first labeled dataset is used to estimate both regression function and conditional variance function while a second unlabeled dataset is exploited to calibrate the desired rejection rate. The resulting predictor with reject option is shown to be almost as good as the optimal predictor with reject option both in terms of risk and rejection rate. We additionally apply our methodology with kNN algorithm and establish rates of convergence for the resulting kNN predictor under mild conditions. Finally, a numerical study is performed to illustrate the benefit of using the proposed procedure. Ahmed Zaoui, Christophe Denis, Mohamed Hebiri |
NeurIPS | 3 |
| 2019 | Leveraging Labeled and Unlabeled Data for Consistent Fair Binary ClassificationabstractWe study the problem of fair binary classification using the notion of Equal Opportunity. It requires the true positive rate to distribute equally across the sensitive groups. Within this setting we show that the fair optimal classifier is obtained by recalibrating the Bayes classifier by a group-dependent threshold. We provide a constructive expression for the threshold. This result motivates us to devise a plug-in classification procedure based on both unlabeled and labeled datasets. While the latter is used to learn the output conditional probability, the former is used for calibration. The overall procedure can be computed in polynomial time and it is shown to be statistically consistent both in terms of the classification error and fairness measure. Finally, we present numerical experiments which indicate that our method is often superior or competitive with the state-of-the-art methods on benchmark datasets. Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil |
NeurIPS | 3 |
| 2017 | Confidence Sets with Expected Sizes for Multiclass ClassificationabstractMulticlass classification problems such as image annotation can involve a large number of classes. In this context, confusion between classes can occur, and single label classification may be misleading. We provide in the present paper a general device that, given an unlabeled dataset and a score function defined as the minimizer of some empirical and convex risk, outputs a set of class labels, instead of a single one. Interestingly, this procedure does not require that the unlabeled dataset explores the whole classes. Even more, the method is calibrated to control the expected size of the output set while minimizing the classification risk. We show the statistical optimality of the procedure and establish rates of convergence under the Tsybakov margin condition. It turns out that these rates are linear on the number of labels. We apply our methodology to convex aggregation of confidence sets based on the $V$-fold cross validation principle also known as the superlearning principle (van der Laan et al., 2007). We illustrate the numerical performance of the procedure on real data and demonstrate in particular that with moderate expected size, w.r.t. the number of labels, the procedure provides significant improvement of the classification risk. Christophe Denis, Mohamed Hebiri |
J. Mach. Learn. Res. | 2 |
| 2013 | Learning Heteroscedastic Models by Convex Programming under Group SparsityabstractSparse estimation methods based on l1 relaxation, such as the Lasso and the Dantzig selector, require the knowledge of the variance of the noise in order to properly tune the regularization parameter. This constitutes a major obstacle in applying these methods in several frameworks, such as time series, random fields, inverse problems, for which noise is rarely homoscedastic and the noise level is hard to know in advance. In this paper, we propose a new approach to the joint estimation of the conditional mean and the conditional variance in a high-dimensional (auto-) regression setting. An attractive feature of the proposed estimator is that it is efficiently computable even for very large scale problems by solving a second-order cone program (SOCP). We present theoretical analysis and numerical results assessing the performance of the proposed procedure. Arnak S. Dalalyan, Mohamed Hebiri, Katia Meziani, Joseph Salmon |
ICML (3) | 2 |
| 2013 | How Correlations Influence Lasso PredictionabstractWe study how correlations in the design matrix influence Lasso prediction. First, we argue that the higher the correlations, the smaller the optimal tuning parameter. This implies in particular that the standard tuning parameters, that do not depend on the design matrix, are not favorable. Furthermore, we argue that Lasso prediction works well for any degree of correlations if suitable tuning parameters are chosen. We study these two subjects theoretically as well as with simulations. Mohamed Hebiri, Johannes Lederer |
IEEE Trans. Inf. Theory | 1 |