Mohamed Hebiri

dblp:78/8006 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 since 2021Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Trustworthy machine learning · 74% Learning theory · 14% Learning paradigms · 4%
Theoretical computer science
2 papers
Mathematical optimization · 50% Algorithms and data structures · 50%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
2.852024
Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024
Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024
Fair regression via plug-in estimator and recalibration with statistical guarantees · NeurIPS 2020
Machine learning › Trustworthy machine learning › fairness › fairness criteria
demographic parity
2.032024
Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024
Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024
Fair regression via plug-in estimator and recalibration with statistical guarantees · NeurIPS 2020
Machine learning › Trustworthy machine learning › fairness › fair prediction
fair regression
1.632024
Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024
Fair regression via plug-in estimator and recalibration with statistical guarantees · NeurIPS 2020
Fair regression with Wasserstein barycenters · NeurIPS 2020
Machine learning › Trustworthy machine learning › fairness › bias mitigation
post-processing fairness
0.812024
Regression under demographic parity constraints via unlabeled post-processing · NeurIPS 2024
Machine learning › Learning theory
classification
0.522024
Confidence Sets with Expected Sizes for Multiclass Classification · J. Mach. Learn. Res. 2017
Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024
Machine learning › Learning paradigms
semi-supervised learning
0.412020
Regression with reject option and application to kNN · NeurIPS 2020
Machine learning › Learning theory
statistical learning theory
0.412020
Regression with reject option and application to kNN · NeurIPS 2020
Algorithms and data structures › similarity search › nearest neighbor search
k-nearest neighbors
0.412020
Regression with reject option and application to kNN · NeurIPS 2020
Algorithms and data structures › similarity search
nearest neighbor search
0.412020
Regression with reject option and application to kNN · NeurIPS 2020
Mathematical optimization
optimal transport
0.412020
Fair regression with Wasserstein barycenters · NeurIPS 2020
Mathematical optimization › optimal transport
wasserstein barycenter
0.412020
Fair regression with Wasserstein barycenters · NeurIPS 2020
Machine learning › Trustworthy machine learning › fairness › fairness criteria
equal opportunity
0.412019
Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification · NeurIPS 2019
Machine learning › Trustworthy machine learning › fairness › fair classification
fair binary classification
0.412019
Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification · NeurIPS 2019
Machine learning › Learning theory › classification
multiclass classification
0.322024
Fairness guarantees in multi-class classification with demographic parity · J. Mach. Learn. Res. 2024
Confidence Sets with Expected Sizes for Multiclass Classification · J. Mach. Learn. Res. 2017
Machine learning › Reinforcement learning › exploration
confidence sets
0.312017
Confidence Sets with Expected Sizes for Multiclass Classification · J. Mach. Learn. Res. 2017
Machine learning › Optimization for machine learning
convex optimization
0.212013
Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
heteroscedastic modeling
0.212013
Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013
Machine learning › Learning theory
high-dimensional statistics
0.212013
How Correlations Influence Lasso Prediction · IEEE Trans. Inf. Theory 2013
Machine learning › Optimization for machine learning › regularized risk minimization
regularized regression
0.212013
How Correlations Influence Lasso Prediction · IEEE Trans. Inf. Theory 2013
Machine learning › Learning theory
high-dimensional regression
0.012013
Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013
Machine learning › Learning theory › high-dimensional statistics
sparse estimation
0.012013
Learning Heteroscedastic Models by Convex Programming under Group Sparsity · ICML (3) 2013

Methods — techniques the papers use, named apart from their topics

plug-in estimator · 1.6wasserstein barycenter · 0.9post-processing · 0.9convergence rates · 0.9conditional variance thresholding · 0.9recalibration · 0.8stochastic convex optimization · 0.8discretization · 0.8smooth optimization · 0.4group-dependent threshold · 0.4
YearPublicationVenuePosition
2025 Class conditional conformal prediction for multiple inputs by p-value aggregation
abstract
Conformal prediction methods are statistical tools designed to quantify uncertainty and generate predictive sets with guaranteed coverage probabilities. This work introduces an innovative refinement to these methods for classification tasks, specifically tailored for scenarios where multiple observations (multi-inputs) of a single instance are available at prediction time. Our approach is particularly motivated by applications in citizen science, where multiple images of the same plant or animal are captured by individuals. Our method integrates the information from each observation into conformal prediction, enabling a reduction in the size of the predicted label set while preserving the required class-conditional coverage guarantee. The approach is based on the aggregation of conformal p-values computed from each observation of a multi-input. By exploiting the exact distribution of these p-values, we propose a general aggregation framework using an abstract scoring function, encompassing many classical statistical tools. Knowledge of this distribution also enables refined versions of standard strategies, such as majority voting. We evaluate our method on simulated and real data, with a particular focus on Pl@ntNet, a prominent citizen science platform that facilitates the collection and identification of plant species through user-submitted images.
Jean-Baptiste Fermanian, Mohamed Hebiri, Joseph Salmon
NeurIPS2
2025 EERO: Early Exit with Reject Option for Efficient Classification with limited budget
abstract
The increasing complexity of advanced machine learning models requires innovative approaches to manage computational resources effectively. One such method is the Early Exit strategy, which allows for adaptive computation by providing a mechanism to shorten the processing path for simpler data instances. In this paper, we propose EERO, a new methodology to translate the problem of early exiting to a problem of using multiple classifiers with reject option in order to better select the exiting head for each instance. We calibrate the probabilities of exiting at the different heads using aggregation with exponential weights to guarantee a fixed budget. We consider factors such as Bayesian risk, budget constraints, and head-specific budget consumption. Experimental results demonstrate that our method achieves competitive compromise between budget allocation and accuracy.
Florian Valade, Mohamed Hebiri, Paul Gay
UAI2
2024 Regression under demographic parity constraints via unlabeled post-processing
abstract
We address the problem of performing regression while ensuring demographic parity, even without access to sensitive attributes during inference. We present a general-purpose post-processing algorithm that, using accurate estimates of the regression function and a sensitive attribute predictor, generates predictions that meet the demographic parity constraint. Our method involves discretization and stochastic minimization of a smooth convex function. It is suitable for online post-processing and multi-class classification tasks only involving unlabeled data for the post-processing. Unlike prior methods, our approach is fully theory-driven. We require precise control over the gradient norm of the convex function, and thus, we rely on more advanced techniques than standard stochastic gradient descent. Our algorithm is backed by finite-sample analysis and post-processing bounds, with experimental results validating our theoretical findings.
Gayane Taturyan, Evgenii Chzhen, Mohamed Hebiri
NeurIPS3
2024 Fairness guarantees in multi-class classification with demographic parity
abstract
Algorithmic Fairness is an established area of machine learning, willing to reduce the influence of hidden bias in the data. Yet, despite its wide range of applications, very few works consider the multi-class classification setting from the fairness perspective. We focus on this question and extend the definition of approximate fairness in the case of Demographic Parity to multi-class classification. We specify the corresponding expressions of the optimal fair classifiers in the attribute-aware case and both for binary and multi-categorical sensitive attributes. This suggests a plug-in data-driven procedure, for which we establish theoretical guarantees. The enhanced estimator is proved to mimic the behavior of the optimal rule both in terms of fairness and risk. Notably, fairness guarantees are distribution-free. The approach is evaluated on both synthetic and real datasets and reveals very effective in decision making with a preset level of unfairness. In addition, our method is competitive (if not better) with the state-of-the-art in binary and multi-class tasks.
Christophe Denis, Romuald Elie, Mohamed Hebiri, François Hu
J. Mach. Learn. Res.3
2024 Active learning algorithm through the lens of rejection arguments
abstract
Abstract Active learning is a paradigm of machine learning which aims at reducing the amount of labeled data needed to train a classifier. Its overall principle is to sequentially select the most informative data points, which amounts to determining the uncertainty of regions of the input space. The main challenge lies in building a procedure that is computationally efficient and that offers appealing theoretical properties; most of the current methods satisfy only one or the other. In this paper, we use the classification with rejection in a novel way to estimate the uncertain regions. We provide an active learning algorithm and prove its theoretical benefits under classical assumptions. In addition to the theoretical results, numerical experiments are carried out on synthetic and non-synthetic datasets. These experiments provide empirical evidence that the use of rejection arguments in our active learning algorithm is beneficial and allows good performance in various statistical situations.
Christophe Denis, Mohamed Hebiri, Boris Ndjia Njike, Xavier Siebert
Mach. Learn.2
2020 Fair regression with Wasserstein barycenters
abstract
We study the problem of learning a real-valued function that satisfies the Demographic Parity constraint. It demands the distribution of the predicted output to be independent of the sensitive attribute. We consider the case that the sensitive attribute is available for prediction. We establish a connection between fair regression and optimal transport theory, based on which we derive a close form expression for the optimal fair predictor. Specifically, we show that the distribution of this optimum is the Wasserstein barycenter of the distributions induced by the standard regression function on the sensitive groups. This result offers an intuitive interpretation of the optimal fair prediction and suggests a simple post-processing algorithm to achieve fairness. We establish risk and distribution-free fairness guarantees for this procedure. Numerical experiments indicate that our method is very effective in learning fair models, with a relative increase in error rate that is inferior to the relative gain in fairness.
Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil
NeurIPS3
2020 Fair regression via plug-in estimator and recalibration with statistical guarantees
abstract
We study the problem of learning an optimal regression function subject to a fairness constraint. It requires that, conditionally on the sensitive feature, the distribution of the function output remains the same. This constraint naturally extends the notion of demographic parity, often used in classification, to the regression setting. We tackle this problem by leveraging on a proxy-discretized version, for which we derive an explicit expression of the optimal fair predictor. This result naturally suggests a two stage approach, in which we first estimate the (unconstrained) regression function from a set of labeled data and then we recalibrate it with another set of unlabeled data. The recalibration step can be efficiently performed via a smooth optimization. We derive rates of convergence of the proposed estimator to the optimal fair predictor both in terms of the risk and fairness constraint. Finally, we present numerical experiments illustrating that the proposed method is often superior or competitive with state-of-the-art methods.
Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil
NeurIPS3
2020 Regression with reject option and application to kNN
abstract
We investigate the problem of regression where one is allowed to abstain from predicting. We refer to this framework as regression with reject option as an extension of classification with reject option. In this context, we focus on the case where the rejection rate is fixed and derive the optimal rule which relies on thresholding the conditional variance function. We provide a semi-supervised estimation procedure of the optimal rule involving two datasets: a first labeled dataset is used to estimate both regression function and conditional variance function while a second unlabeled dataset is exploited to calibrate the desired rejection rate. The resulting predictor with reject option is shown to be almost as good as the optimal predictor with reject option both in terms of risk and rejection rate. We additionally apply our methodology with kNN algorithm and establish rates of convergence for the resulting kNN predictor under mild conditions. Finally, a numerical study is performed to illustrate the benefit of using the proposed procedure.
Ahmed Zaoui, Christophe Denis, Mohamed Hebiri
NeurIPS3
2019 Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification
abstract
We study the problem of fair binary classification using the notion of Equal Opportunity. It requires the true positive rate to distribute equally across the sensitive groups. Within this setting we show that the fair optimal classifier is obtained by recalibrating the Bayes classifier by a group-dependent threshold. We provide a constructive expression for the threshold. This result motivates us to devise a plug-in classification procedure based on both unlabeled and labeled datasets. While the latter is used to learn the output conditional probability, the former is used for calibration. The overall procedure can be computed in polynomial time and it is shown to be statistically consistent both in terms of the classification error and fairness measure. Finally, we present numerical experiments which indicate that our method is often superior or competitive with the state-of-the-art methods on benchmark datasets.
Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto, Massimiliano Pontil
NeurIPS3
2017 Confidence Sets with Expected Sizes for Multiclass Classification
abstract
Multiclass classification problems such as image annotation can involve a large number of classes. In this context, confusion between classes can occur, and single label classification may be misleading. We provide in the present paper a general device that, given an unlabeled dataset and a score function defined as the minimizer of some empirical and convex risk, outputs a set of class labels, instead of a single one. Interestingly, this procedure does not require that the unlabeled dataset explores the whole classes. Even more, the method is calibrated to control the expected size of the output set while minimizing the classification risk. We show the statistical optimality of the procedure and establish rates of convergence under the Tsybakov margin condition. It turns out that these rates are linear on the number of labels. We apply our methodology to convex aggregation of confidence sets based on the $V$-fold cross validation principle also known as the superlearning principle (van der Laan et al., 2007). We illustrate the numerical performance of the procedure on real data and demonstrate in particular that with moderate expected size, w.r.t. the number of labels, the procedure provides significant improvement of the classification risk.
Christophe Denis, Mohamed Hebiri
J. Mach. Learn. Res.2
2013 Learning Heteroscedastic Models by Convex Programming under Group Sparsity
abstract
Sparse estimation methods based on l1 relaxation, such as the Lasso and the Dantzig selector, require the knowledge of the variance of the noise in order to properly tune the regularization parameter. This constitutes a major obstacle in applying these methods in several frameworks, such as time series, random fields, inverse problems, for which noise is rarely homoscedastic and the noise level is hard to know in advance. In this paper, we propose a new approach to the joint estimation of the conditional mean and the conditional variance in a high-dimensional (auto-) regression setting. An attractive feature of the proposed estimator is that it is efficiently computable even for very large scale problems by solving a second-order cone program (SOCP). We present theoretical analysis and numerical results assessing the performance of the proposed procedure.
Arnak S. Dalalyan, Mohamed Hebiri, Katia Meziani, Joseph Salmon
ICML (3)2
2013 How Correlations Influence Lasso Prediction
abstract
We study how correlations in the design matrix influence Lasso prediction. First, we argue that the higher the correlations, the smaller the optimal tuning parameter. This implies in particular that the standard tuning parameters, that do not depend on the design matrix, are not favorable. Furthermore, we argue that Lasso prediction works well for any degree of correlations if suitable tuning parameters are chosen. We study these two subjects theoretically as well as with simulations.
Mohamed Hebiri, Johannes Lederer
IEEE Trans. Inf. Theory1