VLDB 2026 Research / reviewers in the wild / expert
Jean-Michel Loubes
dblp:44/7859
· DBLP profile ↗
20ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0002-1252-2960ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Theory of computation · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 49% Probabilistic and Bayesian machine learning · 34% Optimization for machine learning · 9% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 100% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.1 | 2 | 2024 | Transport-based Counterfactual Models · J. Mach. Learn. Res. 2024 Obtaining Fairness using Optimal Transport Theory · ICML 2019 |
Privacy and data protection
differential privacy |
0.9 | 1 | 2025 | On the Private Estimation of Smooth Transport Maps · ICML 2025 |
Mathematical optimization
optimal transport |
0.9 | 1 | 2025 | On the Private Estimation of Smooth Transport Maps · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.8 | 1 | 2024 | Transport-based Counterfactual Models · J. Mach. Learn. Res. 2024 |
Machine learning › Trustworthy machine learning › fairness › causal fairness
counterfactual fairness |
0.8 | 1 | 2024 | Transport-based Counterfactual Models · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
bayesian nonparametric regression |
0.6 | 1 | 2022 | Nonparametric Bayesian Regression and Classification on Manifolds, With Applications to 3D Cochlear Shapes · IEEE Trans. Image Process. 2022 |
Machine learning › Optimization for machine learning
optimal transport |
0.5 | 2 | 2021 | Obtaining Fairness using Optimal Transport Theory · ICML 2019 Achieving Robustness in Classification Using Optimal Transport With Hinge Regularization · CVPR 2021 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.5 | 1 | 2021 | Achieving Robustness in Classification Using Optimal Transport With Hinge Regularization · CVPR 2021 |
Machine learning › Trustworthy machine learning › robustness
certified robustness |
0.5 | 1 | 2021 | Achieving Robustness in Classification Using Optimal Transport With Hinge Regularization · CVPR 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Achieving Robustness in Classification Using Optimal Transport With Hinge Regularization · CVPR 2021 |
Machine learning › Trustworthy machine learning › fairness
fair classification |
0.4 | 1 | 2019 | Obtaining Fairness using Optimal Transport Theory · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.3 | 1 | 2018 | A Gaussian Process Regression Model for Distribution Inputs · IEEE Trans. Inf. Theory 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
kernel design |
0.3 | 1 | 2018 | A Gaussian Process Regression Model for Distribution Inputs · IEEE Trans. Inf. Theory 2018 |
Mathematical optimization › regularization
group lasso |
0.3 | 2 | 2014 | Oracle Inequalities for a Group Lasso Procedure Applied to Generalized Linear Models in High Dimension · IEEE Trans. Inf. Theory 2014 Group Lasso Estimation of High-dimensional Covariance Matrices · J. Mach. Learn. Res. 2011 |
Machine learning › Learning theory › excess risk bounds
oracle inequality |
0.2 | 1 | 2014 | Oracle Inequalities for a Group Lasso Procedure Applied to Generalized Linear Models in High Dimension · IEEE Trans. Inf. Theory 2014 |
Machine learning › Learning theory
statistical learning theory |
0.2 | 1 | 2014 | Oracle Inequalities for a Group Lasso Procedure Applied to Generalized Linear Models in High Dimension · IEEE Trans. Inf. Theory 2014 |
Mathematical optimization › statistical estimation › high-dimensional estimation
sparse estimation |
0.2 | 1 | 2014 | Oracle Inequalities for a Group Lasso Procedure Applied to Generalized Linear Models in High Dimension · IEEE Trans. Inf. Theory 2014 |
Machine learning › Optimization for machine learning › optimal transport
wasserstein distance |
0.1 | 1 | 2021 | Achieving Robustness in Classification Using Optimal Transport With Hinge Regularization · CVPR 2021 |
Machine learning › Learning theory
high-dimensional statistics |
0.1 | 1 | 2011 | Group Lasso Estimation of High-dimensional Covariance Matrices · J. Mach. Learn. Res. 2011 |
Machine learning › Learning theory › high-dimensional statistics
sparse covariance estimation |
0.1 | 1 | 2011 | Group Lasso Estimation of High-dimensional Covariance Matrices · J. Mach. Learn. Res. 2011 |
Mathematical optimization › statistical estimation
covariance estimation |
0.1 | 1 | 2011 | Group Lasso Estimation of High-dimensional Covariance Matrices · J. Mach. Learn. Res. 2011 |
Mathematical optimization › regularization
regularized estimation |
0.1 | 1 | 2011 | Group Lasso Estimation of High-dimensional Covariance Matrices · J. Mach. Learn. Res. 2011 |
Mathematical optimization › statistical estimation › regression
generalized linear models |
0.1 | 1 | 2014 | Oracle Inequalities for a Group Lasso Procedure Applied to Generalized Linear Models in High Dimension · IEEE Trans. Inf. Theory 2014 |
Methods — techniques the papers use, named apart from their topics
brenier potential · 1.7optimal transport · 1.2spherical gaussian process decomposition · 1.1bayesian inference · 1.1minimax lower bounds · 0.9minimax lower bound · 0.9optimal transport theory · 0.8group lasso · 0.6lipschitz constraint · 0.5hinge regularization · 0.5wasserstein barycenter · 0.4positive definite kernels · 0.3oracle inequalities · 0.2elastic-net penalty · 0.2elastic net penalty · 0.2convex regularization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fairness-Aware Grouping for Continuous Sensitive Variables: Application for Debiasing Face Analysis with Respect to Skin ToneabstractWithin a legal framework, fairness in datasets and models is typically assessed by dividing observations into predefined groups and then computing fairness measures (e.g., Disparate Impact or Equality of Odds with respect to gender). However, when sensitive attributes such as skin color are continuous, dividing into default groups may overlook or obscure the discrimination experienced by certain minority subpopulations. To address this limitation, we propose a fairness-based grouping approach for continuous (possibly multidimensional) sensitive attributes. By grouping data according to observed levels of discrimination, our method identifies the partition that maximizes a novel criterion based on inter-group variance in discrimination, thereby isolating the most critical subgroups. We validate the proposed approach using multiple synthetic datasets and demonstrate its robustness under changing population distributions—revealing how discrimination is manifested within the space of sensitive attributes. Furthermore, we examine a specialized setting of monotonic fairness for the case of skin color. Our empirical results on both CelebA and FFHQ, leveraging the skin tone as predicted by an industrial proprietary algorithm, show that the proposed segmentation uncovers more nuanced patterns of discrimination than previously reported, and that these findings remain stable across datasets for a given model. Finally, we leverage our grouping model for debiasing purpose, aiming at predicting fair scores with group-by-group post-processing. The results demonstrate that our approach improves fairness while having minimal impact on accuracy, thus confirming our partition method and opening the door for industrial deployment. Veronika Shilova, Emmanuel Malherbe, Giovanni Palma, Laurent Risser, Jean-Michel Loubes |
ECAI | 5 |
| 2025 | On the Private Estimation of Smooth Transport MapsabstractEstimating optimal transport maps between two distributions from respective samples is an important element for many machine learning methods. To do so, rather than extending discrete transport maps, it has been shown that estimating the Brenier potential of the transport problem and obtaining a transport map through its gradient is near minimax optimal for smooth problems. In this paper, we investigate the private estimation of such potentials and transport maps with respect to the distribution samples. We propose a differentially private transport map estimator with $L^2$ error at most $n^{-1} \vee n^{-\frac{2 \alpha}{2 \alpha - 2 + d}} \vee (n\epsilon)^{-\frac{2 \alpha}{2 \alpha + d}} $ up do polylog terms where $n$ is the sample size, $\epsilon$ is the desired level of privacy, $\alpha$ is the smoothness of the true transport map, and $d$ is the dimension of the feature space. We also provide a lower bound for the problem. Clément Lalanne, Franck Iutzeler, Jean-Michel Loubes, Julien Chhor |
ICML | 3 |
| 2025 | When majority rules, minority loses: bias amplification of gradient descentabstractDespite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that neglect minority-specific features. Assuming population and variance imbalance, our analysis reveals three key findings: (i) the close proximity between "full-data" and stereotypical predictors, (ii) the dominance of a region where training the entire model tends to merely learn the majority traits, and (iii) a lower bound on the additional training required. Our results are illustrated through experiments in deep learning for tabular and image classification tasks. François Bachoc, Jérôme Bolte, Ryan Boustany, Jean-Michel Loubes |
NeurIPS | 4 |
| 2024 | Transport-based Counterfactual ModelsabstractCounterfactual frameworks have grown popular in machine learning for both explaining algorithmic decisions but also defining individual notions of fairness, more intuitive than typical group fairness conditions. However, state-of-the-art models to compute counterfactuals are either unrealistic or unfeasible. In particular, while Pearl's causal inference provides appealing rules to calculate counterfactuals, it relies on a model that is unknown and hard to discover in practice. We address the problem of designing realistic and feasible counterfactuals in the absence of a causal model. We define transport-based counterfactual models as collections of joint probability distributions between observable distributions, and show their connection to causal counterfactuals. More specifically, we argue that optimal-transport theory defines relevant transport-based counterfactual models, as they are numerically feasible, statistically-faithful, and can coincide under some assumptions with causal counterfactual models. Finally, these models make counterfactual approaches to fairness feasible, and we illustrate their practicality and efficiency on fair learning. With this paper, we aim at laying out the theoretical foundations for a new, implementable approach to counterfactual thinking. Lucas de Lara, Alberto González-Sanz, Nicholas Asher, Laurent Risser, Jean-Michel Loubes |
J. Mach. Learn. Res. | 5 |
| 2024 | Fairness seen as global sensitivity analysis
Clément Benesse, Fabrice Gamboa, Jean-Michel Loubes, Thibaut Boissin |
Mach. Learn. | 3 |
| 2023 | Gaussian Processes on Distributions based on Regularized Optimal TransportabstractWe present a novel kernel over the space of probability measures based on the dual formulation of optimal regularized transport. We propose an Hilbertian embedding of the space of probabilities using their Sinkhorn potentials, which are solutions of the dual entropic relaxed optimal transport between the probabilities and a reference measure $\mathcal{U}$. We prove that this construction enables to obtain a valid kernel, by using the Hilbert norms. We prove that the kernel enjoys theoretical properties such as universality and some invariances, while still being computationally feasible. Moreover we provide theoretical guarantees on the behaviour of a Gaussian process based on this kernel. The empirical performances are compared with other traditional choices of kernels for processes indexed on distributions. François Bachoc, Louis Béthune, Alberto González-Sanz, Jean-Michel Loubes |
AISTATS | 4 |
| 2023 | Counterfactual Explanation for Multivariate Times Series Using A Contrastive Variational AutoencoderabstractWe tackle the issue of anomaly detection for multivariate functional data in a supervised setting. Deep learning applied to multivariate time series has become common nowadays, especially for medical data such as electrocardiogram (ECG). There are not many explanability techniques that can handle multivariate time series. In this paper, we propose a model to understand abnormal class features on multivariate time series. We present a new variational autoencoder (VAE) training method that focuses on dividing the latent space into general and class-based features using supervised contrastive learning. The Contrastive VAE produces a well-organized latent space that enables us to modify only the class-based features and to use the generative part of the VAE to produce counterfactual examples. This method is able to easily provide plausible counterfactual observations, which highlights the differences between pathological and non-pathological data. We demonstrate the superiority of our approach over other counterfactual methods in terms of validity and performance. William Todo, Merwann Selmani, Béatrice Laurent, Jean-Michel Loubes |
ICASSP | 4 |
| 2023 | Diffeomorphic Registration Using Sinkhorn DivergencesabstractAbstract. The diffeomorphic registration framework enables one to define an optimal matching function between two probability measures with respect to a data-fidelity loss function. The nonconvexity of the optimization problem renders the choice of this loss function crucial to avoid poor local minima. Recent work showed experimentally the efficiency of entropy-regularized optimal transportation costs, as they are computationally fast and differentiable while having few minima. Following this approach, we provide in this paper a new framework based on Sinkhorn divergences, unbiased entropic optimal transportation costs, and prove the statistical consistency with rate of the empirical optimal deformations. Lucas de Lara, Alberto González-Sanz, Jean-Michel Loubes |
SIAM J. Imaging Sci. | 3 |
| 2022 | Nonparametric Bayesian Regression and Classification on Manifolds, With Applications to 3D Cochlear ShapesabstractAdvanced shape analysis studies such as regression and classification need to be performed on curved manifolds, where often, there is a lack of standard statistical formulations. To overcome these limitations, we introduce a novel machine-learning method on the shape space of curves that avoids direct inference on infinite-dimensional spaces and instead performs Bayesian inference with spherical Gaussian processes decomposition. As an application, we study the shape of the cochlear spiral-shaped cavity within the petrous part of the temporal bone. This problem is particularly challenging due to the relationship between shape and gender, especially in children. Experimental results for both synthetic and real data show improved performance compared to state-of-the-art methods. Anis Fradi, Chafik Samir, José Braga, Shantanu H. Joshi, Jean-Michel Loubes |
IEEE Trans. Image Process. | 5 |
| 2021 | Achieving Robustness in Classification Using Optimal Transport With Hinge RegularizationabstractAdversarial examples have pointed out Deep Neural Network’s vulnerability to small local noise. It has been shown that constraining their Lipschitz constant should enhance robustness, but make them harder to learn with classical loss functions. We propose a new framework for binary classification, based on optimal transport, which integrates this Lipschitz constraint as a theoretical requirement. We propose to learn 1-Lipschitz networks using a new loss that is an hinge regularized version of the Kantorovich-Rubinstein dual formulation for the Wasserstein distance estimation. This loss function has a direct interpretation in terms of adversarial robustness together with certifiable robustness bound. We also prove that this hinge regularized version is still the dual formulation of an optimal transportation problem, and has a solution. We also establish several geometrical properties of this optimal solution, and extend the approach to multi-class problems. Experiments show that the proposed approach provides the expected guarantees in terms of robustness without any significant accuracy drop. The adversarial examples, on the proposed models, visibly and meaningfully change the input providing an explanation for the classification. Mathieu Serrurier, Franck Mamalet, Alberto González-Sanz, Thibaut Boissin, Jean-Michel Loubes, Eustasio del Barrio |
CVPR | 5 |
| 2021 | Bayesian regression and classification using Gaussian process priors indexed by probability density functions
Anis Fradi, Yan Feunteun, Chafik Samir, M. Baklouti, François Bachoc, Jean-Michel Loubes |
Inf. Sci. | 6 |
| 2020 | optimalFlow: optimal transport approach to flow cytometry gating and population matchingabstractBACKGROUND: Data obtained from flow cytometry present pronounced variability due to biological and technical reasons. Biological variability is a well-known phenomenon produced by measurements on different individuals, with different characteristics such as illness, age, sex, etc. The use of different settings for measurement, the variation of the conditions during experiments and the different types of flow cytometers are some of the technical causes of variability. This mixture of sources of variability makes the use of supervised machine learning for identification of cell populations difficult. The present work is conceived as a combination of strategies to facilitate the task of supervised gating. RESULTS: We propose optimalFlowTemplates, based on a similarity distance and Wasserstein barycenters, which clusters cytometries and produces prototype cytometries for the different groups. We show that supervised learning, restricted to the new groups, performs better than the same techniques applied to the whole collection. We also present optimalFlowClassification, which uses a database of gated cytometries and optimalFlowTemplates to assign cell types to a new cytometry. We show that this procedure can outperform state of the art techniques in the proposed datasets. Our code is freely available as optimalFlow, a Bioconductor R package at https://bioconductor.org/packages/optimalFlow . CONCLUSIONS: optimalFlowTemplates + optimalFlowClassification addresses the problem of using supervised learning while accounting for biological and technical variability. Our methodology provides a robust automated gating workflow that handles the intrinsic variability of flow cytometry data well. Our main innovation is the methodology itself and the optimal transport techniques that we apply to flow cytometry analysis. Eustasio del Barrio, Hristo Inouzhe, Jean-Michel Loubes, Carlos Matrán, Agustín Mayo-Íscar |
BMC Bioinform. | 3 |
| 2020 | Multiple Testing for Outlier Detection in Space TelemetriesabstractWe propose a novel procedure for outlier detection in space telemetries, in a semi-supervised framework. As the data is functional, we reduce its dimension by considering the coefficients obtained after projecting the observations onto orthonormal bases. A multiple testing procedure based on the two-sample test is defined in order to highlight the levels of the coefficients on which the outliers appear as significantly different from the nominal data. The Local Outlier Factor is computed on the selected coefficients to highlight the outliers. This procedure for selecting the features is applied on simulated data that mimic the behavior of space telemetries and on a real telemetry and then compared with existing dimension reduction techniques. Clémentine Barreyre, Béatrice Laurent, Jean-Michel Loubes, Loïc Boussouf, Bertrand Cabon |
IEEE Trans. Big Data | 3 |
| 2019 | Obtaining Fairness using Optimal Transport TheoryabstractIn the fair classification setup, we recast the links between fairness and predictability in terms of probability metrics. We analyze repair methods based on mapping conditional distributions to the Wasserstein barycenter. We propose a Random Repair which yields a tradeoff between minimal information loss and a certain amount of fairness. Paula Gordaliza, Eustasio del Barrio, Fabrice Gamboa, Jean-Michel Loubes |
ICML | 4 |
| 2019 | Learning a Gaussian Process Model on the Riemannian Manifold of Non-decreasing Distribution Functions
Chafik Samir, Jean-Michel Loubes, Anne-Françoise Yao, François Bachoc |
PRICAI (2) | 2 |
| 2018 | A Gaussian Process Regression Model for Distribution InputsabstractMonge-Kantorovich distances, otherwise known as Wasserstein distances, have received a growing attention in statistics and machine learning as a powerful discrepancy measure for probability distributions. In this paper, we focus on forecasting a Gaussian process indexed by probability distributions. For this, we provide a family of positive definite kernels built using transportation based distances. We provide a probabilistic understanding of these kernels and characterize the corresponding stochastic processes. We prove that the Gaussian processes indexed by distributions corresponding to these kernels can be efficiently forecast, opening new perspectives in Gaussian process modeling. François Bachoc, Fabrice Gamboa, Jean-Michel Loubes, Nil Venet |
IEEE Trans. Inf. Theory | 3 |
| 2018 | Destination Prediction by Trajectory Distribution-Based ModelabstractIn this paper, we propose a new method to predict the final destination of vehicle trips based on their initial partial trajectories. We first review how we obtained clustering of trajectories that describes user behavior. Then, we explain how we model main traffic flow patterns by a mixture of 2-D Gaussian distributions. This yielded a density-based clustering of locations, which produces a data driven grid of similar points within each pattern. We present how this model can be used to predict the final destination of a new trajectory based on their first locations using a two-step procedure: we first assign the new trajectory to the clusters it most likely belongs. Second, we use characteristics from trajectories inside these clusters to predict the final destination. Finally, we present experimental results of our methods for classification of trajectories and final destination prediction on data sets of timestamped GPS-Location of taxi trips. We test our methods on two different data sets, to assess the capacity of our method to adapt automatically to different subsets. Philippe C. Besse, Brendan Guillouet, Jean-Michel Loubes, Francois Royer |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Review and Perspective for Distance-Based Clustering of Vehicle TrajectoriesabstractIn this paper, we tackle the issue of clustering trajectories of geolocalized observations based on the distance between trajectories. We first provide a comprehensive review of the different distances used in the literature to compare trajectories. Then, based on the limitations of these methods, we introduce a new distance: symmetrized segment-path distance (SSPD). We compare this new distance to the others according to their corresponding clustering results obtained using both the hierarchical clustering and affinity propagation methods. We finally present a python package: trajectory distance, which contains the methods for calculating the SSPD distance, and the other distances reviewed in this paper. Philippe C. Besse, Brendan Guillouet, Jean-Michel Loubes, Francois Royer |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2014 | Oracle Inequalities for a Group Lasso Procedure Applied to Generalized Linear Models in High DimensionabstractWe present a group lasso procedure for generalized linear models (GLMs) and we study the properties of this estimator applied to sparse high-dimensional GLMs. Under general conditions on the covariates and on the joint distribution of the pair covariates, we provide oracle inequalities promoting group sparsity of the covariables. We get convergence rates for the prediction and estimation error and we show the ability of this estimator to recover good sparse approximation of the true model. Then, we extend this procedure to the case of an elastic net penalty. At last, we apply these results to the so-called Poisson regression model (the output is modeled as a Poisson process whose intensity relies on a linear combination of the covariables). The group lasso method enables to select few groups of meaningful variables among the set of inputs. Mélanie Blazère, Jean-Michel Loubes, Fabrice Gamboa |
IEEE Trans. Inf. Theory | 2 |
| 2011 | Group Lasso Estimation of High-dimensional Covariance Matrices
Jérémie Bigot, Rolando J. Biscay, Jean-Michel Loubes, Lilian Muñiz-Alvarez |
J. Mach. Learn. Res. | 3 |