Patrice Bertail

dblp:66/3243 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 50% Trustworthy machine learning · 25% Transfer learning and domain adaptation · 25%
Theoretical computer science
2 papers
Mathematical optimization · 70% Information theory · 30%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.512021
Learning from Biased Data: A Semi-Parametric Approach · ICML 2021
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.512021
Learning from Biased Data: A Semi-Parametric Approach · ICML 2021
Machine learning › Learning theory › statistical estimation › robust statistics
median-of-means
0.412019
On Medians of (Randomized) Pairwise Means · ICML 2019
Machine learning › Learning theory
statistical learning theory
0.412019
On Medians of (Randomized) Pairwise Means · ICML 2019
Machine learning › Learning theory › statistical estimation
semiparametric estimation
0.112021
Learning from Biased Data: A Semi-Parametric Approach · ICML 2021
Machine learning › Learning theory
empirical risk minimization
0.112019
On Medians of (Randomized) Pairwise Means · ICML 2019
Mathematical optimization › statistical estimation › regression
nonparametric regression
0.112019
On Medians of (Randomized) Pairwise Means · ICML 2019
Mathematical optimization
bootstrap
0.112008
On Bootstrapping the ROC Curve · NIPS 2008
Information theory › statistical inference
confidence bounds
0.112008
On Bootstrapping the ROC Curve · NIPS 2008

Methods — techniques the papers use, named apart from their topics

u-statistics · 0.8bootstrap · 0.8semiparametric estimation · 0.5importance weighting · 0.5smoothed bootstrap · 0.2resampling · 0.2
YearPublicationVenuePosition
2021 Learning from Biased Data: A Semi-Parametric Approach
abstract
We consider risk minimization problems where the (source) distribution $P_S$ of the training observations $Z_1, \ldots, Z_n$ differs from the (target) distribution $P_T$ involved in the risk that one seeks to minimize. Under the natural assumption that $P_S$ dominates $P_T$, \textit{i.e.} $P_T< \! \! Cite this Paper BibTeX @InProceedings{pmlr-v139-bertail21a, title = {Learning from Biased Data: A Semi-Parametric Approach}, author = {Bertail, Patrice and Cl{\'e}men{\c{c}}on, Stephan and Guyonvarch, Yannick and Noiry, Nathan}, booktitle = {Proceedings of the 38th International Conference on Machine Learning}, pages = {803--812}, year = {2021}, editor = {Meila, Marina and Zhang, Tong}, volume = {139}, series = {Proceedings of Machine Learning Research}, month = {18--24 Jul}, publisher = {PMLR}, pdf = {http://proceedings.mlr.press/v139/bertail21a/bertail21a.pdf}, url = {https://proceedings.mlr.press/v139/bertail21a.html}, abstract = {We consider risk minimization problems where the (source) distribution $P_S$ of the training observations $Z_1, \ldots, Z_n$ differs from the (target) distribution $P_T$ involved in the risk that one seeks to minimize. Under the natural assumption that $P_S$ dominates $P_T$, \textit{i.e.} $P_T< \! \! Copy to Clipboard Download Endnote %0 Conference Paper %T Learning from Biased Data: A Semi-Parametric Approach %A Patrice Bertail %A Stephan Clémençon %A Yannick Guyonvarch %A Nathan Noiry %B Proceedings of the 38th International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2021 %E Marina Meila %E Tong Zhang %F pmlr-v139-bertail21a %I PMLR %P 803--812 %U https://proceedings.mlr.press/v139/bertail21a.html %V 139 %X We consider risk minimization problems where the (source) distribution $P_S$ of the training observations $Z_1, \ldots, Z_n$ differs from the (target) distribution $P_T$ involved in the risk that one seeks to minimize. Under the natural assumption that $P_S$ dominates $P_T$, \textit{i.e.} $P_T< \! \! Copy to Clipboard Download APA Bertail, P., Clémençon, S., Guyonvarch, Y. & Noiry, N.. (2021). Learning from Biased Data: A Semi-Parametric Approach. Proceedings of the 38th International Conference on Machine Learning, in Proceedings of Machine Learning Research 139:803-812 Available from https://proceedings.mlr.press/v139/bertail21a.html. Copy to Clipboard Download Related Material Download PDF Supplementary ZIP This site last compiled Sun, 05 Jul 2026 14:53:09 +0000 Github Account Copyright © The authors and PMLR 2026. MLResearchPress
Patrice Bertail, Stéphan Clémençon, Yannick Guyonvarch, Nathan Noiry
ICML1
2019 On Medians of (Randomized) Pairwise Means
abstract
Tournament procedures, recently introduced in the literature, offer an appealing alternative, from a theoretical perspective at least, to the principle of Empirical Risk Minimization in machine learning. Statistical learning by Median-of-Means (MoM) basically consists in segmenting the training data into blocks of equal size and comparing the statistical performance of every pair of candidate decision rules on each data block: that with highest performance on the majority of the blocks is declared as the winner. In the context of nonparametric regression, functions having won all their duels have been shown to outperform empirical risk minimizers w.r.t. the mean squared error under minimal assumptions, while exhibiting robustness properties. It is the purpose of this paper to extend this approach, in order to address other learning problems in particular, for which the performance criterion takes the form of an expectation over pairs of observations rather than over one single observation, as may be the case in pairwise ranking, clustering or metric learning. Precisely, it is proved here that the bounds achieved by MoM are essentially conserved when the blocks are built by means of independent sampling without replacement schemes instead of a simple segmentation. These results are next extended to situations where the risk is related to a pairwise loss function and its empirical counterpart is of the form of a $U$-statistic. Beyond theoretical results guaranteeing the performance of the learning/estimation methods proposed, some numerical experiments provide empirical evidence of their relevance in practice.
Stéphan Clémençon, Pierre Laforgue, Patrice Bertail
ICML3
2016 Learning from Survey Training Samples: Rate Bounds for Horvitz-Thompson Risk Minimizers
abstract
The generalization ability of minimizers of the empirical risk in the context of binary classification has been investigated under a wide variety of complexity assumptions for the collection of classifiers over which optimization is performed. In contrast, the vast majority of the works dedicated to this issue stipulate that the training dataset used to compute the empirical risk functional is composed of i.i.d. observations and involve sharp control of uniform deviation of i.i.d. averages from their expectation. Beyond the cases where training data are drawn uniformly without replacement among a large i.i.d. sample or modelled as a realization of a weakly dependent sequence of r.v.’s, statistical guarantees when the data used to train a classifier are drawn by means of a more general sampling/survey scheme and exhibit a complex dependence structure have not been documented in the literature yet. It is the main purpose of this paper to show that the theory of empirical risk minimization can be extended to situations where statistical learning is based on survey samples and knowledge of the related (first order) inclusion probabilities. Precisely, we prove that minimizing a (possibly biased) weighted version of the empirical risk, refered to as the (approximate) Horvitz-Thompson risk (HT risk), over a class of controlled complexity lead to a rate for the excess risk of the order O_\mathbbP((\kappa_N (\log N)/n)^1/2) with \kappa_N=(n/N)/\min_i≤N\pi_i, when data are sampled by means of a rejective scheme of (deterministic) size n within a statistical population of cardinality N≥n, a generalization of basic \it sampling without replacement with unequal probability weights \pi_i > 0. Extension to other sampling schemes are then established by a coupling argument. Beyond theoretical results, numerical experiments are displayed in order to show the relevance of HT risk minimization and that ignoring the sampling scheme used to generate the training dataset may completely jeopardize the learning procedure.
Stéphan Clémençon, Patrice Bertail, Guillaume Papa
ACML2
2014 Scaling up M-estimation via sampling designs: The Horvitz-Thompson stochastic gradient descent
abstract
In certain situations that shall be undoubtedly more and more common in the Big Data era, the datasets available are so massive that computing statistics over the full sample is hardly feasible, if not unfeasible. A natural approach in this context consists in using survey schemes and substituting the “full data” statistics with their counterparts based on the resulting random samples, of manageable size. It is the purpose of this paper to investigate the impact of survey sampling with unequal inclusion probabilities on (stochastic) gradient descent-based M-estimation methods in large-scale statistical-learning problems. We prove that, in presence of some a priori information, one may significantly reduce the number of terms that must be averaged to estimate the gradient at each step with overwhelming probability, while preserving the asymptotic accuracy. These striking results are described here by limit theorems.
Stéphan Clémençon, Patrice Bertail, Emilie Chautru
IEEE BigData2
2008 On Bootstrapping the ROC Curve
abstract
This paper is devoted to thoroughly investigating how to bootstrap the ROC curve, a widely used visual tool for evaluating the accuracy of test/scoring statistics in the bipartite setup. The issue of confidence bands for the ROC curve is considered and a resampling procedure based on a smooth version of the empirical distribution called the smoothed bootstrap" is introduced. Theoretical arguments and simulation results are presented to show that the "smoothed bootstrap" is preferable to a "naive" bootstrap in order to construct accurate confidence bands."
Patrice Bertail, Stéphan Clémençon, Nicolas Vayatis
NIPS1