EDBT 2026 Demo / reviewers in the wild / expert
Rémi Bardenet
dblp:09/8412
· DBLP profile ↗
18ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-1094-9493ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Probabilistic and Bayesian machine learning · 54% Optimization for machine learning · 15% Kernel, tree and ensemble methods · 13% | |
| Theoretical computer science
6 papers |
Algorithms and data structures · 51% Mathematical optimization · 46% Approximation and online algorithms · 3% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
determinantal point process |
3.4 | 8 | 2024 | Small coresets via negative dependence: DPPs, linear statistics, and concentration · NeurIPS 2024 Nonparametric estimation of continuous DPPs with kernel methods · NeurIPS 2021 Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGD · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
1.0 | 3 | 2019 | DPPy: DPP Sampling with Python · J. Mach. Learn. Res. 2019 On two ways to use determinantal point processes for Monte Carlo integration · NeurIPS 2019 Zonotope Hit-and-run for Efficient Sampling from Projection DPPs · ICML 2017 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.9 | 2 | 2021 | Nonparametric estimation of continuous DPPs with kernel methods · NeurIPS 2021 Kernel interpolation with continuous volume sampling · ICML 2020 |
Machine learning › Efficient and distributed learning › data selection
coreset selection |
0.8 | 1 | 2024 | Small coresets via negative dependence: DPPs, linear statistics, and concentration · NeurIPS 2024 |
Mathematical optimization › numerical analysis
numerical integration |
0.8 | 2 | 2019 | On two ways to use determinantal point processes for Monte Carlo integration · NeurIPS 2019 Kernel quadrature with DPPs · NeurIPS 2019 |
Machine learning › Optimization for machine learning
mini-batch sampling |
0.5 | 1 | 2021 | Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGD · NeurIPS 2021 |
Machine learning › Learning theory › statistical estimation
nonparametric estimation |
0.5 | 1 | 2021 | Nonparametric estimation of continuous DPPs with kernel methods · NeurIPS 2021 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
representer theorem |
0.5 | 1 | 2021 | Nonparametric estimation of continuous DPPs with kernel methods · NeurIPS 2021 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.5 | 1 | 2021 | Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGD · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.5 | 2 | 2017 | On Markov chain Monte Carlo methods for tall data · J. Mach. Learn. Res. 2017 Towards scaling up Markov chain Monte Carlo: an adaptive subsampling approach · ICML 2014 |
Machine learning › Learning theory › nonparametric regression › kernel regression
kernel interpolation |
0.4 | 1 | 2020 | Kernel interpolation with continuous volume sampling · ICML 2020 |
Algorithms and data structures › matrix approximation
column subset selection |
0.4 | 1 | 2020 | A determinantal point process for column subset selection · J. Mach. Learn. Res. 2020 |
Algorithms and data structures › randomized algorithms › sampling
determinantal point process |
0.4 | 1 | 2020 | A determinantal point process for column subset selection · J. Mach. Learn. Res. 2020 |
Algorithms and data structures › numerical linear algebra
dimensionality reduction |
0.4 | 1 | 2020 | A determinantal point process for column subset selection · J. Mach. Learn. Res. 2020 |
Algorithms and data structures › randomized algorithms
sampling |
0.4 | 1 | 2020 | Kernel interpolation with continuous volume sampling · ICML 2020 |
Algorithms and data structures › randomized algorithms › sampling › determinantal point process
volume sampling |
0.4 | 1 | 2020 | Kernel interpolation with continuous volume sampling · ICML 2020 |
Machine learning › Probabilistic and Bayesian machine learning
stochastic processes |
0.4 | 1 | 2019 | DPPy: DPP Sampling with Python · J. Mach. Learn. Res. 2019 |
Mathematical optimization › numerical analysis › numerical integration › quadrature rules
kernel quadrature |
0.4 | 1 | 2019 | Kernel quadrature with DPPs · NeurIPS 2019 |
Mathematical optimization › numerical analysis › numerical integration
monte carlo integration |
0.4 | 1 | 2019 | On two ways to use determinantal point processes for Monte Carlo integration · NeurIPS 2019 |
Mathematical optimization › numerical analysis › numerical integration
quadrature rules |
0.4 | 1 | 2019 | Kernel quadrature with DPPs · NeurIPS 2019 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.3 | 2 | 2013 | Collaborative hyperparameter tuning · ICML (2) 2013 Algorithms for Hyper-Parameter Optimization · NIPS 2011 |
Machine learning › Learning theory
concentration inequalities |
0.2 | 1 | 2024 | Small coresets via negative dependence: DPPs, linear statistics, and concentration · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
metropolis-hastings |
0.2 | 1 | 2014 | Towards scaling up Markov chain Monte Carlo: an adaptive subsampling approach · ICML 2014 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
scalable MCMC |
0.2 | 1 | 2014 | Towards scaling up Markov chain Monte Carlo: an adaptive subsampling approach · ICML 2014 |
Mathematical optimization › statistical estimation
maximum likelihood estimation |
0.1 | 1 | 2021 | Nonparametric estimation of continuous DPPs with kernel methods · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
kernel density estimation |
0.1 | 1 | 2020 | Kernel interpolation with continuous volume sampling · ICML 2020 |
Approximation and online algorithms
approximation algorithms |
0.1 | 1 | 2020 | A determinantal point process for column subset selection · J. Mach. Learn. Res. 2020 |
Machine learning › Optimization for machine learning › model-based optimization
sequential model-based optimization |
0.1 | 1 | 2011 | Algorithms for Hyper-Parameter Optimization · NIPS 2011 |
Software maintenance and evolution
software libraries |
0.1 | 1 | 2019 | DPPy: DPP Sampling with Python · J. Mach. Learn. Res. 2019 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.1 | 1 | 2010 | Surrogating the surrogate: accelerating Gaussian-process-based global optimization with a mixture cross-entropy algorithm · ICML 2010 |
Methods — techniques the papers use, named apart from their topics
determinantal point process · 4.1kernel methods · 1.0fixed-point algorithm · 1.0RKHS · 1.0MCMC sampling · 0.9negative dependence · 0.8concentration inequalities · 0.8bayesian quadrature · 0.8orthogonal polynomials · 0.5linear statistics sampling · 0.5volume sampling · 0.4jacobi ensemble sampling · 0.4herding · 0.4exact sampling · 0.4central limit theorem · 0.4approximate sampling · 0.4linear programming · 0.3hit-and-run MCMC · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On two fundamental properties of the zeros of spectrograms of noisy signalsabstractThe spatial distribution of the zeros of the spectrogram is significantly altered when a signal is added to white Gaussian noise. The zeros tend to delineate the support of the signal, and deterministic structures form in the presence of interference, as if the zeros were trapped. While sophisticated methods have been proposed to detect signals as holes in the pattern of spectrogram zeros, few formal arguments have been made to support the delineation and trapping effects. Through detailed computations for simple toy signals, we show that two basic mathematical arguments –the intensity of zeros and Rouché’s theorem– allow discussing delineation and trapping, and the influence of parameters like the signal-to-noise ratio. In particular, interfering chirps, even nearly superimposed, yield an easy-to-detect deterministic structure among zeros. Arnaud Poinas, Rémi Bardenet |
Signal Process. | 2 |
| 2024 | Small coresets via negative dependence: DPPs, linear statistics, and concentrationabstractDeterminantal point processes (DPPs) are random configurations of points with tunable negative dependence.
Because sampling is tractable, DPPs are natural candidates for subsampling tasks, such as minibatch selection or coreset construction.
A \emph{coreset} is a subset of a (large) training set, such that minimizing an empirical loss averaged over the coreset is a controlled replacement for the intractable minimization of the original empirical loss.
Typically, the control takes the form of a guarantee that the average loss over the coreset approximates the total loss uniformly across the parameter space.
Recent work has provided significant empirical support in favor of using DPPs to build randomized coresets, coupled with interesting theoretical results that are suggestive but leave some key questions unanswered.
In particular, the central question of whether the cardinality of a DPP-based coreset is fundamentally smaller than one based on independent sampling remained open.
In this paper, we answer this question in the affirmative, demonstrating that \emph{DPPs can provably outperform independently drawn coresets}.
In this vein, we contribute a conceptual understanding of coreset loss as a \emph{linear statistic} of the (random) coreset.
We leverage this structural observation to connect the coresets problem to a more general problem of concentration phenomena for linear statistics of DPPs, wherein we obtain \emph{effective concentration inequalities that extend well-beyond the state-of-the-art}, encompassing general non-projection, even non-symmetric kernels.
The latter have been recently shown to be of interest in machine learning beyond coresets, but come with a limited theoretical toolbox, to the extension of which our result contributes. Finally, we are also able to address the coresets problem for vector-valued objective functions, a novelty in the coresets literature. Rémi Bardenet, Subhroshekhar Ghosh, Hugo Simon-Onfroy, Hoang Son Tran |
NeurIPS | 1 |
| 2023 | Smoothing Complex-Valued Signals on Graphs with Monte-CarloabstractWe introduce new smoothing estimators for complex signals on graphs, based on a recently studied Determinantal Point Process (DPP). These estimators are built from subsets of edges and nodes drawn according to this DPP, making up trees and unicycles, i.e., connected components containing exactly one cycle. We provide a Julia implementation of these estimators and study their performance when applied to a ranking problem. Hugo Jaquard, Michaël Fanuel, Pierre-Olivier Amblard, Rémi Bardenet, Simon Barthelmé, Nicolas Tremblay |
ICASSP | 4 |
| 2021 | Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGDabstractStochastic gradient descent (SGD) is a cornerstone of machine learning. When the number $N$ of data items is large, SGD relies on constructing an unbiased estimator of the gradient of the empirical risk using a small subset of the original dataset, called a minibatch. Default minibatch construction involves uniformly sampling a subset of the desired size, but alternatives have been explored for variance reduction. In particular, experimental evidence suggests drawing minibatches from determinantal point processes (DPPs), tractable distributions over minibatches that favour diversity among selected items. However, like in recent work on DPPs for coresets, providing a systematic and principled understanding of how and why DPPs help has been difficult. In this work, we contribute an orthogonal polynomial-based determinantal point process paradigm for performing minibatch sampling in SGD. Our approach leverages the specific data distribution at hand, which endows it with greater sensitivity and power over existing data-agnostic methods. We substantiate our method via a detailed theoretical analysis of its convergence properties, interweaving between the discrete data set and the underlying continuous domain. In particular, we show how specific DPPs and a string of controlled approximations can lead to gradient estimators with a variance that decays faster with the batchsize than under uniform sampling. Coupled with existing finite-time guarantees for SGD on convex objectives, this entails that, for a large enough batchsize and a fixed budget of item-level gradients to evaluate, DPP minibatches lead to a smaller bound on the mean square approximation error than uniform minibatches. Moreover, our estimators are amenable to a recent algorithm that directly samples linear statistics of DPPs (i.e., the gradient estimator) without sampling the underlying DPP (i.e., the minibatch), thereby reducing computational overhead. We provide detailed synthetic as well as real data experiments to substantiate our theoretical claims. Rémi Bardenet, Subhroshekhar Ghosh, Meixia Lin |
NeurIPS | 1 |
| 2021 | Nonparametric estimation of continuous DPPs with kernel methodsabstractDeterminantal Point Process (DPPs) are statistical models for repulsive point patterns. Both sampling and inference are tractable for DPPs, a rare feature among models with negative dependence that explains their popularity in machine learning and spatial statistics. Parametric and nonparametric inference methods have been proposed in the finite case, i.e. when the point patterns live in a finite ground set. In the continuous case, only parametric methods have been investigated, while nonparametric maximum likelihood for DPPs -- an optimization problem over trace-class operators -- has remained an open question. In this paper, we show that a restricted version of this maximum likelihood (MLE) problem falls within the scope of a recent representer theorem for nonnegative functions in an RKHS. This leads to a finite-dimensional problem, with strong statistical ties to the original MLE. Moreover, we propose, analyze, and demonstrate a fixed point algorithm to solve this finite-dimensional problem. Finally, we also provide a controlled estimate of the correlation kernel of the DPP, thus providing more interpretability. Michaël Fanuel, Rémi Bardenet |
NeurIPS | 2 |
| 2020 | Kernel interpolation with continuous volume samplingabstractA fundamental task in kernel methods is to pick nodes and weights, so as to approximate a given function from an RKHS by the weighted sum of kernel translates located at the nodes. This is the crux of kernel density estimation, kernel quadrature, or interpolation from discrete samples. Furthermore, RKHSs offer a convenient mathematical and computational framework. We introduce and analyse continuous volume sampling (VS), the continuous counterpart -for choosing node locations- of a discrete distribution introduced in (Deshpande & Vempala, 2006). Our contribution is theoretical: we prove almost optimal bounds for interpolation and quadrature under VS. While similar bounds already exist for some specific RKHSs using ad-hoc node constructions, VS offers bounds that apply to any Mercer kernel and depend on the spectrum of the associated integration operator. We emphasize that, unlike previous randomized approaches that rely on regularized leverage scores or determinantal point processes, evaluating the pdf of VS only requires pointwise evaluations of the kernel. VS is thus naturally amenable to MCMC samplers. Ayoub Belhadji, Rémi Bardenet, Pierre Chainais |
ICML | 2 |
| 2020 | A determinantal point process for column subset selectionabstractTwo popular approaches to dimensionality reduction are principal component analysis, which projects onto a small number of well-chosen but non-interpretable directions, and feature selection, which selects a small number of the original features. Feature selection can be abstracted as selecting the subset of columns of a matrix $X \in \mathbb{R}^{N \times d}$ which minimize the approximation error, i.e., the norm of the residual after projecting $X$ onto the space spanned by the selected columns. Such a combinatorial optimization is usually impractical, and there has been interest in polynomial-cost, random subset selection algorithms that favour small values of this approximation error. We propose sampling from a projection determinantal point process, a repulsive distribution over column indices that favours diversity among the selected columns. We bound the ratio of the expected approximation error over the optimal error of PCA. These bounds improve over the state-of-the-art bounds of volume sampling when some realistic structural assumptions are satisfied for $X$. Numerical experiments suggest that our bounds are tight, and that our algorithms have comparable performance with the double phase algorithm, often considered the practical state-of-the-art. Ayoub Belhadji, Rémi Bardenet, Pierre Chainais |
J. Mach. Learn. Res. | 2 |
| 2019 | Kernel quadrature with DPPsabstractWe study quadrature rules for functions living in an RKHS, using nodes sampled from a projection determinantal point process (DPP). DPPs are parametrized by a kernel, and we use a truncated and saturated version of the RKHS kernel. This natural link between the two kernels, along with DPP machinery, leads to relatively tight bounds on the quadrature error, that depends on the spectrum of the RKHS kernel. Finally, we experimentally compare DPPs to existing kernel-based quadratures such as herding, Bayesian quadrature, or continuous leverage score sampling. Numerical results confirm the interest of DPPs, and even suggest faster rates than our bounds in particular cases. Ayoub Belhadji, Rémi Bardenet, Pierre Chainais |
NeurIPS | 2 |
| 2019 | On two ways to use determinantal point processes for Monte Carlo integrationabstractWhen approximating an integral by a weighted sum of function evaluations, determinantal point processes (DPPs) provide a way to enforce repulsion between the evaluation points. This negative dependence is encoded by a kernel. Fifteen years before the discovery of DPPs, Ermakov & Zolotukhin (EZ, 1960) had the intuition of sampling a DPP and solving a linear system to compute an unbiased Monte Carlo estimator of the integral. In the absence of DPP machinery to derive an efficient sampler and analyze their estimator, the idea of Monte Carlo integration with DPPs was stored in the cellar of numerical integration. Recently, Bardenet & Hardy (BH, 2019) came up with a more natural estimator with a fast central limit theorem (CLT). In this paper, we first take the EZ estimator out of the cellar, and analyze it using modern arguments. Second, we provide an efficient implementation to sample exactly a particular multidimensional DPP called multivariate Jacobi ensemble. The latter satisfies the assumptions of the aforementioned CLT. Third, our new implementation lets us investigate the behavior of the two unbiased Monte Carlo estimators in yet unexplored regimes. We demonstrate experimentally good properties when the kernel is adapted to basis of functions in which the integrand is sparse or has fast-decaying coefficients. If such a basis and the level of sparsity are known (e.g., we integrate a linear combination of kernel eigenfunctions), the EZ estimator can be the right choice, but otherwise it can display an erratic behavior. Guillaume Gautier, Rémi Bardenet, Michal Valko |
NeurIPS | 2 |
| 2019 | DPPy: DPP Sampling with PythonabstractDeterminantal point processes (DPPs) are specific probability distributions over clouds of points that are used as models and computational tools across physics, probability, statistics, and more recently machine learning. Sampling from DPPs is a challenge and therefore we present DPPy, a Python toolbox that gathers known exact and approximate sampling algorithms for both finite and continuous DPPs. The project is hosted on GitHub, and equipped with an extensive documentation. Guillaume Gautier, Guillermo Polito, Rémi Bardenet, Michal Valko |
J. Mach. Learn. Res. | 3 |
| 2017 | Zonotope Hit-and-run for Efficient Sampling from Projection DPPsabstractDeterminantal point processes (DPPs) are distributions over sets of items that model diversity using kernels. Their applications in machine learning include summary extraction and recommendation systems. Yet, the cost of sampling from a DPP is prohibitive in large-scale applications, which has triggered an effort towards efficient approximate samplers. We build a novel MCMC sampler that combines ideas from combinatorial geometry, linear programming, and Monte Carlo methods to sample from DPPs with a fixed sample cardinality, also called projection DPPs. Our sampler leverages the ability of the hit-and-run MCMC kernel to efficiently move across convex bodies. Previous theoretical results yield a fast mixing time of our chain when targeting a distribution that is close to a projection DPP, but not a DPP in general. Our empirical results demonstrate that this extends to sampling projection DPPs, i.e., our sampler is more sample-efficient than previous approaches which in turn translates to faster convergence when dealing with costly-to-evaluate functions, such as summary extraction in our experiments. Guillaume Gautier, Rémi Bardenet, Michal Valko |
ICML | 2 |
| 2017 | On Markov chain Monte Carlo methods for tall dataabstractMarkov chain Monte Carlo methods are often deemed too computationally intensive to be of any practical use for big data applications, and in particular for inference on datasets containing a large number $n$ of individual data points, also known as tall datasets. In scenarios where data are assumed independent, various approaches to scale up the Metropolis- Hastings algorithm in a Bayesian inference context have been recently proposed in machine learning and computational statistics. These approaches can be grouped into two categories: divide-and-conquer approaches and, subsampling-based algorithms. The aims of this article are as follows. First, we present a comprehensive review of the existing literature, commenting on the underlying assumptions and theoretical guarantees of each method. Second, by leveraging our understanding of these limitations, we propose an original subsampling-based approach relying on a control variate method which samples under regularity conditions from a distribution provably close to the posterior distribution of interest, yet can require less than $O(n)$ data point likelihood evaluations at each iteration for certain statistical models in favourable scenarios. Finally, we emphasize that we have only been able so far to propose subsampling-based methods which display good performance in scenarios where the Bernstein-von Mises approximation of the target posterior distribution is excellent. It remains an open challenge to develop such methods in scenarios where the Bernstein-von Mises approximation is poor. Rémi Bardenet, Arnaud Doucet, Christopher C. Holmes |
J. Mach. Learn. Res. | 1 |
| 2015 | Inference for determinantal point processes without spectral knowledgeabstractDeterminantal point processes (DPPs) are point process models thatnaturally encode diversity between the points of agiven realization, through a positive definite kernel $K$. DPPs possess desirable properties, such as exactsampling or analyticity of the moments, but learning the parameters ofkernel $K$ through likelihood-based inference is notstraightforward. First, the kernel that appears in thelikelihood is not $K$, but another kernel $L$ related to $K$ throughan often intractable spectral decomposition. This issue is typically bypassed in machine learning bydirectly parametrizing the kernel $L$, at the price of someinterpretability of the model parameters. We follow this approachhere. Second, the likelihood has an intractable normalizingconstant, which takes the form of large determinant in the case of aDPP over a finite set of objects, and the form of a Fredholm determinant in thecase of a DPP over a continuous domain. Our main contribution is to derive bounds on the likelihood ofa DPP, both for finite and continuous domains. Unlike previous work, our bounds arecheap to evaluate since they do not rely on approximating the spectrumof a large matrix or an operator. Through usual arguments, these bounds thus yield cheap variationalinference and moderately expensive exact Markov chain Monte Carlo inference methods for DPPs. Rémi Bardenet, Michalis K. Titsias |
NIPS | 1 |
| 2015 | Ten Simple Rules for a Successful Cross-Disciplinary CollaborationabstractCross-disciplinary collaborations have become an increasingly important part of science. They are seen as a key factor for finding solutions to pressing societal challenges on a global scale including green technologies, sustainable food production and drug development. This has also been realized by regulators and policy-makers, as it is reflected in the 80 billion Euro "Horizon 2020" EU Framework Programme for Research and Innovation. This programme puts special emphasis at breaking down barriers between fields to create a path breaking environment for knowledge, research and innovation. However, igniting and successfully maintaining cross-disciplinary collaborations can be a delicate task. In this article we focus on the specific challenges associated with cross-disciplinary research in particular from the perspective of the theoretician. As research fellows of the 2020 Science project (http://www.2020science.net) and collaboration partners, we bring broad experience of developing interdisciplinary collaborations [2–12]. We intend this guide for early career computational researchers as well as more senior scientists who are entering a cross disciplinary setting for the first time. We describe the key benefits, as well as some possible pitfalls, arising from collaborations between scientists with backgrounds in very different fields. This paper has inter alia been cited by Times Higher education: http://www.timeshighereducation.co.uk/news/people/the-secrets-to-successful-interdisciplinary-work/2020267.article . Bernhard Knapp, Rémi Bardenet, Miguel O. Bernabeu, Rafel Bordas, Maria Bruna, Ben Calderhead, Jonathan Cooper, Alexander G. Fletcher, Derek Groen, Bram Kuijper, Joanna Lewis, Greg J. McInerny, Timo Minssen, James M. Osborne, Verena Paulitschke, Joe Pitt-Francis, Jelena Todoric, Christian A. Yates, David Gavaghan, Charlotte M. Deane |
PLoS Comput. Biol. | 2 |
| 2014 | Towards scaling up Markov chain Monte Carlo: an adaptive subsampling approachabstractMarkov chain Monte Carlo (MCMC) methods are often deemed far too computationally intensive to be of any practical use for large datasets. This paper describes a methodology that aims to scale up the Metropolis-Hastings (MH) algorithm in this context. We propose an approximate implementation of the accept/reject step of MH that only requires evaluating the likelihood of a random subset of the data, yet is guaranteed to coincide with the accept/reject step based on the full dataset with a probability superior to a user-specified tolerance level. This adaptive subsampling technique is an alternative to the recent approach developed in (Korattikara et al, ICML’14), and it allows us to establish rigorously that the resulting approximate MH algorithm samples from a perturbed version of the target distribution of interest, whose total variation distance to this very target is controlled explicitly. We explore the benefits and limitations of this scheme on several examples. Rémi Bardenet, Arnaud Doucet, Christopher C. Holmes |
ICML | 1 |
| 2013 | Collaborative hyperparameter tuningabstractHyperparameter learning has traditionally been a manual task because of the limited number of trials. Today’s computing infrastructures allow bigger evaluation budgets, thus opening the way for algorithmic approaches. Recently, surrogate-based optimization was successfully applied to hyperparameter learning for deep belief networks and to WEKA classifiers. The methods combined brute force computational power with model building about the behavior of the error function in the hyperparameter space, and they could significantly improve on manual hyperparameter tuning. What may make experienced practitioners even better at hyperparameter optimization is their ability to generalize across similar learning problems. In this paper, we propose a generic method to incorporate knowledge from previous experiments when simultaneously tuning a learning algorithm on new problems at hand. To this end, we combine surrogate-based ranking and optimization techniques for surrogate-based collaborative tuning (SCoT). We demonstrate SCoT in two experiments where it outperforms standard tuning techniques and single-problem surrogate-based optimization. Rémi Bardenet, Mátyás Brendel, Balázs Kégl, Michèle Sebag |
ICML (2) | 1 |
| 2011 | Algorithms for Hyper-Parameter OptimizationabstractSeveral recent advances to the state of the art in image classification benchmarks have come from better configurations of existing techniques rather than novel approaches to feature learning. Traditionally, hyper-parameter optimization has been the job of humans because they can be very efficient in regimes where only a few trials are possible. Presently, computer clusters and GPU processors make it possible to run more trials and we show that algorithmic approaches can find better results. We present hyper-parameter optimization results on tasks of training neural networks and deep belief networks (DBNs). We optimize hyper-parameters using random search and two new greedy sequential methods based on the expected improvement criterion. Random search has been shown to be sufficiently efficient for learning neural networks for several datasets, but we show it is unreliable for training DBNs. The sequential algorithms are applied to the most difficult DBN learning problems from [Larochelle et al., 2007] and find significantly better results than the best previously reported. This work contributes novel techniques for making response surface models P (y|x) in which many elements of hyper-parameter assignment (x) are known to be irrelevant given particular values of other elements. James Bergstra, Rémi Bardenet, Yoshua Bengio, Balázs Kégl |
NIPS | 2 |
| 2010 | Surrogating the surrogate: accelerating Gaussian-process-based global optimization with a mixture cross-entropy algorithm
Rémi Bardenet, Balázs Kégl |
ICML | 1 |