Valerio Perrone

dblp:202/1297 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
3since 2021 · last 2021
0009-0009-6923-7712ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Optimization for machine learning · 62% Probabilistic and Bayesian machine learning · 27% Transfer learning and domain adaptation · 6%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 50% High-performance computing · 50%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
hyperparameter optimization
1.642021
Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization · KDD 2021
A Quantile-based Approach for Hyperparameter Transfer Learning · ICML 2020
Learning search spaces for Bayesian optimization: Another view of hyperparameter transfer learning · NeurIPS 2019
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
1.132020
A Quantile-based Approach for Hyperparameter Transfer Learning · ICML 2020
Learning search spaces for Bayesian optimization: Another view of hyperparameter transfer learning · NeurIPS 2019
Scalable Hyperparameter Transfer Learning · NeurIPS 2018
Machine learning › Optimization for machine learning › black-box optimization
zeroth-order optimization
0.512021
Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization · KDD 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.412020
A Quantile-based Approach for Hyperparameter Transfer Learning · ICML 2020
Machine learning › Optimization for machine learning › hyperparameter optimization
multi-objective hyperparameter optimization
0.412020
A Quantile-based Approach for Hyperparameter Transfer Learning · ICML 2020
Machine learning › Learning paradigms › multi-task learning
multi-task transfer learning
0.312018
Scalable Hyperparameter Transfer Learning · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference
0.312018
A Likelihood-Free Inference Framework for Population Genetic Data using Exchangeable Neural Networks · NeurIPS 2018
Bioinformatics and computational biology
likelihood-free inference
0.312018
A Likelihood-Free Inference Framework for Population Genetic Data using Exchangeable Neural Networks · NeurIPS 2018
Bioinformatics and computational biology
population genetics
0.312018
A Likelihood-Free Inference Framework for Population Genetic Data using Exchangeable Neural Networks · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.312017
Poisson Random Fields for Dynamic Feature Models · J. Mach. Learn. Res. 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
feature allocation
0.312017
Poisson Random Fields for Dynamic Feature Models · J. Mach. Learn. Res. 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
indian buffet process
0.312017
Poisson Random Fields for Dynamic Feature Models · J. Mach. Learn. Res. 2017
High-performance computing › performance optimization
auto-tuning
0.112021
Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization · KDD 2021
Cloud and datacenter computing
machine learning as a service
0.112021
Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization · KDD 2021
Data mining › text mining
topic model
0.112017
Poisson Random Fields for Dynamic Feature Models · J. Mach. Learn. Res. 2017

Methods — techniques the papers use, named apart from their topics

warm-starting · 1.0random search · 1.0early stopping · 1.0bayesian optimization · 1.0summary statistic-free inference · 0.7exchangeable neural networks · 0.7thompson sampling · 0.4quantile regression · 0.4gaussian copula · 0.4search space learning · 0.4wright-fisher model · 0.3markov chain monte carlo · 0.3
YearPublicationVenuePosition
2021 Fair Bayesian Optimization
abstract
Given the increasing importance of machine learning (ML) in our lives, several algorithmic fairness techniques have been proposed to mitigate biases in the outcomes of the ML models. However, most of these techniques are specialized to cater to a single family of ML models and a specific definition of fairness, limiting their adaptibility in practice. We introduce a general constrained Bayesian optimization (BO) framework to optimize the performance of any ML model while enforcing one or multiple fairness constraints. BO is a model-agnostic optimization method that has been successfully applied to automatically tune the hyperparameters of ML models. We apply BO with fairness constraints to a range of popular models, including random forests, gradient boosting, and neural networks, showing that we can obtain accurate and fair solutions by acting solely on the hyperparameters. We also show empirically that our approach is competitive with specialized techniques that enforce model-specific fairness constraints, and outperforms preprocessing methods that learn fair representations of the input data. Moreover, our method can be used in synergy with such specialized fairness techniques to tune their hyperparameters. Finally, we study the relationship between fairness and the hyperparameters selected by BO. We observe a correlation between regularization and unbiased models, explaining why acting on the hyperparameters leads to ML models that generalize well and are fair.
Valerio Perrone, Michele Donini, Muhammad Bilal Zafar, Robin Schmucker, Krishnaram Kenthapadi, Cédric Archambeau
AIES1
2021 Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization
abstract
Tuning complex machine learning systems is challenging. Machine learning typically requires to set hyperparameters, be it regularization, architecture, or optimization parameters, whose tuning is critical to achieve good predictive performance. To democratize access to machine learning systems, it is essential to automate the tuning. This paper presents Amazon SageMaker Automatic Model Tuning (AMT), a fully managed system for gradient-free optimization at scale. AMT finds the best version of a trained machine learning model by repeatedly evaluating it with different hyperparameter configurations. It leverages either random search or Bayesian optimization to choose the hyperparameter values resulting in the best model, as measured by the metric chosen by the user. AMT can be used with built-in algorithms, custom algorithms, and Amazon SageMaker pre-built containers for machine learning frameworks. We discuss the core functionality, system architecture, our design principles, and lessons learned. We also describe more advanced features of AMT, such as automated early stopping and warm-starting, showing in experiments their benefits to users.
Valerio Perrone, Huibin Shen, Aida Zolic, Iaroslav Shcherbatyi, Amr Ahmed 0004, Tanya Bansal, Michele Donini, Fela Winkelmolen, Rodolphe Jenatton, Jean Baptiste Faddoul, Barbara Pogorzelska, Miroslav Miladinovic, Krishnaram Kenthapadi, Matthias W. Seeger, Cédric Archambeau
KDD1
2021 A Nonmyopic Approach to Cost-Constrained Bayesian Optimization
abstract
Bayesian optimization (BO) is a popular method for optimizing expensive-to-evaluate black-box functions. BO budgets are typically given in iterations, which implicitly assumes each evaluation has the same cost. In fact, in many BO applications, evaluation costs vary significantly in different regions of the search space. In hyperparameter optimization, the time spent on neural network training increases with layer size; in clinical trials, the monetary cost of drug compounds vary; and in optimal control, control actions have differing complexities. Cost-constrained BO measures convergence with alternative cost metrics such as time, money, or energy, for which the sample efficiency of standard BO methods is ill-suited. For cost-constrained BO, cost efficiency is far more important than sample efficiency. In this paper, we formulate cost-constrained BO as a constrained Markov decision process (CMDP), and develop an efficient rollout approximation to the optimal CMDP policy that takes both the cost and future iterations into account. We validate our method on a collection of hyperparameter optimization problems as well as a sensor set selection application.
Eric Hans Lee, David Eriksson, Valerio Perrone, Matthias W. Seeger
UAI3
2020 A Quantile-based Approach for Hyperparameter Transfer Learning
abstract
Bayesian optimization (BO) is a popular methodology to tune the hyperparameters of expensive black-box functions. Traditionally, BO focuses on a single task at a time and is not designed to leverage information from related functions, such as tuning performance objectives of the same algorithm across multiple datasets. In this work, we introduce a novel approach to achieve transfer learning across different datasets as well as different objectives. The main idea is to regress the mapping from hyperparameter to objective quantiles with a semi-parametric Gaussian Copula distribution, which provides robustness against different scales or outliers that can occur in different tasks. We introduce two methods to leverage this estimation: a Thompson sampling strategy as well as a Gaussian Copula process using such quantile estimate as a prior. We show that these strategies can combine the estimation of multiple objectives such as latency and accuracy, steering the optimization toward faster predictions for the same level of accuracy. Experiments on an extensive set of hyperparameter tuning tasks demonstrate significant improvements over state-of-the-art methods for both hyperparameter optimization and neural architecture search.
David Salinas, Huibin Shen, Valerio Perrone
ICML3
2019 Learning search spaces for Bayesian optimization: Another view of hyperparameter transfer learning
abstract
Bayesian optimization (BO) is a successful methodology to optimize black-box functions that are expensive to evaluate. While traditional methods optimize each black-box function in isolation, there has been recent interest in speeding up BO by transferring knowledge across multiple related black-box functions. In this work, we introduce a method to automatically design the BO search space by relying on evaluations of previous black-box functions. We depart from the common practice of defining a set of arbitrary search ranges a priori by considering search space geometries that are learnt from historical data. This simple, yet effective strategy can be used to endow many existing BO methods with transfer learning properties. Despite its simplicity, we show that our approach considerably boosts BO by reducing the size of the search space, thus accelerating the optimization of a variety of black-box optimization problems. In particular, the proposed approach combined with random search results in a parameter-free, easy-to-implement, robust hyperparameter optimization strategy. We hope it will constitute a natural baseline for further research attempting to warm-start BO.
Valerio Perrone, Huibin Shen
NeurIPS1
2018 A Likelihood-Free Inference Framework for Population Genetic Data using Exchangeable Neural Networks
abstract
An explosion of high-throughput DNA sequencing in the past decade has led to a surge of interest in population-scale inference with whole-genome data. Recent work in population genetics has centered on designing inference methods for relatively simple model classes, and few scalable general-purpose inference techniques exist for more realistic, complex models. To achieve this, two inferential challenges need to be addressed: (1) population data are exchangeable, calling for methods that efficiently exploit the symmetries of the data, and (2) computing likelihoods is intractable as it requires integrating over a set of correlated, extremely high-dimensional latent variables. These challenges are traditionally tackled by likelihood-free methods that use scientific simulators to generate datasets and reduce them to hand-designed, permutation-invariant summary statistics, often leading to inaccurate inference. In this work, we develop an exchangeable neural network that performs summary statistic-free, likelihood-free inference. Our framework can be applied in a black-box fashion across a variety of simulation-based tasks, both within and outside biology. We demonstrate the power of our approach on the recombination hotspot testing problem, outperforming the state-of-the-art.
Jeffrey Chan, Valerio Perrone, Jeffrey P. Spence, Paul A. Jenkins, Sara Mathieson, Yun S. Song
NeurIPS2
2018 Scalable Hyperparameter Transfer Learning
abstract
Bayesian optimization (BO) is a model-based approach for gradient-free black-box function optimization, such as hyperparameter optimization. Typically, BO relies on conventional Gaussian process (GP) regression, whose algorithmic complexity is cubic in the number of evaluations. As a result, GP-based BO cannot leverage large numbers of past function evaluations, for example, to warm-start related BO runs. We propose a multi-task adaptive Bayesian linear regression model for transfer learning in BO, whose complexity is linear in the function evaluations: one Bayesian linear regression model is associated to each black-box function optimization problem (or task), while transfer learning is achieved by coupling the models through a shared deep neural net. Experiments show that the neural net learns a representation suitable for warm-starting the black-box optimization problems and that BO runs can be accelerated when the target black-box function (e.g., validation loss) is learned together with other related signals (e.g., training loss). The proposed method was found to be at least one order of magnitude faster that methods recently published in the literature.
Valerio Perrone, Rodolphe Jenatton, Matthias W. Seeger, Cédric Archambeau
NeurIPS1
2017 Relativistic Monte Carlo
abstract
Hamiltonian Monte Carlo (HMC) is a popular Markov chain Monte Carlo (MCMC) algorithm that generates proposals for a Metropolis-Hastings algorithm by simulating the dynamics of a Hamiltonian system. However, HMC is sensitive to large time discretizations and performs poorly if there is a mismatch between the spatial geometry of the target distribution and the scales of the momentum distribution. In particular the mass matrix of HMC is hard to tune well. In order to alleviate these problems we propose relativistic Hamiltonian Monte Carlo, a version of HMC based on relativistic dynamics that introduces a maximum velocity on particles. We also derive stochastic gradient versions of the algorithm and show that the resulting algorithms bear interesting relationships to gradient clipping, RMSprop, Adagrad and Adam, popular optimisation methods in deep learning. Based on this, we develop relativistic stochastic gradient descent by taking the zero-temperature limit of relativistic stochastic gradient Hamiltonian Monte Carlo. In experiments we show that the relativistic algorithms perform better than classical Newtonian variants and Adam.
Valerio Perrone, Leonard Hasenclever, Yee Whye Teh, Sebastian J. Vollmer
AISTATS2
2017 Poisson Random Fields for Dynamic Feature Models
abstract
We present the Wright-Fisher Indian buffet process (WF- IBP), a probabilistic model for time-dependent data assumed to have been generated by an unknown number of latent features. This model is suitable as a prior in Bayesian nonparametric feature allocation models in which the features underlying the observed data exhibit a dependency structure over time. More specifically, we establish a new framework for generating dependent Indian buffet processes, where the Poisson random field model from population genetics is used as a way of constructing dependent beta processes. Inference in the model is complex, and we describe a sophisticated Markov Chain Monte Carlo algorithm for exact posterior simulation. We apply our construction to develop a nonparametric focused topic model for collections of time-stamped text documents and test it on the full corpus of NIPS papers published from 1987 to 2015.
Valerio Perrone, Paul A. Jenkins, Dario Spanò, Yee Whye Teh
J. Mach. Learn. Res.1