Leonard Berrada

dblp:190/7725 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Optimization for machine learning · 28% Trustworthy machine learning · 28% Deep learning architectures and training · 22%
Software engineering, system software, and programming languages
1 paper
Program verification · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate
1.022022
A Stochastic Bundle Method for Interpolation · J. Mach. Learn. Res. 2022
Training Neural Networks for and by Interpolation · ICML 2020
Machine learning › Trustworthy machine learning
robustness
0.512021
Make Sure You're Unsure: A Framework for Verifying Probabilistic Specifications · NeurIPS 2021
Program verification
neural network verification
0.512021
Make Sure You're Unsure: A Framework for Verifying Probabilistic Specifications · NeurIPS 2021
Machine learning › Deep learning architectures and training
training optimization
0.412019
Deep Frank-Wolfe For Neural Network Optimization · ICLR (Poster) 2019
Machine learning › Learning theory
loss function
0.312018
Smooth Loss Functions for Deep Top-k Classification · ICLR (Poster) 2018
Computer vision › Image recognition and object detection › image classification
top-k classification
0.312018
Smooth Loss Functions for Deep Top-k Classification · ICLR (Poster) 2018
Machine learning › Deep learning architectures and training › feedforward neural network
piecewise linear network
0.312017
Trusting SVM for Piecewise Linear CNNs · ICLR (Poster) 2017
Mathematical optimization
support vector machine
0.312017
Trusting SVM for Piecewise Linear CNNs · ICLR (Poster) 2017
Machine learning › Learning theory
generalization
0.212022
A Stochastic Bundle Method for Interpolation · J. Mach. Learn. Res. 2022

Methods — techniques the papers use, named apart from their topics

lagrangian duality · 1.0functional multipliers · 1.0stochastic approximation · 0.6convex optimization · 0.6bundle method · 0.6SVM · 0.6stochastic gradient descent · 0.4interpolation · 0.4frank-wolfe · 0.4smooth loss function · 0.3
YearPublicationVenuePosition
2022 A Stochastic Bundle Method for Interpolation
abstract
We propose a novel method for training deep neural networks that are capable of interpolation, that is, driving the empirical loss to zero. At each iteration, our method constructs a stochastic approximation of the learning objective. The approximation, known as a bundle, is a pointwise maximum of linear functions. Our bundle contains a constant function that lower bounds the empirical loss. This enables us to compute an automatic adaptive learning rate, thereby providing an accurate solution. In addition, our bundle includes linear approximations computed at the current iterate and other linear estimates of the DNN parameters. The use of these additional approximations makes our method significantly more robust to its hyperparameters. Based on its desirable empirical properties, we term our method Bundle Optimisation for Robust and Accurate Training (BORAT). In order to operationalise BORAT, we design a novel algorithm for optimising the bundle approximation efficiently at each iteration. We establish the theoretical convergence of BORAT in both convex and non-convex settings. Using standard publicly available data sets, we provide a thorough comparison of BORAT to other single hyperparameter optimisation algorithms. Our experiments demonstrate BORAT matches the state-of-the-art generalisation performance for these methods and is the most robust.
Alasdair Paren, Leonard Berrada, Rudra P. K. Poudel, M. Pawan Kumar
J. Mach. Learn. Res.2
2021 Make Sure You're Unsure: A Framework for Verifying Probabilistic Specifications
abstract
Most real world applications require dealing with stochasticity like sensor noise or predictive uncertainty, where formal specifications of desired behavior are inherently probabilistic. Despite the promise of formal verification in ensuring the reliability of neural networks, progress in the direction of probabilistic specifications has been limited. In this direction, we first introduce a general formulation of probabilistic specifications for neural networks, which captures both probabilistic networks (e.g., Bayesian neural networks, MC-Dropout networks) and uncertain inputs (distributions over inputs arising from sensor noise or other perturbations). We then propose a general technique to verify such specifications by generalizing the notion of Lagrangian duality, replacing standard Lagrangian multipliers with "functional multipliers" that can be arbitrary functions of the activations at a given layer. We show that an optimal choice of functional multipliers leads to exact verification (i.e., sound and complete verification), and for specific forms of multipliers, we develop tractable practical verification algorithms. We empirically validate our algorithms by applying them to Bayesian Neural Networks (BNNs) and MC Dropout Networks, and certifying properties such as adversarial robustness and robust detection of out-of-distribution (OOD) data. On these tasks we are able to provide significantly stronger guarantees when compared to prior work -- for instance, for a VGG-64 MC-Dropout CNN trained on CIFAR-10 in a verification-agnostic manner, we improve the certified AUC (a verified lower bound on the true AUC) for robust OOD detection (on CIFAR-100) from $0 \% \rightarrow 29\%$. Similarly, for a BNN trained on MNIST, we improve on the $\ell_\infty$ robust accuracy from $60.2 \% \rightarrow 74.6\%$. Further, on a novel specification -- distributionally robust OOD detection -- we improve on the certified AUC from $5\% \rightarrow 23\%$.
Leonard Berrada, Sumanth Dathathri, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Jonathan Uesato, Sven Gowal, M. Pawan Kumar
NeurIPS1
2020 Training Neural Networks for and by Interpolation
abstract
In modern supervised learning, many deep neural networks are able to interpolate the data: the empirical loss can be driven to near zero on all samples simultaneously. In this work, we explicitly exploit this interpolation property for the design of a new optimization algorithm for deep learning, which we term Adaptive Learning-rates for Interpolation with Gradients (ALI-G). ALI-G retains the two main advantages of Stochastic Gradient Descent (SGD), which are (i) a low computational cost per iteration and (ii) good generalization performance in practice. At each iteration, ALI-G exploits the interpolation property to compute an adaptive learning-rate in closed form. In addition, ALI-G clips the learning-rate to a maximal value, which we prove to be helpful for non-convex problems. Crucially, in contrast to the learning-rate of SGD, the maximal learning-rate of ALI-G does not require a decay schedule. This makes ALI-G considerably easier to tune than SGD. We prove the convergence of ALI-G in various stochastic settings. Notably, we tackle the realistic case where the interpolation property is satisfied up to some tolerance. We also provide experiments on a variety of deep learning architectures and tasks: (i) learning a differentiable neural computer; (ii) training a wide residual network on the SVHN data set; (iii) training a Bi-LSTM on the SNLI data set; and (iv) training wide residual networks and densely connected networks on the CIFAR data sets. ALI-G produces state-of-the-art results among adaptive methods, and even yields comparable performance with SGD, which requires manually tuned learning-rate schedules. Furthermore, ALI-G is simple to implement in any standard deep learning framework and can be used as a drop-in replacement in existing code.
Leonard Berrada, Andrew Zisserman, M. Pawan Kumar
ICML1
2019 Deep Frank-Wolfe For Neural Network Optimization
Leonard Berrada, Andrew Zisserman, M. Pawan Kumar
ICLR (Poster)1
2018 Smooth Loss Functions for Deep Top-k Classification
Leonard Berrada, Andrew Zisserman, M. Pawan Kumar
ICLR (Poster)1
2017 Trusting SVM for Piecewise Linear CNNs
Leonard Berrada, Andrew Zisserman, M. Pawan Kumar
ICLR (Poster)1