Achraf Bahamou

dblp:255/6416 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0002-8448-6612ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2023 A Mini-Block Fisher Method for Deep Neural Networks
abstract
Deep Neural Networks (DNNs) are currently predominantly trained using first-order methods. Some of these methods (e.g., Adam, AdaGrad, and RMSprop, and their variants) incorporate a small amount of curvature information by using a diagonal matrix to precondition the stochastic gradient. Recently, effective second-order methods, such as KFAC, K-BFGS, Shampoo, and TNT, have been developed for training DNNs, by preconditioning the stochastic gradient by layer-wise block-diagonal matrices. Here we propose a “mini-block Fisher (MBF)” preconditioned stochastic gradient method, that lies in between these two classes of methods. Specifically, our method uses a block-diagonal approximation to the empirical Fisher matrix, where for each layer in the DNN, whether it is convolutional or feed-forward and fully connected, the associated diagonal block is itself block-diagonal and is composed of a large number of mini-blocks of modest size. Our novel approach utilizes the parallelism of GPUs to efficiently perform computations on the large number of matrices in each layer. Consequently, MBF’s per-iteration computational cost is only slightly higher than it is for first-order methods. The performance of MBF is compared to that of several baseline methods, on Autoencoder, Convolutional Neural Network (CNN), and Graph Convolutional Network (GCN) problems, to validate its effectiveness both in terms of time efficiency and generalization power. Finally, it is proved that an idealized version of MBF converges linearly.
Achraf Bahamou, Donald Goldfarb, Yi Ren 0007
AISTATS1
2021 Optimal Pricing with a Single Point
abstract
We study the following fundamental data-driven pricing problem. How can/should a decision-maker price its product based on observations at a single historical price? The decision-maker optimizes over (potentially randomized) pricing policies to maximize the worst-case ratio of the revenue it can garner compared to an oracle with full knowledge of the distribution of values, when the latter is only assumed to belong to broad non-parametric set. In particular, our framework applies to the widely used regular and monotone non-decreasing hazard rate (mhr) classes of distributions. For settings where the seller knows the exact probability of sale associated with one historical price or only a confidence interval for it, we fully characterize optimal performance and near-optimal pricing algorithms that adjust to the information at hand. As examples, against mhr distributions, we show that it is possible to guarantee $85%$ of oracle performance if one knows that half of the customers have bought at the historical price, and if only $1%$ of the customers bought, it still possible to guarantee $51%$ of oracle performance. The framework we develop leads to new insights on the value of information for pricing, as well as the value of randomization. In addition, it is general and allows to characterize optimal deterministic mechanisms and incorporate uncertainty in the probability of sale.
Amine Allouah, Achraf Bahamou, Omar Besbes
EC2
2021 Revenue Maximization from Finite Samples
abstract
In the present paper, we study the following fundamental problem: how should a decision-maker price based on a finite and limited number of samples from the distribution of values of customers. The decision-maker's objective is to select a pricing policy with maximum competitive ratio when the value distribution is only known to belong to some general non-parametric class. We study achievable performance for two central classes, regular and monotone hazard rate (mhr) distributions, through a general framework. To date, only results are available for a single sample and two samples. We improve existing results but also obtain the first results on achievable performance as the number of samples increases. At a higher level, this work also provides insights on the value of samples for pricing purposes. For example, against mhr distributions (resp. regular), two samples suffice to ensure 71% (resp. 61%) of optimal oracle performance, and ten samples guarantee $80%$ (resp. $65%$) of such performance. Our analysis relies on the introduction of a new (simple) class of policies and the derivation of tractable lower bounds on their performance through factor revealing dynamic programs.
Amine Allouah, Achraf Bahamou, Omar Besbes
EC2
2020 Stochastic Flows and Geometric Optimization on the Orthogonal Group
abstract
We present a new class of stochastic, geometrically-driven optimization algorithms on the orthogonal group O(d) and naturally reductive homogeneous manifolds obtained from the action of the rotation group SO(d). We theoretically and experimentally demonstrate that our methods can be applied in various fields of machine learning including deep, convolutional and recurrent neural networks, reinforcement learning, normalizing flows and metric learning. We show an intriguing connection between efficient stochastic optimization on the orthogonal group and graph theory (e.g. matching problem, partition functions over graphs, graph-coloring). We leverage the theory of Lie groups and provide theoretical results for the designed class of algorithms. We demonstrate broad applicability of our methods by showing strong performance on the seemingly unrelated tasks of learning world models to obtain stable policies for the most difficult Humanoid agent from OpenAI Gym and improving convolutional neural networks.
Krzysztof Choromanski, David Cheikhi, Jared Davis, Valerii Likhosherstov, Achille Nazaret, Achraf Bahamou, Xingyou Song, Mrugank Akarte, Jack Parker-Holder, Jacob Bergquist, Yuan Gao 0038, Aldo Pacchiano, Tamás Sarlós, Adrian Weller, Vikas Sindhwani
ICML6
2020 Practical Quasi-Newton Methods for Training Deep Neural Networks
abstract
We consider the development of practical stochastic quasi-Newton, and in particular Kronecker-factored block diagonal BFGS and L-BFGS methods, for training deep neural networks (DNNs). In DNN training, the number of variables and components of the gradient n is often of the order of tens of millions and the Hessian has n^2 elements. Consequently, computing and storing a full n times n BFGS approximation or storing a modest number of (step, change in gradient) vector pairs for use in an L-BFGS implementation is out of the question. In our proposed methods, we approximate the Hessian by a block-diagonal matrix and use the structure of the gradient and Hessian to further approximate these blocks, each of which corresponds to a layer, as the Kronecker product of two much smaller matrices. This is analogous to the approach in KFAC , which computes a Kronecker-factored block diagonal approximation to the Fisher matrix in a stochastic natural gradient method. Because the indefinite and highly variable nature of the Hessian in a DNN, we also propose a new damping approach to keep the upper as well as the lower bounds of the BFGS and L-BFGS approximations bounded. In tests on autoencoder feed-forward network models with either nine or thirteen layers applied to three datasets, our methods outperformed or performed comparably to KFAC and state-of-the-art first-order stochastic methods.
Donald Goldfarb, Yi Ren 0007, Achraf Bahamou
NeurIPS3