VLDB 2026 Research / reviewers in the wild / expert
Jérôme Bolte
dblp:09/1620
· DBLP profile ↗
11ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-1676-8407ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSA: Improving Performance With a Better Scoring FunctionabstractWhile transformer models exhibit strong incontext learning (ICL) abilities, they often fail to generalize under simple distribution shifts.We analyze these failures and identify Softmax, the scoring function in the attention mechanism, as a contributing factor.We propose Scaled Signed Averaging (SSA), a novel attention scoring function that mitigates these failures.SSA significantly improves performance on our ICL tasks and outperforms transformer models with Softmax on several NLP benchmarks and linguistic probing tasks, in both decoder-only and encoder-only architectures. Omar Naim, Swarnadeep Bhar, Jérôme Bolte, Nicholas Asher |
ACL (1) | 3 |
| 2025 | When majority rules, minority loses: bias amplification of gradient descentabstractDespite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that neglect minority-specific features. Assuming population and variance imbalance, our analysis reveals three key findings: (i) the close proximity between "full-data" and stereotypical predictors, (ii) the dominance of a region where training the entire model tends to merely learn the majority traits, and (iii) a lower bound on the additional training required. Our results are illustrated through experiments in deep learning for tabular and image classification tasks. François Bachoc, Jérôme Bolte, Ryan Boustany, Jean-Michel Loubes |
NeurIPS | 2 |
| 2023 | On the complexity of nonsmooth automatic differentiation
Jérôme Bolte, Ryan Boustany, Edouard Pauwels, Béatrice Pesquet-Popescu |
ICLR | 1 |
| 2023 | One-step differentiation of iterative algorithmsabstractIn appropriate frameworks, automatic differentiation is transparent to the user, at the cost of being a significant computational burden when the number of operations is large. For iterative algorithms, implicit differentiation alleviates this issue but requires custom implementation of Jacobian evaluation. In this paper, we study one-step differentiation, also known as Jacobian-free backpropagation, a method as easy as automatic differentiation and as performant as implicit differentiation for fast algorithms (e.g. superlinear optimization methods). We provide a complete theoretical approximation analysis with specific examples (Newton's method, gradient descent) along with its consequences in bilevel optimization. Several numerical examples illustrate the well-foundness of the one-step estimator. Jérôme Bolte, Edouard Pauwels, Samuel Vaiter |
NeurIPS | 1 |
| 2022 | Automatic differentiation of nonsmooth iterative algorithmsabstractDifferentiation along algorithms, i.e., piggyback propagation of derivatives, is now routinely used to differentiate iterative solvers in differentiable programming. Asymptotics is well understood for many smooth problems but the nondifferentiable case is hardly considered. Is there a limiting object for nonsmooth piggyback automatic differentiation (AD)? Does it have any variational meaning and can it be used effectively in machine learning? Is there a connection with classical derivative? All these questions are addressed under appropriate contractivity conditions in the framework of conservative derivatives which has proved useful in understanding nonsmooth AD. For nonsmooth piggyback iterations, we characterize the attractor set of nonsmooth piggyback iterations as a set-valued fixed point which remains in the conservative framework. This has various consequences and in particular almost everywhere convergence of classical derivatives. Our results are illustrated on parametric convex optimization problems with forward-backward, Douglas-Rachford and Alternating Direction of Multiplier algorithms as well as the Heavy-Ball method. Jérôme Bolte, Edouard Pauwels, Samuel Vaiter |
NeurIPS | 1 |
| 2022 | Second-Order Step-Size Tuning of SGD for Non-Convex Optimization
Camille Castera, Jérôme Bolte, Cédric Févotte, Edouard Pauwels |
Neural Process. Lett. | 2 |
| 2021 | Numerical influence of ReLU'(0) on backpropagationabstractIn theory, the choice of ReLU(0) in [0, 1] for a neural network has a negligible influence both on backpropagation and training. Yet, in the real world, 32 bits default precision combined with the size of deep learning problems makes it a hyperparameter of training methods. We investigate the importance of the value of ReLU'(0) for several precision levels (16, 32, 64 bits), on various networks (fully connected, VGG, ResNet) and datasets (MNIST, CIFAR10, SVHN, ImageNet). We observe considerable variations of backpropagation outputs which occur around half of the time in 32 bits precision. The effect disappears with double precision, while it is systematic at 16 bits. For vanilla SGD training, the choice ReLU'(0) = 0 seems to be the most efficient. For our experiments on ImageNet the gain in test accuracy over ReLU'(0) = 1 was more than 10 points (two runs). We also evidence that reconditioning approaches as batch-norm or ADAM tend to buffer the influence of ReLU'(0)’s value. Overall, the message we convey is that algorithmic differentiation of nonsmooth problems potentially hides parameters that could be tuned advantageously. David Bertoin, Jérôme Bolte, Sébastien Gerchinovitz, Edouard Pauwels |
NeurIPS | 2 |
| 2021 | Nonsmooth Implicit Differentiation for Machine-Learning and OptimizationabstractIn view of training increasingly complex learning architectures, we establish a nonsmooth implicit function theorem with an operational calculus. Our result applies to most practical problems (i.e., definable problems) provided that a nonsmooth form of the classical invertibility condition is fulfilled. This approach allows for formal subdifferentiation: for instance, replacing derivatives by Clarke Jacobians in the usual differentiation formulas is fully justified for a wide class of nonsmooth problems. Moreover this calculus is entirely compatible with algorithmic differentiation (e.g., backpropagation). We provide several applications such as training deep equilibrium networks, training neural nets with conic optimization layers, or hyperparameter-tuning for nonsmooth Lasso-type models. To show the sharpness of our assumptions, we present numerical experiments showcasing the extremely pathological gradient dynamics one can encounter when applying implicit algorithmic differentiation without any hypothesis. Jérôme Bolte, Tam Le, Edouard Pauwels, Antonio Silveti-Falls |
NeurIPS | 1 |
| 2021 | An Inertial Newton Algorithm for Deep LearningabstractWe introduce a new second-order inertial optimization method for machine learning called INNA. It exploits the geometry of the loss function while only requiring stochastic approximations of the function values and the generalized gradients. This makes INNA fully implementable and adapted to large-scale optimization problems such as the training of deep neural networks. The algorithm combines both gradient-descent and Newton-like behaviors as well as inertia. We prove the convergence of INNA for most deep learning problems. To do so, we provide a well-suited framework to analyze deep learning loss functions involving tame optimization in which we study a continuous dynamical system together with its discrete stochastic approximations. We prove sublinear convergence for the continuous-time differential inclusion which underlies our algorithm. Additionally, we also show how standard optimization mini-batch methods applied to non-smooth non-convex problems can yield a certain type of spurious stationary points never discussed before. We address this issue by providing a theoretical framework around the new idea of $D$-criticality; we then give a simple asymptotic analysis of INNA. Our algorithm allows for using an aggressive learning rate of $o(1/\log k)$. From an empirical viewpoint, we show that INNA returns competitive results with respect to state of the art (stochastic gradient descent, ADAGRAD, ADAM) on popular deep learning benchmark problems. Camille Castera, Jérôme Bolte, Cédric Févotte, Edouard Pauwels |
J. Mach. Learn. Res. | 2 |
| 2020 | A mathematical model for automatic differentiation in machine learningabstractAutomatic differentiation, as implemented today, does not have a simple mathematical model adapted to the needs of modern machine learning. In this work we articulate the relationships between differentiation of programs as implemented in practice, and differentiation of nonsmooth functions. To this end we provide a simple class of functions, a nonsmooth calculus, and show how they apply to stochastic approximation methods. We also evidence the issue of artificial critical points created by algorithmic differentiation and show how usual methods avoid these points with probability one. Jérôme Bolte, Edouard Pauwels |
NeurIPS | 1 |
| 2010 | Alternating proximal algorithm for blind image recoveryabstractWe consider a variational formulation of blind image recovery problems. A novel iterative proximal algorithm is proposed to solve the associated nonconvex minimization problem. Under suitable assumptions, this algorithm is shown to have better convergence properties than standard alternating minimization techniques. The objective function includes a smooth convex data fidelity term and nonsmooth convex regularization terms modeling prior information on the data and on the unknown linear degradation operator. A novelty of our approach is to bring into play recent nonsmooth analysis results. The pertinence of the proposed method is illustrated in an image restoration example. Jérôme Bolte, Patrick L. Combettes, Jean-Christophe Pesquet |
ICIP | 1 |