Vincent Roulet

dblp:164/6165 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Deep learning architectures and training · 22% Trustworthy machine learning · 19% Optimization for machine learning · 19%
Theoretical computer science
5 papers
Mathematical optimization · 98% Algorithms and data structures · 2%

Topics — the 29 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
continuous optimization
1.222025
Loss Functions and Operators Generated by f-Divergences · ICML 2025
Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019
Mathematical optimization › continuous optimization
convex optimization
1.222025
Loss Functions and Operators Generated by f-Divergences · ICML 2025
Sharpness, Restart and Acceleration · NIPS 2017
Machine learning › Optimization for machine learning
convergence analysis
0.912025
On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control · J. Mach. Learn. Res. 2025
Machine learning › Generative modeling
energy-based model
0.912025
Joint Learning of Energy-based Models and their Partition Function · ICML 2025
Natural language and speech › Language models and text generation
large language model training
0.912025
Loss Functions and Operators Generated by f-Divergences · ICML 2025
Machine learning › Deep learning architectures and training
loss function design
0.912025
Loss Functions and Operators Generated by f-Divergences · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
partition function estimation
0.912025
Joint Learning of Energy-based Models and their Partition Function · ICML 2025
Robotics › Motion planning and robot control
trajectory optimization
0.912025
On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control · J. Mach. Learn. Res. 2025
Mathematical optimization › continuous optimization
nonlinear optimization
0.912025
On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control · J. Mach. Learn. Res. 2025
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization
0.812024
Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.812024
Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024
Machine learning › Deep learning architectures and training › training optimization
learning rate tuning
0.812024
Stepping on the Edge: Curvature Aware Learning Rate Tuners · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024
Machine learning › Optimization for machine learning
stochastic gradient descent
0.812024
Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Stepping on the Edge: Curvature Aware Learning Rate Tuners · NeurIPS 2024
Mathematical optimization › least squares
gauss-newton method
0.412019
Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019
Mathematical optimization › continuous optimization › convex optimization › first-order methods
gradient-based optimization
0.412019
Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.312018
A Smoother Way to Train Structured Prediction Models · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured SVM
0.312018
A Smoother Way to Train Structured Prediction Models · NeurIPS 2018
Machine learning › Optimization for machine learning › variance reduction
SVRG
0.312018
A Smoother Way to Train Structured Prediction Models · NeurIPS 2018
Machine learning › Optimization for machine learning
convergence acceleration
0.312017
Integration Methods and Optimization Algorithms · NIPS 2017
Mathematical optimization
convergence analysis
0.312017
Sharpness, Restart and Acceleration · NIPS 2017
Mathematical optimization
restart strategies
0.312017
Sharpness, Restart and Acceleration · NIPS 2017
Machine learning › Deep learning architectures and training › loss function design
fenchel-young losses
0.312025
Joint Learning of Energy-based Models and their Partition Function · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
0.312025
Joint Learning of Energy-based Models and their Partition Function · ICML 2025
Algorithms and data structures
dynamic programming
0.112019
Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112018
A Smoother Way to Train Structured Prediction Models · NeurIPS 2018
Computer vision › Image recognition and object detection
object localization
0.112018
A Smoother Way to Train Structured Prediction Models · NeurIPS 2018
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
gradient flow
0.112017
Integration Methods and Optimization Algorithms · NIPS 2017

Methods — techniques the papers use, named apart from their topics

tsallis α-negentropy · 1.7stochastic approximation · 1.7gauss-newton method · 1.7f-divergence · 1.7bisection algorithm · 1.7stochastic gradient descent · 1.6min-min formulation · 0.9variance reduction · 0.8hessian eigenvalue analysis · 0.8bias reduction · 0.8gradient backpropagation · 0.4gauss-newton regularization · 0.4łojasiewicz inequality · 0.3multi-step integration schemes · 0.3acceleration · 0.3
YearPublicationVenuePosition
2025 Loss Functions and Operators Generated by f-Divergences
abstract
The logistic loss (a.k.a. cross-entropy loss) is one of the most popular loss functions used for multiclass classification. It is also the loss function of choice for next-token prediction in language modeling. It is associated with the Kullback-Leibler (KL) divergence and the softargmax operator. In this work, we propose to construct new convex loss functions based on $f$-divergences. Our loss functions generalize the logistic loss in two directions: i) by replacing the KL divergence with $f$-divergences and ii) by allowing non-uniform reference measures. We instantiate our framework for numerous $f$-divergences, recovering existing losses and creating new ones. By analogy with the logistic loss, the loss function generated by an $f$-divergence is associated with an operator, that we dub $f$-softargmax. We derive a novel parallelizable bisection algorithm for computing the $f$-softargmax associated with any $f$-divergence. On the empirical side, one of the goals of this paper is to determine the effectiveness of loss functions beyond the classical cross-entropy in a language model setting, including on pre-training, post-training (SFT) and distillation. We show that the loss function generated by the $\alpha$-divergence (which is equivalent to Tsallis $\alpha$-negentropy in the case of unit reference measures) with $\alpha=1.5$ performs well across several tasks.
Vincent Roulet, Tianlin Liu, Nino Vieillard, Michael E. Sander, Mathieu Blondel
ICML1
2025 Joint Learning of Energy-based Models and their Partition Function
abstract
Energy-based models (EBMs) offer a flexible framework for parameterizing probability distributions using neural networks. However, learning EBMs by exact maximum likelihood estimation (MLE) is generally intractable, due to the need to compute the partition function. In this paper, we propose a novel min-min formulation for approximately learning probabilistic EBMs in combinatorially-large discrete spaces, such as sets or permutations. Our key idea is to jointly learn both an energy model and its log-partition, parameterized as a neural network. Our approach not only provides a novel tractable objective criterion to learn EBMs by stochastic gradient descent (without relying on MCMC), but also a novel means to estimate the log-partition function on unseen data points. On the theoretical side, we show that our approach recovers the optimal MLE solution when optimizing in the space of continuous functions. Furthermore, we show that our approach naturally extends to the broader family of Fenchel-Young losses, allowing us to obtain the first tractable method for optimizing the sparsemax loss in combinatorially-large spaces. We demonstrate our approach on multilabel classification and label ranking.
Michael E. Sander, Vincent Roulet, Tianlin Liu, Mathieu Blondel
ICML2
2025 On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control
abstract
A classical approach for solving discrete time nonlinear control on a finite horizon consists in repeatedly minimizing linear quadratic approximations of the original problem around current candidate solutions. While widely popular in many domains, such an approach has mainly been analyzed locally. We provide detailed convergence guarantees to stationary points as well as local linear convergence rates for the Iterative Linear Quadratic Regulator (ILQR) algorithm and its Differential Dynamic Programming (DDP) variant. For problems without costs on control variables, we observe that global convergence to minima can be ensured provided that the linearized discrete time dynamics are surjective, costs on the state variables are gradient dominated. We further detail quadratic local convergence when the costs are self-concordant. We show that surjectivity of the linearized dynamics hold for appropriate discretization schemes given the existence of a feedback linearization scheme. We present complexity bounds of algorithms based on linear quadratic approximations through the lens of generalized Gauss-Newton methods. Our analysis uncovers several convergence phases for regularized generalized Gauss-Newton algorithms.
Vincent Roulet, Siddhartha S. Srinivasa, Maryam Fazel, Zaïd Harchaoui
J. Mach. Learn. Res.1
2024 Distributionally Robust Optimization with Bias and Variance Reduction
abstract
We consider the distributionally robust optimization (DRO) problem, wherein a learner optimizes the worst-case empirical risk achievable by reweighing the observed training examples. We present Prospect, a stochastic gradient-based algorithm that only requires tuning a single learning rate hyperparameter, and prove that it enjoys linear convergence for smooth regularized losses. This contrasts with previous algorithms that either require tuning multiple hyperparameters or potentially fail to converge due to biased gradient estimates or inadequate regularization. Empirically, we show that Prospect can converge 2-3x faster than baselines such as SGD and stochastic saddle-point methods on distribution shift and fairness benchmarks spanning tabular, vision, and language domains.
Ronak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd Harchaoui
ICLR2
2024 Stepping on the Edge: Curvature Aware Learning Rate Tuners
abstract
Curvature information -- particularly, the largest eigenvalue of the loss Hessian, known as the sharpness -- often forms the basis for learning rate tuners. However, recent work has shown that the curvature information undergoes complex dynamics during training, going from a phase of increasing sharpness to eventual stabilization. We analyze the closed-loop feedback effect between learning rate tuning and curvature. We find that classical learning rate tuners may yield greater one-step loss reduction, yet they ultimately underperform in the long term when compared to constant learning rates in the full batch regime. These models break the stabilization of the sharpness, which we explain using a simplified model of the joint dynamics of the learning rate and the curvature. To further investigate these effects, we introduce a new learning rate tuning method, Curvature Dynamics Aware Tuning (CDAT), which prioritizes long term curvature stabilization over instantaneous progress on the objective. In the full batch regime, CDAT shows behavior akin to prefixed warm-up schedules on deep learning objectives, outperforming tuned constant learning rates. In the mini batch regime, we observe that stochasticity introduces confounding effects that explain the previous success of some learning rate tuners at appropriate batch sizes. Our findings highlight the critical role of understanding the joint dynamics of the learning rate and curvature, beyond greedy minimization, to diagnose failures and design effective adaptive learning rate tuners.
Vincent Roulet, Atish Agarwala, Jean-Bastien Grill, Grzegorz Swirszcz, Mathieu Blondel, Fabian Pedregosa
NeurIPS1
2023 Stochastic Optimization for Spectral Risk Measures
abstract
Spectral risk objectives – also called L-risks – allow for learning systems to interpolate between optimizing average-case performance (as in empirical risk minimization) and worst-case performance on a task. We develop LSVRG, a stochastic algorithm to optimize these quantities by characterizing their subdifferential and addressing challenges such as biasedness of subgradient estimates and non-smoothness of the objective. We show theoretically and experimentally that out-of-the-box approaches such as stochastic subgradient and dual averaging can be hindered by bias, whereas our approach exhibits linear convergence.
Ronak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd Harchaoui
AISTATS2
2022 Differentiable Programming A La Moreau
abstract
The notion of a Moreau envelope is central to the analysis of first-order optimization algorithms for machine learning and signal processing. We define a compositional calculus adapted to Moreau envelopes and show how to apply it to deep networks, and, more broadly, to learning systems equipped with automatic differentiation and implemented in the spirit of differentiable programming.
Vincent Roulet, Zaïd Harchaoui
ICASSP1
2019 Iterative Linearized Control: Stable Algorithms and Complexity Guarantees
abstract
We examine popular gradient-based algorithms for nonlinear control in the light of the modern complexity analysis of first-order optimization algorithms. The examination reveals that the complexity bounds can be clearly stated in terms of calls to a computational oracle related to dynamic programming and implementable by gradient back-propagation using machine learning software libraries such as PyTorch or TensorFlow. Finally, we propose a regularized Gauss-Newton algorithm enjoying worst-case complexity bounds and improved convergence behavior in practice. The software library based on PyTorch is publicly available.
Vincent Roulet, Dmitriy Drusvyatskiy, Siddhartha S. Srinivasa, Zaïd Harchaoui
ICML1
2018 A Smoother Way to Train Structured Prediction Models
abstract
We present a framework to train a structured prediction model by performing smoothing on the inference algorithm it builds upon. Smoothing overcomes the non-smoothness inherent to the maximum margin structured prediction objective, and paves the way for the use of fast primal gradient-based optimization algorithms. We illustrate the proposed framework by developing a novel primal incremental optimization algorithm for the structural support vector machine. The proposed algorithm blends an extrapolation scheme for acceleration and an adaptive smoothing scheme and builds upon the stochastic variance-reduced gradient algorithm. We establish its worst-case global complexity bound and study several practical variants. We present experimental results on two real-world problems, namely named entity recognition and visual object localization. The experimental results show that the proposed framework allows us to build upon efficient inference algorithms to develop large-scale optimization algorithms for structured prediction which can achieve competitive performance on the two real-world problems.
Venkata Pillutla, Vincent Roulet, Sham M. Kakade, Zaïd Harchaoui
NeurIPS2
2017 Sharpness, Restart and Acceleration
abstract
The {\L}ojasiewicz inequality shows that H\"olderian error bounds on the minimum of convex optimization problems hold almost generically. Here, we clarify results of \citet{Nemi85} who show that H\"olderian error bounds directly controls the performance of restart schemes. The constants quantifying error bounds are of course unobservable, but we show that optimal restart strategies are robust, and searching for the best scheme only increases the complexity by a logarithmic factor compared to the optimal bound. Overall then, restart schemes generically accelerate accelerated methods.
Vincent Roulet, Alexandre d'Aspremont
NIPS1
2017 Integration Methods and Optimization Algorithms
abstract
We show that accelerated optimization methods can be seen as particular instances of multi-step integration schemes from numerical analysis, applied to the gradient flow equation. Compared with recent advances in this vein, the differential equation considered here is the basic gradient flow, and we derive a class of multi-step schemes which includes accelerated algorithms, using classical conditions from numerical analysis. Multi-step schemes integrate the differential equation using larger step sizes, which intuitively explains the acceleration phenomenon.
Damien Scieur, Vincent Roulet, Francis R. Bach, Alexandre d'Aspremont
NIPS2