VLDB 2026 Research / reviewers in the wild / expert
Vincent Roulet
dblp:164/6165
· DBLP profile ↗
11ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Deep learning architectures and training · 22% Trustworthy machine learning · 19% Optimization for machine learning · 19% | |
| Theoretical computer science
5 papers |
Mathematical optimization · 98% Algorithms and data structures · 2% |
Topics — the 29 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization
continuous optimization |
1.2 | 2 | 2025 | Loss Functions and Operators Generated by f-Divergences · ICML 2025 Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019 |
Mathematical optimization › continuous optimization
convex optimization |
1.2 | 2 | 2025 | Loss Functions and Operators Generated by f-Divergences · ICML 2025 Sharpness, Restart and Acceleration · NIPS 2017 |
Machine learning › Optimization for machine learning
convergence analysis |
0.9 | 1 | 2025 | On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control · J. Mach. Learn. Res. 2025 |
Machine learning › Generative modeling
energy-based model |
0.9 | 1 | 2025 | Joint Learning of Energy-based Models and their Partition Function · ICML 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | Loss Functions and Operators Generated by f-Divergences · ICML 2025 |
Machine learning › Deep learning architectures and training
loss function design |
0.9 | 1 | 2025 | Loss Functions and Operators Generated by f-Divergences · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
partition function estimation |
0.9 | 1 | 2025 | Joint Learning of Energy-based Models and their Partition Function · ICML 2025 |
Robotics › Motion planning and robot control
trajectory optimization |
0.9 | 1 | 2025 | On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control · J. Mach. Learn. Res. 2025 |
Mathematical optimization › continuous optimization
nonlinear optimization |
0.9 | 1 | 2025 | On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control · J. Mach. Learn. Res. 2025 |
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization |
0.8 | 1 | 2024 | Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.8 | 1 | 2024 | Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024 |
Machine learning › Deep learning architectures and training › training optimization
learning rate tuning |
0.8 | 1 | 2024 | Stepping on the Edge: Curvature Aware Learning Rate Tuners · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.8 | 1 | 2024 | Distributionally Robust Optimization with Bias and Variance Reduction · ICLR 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
0.8 | 1 | 2024 | Stepping on the Edge: Curvature Aware Learning Rate Tuners · NeurIPS 2024 |
Mathematical optimization › least squares
gauss-newton method |
0.4 | 1 | 2019 | Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods
gradient-based optimization |
0.4 | 1 | 2019 | Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.3 | 1 | 2018 | A Smoother Way to Train Structured Prediction Models · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured SVM |
0.3 | 1 | 2018 | A Smoother Way to Train Structured Prediction Models · NeurIPS 2018 |
Machine learning › Optimization for machine learning › variance reduction
SVRG |
0.3 | 1 | 2018 | A Smoother Way to Train Structured Prediction Models · NeurIPS 2018 |
Machine learning › Optimization for machine learning
convergence acceleration |
0.3 | 1 | 2017 | Integration Methods and Optimization Algorithms · NIPS 2017 |
Mathematical optimization
convergence analysis |
0.3 | 1 | 2017 | Sharpness, Restart and Acceleration · NIPS 2017 |
Mathematical optimization
restart strategies |
0.3 | 1 | 2017 | Sharpness, Restart and Acceleration · NIPS 2017 |
Machine learning › Deep learning architectures and training › loss function design
fenchel-young losses |
0.3 | 1 | 2025 | Joint Learning of Energy-based Models and their Partition Function · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
0.3 | 1 | 2025 | Joint Learning of Energy-based Models and their Partition Function · ICML 2025 |
Algorithms and data structures
dynamic programming |
0.1 | 1 | 2019 | Iterative Linearized Control: Stable Algorithms and Complexity Guarantees · ICML 2019 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2018 | A Smoother Way to Train Structured Prediction Models · NeurIPS 2018 |
Computer vision › Image recognition and object detection
object localization |
0.1 | 1 | 2018 | A Smoother Way to Train Structured Prediction Models · NeurIPS 2018 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
gradient flow |
0.1 | 1 | 2017 | Integration Methods and Optimization Algorithms · NIPS 2017 |
Methods — techniques the papers use, named apart from their topics
tsallis α-negentropy · 1.7stochastic approximation · 1.7gauss-newton method · 1.7f-divergence · 1.7bisection algorithm · 1.7stochastic gradient descent · 1.6min-min formulation · 0.9variance reduction · 0.8hessian eigenvalue analysis · 0.8bias reduction · 0.8gradient backpropagation · 0.4gauss-newton regularization · 0.4łojasiewicz inequality · 0.3multi-step integration schemes · 0.3acceleration · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Loss Functions and Operators Generated by f-DivergencesabstractThe logistic loss (a.k.a. cross-entropy loss) is one of the most popular loss functions used for multiclass classification. It is also the loss function of choice for next-token prediction in language modeling. It is associated with the Kullback-Leibler (KL) divergence and the softargmax operator. In this work, we propose to construct new convex loss functions based on $f$-divergences. Our loss functions generalize the logistic loss in two directions: i) by replacing the KL divergence with $f$-divergences and ii) by allowing non-uniform reference measures. We instantiate our framework for numerous $f$-divergences, recovering existing losses and creating new ones. By analogy with the logistic loss, the loss function generated by an $f$-divergence is associated with an operator, that we dub $f$-softargmax. We derive a novel parallelizable bisection algorithm for computing the $f$-softargmax associated with any $f$-divergence. On the empirical side, one of the goals of this paper is to determine the effectiveness of loss functions beyond the classical cross-entropy in a language model setting, including on pre-training, post-training (SFT) and distillation. We show that the loss function generated by the $\alpha$-divergence (which is equivalent to Tsallis $\alpha$-negentropy in the case of unit reference measures) with $\alpha=1.5$ performs well across several tasks. Vincent Roulet, Tianlin Liu, Nino Vieillard, Michael E. Sander, Mathieu Blondel |
ICML | 1 |
| 2025 | Joint Learning of Energy-based Models and their Partition FunctionabstractEnergy-based models (EBMs) offer a flexible framework for parameterizing probability distributions using neural networks.
However, learning EBMs by exact maximum likelihood estimation (MLE) is generally intractable, due to the need to compute the partition function.
In this paper, we propose a novel min-min formulation for approximately learning probabilistic EBMs in combinatorially-large discrete spaces, such as sets or permutations. Our key idea is to jointly learn both an energy model and its log-partition, parameterized as a neural network. Our approach not only provides a novel tractable objective criterion to learn EBMs by stochastic gradient descent (without relying on MCMC), but also a novel means to estimate the log-partition function on unseen data points.
On the theoretical side, we show that our approach recovers the optimal MLE solution when optimizing in the space of continuous functions.
Furthermore, we show that our approach naturally extends to the broader family of Fenchel-Young losses, allowing us to obtain
the first tractable method for optimizing the sparsemax loss in combinatorially-large spaces.
We demonstrate our approach on multilabel classification and label ranking. Michael E. Sander, Vincent Roulet, Tianlin Liu, Mathieu Blondel |
ICML | 2 |
| 2025 | On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear ControlabstractA classical approach for solving discrete time nonlinear control on a finite horizon consists in repeatedly minimizing linear quadratic approximations of the original problem around current candidate solutions. While widely popular in many domains, such an approach has mainly been analyzed locally. We provide detailed convergence guarantees to stationary points as well as local linear convergence rates for the Iterative Linear Quadratic Regulator (ILQR) algorithm and its Differential Dynamic Programming (DDP) variant. For problems without costs on control variables, we observe that global convergence to minima can be ensured provided that the linearized discrete time dynamics are surjective, costs on the state variables are gradient dominated. We further detail quadratic local convergence when the costs are self-concordant. We show that surjectivity of the linearized dynamics hold for appropriate discretization schemes given the existence of a feedback linearization scheme. We present complexity bounds of algorithms based on linear quadratic approximations through the lens of generalized Gauss-Newton methods. Our analysis uncovers several convergence phases for regularized generalized Gauss-Newton algorithms. Vincent Roulet, Siddhartha S. Srinivasa, Maryam Fazel, Zaïd Harchaoui |
J. Mach. Learn. Res. | 1 |
| 2024 | Distributionally Robust Optimization with Bias and Variance ReductionabstractWe consider the distributionally robust optimization (DRO) problem, wherein a learner optimizes the worst-case empirical risk achievable by reweighing the observed training examples. We present Prospect, a stochastic gradient-based algorithm that only requires tuning a single learning rate hyperparameter, and prove that it enjoys linear convergence for smooth regularized losses. This contrasts with previous algorithms that either require tuning multiple hyperparameters or potentially fail to converge due to biased gradient estimates or inadequate regularization. Empirically, we show that Prospect can converge 2-3x faster than baselines such as SGD and stochastic saddle-point methods on distribution shift and fairness benchmarks spanning tabular, vision, and language domains. Ronak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd Harchaoui |
ICLR | 2 |
| 2024 | Stepping on the Edge: Curvature Aware Learning Rate TunersabstractCurvature information -- particularly, the largest eigenvalue of the loss
Hessian, known as the sharpness -- often forms the basis for learning rate
tuners. However, recent work has shown that the curvature information undergoes
complex dynamics during training, going from a phase of increasing sharpness to
eventual stabilization. We analyze the closed-loop feedback effect between
learning rate tuning and curvature. We find that classical learning rate tuners
may yield greater one-step loss reduction, yet they ultimately underperform in
the long term when compared to constant learning rates in the full batch regime.
These models break the stabilization of the sharpness, which we explain using a
simplified model of the joint dynamics of the learning rate and the curvature.
To further investigate these effects, we introduce a new learning rate tuning
method, Curvature Dynamics Aware Tuning (CDAT), which prioritizes long term
curvature stabilization over instantaneous progress on the objective. In the
full batch regime, CDAT shows behavior akin to prefixed warm-up schedules on deep
learning objectives, outperforming tuned constant learning rates. In the mini
batch regime, we observe that stochasticity introduces confounding effects that
explain the previous success of some learning rate tuners at appropriate batch
sizes. Our findings highlight the critical role of understanding the joint
dynamics of the learning rate and curvature, beyond greedy minimization, to
diagnose failures and design effective adaptive learning rate tuners. Vincent Roulet, Atish Agarwala, Jean-Bastien Grill, Grzegorz Swirszcz, Mathieu Blondel, Fabian Pedregosa |
NeurIPS | 1 |
| 2023 | Stochastic Optimization for Spectral Risk MeasuresabstractSpectral risk objectives – also called L-risks – allow for learning systems to interpolate between optimizing average-case performance (as in empirical risk minimization) and worst-case performance on a task. We develop LSVRG, a stochastic algorithm to optimize these quantities by characterizing their subdifferential and addressing challenges such as biasedness of subgradient estimates and non-smoothness of the objective. We show theoretically and experimentally that out-of-the-box approaches such as stochastic subgradient and dual averaging can be hindered by bias, whereas our approach exhibits linear convergence. Ronak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd Harchaoui |
AISTATS | 2 |
| 2022 | Differentiable Programming A La MoreauabstractThe notion of a Moreau envelope is central to the analysis of first-order optimization algorithms for machine learning and signal processing. We define a compositional calculus adapted to Moreau envelopes and show how to apply it to deep networks, and, more broadly, to learning systems equipped with automatic differentiation and implemented in the spirit of differentiable programming. Vincent Roulet, Zaïd Harchaoui |
ICASSP | 1 |
| 2019 | Iterative Linearized Control: Stable Algorithms and Complexity GuaranteesabstractWe examine popular gradient-based algorithms for nonlinear control in the light of the modern complexity analysis of first-order optimization algorithms. The examination reveals that the complexity bounds can be clearly stated in terms of calls to a computational oracle related to dynamic programming and implementable by gradient back-propagation using machine learning software libraries such as PyTorch or TensorFlow. Finally, we propose a regularized Gauss-Newton algorithm enjoying worst-case complexity bounds and improved convergence behavior in practice. The software library based on PyTorch is publicly available. Vincent Roulet, Dmitriy Drusvyatskiy, Siddhartha S. Srinivasa, Zaïd Harchaoui |
ICML | 1 |
| 2018 | A Smoother Way to Train Structured Prediction ModelsabstractWe present a framework to train a structured prediction model by performing smoothing on the inference algorithm it builds upon. Smoothing overcomes the non-smoothness inherent to the maximum margin structured prediction objective, and paves the way for the use of fast primal gradient-based optimization algorithms. We illustrate the proposed framework by developing a novel primal incremental optimization algorithm for the structural support vector machine. The proposed algorithm blends an extrapolation scheme for acceleration and an adaptive smoothing scheme and builds upon the stochastic variance-reduced gradient algorithm. We establish its worst-case global complexity bound and study several practical variants. We present experimental results on two real-world problems, namely named entity recognition and visual object localization. The experimental results show that the proposed framework allows us to build upon efficient inference algorithms to develop large-scale optimization algorithms for structured prediction which can achieve competitive performance on the two real-world problems. Venkata Pillutla, Vincent Roulet, Sham M. Kakade, Zaïd Harchaoui |
NeurIPS | 2 |
| 2017 | Sharpness, Restart and AccelerationabstractThe {\L}ojasiewicz inequality shows that H\"olderian error bounds on the minimum of convex optimization problems hold almost generically. Here, we clarify results of \citet{Nemi85} who show that H\"olderian error bounds directly controls the performance of restart schemes. The constants quantifying error bounds are of course unobservable, but we show that optimal restart strategies are robust, and searching for the best scheme only increases the complexity by a logarithmic factor compared to the optimal bound. Overall then, restart schemes generically accelerate accelerated methods. Vincent Roulet, Alexandre d'Aspremont |
NIPS | 1 |
| 2017 | Integration Methods and Optimization AlgorithmsabstractWe show that accelerated optimization methods can be seen as particular instances of multi-step integration schemes from numerical analysis, applied to the gradient flow equation. Compared with recent advances in this vein, the differential equation considered here is the basic gradient flow, and we derive a class of multi-step schemes which includes accelerated algorithms, using classical conditions from numerical analysis. Multi-step schemes integrate the differential equation using larger step sizes, which intuitively explains the acceleration phenomenon. Damien Scieur, Vincent Roulet, Francis R. Bach, Alexandre d'Aspremont |
NIPS | 2 |