VLDB 2026 Research / reviewers in the wild / expert
Damien Ferbach
dblp:322/0995
· DBLP profile ↗
4ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 34% Optimization for machine learning · 28% Deep learning architectures and training · 14% | |
| Theoretical computer science
1 paper |
Computational complexity · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 |
Machine learning › Optimization for machine learning › stochastic gradient descent
stochastic gradient descent with momentum |
0.9 | 1 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.9 | 1 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification |
0.8 | 1 | 2024 | Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences · NeurIPS 2024 |
Machine learning › Generative modeling
self-consuming generative model |
0.8 | 1 | 2024 | Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › data curation
training data curation |
0.8 | 1 | 2024 | Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis |
0.7 | 1 | 2023 | A General Framework For Proving The Equivariant Strong Lottery Ticket Hypothesis · ICLR 2023 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.7 | 1 | 2023 | A General Framework For Proving The Equivariant Strong Lottery Ticket Hypothesis · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
proof framework · 1.3equivariance · 1.3stochastic gradient descent · 0.9nesterov acceleration · 0.9reward model · 0.8preference optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dimension-adapted Momentum Outscales SGDabstractWe investigate scaling laws for stochastic momentum algorithms on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying data-target complexities. While traditional stochastic gradient descent with momentum (SGD-M) yields identical scaling law exponents to SGD, dimension-adapted Nesterov acceleration (DANA) improves these exponents by scaling momentum hyperparameters based on model size and data complexity. This outscaling phenomenon, which also improves compute-optimal scaling behavior, is achieved by DANA across a broad range of data and target complexities, while traditional methods fall short. Extensive experiments on high-dimensional synthetic quadratics validate our theoretical predictions and large-scale text experiments with LSTMs show DANA's improved loss exponents over SGD hold in a practical setting. Damien Ferbach, Katie Everett, Gauthier Gidel, Elliot Paquette, Courtney Paquette |
NeurIPS | 1 |
| 2024 | Proving Linear Mode Connectivity of Neural Networks via Optimal TransportabstractThe energy landscape of high-dimensional non-convex optimization problems is crucial to understanding the effectiveness of modern deep neural network architectures. Recent works have experimentally shown that two different solutions found after two runs of a stochastic training are often connected by very simple continuous paths (e.g., linear) modulo a permutation of the weights. In this paper, we provide a framework theoretically explaining this empirical observation. Based on convergence rates in Wasserstein distance of empirical measures, we show that, with high probability, two wide enough two-layer neural networks trained with stochastic gradient descent are linearly connected. Additionally, we express upper and lower bounds on the width of each layer of two deep neural networks with independent neuron weights to be linearly connected. Finally, we empirically demonstrate the validity of our approach by showing how the dimension of the support of the weight distribution of neurons, which dictates Wasserstein convergence rates is correlated with linear mode connectivity. Damien Ferbach, Baptiste Goujaud, Gauthier Gidel, Aymeric Dieuleveut |
AISTATS | 1 |
| 2024 | Self-Consuming Generative Models with Curated Data Provably Optimize Human PreferencesabstractThe rapid progress in generative models has resulted in impressive leaps in generation quality, blurring the lines between synthetic and real data. Web-scale datasets are now prone to the inevitable contamination by synthetic data, directly impacting the training of future generated models.
Already, some theoretical results on self-consuming generative models (a.k.a., iterative retraining) have emerged in the literature, showcasing that either model collapse or stability could be possible depending on the fraction of generated data used at each retraining step.
However, in practice, synthetic data is often subject to human feedback and curated by users before being used and uploaded online. For instance, many interfaces of popular text-to-image generative models, such as Stable Diffusion or Midjourney, produce several variations of an image for a given query which can eventually be curated by the users.
In this paper, we theoretically study the impact of data curation on iterated retraining of generative models and show that it can be seen as an implicit preference optimization mechanism. However, unlike standard preference optimization, the generative model does not have access to the reward function or negative samples needed for pairwise comparisons. Moreover, our study doesn't require access to the density function, only to samples. We prove that, if the data is curated according to a reward model, then the expected reward of the iterative retraining procedure is maximized. We further provide theoretical results on the stability of the retraining loop when using a positive fraction of real data at each step. Finally, we conduct illustrative experiments on both synthetic datasets and on CIFAR10 showing that such a procedure amplifies biases of the reward model. Damien Ferbach, Quentin Bertrand, Joey Bose, Gauthier Gidel |
NeurIPS | 1 |
| 2023 | A General Framework For Proving The Equivariant Strong Lottery Ticket Hypothesis
Damien Ferbach, Christos Tsirigotis, Gauthier Gidel, Joey Bose |
ICLR | 1 |