EDBT 2026 Demo / reviewers in the wild / expert
Javier Gonzalvo
dblp:238/0236
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Representation and self-supervised learning · 37% Efficient and distributed learning · 18% Deep learning architectures and training · 18% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
neural collapse |
0.8 | 1 | 2024 | The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 1 | 2024 | The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning
representation geometry |
0.8 | 1 | 2024 | The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024 |
Machine learning › Learning paradigms
multi-objective learning |
0.4 | 1 | 2020 | Agnostic Learning with Multiple Objectives · NeurIPS 2020 |
Machine learning › Learning theory › generalization bounds
rademacher complexity |
0.4 | 1 | 2020 | Agnostic Learning with Multiple Objectives · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.2 | 1 | 2024 | The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
pre-trained initialization · 0.8network growing · 0.8loss surface analysis · 0.8backward error analysis · 0.8symbolic gradient computation · 0.4convex optimization · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Deep Fusion: Efficient Network Training via Pre-trained InitializationsabstractTraining deep neural networks for large language models (LLMs) remains computationally very expensive. To mitigate this, network growing algorithms offer potential cost savings, but their underlying mechanisms are poorly understood. In this paper, we propose a theoretical framework using backward error analysis to illuminate the dynamics of mid-training network growth. Furthermore, we introduce Deep Fusion, an efficient network training approach that leverages pre-trained initializations of smaller networks, facilitating network growth from diverse sources. Our experiments validate the power of our theoretical framework in guiding the optimal use of Deep Fusion. With carefully optimized training dynamics, Deep Fusion demonstrates significant reductions in both training time and resource consumption. Importantly, these gains are achieved without sacrificing performance. We demonstrate reduced computational requirements, and improved generalization performance on a variety of NLP tasks and T5 model sizes. Hanna Mazzawi, Javier Gonzalvo, Michael Wunder, Sammy Jerome, Benoit Dherin |
ICML | 2 |
| 2024 | The Impact of Geometric Complexity on Neural Collapse in Transfer LearningabstractMany of the recent advances in computer vision and language models can be attributed to the success of transfer learning via the pre-training of large foundation models. However, a theoretical framework which explains this empirical success is incomplete and remains an active area of research. Flatness of the loss surface and neural collapse have recently emerged as useful pre-training metrics which shed light on the implicit biases underlying pre-training. In this paper, we explore the geometric complexity of a model's learned representations as a fundamental mechanism that relates these two concepts. We show through experiments and theory that mechanisms which affect the geometric complexity of the pre-trained network also influence the neural collapse. Furthermore, we show how this effect of the geometric complexity generalizes to the neural collapse of new classes as well, thus encouraging better performance on downstream tasks, particularly in the few-shot setting. Michael Munn, Benoit Dherin, Javier Gonzalvo |
NeurIPS | 3 |
| 2020 | Agnostic Learning with Multiple ObjectivesabstractMost machine learning tasks are inherently multi-objective. This means that the learner has to come up with a model that performs well across a number of base objectives $\cL_{1}, \ldots, \cL_{p}$, as opposed to a single one. Since optimizing with respect to multiple objectives at the same time is often computationally expensive, the base objectives are often combined in an ensemble $\sum_{k=1}^{p}\lambda_{k}\cL_{k}$, thereby reducing the problem to scalar optimization. The mixture weights $\lambda_{k}$ are set to uniform or some other fixed distribution, based on the learner's preferences. We argue that learning with a fixed distribution on the mixture weights runs the risk of overfitting to some individual objectives and significantly harming others, despite performing well on an entire ensemble. Moreover, in reality, the true preferences of a learner across multiple objectives are often unknown or hard to express as a specific distribution. Instead, we propose a new framework of \emph{Agnostic Learning with Multiple Objectives} ($\almo$), where a model is optimized for \emph{any} weights in the mixture of base objectives. We present data-dependent Rademacher complexity guarantees for learning in the $\almo$ framework, which are used to guide a scalable optimization algorithm and the corresponding regularization. We present convergence guarantees for this algorithm, assuming convexity of the loss functions and the underlying hypothesis space. We further implement the algorithm in a popular symbolic gradient computation framework and empirically demonstrate on a number of datasets the benefits of $\almo$ framework versus learning with a fixed mixture weights distribution. Corinna Cortes, Mehryar Mohri, Javier Gonzalvo, Dmitry Storcheus |
NeurIPS | 3 |