Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Javier Gonzalvo

dblp:238/0236 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 37% Efficient and distributed learning · 18% Deep learning architectures and training · 18%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
neural collapse
0.812024
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning
pre-training
0.812024
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning
representation geometry
0.812024
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024
Machine learning › Learning paradigms
multi-objective learning
0.412020
Agnostic Learning with Multiple Objectives · NeurIPS 2020
Machine learning › Learning theory › generalization bounds
rademacher complexity
0.412020
Agnostic Learning with Multiple Objectives · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.212024
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

pre-trained initialization · 0.8network growing · 0.8loss surface analysis · 0.8backward error analysis · 0.8symbolic gradient computation · 0.4convex optimization · 0.4
YearPublicationVenuePosition
2024 Deep Fusion: Efficient Network Training via Pre-trained Initializations
abstract
Training deep neural networks for large language models (LLMs) remains computationally very expensive. To mitigate this, network growing algorithms offer potential cost savings, but their underlying mechanisms are poorly understood. In this paper, we propose a theoretical framework using backward error analysis to illuminate the dynamics of mid-training network growth. Furthermore, we introduce Deep Fusion, an efficient network training approach that leverages pre-trained initializations of smaller networks, facilitating network growth from diverse sources. Our experiments validate the power of our theoretical framework in guiding the optimal use of Deep Fusion. With carefully optimized training dynamics, Deep Fusion demonstrates significant reductions in both training time and resource consumption. Importantly, these gains are achieved without sacrificing performance. We demonstrate reduced computational requirements, and improved generalization performance on a variety of NLP tasks and T5 model sizes.
Hanna Mazzawi, Javier Gonzalvo, Michael Wunder, Sammy Jerome, Benoit Dherin
ICML2
2024 The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
abstract
Many of the recent advances in computer vision and language models can be attributed to the success of transfer learning via the pre-training of large foundation models. However, a theoretical framework which explains this empirical success is incomplete and remains an active area of research. Flatness of the loss surface and neural collapse have recently emerged as useful pre-training metrics which shed light on the implicit biases underlying pre-training. In this paper, we explore the geometric complexity of a model's learned representations as a fundamental mechanism that relates these two concepts. We show through experiments and theory that mechanisms which affect the geometric complexity of the pre-trained network also influence the neural collapse. Furthermore, we show how this effect of the geometric complexity generalizes to the neural collapse of new classes as well, thus encouraging better performance on downstream tasks, particularly in the few-shot setting.
Michael Munn, Benoit Dherin, Javier Gonzalvo
NeurIPS3
2020 Agnostic Learning with Multiple Objectives
abstract
Most machine learning tasks are inherently multi-objective. This means that the learner has to come up with a model that performs well across a number of base objectives $\cL_{1}, \ldots, \cL_{p}$, as opposed to a single one. Since optimizing with respect to multiple objectives at the same time is often computationally expensive, the base objectives are often combined in an ensemble $\sum_{k=1}^{p}\lambda_{k}\cL_{k}$, thereby reducing the problem to scalar optimization. The mixture weights $\lambda_{k}$ are set to uniform or some other fixed distribution, based on the learner's preferences. We argue that learning with a fixed distribution on the mixture weights runs the risk of overfitting to some individual objectives and significantly harming others, despite performing well on an entire ensemble. Moreover, in reality, the true preferences of a learner across multiple objectives are often unknown or hard to express as a specific distribution. Instead, we propose a new framework of \emph{Agnostic Learning with Multiple Objectives} ($\almo$), where a model is optimized for \emph{any} weights in the mixture of base objectives. We present data-dependent Rademacher complexity guarantees for learning in the $\almo$ framework, which are used to guide a scalable optimization algorithm and the corresponding regularization. We present convergence guarantees for this algorithm, assuming convexity of the loss functions and the underlying hypothesis space. We further implement the algorithm in a popular symbolic gradient computation framework and empirically demonstrate on a number of datasets the benefits of $\almo$ framework versus learning with a fixed mixture weights distribution.
Corinna Cortes, Mehryar Mohri, Javier Gonzalvo, Dmitry Storcheus
NeurIPS3