Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Alessandro Cabodi

dblp:410/6112 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › loss landscape › mode connectivity
linear mode connectivity
0.912025
Generalized Linear Mode Connectivity for Transformers · NeurIPS 2025
Machine learning › Deep learning architectures and training
loss landscape
0.912025
Generalized Linear Mode Connectivity for Transformers · NeurIPS 2025
Machine learning › Deep learning architectures and training › symmetry-aware learning
permutation invariance
0.912025
Generalized Linear Mode Connectivity for Transformers · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
Generalized Linear Mode Connectivity for Transformers · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

orthogonal transformation · 0.9neuron reordering · 0.9invertible maps · 0.9
YearPublicationVenuePosition
2025 Generalized Linear Mode Connectivity for Transformers
abstract
Understanding the geometry of neural network loss landscapes is a central question in deep learning, with implications for generalization and optimization. A striking phenomenon is $\textit{linear mode connectivity}$ (LMC), where independently trained models can be connected by low- or zero-barrier paths, despite appearing to lie in separate loss basins. However, this is often obscured by symmetries in parameter space—such as neuron permutations—which make functionally equivalent models appear dissimilar. Prior work has predominantly focused on neuron reordering through permutations, but such approaches are limited in scope and fail to capture the richer symmetries exhibited by modern architectures such as Transformers. In this work, we introduce a unified framework that captures four symmetry classes—permutations, semi-permutations, orthogonal transformations, and general invertible maps—broadening the set of valid reparameterizations and subsuming many previous approaches as special cases. Crucially, this generalization enables, for the first time, the discovery of low- and zero-barrier linear interpolation paths between independently trained Vision Transformers and GPT-2 models. Furthermore, our framework extends beyond pairwise alignment, to multi-model and width-heterogeneous settings, enabling alignment across architectures of different sizes. These results reveal deeper structure in the loss landscape and underscore the importance of symmetry-aware analysis for understanding model space geometry.
Alexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto, Sidak Pal Singh, Valentina Boeva
NeurIPS2