Clémentine C. J. Dominé

dblp:387/7804 · also Clémentine Carla Juliette Dominé · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Learning theory · 61% Deep learning architectures and training · 19% Learning paradigms · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
learning dynamics
2.332025
Learning dynamics in linear recurrent neural networks · ICML 2025
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks · ICLR 2025
Exact learning dynamics of deep linear networks with prior knowledge · NeurIPS 2022
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
1.422025
Learning dynamics in linear recurrent neural networks · ICML 2025
Exact learning dynamics of deep linear networks with prior knowledge · NeurIPS 2022
Machine learning › Learning paradigms
continual learning
1.022025
A Theory of Initialisation's Impact on Specialisation · ICLR 2025
Exact learning dynamics of deep linear networks with prior knowledge · NeurIPS 2022
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
Learning dynamics in linear recurrent neural networks · ICML 2025
Machine learning › Learning theory
neural network theory
0.812024
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning · NeurIPS 2024
Bioinformatics and computational biology › gene regulation
binding site prediction
0.812024
Geometric epitope and paratope prediction · Bioinform. 2024
Bioinformatics and computational biology › immunoinformatics
epitope prediction
0.812024
Geometric epitope and paratope prediction · Bioinform. 2024
Bioinformatics and computational biology › immunoinformatics
paratope prediction
0.812024
Geometric epitope and paratope prediction · Bioinform. 2024
Bioinformatics and computational biology
protein structure prediction
0.812024
Geometric epitope and paratope prediction · Bioinform. 2024
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks
0.612022
Exact learning dynamics of deep linear networks with prior knowledge · NeurIPS 2022
Machine learning › Deep learning architectures and training
training dynamics
0.212024
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

neural tangent kernel · 1.7specialization analysis · 0.9linear recurrent neural network · 0.9elastic weight consolidation · 0.9spectral geometric descriptors · 0.8geometric deep learning · 0.8exact solutions for linear networks · 0.8conserved quantities analysis · 0.8matrix riccati equation · 0.6
YearPublicationVenuePosition
2025 From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
abstract
Biological and artificial neural networks develop internal representations that enable them to perform complex tasks. In artificial networks, the effectiveness of these models relies on their ability to build task specific representation, a process influenced by interactions among datasets, architectures, initialization strategies, and optimization algorithms. Prior studies highlight that different initializations can place networks in either a lazy regime, where representations remain static, or a rich/feature learning regime, where representations evolve dynamically. Here, we examine how initialization influences learning dynamics in deep linear neural networks, deriving exact solutions for lambda-balanced initializations-defined by the relative scale of weights across layers. These solutions capture the evolution of representations and the Neural Tangent Kernel across the spectrum from the rich to the lazy regimes. Our findings deepen the theoretical understanding of the impact of weight initialization on learning regimes, with implications for continual learning, reversal learning, and transfer learning, relevant to both neuroscience and practical applications.
Clémentine C. J. Dominé, Nicolas Anguita, Alexandra Maria Proca, Lukas Braun, Daniel Kunin, Pedro A. M. Mediano, Andrew M. Saxe
ICLR1
2025 A Theory of Initialisation's Impact on Specialisation
abstract
Prior work has demonstrated a consistent tendency in neural networks engaged in continual learning tasks, wherein intermediate task similarity results in the highest levels of catastrophic interference. This phenomenon is attributed to the network's tendency to reuse learned features across tasks. However, this explanation heavily relies on the premise that neuron specialisation occurs, i.e. the emergence of localised representations. Our investigation challenges the validity of this assumption. Using theoretical frameworks for the analysis of neural networks, we show a strong dependence of specialisation on the initial condition. More precisely, we show that weight imbalance and high weight entropy can favour specialised solutions. We then apply these insights in the context of continual learning, first showing the emergence of a monotonic relation between task-similarity and forgetting in non-specialised networks. Finally, we show that specialization by weight imbalance is beneficial on the commonly employed elastic weight consolidation regularisation technique.
Devon Jarvis, Sebastian Lee, Clémentine C. J. Dominé, Andrew M. Saxe, Stefano Sarao Mannelli
ICLR3
2025 Learning dynamics in linear recurrent neural networks
abstract
Recurrent neural networks (RNNs) are powerful models used widely in both machine learning and neuroscience to learn tasks with temporal dependencies and to model neural dynamics. However, despite significant advancements in the theory of RNNs, there is still limited understanding of their learning process and the impact of the temporal structure of data. Here, we bridge this gap by analyzing the learning dynamics of linear RNNs (LRNNs) analytically, enabled by a novel framework that accounts for task dynamics. Our mathematical result reveals four key properties of LRNNs: (1) Learning of data singular values is ordered by both scale and temporal precedence, such that singular values that are larger and occur later are learned faster. (2) Task dynamics impact solution stability and extrapolation ability. (3) The loss function contains an effective regularization term that incentivizes small weights and mediates a tradeoff between recurrent and feedforward computation. (4) Recurrence encourages feature learning, as shown through a novel derivation of the neural tangent kernel for finite-width LRNNs. As a final proof-of-concept, we apply our theoretical framework to explain the behavior of LRNNs performing sensory integration tasks. Our work provides a first analytical treatment of the relationship between the temporal dependencies in tasks and learning dynamics in LRNNs, building a foundation for understanding how complex dynamic behavior emerges in cognitive models.
Alexandra Maria Proca, Clémentine C. J. Dominé, Murray Shanahan, Pedro A. M. Mediano
ICML2
2024 Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
abstract
While the impressive performance of modern neural networks is often attributed to their capacity to efficiently extract task-relevant features from data, the mechanisms underlying this *rich feature learning regime* remain elusive, with much of our theoretical understanding stemming from the opposing *lazy regime*. In this work, we derive exact solutions to a minimal model that transitions between lazy and rich learning, precisely elucidating how unbalanced *layer-specific* initialization variances and learning rates determine the degree of feature learning. Our analysis reveals that they conspire to influence the learning regime through a set of conserved quantities that constrain and modify the geometry of learning trajectories in parameter and function space. We extend our analysis to more complex linear models with multiple neurons, outputs, and layers and to shallow nonlinear networks with piecewise linear activation functions. In linear networks, rapid feature learning only occurs from balanced initializations, where all layers learn at similar speeds. While in nonlinear networks, unbalanced initializations that promote faster learning in earlier layers can accelerate rich learning. Through a series of experiments, we provide evidence that this unbalanced rich regime drives feature learning in deep finite-width networks, promotes interpretability of early layers in CNNs, reduces the sample complexity of learning hierarchical data, and decreases the time to grokking in modular arithmetic. Our theory motivates further exploration of unbalanced initializations to enhance efficient feature learning.
Daniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen 0046, David A. Klindt, Andrew M. Saxe, Surya Ganguli
NeurIPS3
2024 Geometric epitope and paratope prediction
abstract
MOTIVATION: Identifying the binding sites of antibodies is essential for developing vaccines and synthetic antibodies. In this article, we investigate the optimal representation for predicting the binding sites in the two molecules and emphasize the importance of geometric information. RESULTS: Specifically, we compare different geometric deep learning methods applied to proteins' inner (I-GEP) and outer (O-GEP) structures. We incorporate 3D coordinates and spectral geometric descriptors as input features to fully leverage the geometric information. Our research suggests that different geometrical representation information is useful for different tasks. Surface-based models are more efficient in predicting the binding of the epitope, while graph models are better in paratope prediction, both achieving significant performance improvements. Moreover, we analyze the impact of structural changes in antibodies and antigens resulting from conformational rearrangements or reconstruction errors. Through this investigation, we showcase the robustness of geometric deep learning methods and spectral geometric descriptors to such perturbations. AVAILABILITY AND IMPLEMENTATION: The python code for the models, together with the data and the processing pipeline, is open-source and available at https://github.com/Marco-Peg/GEP.
Marco Pegoraro 0002, Clémentine C. J. Dominé, Emanuele Rodolà, Petar Velickovic, Andreea Deac
Bioinform.2
2022 Exact learning dynamics of deep linear networks with prior knowledge
abstract
Learning in deep neural networks is known to depend critically on the knowledge embedded in the initial network weights. However, few theoretical results have precisely linked prior knowledge to learning dynamics. Here we derive exact solutions to the dynamics of learning with rich prior knowledge in deep linear networks by generalising Fukumizu's matrix Riccati solution \citep{fukumizu1998effect}. We obtain explicit expressions for the evolving network function, hidden representational similarity, and neural tangent kernel over training for a broad class of initialisations and tasks. The expressions reveal a class of task-independent initialisations that radically alter learning dynamics from slow non-linear dynamics to fast exponential trajectories while converging to a global optimum with identical representational similarity, dissociating learning trajectories from the structure of initial internal representations. We characterise how network weights dynamically align with task structure, rigorously justifying why previous solutions successfully described learning from small initial weights without incorporating their fine-scale structure. Finally, we discuss the implications of these findings for continual learning, reversal learning and learning of structured knowledge. Taken together, our results provide a mathematical toolkit for understanding the impact of prior knowledge on deep learning.
Lukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. Saxe
NeurIPS2