Christian H. X. Ali Mehmeti-Göpel

dblp:295/5427 · also Christian Heinrich Xhemal Ali Mehmeti-Göpel · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 43% Optimization for machine learning · 24% Learning theory · 11%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
learning rate schedule
0.812024
On the Weight Dynamics of Deep Normalized Networks · ICML 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
On the Weight Dynamics of Deep Normalized Networks · ICML 2024
Machine learning › Optimization for machine learning › learning rate schedule
warm-up
0.812024
On the Weight Dynamics of Deep Normalized Networks · ICML 2024
Natural language and speech › Language models and text generation › text generation › surface realization
linearization
0.712023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023
Machine learning › Efficient and distributed learning › model compression › sparsity
network sparsification
0.712023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023
Machine learning › Deep learning architectures and training
neural network expressivity
0.712023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023
Machine learning › Deep learning architectures and training
activation function
0.512021
Ringing ReLUs: Harmonic Distortion Analysis of Nonlinear Feedforward Networks · ICLR 2021
Machine learning › Learning theory › neural network theory
neural network analysis
0.512021
Ringing ReLUs: Harmonic Distortion Analysis of Nonlinear Feedforward Networks · ICLR 2021
Machine learning › Deep learning architectures and training › activation function
ReLU
0.512021
Ringing ReLUs: Harmonic Distortion Analysis of Nonlinear Feedforward Networks · ICLR 2021
Machine learning › Deep learning architectures and training › normalization
normalization layers
0.212024
On the Weight Dynamics of Deep Normalized Networks · ICML 2024
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.212023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023

Methods — techniques the papers use, named apart from their topics

weight dynamics analysis · 0.8warm-up method · 0.8sparsity prior · 0.7average path length measure · 0.7
YearPublicationVenuePosition
2024 On the Weight Dynamics of Deep Normalized Networks
abstract
Recent studies have shown that high disparities in effective learning rates (ELRs) across layers in deep neural networks can negatively affect trainability. We formalize how these disparities evolve over time by modeling weight dynamics (evolution of expected gradient and weight norms) of networks with normalization layers, predicting the evolution of layer-wise ELR ratios. We prove that when training with any constant learning rate, ELR ratios converge to 1, despite initial gradient explosion. We identify a "critical learning rate" beyond which ELR disparities widen, which only depends on current ELRs. To validate our findings, we devise a hyper-parameter-free warm-up method that successfully minimizes ELR spread quickly in theory and practice. Our experiments link ELR spread with trainability, a relationship that is most evident in very deep networks with significant gradient magnitude excursions.
Christian H. X. Ali Mehmeti-Göpel, Michael Wand 0001
ICML1
2023 Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think
abstract
We perform an empirical study of the behaviour of deep networks when fully linearizing some of its feature channels through a sparsity prior on the overall number of nonlinear units in the network. In experiments on image classification and machine translation tasks, we investigate how much we can simplify the network function towards linearity before performance collapses. First, we observe a significant performance gap when reducing nonlinearity in the network function early on as opposed to late in training, in-line with recent observations on the time-evolution of the data-dependent NTK. Second, we find that after training, we are able to linearize a significant number of nonlinear units while maintaining a high performance, indicating that much of a network's expressivity remains unused but helps gradient descent in early stages of training. To characterize the depth of the resulting partially linearized network, we introduce a measure called average path length, representing the average number of active nonlinearities encountered along a path in the network graph. Under sparsity pressure, we find that the remaining nonlinear units organize into distinct structures, forming core-networks of near constant effective depth and width, which in turn depend on task difficulty.
Christian H. X. Ali Mehmeti-Göpel, Jan Disselhoff
ICML1
2021 Ringing ReLUs: Harmonic Distortion Analysis of Nonlinear Feedforward Networks
Christian H. X. Ali Mehmeti-Göpel, David Hartmann, Michael Wand 0001
ICLR1