Pascal Tikeng Notsawo Jr.

dblp:350/0550 · also Pascal Notsawo · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Representation and self-supervised learning · 32% Deep learning architectures and training · 18% Learning theory · 18%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
generalization
0.912025
Grokking Beyond the Euclidean Norm of Model Parameters · ICML 2025
Machine learning › Deep learning architectures and training › training dynamics
grokking
0.912025
Grokking Beyond the Euclidean Norm of Model Parameters · ICML 2025
Machine learning › Optimization for machine learning
implicit regularization
0.912025
Grokking Beyond the Euclidean Norm of Model Parameters · ICML 2025
Machine learning › Representation and self-supervised learning › vector quantization
dynamic vector quantization
0.712023
Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational Coarseness · AAAI 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication
0.712023
Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational Coarseness · AAAI 2023
Machine learning › Representation and self-supervised learning
vector quantization
0.712023
Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational Coarseness · AAAI 2023
Machine learning › Representation and self-supervised learning › representation learning
discrete representation learning
0.212023
Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational Coarseness · AAAI 2023

Methods — techniques the papers use, named apart from their topics

weight decay · 0.9regularization · 0.9gradient descent · 0.9vector quantization · 0.7communication bottleneck · 0.7adaptive discretization · 0.7
YearPublicationVenuePosition
2025 Grokking Beyond the Euclidean Norm of Model Parameters
abstract
Grokking refers to a delayed generalization following overfitting when optimizing artificial neural networks with gradient-based methods. In this work, we demonstrate that grokking can be induced by regularization, either explicit or implicit. More precisely, we show that when there exists a model with a property $P$ (e.g., sparse or low-rank weights) that generalizes on the problem of interest, gradient descent with a small but non-zero regularization of $P$ (e.g., $\ell_1$ or nuclear norm regularization) results in grokking. This extends previous work showing that small non-zero weight decay induces grokking. Moreover, our analysis shows that over-parameterization by adding depth makes it possible to grok or ungrok without explicitly using regularization, which is impossible in shallow cases. We further show that the $\ell_2$ norm is not a reliable proxy for generalization when the model is regularized toward a different property $P$, as the $\ell_2$ norm grows in many cases where no weight decay is used, but the model generalizes anyway. We also show that grokking can be amplified solely through data selection, with any other hyperparameter fixed.
Pascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume Rabusseau
ICML1
2023 Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational Coarseness
abstract
Vector Quantization (VQ) is a method for discretizing latent representations and has become a major part of the deep learning toolkit. It has been theoretically and empirically shown that discretization of representations leads to improved generalization, including in reinforcement learning where discretization can be used to bottleneck multi-agent communication to promote agent specialization and robustness. The discretization tightness of most VQ-based methods is defined by the number of discrete codes in the representation vector and the codebook size, which are fixed as hyperparameters. In this work, we propose learning to dynamically select discretization tightness conditioned on inputs, based on the hypothesis that data naturally contains variations in complexity that call for different levels of representational coarseness which is observed in many heterogeneous data sets. We show that dynamically varying tightness in communication bottlenecks can improve model performance on visual reasoning and reinforcement learning tasks with heterogeneity in representations.
Dianbo Liu, Alex Lamb, Pascal Tikeng Notsawo Jr., Michael C. Mozer, Yoshua Bengio, Kenji Kawaguchi
AAAI4