Noam Itzhak Levi

dblp:391/7481 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Learning theory · 29% Deep learning architectures and training · 24% Language models and text generation · 15%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › training dynamics
grokking
1.622025
Grokking at the Edge of Linear Separability · ICML 2025
Grokking in Linear Estimators - A Solvable Model that Groks without Understanding · ICLR 2024
Machine learning › Generative modeling
diffusion model
0.912025
Probing the Latent Hierarchical Structure of Data via Diffusion Models · ICLR 2025
Machine learning › Learning theory › implicit bias
implicit bias of gradient descent
0.912025
Grokking at the Edge of Linear Separability · ICML 2025
Natural language and speech › Language models and text generation › large language model inference
inference scaling laws
0.912025
A Simple Model of Inference Scaling Laws · ICML 2025
Machine learning › Trustworthy machine learning
machine unlearning
0.912025
Ascent Fails to Forget · NeurIPS 2025
Machine learning › Trustworthy machine learning
privacy and data protection
0.912025
Ascent Fails to Forget · NeurIPS 2025
Machine learning › Learning theory
random matrix theory
0.912025
The Underlying Universal Statistical Structure of Natural Datasets · ICML 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
A Simple Model of Inference Scaling Laws · ICML 2025
Machine learning › Learning theory › statistical learning theory
statistical physics of learning
0.912025
The Underlying Universal Statistical Structure of Natural Datasets · ICML 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
A Simple Model of Inference Scaling Laws · ICML 2025
Machine learning › Learning theory
generalization
0.812024
Grokking in Linear Estimators - A Solvable Model that Groks without Understanding · ICLR 2024
Machine learning › Representation and self-supervised learning
hierarchical representation
0.312025
Probing the Latent Hierarchical Structure of Data via Diffusion Models · ICLR 2025
Machine learning › Deep learning architectures and training
training dynamics
0.212024
Grokking in Linear Estimators - A Solvable Model that Groks without Understanding · ICLR 2024

Methods — techniques the papers use, named apart from their topics

statistical ansatz · 0.9shannon entropy · 0.9random matrix theory · 0.9pass@k metric · 0.9logistic regression analysis · 0.9logistic regression · 0.9gradient descent ascent · 0.9gradient ascent · 0.9eigenvalue analysis · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2025 Probing the Latent Hierarchical Structure of Data via Diffusion Models
abstract
High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce. Likewise, accessing the latent variables underlying such a data structure remains a challenge. In this work, we show that forward-backward experiments in diffusion-based models, where data is noised and then denoised to generate new samples, are a promising tool to probe the latent structure of data. We predict in simple hierarchical models that, in this process, changes in data occur by correlated chunks, with a length scale that diverges at a noise level where a phase transition is known to take place. Remarkably, we confirm this prediction in both text and image datasets using state-of-the-art diffusion models. Our results show how latent variable changes manifest in the data and establish how to measure these effects in real data using diffusion models.
Antonio Sclocchi, Alessandro Favero, Noam Itzhak Levi, Matthieu Wyart
ICLR3
2025 Grokking at the Edge of Linear Separability
abstract
We investigate the phenomenon of grokking -- delayed generalization accompanied by non-monotonic test loss behavior -- in a simple binary logistic classification task, for which "memorizing" and "generalizing" solutions can be strictly defined. Surprisingly, we find that grokking arises naturally even in this minimal model when the parameters of the problem are close to a critical point, and provide both empirical and analytical insights into its mechanism. Concretely, by appealing to the implicit bias of gradient descent, we show that logistic regression can exhibit grokking when the training dataset is nearly linearly separable from the origin and there is strong noise in the perpendicular directions. The underlying reason is that near the critical point, "flat" directions in the loss landscape with nearly zero gradient cause training dynamics to linger for arbitrarily long times near quasi-stable solutions before eventually reaching the global minimum. Finally, we highlight similarities between our findings and the recent literature, strengthening the conjecture that grokking generally occurs in proximity to the interpolation threshold, reminiscent of critical phenomena often observed in physical systems.
Alon Beck, Noam Itzhak Levi, Yohai Bar-Sinai
ICML2
2025 A Simple Model of Inference Scaling Laws
abstract
Neural scaling laws have garnered significant interest due to their ability to predict model performance as a function of increasing parameters, data, and compute. In this work, we propose a simple statistical ansatz based on memorization to study scaling laws in the context of inference. Specifically, how performance improves with multiple inference attempts. We explore the coverage, or pass@k metric, which measures the chance of success over repeated attempts and provide a motivation for the observed functional form of the inference scaling behavior of the coverage in large language models (LLMs) on reasoning tasks. We then define an "inference loss", which exhibits a power law decay as the number of trials increases, and connect this result with prompting costs. We further test the universality of our construction by conducting experiments on a simple generative model, and find that our predictions are in agreement with the empirical coverage curves in a controlled setting. Our simple framework sets the ground for incorporating inference scaling with other known scaling laws.
Noam Itzhak Levi
ICML1
2025 The Underlying Universal Statistical Structure of Natural Datasets
abstract
We study universal properties in real-world complex and synthetically generated datasets. Our approach is to analogize data to a physical system and employ tools from statistical physics and Random Matrix Theory (RMT) to reveal their underlying structure. Examining the local and global eigenvalue statistics of feature-feature covariance matrices, we find: (i) bulk eigenvalue power-law scaling vastly differs between uncorrelated Gaussian and real-world data, (ii) this power law behavior is reproducible using Gaussian data with long-range correlations, (iii) all dataset types exhibit chaotic RMT universality, (iv) RMT statistics emerge at smaller dataset sizes than typical training sets, correlating with power-law convergence, (v) Shannon entropy correlates with RMT structure and requires fewer samples in strongly correlated datasets. These results suggest natural image Gram matrices can be approximated by Wishart random matrices with simple covariance structure, enabling rigorous analysis of neural network behavior.
Noam Itzhak Levi, Yaron Oz
ICML1
2025 Ascent Fails to Forget
abstract
Contrary to common belief, we show that gradient ascent-based unconstrained optimization methods frequently fail to perform machine unlearning, a phenomenon we attribute to the inherent statistical dependence between the forget and retain data sets. This dependence, which can manifest itself even as simple correlations, undermines the misconception that these sets can be independently manipulated during unlearning. We provide empirical and theoretical evidence showing these methods often fail precisely due to this overlooked relationship. For random forget sets, this dependence means that degrading forget set metrics (which, for a retrained model, should mirror test set metrics) inevitably harms overall test performance. Going beyond random sets, we consider logistic regression as an instructive example where a critical failure mode emerges: inter-set dependence causes gradient descent-ascent iterations to progressively diverge from the ideal retrained model. Strikingly, these methods can converge to solutions that are not only far from the retrained ideal but are potentially even further from it than the original model itself, rendering the unlearning process actively detrimental. A toy example further illustrates how this dependence can trap models in inferior local minima, inescapable via finetuning. Our findings highlight that the presence of such statistical dependencies, even when manifest only as correlations, can be sufficient for ascent-based unlearning to fail. Our theoretical insights are corroborated by experiments on complex neural networks, demonstrating that these methods do not perform as expected in practice due to this unaddressed statistical interplay.
Ioannis Mavrothalassitis, Pol Puigdemont, Noam Itzhak Levi, Volkan Cevher
NeurIPS3
2024 Grokking in Linear Estimators - A Solvable Model that Groks without Understanding
abstract
Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data. We show both analytically and numerically that grokking can surprisingly occur in linear networks performing linear tasks in a simple teacher-student setup. In this setting, the full training dynamics is derived in terms of the expected training and generalization data covariance matrix. We present exact predictions on how the grokking time depends on input and output dimensionality, train sample size, regularization, and network parameters initialization. The key findings are that late generalization increase may not imply a transition from "memorization" to "understanding", but can simply be an artifact of the accuracy measure. We provide empirical verification for these propositions, along with preliminary results indicating that some predictions also hold for deeper networks, with non-linear activations.
Noam Itzhak Levi, Alon Beck, Yohai Bar-Sinai
ICLR1