Marvin Li

dblp:359/3841 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 50% Robot navigation and mapping · 21% Optimization for machine learning · 18%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
Blink of an eye: a simple theory for feature localization in generative models · ICML 2025
Critical windows: non-asymptotic theory for feature emergence in diffusion models · ICML 2024
Robotics › Robot navigation and mapping › localization
probabilistic localization
0.912025
Blink of an eye: a simple theory for feature localization in generative models · ICML 2025
Machine learning › Optimization for machine learning › convergence analysis
non-asymptotic analysis
0.812024
Critical windows: non-asymptotic theory for feature emergence in diffusion models · ICML 2024
Security and privacy of machine learning
membership inference
0.712023
MoPe: Model Perturbation based Privacy Attacks on Language Models · EMNLP 2023
Security and privacy of machine learning › model privacy
memorization
0.712023
MoPe: Model Perturbation based Privacy Attacks on Language Models · EMNLP 2023
Machine learning › Generative modeling
autoregressive model
0.312025
Blink of an eye: a simple theory for feature localization in generative models · ICML 2025
Machine learning › Trustworthy machine learning › adversarial machine learning
jailbreak attack
0.312025
Blink of an eye: a simple theory for feature localization in generative models · ICML 2025
Machine learning › Generative modeling › diffusion model
diffusion model interpretability
0.212024
Critical windows: non-asymptotic theory for feature emergence in diffusion models · ICML 2024
Machine learning › Trustworthy machine learning
interpretability
0.212024
Critical windows: non-asymptotic theory for feature emergence in diffusion models · ICML 2024

Methods — techniques the papers use, named apart from their topics

stochastic localization · 0.9mathematical analysis · 0.9mixture of log-concave densities · 0.8hierarchical sampler analysis · 0.8model perturbation · 0.7hessian trace approximation · 0.7
YearPublicationVenuePosition
2025 Blink of an eye: a simple theory for feature localization in generative models
abstract
Large language models can exhibit unexpected behavior in the blink of an eye. In a recent computer use demo, a language model switched from coding to Googling pictures of Yellowstone, and these sudden shifts in behavior have also been observed in reasoning patterns and jailbreaks. This phenomenon is not unique to autoregressive models: in diffusion models, key features of the final output are decided in narrow ``critical windows'' of the generation process. In this work we develop a simple, unifying theory to explain this phenomenon. Using the formalism of stochastic localization for generative models, we show that it emerges generically as the generation process localizes to a sub-population of the distribution it models. While critical windows have been studied at length in diffusion models, existing theory heavily relies on strong distributional assumptions and the particulars of Gaussian diffusion. In contrast to existing work our theory (1) applies to autoregressive and diffusion models; (2) makes very few distributional assumptions; (3) quantitatively improves previous bounds even when specialized to diffusions; and (4) requires basic mathematical tools. Finally, we validate our predictions empirically for LLMs and find that critical windows often coincide with failures in problem solving for various math and reasoning benchmarks.
Marvin Li, Aayush Karan, Sitan Chen
ICML1
2024 Critical windows: non-asymptotic theory for feature emergence in diffusion models
abstract
We develop theory to understand an intriguing property of diffusion models for image generation that we term critical windows. Empirically, it has been observed that there are narrow time intervals in sampling during which particular features of the final image emerge, e.g. the image class or background color (Ho et al., 2020b; Meng et al., 2022; Choi et al., 2022; Raya & Ambrogioni, 2023; Georgiev et al., 2023; Sclocchi et al., 2024; Biroli et al., 2024). While this is advantageous for interpretability as it implies one can localize properties of the generation to a small segment of the trajectory, it seems at odds with the continuous nature of the diffusion. We propose a formal framework for studying these windows and show that for data coming from a mixture of strongly log-concave densities, these windows can be provably bounded in terms of certain measures of inter- and intra-group separation. We also instantiate these bounds for concrete examples like well-conditioned Gaussian mixtures. Finally, we use our bounds to give a rigorous interpretation of diffusion models as hierarchical samplers that progressively “decide” output features over a discrete sequence of times. We validate our bounds with experiments on synthetic data and show that critical windows may serve as a useful tool for diagnosing fairness and privacy violations in real-world diffusion models.
Marvin Li, Sitan Chen
ICML1
2023 MoPe: Model Perturbation based Privacy Attacks on Language Models
abstract
Recent work has shown that Large Language Models (LLMs) can unintentionally leak sensitive information present in their training data.In this paper, we present MoPe θ (Model Perturbations), a new method to identify with high confidence if a given text is in the training data of a pre-trained language model, given white-box access to the models parameters.MoPe θ adds noise to the model in parameter space and measures the drop in log-likelihood at a given point x, a statistic we show approximates the trace of the Hessian matrix with respect to model parameters.Across language models ranging from 70M to 12B parameters, we show that MoPe θ is more effective than existing loss-based attacks and recently proposed perturbation-based methods.We also examine the role of training point order and model size in attack success, and empirically demonstrate that MoPe θ accurately approximate the trace of the Hessian in practice.Our results show that the loss of a point alone is insufficient to determine extractability-there are training points we can recover using our method that have average loss.This casts some doubt on prior works that use the loss of a point as evidence of memorization or "unlearning."
Marvin Li, Jeffrey G. Wang, Seth Neel
EMNLP1