VLDB 2026 Research / reviewers in the wild / expert
Sam Buchanan
dblp:226/5790
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Learning theory · 24% Generative modeling · 23% Representation and self-supervised learning · 20% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | Beyond Scores: Proximal Diffusion Models · NeurIPS 2025 On the Edge of Memorization in Diffusion Models · NeurIPS 2025 Masked Completion via Structured Diffusion with White-Box Transformers · ICLR 2024 |
Machine learning › Learning theory
generalization bounds |
1.5 | 3 | 2025 | On the Edge of Memorization in Diffusion Models · NeurIPS 2025 Deep Networks Provably Classify Data on Curves · NeurIPS 2021 Efficient Dictionary Learning with Gradient Descent · ICML 2019 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › sparse feature learning
sparse rate reduction |
1.4 | 2 | 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? · J. Mach. Learn. Res. 2024 White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
transformer |
1.4 | 2 | 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? · J. Mach. Learn. Res. 2024 White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
memorization |
0.9 | 1 | 2025 | On the Edge of Memorization in Diffusion Models · NeurIPS 2025 |
Machine learning › Learning theory › neural network theory
memorization and generalization |
0.9 | 1 | 2025 | On the Edge of Memorization in Diffusion Models · NeurIPS 2025 |
Machine learning › Learning theory
phase transition |
0.9 | 1 | 2025 | On the Edge of Memorization in Diffusion Models · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
0.9 | 1 | 2025 | Beyond Scores: Proximal Diffusion Models · NeurIPS 2025 |
Machine learning › Generative modeling
inverse problem |
0.8 | 1 | 2024 | What's in a Prior? Learned Proximal Networks for Inverse Problems · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.8 | 1 | 2024 | Masked Completion via Structured Diffusion with White-Box Transformers · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › prior modeling
prior learning |
0.8 | 1 | 2024 | What's in a Prior? Learned Proximal Networks for Inverse Problems · ICLR 2024 |
Computer vision › 3D vision › implicit neural representation
neural field |
0.7 | 1 | 2023 | Canonical Factors for Hybrid Neural Fields · ICCV 2023 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
0.5 | 1 | 2021 | Deep Networks and the Multiple Manifold Problem · ICLR 2021 |
Machine learning › Learning theory
neural network theory |
0.5 | 1 | 2021 | Deep Networks and the Multiple Manifold Problem · ICLR 2021 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.5 | 1 | 2021 | Deep Networks Provably Classify Data on Curves · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning
representation geometry |
0.5 | 1 | 2021 | Deep Networks and the Multiple Manifold Problem · ICLR 2021 |
Machine learning › Optimization for machine learning
convergence guarantees |
0.4 | 1 | 2019 | Efficient Dictionary Learning with Gradient Descent · ICML 2019 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
0.4 | 1 | 2019 | Efficient Dictionary Learning with Gradient Descent · ICML 2019 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.4 | 1 | 2019 | Efficient Dictionary Learning with Gradient Descent · ICML 2019 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.4 | 1 | 2019 | Efficient Dictionary Learning with Gradient Descent · ICML 2019 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2023 | White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023 |
Geometric modeling and processing › surface reconstruction › implicit surface reconstruction
signed distance field reconstruction |
0.2 | 1 | 2023 | Canonical Factors for Hybrid Neural Fields · ICCV 2023 |
Algorithms and data structures › numerical linear algebra › dimensionality reduction › nonlinear dimensionality reduction
manifold learning |
0.1 | 1 | 2021 | Deep Networks Provably Classify Data on Curves · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
proximal matching · 1.6sparse coding · 1.4alternating optimization · 1.4gradient descent · 1.4theoretical analysis · 0.9synthetic data experiments · 0.9stochastic differential equation · 0.9proximal operator · 0.9plug-and-play · 0.8deep unrolling · 0.8factored feature volumes · 0.7canonical transformation learning · 0.7neural tangent kernel · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Edge of Memorization in Diffusion ModelsabstractWhen do diffusion models reproduce their training data, and when are they able to generate samples beyond it? A practically relevant theoretical understanding of this interplay between memorization and generalization may significantly impact real-world deployments of diffusion models with respect to issues such as copyright infringement and data privacy. In this paper, we propose a scientific and
mathematical “laboratory” for investigating memorization and generalization in diffusion models trained on fully synthetic or natural image-like structured data. Within this setting, we theoretically characterize a crossover point wherein the weighted training loss of a fully generalizing model becomes greater than that of an underparameterized memorizing model at a critical value of model (under)parameterization. We then demonstrate via carefully-designed experiments that the location of this crossover predicts a phase transition in diffusion models trained via gradient descent, as our theory enables us to analytically predict the model size at which memorization becomes predominant. Our work provides an analytically tractable and practically meaningful setting for future theoretical and empirical investigations. Code for our experiments is available at https://github.com/DruvPai/diffusion_mem_gen. Sam Buchanan, Druv Pai, Yi Ma 0001, Valentin De Bortoli |
NeurIPS | 1 |
| 2025 | Beyond Scores: Proximal Diffusion ModelsabstractDiffusion models have quickly become some of the most popular and powerful generative models for high-dimensional data. The key insight that enabled their development was the realization that access to the score---the gradient of the log-density at different noise levels---allows for sampling from data distributions by solving a reverse-time stochastic differential equation (SDE) via forward discretization, and that popular denoisers allow for unbiased estimators of this score. In this paper, we demonstrate that an alternative, backward discretization of these SDEs, using proximal maps in place of the score, leads to theoretical and practical benefits. We leverage recent results in _proximal matching_ to learn proximal operators of the log-density and, with them, develop Proximal Diffusion Models (`ProxDM`). Theoretically, we prove that $\widetilde{\mathcal O}(d/\sqrt{\varepsilon})$ steps suffice for the resulting discretization to generate an $\varepsilon$-accurate distribution w.r.t. the KL divergence.
Empirically, we show that two variants of `ProxDM` achieve significantly faster convergence within just a few sampling steps compared to conventional score-matching methods. Zhenghan Fang, Mateo Díaz, Sam Buchanan, Jeremias Sulam |
NeurIPS | 3 |
| 2024 | What's in a Prior? Learned Proximal Networks for Inverse ProblemsabstractProximal operators are ubiquitous in inverse problems, commonly appearing as part of algorithmic strategies to regularize problems that are otherwise ill-posed. Modern deep learning models have been brought to bear for these tasks too, as in the framework of plug-and-play or deep unrolling, where they loosely resemble proximal operators. Yet, something essential is lost in employing these purely data-driven approaches: there is no guarantee that a general deep network represents the proximal operator of any function, nor is there any characterization of the function for which the network might provide some approximate proximal. This not only makes guaranteeing convergence of iterative schemes challenging but, more fundamentally, complicates the analysis of what has been learned by these networks about their training data. Herein we provide a framework to develop *learned proximal networks* (LPN), prove that they provide exact proximal operators for a data-driven nonconvex regularizer, and show how a new training strategy, dubbed *proximal matching*, provably promotes the recovery of the log-prior of the true data distribution. Such LPN provide general, unsupervised, expressive proximal operators that can be used for general inverse problems with convergence guarantees. We illustrate our results in a series of cases of increasing complexity, demonstrating that these models not only result in state-of-the-art performance, but provide a window into the resulting priors learned from data. Zhenghan Fang, Sam Buchanan, Jeremias Sulam |
ICLR | 2 |
| 2024 | Masked Completion via Structured Diffusion with White-Box TransformersabstractModern learning frameworks often train deep neural networks with massive amounts of unlabeled data to learn representations by solving simple pretext tasks, then use the representations as foundations for downstream tasks. These networks are empirically designed; as such, they are usually not interpretable, their representations are not structured, and their designs are potentially redundant. White-box deep networks, in which each layer explicitly identifies and transforms structures in the data, present a promising alternative. However, existing white-box architectures have only been shown to work at scale in supervised settings with labeled data, such as classification. In this work, we provide the first instantiation of the white-box design paradigm that can be applied to large-scale unsupervised representation learning. We do this by exploiting a fundamental connection between diffusion, compression, and (masked) completion, deriving a deep transformer-like masked autoencoder architecture, called CRATE-MAE, in which the role of each layer is mathematically fully interpretable: they transform the data distribution to and from a structured representation. Extensive empirical evaluations confirm our analytical insights. CRATE-MAE demonstrates highly promising performance on large-scale imagery datasets while using only ~30% of the parameters compared to the standard masked autoencoder with the same model configuration. The representations learned by CRATE-MAE have explicit structure and also contain semantic meaning. Druv Pai, Sam Buchanan, Ziyang Wu, Yaodong Yu, Yi Ma 0001 |
ICLR | 2 |
| 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?abstractIn this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimensional Gaussian mixture supported on incoherent subspaces. The goodness of such a representation can be evaluated by a principled measure, called sparse rate reduction, that simultaneously maximizes the intrinsic information gain and extrinsic sparsity of the learned representation. From this perspective, popular deep network architectures, including transformers, can be viewed as realizing iterative schemes to optimize this measure. Particularly, we derive a transformer block from alternating optimization on parts of this objective: the multi-head self-attention operator compresses the representation by implementing an approximate gradient descent step on the coding rate of the features, and the subsequent multi-layer perceptron sparsifies the features. This leads to a family of white-box transformer-like deep network architectures, named CRATE, which are mathematically fully interpretable. We show, by way of a novel connection between denoising and compression, that the inverse to the aforementioned compressive encoding can be realized by the same class of CRATE architectures. Thus, the so-derived white-box architectures are universal to both encoders and decoders. Experiments show that these networks, despite their simplicity, indeed learn to compress and sparsify representations of large-scale real-world image and text datasets, and achieve strong performance across different settings: ViT, MAE, DINO, BERT, and GPT2. We believe the proposed computational framework demonstrates great potential in bridging the gap between theory and practice of deep learning, from a unified perspective of data compression. Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Yuexiang Zhai, Benjamin D. Haeffele, Yi Ma 0001 |
J. Mach. Learn. Res. | 2 |
| 2023 | Canonical Factors for Hybrid Neural FieldsabstractFactored feature volumes offer a simple way to build more compact, efficient, and intepretable neural fields, but also introduce biases that are not necessarily beneficial for realworld data. In this work, we (1) characterize the undesirable biases that these architectures have for axis-aligned signals—they can lead to radiance field reconstruction differences of as high as 2 PSNR—and (2) explore how learning a set of canonicalizing transformations can improve representations by removing these biases. We prove in a simple two-dimensional model problem that a hybrid architecture that simultaneously learns these transformations together with scene appearance succeeds with drastically improved efficiency. We validate the resulting architectures, which we call TILTED, using 2D image, signed distance field, and radiance field reconstruction tasks, where we observe improvements across quality, robustness, compactness, and runtime. Results demonstrate that TILTED can enable capabilities comparable to baselines that are 2x larger, while highlighting weaknesses of standard procedures for evaluating neural field representations. Brent Yi, Weijia Zeng, Sam Buchanan, Yi Ma 0001 |
ICCV | 3 |
| 2023 | White-Box Transformers via Sparse Rate ReductionabstractIn this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent subspaces. The quality of the final representation can be measured by a unified objective function called sparse rate reduction. From this perspective, popular deep networks such as transformers can be naturally viewed as realizing iterative schemes to optimize this objective incrementally. Particularly, we show that the standard transformer block can be derived from alternating optimization on complementary parts of this objective: the multi-head self-attention operator can be viewed as a gradient descent step to compress the token sets by minimizing their lossy coding rate, and the subsequent multi-layer perceptron can be viewed as attempting to sparsify the representation of the tokens. This leads to a family of white-box transformer-like deep network architectures which are mathematically fully interpretable. Despite their simplicity, experiments show that these networks indeed learn to optimize the designed objective: they compress and sparsify representations of large-scale real-world vision datasets such as ImageNet, and achieve performance very close to thoroughly engineered transformers such as ViT.
Code is at https://github.com/Ma-Lab-Berkeley/CRATE. Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Benjamin D. Haeffele, Yi Ma 0001 |
NeurIPS | 2 |
| 2021 | Deep Networks and the Multiple Manifold Problem
Sam Buchanan, Dar Gilboa, John Wright 0001 |
ICLR | 1 |
| 2021 | Deep Networks Provably Classify Data on CurvesabstractData with low-dimensional nonlinear structure are ubiquitous in engineering and scientific problems. We study a model problem with such structure---a binary classification task that uses a deep fully-connected neural network to classify data drawn from two disjoint smooth curves on the unit sphere. Aside from mild regularity conditions, we place no restrictions on the configuration of the curves. We prove that when (i) the network depth is large relative to certain geometric properties that set the difficulty of the problem and (ii) the network width and number of samples is polynomial in the depth, randomly-initialized gradient descent quickly learns to correctly classify all points on the two curves with high probability. To our knowledge, this is the first generalization guarantee for deep networks with nonlinear data that depends only on intrinsic data properties. Our analysis proceeds by a reduction to dynamics in the neural tangent kernel (NTK) regime, where the network depth plays the role of a fitting resource in solving the classification problem. In particular, via fine-grained control of the decay properties of the NTK, we demonstrate that when the network is sufficiently deep, the NTK can be locally approximated by a translationally invariant operator on the manifolds and stably inverted over smooth functions, which guarantees convergence and generalization. Tingran Wang, Sam Buchanan, Dar Gilboa, John Wright 0001 |
NeurIPS | 2 |
| 2019 | Efficient Dictionary Learning with Gradient DescentabstractRandomly initialized first-order optimization algorithms are the method of choice for solving many high-dimensional nonconvex problems in machine learning, yet general theoretical guarantees cannot rule out convergence to critical points of poor objective value. For some highly structured nonconvex problems however, the success of gradient descent can be understood by studying the geometry of the objective. We study one such problem – complete orthogonal dictionary learning, and provide converge guarantees for randomly initialized gradient descent to the neighborhood of a global optimum. The resulting rates scale as low order polynomials in the dimension even though the objective possesses an exponential number of saddle points. This efficient convergence can be viewed as a consequence of negative curvature normal to the stable manifolds associated with saddle points, and we provide evidence that this feature is shared by other nonconvex problems of importance as well. Dar Gilboa, Sam Buchanan, John Wright 0001 |
ICML | 2 |
| 2018 | Efficient Model-Free Learning to Overcome Hardware Nonidealities in Analog-to-Information ConvertersabstractThis paper considers compressed sensing (CS) in the context of RF spectrum sensing and presents an efficient approach for learning hardware nonidealities in an analog-to-information converter (A2IC). The proposed methodology is based on the learned iterative shrinkage-thresholding algorithm (LISTA), which enables co-optimization of the hardware and the reconstruction algorithm and leads to a model-free recovery approach that is optimally tuned for the unique computational constraints and hardware nonidealities present in the RF frontend. To achieve this, we devise a training protocol that employs a dataset and neural network of minimal sizes. We demonstrate the effectiveness of our methodology on simulated data from a model of a well-established CS A2IC in the presence of linear impairments and noise. The recovery process extrapolates from training on 1-sparse signals to recovering the support of signals whose sparsity runs up to the theoretical optimum for l1-based algorithms across a range of typical operating SNRs. Sam Buchanan, Tanbir Haque, Peter R. Kinget, John Wright 0001 |
ICASSP | 1 |