VLDB 2026 Research / reviewers in the wild / expert
Subham Sekhar Sahoo
dblp:268/6867 · also Subham S. Sahoo
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Generative modeling · 71% Optimization for machine learning · 10% Probabilistic and Bayesian machine learning · 6% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
5.0 | 6 | 2025 | Remasking Discrete Diffusion Models with Inference-Time Scaling · NeurIPS 2025 The Diffusion Duality · ICML 2025 Simple Guidance Mechanisms for Discrete Diffusion Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model |
2.8 | 4 | 2025 | The Diffusion Duality · ICML 2025 Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models · ICLR 2025 Simple and Effective Masked Diffusion Language Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
discrete diffusion model |
2.6 | 3 | 2025 | Remasking Discrete Diffusion Models with Inference-Time Scaling · NeurIPS 2025 The Diffusion Duality · ICML 2025 Simple Guidance Mechanisms for Discrete Diffusion Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
controllable generation |
0.9 | 1 | 2025 | Simple Guidance Mechanisms for Discrete Diffusion Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion model |
0.9 | 1 | 2025 | Remasking Discrete Diffusion Models with Inference-Time Scaling · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.8 | 1 | 2024 | Diffusion Models With Learned Adaptive Noise · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion language model |
0.8 | 1 | 2024 | Simple and Effective Masked Diffusion Language Models · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
backpropagation |
0.7 | 1 | 2023 | Backpropagation through Combinatorial Algorithms: Identity with Projection Works · ICLR 2023 |
Machine learning › Optimization for machine learning
combinatorial optimization |
0.7 | 1 | 2023 | Backpropagation through Combinatorial Algorithms: Identity with Projection Works · ICLR 2023 |
Machine learning › Optimization for machine learning
differentiable optimization |
0.7 | 1 | 2023 | Backpropagation through Combinatorial Algorithms: Identity with Projection Works · ICLR 2023 |
Machine learning › Optimization for machine learning
implicit differentiation |
0.7 | 1 | 2023 | Backpropagation through Combinatorial Algorithms: Identity with Projection Works · ICLR 2023 |
Machine learning › Generative modeling
normalizing flow |
0.7 | 1 | 2023 | Semi-Autoregressive Energy Flows: Exploring Likelihood-Free Training of Normalizing Flows · ICML 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Scaling Symbolic Methods using Gradients for Neural Model Explanation · ICLR 2021 |
Machine learning › Trustworthy machine learning › interpretability
model explanation |
0.5 | 1 | 2021 | Scaling Symbolic Methods using Gradients for Neural Model Explanation · ICLR 2021 |
Algorithms and data structures
symbolic computation |
0.5 | 1 | 2021 | Scaling Symbolic Methods using Gradients for Neural Model Explanation · ICLR 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge discovery
equation discovery |
0.3 | 1 | 2018 | Learning Equations for Extrapolation and Control · ICML 2018 |
Natural language and speech › Language models and text generation › neural language model
autoregressive language model |
0.2 | 1 | 2024 | Simple and Effective Masked Diffusion Language Models · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
evidence lower bound |
0.2 | 1 | 2024 | Diffusion Models With Learned Adaptive Noise · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.2 | 1 | 2024 | Diffusion Models With Learned Adaptive Noise · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
variational lower bound · 0.9inference-time scaling · 0.9discrete denoising diffusion · 0.9diffusion guidance · 0.9curriculum learning · 0.9consistency distillation · 0.9classifier-free guidance · 0.9classifier-based guidance · 0.9autoregressive modeling · 0.9KV caching · 0.9symbolic reasoning · 0.5gradient-based search · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Block Diffusion: Interpolating Between Autoregressive and Diffusion Language ModelsabstractDiffusion language models offer unique benefits over autoregressive models due to their potential for parallelized generation and controllability, yet they lag in likelihood modeling and are limited to fixed-length generation. In this work, we introduce a class of block diffusion language models that interpolate between discrete denoising diffusion and autoregressive models. Block diffusion overcomes key limitations of both approaches by supporting flexible-length generation and improving inference efficiency with KV caching and parallel token sampling. We propose a recipe for building effective block diffusion models that includes an efficient training algorithm, estimators of gradient variance, and data-driven noise schedules to minimize the variance. Block diffusion sets a new state-of-the-art performance among diffusion models on language modeling benchmarks and enables generation of arbitrary-length sequences. We provide the code, along with the model weights and blog post on the project page: https://m-arriola.com/bd3lms/ Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhixuan Qi, Subham Sekhar Sahoo, Volodymyr Kuleshov |
ICLR | 7 |
| 2025 | Simple Guidance Mechanisms for Discrete Diffusion ModelsabstractDiffusion models for continuous data gained widespread adoption owing to their high quality generation and control mechanisms. However, controllable diffusion on discrete data faces challenges given that continuous guidance methods do not directly apply to discrete diffusion. Here, we provide a straightforward derivation of classifier-free and classifier-based guidance for discrete diffusion, as well as a new class of diffusion models that leverage uniform noise and that are more guidable because they can continuously edit their outputs. We improve the quality of these models with a novel continuous-time variational lower bound that yields state-of-the-art performance, especially in settings involving guidance or fast generation. Empirically, we demonstrate that our guidance mechanisms combined with uniform noise diffusion improve controllable generation relative to autoregressive and diffusion baselines on several discrete data domains, including genomic sequences, small molecule design, and discretized image generation. Yair Schiff, Subham Sekhar Sahoo, Hao Phung, Guanghan Wang, Sam Boshar, Hugo Dalla-torre, Bernardo P. de Almeida, Alexander M. Rush, Thomas Pierrot, Volodymyr Kuleshov |
ICLR | 2 |
| 2025 | The Diffusion DualityabstractUniform-state discrete diffusion models hold the promise of fast text generation due to their inherent ability to self-correct. However, they are typically outperformed by autoregressive models and masked diffusion models. In this work, we narrow this performance gap by leveraging a key insight: Uniform-state diffusion processes naturally emerge from an underlying Gaussian diffusion. Our method, Duo, transfers powerful techniques from Gaussian diffusion to improve both training and sampling. First, we introduce a curriculum learning strategy guided by the Gaussian process, **doubling training speed** by reducing variance. Models trained with curriculum learning surpass autoregressive models in zero-shot perplexity on 3 of 7 benchmarks.
Second, we present Discrete Consistency Distillation, which adapts consistency distillation from the continuous to the discrete setting. This algorithm **unlocks few-step generation in diffusion language models** by accelerating sampling by two orders of magnitude. We provide the code and model checkpoints on the project page: http://s-sahoo.github.io/duo Subham Sekhar Sahoo, Justin Deschenaux, Aaron Gokaslan, Guanghan Wang, Justin T. Chiu, Volodymyr Kuleshov |
ICML | 1 |
| 2025 | Remasking Discrete Diffusion Models with Inference-Time ScalingabstractPart of the success of diffusion models stems from their ability to perform iterative refinement, i.e., repeatedly correcting outputs during generation. However, modern masked discrete diffusion lacks this capability: when a token is generated, it cannot be updated again, even when it introduces an error. Here, we address this limitation by introducing the remasking diffusion model (ReMDM) sampler, a method that can be applied to pretrained masked diffusion models in a principled way and that is derived from a discrete diffusion model with a custom remasking backward process. Most interestingly, ReMDM endows discrete diffusion with a form of inference-time compute scaling. By increasing the number of sampling steps, ReMDM generates natural language outputs that approach the quality of autoregressive models, whereas when the computation budget is limited, ReMDM better maintains quality. ReMDM also improves sample quality of masked diffusion models for discretized images, and in scientific domains such as molecule design, ReMDM facilitates diffusion guidance and pushes the Pareto frontier of controllability relative to classical masking and uniform noise diffusion. When applied to large pretrained diffusion language models, ReMDM boosts the model’s performance on downstream tasks requiring factual knowledge grasp and reasoning ability. Guanghan Wang, Yair Schiff, Subham Sekhar Sahoo, Volodymyr Kuleshov |
NeurIPS | 3 |
| 2024 | Simple and Effective Masked Diffusion Language ModelsabstractWhile diffusion models excel at generating high-quality images, prior work reports a significant performance gap between diffusion and autoregressive (AR) methods in language modeling.
In this work, we show that simple masked discrete diffusion is more performant than previously thought.
We apply an effective training recipe that improves the performance of masked diffusion models and derive a simplified, Rao-Blackwellized objective that results in additional improvements.
Our objective has a simple form—it is a mixture of classical masked language modeling losses—and can be used to train encoder-only language models that admit efficient samplers, including ones that can generate arbitrary lengths of text semi-autoregressively like a traditional language model.
On language modeling benchmarks, a range of masked diffusion models trained with modern engineering practices achieves a new state-of-the-art among diffusion models, and approaches AR perplexity. We provide the code, along with a blog post and video tutorial on the project page: https://s-sahoo.com/mdlm Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander Rush, Volodymyr Kuleshov |
NeurIPS | 1 |
| 2024 | Diffusion Models With Learned Adaptive NoiseabstractDiffusion models have gained traction as powerful algorithms for synthesizing high-quality images. Central to these algorithms is the diffusion process, a set of equations which maps data to noise
in a way that can significantly affect performance.
In this paper, we explore whether the diffusion
process can be learned from data.
Our work is grounded in Bayesian inference and seeks to improve log-likelihood estimation by casting the learned diffusion process as an approximate variational posterior that yields a tighter lower bound (ELBO) on the likelihood.
A widely held assumption is that the ELBO is invariant to the noise process: our work dispels this assumption and proposes multivariate learned adaptive noise (MuLAN), a learned diffusion process that applies noise at different rates across an image. Our method consists of three components: a multivariate noise schedule, adaptive input-conditional diffusion, and auxiliary variables; these components ensure that the ELBO is no longer invariant to the choice of the noise schedule as in previous works. Empirically, MuLAN sets a new **state-of-the-art** in density estimation on CIFAR-10 and ImageNet while matching the performance of previous state-of-the-art models with **50%** fewer steps. We provide the code, along with a blog post and video tutorial on the project page: https://s-sahoo.com/MuLAN Subham Sekhar Sahoo, Aaron Gokaslan, Christopher De Sa, Volodymyr Kuleshov |
NeurIPS | 1 |
| 2023 | Backpropagation through Combinatorial Algorithms: Identity with Projection Works
Subham Sekhar Sahoo, Anselm Paulus, Marin Vlastelica Pogancic, Vít Musil, Volodymyr Kuleshov, Georg Martius |
ICLR | 1 |
| 2023 | Semi-Autoregressive Energy Flows: Exploring Likelihood-Free Training of Normalizing FlowsabstractTraining normalizing flow generative models can be challenging due to the need to calculate computationally expensive determinants of Jacobians. This paper studies the likelihood-free training of flows and proposes the energy objective, an alternative sample-based loss based on proper scoring rules. The energy objective is determinant-free and supports flexible model architectures that are not easily compatible with maximum likelihood training, including semi-autoregressive energy flows, a novel model family that interpolates between fully autoregressive and non-autoregressive models. Energy flows feature competitive sample quality, posterior inference, and generation speed relative to likelihood-based flows; this performance is decorrelated from the quality of log-likelihood estimates, which are generally very poor. Our findings question the use of maximum likelihood as an objective or a metric, and contribute to a scientific study of its role in generative modeling. Code is available at https://github.com/ps789/SAEF. Phillip Si, Zeyi Chen, Subham Sekhar Sahoo, Yair Schiff, Volodymyr Kuleshov |
ICML | 3 |
| 2021 | Scaling Symbolic Methods using Gradients for Neural Model Explanation
Subham Sekhar Sahoo, Subhashini Venugopalan, Li Li 0060, Rishabh Singh, Patrick F. Riley |
ICLR | 1 |
| 2018 | Learning Equations for Extrapolation and ControlabstractWe present an approach to identify concise equations from data using a shallow neural network approach. In contrast to ordinary black-box regression, this approach allows understanding functional relations and generalizing them from observed data to unseen parts of the parameter space. We show how to extend the class of learnable equations for a recently proposed equation learning network to include divisions, and we improve the learning and model selection strategy to be useful for challenging real-world data. For systems governed by analytical expressions, our method can in many cases identify the true underlying equation and extrapolate to unseen domains. We demonstrate its effectiveness by experiments on a cart-pendulum system, where only 2 random rollouts are required to learn the forward dynamics and successfully achieve the swing-up task. Subham Sekhar Sahoo, Christoph H. Lampert, Georg Martius |
ICML | 1 |