VLDB 2026 Research / reviewers in the wild / expert
Shen Nie
dblp:342/3413
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Generative modeling · 76% Language models and text generation · 17% Optimization for machine learning · 5% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 24 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
7.3 | 9 | 2026 | LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026 Large Language Diffusion Models · NeurIPS 2025 Masked Diffusion Models as Energy Minimization · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
discrete diffusion model |
1.7 | 2 | 2025 | Masked Diffusion Models as Energy Minimization · NeurIPS 2025 Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion model |
1.7 | 2 | 2025 | Masked Diffusion Models as Energy Minimization · NeurIPS 2025 Scaling up Masked Diffusion Models on Text · ICLR 2025 |
Natural language and speech › Language models and text generation
preference optimization |
1.0 | 1 | 2026 | LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026 |
Machine learning › Optimization for machine learning
variance reduction |
1.0 | 1 | 2026 | LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
language modeling |
0.9 | 1 | 2025 | Scaling up Masked Diffusion Models on Text · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Large Language Diffusion Models · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion language model |
0.9 | 1 | 2025 | Scaling up Masked Diffusion Models on Text · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 2 | 2023 | One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale · ICML 2023 All are Worth Words: A ViT Backbone for Diffusion Models · CVPR 2023 |
Machine learning › Generative modeling › generative model › continuous-time generative model
bayesian flow network |
0.8 | 1 | 2024 | Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations · ICML 2024 |
Machine learning › Generative modeling › diffusion model › image editing
diffusion-based image editing |
0.8 | 1 | 2024 | The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing · ICLR 2024 |
Machine learning › Generative modeling › diffusion model › score-based generative model
stochastic differential equation diffusion model |
0.8 | 1 | 2024 | Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations · ICML 2024 |
Visual content generation and editing › image editing
diffusion-based image editing |
0.8 | 1 | 2024 | The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing · ICLR 2024 |
Visual content generation and editing
image editing |
0.8 | 1 | 2024 | The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing · ICLR 2024 |
Machine learning › Generative modeling › diffusion model
multimodal diffusion model |
0.7 | 1 | 2023 | One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale · ICML 2023 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.7 | 1 | 2023 | All are Worth Words: A ViT Backbone for Diffusion Models · CVPR 2023 |
Natural language and speech › Language models and text generation
alignment |
0.3 | 1 | 2026 | LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026 |
Machine learning › Generative modeling
autoregressive model |
0.3 | 1 | 2025 | Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data · ICLR 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.3 | 1 | 2025 | Large Language Diffusion Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.3 | 1 | 2025 | Large Language Diffusion Models · NeurIPS 2025 |
Mathematical optimization › optimal transport
discrete optimal transport |
0.3 | 1 | 2025 | Masked Diffusion Models as Energy Minimization · NeurIPS 2025 |
Mathematical optimization
optimal transport |
0.3 | 1 | 2025 | Masked Diffusion Models as Energy Minimization · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.2 | 1 | 2023 | One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale · ICML 2023 |
Machine learning › Generative modeling
image generation |
0.2 | 1 | 2023 | All are Worth Words: A ViT Backbone for Diffusion Models · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
stochastic differential equation · 2.3beta distribution · 1.7transformer · 1.5variance reduction · 1.0preference optimization · 1.0unsupervised classifier-free guidance · 0.9scaling law analysis · 0.9reparameterization · 0.9masked diffusion · 0.9interpolation schedule · 0.9energy minimization · 0.9concrete score · 0.9ordinary differential equation · 0.8SDE-Drag · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion ModelsabstractFengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang, Chunwei Wu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fengqi Zhu, Rongzhen Wang, Shen Nie, Chunwei Wu, Jun Zhou 0011, Yankai Lin 0001, Ji-Rong Wen, Chongxuan Li |
ACL (1) | 3 |
| 2025 | Scaling up Masked Diffusion Models on TextabstractMasked diffusion models (MDMs) have shown promise in language modeling, yet their scalability and effectiveness in core language tasks, such as text generation and language understanding, remain underexplored. This paper establishes the first scaling law for MDMs, demonstrating a scaling rate comparable to autoregressive models (ARMs) and a relatively small compute gap. Motivated by their scalability, we train a family of MDMs with up to 1.1 billion (B) parameters to systematically evaluate their performance against ARMs of comparable or larger sizes. Fully leveraging the probabilistic formulation of MDMs, we propose a simple yet effective *unsupervised classifier-free guidance* that effectively exploits large-scale unpaired data, boosting performance for conditional inference. In language understanding, the 1.1B MDM outperforms the 1.1B TinyLlama model trained on the same data across four of eight zero-shot benchmarks. Notably, it achieves competitive math reasoning ability with the 7B Llama-2 model on the GSM8K dataset.
In text generation, MDMs with 16 times more pre-training time offer a flexible trade-off against ARMs with the accelerated sampling technique KV-Cache: MDMs match ARMs in performance while being 1.4 times faster during sampling.
Moreover, MDMs address challenging tasks for ARMs by effectively handling bidirectional reasoning and adapting to temporal shifts in data. Notably, a 1.1B MDM breaks the *reverse curse* encountered by much larger ARMs with significantly more data and computation, such as 13B Llama-2 and 175B GPT-3. Our code is available at https://github.com/ML-GSAI/SMDM. Shen Nie, Fengqi Zhu, Tianyu Pang, Qian Liu 0033, Guangtao Zeng, Chongxuan Li |
ICLR | 1 |
| 2025 | Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean DataabstractDiscrete diffusion models with absorbing processes have shown promise in language modeling. The key quantities to be estimated are the ratios between the marginal probabilities of two transitive states at all timesteps, called the concrete score. In this paper, we reveal that the concrete score in absorbing diffusion can be expressed as conditional probabilities of clean data, multiplied by a time-dependent scalar in an analytic form. Motivated by this finding, we propose reparameterized absorbing discrete diffusion (RADD), a dedicated diffusion model without time-condition that characterizes the time-independent conditional probabilities. Besides its simplicity, RADD can reduce the number of function evaluations (NFEs) by caching the output of the time-independent network when the noisy sample remains unchanged in a sampling interval, which enables sampling acceleration. Built upon the new perspective of conditional distributions, we further unify absorbing discrete diffusion and any-order autoregressive models (AO-ARMs), showing that the upper bound on the negative log-likelihood for the diffusion model can be interpreted as an expected negative log-likelihood for AO-ARMs. Further, our RADD models achieve SOTA performance among diffusion models on 5 zero-shot language modeling benchmarks (measured by perplexity) at the GPT-2 scale. Our code is available at \url{https://github.com/ML-GSAI/RADD}. Jingyang Ou, Shen Nie, Fengqi Zhu, Zhenguo Li, Chongxuan Li |
ICLR | 2 |
| 2025 | Masked Diffusion Models as Energy MinimizationabstractWe present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specifically, we prove that three distinct energy formulations—kinetic, conditional kinetic, and geodesic energy—are mathematically equivalent under the structure of MDMs, and that MDMs minimize all three when the mask schedule satisfies a closed-form optimality condition. This unification not only clarifies the theoretical foundations of MDMs, but also motivates practical improvements in sampling. By parameterizing interpolation schedules via Beta distributions, we reduce the schedule design space to a tractable 2D search, enabling efficient post-training tuning without model modification. Experiments on synthetic and real-world benchmarks demonstrate that our energy-inspired schedules outperform hand-crafted baselines, particularly in low-step sampling settings. Shen Nie, Zijin Feng, Zhenguo Li, Ji-Rong Wen, Chongxuan Li |
NeurIPS | 2 |
| 2025 | Large Language Diffusion ModelsabstractThe capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing *LLaDA*, a diffusion model trained from scratch under the pre-training and supervised fine-tuning (SFT) paradigm. LLaDA employs a forward data masking process and a reverse generation process, parameterized by a Transformer to predict masked tokens. It provides a principled generative approach for probabilistic inference by optimizing a likelihood lower bound. Across extensive benchmarks on general tasks, math, code, and so on, LLaDA demonstrates strong *scalability* and performs comparably to our self-constructed ARM baselines. Remarkably, LLaDA 8B is competitive with strong LLMs like LLaMA3 8B in *in-context learning* and, after SFT, exhibits impressive *instruction-following* abilities in case studies such as multi-turn dialogue. Moreover, LLaDA addresses the reversal curse, surpassing GPT-4o in a reversal poem completion task. Our findings show the promise of diffusion models for language modeling at scale and challenge the common assumption that core LLM capabilities discussed above inherently depend on ARMs. Project page and codes: \url{https://ml-gsai.github.io/LLaDA-demo/}. Shen Nie, Fengqi Zhu, Zebin You, Jingyang Ou, Jun Zhou 0011, Yankai Lin 0001, Ji-Rong Wen, Chongxuan Li |
NeurIPS | 1 |
| 2024 | The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image EditingabstractWe present a unified probabilistic formulation for diffusion-based image editing, where a latent variable is edited in a task-specific manner and generally deviates from the corresponding marginal distribution induced by the original stochastic or ordinary differential equation (SDE or ODE). Instead, it defines a corresponding SDE or ODE for editing. In the formulation, we prove that the Kullback-Leibler divergence between the marginal distributions of the two SDEs gradually decreases while that for the ODEs remains as the time approaches zero, which shows the promise of SDE in image editing. Inspired by it, we provide the SDE counterparts for widely used ODE baselines in various tasks including inpainting and image-to-image translation, where SDE shows a consistent and substantial improvement. Moreover, we propose \emph{SDE-Drag} -- a simple yet effective method built upon the SDE formulation for point-based content dragging. We build a challenging benchmark (termed \emph{DragBench}) with open-set natural, art, and AI-generated images for evaluation. A user study on DragBench indicates that SDE-Drag significantly outperforms our ODE baseline, existing diffusion-based methods, and the renowned DragGAN. Our results demonstrate the superiority and versatility of SDE in image editing and push the boundary of diffusion-based editing methods. See the project page \url{https://ml-gsai.github.io/SDE-Drag-demo/} for the code and DragBench dataset. Shen Nie, Hanzhong Allan Guo, Cheng Lu 0011, Chenyu Zheng, Chongxuan Li |
ICLR | 1 |
| 2024 | Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential EquationsabstractBayesian flow networks (BFNs) iteratively refine the parameters, instead of the samples in diffusion models (DMs), of distributions at various noise levels through Bayesian inference. Owing to its differentiable nature, BFNs are promising in modeling both continuous and discrete data, while simultaneously maintaining fast sampling capabilities. This paper aims to understand and enhance BFNs by connecting them with DMs through stochastic differential equations (SDEs). We identify the linear SDEs corresponding to the noise-addition processes in BFNs, demonstrate that BFN’s regression losses are aligned with denoise score matching, and validate the sampler in BFN as a first-order solver for the respective reverse-time SDE. Based on these findings and existing recipes of fast sampling in DMs, we propose specialized solvers for BFNs that markedly surpass the original BFN sampler in terms of sample quality with a limited number of function evaluations (e.g., 10) on both image and text datasets. Notably, our best sampler achieves an increase in speed of $5\sim20$ times for free. Shen Nie, Xu Min, Jun Zhou 0011, Chongxuan Li |
ICML | 3 |
| 2023 | All are Worth Words: A ViT Backbone for Diffusion ModelsabstractVision transformers (ViT) have shown promise in various vision tasks while the U-Net based on a convolutional neural network (CNN) remains dominant in diffusion models. We design a simple and general ViT-based architecture (named U-ViT) for image generation with diffusion models. U-ViT is characterized by treating all inputs including the time, condition and noisy image patches as tokens and employing long skip connections between shallow and deep layers. We evaluate U-ViT in unconditional and classconditional image generation, as well as text-to-image generation tasks, where U-ViT is comparable if not superior to a CNN-based U-Net of a similar size. In particular, latent diffusion models with U-ViT achieve record-breaking FID scores of 2.29 in class-conditional image generation on ImageNet 256×256, and 5.48 in text-to-image generation on MS-COCO, among methods without accessing large external datasets during the training of generative models. Our results suggest that, for diffusion-based image modeling, the long skip connection is crucial while the down-sampling and upsampling operators in CNN-based U-Net are not always necessary. We believe that U-ViT can provide insights for future research on backbones in diffusion models and benefit generative modeling on large scale cross-modality datasets. Fan Bao, Shen Nie, Yue Cao 0001, Chongxuan Li, Hang Su 0006, Jun Zhu 0001 |
CVPR | 2 |
| 2023 | One Transformer Fits All Distributions in Multi-Modal Diffusion at ScaleabstractThis paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoken models (e.g., Stable Diffusion and DALL-E 2) in representative tasks (e.g., text-to-image generation). Fan Bao, Shen Nie, Chongxuan Li, Shi Pu 0002, Yaole Wang, Gang Yue, Yue Cao 0001, Hang Su 0006, Jun Zhu 0001 |
ICML | 2 |