Shen Nie

dblp:342/3413 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Generative modeling · 76% Language models and text generation · 17% Optimization for machine learning · 5%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 24 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
7.392026
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026
Large Language Diffusion Models · NeurIPS 2025
Masked Diffusion Models as Energy Minimization · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
discrete diffusion model
1.722025
Masked Diffusion Models as Energy Minimization · NeurIPS 2025
Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data · ICLR 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion model
1.722025
Masked Diffusion Models as Energy Minimization · NeurIPS 2025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Natural language and speech › Language models and text generation
preference optimization
1.012026
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026
Machine learning › Optimization for machine learning
variance reduction
1.012026
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026
Natural language and speech › Language models and text generation
language modeling
0.912025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Large Language Diffusion Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion language model
0.912025
Scaling up Masked Diffusion Models on Text · ICLR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.922023
One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale · ICML 2023
All are Worth Words: A ViT Backbone for Diffusion Models · CVPR 2023
Machine learning › Generative modeling › generative model › continuous-time generative model
bayesian flow network
0.812024
Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations · ICML 2024
Machine learning › Generative modeling › diffusion model › image editing
diffusion-based image editing
0.812024
The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing · ICLR 2024
Machine learning › Generative modeling › diffusion model › score-based generative model
stochastic differential equation diffusion model
0.812024
Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations · ICML 2024
Visual content generation and editing › image editing
diffusion-based image editing
0.812024
The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing · ICLR 2024
Visual content generation and editing
image editing
0.812024
The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing · ICLR 2024
Machine learning › Generative modeling › diffusion model
multimodal diffusion model
0.712023
One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale · ICML 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.712023
All are Worth Words: A ViT Backbone for Diffusion Models · CVPR 2023
Natural language and speech › Language models and text generation
alignment
0.312026
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models · ACL (1) 2026
Machine learning › Generative modeling
autoregressive model
0.312025
Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data · ICLR 2025
Natural language and speech › Language models and text generation
in-context learning
0.312025
Large Language Diffusion Models · NeurIPS 2025
Natural language and speech › Language models and text generation
instruction following
0.312025
Large Language Diffusion Models · NeurIPS 2025
Mathematical optimization › optimal transport
discrete optimal transport
0.312025
Masked Diffusion Models as Energy Minimization · NeurIPS 2025
Mathematical optimization
optimal transport
0.312025
Masked Diffusion Models as Energy Minimization · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.212023
One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale · ICML 2023
Machine learning › Generative modeling
image generation
0.212023
All are Worth Words: A ViT Backbone for Diffusion Models · CVPR 2023

Methods — techniques the papers use, named apart from their topics

stochastic differential equation · 2.3beta distribution · 1.7transformer · 1.5variance reduction · 1.0preference optimization · 1.0unsupervised classifier-free guidance · 0.9scaling law analysis · 0.9reparameterization · 0.9masked diffusion · 0.9interpolation schedule · 0.9energy minimization · 0.9concrete score · 0.9ordinary differential equation · 0.8SDE-Drag · 0.8
YearPublicationVenuePosition
2026 LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
abstract
Fengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang, Chunwei Wu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fengqi Zhu, Rongzhen Wang, Shen Nie, Chunwei Wu, Jun Zhou 0011, Yankai Lin 0001, Ji-Rong Wen, Chongxuan Li
ACL (1)3
2025 Scaling up Masked Diffusion Models on Text
abstract
Masked diffusion models (MDMs) have shown promise in language modeling, yet their scalability and effectiveness in core language tasks, such as text generation and language understanding, remain underexplored. This paper establishes the first scaling law for MDMs, demonstrating a scaling rate comparable to autoregressive models (ARMs) and a relatively small compute gap. Motivated by their scalability, we train a family of MDMs with up to 1.1 billion (B) parameters to systematically evaluate their performance against ARMs of comparable or larger sizes. Fully leveraging the probabilistic formulation of MDMs, we propose a simple yet effective *unsupervised classifier-free guidance* that effectively exploits large-scale unpaired data, boosting performance for conditional inference. In language understanding, the 1.1B MDM outperforms the 1.1B TinyLlama model trained on the same data across four of eight zero-shot benchmarks. Notably, it achieves competitive math reasoning ability with the 7B Llama-2 model on the GSM8K dataset. In text generation, MDMs with 16 times more pre-training time offer a flexible trade-off against ARMs with the accelerated sampling technique KV-Cache: MDMs match ARMs in performance while being 1.4 times faster during sampling. Moreover, MDMs address challenging tasks for ARMs by effectively handling bidirectional reasoning and adapting to temporal shifts in data. Notably, a 1.1B MDM breaks the *reverse curse* encountered by much larger ARMs with significantly more data and computation, such as 13B Llama-2 and 175B GPT-3. Our code is available at https://github.com/ML-GSAI/SMDM.
Shen Nie, Fengqi Zhu, Tianyu Pang, Qian Liu 0033, Guangtao Zeng, Chongxuan Li
ICLR1
2025 Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
abstract
Discrete diffusion models with absorbing processes have shown promise in language modeling. The key quantities to be estimated are the ratios between the marginal probabilities of two transitive states at all timesteps, called the concrete score. In this paper, we reveal that the concrete score in absorbing diffusion can be expressed as conditional probabilities of clean data, multiplied by a time-dependent scalar in an analytic form. Motivated by this finding, we propose reparameterized absorbing discrete diffusion (RADD), a dedicated diffusion model without time-condition that characterizes the time-independent conditional probabilities. Besides its simplicity, RADD can reduce the number of function evaluations (NFEs) by caching the output of the time-independent network when the noisy sample remains unchanged in a sampling interval, which enables sampling acceleration. Built upon the new perspective of conditional distributions, we further unify absorbing discrete diffusion and any-order autoregressive models (AO-ARMs), showing that the upper bound on the negative log-likelihood for the diffusion model can be interpreted as an expected negative log-likelihood for AO-ARMs. Further, our RADD models achieve SOTA performance among diffusion models on 5 zero-shot language modeling benchmarks (measured by perplexity) at the GPT-2 scale. Our code is available at \url{https://github.com/ML-GSAI/RADD}.
Jingyang Ou, Shen Nie, Fengqi Zhu, Zhenguo Li, Chongxuan Li
ICLR2
2025 Masked Diffusion Models as Energy Minimization
abstract
We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specifically, we prove that three distinct energy formulations—kinetic, conditional kinetic, and geodesic energy—are mathematically equivalent under the structure of MDMs, and that MDMs minimize all three when the mask schedule satisfies a closed-form optimality condition. This unification not only clarifies the theoretical foundations of MDMs, but also motivates practical improvements in sampling. By parameterizing interpolation schedules via Beta distributions, we reduce the schedule design space to a tractable 2D search, enabling efficient post-training tuning without model modification. Experiments on synthetic and real-world benchmarks demonstrate that our energy-inspired schedules outperform hand-crafted baselines, particularly in low-step sampling settings.
Shen Nie, Zijin Feng, Zhenguo Li, Ji-Rong Wen, Chongxuan Li
NeurIPS2
2025 Large Language Diffusion Models
abstract
The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing *LLaDA*, a diffusion model trained from scratch under the pre-training and supervised fine-tuning (SFT) paradigm. LLaDA employs a forward data masking process and a reverse generation process, parameterized by a Transformer to predict masked tokens. It provides a principled generative approach for probabilistic inference by optimizing a likelihood lower bound. Across extensive benchmarks on general tasks, math, code, and so on, LLaDA demonstrates strong *scalability* and performs comparably to our self-constructed ARM baselines. Remarkably, LLaDA 8B is competitive with strong LLMs like LLaMA3 8B in *in-context learning* and, after SFT, exhibits impressive *instruction-following* abilities in case studies such as multi-turn dialogue. Moreover, LLaDA addresses the reversal curse, surpassing GPT-4o in a reversal poem completion task. Our findings show the promise of diffusion models for language modeling at scale and challenge the common assumption that core LLM capabilities discussed above inherently depend on ARMs. Project page and codes: \url{https://ml-gsai.github.io/LLaDA-demo/}.
Shen Nie, Fengqi Zhu, Zebin You, Jingyang Ou, Jun Zhou 0011, Yankai Lin 0001, Ji-Rong Wen, Chongxuan Li
NeurIPS1
2024 The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing
abstract
We present a unified probabilistic formulation for diffusion-based image editing, where a latent variable is edited in a task-specific manner and generally deviates from the corresponding marginal distribution induced by the original stochastic or ordinary differential equation (SDE or ODE). Instead, it defines a corresponding SDE or ODE for editing. In the formulation, we prove that the Kullback-Leibler divergence between the marginal distributions of the two SDEs gradually decreases while that for the ODEs remains as the time approaches zero, which shows the promise of SDE in image editing. Inspired by it, we provide the SDE counterparts for widely used ODE baselines in various tasks including inpainting and image-to-image translation, where SDE shows a consistent and substantial improvement. Moreover, we propose \emph{SDE-Drag} -- a simple yet effective method built upon the SDE formulation for point-based content dragging. We build a challenging benchmark (termed \emph{DragBench}) with open-set natural, art, and AI-generated images for evaluation. A user study on DragBench indicates that SDE-Drag significantly outperforms our ODE baseline, existing diffusion-based methods, and the renowned DragGAN. Our results demonstrate the superiority and versatility of SDE in image editing and push the boundary of diffusion-based editing methods. See the project page \url{https://ml-gsai.github.io/SDE-Drag-demo/} for the code and DragBench dataset.
Shen Nie, Hanzhong Allan Guo, Cheng Lu 0011, Chenyu Zheng, Chongxuan Li
ICLR1
2024 Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations
abstract
Bayesian flow networks (BFNs) iteratively refine the parameters, instead of the samples in diffusion models (DMs), of distributions at various noise levels through Bayesian inference. Owing to its differentiable nature, BFNs are promising in modeling both continuous and discrete data, while simultaneously maintaining fast sampling capabilities. This paper aims to understand and enhance BFNs by connecting them with DMs through stochastic differential equations (SDEs). We identify the linear SDEs corresponding to the noise-addition processes in BFNs, demonstrate that BFN’s regression losses are aligned with denoise score matching, and validate the sampler in BFN as a first-order solver for the respective reverse-time SDE. Based on these findings and existing recipes of fast sampling in DMs, we propose specialized solvers for BFNs that markedly surpass the original BFN sampler in terms of sample quality with a limited number of function evaluations (e.g., 10) on both image and text datasets. Notably, our best sampler achieves an increase in speed of $5\sim20$ times for free.
Shen Nie, Xu Min, Jun Zhou 0011, Chongxuan Li
ICML3
2023 All are Worth Words: A ViT Backbone for Diffusion Models
abstract
Vision transformers (ViT) have shown promise in various vision tasks while the U-Net based on a convolutional neural network (CNN) remains dominant in diffusion models. We design a simple and general ViT-based architecture (named U-ViT) for image generation with diffusion models. U-ViT is characterized by treating all inputs including the time, condition and noisy image patches as tokens and employing long skip connections between shallow and deep layers. We evaluate U-ViT in unconditional and classconditional image generation, as well as text-to-image generation tasks, where U-ViT is comparable if not superior to a CNN-based U-Net of a similar size. In particular, latent diffusion models with U-ViT achieve record-breaking FID scores of 2.29 in class-conditional image generation on ImageNet 256×256, and 5.48 in text-to-image generation on MS-COCO, among methods without accessing large external datasets during the training of generative models. Our results suggest that, for diffusion-based image modeling, the long skip connection is crucial while the down-sampling and upsampling operators in CNN-based U-Net are not always necessary. We believe that U-ViT can provide insights for future research on backbones in diffusion models and benefit generative modeling on large scale cross-modality datasets.
Fan Bao, Shen Nie, Yue Cao 0001, Chongxuan Li, Hang Su 0006, Jun Zhu 0001
CVPR2
2023 One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale
abstract
This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoken models (e.g., Stable Diffusion and DALL-E 2) in representative tasks (e.g., text-to-image generation).
Fan Bao, Shen Nie, Chongxuan Li, Shi Pu 0002, Yaole Wang, Gang Yue, Yue Cao 0001, Hang Su 0006, Jun Zhu 0001
ICML2