Weimin Bai

dblp:121/0917 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 67% 3D vision · 19% Deep learning architectures and training · 9%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.632026
Blind Inversion Using Latent Diffusion Priors · IEEE Trans. Image Process. 2026
Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction · NeurIPS 2025
An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted Observations · NeurIPS 2024
Image and video processing
image restoration
1.822026
Blind Inversion Using Latent Diffusion Priors · IEEE Trans. Image Process. 2026
An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted Observations · NeurIPS 2024
Computer vision › 3D vision
3d reconstruction
1.012026
Blind Inversion Using Latent Diffusion Priors · IEEE Trans. Image Process. 2026
Machine learning › Generative modeling › diffusion model › diffusion prior
latent diffusion prior
1.012026
Blind Inversion Using Latent Diffusion Priors · IEEE Trans. Image Process. 2026
Image and video processing › image restoration › image deblurring
blind image deblurring
1.012026
Blind Inversion Using Latent Diffusion Priors · IEEE Trans. Image Process. 2026
Image and video processing › image restoration › multi-task image restoration
denoising and deblurring
0.812024
An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted Observations · NeurIPS 2024
Machine learning › Deep learning architectures and training
training optimization
0.512021
Learning Deeper Non-Monotonic Networks by Softly Transferring Solution Space · IJCAI 2021
Machine learning › Learning theory › probability metric
f-divergences
0.312025
Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

expectation-maximization · 3.5inverse problem solving · 2.0diffusion model training · 1.5knowledge distillation · 0.9f-divergence expansion · 0.9variational bayesian learning · 0.5soft transfer · 0.5rectified concrete gate · 0.5
YearPublicationVenuePosition
2026 Blind Inversion Using Latent Diffusion Priors
abstract
Diffusion models have emerged as powerful tools for solving inverse problems due to their exceptional ability to model complex prior distributions. However, existing methods predominantly assume known forward operators (i.e., non-blind), limiting their applicability in practical settings where acquiring such operators is costly. Additionally, many current approaches rely on pixel-space diffusion models, leaving the potential of more powerful latent diffusion models (LDMs) underexplored. In this paper, we introduce LatentDEM, an innovative technique that addresses more challenging blind inverse problems using latent diffusion priors. At the core of our method is solving blind inverse problems within an iterative Expectation-Maximization (EM) framework: (1) the E-step recovers clean images from corrupted observations using LDM priors and a known forward model, and (2) the M-step estimates the forward operator based on the recovered images. Additionally, we propose two novel optimization techniques tailored for LDM priors and EM frameworks, yielding more accurate and efficient blind inversion results. As a general framework, LatentDEM supports both linear and non-linear inverse problems. Beyond common 2D image restoration tasks, it enables new capabilities in non-linear 3D inverse rendering problems. We validate LatentDEM's performance on representative 2D blind deblurring and 3D pose-free sparse-view reconstruction tasks, demonstrating its superior efficacy over prior arts. The project page can be found at https://ai4imaging.github.io/latentdem/.
Weimin Bai, Wenzheng Chen, He Sun 0010
IEEE Trans. Image Process.1
2025 Learning Diffusion Model from Noisy Measurement using Principled Expectation-Maximization Method
abstract
Diffusion models have demonstrated exceptional ability in modeling complex image distributions, making them versatile plug-and-play priors for solving imaging inverse problems. However, their reliance on large-scale clean datasets for training limits their applicability in scenarios where acquiring clean data is costly or impractical. Recent approaches have attempted to learn diffusion models directly from corrupted measurements, but these methods either lack theoretical convergence guarantees or are restricted to specific types of data corruption. In this paper, we propose a principled expectation-maximization (EM) framework that iteratively learns diffusion models from noisy data with arbitrary corruption types. Our framework employs a plug-and-play Monte Carlo method to accurately estimate clean images from noisy measurements, followed by training the diffusion model using the reconstructed images. This process alternates between estimation and training until convergence. We evaluate the performance of our method across various imaging tasks, including inpainting, denoising, and deblurring. Experimental results demonstrate that our approach enables the learning of high-fidelity diffusion priors from noisy data, significantly enhancing reconstruction quality in imaging inverse problems.
Weimin Bai, Weiheng Tang, Enze Ye, Wenzheng Chen, He Sun 0010
ICASSP1
2025 Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
abstract
In this paper, we unify more than 10 existing one-step diffusion distillation approaches, such as Diff-Instruct, DMD, SIM, SiD, $f$-distill, etc, inside a theory-driven framework which we name the \textbf{\emph{Uni-Instruct}}. Uni-Instruct is motivated by our proposed diffusion expansion theory of the $f$-divergence family. Then we introduce key theories that overcome the intractability issue of the original expanded $f$-divergence, resulting in an equivalent yet tractable loss that effectively trains one-step diffusion models by minimizing the expanded $f$-divergence family. The novel unification introduced by Uni-Instruct not only offers new theoretical contributions that help understand existing approaches from a high-level perspective but also leads to state-of-the-art one-step diffusion generation performances. On the CIFAR10 generation benchmark, Uni-Instruct achieves record-breaking Frechet Inception Distance (FID) values of \textbf{\emph{1.46}} for unconditional generation and \textbf{\emph{1.38}} for conditional generation. On the ImageNet-$64\times 64$ generation benchmark, Uni-Instruct achieves a new SoTA one-step diffusion FID value of \textbf{\emph{1.06}}, which outperforms its 79-step teacher diffusion with a significant improvement margin of 1.29 (1.06 vs 2.35). We also apply Uni-Instruct on broader tasks like text-to-3D generation. For text-to-3D generation, Uni-Instruct gives decent results, which slightly outperforms previous methods, such as SDS and VSD, in terms of both generation quality and diversity. Both the solid theoretical and empirical contributions of Uni-Instruct will potentially help future studies on one-step diffusion distillation and knowledge transferring of diffusion models.
Weimin Bai, Colin Zhang, Debing Zhang, Weijian Luo, He Sun 0010
NeurIPS2
2025 A temporal kernel for time series forecasting
Weimin Bai
Knowl. Based Syst.1
2024 An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted Observations
abstract
Diffusion models excel in solving imaging inverse problems due to their ability to model complex image priors. However, their reliance on large, clean datasets for training limits their practical use where clean data is scarce. In this paper, we propose EMDiffusion, an expectation-maximization (EM) approach to train diffusion models from corrupted observations. Our method alternates between reconstructing clean images from corrupted data using a known diffusion model (E-step) and refining diffusion model weights based on these reconstructions (M-step). This iterative process leads the learned diffusion model to gradually converge to a local optimum, that is, to approximate the true clean data distribution. We validate our method through extensive experiments on diverse computational imaging tasks, including random inpainting, denoising, and deblurring, achieving new state-of-the-art performance.
Weimin Bai, Wenzheng Chen, He Sun 0010
NeurIPS1
2021 Learning Deeper Non-Monotonic Networks by Softly Transferring Solution Space
abstract
Different from popular neural networks using quasiconvex activations, non-monotonic networks activated by periodic nonlinearities have emerged as a more competitive paradigm, offering revolutionary benefits: 1) compactly characterizing high-frequency patterns; 2) precisely representing high-order derivatives. Nevertheless, they are also well-known for being hard to train, due to easily over-fitting dissonant noise and only allowing for tiny architectures (shallower than 5 layers). The fundamental bottleneck is that the periodicity leads to many poor and dense local minima in solution space. The direction and norm of gradient oscillate continually during error backpropagation. Thus non-monotonic networks are prematurely stuck in these local minima, and leave out effective error feedback. To alleviate the optimization dilemma, in this paper, we propose a non-trivial soft transfer approach. It smooths their solution space close to that of monotonic ones in the beginning, and then improve their representational properties by transferring the solutions from the neural space of monotonic neurons to the Fourier space of non-monotonic neurons as the training continues. The soft transfer consists of two core components: 1) a rectified concrete gate is constructed to characterize the state of each neuron; 2) a variational Bayesian learning framework is proposed to dynamically balance the empirical risk and the intensity of transfer. We provide comprehensive empirical evidence showing that the soft transfer not only reduces the risk of non-monotonic networks on over-fitting noise, but also helps them scale to much deeper architectures (more than 100 layers) achieving the new state-of-the-art performance.
Zheng-Fan Wu, Hui Xue 0002, Weimin Bai
IJCAI3