EDBT 2026 Demo / reviewers in the wild / expert
Weijian Luo
dblp:253/4234
· DBLP profile ↗
19ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the local constraints: Geometric continuity regularization for image alignment
Yinqi Chen, Yangting Zheng, Peiwen Li, Weijian Luo, Xiang Gao 0015 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | One-Step Diffusion and Flow Distillation Through Implicit Generator MatchingabstractDespite strong performances on many generative tasks, diffusion and flow matching models require a large number of sampling steps to generate high-quality images. This has motivated the community to develop effective methods to distill pre-trained models into more efficient models. In this paper, we present Implicit Generator Matching (IGM), a systematic approach to distill both pre-trained diffusion/flow matching models into one-step generator models, while maintaining almost the same sample generation ability as the original model, as well as being data-free with no need for training images. The key challenge is that the traditional diffusion/flow-matching loss is intractable to distill a teacher diffusion/flow model with an explicitly defined field into a student generator, whose field is defined implicitly. The main breakthrough, our Implicit Gradient Theorem, provides an exact and efficient gradient to directly optimize the student by aligning this implicit field with the teacher's. IGM shows strong empirical performance for one-step generators, setting new standards. On CIFAR10, our diffusion-based SIM achieves an FID score of 2.06, while flow-based FGM sets a flow-model record with a 3.08 FID. Scaling to text-to-image models, SIM distillation of PixArt-$\alpha$α yields a leading 6.42 aesthetic score, surpassing SDXL-TURBO (5.33), and FGM distillation of SD3 achieves a competitive 0.65 GenEval score against multi-step accelerators like Hyper-SD3 (0.63). Zemin Huang, Weijian Luo, Zhengyang Geng, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Self-Guidance: Boosting Flow and Diffusion Generation on Their OwnabstractProper guidance strategies are essential to achieve high-quality generation results without retraining diffusion and flow-based text-to-image models. Existing guidance either requires specific training or strong inductive biases of diffusion model networks, which potentially limits their ability and application scope. Motivated by the observation that artifact outliers can be detected by a significant decline in the density from a noisier to a cleaner noise level, we propose Self-Guidance (SG), which can significantly improve the quality of the generated image by suppressing the generation of low-quality samples. The biggest difference from existing guidance is that SG only relies on the sampling score function of the original diffusion or flow model at different noise levels, with no need for any tricky and expensive guidance-specific training. This makes SG highly flexible to be used in a plug-and-play manner by any diffusion or flow models. We also introduce an efficient variant of SG, named SG-prev, which reuses the output from the immediately previous diffusion step to avoid additional forward passes of the diffusion network. We conduct extensive experiments on text-to-image and text-to-video generation with different architectures, including UNet and transformer models. With open-sourced diffusion models such as Stable Diffusion 3.5 and FLUX, SG exceeds existing algorithms on multiple metrics, including both FID and Human Preference Score. SG-prev also achieves strong results over both the baseline and the SG, with 50 percent more efficiency. Moreover, we find that SG and SG-prev both have a surprisingly positive effect on the generation of physiologically correct human body structures such as hands, faces, and arms, showing their ability to eliminate human body artifacts with minimal efforts. Weijian Luo, Zhiyang Chen 0002, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image GenerationabstractDiffusion and flow matching models have achieved remarkable success in text-to-image generation. However, these models typically rely on the predetermined denoising schedules for all prompts. The multi-step reverse diffusion process can be regarded as a kind of chain-of-thought for generating high-quality images step by step. Therefore, diffusion models should reason for each instance to determine the optimal noise schedule adaptively, achieving high generation quality with sampling efficiency. In this paper, we introduce the Time Prediction Diffusion Model (TPDM) for this. TPDM employs a plug-and-play Time Prediction Module (TPM) that predicts the next noise level based on current latent features at each denoising step. We train the TPM using reinforcement learning to maximize a reward that encourages high final image quality while penalizing excessive denoising steps. With such an adaptive scheduler, TPDM not only generates high-quality images that are aligned closely with human preferences but also adjusts diffusion time and the number of denoising steps on the fly, enhancing both performance and efficiency. With Stable Diffusion 3 Medium architecture, TPDM achieves an aesthetic score of 5.44 and a human preference score (HPS) of 29.59, while using around 50% fewer denoising steps to achieve better performance. Zilyu Ye, Zhiyang Chen 0002, Zemin Huang, Weijian Luo, Guo-Jun Qi |
CVPR | 5 |
| 2025 | Consistency Models Made EasyabstractConsistency models (CMs) offer faster sampling than traditional diffusion models, but their training is resource-intensive. For example, as of 2024, training a state-of-the-art CM on CIFAR-10 takes one week on 8 GPUs. In this work, we propose an effective scheme for training CMs that largely improves the efficiency of building such models. Specifically, by expressing CM trajectories via a particular differential equation, we argue that diffusion models can be viewed as a special case of CMs. We can thus fine-tune a consistency model starting from a pretrained diffusion model and progressively approximate the full consistency condition to stronger degrees over the training process. Our resulting method, which we term Easy Consistency Tuning (ECT), achieves vastly reduced training times while improving upon the quality of previous methods: for example, ECT achieves a 2-step FID of 2.73 on CIFAR10 within 1 hour on a single A100 GPU, matching Consistency Distillation trained for hundreds of GPU hours. Owing to this computational efficiency, we investigate the scaling laws of CMs under ECT, showing that they obey the classic power law scaling, hinting at their ability to improve efficiency and performance at larger scales. Our [code](https://github.com/locuslab/ect) is available. Zhengyang Geng, Ashwini Pokle, Weijian Luo, Justin Lin, J. Zico Kolter |
ICLR | 3 |
| 2025 | David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-trainingabstractWe propose Diff-Instruct(DI), a data-efficient post-training approach to one-step text-to-image generative models to improve its human preferences without requiring image data. Our method frames alignment as online reinforcement learning from human feedback (RLHF), which optimizes a human reward function while regularizing the generator to stay close to a reference diffusion process. Unlike traditional RLHF approaches, which rely on the KL divergence for regularization, we introduce a novel score-based divergence regularization that substantially improves performance. Although such a score-based RLHF objective seems intractable when optimizing, we derive a strictly equivalent tractable loss function in theory that can efficiently compute its gradient for optimizations. Building upon this framework, we train DI-SDXL-1step, a 1-step text-to-image model based on Stable Diffusion-XL (2.6B parameters), capable of generating 1024x1024 resolution images in a single step. The 2.6B DI-SDXL-1step model outperforms the 12B FLUX-dev model in ImageReward, PickScore, and CLIP score on the Parti prompts benchmark while using only 1.88% of the inference time. This result strongly supports the thought that with proper post-training, the small one-step model is capable of beating huge multi-step models. We will open-source our industry-ready model to the community. Weijian Luo, Colin Zhang, Debing Zhang, Zhengyang Geng |
ICML | 1 |
| 2025 | Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image GenerationabstractThis paper addresses the challenge of achieving high-quality and fast image generation that aligns with complex human preferences. While recent advancements in diffusion models and distillation have enabled rapid generation, the effective integration of reward feedback for improved abilities like controllability and preference alignment remains a key open problem.
Existing reward-guided post-training approaches targeting accelerated few-step generation often deem diffusion distillation losses indispensable.
However, in this paper, we identify an interesting yet fundamental paradigm shift: as conditions become more specific, well-designed reward functions emerge as the primary driving force in training strong, few-step image generative models. Motivated by this insight, we introduce Reward-Instruct, a novel and surprisingly simple reward-centric approach for converting pre-trained base diffusion models into reward-enhanced few-step generators. Unlike existing methods, Reward-Instruct does not rely on expensive yet tricky diffusion distillation losses. Instead, it iteratively updates the few-step generator's parameters by directly sampling from a reward-tilted parameter distribution. Such a training approach entirely bypasses the need for expensive diffusion distillation losses, making it favorable to scale in high image resolutions. Despite its simplicity, Reward-Instruct yields surprisingly strong performance. Our extensive experiments on text-to-image generation have demonstrated that Reward-Instruct achieves state-of-the-art results in visual quality and quantitative metrics compared to distillation-reliant methods, while also exhibiting greater robustness to the choice of reward function. Yihong Luo, Tianyang Hu 0001, Weijian Luo, Kenji Kawaguchi, Jing Tang 0004 |
NeurIPS | 3 |
| 2025 | Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence InstructionabstractIn this paper, we unify more than 10 existing one-step diffusion distillation approaches, such as Diff-Instruct, DMD, SIM, SiD, $f$-distill, etc, inside a theory-driven framework which we name the \textbf{\emph{Uni-Instruct}}. Uni-Instruct is motivated by our proposed diffusion expansion theory of the $f$-divergence family. Then we introduce key theories that overcome the intractability issue of the original expanded $f$-divergence, resulting in an equivalent yet tractable loss that effectively trains one-step diffusion models by minimizing the expanded $f$-divergence family. The novel unification introduced by Uni-Instruct not only offers new theoretical contributions that help understand existing approaches from a high-level perspective but also leads to state-of-the-art one-step diffusion generation performances. On the CIFAR10 generation benchmark, Uni-Instruct achieves record-breaking Frechet Inception Distance (FID) values of \textbf{\emph{1.46}} for unconditional generation and \textbf{\emph{1.38}} for conditional generation. On the ImageNet-$64\times 64$ generation benchmark, Uni-Instruct achieves a new SoTA one-step diffusion FID value of \textbf{\emph{1.06}}, which outperforms its 79-step teacher diffusion with a significant improvement margin of 1.29 (1.06 vs 2.35). We also apply Uni-Instruct on broader tasks like text-to-3D generation. For text-to-3D generation, Uni-Instruct gives decent results, which slightly outperforms previous methods, such as SDS and VSD, in terms of both generation quality and diversity. Both the solid theoretical and empirical contributions of Uni-Instruct will potentially help future studies on one-step diffusion distillation and knowledge transferring of diffusion models. Weimin Bai, Colin Zhang, Debing Zhang, Weijian Luo, He Sun 0010 |
NeurIPS | 5 |
| 2025 | Deep hashing with mutual information: A comprehensive strategy for image retrieval
Yinqi Chen, Zhiyi Lu, Yangting Zheng, Peiwen Li, Weijian Luo |
Expert Syst. Appl. | 5 |
| 2025 | Content-activating for artistic style transfer with ambiguous sketchy content image
Yinqi Chen, Yangting Zheng, Peiwen Li, Weijian Luo |
Neurocomputing | 4 |
| 2025 | Liberating the expressive capacity of deep hashing for image retrieval
Yinqi Chen, Yangting Zheng, Peiwen Li, Weijian Luo |
Inf. Process. Manag. | 4 |
| 2025 | Deep Hashing With Walsh Domain for Multi-Label Image RetrievalabstractThe existing deep hashing methods for image retrieval typically modeling hash-coding layer in real number space. However, these methods frequently overlook the intrinsic information loss that occurs during the hash-coding process, as the hash layer performs two tasks simultaneously: spatial transformation and dimensionality reduction. Especially in multi-label image retrieval, the exponential increase in the number of combinations with the labels further amplifies the information loss in Hamming space. Consequently, the efficiency of the hash-coding is unsatisfactory. To mitigate this limitation, we introduce a novel approach termed WalshHash, which is grounded in the principles of Walsh transformation in signal processing. Unlike conventional techniques, WalshHash formulates the hash-coding layer as a filtering process based on Kolmogorov-Arnold Networks (KANs) in the Walsh domain accompanied by constraint loss functions on multiple domains. It ensures the dimensionality reduction in the Walsh domain can be effectively projected onto the real number domain with minimal information loss, because the Walsh space encapsulates the critical information components. As a result, WalshHash demonstrates superior performance in multi-label image retrieval compared to State-of-the-Art (SOTA) methods. Yinqi Chen, Peiwen Li, Yangting Zheng, Weijian Luo, Xiang Gao 0015 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Variational Schrödinger Diffusion ModelsabstractSchrödinger bridge (SB) has emerged as the go-to method for optimizing transportation plans in diffusion models. However, SB requires estimating the intractable forward score functions, inevitably resulting in the (costly) implicit training loss based on simulated trajectories. To improve the scalability while preserving efficient transportation plans, we leverage variational inference to linearize the forward score functions (variational scores) of SB and restore *simulation-free* properties in training backward scores. We propose the variational Schrödinger diffusion model (VSDM), where the forward process is a multivariate diffusion and the variational scores are adaptively optimized for efficient transport. Theoretically, we use stochastic approximation to prove the convergence of the variational scores and show the convergence of the adaptively generated samples based on the optimal variational scores. Empirically, we test the algorithm in simulated examples and observe that VSDM is efficient in generations of anisotropic shapes and yields straighter sample trajectories compared to the single-variate diffusion. We also verify the scalability of the algorithm in real-world data and achieve competitive unconditional generation performance in CIFAR10 and conditional generation in time series modeling. Notably, VSDM no longer depends on warm-up initializations required by SB. Wei Deng 0002, Weijian Luo, Yixin Tan, Marin Bilos, Yuriy Nevmyvaka, Ricky T. Q. Chen |
ICML | 2 |
| 2024 | One-Step Diffusion Distillation through Score Implicit MatchingabstractDespite their strong performances on many generative tasks, diffusion models require a large number of sampling steps in order to generate realistic samples. This has motivated the community to develop effective methods to distill pre-trained diffusion models into more efficient models, but these methods still typically require few-step inference or perform substantially worse than the underlying model. In this paper, we present Score Implicit Matching (SIM) a new approach to distilling pre-trained diffusion models into single-step generator models, while maintaining almost the same sample generation ability as the original model as well as being data-free with no need of training samples for distillation. The method rests upon the fact that, although the traditional score-based loss is intractable to minimize for generator models, under certain conditions we \emph{can} efficiently compute the \emph{gradients} for a wide class of score-based divergences between a diffusion model and a generator. SIM shows strong empirical performances for one-step generators: on the CIFAR10 dataset, it achieves an FID of 2.06 for unconditional generation and 1.96 for class-conditional generation. Moreover, by applying SIM to a leading transformer-based diffusion model, we distill a single-step generator for text-to-image (T2I) generation that attains an aesthetic score of 6.42 with no performance decline over the original multi-step counterpart, clearly outperforming the other one-step generators including SDXL-TURBO of 5.33, SDXL-LIGHTNING of 5.34 and HYPER-SDXL of 5.85. We will release this industry-ready one-step transformer-based T2I generator along with this paper. Weijian Luo, Zemin Huang, Zhengyang Geng, J. Zico Kolter, Guo-Jun Qi |
NeurIPS | 1 |
| 2023 | Diff-Instruct: A Universal Approach for Transferring Knowledge From Pre-trained Diffusion ModelsabstractDue to the ease of training, ability to scale, and high sample quality, diffusion models (DMs) have become the preferred option for generative modeling, with numerous pre-trained models available for a wide variety of datasets. Containing intricate information about data distributions, pre-trained DMs are valuable assets for downstream applications. In this work, we consider learning from pre-trained DMs and transferring their knowledge to other generative models in a data-free fashion. Specifically, we propose a general framework called Diff-Instruct to instruct the training of arbitrary generative models as long as the generated samples are differentiable with respect to the model parameters. Our proposed Diff-Instruct is built on a rigorous mathematical foundation where the instruction process directly corresponds to minimizing a novel divergence we call Integral Kullback-Leibler (IKL) divergence. IKL is tailored for DMs by calculating the integral of the KL divergence along a diffusion process, which we show to be more robust in comparing distributions with misaligned supports. We also reveal non-trivial connections of our method to existing works such as DreamFusion \citep{poole2022dreamfusion}, and generative adversarial training. To demonstrate the effectiveness and universality of Diff-Instruct, we consider two scenarios: distilling pre-trained diffusion models and refining existing GAN models. The experiments on distilling pre-trained diffusion models show that Diff-Instruct results in state-of-the-art single-step diffusion-based models. The experiments on refining GAN models show that the Diff-Instruct can consistently improve the pre-trained generators of GAN models across various settings. Our official code is released through \url{https://github.com/pkulwj1994/diff_instruct}. Weijian Luo, Tianyang Hu 0001, Zhenguo Li |
NeurIPS | 1 |
| 2023 | Entropy-based Training Methods for Scalable Neural Implicit SamplersabstractEfficiently sampling from un-normalized target distributions is a fundamental problem in scientific computing and machine learning. Traditional approaches such as Markov Chain Monte Carlo (MCMC) guarantee asymptotically unbiased samples from such distributions but suffer from computational inefficiency, particularly when dealing with high-dimensional targets, as they require numerous iterations to generate a batch of samples. In this paper, we introduce an efficient and scalable neural implicit sampler that overcomes these limitations. The implicit sampler can generate large batches of samples with low computational costs by leveraging a neural transformation that directly maps easily sampled latent vectors to target samples without the need for iterative procedures. To train the neural implicit samplers, we introduce two novel methods: the KL training method and the Fisher training method. The former method minimizes the Kullback-Leibler divergence, while the latter minimizes the Fisher divergence between the sampler and the target distributions. By employing the two training methods, we effectively optimize the neural implicit samplers to learn and generate from the desired target distribution. To demonstrate the effectiveness, efficiency, and scalability of our proposed samplers, we evaluate them on three sampling benchmarks with different scales. These benchmarks include sampling from 2D targets, Bayesian inference, and sampling from high-dimensional energy-based models (EBMs). Notably, in the experiment involving high-dimensional EBMs, our sampler produces samples that are comparable to those generated by MCMC-based methods while being more than 100 times more efficient, showcasing the efficiency of our neural sampler. Besides the theoretical contributions and strong empirical performances, the proposed neural samplers and corresponding training methods will shed light on further research on developing efficient samplers for various applications beyond the ones explored in this study. Weijian Luo |
NeurIPS | 1 |
| 2023 | SA-Solver: Stochastic Adams Solver for Fast Sampling of Diffusion ModelsabstractDiffusion Probabilistic Models (DPMs) have achieved considerable success in generation tasks. As sampling from DPMs is equivalent to solving diffusion SDE or ODE which is time-consuming, numerous fast sampling methods built upon improved differential equation solvers are proposed. The majority of such techniques consider solving the diffusion ODE due to its superior efficiency. However, stochastic sampling could offer additional advantages in generating diverse and high-quality data. In this work, we engage in a comprehensive analysis of stochastic sampling from two aspects: variance-controlled diffusion SDE and linear multi-step SDE solver. Based on our analysis, we propose SA-Solver, which is an improved efficient stochastic Adams method for solving diffusion SDE to generate data with high quality. Our experiments show that SA-Solver achieves: 1) improved or comparable performance compared with the existing state-of-the-art (SOTA) sampling methods for few-step sampling; 2) SOTA FID on substantial benchmark datasets under a suitable number of function evaluations (NFEs). Shuchen Xue, Mingyang Yi, Weijian Luo, Zhenguo Li, Zhiming Ma |
NeurIPS | 3 |
| 2023 | Enhancing Adversarial Robustness via Score-Based OptimizationabstractAdversarial attacks have the potential to mislead deep neural network classifiers by introducing slight perturbations. Developing algorithms that can mitigate the effects of these attacks is crucial for ensuring the safe use of artificial intelligence. Recent studies have suggested that score-based diffusion models are effective in adversarial defenses. However, existing diffusion-based defenses rely on the sequential simulation of the reversed stochastic differential equations of diffusion models, which are computationally inefficient and yield suboptimal results. In this paper, we introduce a novel adversarial defense scheme named ScoreOpt, which optimizes adversarial samples at test-time, towards original clean data in the direction guided by score-based priors. We conduct comprehensive experiments on multiple datasets, including CIFAR10, CIFAR100 and ImageNet. Our experimental results demonstrate that our approach outperforms existing adversarial defenses in terms of both robustness performance and inference speed. Weijian Luo |
NeurIPS | 2 |
| 2019 | Recalibration Of Offshore Chlorophyll Content Based On Virtual Satellite ConstellationabstractIn order to resolve the conflict between the temporal and spatial resolution of remote sensing, we plan to construct a virtual satellite constellation by combining multiple satellites with high spatial resolution sensors to improve the temporal resolution. By using satellite-land synchronized data obtained in 4 different offshore regions near China between 2015 and 2017, we evaluated the remote sensing chlorophyll inverse algorithms of different sensors, and further determined a remote sensing inverse model, which is suitable for data of multiple medium-high resolution satellites. By using GOMS/GOCI chlorophyll concentration product as reference, we first projected the chlorophyll concentration values inverted from satellites HJ-1A, HJ-1B, GF-1,GF-4, Landsat 5, Landsat 8 to a 50-m grid by applying flux conservation resampling method. We further applied a recalibration method to obtain recalibration equations, which make the chlorophyll concentration values inverted from different sensors comparable. Miaofen Huang, Xufeng Xing, Weijian Luo, Zhonglin Wang |
IGARSS | 3 |