Lijiang Li

dblp:310/1558 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0003-3413-3429ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 PFDiff: Training-Free Acceleration of Diffusion Models Combining Past and Future Scores
abstract
Diffusion Probabilistic Models (DPMs) have shown remarkable potential in image generation, but their sampling efficiency is hindered by the need for numerous denoising steps. Most existing solutions accelerate the sampling process by proposing fast ODE solvers. However, the inevitable discretization errors of the ODE solvers are significantly magnified when the number of function evaluations (NFE) is fewer. In this work, we propose PFDiff, a novel training-free and orthogonal timestep-skipping strategy, which enables existing fast ODE solvers to operate with fewer NFE. Specifically, PFDiff initially utilizes score replacement from past time steps to predict a springboard. Subsequently, it employs this ``springboard" along with foresight updates inspired by Nesterov momentum to rapidly update current intermediate states. This approach effectively reduces unnecessary NFE while correcting for discretization errors inherent in first-order ODE solvers. Experimental results demonstrate that PFDiff exhibits flexible applicability across various pre-trained DPMs, particularly excelling in conditional DPMs and surpassing previous state-of-the-art training-free methods. For instance, using DDIM as a baseline, we achieved 16.46 FID (4 NFE) compared to 138.81 FID with DDIM on ImageNet 64x64 with classifier guidance, and 13.06 FID (10 NFE) on Stable Diffusion with 7.5 guidance scale. Code is available at https://github.com/onefly123/PFDiff.
Guangyi Wang, Yuren Cai, Lijiang Li, Wei Peng 0009, Songzhi Su
ICLR3
2025 Diffusion Sampling Correction via Approximately 10 Parameters
abstract
While powerful for generation, Diffusion Probabilistic Models (DPMs) face slow sampling challenges, for which various distillation-based methods have been proposed. However, they typically require significant additional training costs and model parameter storage, limiting their practicality. In this work, we propose **P**CA-based **A**daptive **S**earch (PAS), which optimizes existing solvers for DPMs with minimal additional costs. Specifically, we first employ PCA to obtain a few basis vectors to span the high-dimensional sampling space, which enables us to learn just a set of coordinates to correct the sampling direction; furthermore, based on the observation that the cumulative truncation error exhibits an ``S"-shape, we design an adaptive search strategy that further enhances the sampling efficiency and reduces the number of stored parameters to approximately 10. Extensive experiments demonstrate that PAS can significantly enhance existing fast solvers in a plug-and-play manner with negligible costs. E.g., on CIFAR10, PAS optimizes DDIM's FID from 15.69 to 4.37 (NFE=10) using only **12 parameters and sub-minute training** on a single A100 GPU. Code is available at https://github.com/onefly123/PAS.
Guangyi Wang, Wei Peng 0009, Lijiang Li, Yuren Cai, Songzhi Su
ICML3
2025 Breaking Static Barriers: Dynamic Post-Training Quantization for Diffusion Models
abstract
Current Post-Training Quantization (PTQ) schemes have been extensively studied for traditional convolutional neural networks and language models; however, PTQ application in diffusion models has shown significant performance degradation due to static settings of PTQ. Existing methods only uniformly and statically sample during each denoising step to construct calibration sets, neglecting the different importance of different steps in diffusion models. Furthermore, diffusion models exhibit a large number of activations with skewed distributions, and maintaining a static zero-point during the reconstruction process causes the model to converge only to local optima. To solve these limitations, it is necessary to dynamically design calibration dataset construction methods for different quantization scenarios and develop specialized optimization strategies tailored to specific activation distributions. Thus we proposed a unified framework, termed Dynamic PTQ, to achieve the aforementioned purposes. The framework first applies an evolutionary search algorithm to dynamically construct calibration sets for different quantization scenarios. Then, we design a dynamic zero-point update strategy for the quantizer, significantly reducing the loss during the reconstruction process. Extensive experiments demonstrate that our method outperforms current PTQ methods for diffusion models in generating high-quality samples. In particular, for the LSUN-bedrooms 256×256 task, our method quantizes the corresponding full-precision LDM-4 to W4A6 with only a 0.84 increase in FID.
Huixia Li, Lijiang Li, Xiawu Zheng, Yuexiao Ma, Jie Wu 0001, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001
IJCNN3
2025 VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
abstract
With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience high latency when generating the first audio token during streaming, which poses a significant bottleneck for deployment. To address this issue, we propose VITA-Audio, an end-to-end large speech model with fast audio-text token generation. Specifically, we introduce a lightweight Multiple Cross-modal Token Prediction (MCTP) module that efficiently generates multiple audio tokens within a single model forward pass, which not only accelerates the inference but also significantly reduces the latency for generating the first audio in streaming scenarios. In addition, a four-stage progressive training strategy is explored to achieve model acceleration with minimal loss of speech quality. To our knowledge, VITA-Audio is the first multi-modal large language model capable of generating audio output during the first forward pass, enabling real-time conversational capabilities with minimal latency. VITA-Audio is fully reproducible and is trained on open-source data only. Experimental results demonstrate that our model achieves an inference speedup of 3~5x at the 7B parameter scale, but also significantly outperforms open-source models of similar model size on multiple benchmarks for automatic speech recognition (ASR), text-to-speech (TTS), and spoken question answering (SQA) tasks.
Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Jian Li 0062, Jinlong Peng, Haoyu Cao 0001, Ke Li 0015, Rongrong Ji, Xing Sun 0001
NeurIPS5
2025 NADM: Noise-Aware Diffusion Model for Landscape Painting Video Generation
abstract
Landscape painting is a gem of cultural and artistic heritage that showcases the splendor of nature through the deep observations and imaginations of its painters. Limited by traditional techniques, these artworks were confined to static imagery in ancient times, leaving the dynamism of landscapes and the subtleties of artistic sentiment to the viewer's imagination. Recently, emerging text-to-video (T2V) diffusion methods have shown significant promise in video generation, providing hope for the creation of dynamic landscape paintings. However, current T2V methods focus on generating natural videos, emphasizing the capture of details and the authenticity of physical laws. In contrast, landscape painting videos emphasize the overall dynamic aesthetic. Besides, challenges, such as the lack of specific datasets, the intricacy of artistic styles, and the creation of extensive, high-quality videos pose difficulties for these models in generating landscape painting videos. In this article, we propose landscape painting videos-high definition (LPV-HD), a novel T2V dataset for landscape painting videos, and noise-aware diffusion model (NADM), a T2V model that utilizes Stable Diffusion. Specifically, we present a motion module featuring a dual attention mechanism to capture the dynamic transformations of landscape imageries, alongside a noise adapter to leverage unsupervised contrastive learning in the latent space to ensure the overall beauty of the landscape painting video. Following the generation of keyframes, we employ optical flow for frame interpolation to enhance video smoothness. Our method not only retains the essence of the landscape painting imageries but also achieves dynamic transitions, significantly advancing the field of artistic video generation. Source code and dataset are available at https://github.com/llzlh21/NADM.
Ding-Ming Liu, Shao-Wei Li, Ruo-Yan Zhou, Lili Liang, Yongguan Hong, Yuan-Ze Zeng, Xiang Chang, Lijiang Li, Tianshuo Xu, Fei Chao 0001, Changjing Shang, Qiang Shen 0001
IEEE Trans. Cybern.8
2025 Self-Organizing Type-2 Fuzzy Double Loop Recurrent Neural Network for Uncertain Nonlinear System Control
abstract
Nonlinear systems, such as robotic systems, play an increasingly important role in our modern daily life and have become more dominant in many industries; however, robotic control still faces various challenges due to diverse and unstructured work environments. This article proposes a double-loop recurrent neural network (DLRNN) with the support of a Type-2 fuzzy system and a self-organizing mechanism for improved performance in nonlinear dynamic robot control. The proposed network has a double-loop recurrent structure, which enables better dynamic mapping. In addition, the network combines a Type-2 fuzzy system with a double-loop recurrent structure to improve the ability to deal with uncertain environments. To achieve an efficient system response, a self-organizing mechanism is proposed to adaptively adjust the number of layers in a DLRNN. This work integrates the proposed network into a conventional sliding mode control (SMC) system to theoretically and empirically prove its stability. The proposed system is applied to a three-joint robot manipulator, leading to a comparative study that considers several existing control approaches. The experimental results confirm the superiority of the proposed system and its effectiveness and robustness in response to various external system disturbances.
Lijiang Li, Xiang Chang, Fei Chao 0001, Chih-Min Lin, Tuan-Tu Huynh, Longzhi Yang, Changjing Shang, Qiang Shen 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive Optimization
abstract
PSNR-oriented models are a critical class of super-resolution models with applications across various fields. However, these models tend to generate over-smoothed images, a problem that has been analyzed previously from the perspectives of models or loss functions, but without taking into account the impact of data properties. In this paper, we present a novel phenomenon that we term the center-oriented optimization (COO) problem, where a model's output converges towards the center point of similar high-resolution images, rather than towards the ground truth. We demonstrate that the strength of this problem is related to the uncertainty of data, which we quantify using entropy. We prove that as the entropy of high-resolution images increases, their center point will move further away from the clean image distribution, and the model will generate over-smoothed images. Implicitly optimizing the COO problem, perceptual-driven approaches such as perceptual loss, model structure optimization, or GAN-based methods can be viewed. We propose an explicit solution to the COO problem, called Detail Enhanced Contrastive Loss (DECLoss). DECLoss utilizes the clustering property of contrastive learning to directly reduce the variance of the potential high-resolution distribution and thereby decrease the entropy. We evaluate DECLoss on multiple super-resolution benchmarks and demonstrate that it improves the perceptual quality of PSNR-oriented models. Moreover, when applied to GAN-based methods, such as RaGAN, DECLoss helps to achieve state-of-the-art performance, such as 0.093 LPIPS with 24.51 PSNR on 4× downsampled Urban100, validating the effectiveness and generalization of our approach.
Tianshuo Xu, Lijiang Li, Peng Mi, Xiawu Zheng, Fei Chao 0001, Rongrong Ji, Yonghong Tian 0001, Qiang Shen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Meta Architecture for Point Cloud Analysis
abstract
Recent advances in 3D point cloud analysis bring a diverse set of network architectures to the field. However, the lack of a unified framework to interpret those networks makes any systematic comparison, contrast, or analysis challenging, and practically limits healthy development of the field. In this paper, we take the initiative to explore and propose a unified framework called PointMeta, to which the popular 3D point cloud analysis approaches could fit. This brings three benefits. First, it allows us to compare different approaches in a fair manner, and use quick experiments to verify any empirical observations or assumptions summarized from the comparison. Second, the big picture brought by PointMeta enables us to think across different components, and revisit common beliefs and key design decisions made by the popular approaches. Third, based on the learnings from the previous two analyses, by doing simple tweaks on the existing approaches, we are able to derive a basic building block, termed PointMetaBase. It shows very strong performance in efficiency and effectiveness through extensive experiments on challenging benchmarks, and thus verifies the necessity and benefits of high-level interpretation, contrast, and comparison like PointMeta. In particular, PointMetaBase surpasses the previous state-of-the-art method by 0.7%/1.4/%2.1% mIoU with only 2%/11%/13% of the computation cost on the S3DIS datasets. The code and models are available at https://github.com/linhaojia13/PointMetaBase.
Haojia Lin, Xiawu Zheng, Lijiang Li, Fei Chao 0001, Shanshan Wang 0002, Yan Wang 0059, Yonghong Tian 0001, Rongrong Ji
CVPR3
2023 AutoDiffusion: Training-Free Optimization of Time Steps and Architectures for Automated Diffusion Model Acceleration
abstract
Diffusion models are emerging expressive generative models, in which a large number of time steps (inference steps) are required for a single image generation. To accelerate such tedious process, reducing steps uniformly is considered as an undisputed principle of diffusion models. We consider that such a uniform assumption is not the optimal solution in practice; i.e., we can find different optimal time steps for different models. Therefore, we propose to search the optimal time steps sequence and compressed model architecture in a unified framework to achieve effective image generation for diffusion models without any further training. Specifically, we first design a unified search space that consists of all possible time steps and various architectures. Then, a two stage evolutionary algorithm is introduced to find the optimal solution in the designed search space. To further accelerate the search process, we employ FID score between generated and real samples to estimate the performance of the sampled examples. As a result, the proposed method is (i).training-free, obtaining the optimal time steps and model architecture without any training process; (ii). orthogonal to most advanced diffusion samplers and can be integrated to gain better sample quality. (iii). generalized, where the searched time steps and architectures can be directly applied on different diffusion models with the same guidance scale. Experimental results show that our method achieves excellent performance by using only a few time steps, e.g. 17.86 FID score on ImageNet 64 × 64 with only four steps, compared to 138.66 with DDIM.
Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu 0032, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001, Rongrong Ji
ICCV1
2023 Model compression optimized neural network controller for nonlinear systems
Lijiang Li, Sheng-Lin Zhou, Fei Chao 0001, Xiang Chang, Longzhi Yang, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.1
2022 Searching Lightweight Neural Network for Image Signal Processing
abstract
Recently, it has been shown that the traditional Image Signal Processing (ISP) can be replaced by deep neural networks due to their superior performance. However, most of these networks require heavy computation burden and thus are far from sufficient to be deployed on resource-limited platforms, including but not limited to mobile devices and FPGA. To tackle this challenge, we propose an automated search framework that derives ISP models with high image quality while satisfying the low-computation requirement. To reduce the search cost, we adopt the weight-sharing strategy by introducing a supernet and decouple the architecture search into two stages, supernet training and hard-aware evolutionary search. With the proposed framework, we can train the ISP model once and quickly find high-performance but low-computation models on multiple devices. Experiments demonstrate that the searched ISP models have an excellent trade-off between image quality and model complexity, i.e., achieve compelling reconstruction quality with more than 90% reduction in FLOPs as compared to the state-of-the-art networks.
Haojia Lin, Lijiang Li, Xiawu Zheng, Fei Chao 0001, Rongrong Ji
ACM Multimedia2
2022 A Type 2 wavelet brain emotional learning network with double recurrent loops based controller for nonlinear systems
Zi-Qi Wang, Lijiang Li, Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Changle Zhou, Xiang Chang, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.2