Zhen Li 0031

dblp:74/2397-31 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-3338-228XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
abstract
Previous works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance. Still, the computation overhead is also considerable when the window size gradually increases. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86 dB PSNR score on the Urban100 dataset, which is 0.46 dB higher than that of SwinIR but uses fewer parameters and computations. In addition, we also attempt to scale up the model by further enlarging the window size and channel numbers to explore the potential of Transformer-based models. Experiments show that our scaled model, named SRFormerV2, can further improve the results and achieves state-of-the-art. We hope our simple and effective approach could be useful for future research in super-resolution model design.
Yupeng Zhou, Zhen Li 0031, Chunle Guo, Li Liu 0002, Ming-Ming Cheng, Qibin Hou
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 K-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAs
abstract
Recent studies have explored the combination of different LoRAs to jointly generate learned style and content. However, existing methods either fail to effectively preserve both the original subject and style simultaneously or require additional training. In this paper, we argue that the intrinsic properties of LoRA can effectively guide diffusion models in merging learned subject and style. Based on this insight, we propose K-LoRA, a simple yet effective training-free LoRA fusion approach. In each attention layer, K-LoRA compares the Top-K elements in each LoRA to be fused, determining which LoRA to select for optimal fusion. This selection mechanism ensures that the most representative features of both subject and style are retained during the fusion process, effectively balancing their contributions. Experiments demonstrate that K-LoRA can effectively integrates the subject and style information learned by the original LoRAs, outperforming state-of-the-art training-based approaches in both qualitative and quantitative results.
Ziheng Ouyang, Zhen Li 0031, Qibin Hou
CVPR2
2025 AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang 0001, Zhen Li 0031, Shaohui Jiao, Daquan Zhou, Qibin Hou, Ming-Ming Cheng
ICCV5
2024 PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
abstract
Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However, existing per-sonalized generation methods cannot simultaneously sat-isfy the requirements of high efficiency, promising identity (ID) fidelity, and flexible text controllability. In this work, we introduce PhotoMaker, an efficient personalized text-to-image generation method, which mainly encodes an arbitrary number of input ID images into a stack ID embed-ding for preserving ID information. Such an embedding, serving as a unified ID representation, can not only encap-sulate the characteristics of the same input ID comprehen-sively, but also accommodate the characteristics of differ-ent IDs for subsequent integration. This paves the way for more intriguing and practically valuable applications. Be-sides, to drive the training of our PhotoMaker, we propose an ID-oriented data construction pipeline to assemble the training data. Under the nourishment of the dataset constructed through the proposed pipeline, our PhotoMaker demonstrates better ID preservation ability than test-time fine-tuning based methods, yet provides significant speed improvements, high-quality generation results, strong gen-eralization capabilities, and a wide range of applications.
Zhen Li 0031, Mingdeng Cao, Xintao Wang 0002, Zhongang Qi, Ming-Ming Cheng, Ying Shan
CVPR1
2023 DNF: Decouple and Feedback Network for Seeing in the Dark
abstract
The exclusive properties of RAW data have shown great potential for low-light image enhancement. Nevertheless, the performance is bottlenecked by the inherent limitations of existing architectures in both single-stage and multi-stage methods. Mixed mapping across two different domains, noise-to-clean and RAW-to-sRGB, misleads the single-stage methods due to the domain ambiguity. The multi-stage methods propagate the information merely through the resulting image of each stage, neglecting the abundant features in the lossy image-level dataflow. In this paper, we probe a generalized solution to these bottlenecks and propose a Decouple aNd Feedback framework, abbreviated as DNF. To mitigate the domain ambiguity, domain-specific subtasks are decoupled, along with fully utilizing the unique properties in RAW and sRGB domains. The feature propagation across stages with a feedback mechanism avoids the information loss caused by image-level dataflow. The two key insights of our method resolve the inherent limitations of RAW data-based low-light image enhancement satisfactorily, empowering our method to outperform the previous state-of-the-art method by a large margin with only 19% parameters, achieving 0.97dB and 1.30dB PSNR improvements on the Sony and Fuji subsets of SID.
Xin Jin 0005, Linghao Han, Zhen Li 0031, Chunle Guo, Chongyi Li
CVPR3
2023 AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation
abstract
We present All-Pairs Multi-Field Transforms (AMT), a new network architecture for video frame interpolation. It is based on two essential designs. First, we build bidirectional correlation volumes for all pairs of pixels, and use the predicted bilateral flows to retrieve correlations for updating both flows and the interpolated content feature. Second, we derive multiple groups of fine-grained flow fields from one pair of updated coarse flows for performing backward warping on the input frames separately. Combining these two designs enables us to generate promising task-oriented flows and reduce the difficulties in modeling large motions and handling occluded areas during frame interpolation. These qualities promote our model to achieve state-of-the-art performance on various benchmarks with high efficiency. Moreover, our convolution-based model competes favorably compared to Transformer-based models in terms of accuracy and efficiency. Our code is available at https://github.com/MCG-NKU/AMT.
Zhen Li 0031, Zuo-Liang Zhu, Linghao Han, Qibin Hou, Chunle Guo, Ming-Ming Cheng
CVPR1
2023 SRFormer: Permuted Self-Attention for Single Image Super-Resolution
abstract
Previous works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance but the computation overhead is also considerable. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Our PSA is simple and can be easily applied to existing super-resolution networks based on window self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86dB PSNR score on the Urban100 dataset, which is 0.46dB higher than that of SwinIR but uses fewer parameters and computations. We hope our simple and effective approach can serve as a useful tool for future research in super-resolution model design. Our code is available at https://github.com/HVision-NKU/SRFormer.
Yupeng Zhou, Zhen Li 0031, Chunle Guo, Song Bai 0001, Ming-Ming Cheng, Qibin Hou
ICCV2
2022 Towards An End-to-End Framework for Flow-Guided Video Inpainting
abstract
Optical flow, which captures motion information across frames, is exploited in recent video inpainting methods through propagating pixels along its trajectories. However, the hand-crafted flow-based processes in these methods are applied separately to form the whole inpainting pipeline. Thus, these methods are less efficient and rely heavily on the intermediate results from earlier stages. In this paper, we propose an End-to-End framework for Flow-Guided Video Inpainting (E2FGVI) through elaborately designed three trainable modules, namely, flow completion, feature propagation, and content hallucination modules. The three modules correspond with the three stages of previous flow-based methods but can be Jointly optimized, leading to a more efficient and effective inpainting process. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods both qualitatively and quantitatively and shows promising efficiency. The code is available at https://github.com/MCG-NKU/E2FGVI.
Zhen Li 0031, Chengze Lu, Jianhua Qin, Chunle Guo, Ming-Ming Cheng
CVPR1
2022 Wavelet-Based Texture Reformation Network for Image Super-Resolution
abstract
Most reference-based image super-resolution (RefSR) methods directly leverage the raw features extracted from a pretrained VGG encoder to transfer the matched texture information from a reference image to a low-resolution image. We argue that simply operating on these raw features neglects the influence of irrelevant and redundant information and the importance of abundant high-frequency representations, leading to undesirable texture matching and transfer results. Taking the advantages of wavelet transformation, which represents the contextual and textural information of features at different scales, we propose a Wavelet-based Texture Reformation Network (WTRN) for RefSR. We first decompose the extracted texture features into low-frequency and high-frequency sub-bands and conduct feature matching on the low-frequency component. Based on the correlation map obtained from the feature matching process, we then separately swap and transfer wavelet-domain features at different stages of the network. Furthermore, a wavelet-based texture adversarial loss is proposed to make the network generate more visually plausible textures. Experiments on four benchmark datasets demonstrate that our proposed method outperforms previous RefSR methods both quantitatively and qualitatively. The source code is available at https://github.com/zskuang58/WTRN-TIP.
Zhen Li 0031, Zengsheng Kuang, Zuo-Liang Zhu, Hongpeng Wang 0001, Xiuli Shao
IEEE Trans. Image Process.1
2022 Designing an Illumination-Aware Network for Deep Image Relighting
abstract
Lighting is a determining factor in photography that affects the style, expression of emotion, and even quality of images. Creating or finding satisfying lighting conditions, in reality, is laborious and time-consuming, so it is of great value to develop a technology to manipulate illumination in an image as post-processing. Although previous works have explored techniques based on the physical viewpoint for relighting images, extensive supervisions and prior knowledge are necessary to generate reasonable images, restricting the generalization ability of these works. In contrast, we take the viewpoint of image-to-image translation and implicitly merge ideas of the conventional physical viewpoint. In this paper, we present an Illumination-Aware Network (IAN) which follows the guidance from hierarchical sampling to progressively relight a scene from a single image with high efficiency. In addition, an Illumination-Aware Residual Block (IARB) is designed to approximate the physical rendering process and to extract precise descriptors of light sources for further manipulations. We also introduce a depth-guided geometry encoder for acquiring valuable geometry- and structure-related representations once the depth information is available. Experimental results show that our proposed method produces better quantitative and qualitative relighting results than previous state-of-the-art methods. The code and models are publicly available on https://github.com/NK-CS-ZZL/IAN.
Zuo-Liang Zhu, Zhen Li 0031, Ruixun Zhang, Chunle Guo, Ming-Ming Cheng
IEEE Trans. Image Process.2
2021 Temporal Modulation Network for Controllable Space-Time Video Super-Resolution
abstract
Space-time video super-resolution (STVSR) aims to increase the spatial and temporal resolutions of low-resolution and low-frame-rate videos. Recently, deformable convolution based methods have achieved promising STVSR performance, but they could only infer the intermediate frame pre-defined in the training stage. Besides, these methods undervalued the short-term motion cues among adjacent frames. In this paper, we propose a Temporal Modulation Network (TMNet) to interpolate arbitrary intermediate frame(s) with accurate high-resolution reconstruction. Specifically, we propose a Temporal Modulation Block (TMB) to modulate deformable convolution kernels for controllable feature interpolation. To well exploit the temporal information, we propose a Locally-temporal Feature Comparison (LFC) module, along with the Bi-directional Deformable ConvLSTM, to extract short-term and long-term motion cues in videos. Experiments on three benchmark datasets demonstrate that our TMNet outperforms previous STVSR methods. The code is available at https://github.com/CS-GangXu/TMNet.
Jun Xu 0019, Zhen Li 0031, Liang Wang 0001, Xing Sun 0001, Ming-Ming Cheng
CVPR3
2021 Interactive Knowledge Distillation for image classification
Shipeng Fu, Zhen Li 0031, Zitao Liu 0001, Xiaomin Yang
Neurocomputing2
2021 Delving Deep Into Label Smoothing
abstract
Label smoothing is an effective regularization tool for deep neural networks (DNNs), which generates soft labels by applying a weighted average between the uniform distribution and the hard label. It is often used to reduce the overfitting problem of training DNNs and further improve classification performance. In this paper, we aim to investigate how to generate more reliable soft labels. We present an Online Label Smoothing (OLS) strategy, which generates soft labels based on the statistics of the model prediction for the target category. The proposed OLS constructs a more reasonable probability distribution between the target categories and non-target categories to supervise DNNs. Experiments demonstrate that based on the same classification models, the proposed approach can effectively improve the classification performance on CIFAR-100, ImageNet, and fine-grained datasets. Additionally, the proposed method can significantly improve the robustness of DNN models to noisy labels compared to current label smoothing approaches. The source code is available at our project page: https://mmcheng.net/ols/.
Chang-Bin Zhang, Peng-Tao Jiang, Qibin Hou, Yunchao Wei, Qi Han 0007, Zhen Li 0031, Ming-Ming Cheng
IEEE Trans. Image Process.6
2020 Deep recursive up-down sampling networks for single image super-resolution
Zhen Li 0031, Qilei Li, Wei Wu 0002, Jinglei Yang, Xiaomin Yang
Neurocomputing1
2020 Clustering based multiple branches deep networks for single image super-resolution
Zhen Li 0031, Qilei Li, Wei Wu 0002, Zongjun Wu, Lu Lu 0005, Xiaomin Yang
Multim. Tools Appl.1
2020 Model Compression for IoT Applications in Industry 4.0 via Multiscale Knowledge Transfer
abstract
Recently, Industry 4.0 has attracted much attention. It has close relations with the Internet of Things (IoT). On the other hand, convolutional neural networks (CNNs) have shown promising performance in many foundational services of the IoT applications. For the IoT applications with high-speed data streams and the requirement of time-sensitive actions, fast processing is demanded on small-scale platforms or even on IoT devices themselves. Therefore, it is inappropriate to employ cumbersome CNNs in IoT applications, making the study of model compression necessary. In knowledge transfer, it is common to employ a deep, well-trained network, called teacher, to guide a shallow, untrained network, called student, to have better performance. Previous works have made many attempts to transfer single-scale knowledge from teacher to student, leading to degradation of generalization ability. In this article, we introduce multiscale representations to knowledge transfer, which facilitates the generalization ability of student. We divide student and teacher into several stages. Student learns from multiscale knowledge provided by teacher at the end of each stage. Extensive experiments demonstrate the effectiveness of our proposed method both on image classification and on single image super-resolution. The huge performance gap between student and teacher is significantly narrowed down by our proposed method, making student suitable for IoT applications.
Shipeng Fu, Zhen Li 0031, Kai Liu 0012, Sadia Din, Muhammad Imran 0001, Xiaomin Yang
IEEE Trans. Ind. Informatics2
2019 Gated Multiple Feedback Network for Image Super-Resolution
Qilei Li, Zhen Li 0031, Lu Lu 0005, Gwanggil Jeon, Kai Liu 0012, Xiaomin Yang
BMVC2
2019 Feedback Network for Image Super-Resolution
abstract
Recent advances in image super-resolution (SR) explored the power of deep learning to achieve a better reconstruction performance. However, the feedback mechanism, which commonly exists in human visual system, has not been fully exploited in existing deep learning based image SR methods. In this paper, we propose an image super-resolution feedback network (SRFBN) to refine low-level representations with high-level information. Specifically, we use hidden states in a recurrent neural network (RNN) with constraints to achieve such feedback manner. A feedback block is designed to handle the feedback connections and to generate powerful high-level representations. The proposed SRFBN comes with a strong early reconstruction ability and can create the final high-resolution image step by step. In addition, we introduce a curriculum learning strategy to make the network well suitable for more complicated tasks, where the low-resolution images are corrupted by multiple types of degradation. Extensive experimental results demonstrate the superiority of the proposed SRFBN in comparison with the state-of-the-art methods. Code is avaliable at https://github.com/Paper99/SRFBN_CVPR19.
Zhen Li 0031, Jinglei Yang, Zheng Liu 0002, Xiaomin Yang, Gwanggil Jeon, Wei Wu 0002
CVPR1