Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shuaizheng Liu

dblp:201/7482 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-2358-6713ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
6 papers
Image and video processing · 86% Visual content generation and editing · 8% Computational photography and imaging · 6%
Artificial intelligence
5 papers
Generative modeling · 84% Efficient and distributed learning · 8% Vision and language · 4%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.032025
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution · NeurIPS 2025
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach · CVPR 2025
InstructRestore: Region-Customized Image Restoration with Human Instructions · NeurIPS 2025
Image and video processing
image restoration
1.722025
InstructRestore: Region-Customized Image Restoration with Human Instructions · NeurIPS 2025
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models · ICCV 2025
Image and video processing › super-resolution › image super-resolution
real-world image super-resolution
1.722025
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models · ICCV 2025
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach · CVPR 2025
Image and video processing
super-resolution
1.622025
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models · ICCV 2025
Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective · IEEE Trans. Image Process. 2024
Machine learning › Generative modeling
autoregressive model
0.912025
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models · ICCV 2025
Machine learning › Generative modeling › diffusion model › image restoration
diffusion-based image restoration
0.912025
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach · CVPR 2025
Machine learning › Generative modeling › autoregressive model
multimodal autoregressive generation
0.912025
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models · ICCV 2025
Machine learning › Generative modeling › diffusion model › few-step generation
one-step diffusion
0.912025
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution · NeurIPS 2025
Image and video processing › super-resolution
image super-resolution
0.912025
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach · CVPR 2025
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing
0.912025
InstructRestore: Region-Customized Image Restoration with Human Instructions · NeurIPS 2025
Image and video processing › super-resolution › video super-resolution
real-world video super-resolution
0.912025
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution · NeurIPS 2025
Image and video processing › super-resolution
video super-resolution
0.912025
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution · NeurIPS 2025
Image and video processing › image restoration
perception-distortion tradeoff
0.812024
Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective · IEEE Trans. Image Process. 2024
Mathematical optimization › multi-objective optimization
evolutionary algorithm
0.812024
Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective · IEEE Trans. Image Process. 2024
Mathematical optimization
multi-objective optimization
0.812024
Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective · IEEE Trans. Image Process. 2024
Computational photography and imaging
high dynamic range imaging
0.712023
Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image Dataset · CVPR 2023
Image and video processing › image restoration
image denoising
0.712023
Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image Dataset · CVPR 2023
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.312025
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach · CVPR 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.312025
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach · CVPR 2025
Image and video processing › video processing
temporal consistency
0.312025
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.212024
Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective · IEEE Trans. Image Process. 2024

Methods — techniques the papers use, named apart from their topics

LoRA · 3.5instruction tuning · 1.7entropy-based top-k sampling · 1.7diffusion prior · 1.7diffusion model · 1.7data generation engine · 1.7cross-frame retrieval · 1.7controlnet · 1.7classifier score distillation · 1.7LPIPS loss · 1.7gradient-free evolutionary algorithm · 0.8fusion network · 0.8adam · 0.8
YearPublicationVenuePosition
2025 Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach
abstract
Diffusion prior-based methods have shown impressive results in real-world image super-resolution (SR). However, most existing methods entangle pixel-level and semantic-level SR objectives in the training process, struggling to balance pixel-wise fidelity and perceptual quality. Meanwhile, users have varying preferences on SR results, thus it is demanded to develop an adjustable SR model that can be tailored to different fidelity-perception preferences during inference without re-training. We present Pixel-level and Semantic-level Adjustable SR (PiSA-SR), which learns two LoRA modules upon the pre-trained stable-diffusion (SD) model to achieve improved and adjustable SR results. We first formulate the SD-based SR problem as learning the residual between the low-quality input and the high-quality output, then show that the learning objective can be decoupled into two distinct LoRA weight spaces: one is characterized by the ℓ2-loss for pixel-level regression, and another is characterized by the LPIPS and classifier score distillation losses to extract semantic information from pre-trained classification and SD models. In its default setting, PiSA-SR can be performed in a single diffusion step, achieving leading real-world SR results in both quality and efficiency. By introducing two adjustable guidance scales on the two LoRA modules to control the strengths of pixel-wise fidelity and semantic-level details during inference, PiSA-SR can offer flexible SR results according to user preference without re-training. The source code of our method can be found at https://github.com/csslc/PiSA-SR.
Lingchen Sun, Rongyuan Wu, Zhiyuan Ma 0002, Shuaizheng Liu, Qiaosi Yi, Lei Zhang 0006
CVPR4
2025 Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
abstract
By leveraging the generative priors from pre-trained text-to-image diffusion models, significant progress has been made in real-world image super-resolution (Real-ISR). However, these methods tend to generate inaccurate and unnatural reconstructions in complex and/or heavily degraded scenes, primarily due to their limited perception and understanding capability of the input low-quality image. To address these limitations, we propose, for the first time to our knowledge, to adapt the pre-trained autoregressive multimodal model such as Lumina-mGPT into a robust Real-ISR model, namely PURE, which Perceives and Understands the input low-quality image, then REstores its high-quality counterpart. Specifically, we implement instruction tuning on Lumina-mGPT to perceive the image degradation level and the relationships between previously generated image tokens and the next token, understand the image content by generating image semantic descriptions, and consequently restore the image by generating high-quality image tokens autoregressively with the collected information. In addition, we reveal that the image token entropy reflects the image structure and present a entropy-based Top-k sampling strategy to optimize the local structure of the image during inference. Experimental results demonstrate that PURE preserves image content while generating realistic details, especially in complex scenes with multiple objects, showcasing the potential of autoregressive multimodal generative models for robust Real-ISR. The model and code will be available at https://github.com/nonwhy/PURE.
Hongyang Wei, Shuaizheng Liu, Chun Yuan 0003
ICCV2
2025 InstructRestore: Region-Customized Image Restoration with Human Instructions
abstract
Despite the significant progress in diffusion prior-based image restoration for real-world scenarios, most existing methods apply uniform processing to the entire image, lacking the capability to perform region-customized image restoration according to user preferences. In this work, we propose a new framework, namely InstructRestore, to perform region-adjustable image restoration following human instructions. To achieve this, we first develop a data generation engine to produce training triplets, each consisting of a high-quality image, the target region description, and the corresponding region mask. With this engine and careful data screening, we construct a comprehensive dataset comprising 536,945 triplets to support the training and evaluation of this task. We then examine how to integrate the low-quality image features under the ControlNet architecture to adjust the degree of image details enhancement. Consequently, we develop a ControlNet-like model to identify the target region and allocate different integration scales to the target and surrounding regions, enabling region-customized image restoration that aligns with user instructions. Experimental results demonstrate that our proposed InstructRestore approach enables effective human-instructed image restoration, including restoration with controllable bokeh blur effects and region-specific restoration with continuous intensity control. Our work advances the investigation of interactive image restoration and enhancement techniques. Data, code, and models are publicly available at https://github.com/shuaizhengliu/InstructRestore.git.
Shuaizheng Liu, Jianqi Ma, Lingchen Sun, Xiangtao Kong, Lei Zhang 0006
NeurIPS1
2025 One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
abstract
It is a challenging problem to reproduce rich spatial details while maintaining temporal consistency in real-world video super-resolution (Real-VSR), especially when we leverage pre-trained generative models such as stable diffusion (SD) for realistic details synthesis. Existing SD-based Real-VSR methods often compromise spatial details for temporal coherence, resulting in suboptimal visual quality. We argue that the key lies in how to effectively extract the degradation-robust temporal consistency priors from the low-quality (LQ) input video and enhance the video details while maintaining the extracted consistency priors. To achieve this, we propose a Dual LoRA Learning (DLoRAL) paradigm to train an effective SD-based one-step diffusion model, achieving realistic frame details and temporal consistency simultaneously. Specifically, we introduce a Cross-Frame Retrieval (CFR) module to aggregate complementary information across frames, and train a Consistency-LoRA (C-LoRA) to learn robust temporal representations from degraded inputs. After consistency learning, we fix the CFR and C-LoRA modules and train a Detail-LoRA (D-LoRA) to enhance spatial details while aligning with the temporal space defined by C-LoRA to keep temporal coherence. The two phases alternate iteratively for optimization, collaboratively delivering consistent and detail-rich outputs. During inference, the two LoRA branches are merged into the SD model, allowing efficient and high-quality video restoration in a single diffusion step. Experiments show that DLoRAL achieves strong performance in both accuracy and speed. Code and models will be released.
Lingchen Sun, Shuaizheng Liu, Rongyuan Wu, Zhengqiang Zhang, Lei Zhang 0006
NeurIPS3
2024 Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective
abstract
High perceptual quality and low distortion degree are two important goals in image restoration tasks such as super-resolution (SR). Most of the existing SR methods aim to achieve these goals by minimizing the corresponding yet conflicting losses, such as the$\ell _{1}$loss and the adversarial loss. Unfortunately, the commonly used gradient-based optimizers, such as Adam, are hard to balance these objectives due to the opposite gradient decent directions of the contradictory losses. In this paper, we formulate the perception-distortion trade-off in SR as a multi-objective optimization problem and develop a new optimizer by integrating the gradient-free evolutionary algorithm (EA) with gradient-based Adam, where EA and Adam focus on the divergence and convergence of the optimization directions respectively. As a result, a population of optimal models with different perception-distortion preferences is obtained. We then design a fusion network to merge these models into a single stronger one for an effective perception-distortion trade-off. Experiments demonstrate that with the same backbone network, the perception-distortion balanced SR model trained by our method can achieve better perceptual quality than its competitors while attaining better reconstruction fidelity. Codes and models can be found athttps://github.com/csslc/EA-Adam.
Lingchen Sun, Jie Liang 0007, Shuaizheng Liu, Hongwei Yong, Lei Zhang 0006
IEEE Trans. Image Process.3
2023 Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image Dataset
abstract
Mobile phones have become a ubiquitous and indispensable photographing device in our daily life, while the small aperture and sensor size make mobile phones more susceptible to noise and over-saturation, resulting in low dynamic range (LDR) and low image quality. It is thus crucial to develop high dynamic range (HDR) imaging techniques for mobile phones. Unfortunately, the existing HDR image datasets are mostly constructed by DSLR cameras in daytime, limiting their applicability to the study of HDR imaging for mobile phones. In this work, we develop, for the first time to our best knowledge, an HDR image dataset by using mobile phone cameras, namely Mobile-HDR dataset. Specifically, we utilize three mobile phone cameras to collect paired LDR-HDR images in the raw image domain, covering both daytime and night-time scenes with different noise levels. We then propose a transformer based model with a pyramid cross-attention alignment module to aggregate highly correlated features from different exposure frames to perform joint HDR denoising and fusion. Experiments validate the advantages of our dataset and our method on mobile HDR imaging. Dataset and codes are available at https://github.com/shuaizhengliu/Joint-HDRDN.
Shuaizheng Liu, Lingchen Sun, Zhetong Liang, Hui Zeng 0001, Lei Zhang 0006
CVPR1
2017 RoDLSR: Robust discriminative least squares regression model for multi-category classification
abstract
Discriminative least squares regression (DLSR) is a simple yet effective method for multi-class classification. One problem of DLSR is that it is lack of robustness to outliers. In order to tackle this difficulty, in this paper, we propose a novel Robust DLSR (RoDLSR) model. The core idea behind RoDLSR is to find and further ignore the outliers among the support vector set. Specifically, we modify the regression targets of outliers by adding an additional item. As a result, the range of regression residuals can be controlled within predefined threshold. Extensive experiments evaluate the effectiveness of RoDLSR, especially on the corrupted databases.
Lingfeng Wang 0002, Shuaizheng Liu, Chunhong Pan
ICASSP2