EDBT 2026 Demo / reviewers in the wild / expert
Lingchen Sun
dblp:273/8821
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-2254-7472ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D GenerationabstractSingle-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly degrading 3D reconstruction performance. While recent methods attempt to refine CVC by feeding reconstruction results back into the multi-view generator, these approaches struggle with noisy and unstable reconstruction outputs that limit effective CVC improvement. We introduce AlignCVC, a novel framework that fundamentally re-frames single-image-to-3D generation through distribution alignment rather than relying on strict regression losses. Our key insight is to align both generated and reconstructed multi-view distributions toward the ground-truth multi-view distribution, establishing a principled foundation for improved CVC. Observing that generated images exhibit weak CVC while reconstructed images display strong CVC due to explicit rendering, we propose a soft-hard alignment strategy with distinct objectives for generation and reconstruction models. This approach not only enhances generation quality but also dramatically accelerates inference to as few as 4 steps. As a plug-and-play paradigm, our method, namely AlignCVC, seamlessly integrates various combinations of multiview generation models with 3D reconstruction models. Extensive experiments demonstrate the effectiveness and efficiency of AlignCVC for single-image-to-3D generation. Zhiyuan Ma 0002, Lingchen Sun, Lei Zhang 0006 |
AAAI | 3 |
| 2026 | Online-updating neural network with decentralized prediction for dynamic multi-objective optimization
Ru Lei, Lin Li 0016, Yinan Guo 0001, Yiqi Feng, Lingchen Sun, Rustam Stolkin, Mohammed Eesa Asif |
Expert Syst. Appl. | 6 |
| 2026 | Neural network-based framework for wide visibility dehazing with synthetic benchmarks
Lin Li 0016, Ru Lei, Lingchen Sun, Rustam Stolkin |
Pattern Recognit. | 4 |
| 2025 | Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA ApproachabstractDiffusion prior-based methods have shown impressive results in real-world image super-resolution (SR). However, most existing methods entangle pixel-level and semantic-level SR objectives in the training process, struggling to balance pixel-wise fidelity and perceptual quality. Meanwhile, users have varying preferences on SR results, thus it is demanded to develop an adjustable SR model that can be tailored to different fidelity-perception preferences during inference without re-training. We present Pixel-level and Semantic-level Adjustable SR (PiSA-SR), which learns two LoRA modules upon the pre-trained stable-diffusion (SD) model to achieve improved and adjustable SR results. We first formulate the SD-based SR problem as learning the residual between the low-quality input and the high-quality output, then show that the learning objective can be decoupled into two distinct LoRA weight spaces: one is characterized by the ℓ2-loss for pixel-level regression, and another is characterized by the LPIPS and classifier score distillation losses to extract semantic information from pre-trained classification and SD models. In its default setting, PiSA-SR can be performed in a single diffusion step, achieving leading real-world SR results in both quality and efficiency. By introducing two adjustable guidance scales on the two LoRA modules to control the strengths of pixel-wise fidelity and semantic-level details during inference, PiSA-SR can offer flexible SR results according to user preference without re-training. The source code of our method can be found at https://github.com/csslc/PiSA-SR. Lingchen Sun, Rongyuan Wu, Zhiyuan Ma 0002, Shuaizheng Liu, Qiaosi Yi, Lei Zhang 0006 |
CVPR | 1 |
| 2025 | Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu 0009, Rongyuan Wu, Lingchen Sun, Yuhui Wu 0001, Lei Zhang 0006 |
ICCV | 4 |
| 2025 | InstructRestore: Region-Customized Image Restoration with Human InstructionsabstractDespite the significant progress in diffusion prior-based image restoration for real-world scenarios, most existing methods apply uniform processing to the entire image, lacking the capability to perform region-customized image restoration according to user preferences. In this work, we propose a new framework, namely InstructRestore, to perform region-adjustable image restoration following human instructions. To achieve this, we first develop a data generation engine to produce training triplets, each consisting of a high-quality image, the target region description, and the corresponding region mask. With this engine and careful data screening, we construct a comprehensive dataset comprising 536,945 triplets to support the training and evaluation of this task. We then examine how to integrate the low-quality image features under the ControlNet architecture to adjust the degree of image details enhancement. Consequently, we develop a ControlNet-like model to identify the target region and allocate different integration scales to the target and surrounding regions, enabling region-customized image restoration that aligns with user instructions. Experimental results demonstrate that our proposed InstructRestore approach enables effective human-instructed image restoration, including restoration with controllable bokeh blur effects and region-specific restoration with continuous intensity control. Our work advances the investigation of interactive image restoration and enhancement techniques. Data, code, and models are publicly available at https://github.com/shuaizhengliu/InstructRestore.git. Shuaizheng Liu, Jianqi Ma, Lingchen Sun, Xiangtao Kong, Lei Zhang 0006 |
NeurIPS | 3 |
| 2025 | One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-ResolutionabstractIt is a challenging problem to reproduce rich spatial details while maintaining temporal consistency in real-world video super-resolution (Real-VSR), especially when we leverage pre-trained generative models such as stable diffusion (SD) for realistic details synthesis. Existing SD-based Real-VSR methods often compromise spatial details for temporal coherence, resulting in suboptimal visual quality.
We argue that the key lies in how to effectively extract the degradation-robust temporal consistency priors from the low-quality (LQ) input video and enhance the video details while maintaining the extracted consistency priors.
To achieve this, we propose a Dual LoRA Learning (DLoRAL) paradigm to train an effective SD-based one-step diffusion model, achieving realistic frame details and temporal consistency simultaneously.
Specifically, we introduce a Cross-Frame Retrieval (CFR) module to aggregate complementary information across frames, and train a Consistency-LoRA (C-LoRA) to learn robust temporal representations from degraded inputs.
After consistency learning, we fix the CFR and C-LoRA modules and train a Detail-LoRA (D-LoRA) to enhance spatial details while aligning with the temporal space defined by C-LoRA to keep temporal coherence.
The two phases alternate iteratively for optimization, collaboratively delivering consistent and detail-rich outputs. During inference, the two LoRA branches are merged into the SD model, allowing efficient and high-quality video restoration in a single diffusion step. Experiments show that DLoRAL achieves strong performance in both accuracy and speed. Code and models will be released. Lingchen Sun, Shuaizheng Liu, Rongyuan Wu, Zhengqiang Zhang, Lei Zhang 0006 |
NeurIPS | 2 |
| 2025 | DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-ResolutionabstractBenefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Although this randomness is sometimes seen as a limitation, it also introduces a wider perceptual quality range, which can be exploited to improve Real-ISR performance. To this end, we introduce Direct Perceptual Preference Optimization for Real-ISR (DP²O-SR), a framework that aligns generative models with perceptual preferences without requiring costly human annotations. We construct a hybrid reward signal by combining full-reference and no-reference image quality assessment (IQA) models trained on large-scale human preference datasets. This reward encourages both structural fidelity and natural appearance. To better utilize perceptual diversity, we move beyond the standard best-vs-worst selection and construct multiple preference pairs from outputs of the same model. Our analysis reveals that the optimal selection ratio depends on model capacity: smaller models benefit from broader coverage, while larger models respond better to stronger contrast in supervision. Furthermore, we propose hierarchical preference optimization, which adaptively weights training pairs based on intra-group reward gaps and inter-group diversity, enabling more efficient and stable learning. Extensive experiments across both diffusion- and flow-based T2I backbones demonstrate that DP²O-SR significantly improves perceptual quality and generalizes well to real-world benchmarks. Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Tianhe Wu, Qiaosi Yi, Shuai Li 0014, Lei Zhang 0006 |
NeurIPS | 2 |
| 2025 | GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and GenerationabstractEffective and efficient tokenization plays an important role in image representation and generation. Conventional methods, constrained by uniform 2D/1D grid tokenization, are inflexible to represent regions with varying shapes and textures and at different locations, limiting their efficacy of feature representation. In this work, we propose **GPSToken**, a novel **G**aussian **P**arameterized **S**patially-adaptive **Token**ization framework, to achieve non-uniform image tokenization by leveraging parametric 2D Gaussians to dynamically model the shape, position, and textures of different image regions. We first employ an entropy-driven algorithm to partition the image into texture-homogeneous regions of variable sizes. Then, we parameterize each region as a 2D Gaussian (mean for position, covariance for shape) coupled with texture features. A specialized transformer is trained to optimize the Gaussian parameters, enabling continuous adaptation of position/shape and content-aware feature extraction. During decoding, Gaussian parameterized tokens are reconstructed into 2D feature maps through a differentiable splatting-based renderer, bridging our adaptive tokenization with standard decoders for end-to-end training. GPSToken disentangles spatial layout (Gaussian parameters) from texture features to enable efficient two-stage generation: structural layout synthesis using lightweight networks, followed by structure-conditioned texture generation. Experiments demonstrate the state-of-the-art performance of GPSToken, which achieves rFID and FID scores of 0.65 and 1.50 on image reconstruction and generation tasks using 128 tokens, respectively. Codes and models of GPSToken can be found at https://github.com/xtudbxk/GPSToken. Zhengqiang Zhang, Rongyuan Wu, Lingchen Sun, Lei Zhang 0006 |
NeurIPS | 3 |
| 2025 | Improving the Stability and Efficiency of Diffusion Models for Content Consistent Super-ResolutionabstractThe generative priors of pre-trained latent diffusion models (DMs) have demonstrated great potential to enhance the visual quality of image super-resolution (SR) results. However, the noise sampling process in DMs introduces randomness in the SR outputs, and the generated contents can differ a lot with different noise samples. The multi-step diffusion process can be accelerated by distilling methods, but the generative capacity is difficult to control. To address these issues, we analyze the respective advantages of DMs and generative adversarial networks (GANs) and propose to partition the generative SR process into two stages, where the DM is employed for reconstructing image structures and the GAN is employed for improving fine-grained details. Specifically, we propose a non-uniform timestep sampling strategy in the first stage. A single timestep sampling is first applied to extract the coarse information from the input image, then a few reverse steps are used to reconstruct the main structures. In the second stage, we finetune the decoder of the pre-trained variational auto-encoder by adversarial GAN training for deterministic detail enhancement. Once trained, our proposed method, namely content consistent super-resolution (CCSR), allows flexible use of different diffusion steps in the inference stage without re-training. Extensive experiments show that with 2 or even 1 diffusion step, CCSR can significantly improve the content consistency of SR outputs while keeping high perceptual quality. Codes and models can be found at https://github.com/csslc/CCSR. Lingchen Sun, Rongyuan Wu, Jie Liang 0007, Zhengqiang Zhang, Hongwei Yong, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2024 | SeeSR: Towards Semantics-Aware Real-World Image Super-ResolutionabstractOwe to the powerful generative priors, the pretrained text-to-image (T2I) diffusion models have become increasingly popular in solving the real-world image super-resolution problem. However, as a consequence of the heavy quality degradation of input low-resolution (LR) images, the destruction of local structures can lead to ambiguous image semantics. As a result, the content of reproduced high-resolution image may have semantic errors, deteriorating the super-resolution performance. To address this issue, we present a semantics-aware approach to better preserve the semantic fidelity of generative real-world image super-resolution. First, we train a degradation-aware prompt extractor, which can generate accurate soft and hard semantic prompts even under strong degradation. The hard semantic prompts refer to the image tags, aiming to enhance the local perception ability of the T2I model, while the soft semantic prompts compensate for the hard ones to provide additional representation information. These semantic prompts encourage the T2I model to generate detailed and semantically accurate results. Further-more, during the inference process, we integrate the LR images into the initial sampling noise to mitigate the diffusion model's tendency to generate excessive random details. The experiments show that our method can reproduce more realistic image details and hold better the semantics. The source code of our method can be found at https://github.com/cswry/SeeSR. Rongyuan Wu, Tao Yang 0042, Lingchen Sun, Zhengqiang Zhang, Shuai Li 0014, Lei Zhang 0006 |
CVPR | 3 |
| 2024 | One-Step Effective Diffusion Network for Real-World Image Super-ResolutionabstractThe pre-trained text-to-image diffusion models have been increasingly employed to tackle the real-world image super-resolution (Real-ISR) problem due to their powerful generative image priors. Most of the existing methods start from random noise to reconstruct the high-quality (HQ) image under the guidance of the given low-quality (LQ) image. While promising results have been achieved, such Real-ISR methods require multiple diffusion steps to reproduce the HQ image, increasing the computational cost. Meanwhile, the random noise introduces uncertainty in the output, which is unfriendly to image restoration tasks. To address these issues, we propose a one-step effective diffusion network, namely OSEDiff, for the Real-ISR problem.
We argue that the LQ image contains rich information to restore its HQ counterpart, and hence the given LQ image can be directly taken as the starting point for diffusion, eliminating the uncertainty introduced by random noise sampling. We finetune the pre-trained diffusion network with trainable layers to adapt it to complex image degradations. To ensure that the one-step diffusion model could yield HQ Real-ISR output, we apply variational score distillation in the latent space to conduct KL-divergence regularization. As a result, our OSEDiff model can efficiently and effectively generate HQ images in just one diffusion step.
Our experiments demonstrate that OSEDiff achieves comparable or even better Real-ISR results, in terms of both objective metrics and subjective evaluations, than previous diffusion model-based Real-ISR methods that require dozens or hundreds of steps. The source codes are released at https://github.com/cswry/OSEDiff. Rongyuan Wu, Lingchen Sun, Zhiyuan Ma 0002, Lei Zhang 0006 |
NeurIPS | 2 |
| 2024 | Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization PerspectiveabstractHigh perceptual quality and low distortion degree are two important goals in image restoration tasks such as super-resolution (SR). Most of the existing SR methods aim to achieve these goals by minimizing the corresponding yet conflicting losses, such as the$\ell _{1}$loss and the adversarial loss. Unfortunately, the commonly used gradient-based optimizers, such as Adam, are hard to balance these objectives due to the opposite gradient decent directions of the contradictory losses. In this paper, we formulate the perception-distortion trade-off in SR as a multi-objective optimization problem and develop a new optimizer by integrating the gradient-free evolutionary algorithm (EA) with gradient-based Adam, where EA and Adam focus on the divergence and convergence of the optimization directions respectively. As a result, a population of optimal models with different perception-distortion preferences is obtained. We then design a fusion network to merge these models into a single stronger one for an effective perception-distortion trade-off. Experiments demonstrate that with the same backbone network, the perception-distortion balanced SR model trained by our method can achieve better perceptual quality than its competitors while attaining better reconstruction fidelity. Codes and models can be found athttps://github.com/csslc/EA-Adam. Lingchen Sun, Jie Liang 0007, Shuaizheng Liu, Hongwei Yong, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2023 | Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image DatasetabstractMobile phones have become a ubiquitous and indispensable photographing device in our daily life, while the small aperture and sensor size make mobile phones more susceptible to noise and over-saturation, resulting in low dynamic range (LDR) and low image quality. It is thus crucial to develop high dynamic range (HDR) imaging techniques for mobile phones. Unfortunately, the existing HDR image datasets are mostly constructed by DSLR cameras in daytime, limiting their applicability to the study of HDR imaging for mobile phones. In this work, we develop, for the first time to our best knowledge, an HDR image dataset by using mobile phone cameras, namely Mobile-HDR dataset. Specifically, we utilize three mobile phone cameras to collect paired LDR-HDR images in the raw image domain, covering both daytime and night-time scenes with different noise levels. We then propose a transformer based model with a pyramid cross-attention alignment module to aggregate highly correlated features from different exposure frames to perform joint HDR denoising and fusion. Experiments validate the advantages of our dataset and our method on mobile HDR imaging. Dataset and codes are available at https://github.com/shuaizhengliu/Joint-HDRDN. Shuaizheng Liu, Lingchen Sun, Zhetong Liang, Hui Zeng 0001, Lei Zhang 0006 |
CVPR | 3 |
| 2021 | An Automatic and Optimal MPA Design MethodabstractRaw polarimetric images are captured by a focal plane polarimeter which is covered by a micro-polarizer array (MPA). The design of the MPA plays a crucial role in polarimetric imaging. MPAs are predominantly designed according to expert engineering experience and rules of thumb. Typically, only one optimization criterion, maximizing bandwidth, is used to design the MPA. To select a design, an exhaustive search is usually performed on a very limited set of available polarizing patterns, which must be constrained in order to make the search tractable. In contrast, this paper proposes a fully automated and optimal MPA design method (AO-MPA) which generates significantly improved MPAs. Instead of the single criterion of bandwidth, we propose six design principles, and show how they can be utilized to mutually optimize the MPA design by formulating a tri-objective optimization problem with multiple constraints. A much larger set of possible MPA patterns is rapidly and automatically searched by applying advanced multi-objective optimization techniques. We have tested AO-MPA using two groups of experiments, in which AO-MPA is compared against several other leading MPA design methods, and the patterns generated by AO-MPA are compared against state-of-the-art patterns from the literature. The results, obtained using a public benchmark dataset, show that the AO-MPA method is very computationally efficient, and can find all optimal MPA patterns for all array sizes. Moreover, for each size, AO-MPA obtains all optimal layouts simultaneously. AO-MPA generates designs which require fewer polarization orientations, while also yielding better performance in estimating intensity measurements, Stokes vector and the degree of linear polarization. This results in MPAs which are easier to manufacture while also being more robust to noise. Lin Li 0016, Lingchen Sun, Rustam Stolkin, Zhunga Liu |
IEEE Trans. Image Process. | 2 |
| 2020 | Novel Multi-objecitve Evolutionary Algorithm for Color Filter Arrays DesignabstractMost digital cameras use a single sensor covered with a Color Filter Array (CFA) in order to reduce the size, complexity and cost. Now, more and more researchers have paid attentions to the representation of CFA since its crucial role in the process of reconstructing the full color image. The representation of CFA in frequency domain records all the frequency information of image mosaicked with the CFA, thus provides a theoretical approach to design the CFA. However, almost all the existing CFA design methods in frequency domain have limitations. In this paper, we propose a new automatic CFA design method in frequency domain. We choose the optimal frequency structures which satisfy all the design principles through a novel coded multi-objective evolutionary algorithm (MOEA) at first. Then we convert the parameters optimization of the given frequency structures into a constraint single optimization problem, and we solve it by converting it into a multi-objective optimization problem without constraints which is hard to handle. So a MOEA with novel selection and crossover-mutate strategies is proposed to solve the model. The experimental results show that the proposed CFA design method has more advantage when compared with the other existing method. Lingchen Sun, Lin Li 0016, Quan Pan 0001 |
CEC | 1 |