EDBT 2026 Demo / reviewers in the wild / expert
Long Peng 0003
dblp:24/11239-3
· DBLP profile ↗
19ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-1012-5109ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Codebook Knowledge with Mamba-Transformer For Low-Light Image EnhancementabstractLow-light image enhancement is a critical task which aims to improve the quality of images captured in bad lighting conditions, and contribute to more robust and reliable computer vision systems. Existing methods failed to account for multifaceted and intertwined degradations typically encountered in low-light scenarios. In this paper, we have reconsider the application of vector-quantized codebook in low-light image enhancement task as a domain adaptation paradigm and proposed an effective method called CodeMTNet to solve aforementioned issues. Specifically, we leverage codebook learning from collections of norm-light images to provide unified high-quality knowledge guidance. We have further developed two learning schemes, namely Domain Adaptation Encoder with implicit neural representation regularization across multiple scales, and hybrid Mamba-Transformer blocks for nearest neighbor matching, to tackle distribution mismatch between features of low-quality low-light images and high-quality normal-light images. Additionally, to solve structural information loss during codebook retrieval, we have introduced a controllable feature fusion modules for well texture detail preservation. Experiments conducted on public datasets have demonstrated that CodeMTNet consistently outperforms many state-of-the-art methods and restore images better in line with human perception. Related source codes and pretrained parameters are in https://github.com/YunHDan/CodeMTNet.git. Runhua Deng, Aiwen Jiang, Long Peng 0003, Qiuhai Yan |
WACV | 3 |
| 2026 | Efficient Real-World Image Super-Resolution Via Adaptive Directional Gradient Convolution
Long Peng 0003, Zhanfeng Feng, Renjing Pei, Wenbo Li 0002, Jiaming Guo, Xueyang Fu, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
Int. J. Comput. Vis. | 1 |
| 2026 | DaLPSR: Leverage degradation-aligned language prompt for real-world image super-resolution
Aiwen Jiang, Zhi Wei 0004, Long Peng 0003, Feiqiang Liu, Mingwen Wang 0001 |
Image Vis. Comput. | 3 |
| 2025 | Boosting Image De-Raining via Central-Surrounding Synergistic ConvolutionabstractRainy images suffer from quality degradation due to the synergistic effect of rain streaks and accumulation. The rain streaks are anisotropic and show a specific directional arrangement, while the rain accumulation is isotropic and shows a consistent concentration distribution in local regions. This distribution difference makes unified representation learning for rain streaks and accumulation challenging, which may lead to structure distortion and contrast degradation in the deraining results. To address this problem, a central-surrounding mechanism inspired Synergistic Convolution (SC) is proposed to extract rain streaks and accumulation features simultaneously. Specifically, the SC consists of two parallel novel convolutions: Central-Surrounding Difference Convolution (CSD) and Central-Surrounding Addition Convolution (CSA). In CSD, the difference operation between central and surrounding pixels is injected into the feature extraction process of convolution to perceive the direction distribution of rain streaks. In CSA, the addition operation between central and surrounding pixels is injected into the feature extraction process of convolution to facilitate the modeling of rain accumulation properties. The SC can be used as a general unit to substitute Vanilla Convolution (VC) in current de-raining networks to boost performance. To reduce computational costs, CSA and CSD in SC are merged into a single VC kernel by our parameter equivalent transformation before inferencing. Evaluations of twelve de-raining methods on nine public datasets demonstrate that our proposed SC can comprehensively improve the performance of twelve de-raining networks under various rainy conditions without changing the original network structure or introducing extra computational costs. Even for the current SOTA methods, SC can further achieve SOTA++ performance. The source codes will be publicly available. Long Peng 0003, Yang Wang 0015, Xin Di, Peizhe Xia, Xueyang Fu, Yang Cao 0010, Zhengjun Zha |
AAAI | 1 |
| 2025 | QMambaBSR: Burst Image Super-Resolution with Query State Space ModelabstractBurst super-resolution (BurstSR) aims to reconstruct high-resolution images by fusing subpixel details from multiple low-resolution burst frames. The primary challenge lies in effectively extracting useful information while mitigating the impact of high-frequency noise. Most existing methods rely on frame-by-frame fusion, which often struggles to distinguish informative subpixels from noise, leading to suboptimal performance. To address these limitations, we introduce a novel Query Mamba Burst Super-Resolution (QMambaBSR) network. Specifically, we observe that sub-pixels have consistent spatial distribution while noise appears randomly. Considering the entire burst sequence during fusion allows for more reliable extraction of consistent subpixels and better suppression of noise outliers. Based on this, a Query State Space Model (QSSM) is proposed for both inter-frame querying and intra-frame scanning, enabling a more efficient fusion of useful subpixels. Additionally, to overcome the limitations of static upsampling methods that often result in over-smoothing, we propose an Adaptive Upsampling (AdaUp) module that dynamically adjusts the upsampling kernel to suit the characteristics of different burst scenes, achieving superior detail reconstruction. Extensive experiments on four benchmark datasets—spanning both synthetic and real-world images—demonstrate that QMambaBSR outperforms existing state-of-the-art methods. Xin Di, Long Peng 0003, Peizhe Xia, Wenbo Li 0002, Renjing Pei, Yang Cao 0010, Yang Wang 0015, Zhengjun Zha |
CVPR | 2 |
| 2025 | Fast Image Super-Resolution via Consistency Rectified Flow
Wenbo Li 0002, Haoze Sun, Zhixin Wang, Long Peng 0003, Xiaowei Hu 0001, Renjing Pei, Pheng-Ann Heng |
ICCV | 6 |
| 2025 | Towards Realistic Data Generation for Real-World Super-ResolutionabstractExisting image super-resolution (SR) techniques often fail to generalize effectively in complex real-world settings due to the significant divergence between training data and practical scenarios. To address this challenge, previous efforts have either manually simulated intricate physical-based degradations or utilized learning-based techniques, yet these approaches remain inadequate for producing large-scale, realistic, and diverse data simultaneously. In this paper, we introduce a novel Realistic Decoupled Data Generator (RealDGen), an unsupervised learning data generation framework designed for real-world super-resolution. We meticulously develop content and degradation extraction strategies, which are integrated into a novel content-degradation decoupled diffusion model to create realistic low-resolution images from unpaired real LR and HR images. Extensive experiments demonstrate that RealDGen excels in generating large-scale, high-quality paired data that mirrors real-world degradations, significantly advancing the performance of popular SR models on various real-world benchmarks. Long Peng 0003, Wenbo Li 0002, Renjing Pei, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
ICLR | 1 |
| 2025 | Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image RestorationabstractImage restoration aims to recover details and enhance contrast in degraded images. With the growing demand for high-quality imaging (e.g., 4K and 8K), achieving a balance between restoration quality and computational efficiency has become increasingly critical. Existing methods, primarily based on CNNs, Transformers, or their hybrid approaches, apply uniform deep representation extraction across the image. However, these methods often struggle to effectively model long-range dependencies and largely overlook the spatial characteristics of image degradation (regions with richer textures tend to suffer more severe damage), making it hard to achieve the best trade-off between restoration quality and efficiency. To address these issues, we propose a novel texture-aware image restoration method, TAMambaIR, which simultaneously perceives image textures and achieves a trade-off between performance and efficiency. Specifically, we introduce a novel Texture-Aware State Space Model, which enhances texture awareness and improves efficiency by modulating the transition matrix of the state-space equation and focusing on regions with complex textures. Additionally, we design a Multi-Directional Perception Block to improve multi-directional receptive fields while maintaining low computational overhead. Extensive experiments on benchmarks for image super-resolution, deraining, and low-light image enhancement demonstrate that TAMambaIR achieves state-of-the-art performance with significantly improved efficiency, establishing it as a robust and efficient framework for image restoration. Long Peng 0003, Xin Di, Zhanfeng Feng, Wenbo Li 0002, Renjing Pei, Yang Wang 0015, Xueyang Fu, Yang Cao 0010, Zhengjun Zha |
IJCAI | 1 |
| 2025 | PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementabstractMulti-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which are widely used in streaming video processing, surveillance, and generation. Although numerous Transformer-based enhancement methods have achieved impressive performance, their computational and memory demands hinder deployment on edge devices. Quantization offers a practical solution by reducing the bit-width of weights and activations to improve efficiency. However, directly applying existing quantization methods to video enhancement tasks often leads to significant performance degradation and loss of fine details. This stems from two limitations: (a) inability to allocate varying representational capacity across frames, which results in suboptimal dynamic range adaptation; (b) over-reliance on full-precision teachers, which limits the learning of low-bit student models. To tackle these challenges, we propose a novel quantization method for video enhancement: Progressive Multi-Frame Quantization for Video Enhancement (PMQ-VE). This framework features a coarse-to-fine two-stage process: Backtracking-based Multi-Frame Quantization (BMFQ) and Progressive Multi-Teacher Distillation (PMTD). BMFQ utilizes a percentile-based initialization and iterative search with pruning and backtracking for robust clipping bounds. PMTD employs a progressive distillation strategy with both full-precision and multiple high-bit (INT) teachers to enhance low-bit models' capacity and quality. Extensive experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance across multiple tasks and benchmarks. The code will be made publicly available. Zhanfeng Feng, Long Peng 0003, Xin Di, Wenbo Li 0002, Yulun Zhang 0001, Renjing Pei, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
NeurIPS | 2 |
| 2025 | Dropout the High-Rate Downsampling: A Novel Design Paradigm for UHD Image RestorationabstractWith the popularization of high-end mobile devices, Ultra-high-definition (UHD) images have become ubiquitous in our lives. The restoration of UHD images is a highly challenging problem due to the exaggerated pixel count, which often leads to memory overflow during processing. Existing methods either downsample UHD images at a high rate before processing or split them into multiple patches for separate processing. However, high-rate downsampling leads to significant information loss, while patch-based approaches inevitably introduce boundary artifacts. In this paper, we propose a novel design paradigm to solve the UHD image restoration problem, called D2Net. D2Net enables direct full-resolution inference on UHD images without the need for high-rate downsampling or dividing the images into several patches. Specifically, we ingeniously utilize the characteristics of the frequency domain to establish long-range dependencies of features. Taking into account the richer local patterns in UHD images, we also design a multi-scale convolutional group to capture local features. Additionally, during the decoding stage, we dynamically incorporate features from the encoding stage to reduce the flow of irrelevant information. Extensive experiments on three UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring, show that our model achieves better quantitative and qualitative results than state-of-the-art methods. Chen Wu 0006, Long Peng 0003, Dianjie Lu, Zhuoran Zheng |
WACV | 3 |
| 2025 | Textual prompt guided image restoration
Qiuhai Yan, Aiwen Jiang, Long Peng 0003, Qiaosi Yi |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Latent Degradation Representation Constraint for Single Image DerainingabstractSince rain shows a variety of shapes and directions, learning the degradation representation is extremely challenging for single image deraining. Existing methods mainly propose to designing complicated modules to implicitly learn latent degradation representation from rainy images. However, it is hard to decouple the content-independent degradation representation due to the lack of explicit constraint, resulting in over- or under-enhancement problems. To tackle this issue, we propose a novel Latent Degradation Representation Constraint Network (LDRCNet) that consists of the Direction-Aware Encoder (DAEncoder), Deraining Network, and Multi-Scale Interaction Block (MSIBlock). Specifically, the DAEncoder is proposed to extract latent degradation representation adaptively by first using the deformable convolutions to exploit the direction property of rain streaks. Next, a constraint loss is introduced to explicitly constraint the degradation representation learning during training. Last, we propose an MSIBlock to fuse with the learned degradation representation and decoder features of the deraining network for adaptive information interaction to remove various complicated rainy patterns and reconstruct image details. Experimental results on five synthetic and four real datasets demonstrate that our method achieves state-of-the-art performance. The source code is available at https://github.com/Madeline-hyh/LDRCNet. Long Peng 0003, Lu Wang 0001, Jun Cheng 0003 |
ICASSP | 2 |
| 2024 | FC3DNET: A Fully Connected Encoder-Decoder for Efficient DemoiréingabstractMoiré patterns are commonly seen when taking photos of screens. Camera devices usually have limited hardware performance but take high-resolution photos. However, users are sensitive to the photo processing time, which presents a hardly considered challenge of efficiency for demoiréing methods. To balance the network speed and quality of results, we propose a Fully Connected enCoder-deCoder based Demoiréing Network (FC3DNet). FC3DNet utilizes features with multiple scales in each stage of the decoder for comprehensive information, which contains long-range patterns as well as various local moiré styles that both are crucial aspects in demoiréing. Besides, to make full use of multiple features, we design a Multi-Feature Multi-Attention Fusion (MFMAF) module to weigh the importance of each feature and compress them for efficiency. These designs enable our network to achieve performance comparable to state-of-the-art (SOTA) methods in real-world datasets while utilizing only a fraction of parameters, FLOPs, and runtime. Zhibo Du, Long Peng 0003, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
ICIP | 2 |
| 2024 | Dual-Path Coupled Image Deraining Network Via Spatial-Frequency InteractionabstractTransformers have recently emerged as a significant force in the field of image deraining. Existing image deraining methods utilize extensive research on self-attention. Though showcasing impressive results, they tend to neglect critical frequency information, as self-attention is generally less adept at capturing high-frequency details. To overcome this shortcoming, we have developed an innovative Dual-Path Coupled Deraining Network (DPCNet) that integrates information from both spatial and frequency domains through Spatial Feature Extraction Block (SFEBlock) and Frequency Feature Extraction Block (FFEBlock). We have further introduced an effective Adaptive Fusion Module (AFM) for the dual-path feature aggregation. Extensive experiments on six public deraining benchmarks and downstream vision tasks have demonstrated that our proposed method not only outperforms the existing state-of-the-art deraining method but also achieves visually pleasuring results with excellent robustness on downstream vision tasks. The source code is available at https://github.com/Madeline-hyh/DPCNet. Aiwen Jiang, Lingfang Jiang, Long Peng 0003, Zhifeng Wang 0006, Lu Wang 0001 |
ICIP | 4 |
| 2024 | UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New PeaksabstractUltra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality images at multiple resolutions (\textit{e.g.}, 1K, 2K, and 4K) within a single model, while maintaining computational efficiency. UltraPixel leverages semantics-rich representations of lower-resolution images in a later denoising stage to guide the whole generation of highly detailed high-resolution images, significantly reducing complexity. Specifically, we introduce implicit neural representations for continuous upsampling and scale-aware normalization layers adaptable to various resolutions. Notably, both low- and high-resolution processes are performed in the most compact space, sharing the majority of parameters with less than 3$\%$ additional parameters for high-resolution outputs, largely enhancing training and inference efficiency. Our model achieves fast training with reduced data requirements, producing photo-realistic high-resolution images and demonstrating state-of-the-art performance in extensive experiments. Wenbo Li 0002, Haoyu Chen 0003, Renjing Pei, Long Peng 0003, Fenglong Song, Lei Zhu 0003 |
NeurIPS | 7 |
| 2024 | Lightweight Adaptive Feature De-Drifting for Compressed Image ClassificationabstractJPEG is a widely used compression scheme to efficiently reduce the volume of the transmitted images at the expense of visual perception drop. The artifacts appear among blocks due to the information loss in the compression process, which not only affects the quality of images but also harms the subsequent high-level tasks in terms of feature drifting. High-level vision models trained on high-quality images will suffer performance degradation when dealing with compressed images, especially on mobile devices. In recent years, numerous learning-based JPEG artifacts removal methods have been proposed to handle visual artifacts. However, it is not an ideal choice to use these JPEG artifacts removal methods as a pre-processing for compressed image classification for the following reasons: 1) These methods are designed for human vision rather than high-level vision models. 2) These methods are not efficient enough to serve as a pre-processing on resource-constrained devices. To address these issues, this paper proposes a novel lightweight adaptive feature de-drifting module (AFD-Module) to boost the performance of pre-trained image classification models when facing compressed images. First, a Feature Drifting Estimation Network (FDE-Net) is devised to generate the spatial-wise Feature Drifting Map (FDM) in the DCT domain. Next, the estimated FDM is transmitted to the Feature Enhancement Network (FE-Net) to generate the mapping relationship between degraded features and corresponding high-quality features. Specially, a simple but effective RepConv block equipped with structural re-parameterization is utilized in FE-Net, which enriches feature representation in the training phase while keeping efficiency in the deployment phase. After training on limited compressed images, the AFD-Module can serve as a “plug-and-play” module for pre-trained classification models to improve their performance on compressed images. Experiments on images compressed once (i.e.ImageNet-C) and multiple times demonstrate that our proposed AFD-Module can comprehensively improve the accuracy of the pre-trained classification models and significantly outperform the existing methods. Long Peng 0003, Yang Cao 0010, Yuejin Sun, Yang Wang 0015 |
IEEE Trans. Multim. | 1 |
| 2023 | Decoupling-and-Aggregating for Image Exposure CorrectionabstractThe images captured under improper exposure conditions often suffer from contrast degradation and detail distortion. Contrast degradation will destroy the statistical properties of low-frequency components, while detail distortion will disturb the structural properties of high-frequency components, leading to the low-frequency and high-frequency components being mixed and inseparable. This will limit the statistical and structural modeling capacity for exposure correction. To address this issue, this paper proposes to decouple the contrast enhancement and detail restoration within each convolution process. It is based on the observation that, in the local regions covered by convolution kernels, the feature response of low-/high-frequency can be decoupled by addition/difference operation. To this end, we inject the addition/difference operation into the convolution process and devise a Contrast Aware (CA) unit and a Detail Aware (DA) unit to facilitate the statistical and structural regularities modeling. The proposed CA and DA can be plugged into existing CNN-based exposure correction networks to substitute the Traditional Convolution (TConv) to improve the performance. Furthermore, to maintain the computational costs of the network without changing, we aggregate two units into a single TConv kernel using structural re-parameterization. Evaluations of nine methods and five benchmark datasets demonstrate that our proposed method can comprehensively improve the performance of existing methods without introducing extra computational costs compared with the original networks. The codes will be publicly available. Yang Wang 0015, Long Peng 0003, Liang Li 0003, Yang Cao 0010, Zhengjun Zha |
CVPR | 2 |
| 2021 | Ensemble single image deraining network via progressive structural boosting constraints
Long Peng 0003, Aiwen Jiang, Bo Liu 0006, Mingwen Wang 0001 |
Signal Process. Image Commun. | 1 |
| 2020 | Cumulative Rain Density Sensing Network for Single Image DerainabstractThis paper focuses on single image derain, which aims to restore clear image from single rain image. Through full consideration of different frequency information preservation and the complicated interactions between rain-streaks and background, a novel end-to-end cumulative rain-density sensing network (CRDNet) is proposed for adaptive rain-streaks removal. An effective W-Net with powerful learning ability is proposed as a key component to recover rain-invariant low-frequency signals. A cumulative rain-density classifier with a novel cost-sensitive label encoding strategy is proposed as an auxiliary network to improve discriminative power of extracted high-frequency rain-streaks through multi-task training. The proposed CRDNet has been compared with state-of-the-art methods on two public datasets. The quantitative and visual experimental results demonstrate that it can achieve excellent performance with great improvement. Related source code and models are available on github https://github.com/peylnog/CRDNet. Long Peng 0003, Aiwen Jiang, Qiaosi Yi, Mingwen Wang 0001 |
IEEE Signal Process. Lett. | 1 |