VLDB 2026 Research / reviewers in the wild / expert
Feng Zhang 0039
dblp:48/1294-39
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-7668-8580ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Resolution Photo Enhancement in Real-Time: A Laplacian Pyramid NetworkabstractPhoto enhancement plays a crucial role in augmenting the visual aesthetics of a photograph. In recent years, photo enhancement methods have either focused on enhancement performance, producing powerful models that cannot be deployed on edge devices, or prioritized computational efficiency, resulting in inadequate performance for real-world applications. To this end, this paper introduces a pyramid network called LLF-LUT++, which integrates global and local operators through closed-form Laplacian pyramid decomposition and reconstruction. This approach enables fast processing of high-resolution images while also achieving excellent performance. Specifically, we utilize an image-adaptive 3D LUT that capitalizes on the global tonal characteristics of downsampled images, while incorporating two distinct weight fusion strategies to achieve coarse global image enhancement. To implement this strategy, we designed a spatial-frequency transformer weight predictor that effectively extracts the desired distinct weights by leveraging frequency features. Additionally, we apply local Laplacian filters to adaptively refine edge details in high-frequency components. After meticulously redesigning the network structure and transformer model, LLF-LUT++ not only achieves a 2.64 dB improvement in PSNR on the HDR+ dataset, but also further reduces runtime, with 4 K resolution images processed in just 13 ms on a single GPU. Extensive experimental results on two benchmark datasets further show that the proposed approach performs favorably compared to state-of-the-art methods. Feng Zhang 0039, Haoyou Deng, Lida Li, Qingbo Lu, Zisheng Cao, Minchen Wei, Changxin Gao, Nong Sang, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Toward Robust Alignment for Video Dehazing With Temporal Lookup TableabstractVideo dehazing aims to restore clean scenarios from a sequence of hazy frames, where frame alignment is a critical stage for leveraging temporal information. However, haze degrades contrast and obscures details, making alignment challenging. Existing methods ignore the impairment of haze on alignment and thus struggle to align frames accurately. To address this challenge, we propose an alignment network with the temporal lookup table (temporal-LUT), which effectively enhances the haze-degraded frames and provides vivid cues for precise alignment. Specifically, to tackle the color degradation of haze, we employ a learnable lookup table (LUT) to enhance hazy color. The color mapping nature of LUT favorably preserves the naturalness of enhanced outcomes. Besides, we introduce a temporal weight prediction strategy to strengthen inter-frame interaction, which ensures temporal consistency across enhanced results and thereby benefits alignment. Extensive experimental results on two widely used benchmarks and real-world scenes demonstrate the superiority of our method. Haoyou Deng, Feng Zhang 0039, Qingbo Lu, Changxin Gao, Nong Sang |
IEEE Trans. Image Process. | 3 |
| 2024 | Real-Time Exposure Correction via Collaborative Transformations and Adaptive SamplingabstractMost of the previous exposure correction methods learn dense pixel-wise transformations to achieve promising results, but consume huge computational resources. Recently, Learnable 3D lookup tables (3D LUTs) have demon-strated impressive performance and efficiency for image enhancement. However, these methods can only perform global transformations and fail to finely manipulate local regions. Moreover, they uniformly downsample the input image, which loses the rich color information and limits the learning of color transformation capabilities. In this paper, we present a collaborative transformation framework (CoTF) for real-time exposure correction, which integrates global transformation with pixel-wise transformations in an efficient manner. Specifically, the global transformation adjusts the overall appearance using image-adaptive 3D LUTs to provide decent global contrast and sharp details, while the pixel transformation compensates for local context. Then, a relation-aware modulation module is designed to combine these two components effectively. In addition, we propose an adaptive sampling strategy to preserve more color information by predicting the sampling intervals, thus providing higher quality input data for the learning of 3D LUTs. Extensive experiments demonstrate that our method can process high-resolution images in real-time on GPUs while achieving comparable performance against current state-of-the-art methods. The code is avail-able at https://github.com/HUST-IAL/CoTF. Ziwen Li 0005, Feng Zhang 0039, Jinpu Zhang, Yuanjie Shao, Yuehuan Wang, Nong Sang |
CVPR | 2 |
| 2024 | UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityabstractExisting text-based person retrieval datasets often have relatively coarse-grained text annotations. This hinders the model to comprehend the fine-grained semantics of query texts in real scenarios. To address this problem, we con-tribute a new benchmark named UFineBench for text-based person retrieval with ultra-fine granularity. Firstly, we construct a new dataset named UFine6926. We collect a large number of person images and manually annotate each image with two detailed textual descriptions, averaging 80.8 words each. The average word count is three to four times that of the previous datasets. In addition of standard in-domain evaluation, we also propose a spe-cial evaluation paradigm more representative of real sce-narios. It contains a new evaluation set with cross domains, cross textual granularity and cross textual styles, named UFine3C, and a new evaluation metric for accurately mea-suring retrieval ability, named mean Similarity Distribution (mSD). Moreover, we propose CFAM, a more efficient al-gorithm especially designed for text-based person retrieval with ultra fine-grained texts. It achieves fine granularity mining by adopting a shared cross-modal granularity de-coder and hard negative match mechanism. With standard in-domain evaluation, CFAM establishes competitive performance across various datasets, espe-cially on our ultra fine-grained UFine6926. Furthermore, by evaluating on UFine3C, we demonstrate that training on our UFine6926 significantly improves generalization to real scenarios compared with other coarse-grained datasets. The dataset and code will be made publicly available at https://github.com/Zplusdragon/UFineBench. Jialong Zuo, Hanyu Zhou, Feng Zhang 0039, Tianyu Guo 0001, Nong Sang, Yunhe Wang 0001, Changxin Gao |
CVPR | 4 |
| 2024 | PLIP: Language-Image Pre-training for Person Representation LearningabstractLanguage-image pre-training is an effective technique for learning powerful representations in general domains. However, when directly turning to person representation learning, these general pre-training methods suffer from unsatisfactory performance. The reason is that they neglect critical person-related characteristics, i.e., fine-grained attributes and identities. To address this issue, we propose a novel language-image pre-training framework for person representation learning, termed PLIP. Specifically, we elaborately design three pretext tasks: 1) Text-guided Image Colorization, aims to establish the correspondence between the person-related image regions and the fine-grained color-part textual phrases. 2) Image-guided Attributes Prediction, aims to mine fine-grained attribute information of the person body in the image; and 3) Identity-based Vision-Language Contrast, aims to correlate the cross-modal representations at the identity level rather than the instance level. Moreover, to implement our pre-train framework, we construct a large-scale person dataset with image-text pairs named SYNTH-PEDES by automatically generating textual annotations. We pre-train PLIP on SYNTH-PEDES and evaluate our models by spanning downstream person-centric tasks. PLIP not only significantly improves existing methods on all these tasks, but also shows great ability in the zero-shot and domain generalization settings. The code, dataset and weight will be made publicly available. Jialong Zuo, Jiahao Hong, Feng Zhang 0039, Changqian Yu, Hanyu Zhou, Changxin Gao, Nong Sang, Jingdong Wang 0001 |
NeurIPS | 3 |
| 2024 | Difficulty-Aware Dynamic Network for Lightweight Exposure CorrectionabstractRecently, deep learning-based methods have been successfully applied to the field of exposure correction. However, most of the existing methods treat different locations of an image in the same way, ignoring the inhomogeneous recovery difficulty and spatially-varying visual patterns in the image, which is sub-optimal and not perfectly efficient. In this paper, we propose a difficulty-aware dynamic network (DDNet) for lightweight exposure correction. Specifically, we propose a difficulty-aware strategy that determines the difficulty of feature patches according to a difficulty mask. Then, only the difficult patches are further refined instead of the whole features, which greatly reduces the overall computational complexity. Moreover, in order to achieve spatially-varying processing with a minimal computational burden, we design a spatial-aware dynamic convolution (SDConv), which is generated by predicting a set of basic kernels and a spatial-aware weight map. Benefiting from these designs, our method can strike a good trade-off between performance and complexity. Extensive experiments on several datasets demonstrate that our approach outperforms the state-of-the-art methods both qualitatively and quantitatively while requiring cheaper computational costs. Ziwen Li 0005, Yuanjie Shao, Feng Zhang 0039, Jinpu Zhang, Yuehuan Wang, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Toward Blind Flare Removal Using Knowledge-Driven Flare-Level EstimatorabstractLens flare is a common phenomenon when strong light rays arrive at the camera sensor and a clean scene is consequently mixed up with various opaque and semi-transparent artifacts. Existing deep learning methods are always constrained with limited real image pairs for training. Though recent synthesis-based approaches are found effective, synthesized pairs still deviate from the real ones as the mixing mechanism of flare artifacts and scenes in the wild always depends on a line of undetermined factors, such as lens structure, scratches, etc. In this paper, we present a new perspective from the blind nature of the flare removal task in a knowledge-driven manner. Specifically, we present a simple yet effective flare-level estimator to predict the corruption level of a flare-corrupted image. The estimated flare-level can be interpreted as additive information of the gap between corrupted images and their flare-free correspondences to facilitate a network at both training and testing stages adaptively. Besides, we utilize a flare-level modulator to better integrate the estimations into networks. We also devise a flare-aware block for more accurate flare recognition and reconstruction. Additionally, we collect a new real-world flare dataset for benchmarking, namely WiderFlare. Extensive experiments on three benchmark datasets demonstrate that our method outperforms state-of-the-art methods quantitatively and qualitatively. Haoyou Deng, Lida Li, Feng Zhang 0039, Qingbo Lu, Changxin Gao, Nong Sang |
IEEE Trans. Image Process. | 3 |
| 2023 | Towards General Low-Light Raw Noise Synthesis and ModelingabstractModeling and synthesizing low-light raw noise is a fundamental problem for computational photography and image processing applications. Although most recent works have adopted physics-based models to synthesize noise, the signal-independent noise in low-light conditions is far more complicated and varies dramatically across camera sensors, which is beyond the description of these models. To address this issue, we introduce a new perspective to synthesize the signal-independent noise by a generative model. Specifically, we synthesize the signal-dependent and signal-independent noise in a physics-and learning-based manner, respectively. In this way, our method can be considered as a general model, that is, it can simultaneously learn different noise characteristics for different ISO levels and generalize to various sensors. Subsequently, we present an effective multi-scale discriminator termed Fourier transformer discriminator (FTD) to distinguish the noise distribution accurately. Additionally, we collect a new low-light raw denoising (LRD) dataset for training and benchmarking. Qualitative validation shows that the noise generated by our proposed noise model can be highly similar to the real noise in terms of distribution. Furthermore, extensive denoising experiments demonstrate that our method performs favorably against state-of-the-art methods on different sensors. Feng Zhang 0039, Qingbo Lu, Changxin Gao, Nong Sang |
ICCV | 1 |
| 2023 | Lookup Table meets Local Laplacian Filter: Pyramid Reconstruction Network for Tone MappingabstractTone mapping aims to convert high dynamic range (HDR) images to low dynamic range (LDR) representations, a critical task in the camera imaging pipeline. In recent years, 3-Dimensional LookUp Table (3D LUT) based methods have gained attention due to their ability to strike a favorable balance between enhancement performance and computational efficiency. However, these methods often fail to deliver satisfactory results in local areas since the look-up table is a global operator for tone mapping, which works based on pixel values and fails to incorporate crucial local information. To this end, this paper aims to address this issue by exploring a novel strategy that integrates global and local operators by utilizing closed-form Laplacian pyramid decomposition and reconstruction. Specifically, we employ image-adaptive 3D LUTs to manipulate the tone in the low-frequency image by leveraging the specific characteristics of the frequency information. Furthermore, we utilize local Laplacian filters to refine the edge details in the high-frequency components in an adaptive manner. Local Laplacian filters are widely used to preserve edge details in photographs, but their conventional usage involves manual tuning and fixed implementation within camera imaging pipelines or photo editing tools. We propose to learn parameter value maps progressively for local Laplacian filters from annotated data using a lightweight network. Our model achieves simultaneous global tone manipulation and local edge detail preservation in an end-to-end manner. Extensive experimental results on two benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art methods. Feng Zhang 0039, Ming Tian, Qingbo Lu, Changxin Gao, Nong Sang |
NeurIPS | 1 |
| 2023 | Self-supervised Low-Light Image Enhancement via Histogram Equalization Prior
Feng Zhang 0039, Yuanjie Shao, Yishi Sun, Changxin Gao, Nong Sang |
PRCV (11) | 1 |