EDBT 2026 Demo / reviewers in the wild / expert
Zisheng Cao
dblp:06/3607
· DBLP profile ↗
9ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-7037-8039ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
9 papers |
Image and video processing · 48% Computational photography and imaging · 24% Image and video coding · 19% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 20 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image enhancement |
2.4 | 4 | 2026 | High-Resolution Photo Enhancement in Real-Time: A Laplacian Pyramid Network · IEEE Trans. Pattern Anal. Mach. Intell. 2026 Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-Time · IEEE Trans. Pattern Anal. Mach. Intell. 2022 CameraNet: A Two-Stage Framework for Effective Camera ISP Learning · IEEE Trans. Image Process. 2021 |
Computational photography and imaging
image aesthetics |
1.0 | 2 | 2022 | Grid Anchor Based Image Cropping: A New Benchmark and An Efficient Model · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Reliable and Efficient Image Cropping: A Grid Anchor Based Approach · CVPR 2019 |
Visual content generation and editing
image cropping |
1.0 | 2 | 2022 | Grid Anchor Based Image Cropping: A New Benchmark and An Efficient Model · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Reliable and Efficient Image Cropping: A Grid Anchor Based Approach · CVPR 2019 |
Image and video processing › image enhancement
color and tone enhancement |
0.6 | 1 | 2022 | Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-Time · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computational photography and imaging
image-adaptive 3d lookup table |
0.6 | 1 | 2022 | Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-Time · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Image and video processing
image restoration |
0.5 | 1 | 2021 | CameraNet: A Two-Stage Framework for Effective Camera ISP Learning · IEEE Trans. Image Process. 2021 |
Computational photography and imaging
image signal processing |
0.5 | 1 | 2021 | CameraNet: A Two-Stage Framework for Effective Camera ISP Learning · IEEE Trans. Image Process. 2021 |
Image and video processing › image restoration › degradation removal
low-light image restoration |
0.5 | 1 | 2021 | CameraNet: A Two-Stage Framework for Effective Camera ISP Learning · IEEE Trans. Image Process. 2021 |
Image and video coding › image quality assessment
image aesthetics assessment |
0.4 | 1 | 2020 | A Unified Probabilistic Formulation of Image Aesthetic Assessment · IEEE Trans. Image Process. 2020 |
Image and video coding
image quality assessment |
0.4 | 1 | 2020 | A Unified Probabilistic Formulation of Image Aesthetic Assessment · IEEE Trans. Image Process. 2020 |
Image and video coding › image compression
learned image compression |
0.4 | 1 | 2020 | Learning a Single Tucker Decomposition Network for Lossy Image Compression With Multiple Bits-per-Pixel Rates · IEEE Trans. Image Process. 2020 |
Image and video coding › image compression
lossy image compression |
0.4 | 1 | 2020 | Learning a Single Tucker Decomposition Network for Lossy Image Compression With Multiple Bits-per-Pixel Rates · IEEE Trans. Image Process. 2020 |
Image and video coding
variable-rate coding |
0.4 | 1 | 2020 | Learning a Single Tucker Decomposition Network for Lossy Image Compression With Multiple Bits-per-Pixel Rates · IEEE Trans. Image Process. 2020 |
Image and video processing › super-resolution
image super-resolution |
0.4 | 1 | 2019 | Toward Real-World Single Image Super-Resolution: A New Benchmark and a New Model · ICCV 2019 |
Image and video processing › super-resolution › image super-resolution
real-world image super-resolution |
0.4 | 1 | 2019 | Toward Real-World Single Image Super-Resolution: A New Benchmark and a New Model · ICCV 2019 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.4 | 1 | 2019 | Toward Real-World Single Image Super-Resolution: A New Benchmark and a New Model · ICCV 2019 |
Computational photography and imaging
high dynamic range imaging |
0.3 | 1 | 2018 | A Hybrid l1-l0 Layer Decomposition Model for Tone Mapping · CVPR 2018 |
Image and video processing › image decomposition › image separation
layer separation |
0.3 | 1 | 2018 | A Hybrid l1-l0 Layer Decomposition Model for Tone Mapping · CVPR 2018 |
Computational photography and imaging
tone mapping |
0.3 | 1 | 2018 | A Hybrid l1-l0 Layer Decomposition Model for Tone Mapping · CVPR 2018 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2026 | High-Resolution Photo Enhancement in Real-Time: A Laplacian Pyramid Network · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Methods — techniques the papers use, named apart from their topics
spatial-frequency transformer · 2.0laplacian pyramid · 2.03D LUT · 2.0convolutional neural network · 1.9unpaired learning · 0.6pairwise learning · 0.6multi-scale feature learning · 0.6two-stage network · 0.5joint fine-tuning · 0.5nonuniform quantization · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Resolution Photo Enhancement in Real-Time: A Laplacian Pyramid NetworkabstractPhoto enhancement plays a crucial role in augmenting the visual aesthetics of a photograph. In recent years, photo enhancement methods have either focused on enhancement performance, producing powerful models that cannot be deployed on edge devices, or prioritized computational efficiency, resulting in inadequate performance for real-world applications. To this end, this paper introduces a pyramid network called LLF-LUT++, which integrates global and local operators through closed-form Laplacian pyramid decomposition and reconstruction. This approach enables fast processing of high-resolution images while also achieving excellent performance. Specifically, we utilize an image-adaptive 3D LUT that capitalizes on the global tonal characteristics of downsampled images, while incorporating two distinct weight fusion strategies to achieve coarse global image enhancement. To implement this strategy, we designed a spatial-frequency transformer weight predictor that effectively extracts the desired distinct weights by leveraging frequency features. Additionally, we apply local Laplacian filters to adaptively refine edge details in high-frequency components. After meticulously redesigning the network structure and transformer model, LLF-LUT++ not only achieves a 2.64 dB improvement in PSNR on the HDR+ dataset, but also further reduces runtime, with 4 K resolution images processed in just 13 ms on a single GPU. Extensive experimental results on two benchmark datasets further show that the proposed approach performs favorably compared to state-of-the-art methods. Feng Zhang 0039, Haoyou Deng, Lida Li, Qingbo Lu, Zisheng Cao, Minchen Wei, Changxin Gao, Nong Sang, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-TimeabstractRecent years have witnessed the increasing popularity of learning based methods to enhance the color and tone of photos. However, many existing photo enhancement methods either deliver unsatisfactory results or consume too much computational and memory resources, hindering their application to high-resolution images (usually with more than 12 megapixels) in practice. In this paper, we learn image-adaptive 3-dimensional lookup tables (3D LUTs) to achieve fast and robust photo enhancement. 3D LUTs are widely used for manipulating color and tone of photos, but they are usually manually tuned and fixed in camera imaging pipeline or photo editing tools. We, for the first time to our best knowledge, propose to learn 3D LUTs from annotated data using pairwise or unpaired learning. More importantly, our learned 3D LUT is image-adaptive for flexible photo enhancement. We learn multiple basis 3D LUTs and a small convolutional neural network (CNN) simultaneously in an end-to-end manner. The small CNN works on the down-sampled version of the input image to predict content-dependent weights to fuse the multiple basis 3D LUTs into an image-adaptive one, which is employed to transform the color and tone of source images efficiently. Our model contains less than 600K parameters and takes less than 2 ms to process an image of 4K resolution using one Titan RTX GPU. While being highly efficient, our model also outperforms the state-of-the-art photo enhancement methods by a large margin in terms of PSNR, SSIM and a color difference metric on two publically available benchmark datasets. Code will be released at https://github.com/HuiZeng/Image-Adaptive-3DLUT. Hui Zeng 0001, Jianrui Cai, Lida Li, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Grid Anchor Based Image Cropping: A New Benchmark and An Efficient ModelabstractImage cropping aims to improve the composition as well as aesthetic quality of an image by removing extraneous content from it. Most of the existing image cropping databases provide only one or several human-annotated bounding boxes as the groundtruths, which can hardly reflect the non-uniqueness and flexibility of image cropping in practice. The employed evaluation metrics such as intersection-over-union cannot reliably reflect the real performance of a cropping model, either. This work revisits the problem of image cropping, and presents a grid anchor based formulation by considering the special properties and requirements (e.g., local redundancy, content preservation, aspect ratio) of image cropping. Our formulation reduces the searching space of candidate crops from millions to no more than ninety. Consequently, a grid anchor based cropping benchmark is constructed, where all crops of each image are annotated and more reliable evaluation metrics are defined. To meet the practical demands of robust performance and high efficiency, we also design an effective and lightweight cropping model. By simultaneously considering the region of interest and region of discard, and leveraging multi-scale information, our model can robustly output visually pleasing crops for images of different scenes. With less than 2.5M parameters, our model runs at a speed of 200 FPS on one single GTX 1080Ti GPU and 12 FPS on one i7-6800K CPU. The code is available at: https://github.com/HuiZeng/Grid-Anchor-based-Image-Cropping-Pytorch. Hui Zeng 0001, Lida Li, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | CameraNet: A Two-Stage Framework for Effective Camera ISP LearningabstractTraditional image signal processing (ISP) pipeline consists of a set of cascaded image processing modules onboard a camera to reconstruct a high-quality sRGB image from the sensor raw data. Recently, some methods have been proposed to learn a convolutional neural network (CNN) to improve the performance of traditional ISP. However, in these works usually a CNN is directly trained to accomplish the ISP tasks without considering much the correlation among the different components in an ISP. As a result, the quality of reconstructed images is barely satisfactory in challenging scenarios such as low-light imaging. In this paper, we firstly analyze the correlation among the different tasks in an ISP, and categorize them into two weakly correlated groups: restoration and enhancement. Then we design a two-stage network, called CameraNet, to progressively learn the two groups of ISP tasks. In each stage, a ground truth is specified to supervise the subnetwork learning, and the two subnetworks are jointly fine-tuned to produce the final output. Experiments on three benchmark datasets show that the proposed CameraNet achieves consistently compelling reconstruction quality and outperforms the recently proposed ISP learning methods. Zhetong Liang, Jianrui Cai, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2020 | Learning a Single Tucker Decomposition Network for Lossy Image Compression With Multiple Bits-per-Pixel RatesabstractLossy image compression (LIC), which aims to utilize inexact approximations to represent an image more compactly, is a classical problem in image processing. Recently, deep convolutional neural networks (CNNs) have achieved interesting results in LIC by learning an encoder-quantizer-decoder network from a large amount of data. However, existing CNN-based LIC methods generally train a network for a specific bits-perpixel (bpp). Such a "one-network-per-bpp" problem limits the generality and flexibility of CNNs to practical LIC applications. In this paper, we propose to learn a single CNN which can perform LIC at multiple bpp rates. A simple yet effective Tucker Decomposition Network (TDNet) is developed, where there is a novel tucker decomposition layer (TDL) to decompose a latent image representation into a set of projection matrices and a core tensor. By changing the rank of core tensor and its quantization, we can easily adjust the bpp rate of latent image representation within a single CNN. Furthermore, an iterative non-uniform quantization scheme is presented to optimize the quantizer, and a coarse-to-fine training strategy is introduced to reconstruct the decompressed images. Extensive experiments demonstrate the state-of-the-art compression performance of TDNet in terms of both PSNR and MS-SSIM indices. Jianrui Cai, Zisheng Cao, Lei Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2020 | A Unified Probabilistic Formulation of Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has been attracting considerable attention in recent years due to the explosive growth of digital photography in Internet and social networks. The IAA problem is inherently challenging, owning to the ineffable nature of the human sense of aesthetics and beauty, and its close relationship to understanding pictorial content. Three different approaches to framing and solving the problem have been posed: binary classification, average score regression and score distribution prediction. Solutions that have been proposed have utilized different types of aesthetic labels and loss functions to train deep IAA models. However, these studies ignore the fact that the three different IAA tasks are inherently related. Here, we reveal that the use of the different types of aesthetic labels can be developed within the same statistical framework, which we use to create a unified probabilistic formulation of all the three IAA tasks. This unified formulation motivates the use of an efficient and effective loss function for training deep IAA models to conduct different tasks. We also discuss the problem of learning from a noisy raw score distribution which hinders network performance. We then show that by fitting the raw score distribution to a more stable and discriminative score distribution, we are able to train a single model which is able to obtain highly competitive performance on all three IAA tasks. Extensive qualitative analysis and experimental results on image aesthetic benchmarks validate the superior performance afforded by the proposed formulation. The source code is available at. Hui Zeng 0001, Zisheng Cao, Lei Zhang 0006, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2019 | Reliable and Efficient Image Cropping: A Grid Anchor Based ApproachabstractImage cropping aims to improve the composition as well as aesthetic quality of an image by removing extraneous content from it. Existing image cropping databases provide only one or several human-annotated bounding boxes as the groundtruth, which cannot reflect the non-uniqueness and flexibility of image cropping in practice. The employed evaluation metrics such as intersection-over-union cannot reliably reflect the real performance of cropping models, either. This work revisits the problem of image cropping, and presents a grid anchor based formulation by considering the special properties and requirements (e.g., local redundancy, content preservation, aspect ratio) of image cropping. Our formulation reduces the searching space of candidate crops from millions to less than one hundred. Consequently, a grid anchor based cropping benchmark is constructed, where all crops of each image are annotated and more reliable evaluation metrics are defined. We also design an effective and lightweight network module, which simultaneously considers the region of interest and region of discard for more accurate image cropping. Our model can stably output visually pleasing crops for images of different scenes and run at a speed of 125 FPS. Hui Zeng 0001, Lida Li, Zisheng Cao, Lei Zhang 0006 |
CVPR | 3 |
| 2019 | Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelabstractMost of the existing learning-based single image super-resolution (SISR) methods are trained and evaluated on simulated datasets, where the low-resolution (LR) images are generated by applying a simple and uniform degradation (i.e., bicubic downsampling) to their high-resolution (HR) counterparts. However, the degradations in real-world LR images are far more complicated. As a consequence, the SISR models trained on simulated data become less effective when applied to practical scenarios. In this paper, we build a real-world super-resolution (RealSR) dataset where paired LR-HR images on the same scene are captured by adjusting the focal length of a digital camera. An image registration algorithm is developed to progressively align the image pairs at different resolutions. Considering that the degradation kernels are naturally non-uniform in our dataset, we present a Laplacian pyramid based kernel prediction network (LP-KPN), which efficiently learns per-pixel kernels to recover the HR image. Our extensive experiments demonstrate that SISR models trained on our RealSR dataset deliver better visual quality with sharper edges and finer textures on real-world scenes than those trained on simulated datasets. Though our RealSR dataset is built by using only two cameras (Canon 5D3 and Nikon D810), the trained model generalizes well to other camera devices such as Sony a7II and mobile phones. Jianrui Cai, Hui Zeng 0001, Hongwei Yong, Zisheng Cao, Lei Zhang 0006 |
ICCV | 4 |
| 2018 | A Hybrid l1-l0 Layer Decomposition Model for Tone MappingabstractTone mapping aims to reproduce a standard dynamic range image from a high dynamic range image with visual information preserved. State-of-the-art tone mapping algorithms mostly decompose an image into a base layer and a detail layer, and process them accordingly. These methods may have problems of halo artifacts and over-enhancement, due to the lack of proper priors imposed on the two layers. In this paper, we propose a hybrid ℓ1-ℓ0decomposition model to address these problems. Specifically, an ℓ1sparsity term is imposed on the base layer to model its piecewise smoothness property. An ℓ0sparsity term is imposed on the detail layer as a structural prior, which leads to piecewise constant effect. We further propose a multiscale tone mapping scheme based on our layer decomposition model. Experiments show that our tone mapping algorithm achieves visually compelling results with little halo artifacts, outperforming the state-of-the-art tone mapping algorithms in both subjective and objective evaluations. Zhetong Liang, Jun Xu 0019, David Zhang 0001, Zisheng Cao, Lei Zhang 0006 |
CVPR | 4 |