EDBT 2026 Demo / reviewers in the wild / expert
Zhen Cheng 0002
dblp:70/3835-2
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0002-7655-9930ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Toward DNN of LUTs: Learning Efficient Image Restoration With Multiple Look-Up TablesabstractThe widespread usage of high-definition screens on edge devices stimulates a strong demand for efficient image restoration algorithms. The way of caching deep learning models in a look-up table (LUT) is recently introduced to respond to this demand. However, the size of a single LUT grows exponentially with the increase of its indexing capacity, which restricts its receptive field and thus the performance. To overcome this intrinsic limitation of the single-LUT solution, we propose a universal method to construct multiple LUTs like a neural network, termed MuLUT. First, we devise novel complementary indexing patterns, as well as a general implementation for arbitrary patterns, to construct multiple LUTs in parallel. Second, we propose a re-indexing mechanism to enable hierarchical indexing between cascaded LUTs. Finally, we introduce channel indexing to allow cross-channel interaction, enabling LUTs to process color channels jointly. In these principled ways, the total size of MuLUT is linear to its indexing capacity, yielding a practical solution to obtain superior performance with the enlarged receptive field. We examine the advantage of MuLUT on various image restoration tasks, including super-resolution, demosaicing, denoising, and deblocking. MuLUT achieves a significant improvement over the single-LUT solution, e.g., up to 1.1 dB PSNR for super-resolution and up to 2.8 dB PSNR for grayscale denoising, while preserving its efficiency, which is 100× less in energy cost compared with lightweight deep neural networks. Jiacheng Li 0004, Chang Chen 0004, Zhen Cheng 0002, Zhiwei Xiong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Light Field Super-Resolution Using Decoupled Selective MatchingabstractNon-local self-similarity has been well exploited in the single image super-resolution task as an effective prior. However, due to the difficulty of modeling the 4D correspondence globally, the potential of the non-local prior is less revealed for light field (LF) super-resolution. Meanwhile, existing non-local models only utilize the global spatial correspondence, but largely neglect the global geometric correspondence. To address the aforementioned problems, we propose a Decoupled Selective Matching Network (DSMNet) for LF super-resolution, by designing a novel selective matching mechanism to flexibly extract non-local information from specific 4D positions in an LF. Such a mechanism matches the reference patch with several auxiliary patches dynamically searched from predefined windows, which promotes efficiency while improving performance compared to the existing non-local models. Specifically, our DSMNet decouples the whole LF into Sub-Aperture Images (SAIs) and Epipolar Plane Images (EPIs). For each SAI patch, we separately perform the selective matching inside the current SAI and cross different SAIs to exploit the global spatial correspondence efficiently. For each EPI patch, we separately perform the selective matching in EPIs of different orientations to embed robust LF geometric information into features by enhancing EPI textures, which exploits the global geometric correspondence in an efficient manner. Comprehensive experiments validate that DSMNet outperforms state-of-the-art LF super-resolution methods both quantitatively and qualitatively. Zhen Cheng 0002, Zeyu Xiao 0002, Zhiwei Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Space-Time Super-Resolution for Light Field VideosabstractLight field (LF) cameras suffer from a fundamental trade-off between spatial and angular resolutions. Additionally, due to the significant amount of data that needs to be recorded, the Lytro ILLUM, a modern LF camera, can only capture three frames per second. In this paper, we consider space-time super-resolution (SR) for LF videos, aiming at generating high-resolution and high-frame-rate LF videos from low-resolution and low-frame-rate observations. Extending existing space-time video SR methods to this task directly will meet two key challenges: 1) how to re-organize sub-aperture images (SAIs) efficiently and effectively given highly redundant LF videos, and 2) how to aggregate complementary information between multiple SAIs and frames considering the coherence in LF videos. To address the above challenges, we propose a novel framework for space-time super-resolving LF videos for the first time. First, we propose a novel Multi-Scale Dilated SAI Re-organization strategy for re-organizing SAIs into auxiliary view stacks with decreasing resolution as the Chebyshev distance in the angular dimension increases. In particular, the auxiliary view stack with original resolution preserves essential visual details, while the down-scaled view stacks capture long-range contextual information. Second, we propose the Multi-Scale Aggregated Feature extractor and the Angular-Assisted Feature Interpolation module to utilize and aggregate information from the spatial, angular, and temporal dimensions in LF videos. The former aggregates similar contents from different SAIs and frames for subsequent reconstruction in a disparity-free manner at the feature level, whereas the latter interpolates intermediate frames temporally by implicitly aggregating geometric information. Compared to other potential approaches, experimental results demonstrate that the reconstructed LF videos generated by our framework achieve higher reconstruction quality and better preserve the LF parallax structure and temporal consistency. The implementation code is available at https://github.com/zeyuxiao1997/LFSTVSR. Zeyu Xiao 0002, Zhen Cheng 0002, Zhiwei Xiong |
IEEE Trans. Image Process. | 2 |
| 2022 | Degradation-agnostic Correspondence from Resolution-asymmetric StereoabstractIn this paper, we study the problem of stereo matching from a pair of images with different resolutions, e.g., those acquired with a tele-wide camera system. Due to the difficulty of obtaining ground-truth disparity labels in diverse real-world systems, we start from an unsupervised learning perspective. However, resolution asymmetry caused by unknown degradations between two views hinders the effectiveness of the generally assumed photometric consistency. To overcome this challenge, we propose to impose the consistency between two views in a feature space instead of the image space, named feature-metric consistency. Interestingly, we find that, although a stereo matching network trained with the photometric loss is not optimal, its feature extractor can produce degradation-agnostic and matching-specific features. These features can then be utilized to formulate a feature-metric loss to avoid the photometric inconsistency. Moreover, we introduce a self-boosting strategy to optimize the feature extractor progressively, which further strengthens the feature-metric consistency. Experiments on both simulated datasets with various degradations and a self-collected real-world dataset validate the superior performance of the proposed method over existing solutions. Xihao Chen, Zhiwei Xiong, Zhen Cheng 0002, Jiayong Peng, Yueyi Zhang 0001, Zhengjun Zha |
CVPR | 3 |
| 2022 | Towards Real-World HDRTV Reconstruction: A Data Synthesis-Based Approach
Zhen Cheng 0002, Fenglong Song, Chang Chen 0004, Zhiwei Xiong |
ECCV (19) | 1 |
| 2022 | MuLUT: Cooperating Multiple Look-Up Tables for Efficient Image Super-Resolution
Jiacheng Li 0004, Chang Chen 0004, Zhen Cheng 0002, Zhiwei Xiong |
ECCV (18) | 3 |
| 2022 | Domain Adaptive Mitochondria Segmentation via Enforcing Inter-Section Consistency
Wei Huang 0036, Xiaoyu Liu 0006, Zhen Cheng 0002, Yueyi Zhang 0001, Zhiwei Xiong |
MICCAI (4) | 3 |
| 2021 | Light Field Super-Resolution With Zero-Shot LearningabstractDeep learning provides a new avenue for light field super-resolution (SR). However, the domain gap caused by drastically different light field acquisition conditions poses a main obstacle in practice. To fill this gap, we propose a zero-shot learning framework for light field SR, which learns a mapping to super-resolve the reference view with examples extracted solely from the input low-resolution light field itself. Given highly limited training data under the zero-shot setting, however, we observe that it is difficult to train an end-to-end network successfully. Instead, we divide this challenging task into three sub-tasks, i.e., pre-upsampling, view alignment, and multi-view aggregation, and then conquer them separately with simple yet efficient CNNs. Moreover, the proposed framework can be readily extended to finetune the pre-trained model on a source dataset to better adapt to the target input, which further boosts the performance of light field SR in the wild. Experimental results validate that our method not only outperforms classic non-learning-based methods, but also generalizes better to unseen light fields than state-of-the-art deep-learning-based methods when the domain gap is large. Zhen Cheng 0002, Zhiwei Xiong, Chang Chen 0004, Dong Liu 0002, Zhengjun Zha |
CVPR | 1 |
| 2021 | Space-Time Distillation for Video Super-ResolutionabstractCompact video super-resolution (VSR) networks can be easily deployed on resource-limited devices, e.g., smartphones and wearable devices, but have considerable performance gaps compared with complicated VSR networks that require a large amount of computing resources. In this paper, we aim to improve the performance of compact VSR networks without changing their original architectures, through a knowledge distillation approach that transfers knowledge from a complicated VSR network to a compact one. Specifically, we propose a space-time distillation (STD) scheme to exploit both spatial and temporal knowledge in the VSR task. For space distillation, we extract spatial attention maps that hint the high-frequency video content from both networks, which are further used for transferring spatial modeling capabilities. For time distillation, we narrow the performance gap between compact models and complicated models by distilling the feature similarity of the temporal memory cells, which are encoded from the sequence of feature maps generated in the training clips using ConvLSTM. During the training process, STD can be easily incorporated into any network without changing the original network architecture. Experimental results on standard benchmarks demonstrate that, in resource-constrained situations, the proposed method notably improves the performance of existing VSR networks without increasing the inference time. Zeyu Xiao 0002, Xueyang Fu, Jie Huang 0017, Zhen Cheng 0002, Zhiwei Xiong |
CVPR | 4 |
| 2020 | Light Field Super-Resolution By Jointly Exploiting Internal and External SimilaritiesabstractLight field images taken by plenoptic cameras often have a tradeoff between spatial and angular resolutions. In this paper, we propose a novel spatial super-resolution approach for light field images by jointly exploiting internal and external similarities. The internal similarity refers to the correlations across the angular dimensions of the 4D light field itself, while the external similarity refers to the cross-scale correlations learned from an external light field dataset. Specifically, we advance the classic projection-based method that exploits the internal similarity by introducing the intensity consistency checking criterion and a back-projection refinement, while the external correlation is learned by a CNN-based method which aggregates all warped high-resolution sub-aperture images upsampled from the low-resolution input using a single image super-resolution method. By analyzing the error distributions of the above two methods and investigating the upperbound of combining them, we find that the internal and external similarities are complementary to each other. Accordingly, we further propose a pixel-wise adaptive fusion network to take advantage of both their merits by learning a weighting matrix. Experimental results on both synthetic and real-world light field datasets validate the superior performance of the proposed approach over the state-of-the-arts. Zhen Cheng 0002, Zhiwei Xiong, Dong Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Tensor-Based Light Field Denoising by Integrating Super-ResolutionabstractLight field, a promising representation to describe the scene appearance, is susceptible to various noise due to the current sensor design. This paper proposes a novel tensor-based denoising method for the 4D light field that consists of two main steps. First, we generalize the intrinsic tensor sparsity measure to light field images by exploiting the nonlocal similarity across the spatial and angular dimensions. Second, we further exploit the spatial-angular correlation by integrating light field super-resolution into the denoising process to eliminate the sub-pixel misalignment of different views. After a back-projection from the refined high-resolution central view under an intensity consistency criteria, the denoising performance for the light field can be boosted. Experimental results validate the superior performance of the proposed method in terms of both PSNR and visual quality on the HCI light field dataset. Na Qi, Zhen Cheng 0002, Dong Liu 0002, Qing Ling 0001, Zhiwei Xiong |
ICIP | 3 |
| 2017 | Light field super-resolution using internal and external similaritiesabstractThis paper presents a novel super-resolution method for light field images by jointly exploiting internal and external similarities. The internal similarity refers to the correlations that exist across the angular dimensions of the 4D light field itself, while the external similarity refers to the correlations learned from a conventional 2D image dataset. Our key observation is that the internal and external similarities are complementary to each other, and we propose a depth-adaptive fusion scheme to take advantage of both their merits. Moreover, we improve the traditional projection-based method that exploits the internal similarity, by introducing a back-projection refinement and getting rid of the dependency on camera parameters. Experimental results on a variety of light field images validate the superior performance of the proposed method. Zhiwei Xiong, Zhen Cheng 0002, Jiayong Peng, Hanzhi Fan, Dong Liu 0002, Feng Wu 0001 |
ICIP | 2 |