EDBT 2026 Demo / reviewers in the wild / expert
Zhe Zhang 0041
dblp:87/5809-41
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-8772-2107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lightweight stereo image super-resolution via adaptive pruning and bridge distillation
Zhe Zhang 0041, Bingzheng Liu, Lei Chen 0091, Pengzhi Li, Yidan Zhang 0002, Jianjun Lei 0001 |
Knowl. Based Syst. | 1 |
| 2026 | Depth-Aware Transformer for Aerial LocalizationabstractRecently, deep learning-based visual localization has gained significant attention and made remarkable advancements. Although previous visual localization methods have obtained promising performance on indoor or outdoor street scenes, there have been few attempts at visual localization on aerial scenes. In this article, a depth-aware aerial localization transformer (DALTR) is proposed to learn camera poses in real-world aerial scenes assisted by the depth map. To improve the ability of network to perceive on aerial scenes, a multi-level depth embedding transformer module is presented by adaptively incorporating depth information into multiple levels of transformer. In addition, to encourage the piece-wise smooth geometric characteristic of the scene coordinates, a depth-guided smoothness constraint is developed to provide additional supervision for scene coordinate regression. Extensive experimental results on aerial localization benchmark datasets demonstrate that the proposed DALTR achieves superior aerial localization performance. Jianjun Lei 0001, Duohui Tu, Bo Peng 0007, Zhe Zhang 0041, Chong Wu 0004, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Mamba-Based Global Correlation Learning for Light Field Spatial Super-ResolutionabstractLight field spatial super-resolution is a challenging task due to complex 4D light field structure. How to explore the global correlation of the light field to improve the reconstruction quality remains a key problem in LF-SSR. In this paper, a novel light field spatial super-resolution method with Mamba-based global correlation learning is proposed, which effectively explores pixel-wise correlation within and across different subspaces based on Mamba block. Specifically, a tri-subspace scanned feature enhancement module is proposed to capture the long-range dependency within each 2D light field subspace with the specially-designed scanning schemes. Besides, a Mamba-based spatial-angular information fusion module is designed to mine the global complementary information among the spatial, angular, and EPI features for light field image reconstruction. Extensive experiments show that the proposed method is superior to other advanced light field super-resolution methods. Ruoxi Li, Zhe Zhang 0041 |
ICIP | 4 |
| 2025 | Advancing Real-World Stereoscopic Image Super-Resolution via Vision-Language ModelabstractRecent years have witnessed the remarkable success of the vision-language model in various computer vision tasks. However, how to exploit the semantic language knowledge of the vision-language model to advance real-world stereoscopic image super-resolution remains a challenging problem. This paper proposes a vision-language model-based stereoscopic image super-resolution (VLM-SSR) method, in which the semantic language knowledge in CLIP is exploited to facilitate stereoscopic image SR in a training-free manner. Specifically, by designing visual prompts for CLIP to infer the region similarity, a prompt-guided information aggregation mechanism is presented to capture inter-view information among relevant regions between the left and right views. Besides, driven by the prior knowledge of CLIP, a cognition prior-driven iterative enhancing mechanism is presented to optimize fuzzy regions adaptively. Experimental results on four datasets verify the effectiveness of the proposed method. Zhe Zhang 0041, Jianjun Lei 0001, Bo Peng 0007, Liying Xu, Qingming Huang |
IEEE Trans. Image Process. | 1 |
| 2025 | Advancing Generalizable Occlusion Modeling for Neural Human Radiance FieldabstractGeneralizable human neural rendering aims to render the target views of the human body by leveraging source views and the skinned multi-person linear (SMPL) model. Despite exhibiting promising performance, the target views rendered by previous methods usually contain corrupted parts of the human body. Two primary challenges hinder high-quality human neural rendering. These challenges involve non-correspondences between 2D pixels and 3D SMPL vertices induced by self-occlusion of the human body and erroneous appearance predictions caused by occlusion between the source and target views. To solve these two challenges, we propose an advancing generalizable occlusion modeling method for the neural human radiance field, in which the hurdles from the self-occlusion of the human body and the occlusion between source and target views are explored and solved. Specifically, to alleviate the non-correspondence problem induced by self-occlusion, a geometry perception module is designed to obtain 3D geometric representations of SMPL vertices, enabling the prediction of accurate density values. Furthermore, a visibility aggregation module is designed to estimate the visibility maps with respect to different source views by utilizing the predicted density. Then, the complementary information among multiple source views is integrated with the support of the visibility maps in the visibility aggregation module, thus effectively addressing the occlusion between views. Experiments on the ZJU-MoCap and THUman datasets show that the proposed method achieves promising performance compared with the existing state-of-the-art methods. Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang |
IEEE Trans. Multim. | 4 |
| 2025 | Adaptive Multi-Exposure Image Correction via Joint Lightness and Structure AwarenessabstractIn order to alleviate the impact of ambient light on the quality of captured images, correcting multi-exposure images has become a popular topic. Most existing multi-exposure image correction methods mainly focus on the adjustment of lightness levels, but ignore the significant issue of structural information loss in incorrectly exposed images. Taking into consideration both lightness adjustment and structural reconstruction, this article proposes an adaptive multi-exposure image correction network by jointly exploring the lightness and structure information, named LSANet. Specifically, the proposed LSANet first extracts lightness and structure representations of the input image in the frequency domain, and then performs exposure level adjustment and structure detail reconstruction based on the lightness and structure representations. In the proposed network, the lightness- and structure-aware adaptive module is designed to achieve adaptive correction by predicting dynamic kernels under the guidance of the lightness and structure representations. Experimental results on the widely used ME and SICE datasets demonstrate that the proposed LSANet achieves excellent performance and generates images with well-exposed levels and rich structural details. Bo Peng 0007, Jia Zhang 0025, Zhe Zhang 0041, Liying Xu, Qingming Huang, Tao Wang 0119, Jianjun Lei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Unsupervised Single-View Synthesis Network via Style Guidance and Prior DistillationabstractView synthesis aims to learn a view transformation and synthesize the target views from a single or multiple source views. Although previous view synthesis methods have obtained promising performance, they heavily rely on the supervision of the target view. In this paper, we propose an unsupervised single-view synthesis network (USVS-Net) to learn the view transformation without the supervision of the target view. Specifically, with the usage of only a single source view, a style-guidance view synthesis model is proposed to learn an intrinsic representation, which intends to describe the object from a reference pose. With the intrinsic representation, the view transformation is learned to boost the learning of the unsupervised single-view synthesis. Then, taking the style-guidance view synthesis model as the teacher, a prior-distillation view synthesis model is further presented as the student to learn a more direct view transformation. By utilizing the proposed method, high-quality target views are synthesized in a time-efficient manner. Experiments on both synthetic and real-scene datasets show that despite the lack of supervision of the target view, the proposed method achieves promising results compared with the existing view synthesis methods. Bingzheng Liu, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang, Nam Ling, Jianjun Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Recurrent Interaction Network for Stereoscopic Image Super-ResolutionabstractRecently, deep learning-based stereoscopic image super-resolution has attracted extensive attention and made great progress. However, existing methods have not adequately explored the inter-view dependency among two-view multi-level features. In this paper, a recurrent interaction network for stereoscopic image super-resolution (RISSRnet) is proposed to learn the inter-view dependency. To efficiently utilize the relationship between the two views, a recurrent interaction module is designed to achieve recurrent interaction among two-view multi-level features from the regrouped sequences, which are generated by a coupled queue-regroup mechanism. In addition, to recursively enhance features in the recurrent interaction module, an iterative propagation strategy is developed for sufficient interaction. Extensive experimental results demonstrate the effectiveness and superiority of the proposed RISSRnet. Zhe Zhang 0041, Bo Peng 0007, Jianjun Lei 0001, Haifeng Shen, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Transferring knowledge from monocular completion for self-supervised monocular depth estimation
Bingzheng Liu, Liying Xu, Zhe Zhang 0041 |
Multim. Tools Appl. | 5 |
| 2022 | Deep region segmentation-based intra prediction for depth video coding
Jing Zhang 0017, Yonghong Hou, Zhe Zhang 0041, Dengchao Jin, Peihan Zhang, Ge Li 0002 |
Multim. Tools Appl. | 3 |
| 2022 | LVE-S2D: Low-Light Video Enhancement From Static to DynamicabstractRecently, deep-learning-based low-light video enhancement methods have drawn wide attention and achieved remarkable performance. However, limited by the difficulty in collecting dynamic low-light and well-lighted video pairs in real scenes, how to construct video sequences for supervised learning and design a low-light enhancement network for real dynamic video remains a challenge. In this paper, we propose a simple yet effective low-light video enhancement method (LVE-S2D), which generates dynamic video training pairs from static videos, and enhances the low-light video by mining dynamic temporal information. To obtain low-light and well-lighted video pairs, a sliding window-based dynamic video generation mechanism is designed to produce pseudo videos with rich dynamic temporal information. Then, a siamese dynamic low-light video enhancement network is presented, which effectively utilizes temporal correlation between adjacent frames to enhance the video frames. Extensive experimental results demonstrate that the proposed method not only achieves superior performance on static low-light videos, but also outperforms the state-of-the-art methods on real dynamic low-light videos. Bo Peng 0007, Jianjun Lei 0001, Zhe Zhang 0041, Nam Ling, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Deep Stereoscopic Image Super-Resolution via Interaction ModuleabstractDeep learning-based methods have achieved remarkable performance in single image super-resolution. However, these methods cannot be effectively applied in stereoscopic image super-resolution without considering the characteristics of stereoscopic images. In this article, an interaction module-based stereoscopic image super-resolution network (IMSSRnet) is proposed to effectively utilize the correlation information in stereoscopic images. The key insight of the network lies with how to explore the complementary information of one view to help the reconstruction of another view. Thus, an interaction module is designed to acquire the enhanced features by utilizing complementary information between different views. Specifically, the interaction module is composed of a series of interaction units with a residual structure. In addition, the single image features of left and right views are obtained by a spatial feature extraction module, which can be realized by any existing single image super-resolution models. In order to obtain high-quality stereoscopic images, a gradient loss is introduced to preserve the texture details in a view, and a disparity loss is developed to constrain the disparity relationship between different views. Experimental results demonstrate that the proposed method achieves a promising performance and outperforms the state-of-the-art methods. Jianjun Lei 0001, Zhe Zhang 0041, Xiaoting Fan, Bolan Yang, Ying Chen 0011, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |