Han Zhu 0003

dblp:02/1555-3 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-7763-8342ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Contrastive Online Distillation based on Information Bottleneck for Image Super-Resolution
Han Zhu 0003, Jingyun Liu, Zhenzhong Chen 0001
PCS1
2025 Meta-RawResampler: Raw image rescaling based on pattern guidance
Jingyun Liu, Han Zhu 0003, Daiqin Yang, Zhenzhong Chen 0001, Shan Liu 0001
J. Vis. Commun. Image Represent.2
2025 Information Bottleneck Based Self-Distillation: Boosting Lightweight Network for Real-World Super-Resolution
abstract
Most existing single-image super-resolution (SISR) methods focus on addressing predefined uniform degradations, such as bicubic. However, these methods often perform poorly in real-world scenarios due to complicated and varying realistic degradations. In this paper, we propose a novel information bottleneck-based self-distillation method (IBSD) to boost lightweight networks for real-world image super-resolution. The proposed IBSD leverages the principle of information bottleneck to guide SR networks to learn invariant correlations from low-resolution (LR) to high-resolution (HR) across various degradations, thereby improving their generalization capacity. Specifically, the target super-resolution network (i.e., student) is interpreted as a Markov chain, and the distillation process is carried out through two modules. Mutual information (MI) estimation networks are used to quantify the mutual information between adjacent nodes within the Markov chain. To enhance robustness against blur and noise in real-world scenarios, an auxiliary loss with a progressive soft target is employed to better identify what is effective for reconstruction in the high-frequency domain. Minimizing the mutual information while preserving task-relevant features can help remove information that reflects spurious correlations between specific degradations and reconstructed targets. Experiments conducted on real-world image super-resolution datasets demonstrate that our proposed method can significantly improve the performance of recent lightweight SR models without adding any extra inference complexity, and it outperforms existing self-distillation approaches. Code is publicly available athttps://github.com/hanzhu1121/IBSD.
Han Zhu 0003, Zhenzhong Chen 0001, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Mutual Guidance Distillation for Joint Demosaicking and Denoising of Raw Images
abstract
Joint demoisaicking and denoising (JDD) serves as an initial step of image signal processing (ISP), whose performance significantly influences the subsequent operations like image processing and compression. Although deep learning-based methods have demonstrated remarkable performance in JDD, they suffer from heavy computational cost and memory occupation, hindering their deployment of resource-constrained devices. To reach a compromise between performance and complexity, we propose a novel knowledge distillation method named Mutual Guidance Distillation (MGD). It aims at promoting the accuracy of lightweight JDD networks (student) by imitating the mutual guidance procedure between color components from a cumbersome network (teacher). The procedure is achieved by computing spatial correlation between representations of red, blue and green components from different layers. The higher the correlation, the greater the influences of each component representation on the restoration of other components. Then the single-branch student network is trained to mimic the correlation of a multi-branch teacher network. Since the multi-branch teacher derives advantages from component mutual guidance and achieves outstanding performance, student networks can be enhanced under the instruction of the teacher. Experimental results on various joint demosaicking and denoising datasets demonstrate that MGD outperforms several state-of-the-art distillation methods quantitatively and qualitatively, and visual results illustrate that MGD effectively mitigates color artifacts, even on hard cases from MIT moiré dataset.
Jingyun Liu, Han Zhu 0003, Zhenzhong Chen 0001, Shan Liu 0001
PCS2
2024 Swin Transformer-Based In-Loop Filter for VVC Intra Coding
abstract
As an emerging video coding standard, H.266NVC is widely recognized for efficiently reducing bit rate and im-proving compression ratio. However, adopting a block-based hybrid coding framework, VVC still encounters the challenge of compression artifacts, which cause a serious impact on both subjective and objective quality. Recently, the emergence of deep learning techniques has boosted the development of neural network-based in-loop filter methods, most of which adopt convolutional neural networks (CNNs) as their backbone. However, the local receptive field of CNN limits its ability to extract global features, which leads to a performance bottleneck in the CNN-base in-loop filter. In contrast, this paper offers a swin transformer-based in-loop filter for VVC, which has a more flexible and global receptive field than CNN. Specifically, we first introduce and optimize the swin transformer-based network for the video compression task. In addition, to ensure the rate-distortion performance, a coding tree unit-level flag is designed to involve the network in the RDO decision-making process at the block level. The experimental results show that compared with VTM-ll.0, the proposed swin transformer-based in-loop filter method can achieve an average of 6.77%, 14.62%, 14.97% and 6.88%, 17.24%, 17.50% Bjontegaard Delta (BD)-Bitrate savings under All-intra (AI) configurations for Y, U and V components when using PSNR and MS-SSIM as the quality metric, respectively.
Tong Ouyang, Huairui Wang, Han Zhu 0003, Zhenzhong Chen 0001
PCS4
2024 Deep Reference Frame Generation Method for VVC Inter Prediction Enhancement
abstract
In video coding, inter prediction aims to reduce temporal redundancy by using previously encoded frames as references. The quality of reference frames is crucial to the performance of inter prediction. This paper presents a deep reference frame generation method to optimize the inter prediction in Versatile Video Coding (VVC). Specifically, reconstructed frames are sent to a well-designed frame generation network to synthesize a picture similar to the current encoding frame. The synthesized picture serves as an additional reference frame inserted into the reference picture list (RPL) to provide a more reliable reference for subsequent motion estimation (ME) and motion compensation (MC). The frame generation network employs optical flow to predict motion precisely. Moreover, an optical flow reorganization strategy is proposed to enable bi-directional and uni-directional predictions with only a single network architecture. To reasonably apply our method to VVC, we further introduce a normative modification of the temporal motion vector prediction (TMVP). Integrated into the VVC reference software VTM-15.0, the deep reference frame generation method achieves coding efficiency improvements of 5.22%, 3.61%, and 3.83% for the Y component under random access (RA), low delay B (LDB), and low delay P (LDP) configurations, respectively. The proposed method has been discussed in Joint Video Exploration Team (JVET) meeting and is currently part of Exploration Experiments (EE) for further study.
Jianghao Jia, Yuantong Zhang, Han Zhu 0003, Zhenzhong Chen 0001, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Learning knowledge representation with meta knowledge distillation for single image super-resolution
Han Zhu 0003, Zhenzhong Chen 0001, Shan Liu 0001
J. Vis. Commun. Image Represent.1
2023 Optical Flow Reusing for High-Efficiency Space-Time Video Super Resolution
abstract
In this paper, we consider the task of space-time video super-resolution (ST-VSR), which can increase the spatial resolution and frame rate for a given video simultaneously. Despite the remarkable progress of recent methods, most of them still suffer from high computational costs and inefficient long-range information usage. To alleviate these problems, we propose a Bidirectional Recurrence Network (BRN) with the optical-flow-reuse strategy to better use temporal knowledge from long-range neighboring frames for high-efficiency reconstruction. Specifically, an efficient and memory-saving multi-frame motion utilization strategy is proposed by reusing the intermediate flow of adjacent frames, which considerably reduces the computation burden of frame alignment compared with traditional LSTM-based designs. In addition, the proposed hidden state in BRN is updated by the reused optical flow and refined by the Feature Refinement Module (FRM) for further optimization. Moreover, by utilizing intermediate flow estimation, the proposed method can inference non-linear motion and restore details better. Extensive experiments demonstrate that our optical-flow-reuse-based bidirectional recurrent network (OFR-BRN) is superior to state-of-the-art methods in accuracy and efficiency. Codes are available on URL:https://github.com/hahazh/OFR-BRN
Yuantong Zhang, Huairui Wang, Han Zhu 0003, Zhenzhong Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Visual Quality Evaluation for Semantic Segmentation: Subjective Assessment Database and Objective Assessment Measure
abstract
To promote the applications of semantic segmentation, quality evaluation is important to assess different algorithms and guide their development and optimization. In this paper, we establish a subjective semantic segmentation quality assessment database based on the stimulus-comparison method. Given that the database reflects the relative quality of semantic segmentation result pairs, we adopt a robust regression mapping model to explore the relationship between subjective assessment and objective distance. With the help of the regression model, we can examine whether objective metrics coincide with subjective judgement. In addition, we propose a novel relative quality prediction network (RQPN) based on Siamese CNN as a new objective metric. The metric is trained by our subjective assessment database and can be applied to evaluate the performances of semantic segmentation algorithms, even if the algorithms were not used to build the database. Experiments are conducted to show the advance and the reliability of our database and demonstrate that results predicted by RQPN are more consistent to subjective assessment than existing objective metrics.
Zhenzhong Chen 0001, Han Zhu 0003
IEEE Trans. Image Process.2
2018 Dense Residual Convolutional Neural Network based In-Loop Filter for HEVC
abstract
In-loop filtering is a key technique in the state of art video coding standard that plays a significant role in suppressing compression artifacts. Motivated by the latest advances of deep learning, in this paper, we design a dense residual convolutional neural network (DRN) based in-loop filter for High Efficiency Video Coding (HEVC). Taking advantage of both dense shortcuts and residual learning, DRN efficiently exploits the multi-level features to restore the high quality image from the degraded one. Bottleneck layers are employed in DRN in order to adaptively fuse the hierarchical features and saving the computational resources at the same time. Experimental results show that the proposed DRN based in-loop filter can further boost the coding performance, which provides 6.9% BD-rate reduction on average compared to the HEVC baseline. In addition, the proposed DRN outperforms previous CNN based in-loop filters.
Yingbin Wang, Han Zhu 0003, Yiming Li 0001, Zhenzhong Chen 0001, Shan Liu 0001
VCIP2