Huairui Wang

dblp:283/9057 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2025
0009-0004-2870-6117ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic kernel-based adaptive spatial aggregation for learned image compression
abstract
Learned image compression methods have shown remarkable performance and expansion potential compared to traditional codecs. Currently, there are two mainstream image compression frameworks: one uses stacked convolution and other uses window-based self-attention for transform coding, most of which aggregate valuable dependencies in a fixed spatial range. In this paper, we focus on extending content-adaptive aggregation capability and propose a dynamic kernel-based transform coding. The proposed adaptive aggregation generates kernel offsets to capture valuable information with dynamic sampling convolution to help transform. With the adaptive aggregation strategy and the sharing weights mechanism, our method can achieve promising transform capability with acceptable model complexity. Besides, considering the coarse hyper prior, the channel-wise, and the spatial context, we formulate a generalized entropy model. Based on it, we introduce dynamic kernel in hyper-prior to generate more expressive side information context. Furthermore, we propose an asymmetric sparse entropy model according to the investigation of the spatial and variance characteristics of the grouped latents. The proposed entropy model can facilitate entropy coding to reduce statistical redundancy while maintaining inference efficiency. Experimental results demonstrate that our method achieves superior rate–distortion performance on three benchmarks compared to the state-of-the-art learning-based methods.
Huairui Wang, Nianxiang Fu, Zhenzhong Chen 0001, Shan Liu 0001
J. Vis. Commun. Image Represent.1
2024 Learned Lossless Image Compression Based on Bit Plane Slicing
abstract
Autoregressive Initial Bits (ArIB), a framework that combines subimage autoregression and latent variable models, has shown its advantages in lossless image compression. However, in current methods, the image splitting makes the information of latent variables being uniformly distributed in each subimage, and causes inadequate use of latent variables in addition to posterior collapse. To tackle these issues, we introduce Bit Plane Slicing (BPS), splitting images in the bit plane dimension with the considerations on different importance for latent variables. Thus, BPS provides a more effective representation by arranging subimages with decreasing importance for latent variables. To solve the problem of the increased number of dimensions caused by BPS, we further propose a dimension-tailored autoregressive model that tailors autoregression methods for each dimension based on their characteristics, efficiently capturing the dependencies in plane, space, and color dimensions. As shown in the extensive experimental results, our method demonstrates the superior compression performance with comparable inference speed, when compared to the state-of-the-art normalizing-flow-based methods. The code is at https://github.com/ZZ022/ArIB-BPS.
Huairui Wang, Zhenzhong Chen 0001, Shan Liu 0001
CVPR2
2024 Perceptual-oriented Learned Image Compression with Dynamic Kernel
abstract
In this paper, we extend our prior research named DKIC [1] and propose the perceptual-oriented learned image compression method, PO-DKIC, which is shown in Figure 1 . Specifically, DKIC adopts a dynamic kernel-based dynamic residual block group to enhance the transform coding and an asymmetric space-channel context entropy model to facilitate the estimation of Gaussian parameters. Based on DKIC, PO-DKIC introduces PatchGAN and LPIPS loss to enhance visual quality. Furthermore, to maximize the overall perceptual quality under a rate constraint, we formulate this challenge into a constrained programming problem and use the Linear Integer Programming method for resolution. The experiments demonstrate that our proposed method can generate realistic images with richer textures and finer details when compared to state-of-the-art image compression techniques.
Nianxiang Fu, Junxi Zhang, Huairui Wang, Zhenzhong Chen 0001
DCC3
2024 Swin Transformer-Based In-Loop Filter for VVC Intra Coding
abstract
As an emerging video coding standard, H.266NVC is widely recognized for efficiently reducing bit rate and im-proving compression ratio. However, adopting a block-based hybrid coding framework, VVC still encounters the challenge of compression artifacts, which cause a serious impact on both subjective and objective quality. Recently, the emergence of deep learning techniques has boosted the development of neural network-based in-loop filter methods, most of which adopt convolutional neural networks (CNNs) as their backbone. However, the local receptive field of CNN limits its ability to extract global features, which leads to a performance bottleneck in the CNN-base in-loop filter. In contrast, this paper offers a swin transformer-based in-loop filter for VVC, which has a more flexible and global receptive field than CNN. Specifically, we first introduce and optimize the swin transformer-based network for the video compression task. In addition, to ensure the rate-distortion performance, a coding tree unit-level flag is designed to involve the network in the RDO decision-making process at the block level. The experimental results show that compared with VTM-ll.0, the proposed swin transformer-based in-loop filter method can achieve an average of 6.77%, 14.62%, 14.97% and 6.88%, 17.24%, 17.50% Bjontegaard Delta (BD)-Bitrate savings under All-intra (AI) configurations for Y, U and V components when using PSNR and MS-SSIM as the quality metric, respectively.
Tong Ouyang, Huairui Wang, Han Zhu 0003, Zhenzhong Chen 0001
PCS3
2024 Learned Image Compression with Quantization Error Compensator
abstract
Recent advancements in learned image compression methods have demonstrated superior rate-distortion performance and remarkable potential compared to traditional compression techniques. However, the core operation of quantization, inherent to lossy image compression, introduces errors that can degrade the quality of the reconstructed image. To address this challenge, we propose a novel Quantization Error Compensator (QEC), which leverages spatial context within latent representations and hyperprior information to effectively mitigate the impact of quantization error. Moreover, we propose a tailored quantization error optimization training strategy to further improve rate-distortion performance. Notably, QEC serves as a lightweight, plug-and-play module, offering high flexibility and seamless integration into various learned image compression methods. Extensive experimental results consistently demonstrate significant coding efficiency improvements achievable by incorporating the proposed QEC into state-of-the-art methods, with a slight increase in runtime.
Nianxiang Fu, Zhenzhong Chen 0001, Huairui Wang, Shan Liu 0001
VCIP3
2024 Deep Reference Frame for Versatile Video Coding with Structural Re-parameterization
abstract
In video coding, inter-prediction leverages neigh-boring frames to reduce temporal redundancy. The quality of these reference frames is essential for effective inter-prediction. Although many neural network-based methods have been proposed to improve the quality of reference frames, there is still room for the performance and efficiency trade-off. In this paper, we propose an interpolation diverse branch block (InterDBB) suitable for lightweight frame interpolation networks, which optimizes deep reference frame interpolation networks to improve performance without sacrificing speed and increasing complexity. Specifically, we propose a multi-branch structural reparameterization block without batch normalization. This straightforward yet effective modification ensures training stability and performance improvement. Moreover, we propose a parameterized motion estimation strategy based on different input resolution, to achieve a better trade-off between performance and computational complexity. Experimental results demonstrate that our method achieves -2.01%/-2.87%/-2.44% coding efficiency improvements for Y/U/V components under random access (RA) configuration compared to VTM-11.0_NNVC-5.0.
Chengzhuo Gui, Yuantong Zhang, Weijie Bao, Zhenzhong Chen 0001, Huairui Wang, Shan Liu 0001
VCIP5
2024 Lightweight Arbitrary-Scale Super-Resolution of Remote Sensing Images via Super-Scale Feature
abstract
Remote sensing image (RSI) super-resolution (SR) demands lightweight and efficient methods due to required rapid response in practical applications. Integrating RSIs with different resolutions for diverse applications also requires arbitrary-scale SR, making fix-scaled SR scale inflexible. Therefore, a lightweight SR algorithm capable of arbitrary-scale is necessary for RSIs. To address the above issue, a super-scale feature-based lightweight arbitrary-scale (SFLA) SR network is proposed in this paper. The network consists of two modules: 1) A super-scale feature extraction (SSFE) module that extracts features at both the initial low-resolution (LR) and an integer super-scale resolution, 2) A self-attention implicit function reconstruction (SIFR) module that utilizes multi-layer perceptron (MLP) network and self-attention mechanism for pixel-wise feature mapping to achieve superior SR results. Comparative experiments and ablation results demonstrate that the proposed SFLA algorithm effectively strikes a good balance between performance and complexity.
Yifei Long, Yuantong Zhang, Daiqin Yang, Zhenzhong Chen 0001, Huairui Wang, Shan Liu 0001
VCIP5
2024 Exploring Long- and Short-Range Temporal Information for Learned Video Compression
abstract
Learned video compression methods have gained various interests in the video coding community. Most existing algorithms focus on exploring short-range temporal information and developing strong motion compensation. Still, the ignorance of long-range temporal information utilization constrains the potential of compression. In this paper, we are dedicated to exploiting both long- and short-range temporal information to enhance video compression performance. Specifically, for long-range temporal information exploration, we propose a temporal prior that can be continuously supplemented and updated during compression within the group of pictures (GOP). With the updating scheme, the temporal prior can provide richer mutual information between the overall prior and the current frame for the entropy model, thus facilitating Gaussian parameter prediction. As for the short-range temporal information, we propose a progressive guided motion compensation to achieve robust and accurate compensation. In particular, we design a hierarchical structure to build multi-scale compensation, and by employing optical flow guidance, we generate pixel offsets as motion information at each scale. Additionally, the compensation results at each scale will guide the next scale's compensation, forming a flow-to-kernel and scale-by-scale stable guiding strategy. Extensive experimental results demonstrate that our method can obtain advanced rate-distortion performance compared to the state-of-the-art learned video compression approaches and the latest standard reference software in terms of PSNR and MS-SSIM. The codes are publicly available on: https://github.com/Huairui/LSTVC.
Huairui Wang, Zhenzhong Chen 0001
IEEE Trans. Image Process.1
2024 Learned Video Compression via Heterogeneous Deformable Compensation Network
abstract
Learned video compression has recently emerged as an essential research topic in developing advanced video compression technologies, where motion compensation is considered one of the most challenging issues. In this article, we propose a learned video compression framework via heterogeneous deformable compensation strategy (HDCVC) to tackle the problems of unstable compression performance caused by single-size deformable kernels in downsampled feature domain. More specifically, instead of utilizing optical flow warping or single-size-kernel deformable alignment, the proposed algorithm extracts features from the two adjacent frames to estimate content-adaptive heterogeneous deformable (HetDeform) kernel offsets. Then we align the features extracted from the reference frames with the HetDeform convolution to accomplish motion compensation. Moreover, we design a Spatial-Neighborhood-Conditioned Divisive Normalization (SNCDN) to reduce spatial statistic dependencies and achieve more effective data Gaussianization combined with the Generalized Divisive Normalization. Furthermore, we propose a multi-frame enhanced reconstruction module for exploiting context and temporal information for final quality enhancement. Experimental results indicate that HDCVC achieves superior performance than the recent state-of-the-art learned video compression approaches.
Huairui Wang, Zhenzhong Chen 0001, Chang Wen Chen
IEEE Trans. Multim.1
2023 Efficient Learned Video Compression via Bidirectional Temporal Information Exploration
abstract
In recent years, learned video compression methods have improved substantially. However, most existing algorithms focus on exploring short-term temporal information, thus constraining the compression capability. In this paper, to further boost video compression performance, we exploit both long-and short-range temporal information and consider bidirectional temporal information. For long- and short-range temporal information exploration, we adopt temporal prior and progressive guided motion compensation. Specifically, with the continuously updating strategy, the temporal prior can provide rich mutual information between the overall prior and the current frame, facilitating Gaussian parameter prediction in the entropy model. Besides, the progressive guided motion compensation utilizes flow-to-kernel and scale-by-scale stable guiding strategy, thus achieving robust and effective inter coding. Furthermore, existing low-latency-oriented methods often suffer from strong error propagation, so we extend the framework with a bidirectional prediction scheme and propose the bidirectional temporal prior. Extensive experimental results demonstrate that our method can obtain competitive performance compared to the state-of-the-art learned video compression approaches and the standard reference software HM-16.22.
Huairui Wang, Nianxiang Fu, Zhenzhong Chen 0001
ISCAS1
2023 Learned Image Compression with Enhanced Dynamic Spatial Aggregation and Asymmetric Entropy model
abstract
Learned image compression (LIC) has shown significant potential and better rate-distortion performance than traditional techniques. However, existing CNN-based approaches or window-based self-attention methods can only capture spatial information within fixed ranges. To tackle this limitation, we propose a novel method with dynamic spatial aggregation for transform coding. Our approach introduces enhanced adaptive aggregation, which generates kernel offsets to capture relevant information within content-dependent ranges, improving the transform process. Furthermore, we define a generalized coarse-to-fine entropy model that considers global context, channel-wise information, and spatial context in a coarse-to-fine manner. Additionally, our method takes into full consideration the model’s efficiency and complexity. By introducing the asymmetric entropy model structure and efficient heterogeneous convolution, our approach maintains lower coding complexity and higher decoding speed while ensuring performance. Experimental results demonstrate a substantial improvement in rate-distortion performance achieved by our method when compared to traditional compression methods like VTM and BPG, as well as some LIC methods.
Yuanton Zhang, Nianxiang Fu, Xiangdong Lv, Huairui Wang, Zhenzhong Chen 0001
VCIP5
2023 Optical Flow Reusing for High-Efficiency Space-Time Video Super Resolution
abstract
In this paper, we consider the task of space-time video super-resolution (ST-VSR), which can increase the spatial resolution and frame rate for a given video simultaneously. Despite the remarkable progress of recent methods, most of them still suffer from high computational costs and inefficient long-range information usage. To alleviate these problems, we propose a Bidirectional Recurrence Network (BRN) with the optical-flow-reuse strategy to better use temporal knowledge from long-range neighboring frames for high-efficiency reconstruction. Specifically, an efficient and memory-saving multi-frame motion utilization strategy is proposed by reusing the intermediate flow of adjacent frames, which considerably reduces the computation burden of frame alignment compared with traditional LSTM-based designs. In addition, the proposed hidden state in BRN is updated by the reused optical flow and refined by the Feature Refinement Module (FRM) for further optimization. Moreover, by utilizing intermediate flow estimation, the proposed method can inference non-linear motion and restore details better. Extensive experiments demonstrate that our optical-flow-reuse-based bidirectional recurrent network (OFR-BRN) is superior to state-of-the-art methods in accuracy and efficiency. Codes are available on URL:https://github.com/hahazh/OFR-BRN
Yuantong Zhang, Huairui Wang, Han Zhu 0003, Zhenzhong Chen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Controllable Space-Time Video Super-Resolution via Enhanced Bidirectional Flow Warping
abstract
Space-time video super-resolution targets to increase a given video's frame rate and resolution simultaneously. Al-though existing approaches have made great progress, most of them still suffer from the inaccurate approximation of large motions or fail to generate temporal consistent motion trajectory. To alleviate these problems, we carefully review the characteris-tics of different optical flow warping strategies, integrating and enhancing them to achieve more robust capabilities for handling extreme motions and time-modulated interpolation. Specifically, we utilize enhanced backward warping to perform alignment, mine space-time information across low resolution input frames, and propose an enhanced forward warping strategy to interpolate arbitrary intermediate frames. Furthermore, the proposed model can be trained end-to-end and produce intermediate results at any time by merely supervising the center moment. Experimental results show that the proposed algorithm performs favorably against the state-of-the-art methods in objective metrics and subjective visual effects.
Yuantong Zhang, Huairui Wang, Zhenzhong Chen 0001
VCIP2
2022 Multi-objective optimization based perceptual bit allocation for gaming video coding in VVC
Huairui Wang, Daiqin Yang
Signal Process.3
2020 DOVE: Decomposition Oriented Video super-rEsolution
abstract
Video super-resolution (VSR) has attracted a lot of attention that converts a low resolution (LR) video into a high resolution (HR) one. The original LR video is typically produced either by the downscaling processing or low-resolution sensor. Considering that the resolution degradation or limitation makes different impacts on different low-frequency (LF) and high-frequency (HF) components of the LR video signal, we propose a Decomposition Oriented Video super-rEsolution (DOVE) method in this paper. More specifically, a three-stream VSR network is designed in which the proposed LF and HF stream is responsible for modeling LF and HF components in the feature space. Moreover, a multi-frame refinement stream takes features of coarsely aligned frames as input and generates finely aligned counterparts progressively to guide the learning of LF and HF streams at the intermediate feature level. Furthermore, non-local channel attention is devised to capture long-range dependencies on a global scale both in the channel domain. Experimental results indicate that separating the learning of LF and HF components helps better estimate the HR frame from LR frames and superior VSR performance is achieved when compared with that of recent state-of-the-art methods.
Huairui Wang, Wanjie Sun, Daiqin Yang
VCIP1