Hewei Liu

dblp:54/10296 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep Network-Based Adaptive Quantization for Practical Video Coding
abstract
The optimization of block-level quantization parameters (QP) is critical to improving the performance of practical block-based video compression encoders, but the extremely large optimization space makes it challenging to solve. Existing solutions, e.g. HEVC encoder x265, usually add some optimization constraints of the block-independent assumption and linear distortion propagation model, which limits compression efficiency improvement to a certain extent. To address this problem, a deep learning-based encoder-only adaptive quantization method (DAQ) is proposed in this paper, where a deep network is designed to adaptively model the joint temporal propagation relationship of quantization among blocks. Specifically, DAQ consists of two phases: in the training phase, considering the heavy searching cost of the traditional codec, we introduce a well-designed end-to-end learned block-based video compression network as an effective training proxy tool for the deep encoder-side network. While in the deployment phase, the trained deep network is applied to jointly predict all block QPs in a frame for the traditional encoder. Besides, our network deploys only on the encoder side without changing the standard decoder and has very low inference complexity, making it able to apply in practice. At last, we deploy DAQ in HEVC and VVC encoder for performance comparison, and the experimental results demonstrate that DAQ significantly outperforms practically used x265 with on average 15.0%, 10.9% BD-rate reduction under the SSIM and PSNR, and also achieves 12.5%, 5.0% coding gain than VTM. Moreover, for deploying deep video codec in practice, this work provides a new insight for optimizing the encoder parameters with a large space.
Hewei Liu, Jiawen Gu, Dengchao Jin, Meng Lei, Chao Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.2
2025 Deep Adaptive Quantization for Practical Video Compression
abstract
In this work, we propose a deep learning-based adaptive quantization method to promote video coding performance. Due to inter-prediction and reference mechanism, the block-level quantization parameter (QP) not only influences current block distortion but also has complex temporal propagation effects on subsequent coding frames. Our idea is to utilize a deep network to model the complex temporal propagation relationship of quantization. As shown in Fig. 1, the deep network directly predicts all block-level QPs of the frame for the traditional encoder without changing the standard decoder. Since our network deploys only on the encoder side and has low inference complexity, it can be easily applied in practice. In addition, we use a learned coding network as a proxy of the traditional codec to train our network.
Hewei Liu, Jiawen Gu, Dengchao Jin, Meng Lei, Chao Zhou 0003
DCC2
2023 Efficient online real-time video stabilization with a novel least squares formulation and parallel AC-RANSAC
Jianwei Ke, Alex J. Watras, Jae-Jun Kim, Hewei Liu, Hongrui Jiang, Yu Hen Hu
J. Vis. Commun. Image Represent.4
2023 DFCE: Decoder-Friendly Chrominance Enhancement for HEVC Intra Coding
abstract
We propose a decoder-friendly chrominance enhancement method for the compressed images. Our proposed method is developed based on the luminance-guided chrominance enhancement network (LGCEN) and online learning. With LGCEN, the textures of the compressed chrominance components are enhanced by the guidance of luminance component. Moreover, LGCEN is constructed with the recursive design and the light-weight channel attention mechanism to achieve high performance as well as low complexity. It is arranged at both encoder and decoder sides. Given the input image, we train LGCEN at encoder side by using online learning. With online learning, we partially update network parameters and transmit them to decoder to update LGCEN arranged there. The adoption of online learning effectively reduces the workload of decoder and guarantee high robustness. Compared with the state-of-the-art methods, our proposed approach achieves superior performance.
Renwei Yang, Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Luminance-Guided Chrominance Image Enhancement for HEVC Intra Coding
abstract
In this paper, we propose a luminance-guided chrominance image enhancement convolutional neural network for HEVC intra coding. Specifically, we firstly develop a gated recursive asymmetric-convolution block to restore each degraded chrominance image, which generates an intermediate output. Then, guided by the luminance image, the quality of this intermediate output is further improved, which finally produces the high-quality chrominance image. When our proposed method is adopted in the compression of color images with HEVC intra coding, it achieves 28.96% and 16.74% BD-rate gains over HEVC for the U and V images, respectively, which accordingly demonstrate its superiority. The code is available online: https://github.com/Nickyang4900/Luminance-Guided-Chrominance-Enhancement-for-HEVC-Intra-Coding.
Hewei Liu, Renwei Yang, Shuyuan Zhu, Bing Zeng 0001
ISCAS1
2022 Deep Feature Compression with Collaborative Coding of Image Texture
abstract
In this paper, we propose a coding scheme for the deep intermediate feature and it is implemented with the collaborative compression of image texture. More specifically, we separately compress the feature and texture of the image to form two data layers. The first one is the intermediate feature layer and the second one is the texture layer. The texture layer can provide an image for users and the feature layer can be used to implement the computer vision (CV) task. With our proposed deep reconstruction network (RecNet), the texture and features cooperate to achieve a high-quality visual output as well as a high-efficiency CV task. The experimental results demonstrate the excellent performance by using our proposed method to compress the deep features.
Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Ruiqin Xiong, Bing Zeng 0001
ISCAS2
2022 Inter-Frame Dependency-Based Rate Control for VVC Low-Delay Coding
abstract
In this letter, we propose two solutions for the rate control of the VVC low-delay coding. Both solutions are developed by determining the bit allocation factors for video frames based on their dependency. Specifically, we design the first solution according to the distortion correlation between the key-frame and its subsequent frames. With this solution, the bit allocation factors are determined by applying multi-pass coding on the video to build up the cross-frame distortion model. This model offers us a straightforward way to achieve better rate control performance but results in a rather high complexity. To solve the complexity problem, we propose the second solution based on the difference between frames. In this solution, we construct the bit allocation model and apply it to frames so that we can adaptively determine the allocation factors with a low complexity. The experimental results demonstrate that our proposed two solutions can offer better rate-distortion performances than the state-of-the-art method.
Hewei Liu, Shuyuan Zhu, Bing Zeng 0001
IEEE Signal Process. Lett.1
2021 Efficient Real-Time Video Stabilization with a Novel Least Squares Formulation
abstract
We present a novel video stabilization algorithm (LSstab) that removes unwanted motions in real-time. LSstab is based on a novel least squares formulation of the smoothing cost function to alleviate the undesirable camera jitter. A recursive least square solver is derived to minimize the smoothing cost function with an O(N) computation complexity. LSstab is evaluated using a suite of publicly available videos against the state of the art video stabilization methods. Results show LSstab reaches comparable or better performance, achieving real-time processing speed when a GPU is used.
Jianwei Ke, Alex J. Watras, Jae-Jun Kim, Hewei Liu, Hongrui Jiang, Yu Hen Hu
ICASSP4
2021 Cross-Block Difference Guided Fast CU Partition for VVC Intra Coding
abstract
In this paper, we propose a new fast CU partition method for VVC intra coding based on the cross-block difference. This difference is measured by the gradient and the content of sub-blocks obtained from partition and is employed to guide the skipping of unnecessary horizontal and vertical partition modes. With this guidance, a fast determination of block partitions is accordingly achieved. Compared with VVC, our proposed method can save 41.64% (on average) encoding time with only 0.97% (on average) increase of BD-rate.
Hewei Liu, Shuyuan Zhu, Ruiqin Xiong, Guanghui Liu 0001, Bing Zeng 0001
VCIP1
2020 Towards Real-Time, Multi-View Video Stereopsis
abstract
We present a real-time, multi-view video stereopsis (RTMVS) algorithm. This algorithm processes five synchronized video streams from cameras of a stationary camera array using a commodity laptop computer equipped with an Nvidia GPU. It provides 3D visualization of a dynamic scene from a chosen viewpoint at the video frame rate. In RTMVS, 3D surfaces are represented as a set of triangles anchored on a sparse set of 3D feature points. The computationally intensive Structure-from-Motion (SfM) algorithm is executed as an initial step. Feature points in each video stream are tracked using a KLT tracker. Triangles will be updated only when at least one vertex moves from its current position. The algorithm will redetect features every X frame. Epipolar geometry and Trifocal tensor are also exploited to accelerate sparse feature point matching and track filtering. Compared to a dense point cloud multi-view stereopsis baseline algorithm, RTMVS reduces the processing time per frame from 30s down to less than 44 ms.
Jianwei Ke, Alex J. Watras, Jae-Jun Kim, Hewei Liu, Hongrui Jiang, Yu Hen Hu
ICASSP4