Lei Zhao 0032

dblp:87/734-32 · DBLP profile ↗
← Back
11ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0003-4497-2197ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Template Matching Based Motion Refinement on Subblock Merge Mode
abstract
The subblock merge mode, in which the current coding block is split into multiple subblocks for motion compensation but still inherits the motion at the coding block level, improves the accuracy of the inter prediction and reduces the signaling overhead of motion information at the same time. And thus, it was adopted into versatile video coding (VVC) due to its high efficiency and continually improved in the enhanced compression model (ECM). However, the motion used in subblock merge mode was inherited from the previously coded blocks and may not match well with the current coding block. Thus, to improve the accuracy of the motion for the subblock merge mode, it is proposed to apply template matching (TM) based motion refinement. For subblock temporal motion vector predictor (SbTMVP) candidates, it is proposed to refine both the motion displacement and subblock motion vectors (MVs) based on TM; for affine motion candidates, it is proposed to refine the affine model, including base MV and non-translation parameters, based on TM. The proposed method was implemented on top of ECM, and the experiment results show that by applying the proposed method, it achieves {−0.23%(Y), −0.23%(U), −0.20%(V)} and {−0.09%(Y), −0.36%(U), −0.01%(V)} BD-rate reduction on random access (RA) and low delay B (LDB) configurations, respectively. Due to the good trade-off between performance and complexity, the proposed method was adopted into ECM.
Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003, Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0136
DCC5
2024 Subblock-Based Combined Inter and Intra Prediction Beyond VVC
abstract
Combined Inter and Intra Prediction (CIIP) is a coding tool that blends an inter prediction and an intra prediction to generate a hybrid prediction, which is rather efficient when neither intra nor inter prediction alone can well capture the content characteristic of a coding block. Due to the considerable coding performance, CIIP has been adopted in versatile video coding (VVC), and has been enhanced with further extensions in the recent enhanced compression model (ECM). However, CIIP is still less effective to handle the scenarios with complex motions since only translational motion model is leveraged. To tackle such issue, this paper investigates a method of subblock-based CIIP (subblock-CIIP), where the motion compensation signal of affine or subblock-based TMVP (SbTMVP) mode can be combined with an intra prediction to produce the CIIP prediction. Simulation results on ECM-11.0 indicate subblock-CIIP can achieve average 0.24% coding gain on sequences with abundant affine motions. The proposed subblock-CIIP method has been adopted in ECM-12.0.
Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0136
ICIP1
2024 Template Matching-Based Subblock Motion Refinement Towards Next Generation Video Coding
abstract
Subblock-based motion compensation (SMC) has shown re-markable effectiveness in handling non-uniform motions for inter video coding. However, existing strategies typically derive subblock motion rely solely on the inherited information from an already coded region, lacking content-adaptive optimization and potentially lead to deviations from the real motion. In this paper, we propose an overhead-free adaptive refinement scheme towards SMC, which leverages template matching (TM) to enhance the fidelity of subblock motion derivation for affine and subblock-based TMVP (SbTMVP) prediction. In particular, TM-based control point motion vector (CPMV) refinement is proposed to facilitate affine MC, where TM cost serves as a fidelity metric during the seeking of optimal translational affine motion offset. Additionally, the motion shift (MS) to locate SbTMVP is also refined by TM before specifying the subblock motion. The proposed method has been adopted into ECM-11.0. Experimental results on ECM-10.0 demonstrate overall 0.09% coding gains under random-access (RA) configuration, with negligible runtime increase.
Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0006
PCS1
2023 Enhanced Temporal Motion Derivation Beyond VVC
abstract
Temporal motion information has been widely exploited in advanced video coding standards to facilitate the efficiency of inter coding. In prior-arts, temporal cues are derived from fixed locations in a single collocated frame without considering trajectory consistency. This paper proposes to enhance temporal motion derivation in two aspects. On one hand, two collocated frames are utilized to provide more abundant temporal motion information from both forward and backward directions. On the other hand, motion shifts to locate regular temporal motion vector prediction (TMVP) or sub-block based TMVP (SbTMVP) are adaptively determined from multiple locations according to template matching costs. Simulation results indicate the proposed methods can achieve 0.11% and 0.15% coding gains respectively for RA and LDB configurations on ECM-7.0, with negligible complexity increase. The proposed method has been adopted into ECM-8.0.
Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0006
ICIP1
2022 Quality-Aware Merge Candidate Construction For Video Coding
abstract
Efficient construction of merge candidate list is a critical issue in terms of the advanced inter coding technologies. Existing merge candidate list construction methods are typically realized by traversing the potential candidates in a predefined order, without considering different qualities of the potential candidates. In this paper, we propose a quality-aware merge candidate list construction method based on template matching (TM) cost sorting, termed as QAMC. Instead of checking and appending potential candidates based on a predefined traversing order, QAMC is proposed to select and rank potential merge candidates according to the quality, where the quality is measured by the distortion between a reconstructed template region and its reference region. Simulation results on ECM-2.0 demonstrate that QAMC provides 0.27% and 0.32% coding gains respectively for RA and LDB configurations, with reasonable complexity increase.
Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0006
ISCAS1
2022 Enhanced Surveillance Video Compression With Dual Reference Frames Generation
abstract
In this paper, we improve the inter coding performance of surveillance videos by simultaneously investigating the distinct characteristics of background and foreground redundancy, and introduce two novel reference frames in a complementary manner. On one hand, a block level background reference frame (BRF) is proposed to reduce the background redundancy. The proposed scheme incorporates semantic information into the compression process, and makes use of instance segmentation to facilitate the background block decision, making the generated BRF free from foreground pollution. On the other hand, in order to handle foreground redundancy, a foreground reference frame (FRF) is generated based on Surveillance Prediction Generative Adversarial Network (SP-GAN), which utilizes previous reconstructed frames, optical flow based prediction, as well as BRF to infer the foreground objects of the to-be-coded frame. We integrate the proposed scheme into HM-16.6 software and append BRF and FRF to the reference pictures list (RPS). Simulation results demonstrate considerable superiority of the proposed scheme. In particular, by adding the proposed BRF to RPS, 3% coding gains are observed compared with the state-of-the-art BRF method. When both BRF and FRF are incorporated into RPS, 5.8% gains are achieved for surveillance video coding.
Lei Zhao 0032, Shiqi Wang 0001, Shanshe Wang, Yan Ye 0003, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Instance Segmentation Based Background Reference Frame Generation for Surveillance Video Coding
abstract
Efficient intelligent analysis and video compression are critical modules in terms of the advanced surveillance system. However, existing solutions always deal each task with independent strategies, leading to low-efficiency of the surveillance system. In this paper, we propose to handle these two tasks in a hybrid manner. In particular, a hybrid surveillance processing scheme towards efficient analysis and compression is presented, where the extracted semantic information can not only be utilized in intelligent analysis tasks, but also used to improve compression efficiency by facilitating the background reference frame (BRF) generation. Moreover, we propose to remove background redundancy of surveillance video by introducing the high quality BRF, where motion metric and semantic metric work in a complementary way to ensure the accurate detection of background blocks. Experimental results manifest considerable advantages of the proposed BRF. When the proposed BRF is integrated into reference picture set (RPS), 3% coding gains are obtained compared with state-of-the-art method.
Lei Zhao 0032, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Yan Ye 0003, Siwei Ma 0001, Wen Gao 0001
PCS1
2019 Enhanced Motion-Compensated Video Coding With Deep Virtual Reference Frame Generation
abstract
In this paper, we propose an efficient inter prediction scheme by introducing the deep virtual reference frame (VRF), which serves better reference in the temporal redundancy removal process of video coding. In particular, the high quality VRF is generated with the deep learning-based frame rate up conversion (FRUC) algorithm from two reconstructed bi-directional frames, which is subsequently incorporated into the reference list serving as the high quality reference. Moreover, to alleviate the compression artifacts of VRF, we develop a convolutional neural network (CNN)-based enhancement model to further improve its quality. To facilitate better utilization of the VRF, a CTU level coding mode termed as direct virtual reference frame (DVRF) is devised, which achieves better trade-off between compression performance and complexity. The proposed scheme is integrated into HM-16.6 and JEM-7.1 software platforms, and the simulation results under random access (RA) configuration demonstrate significant superiority of the proposed method. When adding VRF to RPS, more than 6% average BD-rate gain is achieved for HEVC test sequences on HM-16.6, and 0.8% BD-rate gain is observed based on JEM-7.1 software. Regarding the DVRF mode, 3.6% bitrate saving is achieved on HM-16.6 with the computational complexity effectively reduced.
Lei Zhao 0032, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Image Process.1
2018 Enhanced Ctu-Level Inter Prediction with Deep Frame Rate Up-Conversion for High Efficiency Video Coding
abstract
Inter prediction serves as the foundation of prediction based hybrid video coding framework. The state-of-the-art video coding standards employ the reconstructed frames as the references, and the motion vectors which convey the relative position shift between the current block and the prediction block are explicitly signalled in the bitstream. In this paper, we propose a high efficient inter prediction scheme by introducing a new methodology based on virtual reference frame, which is effectively generated with the deep neural network such that the motion data does not need to be explicitly signalled. In particular, the high quality virtual reference frame is generated with the deep learning based frame rate up-conversion (FRUC) algorithm from two reconstructed bi-prediction frames. Subsequently, a novel CTU level coding mode termed as direct virtual reference frame (DVRF) mode, is proposed to adaptively compensate for the current to-be-coded block in the sense of rate-distortion optimization (RDO). The proposed scheme is integrated into the HM-16.6 software, and experimental results demonstrate significant superiority of the proposed method, which provides more than 3% coding gains on average for HEVC test sequences.
Lei Zhao 0032, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICIP1
2017 Intelligent analysis oriented surveillance video coding
abstract
The fast growth of surveillance video big data presents great challenges to the video coding technology. Most existing video coding techniques target for visual quality optimization, while the ultimate utility of surveillance videos mainly lies in intelligent analyses, e.g., pedestrian detection and vehicle tracking. In view of this, we aim at proposing an efficient, standard-compatible and simultaneously analysis-friendly coding framework for intelligent surveillance videos. In particular, the foreground objects are first extracted from the video sequence by accurate background modeling. Subsequently, the foregrounds can be constructed as a sequence and compressed in higher quality while very few background pictures are required to signal in lower quality. At the decoder side, the foreground objects can be directly used for efficient analysis tasks and the surveillance videos can be also reconstructed by synthesizing background and foreground frames. The effectiveness and potential of the proposed framework have been demonstrated in the pedestrian detection application, where the coding bits can be greatly saved with the detection accuracy being well maintained.
Lei Zhao 0032, Xiang Zhang 0004, Xinfeng Zhang 0001, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICME1
2016 Mode-dependent pixel-wise motion refinement for HEVC
abstract
High efficiency video coding (HEVC) standard is within the block-based hybrid coding framework, which essentially adopts prediction unit (PU) as the basic motion compensation unit. However, in the case of tiny motion, the actual motion vectors (MVs) for each sample may differ from the PU's MV, thus resulting in more residual energy. In this paper, a novel pixel-wise motion refinement method (PMR) is presented by extending the traditional bidirectional optical flow (BIO) to PUs with two unidirectional reference blocks. To get the robust MV for each sample, a median filtering process is introduced to MV shifting values of the neighboring pixels. Furthermore, a PU level mode-dependent pixel-wise motion refinement (MPMR) scheme is also presented to improve the coding performance. Simulation results demonstrate that the proposed method achieves 0.83% bitrate reduction on average for HEVC test sequences under HM12.0 low delay B (LDB) configuration, and achieves up to 4.92% bitrate reduction for sequences with complex motion, e.g. affine motion.
Lei Zhao 0032, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001
ICIP1