VLDB 2026 Research / reviewers in the wild / expert
Han Gao 0001
dblp:56/1065-1
· DBLP profile ↗
13ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-6547-1557ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video Coding With Cross-Component Sample OffsetabstractBeyond the exploration of traditional spatial, temporal and subjective visual signal redundancy in image and video compression, recent research has focused on leveraging cross-color component redundancy to enhance coding efficiency. Cross-component coding approaches are motivated by the statistical correlations among different color components, such as those in the Y'CbCr color space, where luma (Y) color component typically exhibits finer details than chroma (Cb/Cr) color components. Inspired by previous cross-component coding algorithms, this paper introduces a novel in-loop filtering approach named Cross-Component Sample Offset (CCSO). CCSO utilizes co-located and neighboring luma samples to generate correction signals for both luma and chroma reconstructed samples. It is a multiplication-free, non-linear mapping process implemented using a look-up-table. The input to the mapping is a group of reconstructed luma samples, and the output is an offset value applied on the center luma or co-located chroma sample. Experimental results demonstrate that the proposed CCSO can be applied to both image and video coding, resulting in improved coding efficiency and visual quality. The method has been adopted into an experimental next-generation video codec beyond AV1 developed by the Alliance for Open Media (AOMedia), demonstrating average -0.81% and -0.69% coding gain on PSNR and VMAF quality metric, respectively, under random access configuration. Additionally, CCSO notably improves the subjective visual quality. Han Gao 0001, Xin Zhao 0003, Shan Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Content Adaptive Weighted Prediction for Video Coding Beyond AV1abstractIn AV1, the compound prediction supports three methods for averaging two prediction blocks. The weighting factors for two prediction blocks can be based on either the wedge mask, the difference between prediction samples, or predefined weighting factors. However, for single prediction, only one predefined weighting factor is used, which is suboptimal when there are illumination changes between reference frames and the current frame. To address this limitation, this paper proposes a content adaptive weighted prediction method. This approach aims to enhance the prediction accuracy for single prediction. It involves the use of multiple predefined weighting factor look-up tables, and the selection among these different look-up tables are implicitly determined based on coded information of the current block, and the index of scaling factor in the look-up table is signaled in bitstream and parsed at the decoder side. Experimental results show that the proposed method can achieve an average 0.31% and 0.24% coding gain in terms of BD-rate with random access and low delay configurations, respectively. Xin Zhao 0003, Han Gao 0001, Shan Liu 0001 |
VCIP | 4 |
| 2022 | Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
Ahmet Burakhan Koyuncu, Han Gao 0001, Atanas Boev, Georgii Gaikov, Elena Alshina, Eckehard G. Steinbach |
ECCV (19) | 2 |
| 2021 | Motion Vector Coding and Block Merging in the Versatile Video Coding StandardabstractThis paper overviews the motion vector coding and block merging techniques in the Versatile Video Coding (VVC) standard developed by the Joint Video Experts Team (JVET). In general, inter-prediction techniques in VVC can be classified into two major groups: “whole block-based inter prediction” and “subblock-based inter prediction”. In this paper, we focus on techniques for whole block-based inter prediction. As in its predecessor, High Efficiency Video Coding (HEVC), whole block-based inter prediction in VVC is represented by adaptive motion vector prediction (AMVP) mode or merge mode. Newly introduced features purely for AMVP mode include symmetric motion vector difference and adaptive motion vector resolution. The features purely for merge mode include pairwise average merge, merge with motion vector difference, combined inter-intra prediction and geometric partitioning mode. Coding tools such as history-based motion vector prediction and bidirectional prediction with coding unit weights can be applied on both AMVP mode and merge mode. This paper discusses the design principles and the implementation of the new inter-prediction methods. Using objective metrics, simulation results show that the methods overviewed in the paper can jointly achieve 6.2% and 4.7% BD-rate savings on average with the random access and low-delay configurations, respectively. Significant subjective picture quality improvements of some tools are also reported when comparing the resulting pictures at same bitrates. Wei-Jung Chien, Li Zhang 0006, Martin Winken, Xiang Li 0003, Ru-Ling Liao, Han Gao 0001, Hongbin Liu 0004, Chun-Chi Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Decoder-Side Motion Vector Refinement in VVC: Algorithm and Hardware Implementation ConsiderationsabstractThis paper presents an overview of the decoder-side motion vector refinement (DMVR) algorithm in the Versatile Video Coding (VVC) standard. The proposed DMVR algorithm aims to increase the prediction accuracy of the blocks coded in merge mode using the bilateral matching-based refinement method. Compared with previous decoder-side motion vector derivation approaches, the proposed method significantly increases the coding efficiency without signaling additional side information. Furthermore, the hardware implementation considerations of the DMVR design are particularly focused in this study. This paper details and analyzes the novel features of DMVR contributing to the increase in coding efficiency and the reduction in computational complexity and implementation difficulty. Experimental results based on the VVC test model version 8.0 demonstrate that average Bjøntegaard Delta rate savings of 0.80 % and 2.81 % are achieved for the “tool-off” and “tool-on” test configurations, respectively. Moreover, 4 % additional decoding time and negligible additional external memory bandwidth requirements of DMVR based on the common test conditions for VVC are reported. Han Gao 0001, Semih Esenlik, Jianle Chen, Eckehard G. Steinbach |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Geometric Partitioning Mode in Versatile Video Coding: Algorithm Review and AnalysisabstractThis paper presents an overview of the geometric partitioning mode (GPM) algorithm that is a part of the most recent Versatile Video Coding (VVC) standard. The GPM algorithm aims to increase the partitioning precision of moving objects using non-rectangular and asymmetric rectangular partitions on top of the conventional rectangular block partitioning structure of VVC. Novel features of GPM contributing to the increase in coding efficiency and the reduction in encoder and decoder complexity are detailed and analyzed in this paper. Evaluated with VVC test model version 8.0 under the joint video experts team common test conditions, experimental results show that the presented GPM algorithm provides luma Bjøntegaard Delta rate reduction of 0.70% for random access and of 1.55% for low delay with B slices configurations, with roughly 3% to 5% additional encoding time and negligible decoder runtime change. Furthermore, as GPM provides more precise partitions for the boundaries of the moving objects, an improvement of visual quality is seen in GPM coded sequences. Han Gao 0001, Semih Esenlik, Elena Alshina, Eckehard G. Steinbach |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Block Partitioning Structure in the VVC StandardabstractVersatile Video Coding (VVC) is the latest video coding standard jointly developed by ITU-T VCEG and ISO/IEC MPEG. In this paper, technical details and experimental results for the VVC block partitioning structure are provided. Among all the new technical aspects of VVC, the block partitioning structure is identified as one of the most substantial changes relative to the previous video coding standards and provides the most significant coding gains. The new partitioning structure is designed using a more flexible scheme. Each coding tree unit (CTU) is either treated as one coding unit or split into multiple coding units by one or more recursive quaternary tree partitions followed by one or more recursive multi-type tree splits. The latter can be horizontal binary tree split, vertical binary tree split, horizontal ternary tree split, or vertical ternary tree split. A CTU dual tree for intra-coded slices is described on top of the new block partitioning structure, allowing separate coding trees for luma and chroma. Also, a new way of handling picture boundaries is presented. Additionally, to reduce hardware decoder complexity, virtual pipeline data unit constraints are introduced, which forbid certain multi-type tree splits. Finally, a local dual tree is described, which reduces the number of small chroma intra blocks. Yu-Wen Huang, Jicheng An, Han Huang 0001, Xiang Li 0003, Shih-Ta Hsiang, Kai Zhang 0007, Han Gao 0001, Jackie Ma, Olena Chubach |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | Advanced Geometric-Based Inter Prediction for Versatile Video CodingabstractBlock-based partitioning is one of the fundamental techniques in video coding. Geometric-based block partitioning is a well-studied method to enable better spatial adaptation to the signal properties. This paper introduces the most recent proposal of advanced geometric-based inter prediction (GIP) made to the state-of-the-art are video coding standard - Versatile Video Coding (VVC). Implemented in the latest test model VTM-6.0 to generalize the existing triangle partition mode (TPM) and evaluated with the Joint Video Experts Team (JVET) Common Test Conditions (CTC) sequences, the proposed advanced GIP scheme provides luma BD-rate reduction of 0.56% for random access (RA) and 1.37% for low-delay (LB) test cases with 2% encoder runtime increase and negligible decoder runtime increase. Furthermore, BD-rate reductions up to 2.92% and 3.49% for RA and LB test cases can be achieved in the absence of multiple related VVC inter prediction tools. Han Gao 0001, Ru-Ling Liao, Kevin Reuze, Semih Esenlik, Elena Alshina, Yan Ye 0003, Jie Chen 0006, Jiancong Luo, Chun-Chi Chen, Han Huang 0001, Wei-Jung Chien, Vadim Seregin, Marta Karczewicz |
DCC | 1 |
| 2020 | Video Codec Using Flexible Block Partitioning and Advanced Prediction, Transform and Loop Filtering TechnologiesabstractThis paper describes a joint response to the Call for Proposals by Samsung, Huawei, GoPro, and HiSilicon on Video Compression with Capability beyond HEVC/H.265, jointly issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG). In the proposed codec, the coding framework supports hierarchical splitting with binary and ternary trees and flexible coding order representations. Additionally, novel compression tools on inter/intra prediction, in-loop filtering, and entropy coding have been proposed. The proposed compression scheme provides significantly higher compression capability than the state-of-the-art HEVC/H.265 standard for SDR (Standard Dynamic Range) category while maintaining complexity acceptable for emerging applications. When all the proposed algorithmic tools are used, the proposed video codec achieves approximately 40% bit-saving for the SDR cetegory on average compared to HEVC/H.265 anchor. Kiho Choi, Jianle Chen, Haitao Yang 0001, Woongil Choi, Sergey Ikonin, Yinji Piao, Semih Esenlik, Minsoo Park, Ye-Kui Wang, Narae Choi, Yin Zhao, Seungsoo Jeong, Anish Tamse, Alexey Filippov, Heechul Yang, Junghye Min, Roman Chernyak, Bora Jin, Anand Meher Kotra, Sunil Lee, Han Gao 0001, Chanyul Kim, Timofey Solovyev, Kwangpyo Choi, Vasily Rufitskiy, Maxim Sychev, Jeonghoon Park |
IEEE Trans. Circuits Syst. Video Technol. | 24 |
| 2019 | Decoder Side Motion Vector Refinement for Versatile Video CodingabstractInter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the bitrate required for motion vector (MV) signaling, the High Efficiency Video Coding (HEVC) standard and the latest Versatile Video Coding (VVC) draft utilize a merge mode to signal the MV. While the merge mode saves the bits for MV indication, it generates inaccurate MVs, which lead to imprecise prediction. To improve the coding performance, novel decoder side motion vector refinement (DMVR) schemes are currently being proposed. The DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. We present two variants of block matching-based DMVR, namely template matching and bilateral matching. Experimental results obtained with the VTM-2.0 reference software, after integrating our approaches, demonstrate that our proposed methods provide an average luma BD-rate reduction of 4.71% for the template matching-based DMVR and 4.92% for the bilateral matching-based DMVR when using the random access configuration. Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen |
MMSP | 1 |
| 2019 | Low-Complexity Geometric Inter-Prediction for Versatile Video CodingabstractNon-rectangular block partitioning is a well-known method for improved inter-picture prediction in video coding, enabling better spatial adaptation to the signal properties. This contribution presents the most recent proposal of geometric inter-prediction (GIP) made to the Versatile Video Coding (VVC) standardization activity led by the Joint Video Experts Team (JVET). Implemented in the latest test model VTM-5.0 and evaluated according to the JVET Common Test Conditions, the proposed low-complexity GIP scheme provides objective luma BD-rate reductions of 0.22 % for random access and 0.44 % for low-delay test cases at 7% encoder runtime increase and negligible decoder runtime increase. The coding gain is provided by non-triangular partitioned blocks and in the presence of multiple other VVC coding tools. Furthermore, BD-rate reductions of 2.58 % and 2.78 % can be achieved specifically for pure screen content by employing an adaptive blending filter. Max Bläser, Han Gao 0001, Semih Esenlik, Elena Alshina, Zhijie Zhao, Christian Rohlfing, Eckehard G. Steinbach |
PCS | 2 |
| 2019 | Low Complexity Decoder Side Motion Vector Refinement for VVCabstractInter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the motion vector signaling overhead, a merge mode with subsequent decoder side motion vector refinement (DMVR) is current under investigation for the first working draft of the standardization activity on Versatile Video Coding (VVC). While the DMVR method searches the refined MVs at the decoder side, it heavily increases the decoding complexity and the memory bandwidth requirements. To address these issues, a novel low complexity DMVR scheme is proposed in this paper. The low complexity DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. The proposed low complexity improvements are added to a previously proposed bilateral matching-based DMVR approach. Experimental results obtained with the VTM 2.0 reference software, after integrating our approaches, show that the previously proposed DMVR provides an average luma BD-rate reduction of 4.59% with 32% additional decoding time and the proposed low complexity DMVR provides an average luma BD-rate reduction of 1.67% with only 6% additional decoding time when using the random access configuration. Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen |
PCS | 1 |
| 2018 | Improving Picture Boundary Handling for Video Coding Beyond HEVCabstractBlock partitioning is an essential component of today's hybrid video codecs. In order to support video sequences with arbitrary spatial resolution, the HEVC standard and the JEM reference software utilize a forced quadtree (QT) method to process CTUs/CUs located at the picture boundaries. While the forced QT method is simple to implement, it is unsatisfactory from a coding efficiency perspective. To improve the compression performance, novel block partitioning methods are currently being proposed. Our approach for picture boundary handling uses a binary tree (BT) based splitting scheme to deal with the boundary located CTUs/CUs. We present two variants with the first one combining forced BT with forced QT partitioning (forced QTBT) and the second one adaptively choosing between BT and QT (adaptive QTBT). Experimental results obtained with the JEM-6.0 reference software after integrating our approaches demonstrate that our proposed methods provide an average luma BD-rate reduction of 0.85% for the forced QTBT scheme and 0.96% for the adaptive QTBT scheme when using the random access configuration. Han Gao 0001, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen |
VCIP | 1 |