Semih Esenlik

dblp:241/0256 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0003-5573-0326ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2026 An Overview of the JPEG AI Learning-Based Image Coding Standard
abstract
JPEG AI is an emerging learning-based image coding standard developed by Joint Photographic Experts Group (JPEG). The scope of the JPEG AI is the creation of a practical learning-based image coding standard offering a single-stream, compact compressed domain representation, targeting both human visualization and machine consumption. Scheduled for completion in early 2025, the first version of JPEG AI focuses on human vision tasks, demonstrating significant BD-rate reductions compared to existing standards, in terms of MS-SSIM, FSIM, VIF, VMAF, PSNR-HVS, IW-SSIM and NLPD quality metrics. Designed to ensure broad interoperability, JPEG AI incorporates various design features to support deployment across diverse devices and applications. This paper provides an overview of the technical features and characteristics of the JPEG AI standard.
Semih Esenlik, Yaojun Wu 0001, Zhaobin Zhang, Ye-Kui Wang, Kai Zhang 0007, Li Zhang 0006, João Ascenso, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Neural Video Compression with In-Loop Contextual Filtering and Out-of-Loop Reconstruction Enhancement
abstract
This paper explores the application of enhancement filtering techniques in neural video compression. Specifically, we categorize these techniques into in-loop contextual filtering and out-of-loop reconstruction enhancement based on whether the enhanced representation affects the subsequent coding loop. In-loop contextual filtering refines the temporal context by mitigating error propagation during frame-by-frame encoding. However, its influence on both the current and subsequent frames poses challenges in adaptively applying filtering throughout the sequence. To address this, we introduce an adaptive coding decision strategy that dynamically determines filtering application during encoding. Additionally, out-of-loop reconstruction enhancement is employed to refine the quality of reconstructed frames, providing a simple yet effective improvement in coding efficiency. To the best of our knowledge, this work presents the first systematic study of enhancement filtering in the context of conditional-based neural video compression. Extensive experiments demonstrate a 7.71% reduction in bit rate compared to state-of-the-art neural video codecs, validating the effectiveness of the proposed approach.
Yaojun Wu 0001, Chaoyi Lin, Yiming Wang 0008, Semih Esenlik, Zhaobin Zhang, Kai Zhang 0007, Li Zhang 0006
ACM Multimedia4
2024 Leveraging Conv-Attention for Efficient and High-Quality JPEG AI Image Coding
abstract
In this paper, we present a Conv-Attention, a decoder-friendly attention mechanism, in an effort to advancing the practical application of the artificial intelligence-based image coding. More specifically, the proposed method is tailored for JPEG AI, which is the latest advanced neural-network based image coding standard. By identifying the obstacles by profiling the decoding complexity of JPEG AI, the attention module accounts for a significant proportion, which mainly attributes to the intricate network structure and involvement of less efficient operations. Conv-Attention model is composed with plain convolution and activation computations, equipping with sub-scaling and up-scaling design, such that the non-adjacent features can be well captured, leading to the reduction of decoding complexity and maintenance of the synthesis and attentive capability. Simulation results verify the effectiveness of the proposed method with JPEG AI reference software, wherein the decoding complexity is reduced by 80% with negligible coding performance loss. The proposed method was adopted in the 100th JPEG meeting.
Meng Wang 0017, Semih Esenlik, Zhaobin Zhang, Yaojun Wu 0001, Kai Zhang 0007, Li Zhang 0006, Shiqi Wang 0001
DCC2
2024 Optimized Decoupled Structure with Non-Local Attention for Deep Image Compression
abstract
Recently, a decoupled framework for learning-based image compression has been proposed and adopted into the JPEG AI image coding standard developed by ISO/IEC WG1. The decoupled structure disentangles the sample reconstruction process and the entropy decoding process, making the decoding extremely fast. The corresponding techniques constitute the essential parts of the JPEG AI verification model software. However, its analysis transform and synthesis transform are relatively simple, which are built with stacked convolution layers, thereby may lack the capability to interpret data correlations. In this work, we enhance the transform networks by introducing the non-local attention mechanism, which has proven efficient in image compression tasks. The proposed framework thus shares the merits of the fast decoding from the decoupled architecture and the strong transform capabilities from the non-local attention, making it a stronger candidate for practical end-to-end image codec deployment. Experimental results on the Kodak test set and JPEG AI CfP test set show that our method achieves better BDRate performance compared to the original Decoupled-anchor and significantly faster decoding speed compared to NIC. The proposed solution has been adopted by the IEEE 1857.11 Working Subgroup (1857.11 WSG) in developing neural network-based image coding standards in the 10th Meeting.
Xuanye Zhang, Zhaobin Zhang, Yaojun Wu 0001, Semih Esenlik, Xiaoyan Sun 0001, Kai Zhang 0007, Li Zhang 0006
ICIP4
2024 Wavelet-like Transform with Subbands Fusion in Decoupled Structure for Deep Image Compression
abstract
Wavelet-like transform, based on convolutional neural network (CNN), is content-adaptive and has made remarkable achievements in end-to-end image compression. However, the subsequent sequential processing of each subband in the entropy module takes a relatively long decoding time, resulting in incon-venience for real-world applications. In this work, for lossy image compression, the wavelet-like transform is transplanted into the prevailing autoencoder structure to enhance the analysis and synthesis transform due to its excellent decomposition capability. The obtained subbands of different frequencies will undergo a hierarchical decorrelation architecture for subband fusion, also called cross fusing module. The specialized treatment will be applied to different subbands according to their spatial resolution to attain a more compact latent representation. In addition, the proposed solution features an architecture that decouples the arithmetic decoding process from the sample prediction process, which significantly reduces the decoding complexity. Experiments on the Kodak test set show that the proposed method achieves −3.04% BD-Rate compared to existing decoupled end-to-end structure in RGB Peak Signal-to-Noise Ratio (PSNR).
Yaojun Wu 0001, Zhaobin Zhang, Semih Esenlik, Xiaoyan Sun 0001, Kai Zhang 0007, Li Zhang 0006
PCS4
2024 End-to-End Learning-Based Image Compression With a Decoupled Framework
abstract
The autoregressive model has been widely used in learning-based image compression due to its superior context modeling capability. However, its sequential processing nature also undermines the ability of decoding in parallel and hinders the deployment in real applications. In this paper, we propose a decoupled framework to resolve this issue. With the decoupled architecture, the entropy decoding process is independent of the latent sample reconstruction process. The entropy decoding process thus can be finished before the latent sample prediction process begins, which leads to significant decoding time savings by enabling the two processes to be conducted in parallel. To further reduce the decoding time, we introduce wavefront processing, where multiple rows can be processed simultaneously when reconstructing the latent samples. On top of that, we design a series of coding tools to improve the rate-distortion efficiency and reduce the decoding complexity. Device interoperability is also supported by the proposed solution, where the same bitstream can be successfully decoded on different CPU/GPU devices. Comprehensive experiments are conducted to validate the effectiveness of the proposed method. Using objective evaluation metrics required by JPEG AI Call for Proposals (CfP), the proposed method achieves an average of -29.6% BD-rate changes with 2.44 times faster decoding speed compared to VVC image coding. When compared to the commonly used benchmark learning-based methods, the proposed method achieves -30.5% BD-rate changes and 101 times faster decoding speed over cheng2020attn. The proposed solution has been proposed to JPEG AI and IEEE 1857.11 as a response to CfP and the core techniques have been adopted to build the verification model. The software and the instructions can be accessed at https://github.com/bytedance/BEE.
Zhaobin Zhang, Semih Esenlik, Yaojun Wu 0001, Meng Wang 0017, Kai Zhang 0007, Li Zhang 0006
IEEE Trans. Circuits Syst. Video Technol.2
2021 Decoder-Side Motion Vector Refinement in VVC: Algorithm and Hardware Implementation Considerations
abstract
This paper presents an overview of the decoder-side motion vector refinement (DMVR) algorithm in the Versatile Video Coding (VVC) standard. The proposed DMVR algorithm aims to increase the prediction accuracy of the blocks coded in merge mode using the bilateral matching-based refinement method. Compared with previous decoder-side motion vector derivation approaches, the proposed method significantly increases the coding efficiency without signaling additional side information. Furthermore, the hardware implementation considerations of the DMVR design are particularly focused in this study. This paper details and analyzes the novel features of DMVR contributing to the increase in coding efficiency and the reduction in computational complexity and implementation difficulty. Experimental results based on the VVC test model version 8.0 demonstrate that average Bjøntegaard Delta rate savings of 0.80 % and 2.81 % are achieved for the “tool-off” and “tool-on” test configurations, respectively. Moreover, 4 % additional decoding time and negligible additional external memory bandwidth requirements of DMVR based on the common test conditions for VVC are reported.
Han Gao 0001, Semih Esenlik, Jianle Chen, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.3
2021 Geometric Partitioning Mode in Versatile Video Coding: Algorithm Review and Analysis
abstract
This paper presents an overview of the geometric partitioning mode (GPM) algorithm that is a part of the most recent Versatile Video Coding (VVC) standard. The GPM algorithm aims to increase the partitioning precision of moving objects using non-rectangular and asymmetric rectangular partitions on top of the conventional rectangular block partitioning structure of VVC. Novel features of GPM contributing to the increase in coding efficiency and the reduction in encoder and decoder complexity are detailed and analyzed in this paper. Evaluated with VVC test model version 8.0 under the joint video experts team common test conditions, experimental results show that the presented GPM algorithm provides luma Bjøntegaard Delta rate reduction of 0.70% for random access and of 1.55% for low delay with B slices configurations, with roughly 3% to 5% additional encoding time and negligible decoder runtime change. Furthermore, as GPM provides more precise partitions for the boundaries of the moving objects, an improvement of visual quality is seen in GPM coded sequences.
Han Gao 0001, Semih Esenlik, Elena Alshina, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.2
2021 Subblock-Based Motion Derivation and Inter Prediction Refinement in the Versatile Video Coding Standard
abstract
Efficient representation and coding of fine-granular motion information is one of the key research areas for exploiting inter-frame correlation in video coding. Representative techniques towards this direction are affine motion compensation (AMC), decoder-side motion vector refinement (DMVR), and subblock-based temporal motion vector prediction (SbTMVP). Fine-granular motion information is derived at subblock level for all the three coding tools. In addition, the obtained inter prediction can be further refined by two optical flow-based coding tools, the bi-directional optical flow (BDOF) for bi-directional inter prediction and the prediction refinement with optical flow (PROF) exclusively used in combination with AMC. The aforementioned five coding tools have been extensively studied and finally adopted in the Versatile Video Coding (VVC) standard. This paper presents technical details of each tool and highlights the design elements with the consideration of typical hardware implementations. Following the common test conditions defined by Joint Video Experts Team (JVET) for the development of VVC, 5.7% bitrate reduction on average is achieved by the five tools. For test sequences characterized by large and complex motion, up to 13.4% bitrate reduction is observed. Additionally, visual quality improvement is demonstrated and analyzed.
Haitao Yang 0001, Huanbang Chen, Jianle Chen, Semih Esenlik, Sriram Sethuraman, Xiaoyu Xiu, Elena Alshina, Jiancong Luo
IEEE Trans. Circuits Syst. Video Technol.4
2020 Advanced Geometric-Based Inter Prediction for Versatile Video Coding
abstract
Block-based partitioning is one of the fundamental techniques in video coding. Geometric-based block partitioning is a well-studied method to enable better spatial adaptation to the signal properties. This paper introduces the most recent proposal of advanced geometric-based inter prediction (GIP) made to the state-of-the-art are video coding standard - Versatile Video Coding (VVC). Implemented in the latest test model VTM-6.0 to generalize the existing triangle partition mode (TPM) and evaluated with the Joint Video Experts Team (JVET) Common Test Conditions (CTC) sequences, the proposed advanced GIP scheme provides luma BD-rate reduction of 0.56% for random access (RA) and 1.37% for low-delay (LB) test cases with 2% encoder runtime increase and negligible decoder runtime increase. Furthermore, BD-rate reductions up to 2.92% and 3.49% for RA and LB test cases can be achieved in the absence of multiple related VVC inter prediction tools.
Han Gao 0001, Ru-Ling Liao, Kevin Reuze, Semih Esenlik, Elena Alshina, Yan Ye 0003, Jie Chen 0006, Jiancong Luo, Chun-Chi Chen, Han Huang 0001, Wei-Jung Chien, Vadim Seregin, Marta Karczewicz
DCC4
2020 Video Codec Using Flexible Block Partitioning and Advanced Prediction, Transform and Loop Filtering Technologies
abstract
This paper describes a joint response to the Call for Proposals by Samsung, Huawei, GoPro, and HiSilicon on Video Compression with Capability beyond HEVC/H.265, jointly issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG). In the proposed codec, the coding framework supports hierarchical splitting with binary and ternary trees and flexible coding order representations. Additionally, novel compression tools on inter/intra prediction, in-loop filtering, and entropy coding have been proposed. The proposed compression scheme provides significantly higher compression capability than the state-of-the-art HEVC/H.265 standard for SDR (Standard Dynamic Range) category while maintaining complexity acceptable for emerging applications. When all the proposed algorithmic tools are used, the proposed video codec achieves approximately 40% bit-saving for the SDR cetegory on average compared to HEVC/H.265 anchor.
Kiho Choi, Jianle Chen, Haitao Yang 0001, Woongil Choi, Sergey Ikonin, Yinji Piao, Semih Esenlik, Minsoo Park, Ye-Kui Wang, Narae Choi, Yin Zhao, Seungsoo Jeong, Anish Tamse, Alexey Filippov, Heechul Yang, Junghye Min, Roman Chernyak, Bora Jin, Anand Meher Kotra, Sunil Lee, Han Gao 0001, Chanyul Kim, Timofey Solovyev, Kwangpyo Choi, Vasily Rufitskiy, Maxim Sychev, Jeonghoon Park
IEEE Trans. Circuits Syst. Video Technol.8
2019 New Video Codec for High-Quality Video Service and Emerging Applications
abstract
This paper proposes a novel video compression scheme for high-quality video service and emerging applications such as 360-degree omnidirectional and high dynamic range video coding. The coding framework supports hierarchical splitting of blocks with binary and ternary-split trees and flexible coding order representations. Moreover, minimal tool set to obtain high precision prediction and compression enhancement has been proposed. Compared to HEVC, bit-rate reduction of around 40% based on objective measures has been shown. This was one of the responses to the Call for Proposals (CfP) for VVC standardization.
Kiho Choi, Jianle Chen, Anish Tamse, Haitao Yang 0001, Sergey Ikonin, Woongil Choi, Semih Esenlik
DCC8
2019 Decoder Side Motion Vector Refinement for Versatile Video Coding
abstract
Inter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the bitrate required for motion vector (MV) signaling, the High Efficiency Video Coding (HEVC) standard and the latest Versatile Video Coding (VVC) draft utilize a merge mode to signal the MV. While the merge mode saves the bits for MV indication, it generates inaccurate MVs, which lead to imprecise prediction. To improve the coding performance, novel decoder side motion vector refinement (DMVR) schemes are currently being proposed. The DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. We present two variants of block matching-based DMVR, namely template matching and bilateral matching. Experimental results obtained with the VTM-2.0 reference software, after integrating our approaches, demonstrate that our proposed methods provide an average luma BD-rate reduction of 4.71% for the template matching-based DMVR and 4.92% for the bilateral matching-based DMVR when using the random access configuration.
Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
MMSP2
2019 Low-Complexity Geometric Inter-Prediction for Versatile Video Coding
abstract
Non-rectangular block partitioning is a well-known method for improved inter-picture prediction in video coding, enabling better spatial adaptation to the signal properties. This contribution presents the most recent proposal of geometric inter-prediction (GIP) made to the Versatile Video Coding (VVC) standardization activity led by the Joint Video Experts Team (JVET). Implemented in the latest test model VTM-5.0 and evaluated according to the JVET Common Test Conditions, the proposed low-complexity GIP scheme provides objective luma BD-rate reductions of 0.22 % for random access and 0.44 % for low-delay test cases at 7% encoder runtime increase and negligible decoder runtime increase. The coding gain is provided by non-triangular partitioned blocks and in the presence of multiple other VVC coding tools. Furthermore, BD-rate reductions of 2.58 % and 2.78 % can be achieved specifically for pure screen content by employing an adaptive blending filter.
Max Bläser, Han Gao 0001, Semih Esenlik, Elena Alshina, Zhijie Zhao, Christian Rohlfing, Eckehard G. Steinbach
PCS3
2019 Low Complexity Decoder Side Motion Vector Refinement for VVC
abstract
Inter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the motion vector signaling overhead, a merge mode with subsequent decoder side motion vector refinement (DMVR) is current under investigation for the first working draft of the standardization activity on Versatile Video Coding (VVC). While the DMVR method searches the refined MVs at the decoder side, it heavily increases the decoding complexity and the memory bandwidth requirements. To address these issues, a novel low complexity DMVR scheme is proposed in this paper. The low complexity DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. The proposed low complexity improvements are added to a previously proposed bilateral matching-based DMVR approach. Experimental results obtained with the VTM 2.0 reference software, after integrating our approaches, show that the previously proposed DMVR provides an average luma BD-rate reduction of 4.59% with 32% additional decoding time and the proposed low complexity DMVR provides an average luma BD-rate reduction of 1.67% with only 6% additional decoding time when using the random access configuration.
Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
PCS2