EDBT 2026 Demo / reviewers in the wild / expert
Zhaobin Zhang
dblp:205/7225
· DBLP profile ↗
19ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Overview of the JPEG AI Learning-Based Image Coding StandardabstractJPEG AI is an emerging learning-based image coding standard developed by Joint Photographic Experts Group (JPEG). The scope of the JPEG AI is the creation of a practical learning-based image coding standard offering a single-stream, compact compressed domain representation, targeting both human visualization and machine consumption. Scheduled for completion in early 2025, the first version of JPEG AI focuses on human vision tasks, demonstrating significant BD-rate reductions compared to existing standards, in terms of MS-SSIM, FSIM, VIF, VMAF, PSNR-HVS, IW-SSIM and NLPD quality metrics. Designed to ensure broad interoperability, JPEG AI incorporates various design features to support deployment across diverse devices and applications. This paper provides an overview of the technical features and characteristics of the JPEG AI standard. Semih Esenlik, Yaojun Wu 0001, Zhaobin Zhang, Ye-Kui Wang, Kai Zhang 0007, Li Zhang 0006, João Ascenso, Shan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Dual-Scale Transformer with Variable Bitrate Synchronization for Neural Video CompressionabstractNeural video compression (NVC) has emerged as a promising paradigm for improving rate-distortion performance. However, existing neural video codecs predominantly rely on convolutional neural networks (CNNs) with limited local receptive fields to generate the latent representations, often neglecting global–local spatial correlations. This leads to suboptimal feature modeling and redundancy in the latent space. To address this limitation, we propose a novel Dual-Scale Transformer (DST) block specifically tailored for NVC, which effectively enhances coding efficiency. The DST block incorporates a Global–Local (Shifted) Window-based Self-Attention (GL(S)WSA) mechanism to jointly capture global structure information and local texture details. Moreover, we design a Cross-Gated Feed-Forward Network (CGFFN) to adaptively modulate complementary components, producing more compact and expressive latent representations. Furthermore, to overcome the drawbacks of traditional asynchronous training and further boost rate-distortion performance, we introduce a Variable Bitrate Synchronization (VBRS) strategy that leverages multi-GPU parallel training, with each GPU dedicated to a specific bitrate and synchronized via gradient backpropagation for joint optimization. Experimental results demonstrate that our proposed method achieves the higher coding performance compared to the previous state-of-the-art (SOTA) methods and significantly outperforms H.266/VVC (VTM-13.2) under various low delay B (LDB) coding configurations. Yiming Wang 0008, Yaojun Wu 0001, Zhaobin Zhang, Qian Huang 0008, Bin Tang 0002, Zhangjing Yang, Kai Zhang 0007, Li Zhang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Efficient Local-Global Collaboration Transcoding for JPEG AIabstractIn the past decade, learning-based image compression has made significant advancements, with the Joint Photographic Experts Group (JPEG) working towards the launch of the first neural network-based image coding standard, JPEG AI. However, traditional codecs like JPEG and HEVC intra coding remain the most widely used. A key challenge arises when attempting to compress images pre-encoded with these traditional codecs using JPEG AI, as such images often carry artifacts that JPEG AI is not fully optimized to handle, leading to significant loss in performance. This paper addresses this issue by proposing a novel transcoding framework designed to optimize pre-encoded images for JPEG AI, bridging the gap between traditional and learning-based compression methods. The framework employs two branches for processing luminance and chrominance components, with a local-global modulation module (LGMM) for luminance and a dynamic fusion module (DFM) for chrominance. Experimental results demonstrate that the proposed scheme achieves up to 22.36% rate savings on images pre-encoded with HEVC intra and up to 22.29% rate savings on images pre-encoded with JPEG using JPEG AI reference software, highlighting its effectiveness in enhancing the performance of JPEG AI for real-world applications. Yiming Wang 0008, Zhaobin Zhang, Yaojun Wu 0001, Qian Huang 0008, Bin Tang 0002, Kai Zhang 0007, Li Zhang 0136 |
ICME | 2 |
| 2025 | Neural Video Compression with In-Loop Contextual Filtering and Out-of-Loop Reconstruction EnhancementabstractThis paper explores the application of enhancement filtering techniques in neural video compression. Specifically, we categorize these techniques into in-loop contextual filtering and out-of-loop reconstruction enhancement based on whether the enhanced representation affects the subsequent coding loop. In-loop contextual filtering refines the temporal context by mitigating error propagation during frame-by-frame encoding. However, its influence on both the current and subsequent frames poses challenges in adaptively applying filtering throughout the sequence. To address this, we introduce an adaptive coding decision strategy that dynamically determines filtering application during encoding. Additionally, out-of-loop reconstruction enhancement is employed to refine the quality of reconstructed frames, providing a simple yet effective improvement in coding efficiency. To the best of our knowledge, this work presents the first systematic study of enhancement filtering in the context of conditional-based neural video compression. Extensive experiments demonstrate a 7.71% reduction in bit rate compared to state-of-the-art neural video codecs, validating the effectiveness of the proposed approach. Yaojun Wu 0001, Chaoyi Lin, Yiming Wang 0008, Semih Esenlik, Zhaobin Zhang, Kai Zhang 0007, Li Zhang 0006 |
ACM Multimedia | 5 |
| 2024 | Leveraging Conv-Attention for Efficient and High-Quality JPEG AI Image CodingabstractIn this paper, we present a Conv-Attention, a decoder-friendly attention mechanism, in an effort to advancing the practical application of the artificial intelligence-based image coding. More specifically, the proposed method is tailored for JPEG AI, which is the latest advanced neural-network based image coding standard. By identifying the obstacles by profiling the decoding complexity of JPEG AI, the attention module accounts for a significant proportion, which mainly attributes to the intricate network structure and involvement of less efficient operations. Conv-Attention model is composed with plain convolution and activation computations, equipping with sub-scaling and up-scaling design, such that the non-adjacent features can be well captured, leading to the reduction of decoding complexity and maintenance of the synthesis and attentive capability. Simulation results verify the effectiveness of the proposed method with JPEG AI reference software, wherein the decoding complexity is reduced by 80% with negligible coding performance loss. The proposed method was adopted in the 100th JPEG meeting. Meng Wang 0017, Semih Esenlik, Zhaobin Zhang, Yaojun Wu 0001, Kai Zhang 0007, Li Zhang 0006, Shiqi Wang 0001 |
DCC | 3 |
| 2024 | Optimized Decoupled Structure with Non-Local Attention for Deep Image CompressionabstractRecently, a decoupled framework for learning-based image compression has been proposed and adopted into the JPEG AI image coding standard developed by ISO/IEC WG1. The decoupled structure disentangles the sample reconstruction process and the entropy decoding process, making the decoding extremely fast. The corresponding techniques constitute the essential parts of the JPEG AI verification model software. However, its analysis transform and synthesis transform are relatively simple, which are built with stacked convolution layers, thereby may lack the capability to interpret data correlations. In this work, we enhance the transform networks by introducing the non-local attention mechanism, which has proven efficient in image compression tasks. The proposed framework thus shares the merits of the fast decoding from the decoupled architecture and the strong transform capabilities from the non-local attention, making it a stronger candidate for practical end-to-end image codec deployment. Experimental results on the Kodak test set and JPEG AI CfP test set show that our method achieves better BDRate performance compared to the original Decoupled-anchor and significantly faster decoding speed compared to NIC. The proposed solution has been adopted by the IEEE 1857.11 Working Subgroup (1857.11 WSG) in developing neural network-based image coding standards in the 10th Meeting. Xuanye Zhang, Zhaobin Zhang, Yaojun Wu 0001, Semih Esenlik, Xiaoyan Sun 0001, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 2 |
| 2024 | Wavelet-like Transform with Subbands Fusion in Decoupled Structure for Deep Image CompressionabstractWavelet-like transform, based on convolutional neural network (CNN), is content-adaptive and has made remarkable achievements in end-to-end image compression. However, the subsequent sequential processing of each subband in the entropy module takes a relatively long decoding time, resulting in incon-venience for real-world applications. In this work, for lossy image compression, the wavelet-like transform is transplanted into the prevailing autoencoder structure to enhance the analysis and synthesis transform due to its excellent decomposition capability. The obtained subbands of different frequencies will undergo a hierarchical decorrelation architecture for subband fusion, also called cross fusing module. The specialized treatment will be applied to different subbands according to their spatial resolution to attain a more compact latent representation. In addition, the proposed solution features an architecture that decouples the arithmetic decoding process from the sample prediction process, which significantly reduces the decoding complexity. Experiments on the Kodak test set show that the proposed method achieves −3.04% BD-Rate compared to existing decoupled end-to-end structure in RGB Peak Signal-to-Noise Ratio (PSNR). Yaojun Wu 0001, Zhaobin Zhang, Semih Esenlik, Xiaoyan Sun 0001, Kai Zhang 0007, Li Zhang 0006 |
PCS | 3 |
| 2024 | End-to-End Learning-Based Image Compression With a Decoupled FrameworkabstractThe autoregressive model has been widely used in learning-based image compression due to its superior context modeling capability. However, its sequential processing nature also undermines the ability of decoding in parallel and hinders the deployment in real applications. In this paper, we propose a decoupled framework to resolve this issue. With the decoupled architecture, the entropy decoding process is independent of the latent sample reconstruction process. The entropy decoding process thus can be finished before the latent sample prediction process begins, which leads to significant decoding time savings by enabling the two processes to be conducted in parallel. To further reduce the decoding time, we introduce wavefront processing, where multiple rows can be processed simultaneously when reconstructing the latent samples. On top of that, we design a series of coding tools to improve the rate-distortion efficiency and reduce the decoding complexity. Device interoperability is also supported by the proposed solution, where the same bitstream can be successfully decoded on different CPU/GPU devices. Comprehensive experiments are conducted to validate the effectiveness of the proposed method. Using objective evaluation metrics required by JPEG AI Call for Proposals (CfP), the proposed method achieves an average of -29.6% BD-rate changes with 2.44 times faster decoding speed compared to VVC image coding. When compared to the commonly used benchmark learning-based methods, the proposed method achieves -30.5% BD-rate changes and 101 times faster decoding speed over cheng2020attn. The proposed solution has been proposed to JPEG AI and IEEE 1857.11 as a response to CfP and the core techniques have been adopted to build the verification model. The software and the instructions can be accessed at https://github.com/bytedance/BEE. Zhaobin Zhang, Semih Esenlik, Yaojun Wu 0001, Meng Wang 0017, Kai Zhang 0007, Li Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Learning-Based End-to-End Video Compression with Spatial-Temporal AdaptationabstractThe learning-based end-to-end video compression exhibits a fast development with continuous improvements. In previous works, a key frame in random-access scenarios is typically compressed by an image compressor, and the remaining frames are reconstructed by interpolations. But solely using image compression fails to leverage the temporal correlations, which is critical to achieve a substantial gain in video compression. To exploit temporal correlations among key frames, we introduce a learning-based end-to-end spatial-temporal adaptive (e2e-STA) compression solution to offer flexible options for the key frames. First of all, we design an extrapolation-based key frame compression scheme. Given a key frame, e2e-STA can switch between an image compressor and an extrapolative compressor. A key frame is thereby able to select the optimal solution adaptively according to the rate-distortion optimization criteria and the optimal selection is sent to the decoder. The proposed approach is optimized end-to-end with all the networks. The experimental results validate the effectiveness of the proposed mechanism. The proposed method outperforms existing learning-based video compression methods by a noticeable margin and provides promising performance compared to traditional benchmark video codecs in terms of PSNR and MS-SSIM. Zhaobin Zhang, Yue Li 0015, Kai Zhang 0007, Li Zhang 0006, Yuwen He |
ICIP | 1 |
| 2022 | Optimized Bit Allocation for Learning-based Video CompressionabstractThe optimized bit allocation among frames has been intensively explored and improved the compression performance significantly in conventional video coding. However, the optimized bit allocation is still in its infant stage for learning-based video coding. Most existing learning-based video compression methods either use uniform bit allocation or empirically determined bit allocation weights among frames. In this paper, we develop an optimized bit allocation scheme for learning-based end-to-end video compression. In particular, we realize a hierarchical quality-control mechanism based on the importance of different frames under random-access scenarios. Considering the varying importance of frames on different temporal layers, we propose an efficient yet simple scheme, in which a set of optimized bit allocation weights are introduced to the rate-distortion (R-D) loss function. Experimental results demonstrate the effectiveness of the proposed scheme. In addition, the proposed scheme can be easily applied to most existing learning-based video compression frameworks under random-access scenarios. Zhaobin Zhang, Yue Li 0015, Kai Zhang 0007, Li Zhang 0006, Yuwen He |
ISCAS | 1 |
| 2021 | CNN-based Super Resolution for Video Coding Using Decoded InformationabstractDown-sampling followed by an up-sampling is a well-known strategy to compress high-resolution pictures given a limited bandwidth in image as well as video coding. Recently, inspired by the latest advances of image super resolution (SR) technologies using convolutional neural network (CNN), CNN-based SR has been explored for resampling-based coding. However, the side information generated during the compression process is not utilized efficiently by the SR network in prior arts. In this paper, we propose a CNN-based SR method for video coding, where more side information is leveraged as a supplement to reconstruction samples. Specifically, we introduce prediction samples to be the auxiliary information, as it can provide the texture and directional information about the original picture. Considering the different characteristics, we design different networks for the luma and chroma components. When designing the chroma up-sampling CNN, the luma reconstruction is used as the auxiliary information of the chroma network, which exploits the cross-component correlation. Experimental results show that the proposed method achieves 11.07% BD-rate savings in all-intra configuration compared with VTM-11.0. Further experiments validate the effectiveness of using luma information to aid the chroma up-sampling process. Chaoyi Lin, Yue Li 0015, Kai Zhang 0007, Zhaobin Zhang, Li Zhang 0006 |
VCIP | 4 |
| 2021 | Fast DST-VII/DCT-VIII With Dual Implementation Support for Versatile Video CodingabstractThe Joint Video Exploration Team (JVET) recently launched the standardization of the next-generation video coding named Versatile Video Coding (VVC) with the inherited technical framework from its predecessor High-Efficiency Video Coding (HEVC). The simplified Enhanced Multiple Transform (EMT) has been adopted as the primary residual coding transform solution, termed Multiple Transform Selection (MTS). In MTS, only the transform set consisting of DST-VII and DCT-VIII remains, excluding the other transform sets and the dependency on intra prediction modes. Significant coding gains are achieved by introducing new DST/DCT transforms, but the full matrix implementation is relatively costly compared to partial butterfly in terms of both software run-time and operation counts. In this work, we exploit the inherent features existing in DST-VII and DCT-VIII. Instead of repeating the element-wise additions and multiplications in full matrix operation, these features can be leveraged to achieve more efficient implementations which only use partial elements to derive the identical results. Existing transform matrices are further tuned to utilize these (anti-)symmetric features. A partial butterfly-type fast algorithm with dual-implementation support is proposed for DST-VII/DCT-VIII transform in VVC. Complexity analysis including operation counts and software run-time are conducted to validate the effectiveness. In addition, we prove the features are perfectly supported by theory. The proposed fast methods achieve noticeable software run-time savings without compromising on coding performance by comparing with the VVC Test Model VTM-3.0. It is shown that under Common Test Condition (CTC) with inter MTS enabled, an average of 9%, 0%, and 3% decoding time savings are achieved for All Intra (AI), Random Access (RA) and Low Delay B (LDB), respectively. Under low QP test condition with inter MTS enabled, the proposed fast methods achieve 1%, 2% and 4% decoding time savings on average for AI, RA, and LDB, respectively. Zhaobin Zhang, Xin Zhao 0003, Xiang Li 0003, Li Li 0040, Shan Liu 0001, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Fast Adaptive Multiple Transform for Versatile Video CodingabstractThe Joint Video Exploration Team (JVET) recently launched the standardization of next-generation video coding named Versatile Video Coding (VVC) in which the Adaptive Multiple Transforms (AMT) is adopted as the primary residual coding transform solution. AMT introduces multiple transforms selected from the DST/DCT families and achieves noticeable coding gains. However, the set of transforms are calculated using direct matrix multiplication which induces higher run-time complexity and limits the application for practical video codec. In this paper, a fast DST-VII/DCT-VIII algorithm based on partial butterfly with dual implementation support is proposed, which aims at achieving reduced operation counts and run-time cost meanwhile yield almost the same coding performance. The proposed method has been implemented on top of the VTM-1.1 and experiments have been conducted using Common Test Conditions (CTC) to validate the efficacy. The experimental results show that the proposed methods, in the state-of-the-art codec, can provide an average of 7%, 5% and 8% overall decoding time savings under All Intra (AI), Random Access (RA) and Low Delay B (LDB) configuration, respectively yet still maintains coding performance. Zhaobin Zhang, Xin Zhao 0003, Xiang Li 0003, Zhu Li 0001, Shan Liu 0001 |
DCC | 1 |
| 2019 | Multiple Linear Regression for High Efficiency Video Intra CodingabstractIn video coding frameworks, the essence of intra coding is leveraging the spatial correlation within a frame to remove redundancy thus achieving compact transmitting data. With modern video acquisition devices improvement, more high-definition videos emerge into people's lives which has set a new challenge for high efficiency video coding. In this paper, we propose a novel intra video coding scheme based on Multiple Linear Regression (MLR), named Multiple linear regression Intra Prediction (MIP). Instead of predicting pixel values by extrapolating, we try to exploit the potential capability of homogeneous regression method. The proposed method has a very concise and neat design yet achieves better performance compared with High Efficiency Video Coding (HEVC) reference software anchor. The experimental results demonstrate the effectiveness of the proposed method and provide interesting insights for further exploiting the capability of conventional algorithms for video coding when many people favor deep learning-based approaches. Zhaobin Zhang, Yue Li 0015, Li Li 0040, Zhu Li 0001, Shan Liu 0001 |
ICASSP | 1 |
| 2019 | Mobile Visual Search Compression With Grassmann Manifold EmbeddingabstractWith the increasing popularity of mobile phones and tablets, the explosive growth of query-by-capture applications calls for a compact representation of the query image feature. Compact descriptors for visual search (CDVS) is a recently released standard from the ISO/IEC moving pictures experts group, which achieves state-of-the-art performance in the context of image retrieval applications. However, they did not consider the matching characteristics in local space in a large-scale database, which might deteriorate the performance. In this paper, we propose a more compact representation with scale invariant feature transform (SIFT) descriptors for the visual query based on Grassmann manifold. Due to the drastic variations in image content, it is not sufficient to capture all the information using a single transform. To achieve more efficient representations, a SIFT manifold partition tree (SMPT) is initially constructed to divide the large dataset into small groups at multiple scales, which aims at capturing more discriminative information. Grassmann manifold is then applied to prune the SMPT and search for the most distinctive transforms. The experimental results demonstrate that the proposed framework achieves state-of-the-art performance on the standard benchmark CDVS dataset. Zhaobin Zhang, Li Li 0040, Zhu Li 0001, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Combining Intra Block Copy and Neighboring Samples Using Convolutional Neural Network for Image CodingabstractIntra prediction is an essential step to remove spatial redundancy in image coding. How to best predict current block given surrounding pixels is the key efficiency challenge. Inspired by the recent success in applying deep learning to image/video coding systems, we propose an intra prediction method by combining intra block copy and neighboring samples using convolutional neural networks. A novel CNN is developed to further exploit the spatial correlation. Instead of only considering local information, the proposed method can infer the current block via fusing the non-local recurrent features, which is captured by intra block copy, with the local samples located at the left and above boundaries of current block. We also investigate how the performance is affected by the way of fusing IBC and reference boundary pixels. In additional, training data pre-processing is studied to enable the CNN with a better learning capability. Simulation results yield promising coding gain and indicate great potential ability that CNN can be used for next generation video coding framework. Zhaobin Zhang, Yue Li 0015, Li Li 0040, Zhu Li 0001, Shan Liu 0001 |
VCIP | 1 |
| 2017 | Robust emotion recognition from low quality and low bit rate video: A deep learning approachabstractEmotion recognition from facial expressions is tremendously useful, especially when coupled with smart devices and wireless multimedia applications. However, the inadequate network bandwidth often limits the spatial resolution of the transmitted video, which will heavily degrade the recognition reliability. We develop a novel framework to achieve robust emotion recognition from low bit rate video. While video frames are downsampled at the encoder side, the decoder is embedded with a deep network model for joint super-resolution (SR) and recognition. Notably, we propose a novel max-mix training strategy, leading to a single “One-for-All” model that is remarkably robust to a vast range of downsampling factors. That makes our framework well adapted for the varied bandwidths in real transmission scenarios, without hampering scalability or efficiency. The proposed framework is evaluated on the AVEC 2016 benchmark, and demonstrates significantly improved stand-alone recognition performance, as well as rate-distortion (R-D) performance, than either directly recognizing from LR frames, or separating SR and recognition. Bowen Cheng, Zhangyang Wang, Zhaobin Zhang, Zhu Li 0001, Ding Liu 0001, Jianchao Yang, Shuai Huang 0001, Thomas S. Huang |
ACII | 3 |
| 2017 | Visual query compression with locality preserving projection on Grassmann manifoldabstractFor a variety of visual search and visual key points based navigation applications, compression of visual key point features like SIFT is an important part of the overall system that can directly affect the efficiency and latency. In this work, we examine a new approach in visual key points compression, that utilizes subspaces that optimized for preserving key point feature matching properties than the reconstruction performance, and allows for a set of optimal subspaces on Grassmann manifold that can better adapt to the local manifold geometry. The simulation demonstrates that such scheme has very low overhead in signaling subspaces, and has very much improved performance on the repeatability of the keypoint matching subject to bit rate constraints. Zhaobin Zhang, Li Li 0040, Zhu Li 0001, Houqiang Li |
ICIP | 1 |
| 2017 | Attribute compression of 3D point clouds using Laplacian sparsity optimized graph transformabstract3D sensing and content capturing have made significant progress in recent years and the MPEG standardization organization is launching a new project on immersive media with point cloud compression (PCC) as one key corner stone. In this work, we introduce a new binary tree based point cloud partition and explore the graph signal processing tools, especially the graph transform with optimized Laplacian sparsity, to achieve better energy compaction and compression efficiency. The resulting rate-distortion operating points are convex-hull optimized over the existing Lagrangian solutions. Simulation results on the latest high quality point cloud content from the MPEG PCC demonstrate the transform efficiency and rate-distortion (R-D) optimal potential of the proposed solutions. Yiting Shao, Zhaobin Zhang, Zhu Li 0001, Kui Fan, Ge Li 0002 |
VCIP | 2 |