EDBT 2026 Demo / reviewers in the wild / expert
Haotian Zhang 0009
dblp:83/4184-9
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0004-7193-9127ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical Image Compression with Energy-Guided Asymmetric Entropy ModelingabstractRecent learned image compression (LIC) methods have demonstrated superior ratedistortion performance compared to classical image compression standards. However, their substantially higher computational complexity and inefficient entropy modeling designs raise challenges to practical deployment. To address this issue, we propose Energy-Guided Asymmetric Entropy Modeling for ultra-low-complexity learned image compression. Building on the observation that low-energy channels can be effectively modeled with less complexity, we propose a novel Asymmetric Hyperprior Transform (AHT). AHT evenly split the latent features into channel groups. Low-energy groups are processed by lightweight subnetworks, achieving reduced complexity. To further ensure proper energy allocation on channel groups, we design an Energy Allocation (EA) Loss that constrains latent features with respect to their estimated mean, thus enabling flexible control over channel energy distribution. Overall, our entropy modeling design is extremely lightweight and can be efficiently deployed on CPUs, making it well-suited for practical deployment. Experiments demonstrate that our proposed model achieves a 5% reduction in BD-rate over BPG with decoding complexity below$10 \text{kMACs} /$pixel, and achieves a BD-rate gain per decoding MACs/pixel of -645.89, striking a favorable balance between rate-distortion performance and computational cost. Yiheng Jiang, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
DCC | 4 |
| 2026 | USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020sabstractImage/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD. Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 4 |
| 2025 | Few-Shot Domain Adaptation for Learned Image CompressionabstractLearned image compression (LIC) has achieved state-of-the-art rate-distortion performance, deemed promising for next-generation image compression techniques. However, pre-trained LIC models usually suffer from significant performance degradation when applied to out-of-training-domain images, implying their poor generalization capabilities. To tackle this problem, we propose a few-shot domain adaptation method for LIC by integrating plug-and-play adapters into pre-trained models. Drawing inspiration from the analogy between latent channels and frequency components, we examine domain gaps in LIC and observe that out-of-training-domain images disrupt pre-trained channel-wise decomposition. Consequently, we introduce a method for channel-wise re-allocation using convolution-based adapters and low-rank adapters, which are lightweight and compatible to mainstream LIC schemes. Extensive experiments across multiple domains and multiple representative LIC schemes demonstrate that our method significantly enhances pre-trained models, achieving comparable performance to H.266/VVC intra coding with merely 25 target-domain samples. Additionally, our method matches the performance of full-model finetune while transmitting fewer than 2% of the parameters. Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
AAAI | 2 |
| 2025 | Learned Image Compression with Hierarchical Progressive Context ModelingabstractContext modeling is essential in learned image compression for accurately estimating the distribution of latents. While recent advanced methods have expanded context modeling capacity, they still struggle to efficiently exploit long-range dependency and diverse context information across different coding steps. In this paper, we introduce a novel Hierarchical Progressive Context Model (HPCM) for more efficient context information acquisition. Specifically, HPCM employs a hierarchical coding schedule to sequentially model the contextual dependencies among latents at multiple scales, which enables more efficient long-range context modeling. Furthermore, we propose a progressive context fusion mechanism that incorporates contextual information from previous coding steps into the current step, effectively exploiting diverse contextual information. Experimental results demonstrate that our method achieves state-of-the-art rate-distortion performance and strikes a better balance between compression performance and computational complexity. The code is available at https://github.com/lyq133/LIC-HPCM. Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
ICCV | 2 |
| 2025 | Transformer-Based Channel Autoregressive with Space-to-Channel Context Ordering for Learned Image CompressionabstractIn recent years, significant advancements have been made in learned image compression (LIC). In LIC frameworks, the entropy model plays a vital role in predicting the probability distribution of latent features. The transformer–based channel autoregressive entropy model has recently demonstrated impressive performance. However, it solely exploits channel–wise context while ignoring rich spatial dependencies. To address this limitation, we propose a space–to–channel context ordering approach. Specifically, we introduce a space–to–channel operation that enables transformers to jointly utilize spatial and channel context. For more effective context modeling, we further analyze the context correlations between latent feature groups and present a mixed space–channel context ordering strategy. Moreover, we incorporate zero–centered quantization and optimize the entropy coder to further enhance performance. Experimental results demonstrate that our method achieves approximately 4% and 6% rate saving on the Kodak and the Tecnick datasets compared to the original transformer–based channel autoregressive entropy model. Haotian Zhang 0009, Xiongzhuang Liang, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
VCIP | 2 |
| 2025 | Sparse Point Clouds Assisted Learned Image CompressionabstractIn the field of autonomous driving, a variety of sensor data types exist, each representing different modalities of the same scene. Therefore, it is feasible to utilize data from other sensors to facilitate image compression. However, few techniques have explored the potential benefits of utilizing inter-modality correlations to enhance the image compression performance. In this paper, motivated by the recent success of learned image compression, we propose a new framework that uses sparse point clouds to assist in learned image compression in the autonomous driving scenario. We first project the 3D sparse point cloud onto a 2D plane, resulting in a sparse depth map. Utilizing this depth map, we proceed to predict camera images. Subsequently, we use these predicted images to extract multi-scale structural features. These features are then incorporated into learned image compression pipeline as additional information to improve the compression performance. Our proposed framework is compatible with various mainstream learned image compression models, and we validate our approach using different existing image compression methods. The experimental results show that incorporating point cloud assistance into the compression pipeline consistently enhances the performance. Yiheng Jiang, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Learning Switchable Priors for Neural Image CompressionabstractNeural image compression (NIC) usually adopts a predefined family of probabilistic distributions as the prior of the latent variables, and meanwhile relies on entropy models to estimate the parameters for the probabilistic family. More complex probabilistic distributions may fit the latent variables more accurately, but also incur higher complexity of the entropy models, limiting their practical value. To address this dilemma, we propose a solution to decouple the entropy model complexity from the prior distributions. We use a finite set of trainable priors that correspond to samples of the parametric probabilistic distributions. We train the entropy model to predict the index of the appropriate prior within the set, rather than the specific parameters. Switching between the trained priors further enables us to embrace a skip mode into the prior set, which simply omits a latent variable during the entropy coding. To demonstrate the practical value of our solution, we present a lightweight NIC model, namely FastNIC, together with the learning of switchable priors. FastNIC obtains a better trade-off between compression efficiency and computational complexity for neural image compression. We also implanted the switchable priors into state-of-the-art NIC models and observed improved compression efficiency with a significant reduction of entropy coding complexity. Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Generalized Gaussian Model for Learned Image CompressionabstractIn learned image compression, probabilistic models play an essential role in characterizing the distribution of latent variables. The Gaussian model with mean and scale parameters has been widely used for its simplicity and effectiveness. Probabilistic models with more parameters, such as the Gaussian mixture models, can fit the distribution of latent variables more precisely, but the corresponding complexity is higher. To balance the compression performance and complexity, we extend the Gaussian model to the generalized Gaussian family for more flexible latent distribution modeling, introducing only one additional shape parameter than the Gaussian model. To enhance the performance of the generalized Gaussian model by alleviating the train-test mismatch, we propose improved training methods, including -dependent lower bounds for scale parameters and gradient rectification. Our proposed generalized Gaussian model, coupled with the improved training methods, is demonstrated to outperform the Gaussian and Gaussian mixture models on a variety of learned image compression networks. Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Image Process. | 1 |
| 2024 | Offline and Online Optical Flow Enhancement for Deep Video CompressionabstractVideo compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these networks often adopt pre-trained optical flow estimation networks for motion estimation. The optical flows, however, may be less suitable for video compression due to the following two factors. First, the optical flow estimation networks were trained to perform inter-frame prediction as accurately as possible, but the optical flows themselves may cost too many bits to encode. Second, the optical flow estimation networks were trained on synthetic data, and may not generalize well enough to real-world videos. We address the twofold limitations by enhancing the optical flows in two stages: offline and online. In the offline stage, we fine-tune a trained optical flow estimation network with the motion information provided by a traditional (non-deep) video compression scheme, e.g. H.266/VVC, as we believe the motion information of H.266/VVC achieves a better rate-distortion trade-off. In the online stage, we further optimize the latent features of the optical flows with a gradient descent-based algorithm for the video to be compressed, so as to enhance the adaptivity of the optical flows. We conduct experiments on two state-of-the-art deep video compression schemes, DCVC and DCVC-DC. Experimental results demonstrate that the proposed offline and online enhancement together achieves on average 13.4% bitrate saving for DCVC and 4.1% bitrate saving for DCVC-DC on the tested videos, without increasing the model or computational complexity of the decoder side. Chuanbo Tang, Xihua Sheng, Zhuoyuan Li 0001, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
AAAI | 4 |
| 2024 | Practical Learned Image Compression with Online Encoder OptimizationabstractLearned image compression methods have shown significant advances in performance. However, they often suffer from higher decoding complexity compared to traditional codecs. In this paper, we present an approach toward practical learned image compression that focuses on faster decoding by employing online encoder optimization. To reduce network complexity, we employ a hyperprior structure with smaller convolution kernels, which enables efficient compression. We also introduce a content adaptive skipping algorithm to accelerate entropy decoding. By jointly optimizing entropy decoding complexity and the distortion of the reconstructed image at the encoder side, we achieve faster decoding while maintaining the quality of the reconstructed image. Furthermore, we incorporate online iterative optimization at the encoder side to enhance rate-distortion performance without increasing decoding complexity. This approach allows us to fine-tune the latent and improve overall performance. Our method achieves a notable improvement over BPG on the VCIP 2023 challenge test set, with more than 1 dB increase in PSNR and a 50% reduction in decoding time. Haotian Zhang 0009, Feihong Mei, Junqi Liao, Li Li 0040, Houqiang Li, Dong Liu 0002 |
PCS | 1 |
| 2024 | Deviation Control for Learned Image CompressionabstractMost approaches in learned image compression follow the transform coding scheme. The characteristics of latent variables transformed from images significantly influence the performance of codecs. In this paper, we present visual analyses on latent features of learned image compression and find that the latent variables are spread over a wide range, which may lead to complex entropy coding processes. To address this, we introduce a Deviation Control (DC) method, which applies a constraint loss on latent features and entropy parameter μ. Training with DC loss, we obtain latent features with smaller values of coding symbols and σ, effectively reducing entropy coding complexity. Our experimental results show that the plug-and-play DC loss reduces entropy coding time by 30-40% and improves compression performance. Haotian Zhang 0009, Xiaomin Song, Huiming Zheng, Li Li 0040, Dong Liu 0002 |
VCIP | 2 |
| 2023 | Padding-Aware Learned Image CompressionabstractFor current learned image compression methods, padding input images is necessary to meet the resolution requirements of down-sampling layers. However, the impact of padding has not been studied thoroughly. Most previous studies ignore padded images in the training process. In this paper, we analyze the impact of padding on compression performance. Then, we propose a padding-aware training (PAT) strategy, handling the padding effect during the training. Specifically, our PAT strategy calculates the loss of pre-padding image through a masking operation. Finally, according to our systematic experimental results, we find that images with different resolutions tend to favor different padding modes. Therefore, we further propose to conduct padding mode decision in the encoding process for rate-distortion optimization. Experiments demonstrate that our proposed PAT strategy and padding mode decision effectively compensate for the performance drop caused by padding. Haotian Zhang 0009, Junqi Liao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
ISCAS | 1 |
| 2023 | Flexible Coding Order for Learned Image CompressionabstractLearned image compression (LIC) methods have made significant advances in recent years. In LIC, entropy model is an essential component, which utilizes conditional information to predict the probability distribution over the latent space. In the entropy models, many context models follow a spatially autoregressive paradigm, which leads to sequential coding. The autoregressive coding order, however, may be neither optimal nor efficient. We conduct an elaborate study on coding orders in LIC entropy models, shedding light on the potential for improving compression performance by adapting the coding order. We present Mask Modeling Context Model (MMCM), a transformer-based context model designed with a Patch-restricted Iterative Mask Modeling (PIMM) training strategy. Through training to predict the probability distribution of randomly masked tokens, we are able to use a patch-level arbitrary coding order to encode/decode the latent space in a few iterative steps. Additionally, we employ offline and online rate-distortion optimization to adaptively select the appropriate coding order for each image. Our extensive experimental results demonstrate that the proposed method achieves a better rate-distortion performance than the autoregressive context model while requiring less compression complexity. Haotian Zhang 0009 |
VCIP | 2 |