Chao Li 0071

dblp:66/190-71 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
17since 2021 · last 2026
0009-0007-3639-4223ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Turbo principles meet compression: Rethinking nonlinear transformations in learned image compression
Chao Li 0071, Wen Tan 0001, Fanyang Meng, Runwei Ding, Ye Wang 0002, Wei Liu 0065, Yongsheng Liang 0001
J. Vis. Commun. Image Represent.1
2026 Entropy-aware image representation via 2D Gaussian splatting
Jiacong Chen, Qingyu Mao, Shuai Liu 0022, Chao Li 0071, Jierun Lin, Xiandong Meng, Fanyang Meng, Yongsheng Liang 0001
Signal Process.4
2025 DMSO: A Dynamic Momentum-Smoothing Optimizer for Learned Image Compression
abstract
Learned Image Compression (LIC) has rapidly evolved and recently surpassed traditional methods in Rate-Distortion (R-D) performance. However, most LIC approaches improve network architectures while increasing computational overhead and overlooking using optimizers tailored specifically for LIC. This paper proposes a Dynamic Momentum-Smoothing Optimizer (DMSO) tailored for LIC to achieve faster convergence and better R-D performance. Specifically, DMSO leverages historical gradient information to smooth the optimization process dynamically, thereby reducing in-stability from gradient oscillations. Furthermore, DMSO introduces a novel Enhanced Second-order Momentum mechanism to mitigate cumulative noise and align momentum updates more closely with the true gradient. Experimental results demonstrate that DMSO operates as a universal optimization plugin for LIC methods, achieving faster and more stable convergence while improving R-D performance to varying degrees without additional parameter count or computational cost.
Chao Li 0071, Chuanmin Jia, Fanyang Meng, Siwei Ma 0001, Yongsheng Liang 0001
ICIP1
2025 Structured Sparsity Learning for Efficient Learned Image Compression
abstract
Existing learned image compression (LIC) methods have achieved outstanding performance, but their deployment on resource-constrained devices is hindered by the high computational complexity and large model storage. Sparsity learning can obtain sparse neural networks by applying regularization term and further achieve model compression by pruning. However, it is difficult to achieve a lightweight LIC network by directly applying sparsity learning and pruning due to unstructured sparsity and limitations of entropy model. In this paper, we propose to add L2,1regularization during the network training for image compression task, which generates structured sparsity at both filter and channel level. We further analyze the effect of entropy model capacity, and adopt filter/channel fixing to achieve the alignment of entropy estimation for actual pruning. Moreover, we utilize incremental regularization to improve the sparsity of network and training stability. Experimental results show that our pruned lightweight model can effectively reduce network parameters by an average of 59.73% at the cost of 1.57% BD-rate increase compared with original hyperprior model.
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Lihan Zhu, Yongsheng Liang 0001
ISCAS4
2025 Adaptive cross-channel transformation based on self-modulation for learned image compression
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Yongsheng Liang 0001
Signal Process. Image Commun.4
2025 One is All: A Unified Rate-Distortion-Complexity Framework for Learned Image Compression Under Energy Concentration Criteria
abstract
The learned image compression (LIC) technique has surpassed the state-of-the-art traditional codecs (H.266/VVC) in case of rate-distortion (R-D) performance. Its real-time deployments are far advanced. In order to achieve more flexible deployments, an LIC technique should be flexible in adjusting its computational complexity and rate as demanded by a situation and its environment. In this paper, we propose a unified Rate-Distortion-Complexity (R-D-C) framework for LIC under channel energy concentration criteria. Specifically, we first introduce an Energy Asymptotic Nonlinear Transformation (EANT) designed to directly concentrate on the channel energy of latent representations, thus laying the groundwork for a scalable entropy coding. Next, leveraging this energy concentration characteristic, we propose a corresponding Heterogeneous Scalable Entropy Model (HSEM) for flexibly scaling bitstreams as needed. Finally, utilizing the proposed EANT, we construct a fine-grained scalable codec for formulating, in combination with HSEM, a comprehensive scalable R-D-C framework under the energy concentration criteria. The obtained experimental results demonstrate that the proposed method could enable seamless transitions between 13 different widths of sub-models within a single network, allowing for fine-grained control over the model bitrate, complexity, and hardware inference time. Additionally, the proposed method exhibits competitive R-D performance compared to many existing methods.
Chao Li 0071, Fanyang Meng, Qingyu Mao, Youneng Bao, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Multim.1
2024 Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image Compression
abstract
Deep neural network-based image compression (NIC) has achieved excellent performance, but NIC method models have been shown to be susceptible to backdoor attacks. Adversarial training has been validated in image compression models as a common method to enhance model robustness. However, the improvement effect of adversarial training on model robustness is limited. In this paper, we propose a prior knowledge-guided adversarial training framework for image compression models. Specifically, first, we propose a gradient regularization constraint for training robust teacher models. Subsequently, we design a knowledge distillation-based strategy to generate a priori knowledge from the teacher model to the student model for guiding adversarial training. Experimental results show that our method improves the reconstruction quality by about 9dB when the Kodak dataset is elected as the backdoor attack object for psnr attack. Compared with Ma2023 [1], our method has a 5dB higher PSNR output at high bitrate points.
Youneng Bao, Fanyang Meng, Chao Li 0071, Wen Tan 0001, Genhong Wang, Yongsheng Liang 0001
ICASSP4
2024 Leveraging Redundancy in Feature for Efficient Learned Image Compression
abstract
In recent years, with the development of the field of learned image compression, numerous models with excellent rate-distortion performance have emerged. However, the considerable computational complexity inherent in these models poses challenges for their practical deployment. In this paper, we investigate feature redundancy in learned image compression (LIC) algorithms for efficient feature extraction and introduce an efficient and lightweight LIC framework. Specifically, we explore the existence of a large number of similar features in the network. Subsequently, we design effective feature extraction modules across various levels, such as layer and block. In addition, based on the fact that the role of the codec’s encoder is to remove redundancy and the decoder is to reconstruct, we propose an asynchronous feature fusion block. This fusion block incorporates an "edge smoothing" operator in the encoder and an "edge enhancement" operator in the decoder. Our methodology strikes an ideal balance between rate-distortion performance and efficiency. The experimental results indicate that our approach necessitates only 310KMac/pixel computation and 9.5M parameters, while in terms of performance, our method achieves a 20.7% BD-rate advantage over BPG on Kodak data, mirroring VVC’s performance. Compared to other learned image compression algorithms with SOTA performance, our method has a great advantage in terms of computation/parameter count.
Youneng Bao, Fanyang Meng, Wen Tan 0001, Chao Li 0071, Genhong Wang, Yongsheng Liang 0001
ICASSP5
2024 Fine-Grained Adjustable Entropy Models for Rate-Complexity Jointly Adjustable Image Compression
Chao Li 0071, Shanzhi Yin, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
PRCV (9)2
2024 Multirate Progressive Entropy Model for Learned Image Compression
abstract
This paper proposes a unified and efficient entropy coding method for learned image compression (LIC) from the perspective of traditional signal processing. First, the consistency of structures and optimization objectives are used to interpret the existing split-coded-then-merge entropy coding strategies in LIC as a particular filter banks framework, with feature separation and feature aggregation representing the analysis filter bank and synthesis filter bank, respectively. Thus, we borrow the design from the multirate filter banks and proposed Multirate Progressive Entropy Model (MPEM) to enhance the rate-distortion performance and decoding speed. In particular, we create an analysis filter bank that divides compact features into a few nonuniform subsets based on various spatial and channel sampling rates. Then multi-scale detail and mean coefficients within the current subset are used as prior representations to help generate the prediction parameters of the next subset, and the carefully designed synthetic filter bank performs a near-perfect reconstruction of the features. In addition, we propose a Multi-level Edge Attention Moudal (MEAM) to increase the edge and texture information’s contribution and reduce the high-frequency information loss brought on by MPEM’s inherent multi-rate spatial sampling, which leverages the edge operator and structural reparameterization principles. The results of the experiments show that, in comparison to the effective LIC methods and traditional code, the proposed MPEM can decode data at a cutting-edge speed while also offering comparable rate-distortion performance.
Chao Li 0071, Shanzhi Yin, Chuanmin Jia, Fanyang Meng, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 A Complex-Valued Neural Network Based Robust Image Compression
Can Luo, Youneng Bao, Wen Tan 0001, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001
PRCV (10)4
2023 Nonlinear Transforms in Learned Image Compression From a Communication Perspective
abstract
Recently, remarkable progress has been made in learned image compression (LIC), in which nonlinear transforms (NTs) play a crucial role. Although there are many NT methods for improving the rate distortion performance, all the existing methods sacrifice the computational complexity and the number of parameters of the transformation. This paper provides a fundamental novel viewpoint on nonlinear transforms from a communication perspective, and shows how this idea can be extended to design efficient NT methods. In particular, the nonlinear transforms are inferred as signal modulation modules. Under this extrapolation, the current NTs are generalized as amplitude modulation that only varies the amplitude of the carrier wave. Therefore, a nonlinear modulation-like transform (NMLT) which varies the phase angle of the carrier is proposed. Moreover, this concept is extended by introducing In-phase/Quadrature (IQ) modulation, which is a boosting technique in communication field, in order to enhance NMLT. Furthermore, the Bit-interleaved technique in communication is used to guide the optimization of NTML with IQ. The experimental results on different datasets and backbone architectures verify the efficiency and robustness of the proposed methods. For example, when backbone architecture is hyperprior model, our method achieves 19.37% BD-rate reduction over GDN on the Kodak dataset. In addition, our method with channel wise autoregressive model leads to the state-of-the-art rate-distortion performance.
Youneng Bao, Fanyang Meng, Chao Li 0071, Siwei Ma 0001, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 AdderIC: Towards Low Computation Cost Image Compression
abstract
Recently, learned image compression methods have shown their outstanding rate-distortion performance when compared to traditional frameworks. Although numerous progress has been made in learned image compression, the computation cost is still at a high level. To address this problem, we propose AdderIC, which utilizes adder neural networks (AdderNet) to construct an image compression framework. According to the characteristics of image compression, we introduce several strategies to improve the performance of AdderNet in this field. Specifically, Haar Wavelet Transform is adopted to make AdderIC learn high-frequency information efficiently. In addition, implicit deconvolution with the kernel size of 1 is applied after each adder layer to reduce spatial redundancies. Moreover, we develop a novel Adder-ID-PixelShuffle cascade upsampling structure to remove checkerboard artifacts. Experiments demonstrate that our AdderIC model can largely outperform conventional AdderNet when applied in image compression and achieve comparable rate-distortion performance to that of its CNN baseline with about 80% multiplication FLOPs and 30% energy consumption reduction.
Xin Yao 0001, Chao Li 0071, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
ICASSP3
2022 Universal Efficient Variable-Rate Neural Image Compression
abstract
Recently, Learning-based image compression has reached comparable performance with traditional image codecs(such as JPEG, BPG, WebP). However, computational complexity and rate flexibility are still two major challenges for its practical deployment. To tackle these problems, this paper proposes two universal modules named Energy-based Channel Gating(ECG) and Bit-rate Modulator(BM), which can be directly embedded into existing end-to-end image compression models. ECG uses dynamic pruning to reduce FLOPs for more than 50% in convolution layers, and a BM pair can modulate the latent representation to control the bit-rate in a channel-wise manner. By implementing these two modules, existing learning-based image codecs can obtain ability to output arbitrary bit-rate with a single model and reduced computation.
Shanzhi Yin, Chao Li 0071, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng, Wei Liu 0065
ICASSP2
2022 Exploring Structural Sparsity in Neural Image Compression
abstract
The performance of neural image compression have reached or suppressed traditional methods (such as JPEG, BPG, WebP). However, their sophisticated network structures with cascaded convolution layers bring heavy computational burden for practical deployment. In this paper, we explore structural sparsity in neural image compression network to obtain real-time acceleration without any specialized hardware design or algorithm. We propose a simple plug-in adaptive binary channel masking(ABCM) to judge the importance of each convolution channel and introduce sparsity during training. During inference, the unimportant channels are pruned to obtain slimmer network and less computation. We implement our method into three neural image compression networks with different entropy models to verify its effectiveness and generalization, the experiment results show that up to 7× computation reduction and 3× acceleration can be achieved with negligible performance drop.
Shanzhi Yin, Chao Li 0071, Fanyang Meng, Wen Tan 0001, Youneng Bao, Yongsheng Liang 0001, Wei Liu 0065
ICIP2
2021 Mbb: A Multi-Scale Method For Data Based On Bit Plane Slicing
abstract
Multi-scale methodology can enhance the performance of the model in deep learning. The current multi-scale methodology focuses on changing the formation, which will increase the parameters and calculations of the network. This paper offers a multi-scale method for data based on bit plane slicing(MBB). This expands the receptive field of valid information in image data. It is done by multi-level fusing image with high bit planes. Our experimentation shows that by adding MBB in front of the backbone network, one can achieve a significant performance improvement. The MBB approach is widely applicable because it does not require changes to the structure of the backbone network.
Youneng Bao, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001, Wei Liu 0065, Kaiyu Liu
ICIP2
2021 Improving Convolutional Networks with Boosting Attention Convolutions
abstract
Convolutional neural networks (CNNs) have been widely used in a range of tasks because of its robust convolutional feature transformation ability. In this paper, we propose a novel type of convolution called Boosting Attention Convolution (BAC) to improve the basic convolutional feature transformation process of CNNs. The proposed method is designed based on two principles, boosting and attention mechanism. Specifically, we design a set of simple yet effective Boosting Attention Modules (BAM) within grouped convolution, which progressively recalibrate distribution of feature map and enable the future filters nested in a convolution layer to focus more on the feature regions that are unactivated by previous filters. Thus, it can help CNNs generate more discriminative representations by explicitly incorporating richer information. The experimental results on various datasets verify that BAC outperforms state-of-the-art methods. More importantly, the proposed BAC is a general convolution that can be deployed to various modern networks without introducing much parameters and computational complexity.
Chao Li 0071, Yongsheng Liang 0001, Huo-Xiang Yang, Fanyang Meng, Wei Liu 0065, Handong Wang
ICME1