Wen Tan 0001

dblp:60/2456-1 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-8560-7554ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Turbo principles meet compression: Rethinking nonlinear transformations in learned image compression
Chao Li 0071, Wen Tan 0001, Fanyang Meng, Runwei Ding, Ye Wang 0002, Wei Liu 0065, Yongsheng Liang 0001
J. Vis. Commun. Image Represent.2
2025 Multimodal-Guided Perceptual Image Compression via Joint Text and Audio
Genhong Wang, Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
ICIC (3)2
2025 Towards Robust Text-Guided Image Compression Under Modality Missing
abstract
Text-guided image compression aims to enhance the perceptual quality of reconstructed images by leveraging textual semantic information. However, existing methods struggle to effectively integrate text information and suffer from significant performance degradation when text is unavailable. To address these issues, we propose a Robust Text-Guided Image Compression (RobustTGIC) network that fully utilizes text semantics when available and mitigates performance loss when absent. Specifically, we introduce a Dual-Dimensional Text Modulation (DDTM) module to enhance perceptual quality by accurately fusing textual information. Building on this, we further propose an Intermediate Feature Modulation (IFM) module, which compensates for missing semantics through lightweight adaptation, improving robustness. Experimental results indicate that, at low bitrates (e.g., 0.07 bpp), our method achieves superior perceptual quality reconstruction while significantly reducing bitrates (e.g., 0.5× HiFiC and 0.4× Bpg). Moreover, our method effectively mitigates performance degradation caused by missing text with a parameter increase of less than 1% of the total parameters.
Genhong Wang, Wen Tan 0001, Fanyang Meng, Yongsheng Liang 0001
ICIP2
2025 Grouped Transform for Ultra-Low-Complexity Learned Image Compression
abstract
Existing learned image compression (LIC) methods have shown strong performance advantages but also bring high computational complexity, making it challenging to deploy them on resource-constrained devices. To reduce the high computational and storage cost, we propose a fully grouped image compression network by introducing spatial and channel grouping operations. Grouping operation is helpful to obtain compact representations by aggregating similar features and reducing redundancy between features in LIC task. Specifically, our proposed network consists of two efficient parts, one is the spatial grouping transform for spatial resolution sampling, and the other is the channel grouping transform for nonlinear representation capability enhancement. Moreover, convolutional kernel factorization and inverted bottleneck are used to reduce redundancy and enrich information of each group in the channel grouping transform, which achieve a good balance between computational complexity and network performance. Experimental results show that our method not only achieves competitive rate-distortion performance with fewer KMACs/pixel and model parameters, but also reduces the real-world runtime. In particular, our proposed models provide at least over 84.6% computational complexity reduction when compared with several advanced LIC methods.
Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001
ISCAS1
2025 Structured Sparsity Learning for Efficient Learned Image Compression
abstract
Existing learned image compression (LIC) methods have achieved outstanding performance, but their deployment on resource-constrained devices is hindered by the high computational complexity and large model storage. Sparsity learning can obtain sparse neural networks by applying regularization term and further achieve model compression by pruning. However, it is difficult to achieve a lightweight LIC network by directly applying sparsity learning and pruning due to unstructured sparsity and limitations of entropy model. In this paper, we propose to add L2,1regularization during the network training for image compression task, which generates structured sparsity at both filter and channel level. We further analyze the effect of entropy model capacity, and adopt filter/channel fixing to achieve the alignment of entropy estimation for actual pruning. Moreover, we utilize incremental regularization to improve the sparsity of network and training stability. Experimental results show that our pruned lightweight model can effectively reduce network parameters by an average of 59.73% at the cost of 1.57% BD-rate increase compared with original hyperprior model.
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Lihan Zhu, Yongsheng Liang 0001
ISCAS1
2025 Stable successive Neural Image Compression via coherent demodulation-based transformation
Youneng Bao, Wen Tan 0001, Mu Li 0005, Fanyang Meng, Yongsheng Liang 0001
Signal Process.2
2025 Adaptive cross-channel transformation based on self-modulation for learned image compression
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Yongsheng Liang 0001
Signal Process. Image Commun.1
2025 ShiftLIC: Lightweight Learned Image Compression With Spatial-Channel Shift Operations
abstract
Learned Image Compression (LIC) has attracted considerable attention due to their outstanding rate-distortion (R-D) performance and flexibility. However, the substantial computational cost poses challenges for practical deployment. The issue of feature redundancy in LIC is rarely addressed. Our findings indicate that many features within the LIC backbone network exhibit similarities. This paper introduces ShiftLIC, a novel and efficient LIC framework that employs parameter-free shift operations to replace large-kernel convolutions, significantly reducing the model’s computational burden and parameter count. Specifically, we propose the Spatial Shift Block (SSB), which combines shift operations with small-kernel convolutions to replace large-kernel. This approach maintains feature extraction efficiency while reducing both computational complexity and model size. To further enhance the representation capability in the channel dimension, we propose a channel attention module based on recursive feature fusion. This module enhances feature interaction while minimizing computational overhead. Additionally, we introduce an improved entropy model integrated with the SSB module, making the entropy estimation process more lightweight and thereby comprehensively reducing computational costs. Experimental results demonstrate that ShiftLIC outperforms leading compression methods, such as VVC Intra and GMM, in terms of computational cost, parameter count, and decoding latency. Additionally, ShiftLIC sets a new SOTA benchmark with a BD-rate gain per MACs/pixel of −102.6%, showcasing its potential for practical deployment in resource-constrained environments. The code is released athttps://github.com/baoyu2020/ShiftLIC.
Youneng Bao, Wen Tan 0001, Chuanmin Jia, Mu Li 0005, Yongsheng Liang 0001, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image Compression
abstract
Deep neural network-based image compression (NIC) has achieved excellent performance, but NIC method models have been shown to be susceptible to backdoor attacks. Adversarial training has been validated in image compression models as a common method to enhance model robustness. However, the improvement effect of adversarial training on model robustness is limited. In this paper, we propose a prior knowledge-guided adversarial training framework for image compression models. Specifically, first, we propose a gradient regularization constraint for training robust teacher models. Subsequently, we design a knowledge distillation-based strategy to generate a priori knowledge from the teacher model to the student model for guiding adversarial training. Experimental results show that our method improves the reconstruction quality by about 9dB when the Kodak dataset is elected as the backdoor attack object for psnr attack. Compared with Ma2023 [1], our method has a 5dB higher PSNR output at high bitrate points.
Youneng Bao, Fanyang Meng, Chao Li 0071, Wen Tan 0001, Genhong Wang, Yongsheng Liang 0001
ICASSP5
2024 Leveraging Redundancy in Feature for Efficient Learned Image Compression
abstract
In recent years, with the development of the field of learned image compression, numerous models with excellent rate-distortion performance have emerged. However, the considerable computational complexity inherent in these models poses challenges for their practical deployment. In this paper, we investigate feature redundancy in learned image compression (LIC) algorithms for efficient feature extraction and introduce an efficient and lightweight LIC framework. Specifically, we explore the existence of a large number of similar features in the network. Subsequently, we design effective feature extraction modules across various levels, such as layer and block. In addition, based on the fact that the role of the codec’s encoder is to remove redundancy and the decoder is to reconstruct, we propose an asynchronous feature fusion block. This fusion block incorporates an "edge smoothing" operator in the encoder and an "edge enhancement" operator in the decoder. Our methodology strikes an ideal balance between rate-distortion performance and efficiency. The experimental results indicate that our approach necessitates only 310KMac/pixel computation and 9.5M parameters, while in terms of performance, our method achieves a 20.7% BD-rate advantage over BPG on Kodak data, mirroring VVC’s performance. Compared to other learned image compression algorithms with SOTA performance, our method has a great advantage in terms of computation/parameter count.
Youneng Bao, Fanyang Meng, Wen Tan 0001, Chao Li 0071, Genhong Wang, Yongsheng Liang 0001
ICASSP4
2024 Enhanced Interpretability in Learned Image Compression via Convolutional Sparse Coding
abstract
Compared to traditional image compression methods, learned image compression (LIC) methods have demonstrated increasingly superior rate-distortion performance. However, LIC networks are often regarded as black boxes, still lacking a theoretical understanding. Sparse coding provides the sparse and interpretable modeling for analyzing or synthesizing natural images in various signal and image processing applications. Therefore, we introduce convolutional sparse coding (CSC) into transform network for enhancing the interpretability of LIC methods. In this paper, we first employ CSC layers to achieve certain theoretical modeling for LIC network, and adopt a weight sharing strategy in encoder-decoder pair and attention mechanism to balance the complexity and performance. Additionally, we analyze the model robustness against data input perturbations and consider the impact of sparsity trade-off parameter in the CSC layer optimization process. Experimental results demonstrate that our method achieves comparable performance with the corresponding baseline, and our model is more robust.
Yiwen Tu, Wen Tan 0001, Youneng Bao, Genhong Wang, Fanyang Meng, Yongsheng Liang 0001
ICME2
2023 A Decoupled Spatial-Channel Inverted Bottleneck For Image Compression
abstract
Residual block has achieved great success in deep networks to eliminate accuracy degradation, and there emerges a large number of variants with more competitive performance. However, these blocks are introduced for high-level tasks that only encode the input image into semantic and struc¬tural features but do not need to reconstruct. So for the low-level task like image compression where reconstruction quality contributes significantly to the rate-distortion perfor¬mance, the structure of the residual block needs modification for more suitable implementation. In this paper, we revisit the existing residual blocks and discover two key principles summarized as two decouplings: spatial-channel decoupling and linear-nonlinear decoupling. We propose an efficient nonlinear transform based on the principles dubbed decou¬pled spatial-channe inverted bottleneck(DSCIB), which has a linear-spatial branch for rough reconstruction and a nonlinear¬channel branch to provide detailed featrues. We employ the DSCIB module in the joint auto regression model to build an overall network. Experimental results show that our method achieves comparable performance with the existing learning-based image compression methods at high bitrate while re¬ducing 38% FLOPs.
Wen Tan 0001, Fanyang Meng, Yongsheng Liang 0001
ICIP2
2023 A Complex-Valued Neural Network Based Robust Image Compression
Can Luo, Youneng Bao, Wen Tan 0001, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001
PRCV (10)3
2023 Taylor series based dual-branch transformation for learned image compression
Youneng Bao, Wen Tan 0001, Linfeng Zheng, Fanyang Meng, Wei Liu 0065, Yongsheng Liang 0001
Signal Process.2
2022 Exploring Structural Sparsity in Neural Image Compression
abstract
The performance of neural image compression have reached or suppressed traditional methods (such as JPEG, BPG, WebP). However, their sophisticated network structures with cascaded convolution layers bring heavy computational burden for practical deployment. In this paper, we explore structural sparsity in neural image compression network to obtain real-time acceleration without any specialized hardware design or algorithm. We propose a simple plug-in adaptive binary channel masking(ABCM) to judge the importance of each convolution channel and introduce sparsity during training. During inference, the unimportant channels are pruned to obtain slimmer network and less computation. We implement our method into three neural image compression networks with different entropy models to verify its effectiveness and generalization, the experiment results show that up to 7× computation reduction and 3× acceleration can be achieved with negligible performance drop.
Shanzhi Yin, Chao Li 0071, Fanyang Meng, Wen Tan 0001, Youneng Bao, Yongsheng Liang 0001, Wei Liu 0065
ICIP4