EDBT 2026 Demo / reviewers in the wild / expert
Youneng Bao
dblp:307/3082
· DBLP profile ↗
24ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0003-3781-6938ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image CompressionabstractPrevailing quantization techniques in Learned Image Compression (LIC) typically employ a static, uniform bit-width across all layers, failing to adapt to the highly diverse data distributions and sensitivity characteristics inherent in LIC models. This leads to a suboptimal trade-off between performance and efficiency. In this paper, we introduce DynaQuant, a novel framework for dynamic mixed-precision quantization that operates on two complementary levels. First, we propose content-aware quantization, where learnable scaling and offset parameters dynamically adapt to the statistical variations of latent features. This fine-grained adaptation is trained end-to-end using a novel Distance-aware Gradient Modulator (DGM), which provides a more informative learning signal than the standard Straight-Through Estimator. Second, we introduce a data-driven, dynamic bit-width selector that learns to assign an optimal bit precision to each layer, dynamically reconfiguring the network's precision profile based on the input data. Our fully dynamic approach offers substantial flexibility in balancing rate-distortion (R-D) performance and computational cost. Experiments demonstrate that DynaQuant achieves R-D performance comparable to full-precision models while significantly reducing computational and storage requirements, thereby enabling the practical deployment of advanced LIC on diverse hardware platforms. Youneng Bao, Yulong Cheng, Mu Li 0005, Yongsheng Liang 0001 |
AAAI | 1 |
| 2026 | Gradient descent-driven sampling for multimodal long-term scanpath prediction in panoramic videos
Tianming Zhou, Yulong Cheng, Kanglong Fan, Youneng Bao, Mu Li 0005 |
Pattern Recognit. | 4 |
| 2025 | Dataset Distillation as Data Compression: A Rate-Utility PerspectiveabstractDriven by the ``scale-is-everything'' paradigm, modern machine learning increasingly demands ever-larger datasets and models, yielding prohibitive computational and storage requirements. Dataset distillation mitigates this by compressing an original dataset into a small set of synthetic samples, while preserving its full utility. Yet, existing methods either maximize performance under fixed storage budgets or pursue suitable synthetic data representations for redundancy removal, without jointly optimizing both objectives. In this work, we propose a joint rate-utility optimization method for dataset distillation. We parameterize synthetic samples as optimizable latent codes decoded by extremely lightweight networks. We estimate the Shannon entropy of quantized latents as the rate measure and plug any existing distillation loss as the utility measure, trading them off via a Lagrange multiplier. To enable fair, cross-method comparisons, we introduce bits per class (bpc), a precise storage metric that accounts for sample, label, and decoder parameter costs. On CIFAR-10, CIFAR-100, and ImageNet-128, our method achieves up to $170\times$ greater compression than standard distillation at comparable accuracy. Across diverse bpc budgets, distillation losses, and backbone architectures, our approach consistently establishes better rate-utility trade-offs. Youneng Bao, Yongsheng Liang 0001, Mu Li 0005, Kede Ma |
ICCV | 1 |
| 2025 | Multimodal-Guided Perceptual Image Compression via Joint Text and Audio
Genhong Wang, Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
ICIC (3) | 3 |
| 2025 | Grouped Transform for Ultra-Low-Complexity Learned Image CompressionabstractExisting learned image compression (LIC) methods have shown strong performance advantages but also bring high computational complexity, making it challenging to deploy them on resource-constrained devices. To reduce the high computational and storage cost, we propose a fully grouped image compression network by introducing spatial and channel grouping operations. Grouping operation is helpful to obtain compact representations by aggregating similar features and reducing redundancy between features in LIC task. Specifically, our proposed network consists of two efficient parts, one is the spatial grouping transform for spatial resolution sampling, and the other is the channel grouping transform for nonlinear representation capability enhancement. Moreover, convolutional kernel factorization and inverted bottleneck are used to reduce redundancy and enrich information of each group in the channel grouping transform, which achieve a good balance between computational complexity and network performance. Experimental results show that our method not only achieves competitive rate-distortion performance with fewer KMACs/pixel and model parameters, but also reduces the real-world runtime. In particular, our proposed models provide at least over 84.6% computational complexity reduction when compared with several advanced LIC methods. Wen Tan 0001, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
ISCAS | 2 |
| 2025 | Structured Sparsity Learning for Efficient Learned Image CompressionabstractExisting learned image compression (LIC) methods have achieved outstanding performance, but their deployment on resource-constrained devices is hindered by the high computational complexity and large model storage. Sparsity learning can obtain sparse neural networks by applying regularization term and further achieve model compression by pruning. However, it is difficult to achieve a lightweight LIC network by directly applying sparsity learning and pruning due to unstructured sparsity and limitations of entropy model. In this paper, we propose to add L2,1regularization during the network training for image compression task, which generates structured sparsity at both filter and channel level. We further analyze the effect of entropy model capacity, and adopt filter/channel fixing to achieve the alignment of entropy estimation for actual pruning. Moreover, we utilize incremental regularization to improve the sparsity of network and training stability. Experimental results show that our pruned lightweight model can effectively reduce network parameters by an average of 59.73% at the cost of 1.57% BD-rate increase compared with original hyperprior model. Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Lihan Zhu, Yongsheng Liang 0001 |
ISCAS | 2 |
| 2025 | Motion Matters: Compact Gaussian Streaming for Free-Viewpoint Video Reconstructionabstract3D Gaussian Splatting (3DGS) has emerged as a high-fidelity and efficient paradigm for online free-viewpoint video (FVV) reconstruction, offering viewers rapid responsiveness and immersive experiences. However, existing online methods face challenge in prohibitive storage requirements primarily due to point-wise modeling that fails to exploit the motion properties. To address this limitation, we propose a novel Compact Gaussian Streaming (ComGS) framework, leveraging the locality and consistency of motion in dynamic scene, that models object-consistent Gaussian point motion through keypoint-driven motion representation. By transmitting only the keypoint attributes, this framework provides a more storage-efficient solution. Specifically, we first identify a sparse set of motion-sensitive keypoints localized within motion regions using a viewspace gradient difference strategy. Equipped with these keypoints, we propose an adaptive motion-driven mechanism that predicts a spatial influence field for propagating keypoint motion to neighboring Gaussian points with similar motion. Moreover, ComGS adopts an error-aware correction strategy for key frame reconstruction that selectively refines erroneous regions and mitigates error accumulation without unnecessary overhead. Overall, ComGS achieves a remarkable storage reduction of over 159 × compared to 3DGStream and 14 × compared to the SOTA method QUEEN, while maintaining competitive visual fidelity and rendering speed. Project page: https://chenjiacong-1005.github.io/ComGS/. Jiacong Chen, Qingyu Mao, Youneng Bao, Xiandong Meng, Fanyang Meng, Ronggang Wang, Yongsheng Liang 0001 |
NeurIPS | 3 |
| 2025 | Stable successive Neural Image Compression via coherent demodulation-based transformation
Youneng Bao, Wen Tan 0001, Mu Li 0005, Fanyang Meng, Yongsheng Liang 0001 |
Signal Process. | 1 |
| 2025 | Adaptive cross-channel transformation based on self-modulation for learned image compression
Wen Tan 0001, Youneng Bao, Fanyang Meng, Chao Li 0071, Yongsheng Liang 0001 |
Signal Process. Image Commun. | 2 |
| 2025 | ShiftLIC: Lightweight Learned Image Compression With Spatial-Channel Shift OperationsabstractLearned Image Compression (LIC) has attracted considerable attention due to their outstanding rate-distortion (R-D) performance and flexibility. However, the substantial computational cost poses challenges for practical deployment. The issue of feature redundancy in LIC is rarely addressed. Our findings indicate that many features within the LIC backbone network exhibit similarities. This paper introduces ShiftLIC, a novel and efficient LIC framework that employs parameter-free shift operations to replace large-kernel convolutions, significantly reducing the model’s computational burden and parameter count. Specifically, we propose the Spatial Shift Block (SSB), which combines shift operations with small-kernel convolutions to replace large-kernel. This approach maintains feature extraction efficiency while reducing both computational complexity and model size. To further enhance the representation capability in the channel dimension, we propose a channel attention module based on recursive feature fusion. This module enhances feature interaction while minimizing computational overhead. Additionally, we introduce an improved entropy model integrated with the SSB module, making the entropy estimation process more lightweight and thereby comprehensively reducing computational costs. Experimental results demonstrate that ShiftLIC outperforms leading compression methods, such as VVC Intra and GMM, in terms of computational cost, parameter count, and decoding latency. Additionally, ShiftLIC sets a new SOTA benchmark with a BD-rate gain per MACs/pixel of −102.6%, showcasing its potential for practical deployment in resource-constrained environments. The code is released athttps://github.com/baoyu2020/ShiftLIC. Youneng Bao, Wen Tan 0001, Chuanmin Jia, Mu Li 0005, Yongsheng Liang 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | One is All: A Unified Rate-Distortion-Complexity Framework for Learned Image Compression Under Energy Concentration CriteriaabstractThe learned image compression (LIC) technique has surpassed the state-of-the-art traditional codecs (H.266/VVC) in case of rate-distortion (R-D) performance. Its real-time deployments are far advanced. In order to achieve more flexible deployments, an LIC technique should be flexible in adjusting its computational complexity and rate as demanded by a situation and its environment. In this paper, we propose a unified Rate-Distortion-Complexity (R-D-C) framework for LIC under channel energy concentration criteria. Specifically, we first introduce an Energy Asymptotic Nonlinear Transformation (EANT) designed to directly concentrate on the channel energy of latent representations, thus laying the groundwork for a scalable entropy coding. Next, leveraging this energy concentration characteristic, we propose a corresponding Heterogeneous Scalable Entropy Model (HSEM) for flexibly scaling bitstreams as needed. Finally, utilizing the proposed EANT, we construct a fine-grained scalable codec for formulating, in combination with HSEM, a comprehensive scalable R-D-C framework under the energy concentration criteria. The obtained experimental results demonstrate that the proposed method could enable seamless transitions between 13 different widths of sub-models within a single network, allowing for fine-grained control over the model bitrate, complexity, and hardware inference time. Additionally, the proposed method exhibits competitive R-D performance compared to many existing methods. Chao Li 0071, Fanyang Meng, Qingyu Mao, Youneng Bao, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image CompressionabstractDeep neural network-based image compression (NIC) has achieved excellent performance, but NIC method models have been shown to be susceptible to backdoor attacks. Adversarial training has been validated in image compression models as a common method to enhance model robustness. However, the improvement effect of adversarial training on model robustness is limited. In this paper, we propose a prior knowledge-guided adversarial training framework for image compression models. Specifically, first, we propose a gradient regularization constraint for training robust teacher models. Subsequently, we design a knowledge distillation-based strategy to generate a priori knowledge from the teacher model to the student model for guiding adversarial training. Experimental results show that our method improves the reconstruction quality by about 9dB when the Kodak dataset is elected as the backdoor attack object for psnr attack. Compared with Ma2023 [1], our method has a 5dB higher PSNR output at high bitrate points. Youneng Bao, Fanyang Meng, Chao Li 0071, Wen Tan 0001, Genhong Wang, Yongsheng Liang 0001 |
ICASSP | 2 |
| 2024 | Leveraging Redundancy in Feature for Efficient Learned Image CompressionabstractIn recent years, with the development of the field of learned image compression, numerous models with excellent rate-distortion performance have emerged. However, the considerable computational complexity inherent in these models poses challenges for their practical deployment. In this paper, we investigate feature redundancy in learned image compression (LIC) algorithms for efficient feature extraction and introduce an efficient and lightweight LIC framework. Specifically, we explore the existence of a large number of similar features in the network. Subsequently, we design effective feature extraction modules across various levels, such as layer and block. In addition, based on the fact that the role of the codec’s encoder is to remove redundancy and the decoder is to reconstruct, we propose an asynchronous feature fusion block. This fusion block incorporates an "edge smoothing" operator in the encoder and an "edge enhancement" operator in the decoder. Our methodology strikes an ideal balance between rate-distortion performance and efficiency. The experimental results indicate that our approach necessitates only 310KMac/pixel computation and 9.5M parameters, while in terms of performance, our method achieves a 20.7% BD-rate advantage over BPG on Kodak data, mirroring VVC’s performance. Compared to other learned image compression algorithms with SOTA performance, our method has a great advantage in terms of computation/parameter count. Youneng Bao, Fanyang Meng, Wen Tan 0001, Chao Li 0071, Genhong Wang, Yongsheng Liang 0001 |
ICASSP | 2 |
| 2024 | Enhanced Interpretability in Learned Image Compression via Convolutional Sparse CodingabstractCompared to traditional image compression methods, learned image compression (LIC) methods have demonstrated increasingly superior rate-distortion performance. However, LIC networks are often regarded as black boxes, still lacking a theoretical understanding. Sparse coding provides the sparse and interpretable modeling for analyzing or synthesizing natural images in various signal and image processing applications. Therefore, we introduce convolutional sparse coding (CSC) into transform network for enhancing the interpretability of LIC methods. In this paper, we first employ CSC layers to achieve certain theoretical modeling for LIC network, and adopt a weight sharing strategy in encoder-decoder pair and attention mechanism to balance the complexity and performance. Additionally, we analyze the model robustness against data input perturbations and consider the impact of sparsity trade-off parameter in the CSC layer optimization process. Experimental results demonstrate that our method achieves comparable performance with the corresponding baseline, and our model is more robust. Yiwen Tu, Wen Tan 0001, Youneng Bao, Genhong Wang, Fanyang Meng, Yongsheng Liang 0001 |
ICME | 3 |
| 2024 | Fine-Grained Adjustable Entropy Models for Rate-Complexity Jointly Adjustable Image Compression
Chao Li 0071, Shanzhi Yin, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
PRCV (9) | 4 |
| 2024 | Learning Content-Weighted Pseudocylindrical Representation for 360° Image CompressionabstractLearned 360° image compression methods using equirectangular projection (ERP) often confront a non-uniform sampling issue, inherent to sphere-to-rectangle projection. While uniformly or nearly uniformly sampling representations, along with their corresponding convolution operations, have been proposed to mitigate this issue, these methods often concentrate solely on uniform sampling rates, thus neglecting the content of the image. In this paper, we urge that different contents within 360° images have varying significance and advocate for the adoption of a content-adaptive parametric representation in 360° image compression, which takes into account both the content and sampling rate. We first introduce the parametric pseudocylindrical representation and corresponding convolution operation, upon which we build a learned 360° image codec. Then, we model the hyperparameter of the representation as the output of a network, derived from the image's content and its spherical coordinates. We treat the optimization of hyperparameters for different 360° images as distinct compression tasks and propose a meta-learning algorithm to jointly optimize the codec and the metaknowledge, i.e., the hyperparameter estimation network. A significant challenge is the lack of a direct derivative from the compression loss to the hyperparameter network. To address this, we present a novel method to relax the rate-distortion loss as a function of the hyperparameters, enabling gradient-based optimization of the metaknowledge. Experimental results on omnidirectional images demonstrate that our method achieves state-of-the-art performance and superior visual quality. Mu Li 0005, Youneng Bao, Xiaohang Sui, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | A Complex-Valued Neural Network Based Robust Image Compression
Can Luo, Youneng Bao, Wen Tan 0001, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001 |
PRCV (10) | 2 |
| 2023 | Taylor series based dual-branch transformation for learned image compression
Youneng Bao, Wen Tan 0001, Linfeng Zheng, Fanyang Meng, Wei Liu 0065, Yongsheng Liang 0001 |
Signal Process. | 1 |
| 2023 | Nonlinear Transforms in Learned Image Compression From a Communication PerspectiveabstractRecently, remarkable progress has been made in learned image compression (LIC), in which nonlinear transforms (NTs) play a crucial role. Although there are many NT methods for improving the rate distortion performance, all the existing methods sacrifice the computational complexity and the number of parameters of the transformation. This paper provides a fundamental novel viewpoint on nonlinear transforms from a communication perspective, and shows how this idea can be extended to design efficient NT methods. In particular, the nonlinear transforms are inferred as signal modulation modules. Under this extrapolation, the current NTs are generalized as amplitude modulation that only varies the amplitude of the carrier wave. Therefore, a nonlinear modulation-like transform (NMLT) which varies the phase angle of the carrier is proposed. Moreover, this concept is extended by introducing In-phase/Quadrature (IQ) modulation, which is a boosting technique in communication field, in order to enhance NMLT. Furthermore, the Bit-interleaved technique in communication is used to guide the optimization of NTML with IQ. The experimental results on different datasets and backbone architectures verify the efficiency and robustness of the proposed methods. For example, when backbone architecture is hyperprior model, our method achieves 19.37% BD-rate reduction over GDN on the Kodak dataset. In addition, our method with channel wise autoregressive model leads to the state-of-the-art rate-distortion performance. Youneng Bao, Fanyang Meng, Chao Li 0071, Siwei Ma 0001, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | AdderIC: Towards Low Computation Cost Image CompressionabstractRecently, learned image compression methods have shown their outstanding rate-distortion performance when compared to traditional frameworks. Although numerous progress has been made in learned image compression, the computation cost is still at a high level. To address this problem, we propose AdderIC, which utilizes adder neural networks (AdderNet) to construct an image compression framework. According to the characteristics of image compression, we introduce several strategies to improve the performance of AdderNet in this field. Specifically, Haar Wavelet Transform is adopted to make AdderIC learn high-frequency information efficiently. In addition, implicit deconvolution with the kernel size of 1 is applied after each adder layer to reduce spatial redundancies. Moreover, we develop a novel Adder-ID-PixelShuffle cascade upsampling structure to remove checkerboard artifacts. Experiments demonstrate that our AdderIC model can largely outperform conventional AdderNet when applied in image compression and achieve comparable rate-distortion performance to that of its CNN baseline with about 80% multiplication FLOPs and 30% energy consumption reduction. Xin Yao 0001, Chao Li 0071, Youneng Bao, Fanyang Meng, Yongsheng Liang 0001 |
ICASSP | 4 |
| 2022 | Universal Efficient Variable-Rate Neural Image CompressionabstractRecently, Learning-based image compression has reached comparable performance with traditional image codecs(such as JPEG, BPG, WebP). However, computational complexity and rate flexibility are still two major challenges for its practical deployment. To tackle these problems, this paper proposes two universal modules named Energy-based Channel Gating(ECG) and Bit-rate Modulator(BM), which can be directly embedded into existing end-to-end image compression models. ECG uses dynamic pruning to reduce FLOPs for more than 50% in convolution layers, and a BM pair can modulate the latent representation to control the bit-rate in a channel-wise manner. By implementing these two modules, existing learning-based image codecs can obtain ability to output arbitrary bit-rate with a single model and reduced computation. Shanzhi Yin, Chao Li 0071, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng, Wei Liu 0065 |
ICASSP | 3 |
| 2022 | Exploring Structural Sparsity in Neural Image CompressionabstractThe performance of neural image compression have reached or suppressed traditional methods (such as JPEG, BPG, WebP). However, their sophisticated network structures with cascaded convolution layers bring heavy computational burden for practical deployment. In this paper, we explore structural sparsity in neural image compression network to obtain real-time acceleration without any specialized hardware design or algorithm. We propose a simple plug-in adaptive binary channel masking(ABCM) to judge the importance of each convolution channel and introduce sparsity during training. During inference, the unimportant channels are pruned to obtain slimmer network and less computation. We implement our method into three neural image compression networks with different entropy models to verify its effectiveness and generalization, the experiment results show that up to 7× computation reduction and 3× acceleration can be achieved with negligible performance drop. Shanzhi Yin, Chao Li 0071, Fanyang Meng, Wen Tan 0001, Youneng Bao, Yongsheng Liang 0001, Wei Liu 0065 |
ICIP | 5 |
| 2022 | Spatial-Temporal Asynchronous Normalization for Unsupervised 3D Action Representation LearningabstractUnsupervised 3D action representation learning from skeleton sequences has attracted increasing attention in recent years. Existing methods have successfully applied autoencoder network to learn 3D action representation by reconstructing original skeleton sequence. However, these methods ignore motion cues thus suffer from distinguishing actions especially with similar shape information and slightly different motion information. Instead of reconstructing original skeleton sequence, we learn distinctive 3D action representation with autoencoder network by reconstructing normalized motion sequence extracted from original input. To obtain the normalized motion sequence, we specifically design a novel spatial-temporal asynchronous normalization (STAN) method, which normalizes original skeleton sequence in two steps. First, STAN reduces redundant temporal information and extracts motion sequence by subtracting mean value along the temporal dimension. Second, STAN further normalizes the motion sequence along the spatial dimension and generates normalized motion sequence that suffers less from the effect of different human body shapes. Extensive experiments on large scale NTU RGB+D 60 and NTU RGB+D 120 datasets verify the effectiveness of our proposed STAN method, which achieves comparative results with state-of-the-art methods, and also outperforms alternative normalization methods. Mengyuan Liu 0001, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng |
IEEE Signal Process. Lett. | 2 |
| 2021 | Mbb: A Multi-Scale Method For Data Based On Bit Plane SlicingabstractMulti-scale methodology can enhance the performance of the model in deep learning. The current multi-scale methodology focuses on changing the formation, which will increase the parameters and calculations of the network. This paper offers a multi-scale method for data based on bit plane slicing(MBB). This expands the receptive field of valid information in image data. It is done by multi-level fusing image with high bit planes. Our experimentation shows that by adding MBB in front of the backbone network, one can achieve a significant performance improvement. The MBB approach is widely applicable because it does not require changes to the structure of the backbone network. Youneng Bao, Chao Li 0071, Fanyang Meng, Yongsheng Liang 0001, Wei Liu 0065, Kaiyu Liu |
ICIP | 1 |