Kai Wang 0070

dblp:78/2022-70 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-3132-0495ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Practical Lossless Volumetric Medical Image Compression via Tri-Plane Context Tree Learning
abstract
Lossless compression of volumetric medical images is of paramount importance for clinical and research applications where data fidelity is essential. Traditional compression methods are often limited in efficiency due to rigid, handcrafted models. Conversely, deep neural network (DNN)-based compression methods, while effective, demand substantial computational resources, hindering deployment in resource-constrained settings. To address these challenges, we propose a novel tri-plane context tree (TCT)-based method for lossless volumetric medical image compression that delivers high performance without relying on DNNs or external training data. To exploit intra-slice and inter-slice redundancies, we introduce a compact tri-plane context representation that decomposes complex 3D context modeling into efficient 2D modeling on three orthogonal planes. By integrating this representation with a context tree framework, we develop an input-specific TCT model employing an adaptive binary tree structure. At each tree node, the model dynamically selects from a suite of tri-plane based predictors and contextual feature extractors, enabling data-adaptive context modeling tailored to local structural characteristics. Instead of offline training, we sample a subset of the input volume to learn the TCT model by optimizing the minimum description length (MDL) through iterative construction and pruning. With the learned TCT model, each pixel retrieves its corresponding context, computes the prediction residual using the predictor dictated by the context, and performs entropy encoding based on the associated histograms. Experimental results demonstrate that the proposed method achieves compression performance on par with recent DNN-based methods on multiple datasets, while maintaining low computational cost and fast coding speeds, making it highly applicable in practice.
Yuanchao Bai, Kai Wang 0070, Yuanbo Du, Jie Chen 0001, Teng Fang, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Image Process.3
2026 3D-SLARM: Practical Lossless Volumetric Image Compression via a 3D-Scanning Lightweight Autoregressive Model
abstract
Volumetric images often encapsulate critical information, making it essential to employ lossless compression to preserve data integrity. Although various learned methods have demonstrated effective lossless compression for volumetric images, balancing high compression ratios with rapid coding speeds and lightweight architectures remains challenging. In this paper, we propose a 3D-scanning lightweight autoregressive model (3D-SLARM) for practical lossless volumetric image compression. 3D-SLARM integrates a novel 3D plane scanning module, a lightweight feature extraction (FE) module, and a lightweight distribution parameter and adaptive range predictor (DPARP) module. Initially, 3D-SLARM leverages a 3D plane scanning module to determine the scanning order of each voxel, allowing parallel coding of voxels within the same plane. Next, the lightweight FE module captures both intra-slice and inter-slice dependencies in the receptive field defined by the 3D plane scanning module. By incorporating our proposed serial re-parameterization (SerRep) technology alongside non-centric masked convolution (NCMC), the FE module attains a lightweight design while effectively capturing complex dependencies. Finally, 3D-SLARM employs a lightweight DPARP module to compute distribution parameters for both 8-bit and high bit-depth volumetric images. For high bit-depth images, the module further generates an adaptive probability range for each voxel, resulting in compact, voxel-specific PMF tables that facilitate efficient compression. Extensive experiments demonstrate that our 3D-SLARM achieves state-of-the-art lossless compression performance on majority volumetric image datasets and maintains fast coding speed with a lightweight design, underscoring its practical applicability.
Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005
IEEE Trans. Image Process.1
2025 CALLIC: Content Adaptive Learning for Lossless Image Compression
abstract
Learned lossless image compression has achieved significant advancements in recent years. However, existing methods often rely on training amortized generative models on massive datasets, resulting in sub-optimal probability distribution estimation for specific testing images during encoding process. To address this challenge, we explore the connection between the Minimum Description Length (MDL) principle and Parameter-Efficient Transfer Learning (PETL), leading to the development of a novel content-adaptive approach for learned lossless image compression, dubbed CALLIC. Specifically, we first propose a content-aware autoregressive self-attention mechanism by leveraging convolutional gating operations, termed Masked Gated ConvFormer (MGCF), and pretrain MGCF on training dataset. Cache then Crop Inference (CCI) is proposed to accelerate the coding process. During encoding, we decompose pretrained layers, including depth-wise convolutions, using low-rank matrices and then adapt the incremental weights on testing image by Rate-guided Progressive Fine-Tuning (RPFT). RPFT fine-tunes with gradually increasing patches that are sorted in descending order by estimated entropy, optimizing learning process and reducing adaptation time. Extensive experiments across diverse datasets demonstrate that CALLIC sets a new state-of-the-art (SOTA) for learned lossless image compression.
Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005, Wen Gao 0001
AAAI3
2025 Learning Lossless Compression for High Bit-Depth Volumetric Medical Image
abstract
Recent advances in learning-based methods have markedly enhanced the capabilities of image compression. However, these methods struggle with high bit-depth volumetric medical images, facing issues such as degraded performance, increased memory demand, and reduced processing speed. To address these challenges, this paper presents the Bit-Division based Lossless Volumetric Image Compression (BD-LVIC) framework, which is tailored for high bit-depth medical volume compression. The BD-LVIC framework skillfully divides the high bit-depth volume into two lower bit-depth segments: the Most Significant Bit-Volume (MSBV) and the Least Significant Bit-Volume (LSBV). The MSBV concentrates on the most significant bits of the volumetric medical image, capturing vital structural details in a compact manner. This reduction in complexity greatly improves compression efficiency using traditional codecs. Conversely, the LSBV deals with the least significant bits, which encapsulate intricate texture details. To compress this detailed information effectively, we introduce an effective learning-based compression model equipped with a Transformer-Based Feature Alignment Module, which exploits both intra-slice and inter-slice redundancies to accurately align features. Subsequently, a Parallel Autoregressive Coding Module merges these features to precisely estimate the probability distribution of the least significant bit-planes. Our extensive testing demonstrates that the BD-LVIC framework not only sets new performance benchmarks across various datasets but also maintains a competitive coding speed, highlighting its significant potential and practical utility in the realm of volumetric medical image compression.
Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005
IEEE Trans. Image Process.1
2024 Semantic Ensemble Loss and Latent Refinement for High-Fidelity Neural Image Compression
abstract
Recent advancements in neural compression have surpassed traditional codecs in PSNR and MS-SSIM measurements. However, at low bit-rates, these methods can introduce visually displeasing artifacts, such as blurring, color shifting, and texture loss, thereby compromising perceptual quality of images. To address these issues, this study presents an enhanced neural compression method designed for optimal visual fidelity. We have trained our model with a sophisticated semantic ensemble loss, integrating Charbonnier loss, perceptual loss, style loss, and a non-binary adversarial loss, to enhance the perceptual quality of image reconstructions. Additionally, we have implemented a latent refinement process to generate content-aware latent codes. These codes adhere to bit-rate constraints, and prioritize bit allocation to regions of greater importance. Our empirical findings demonstrate that this approach significantly improves the statistical fidelity of neural image compression.
Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005
VCIP3
2024 Deep Lossy Plus Residual Coding for Lossless and Near-Lossless Image Compression
abstract
Lossless and near-lossless image compression is of paramount importance to professional users in many technical fields, such as medicine, remote sensing, precision engineering and scientific research. But despite rapidly growing research interests in learning-based image compression, no published method offers both lossless and near-lossless modes. In this paper, we propose a unified and powerful deep lossy plus residual (DLPR) coding framework for both lossless and near-lossless image compression. In the lossless mode, the DLPR coding system first performs lossy compression and then lossless coding of residuals. We solve the joint lossy and residual compression problem in the approach of VAEs, and add autoregressive context modeling of the residuals to enhance lossless compression performance. In the near-lossless mode, we quantize the original residuals to satisfy a given ℓ∞error bound, and propose a scalable near-lossless compression scheme that works for variable ℓ∞bounds instead of training multiple networks. To expedite the DLPR coding, we increase the degree of algorithm parallelization by a novel design of coding context, and accelerate the entropy coding with adaptive residual interval. Experimental results demonstrate that the DLPR coding system achieves both the state-of-the-art lossless and near-lossless image compression performance with competitive coding speed.
Yuanchao Bai, Xianming Liu 0005, Kai Wang 0070, Xiangyang Ji, Xiaolin Wu 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 GroupedMixer: An Entropy Model With Group-Wise Token-Mixers for Learned Image Compression
abstract
Transformer-based entropy models have gained prominence in recent years due to their superior ability to capture long-range dependencies in probability distribution estimation compared to convolution-based methods. However, previous transformer-based entropy models suffer from a sluggish coding process due to pixel-wise autoregression or duplicated computation during inference. In this paper, we propose a novel transformer-based entropy model called GroupedMixer, which enjoys both faster coding speed and better compression performance than previous transformer-based methods. Specifically, our approach builds upon group-wise autoregression by first partitioning the latent variables into groups along spatial-channel dimensions, and then entropy coding the groups with the proposed transformer-based entropy model. The global causal self-attention is decomposed into more efficient group-wise interactions, implemented using inner-group and cross-group token-mixers. The inner-group token-mixer incorporates contextual elements within a group while the cross-group token-mixer interacts with previously decoded groups. Alternate arrangement of two token-mixers enables global contextual reference. To further expedite the network inference, we introduce context cache optimization to GroupedMixer, which caches attention activation values in cross-group token-mixers and avoids complex and duplicated computation. Experimental results demonstrate that the proposed GroupedMixer yields the state-of-the-art rate-distortion performance with fast compression speed.
Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Learning Lossless Compression for High Bit-Depth Medical Imaging
abstract
We propose a learned lossless image compression method for high bit-depth medical imaging (up to 16 bit-depths). Instead of compressing a high bit-depth medical image as a whole, we split it into two low bit-depth subimages, i.e., the most significant bytes (MSB) subimage and the least significant bytes (LSB) subimage, respectively. The MSB subimage depicts piece-wise smooth structure information that is relatively easy to compress. We thus use traditional lossless codecs for low complexity. The LSB subimage depicts the complementary texture information that is more challenging to compress. We design an autoregressive entropy model conditioned on the MSB subimage that models the probability distribution of the LSB subimage and effectively reduces the redundancy between the MSB and LSB subimages. We then encode the LSB subimage to bitstreams based on the learned entropy model. The compressed high bit-depth medical image is finally stored including the bitstreams of the MSB and LSB subimages. Experimental results demonstrate the state-of-the-art compression performance of the proposed method on high bit-depth medical images, compared with both existing traditional and learned lossless image codecs.
Kai Wang 0070, Yuanchao Bai, Deming Zhai, Daxin Li, Junjun Jiang, Xianming Liu 0005
ICME1