Gai Zhang

dblp:339/2491 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-0418-0689ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 UAR-NVC: A Unified Autoregressive Framework for Memory-Efficient Neural Video Compression
abstract
Implicit Neural Representations (INRs) have demonstrated significant potential in video compression by representing videos as neural networks. However, as the number of frames increases, the memory consumption for training and inference increases substantially, posing challenges in resource-constrained scenarios. Inspired by the success of traditional video compression frameworks, which process video frame by frame and can efficiently compress long videos, we adopt this modeling strategy for INRs to decrease memory consumption, while aiming to unify the frameworks from the perspective of timeline-based autoregressive modeling. In this work, we present a novel understanding of INR models from an autoregressive (AR) perspective and introduce a Unified AutoRegressive Framework for memory-efficient Neural Video Compression (UAR-NVC). UAR-NVC integrates timeline-based and INR-based neural video compression under a unified autoregressive paradigm. It partitions videos into several clips and processes each clip using a different INR model instance, leveraging the advantages of both compression frameworks while allowing seamless adaptation to either in form. To further reduce temporal redundancy between clips, we treat the corresponding model parameters as proxies for these clips, and design two modules to optimize the initialization, training, and compression of these model parameters. In special, the Residual Quantization and Entropy Constraint (RQEC) module dynamically balances the reconstruction quality of the current clip and the newly introduced bitrate cost using the previously optimized parameters as conditioning. In addition, the Interpolation-based Initialization (II) module flexibly adjusts the degree of reference used during the initialization of neighboring video clips, based on their correlation. UAR-NVC supports adjustable latencies by varying the clip length. Extensive experimental results demonstrate that UAR-NVC, with its flexible video clip setting, can adapt to resource-constrained environments and significantly improve performance compared to different baseline models. The project page: https://wj-inf.github.io/UAR-NVC-page/.
Jia Wang 0018, Xinfeng Zhang 0001, Gai Zhang, Lv Tang, Li Zhang 0006
IEEE Trans. Circuits Syst. Video Technol.3
2025 2D Gaussian Splatting With Pre-Trained Dictionaries for Image Compression
Yifan Gao 0009, Xinfeng Zhang 0001, Gai Zhang
IEEE Signal Process. Lett.3
2024 VQNeRV: Vector Quantization Neural Representation for Video Compression
abstract
The application of Implicit Neural Representations (INR) for video compression represents an evolving area of research. Despite its potential, conventional CNN-based INR methods often encounter difficulties in modeling complex spatiotemporal scenes. This is primarily due to the inherent limitations of CNNs in extracting intricate information, particularly when it comes to capturing temporal dynamics. Consequently, the integration of spatiotemporal contextual information becomes imperative to enhance INR's capacity for scene modeling. In response to these challenges, this paper introduces a novel approach, termed Vector Quantization Neural Representation (VQNeRV), specifically designed for video compression. Our methodology unfolds in three distinct stages: Firstly, we establish a CNN-based INR equipped with spatial embeddings to model the video content. Secondly, a Vector Quantization (VQ) encoder is then utilized to distill spatiotemporal features from the video. These features are subsequently amalgamated with the spatial embeddings through cross-attention mechanisms, facilitating a comprehensive feature fusion. Finally, an INR decoder, leveraging the combined features embodying spatiotemporal contextual information, reconstructs the video. Experimental results demonstrate that our method shows a much better rate-distortion performance compared to the VVC.
Gai Zhang, Lv Tang, Xinfeng Zhang 0001
ISCAS1
2024 Pose-Driven Compression for Dynamic 3D Human via Human Prior Models
abstract
To cost-effectively transmit high-quality dynamic 3D human images in immersive multimedia applications, efficient data compression is crucial. Unlike existing methods that focus on reducing signal-level reconstruction errors, we propose the first dynamic 3D human compression framework based on human priors. The layered coding architecture significantly enhances the perceptual quality while also supporting a variety of downstream tasks, including visual analysis and content editing. Specifically, a high-fidelity pose-driven Avatar is generated from the original frames as the basic structure layer to implicitly represent the human shape. Then, human movements between frames are parameterized via a commonly-used human prior model, i.e., the Skinned Multi-Person Linear Model (SMPL), to form the motion layer and drive the Avatar. Furthermore, the normals are also introduced as an enhancement layer to preserve fine-grained geometric details. Finally, the Avatar, SMPL parameters, and normal maps are efficiently compressed into layered semantic bitstreams. Extensive qualitative and quantitative experiments show that the proposed framework remarkably outperforms other state-of-the-art 3D codecs in terms of subjective quality with only a few bits. More notably, as the size or frame number of the 3D human sequence increases, the superiority of our framework in perceptual quality becomes more significant while saving more bitrates.
Ruoke Yan, Qian Yin 0002, Xinfeng Zhang 0001, Qi Zhang 0042, Gai Zhang, Siwei Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Unified and Scalable Deep Image Compression Framework for Human and Machine
abstract
Image compression aims to minimize the amount of data in image representation while maintaining a certain visual quality for humans, which is an essential technique for storage and transmission. Recently, along with the development of computer vision, machines have become another primary receiver for images and require compressed images at a certain quality level, which may be different from that of human vision. In many scenarios, compressed images should serve both human and machine vision tasks, but few compression methods are designed for both goals simultaneously. In this article, we propose a unified and scalable deep image compression (USDIC) framework that jointly optimizes the image quality according to human and machine vision in an end-to-end style. For the encoder, we propose an information splitting mechanism (ISM) to separate images into semantic and visual features, which mainly aims at machine analysis and human viewing tasks. For the decoder, we design a scalable decoding architecture. The encoded semantic feature is first decoded for machine analysis tasks, and the image is decoded and reconstructed further by leveraging the decoded semantic features. Herein, to further remove the redundancy between the semantic and visual features of images, we propose a scalable entropy model (SEM) with a joint optimization strategy to reconstruct the image using the two kinds of decoded features. Extensive experimental results show that the proposed USDIC achieves much better performance on the image analysis task while maintaining competitive performance on the traditional image reconstruction task compared with popular image compression methods.
Gai Zhang, Xinfeng Zhang 0001, Lv Tang
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Scene Matters: Model-based Deep Video Compression
abstract
Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal prediction theory to enhance compression performance by designing high efficient intra and inter prediction strategies and compressing video frames one by one. In this paper, we propose a novel model-based video compression (MVC) framework that regards scenes as the fundamental units for video sequences. Our proposed MVC directly models the intensity variation of the entire video sequence in one scene, seeking non-redundant representations instead of reducing redundancy through spatio-temporal predictions. To achieve this, we employ implicit neural representation as our basic modeling architecture. To improve the efficiency of video modeling, we first propose context-related spatial positional embedding and frequency domain supervision in spatial context enhancement. For temporal correlation capturing, we design the scene flow constrain mechanism and temporal contrastive loss. Extensive experimental results demonstrate that our method achieves up to a 20% bitrate reduction compared to the latest video coding standard H.266 and is more efficient in decoding than existing video coding strategies.
Lv Tang, Xinfeng Zhang 0001, Gai Zhang, Xiaoqi Ma
ICCV3
2023 Enhanced Quantified Local Implicit Neural Representation for Image Compression
abstract
Recently, implicit neural representation (INR) has been applied to image compression. However, the rate-distortion performance of most existing INR-based image compression methods is still obviously inferior to the state-of-the-art image compression methods. In this paper, we propose an Enhanced Quantified Local Implicit Neural Representation (EQLINR) for image compression by enhancing the utilization of local relationships of INR and narrow the quantization gap between training and encoding to further improve the performance of INR-based image compression. Our framework consists of latent representation and the corresponding implicit neural network consisting of MLP and CNN, which can transform the latent representation into the image space. To enhance local relationships utilization, we design a local enhancement module (LEM) consisted of CNN to capture the neighborhood relationships of the reconstructed image from MLP. Furthermore, to mitigate the performance loss caused by quantization of latent representation, we employ an enhanced quantization scheme (EQS) in our training process. We use uniform noise for network initialization and then use Stochastic Gumbel Annealing (SGA) with dynamic temperature regulation as a proxy function for quantization during training. Extensive experimental results demonstrate that our approach significantly the compression performance of INR-based image compression, and even better than BPG
Gai Zhang, Xinfeng Zhang 0001, Lv Tang
IEEE Signal Process. Lett.1
2022 Local and Global Fusion Network For Learned Image Compression
abstract
Convolution autoencoder is a widely utilized framework for learned image compression. However, it is facing performance bottleneck because the convolution with limited receptive field mainly extracts image local information. In this paper, we propose a novel neural network based image compression framework by removing both local and global redundancy. Herein, convolution autoencoder and Generative flow (Glow) are utilized to extract image local and global information respectively. Glow is a lossless invertible neural network and can facilitate global information extraction. Furthermore, we design a DenseNet module to fuse the local and global information extracted from convolution autoencoder and Glow. Extensive experimental results show that the proposed framework outperform the intra-frame coding of Versatile Video Coding (VVC) and state-of-the-art neural network based image compression methods.
Gai Zhang, Xinfeng Zhang 0001, Shuyuan Zhu
ICIP1