Guo Lu

dblp:76/7805 · DBLP profile ↗
← Back
68ranked-venue papers
10as first author
59since 2021 · last 2026
0000-0001-6951-0090ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 9 first-author · 42 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 26 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LMM-VSC: Ultra-Low Bitrate Video Compression with Semantic Understanding
Chaolei Liu, Li Song 0001, Chuqin Zhou, Guo Lu
ISCAS4
2026 Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
Wenwei Jin, Jintao Tong, Pengda Qin, Guo Lu
KDD (1)6
2026 SMC++: Masked Learning of Unsupervised Video Semantic Compression
abstract
Most video compression methods focus on human visual perception, neglecting semantic preservation. This leads to severe semantic loss during the compression, hampering downstream video analysis tasks. In this paper, we propose a Masked Video Modeling (MVM)-powered compression framework that particularly preserves video semantics, by jointly mining and compressing the semantics in a self-supervised manner. While MVM is proficient at learning generalizable semantics through the masked patch prediction task, it may also encode non-semantic information like trivial textural details, wasting bitcost and bringing semantic noises. To suppress this, we explicitly regularize the non-semantic entropy of the compressed video in the MVM token space. The proposed framework is instantiated as a simple Semantic-Mining-then-Compression (SMC) model. Furthermore, we extend SMC as an advanced SMC++ model from several aspects. First, we equip it with a masked motion prediction objective, leading to better temporal semantic learning ability. Second, we introduce a Transformer-based compression module, to improve the semantic compression efficacy. Considering that directly mining the complex redundancy among heterogeneous features in different coding stages is non-trivial, we introduce a compact blueprint semantic representation to align these features into a similar form, fully unleashing the power of the Transformer-based compression module. Extensive results demonstrate the proposed SMC and SMC++ models show remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets.
Yuan Tian 0017, Xiaoyue Ling, Cong Geng, Qiang Hu 0003, Guo Lu, Guangtao Zhai
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression
abstract
Learning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity design while maintaining compression performance. We introduce a novel Learned Lossless Low-complexity Text Compression method (L3TC). Specifically, we conduct extensive experiments demonstrating that RWKV models achieve the fastest decoding speed with a moderate compression ratio, making it the most suitable backbone for our method. Second, we propose an outlier-aware tokenizer that uses a limited vocabulary to cover frequent tokens while allowing outliers to bypass the prediction and encoding. Third, we propose a novel high-rank reparameterization strategy that enhances the learning capability during training without increasing complexity during inference. Experimental results validate that our method achieves 48% bit saving compared to gzip compressor. Besides, L3TC offers compression performance comparable to other learned compressors, with a 50x reduction in model parameters. More importantly, L3TC is the fastest among all learned compressors, providing real-time decoding speeds up to megabytes per second.
Junxuan Zhang, Zhengxue Cheng, Yan Zhao 0041, Dajiang Zhou, Guo Lu, Li Song 0001
AAAI6
2025 Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression
abstract
Neural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects for a fixed neural image codec. Specifically, we introduce a plug-and-play module at the decoder side that leverages a latent diffusion process to transform the decoded features, enhancing either low distortion or high perceptual quality without altering the original image compression codec. Our approach facilitates fusion of original and transformed features without additional training, enabling users to flexibly adjust the balance between distortion and perception during inference. Extensive experimental results demonstrate that our method significantly enhances the pretrained codecs with a wide, adjustable distortion-perception range while maintaining their original compression capabilities. For instance, we can achieve more than 150% improvement in LPIPS-BDRate without sacrificing more than 1 dB in PSNR.
Chuqin Zhou, Guo Lu, Jiangchuan Li, Zhengxue Cheng, Li Song 0001, Wenjun Zhang 0001
AAAI2
2025 Linear Attention Modeling for Learned Image Compression
abstract
Recent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKV-based Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https://github.com/sjtu-medialab/RwkvCompress.
Donghui Feng 0003, Zhengxue Cheng, Shen Wang 0013, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song 0001
CVPR6
2025 Image Quality Assessment: From Human to Machine Preference
abstract
Image Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference depends on downstream tasks such as segmentation and detection, rather than visual appeal. Considering the huge gap between human and machine visual systems, this paper proposes the topic: Image Quality Assessment for Machine Vision for the first time. Specifically, we (1) defined the subjective preferences of machines, including downstream tasks, test models, and evaluation metrics; (2) established the Machine Preference Database (MPD), which contains 2.25M fine-grained annotations and 30k reference/distorted image pair instances; (3) verified the performance of mainstream IQA algorithms on MPD. Experiments show that current IQA metrics are human-centric and cannot accurately characterize machine preferences. We sincerely hope that MPD can promote the evolution of IQA from human to machine preferences. Project page is on: https://github.com/lcysyzxdxc/MPD.
Chunyi Li 0001, Yuan Tian 0017, Xiaoyue Ling, Haodong Duan, Haoning Wu 0001, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Guo Lu, Weisi Lin, Guangtao Zhai
CVPR10
2025 Multi-Style Facial Sketch Synthesis through Masked Generative Modeling
abstract
The facial sketch synthesis (FSS) model, capable of generating sketch portraits from given facial photographs, holds profound implications across multiple domains, encompassing cross-modal face recognition, entertainment, art, media, among others. However, the production of high-quality sketches remains a formidable task, primarily due to the challenges and flaws associated with three key factors: (1) the scarcity of artist-drawn data, (2) the constraints imposed by limited style types, and (3) the deficiencies of processing input information in existing models. To address these difficulties, we propose a lightweight end-to-end synthesis model that efficiently converts images to corresponding multi-stylized sketches, obviating the necessity for any supplementary inputs (e.g., 3D geometry). In this study, we overcome the issue of data insufficiency by incorporating semi-supervised learning into the training process. Additionally, we employ a feature extraction module and style embeddings to proficiently steer the generative transformer during the iterative prediction of masked image tokens, thus achieving a continuous stylized output that retains facial features accurately in sketches. The extensive experiments demonstrate that our method consistently outperforms previous algorithms across multiple benchmarks, exhibiting a discernible disparity.
Guo Lu, Shibao Zheng
ICASSP2
2025 Knowledge Distillation for Learned Image Compression
Yunuo Chen 0002, Zezheng Lyu, Guo Lu, Wenjun Zhang 0001
ICCV6
2025 QaVA: Query-Aware Video Analysis Framework Based on Data Access Pattern
abstract
With the explosive growth of video data, efficient video analysis technology has garnered widespread attention. Existing online methods train proxy neural networks upon query arrival and use these networks to scan the entire dataset, guiding the invocation of the expensive deep neural network. While index-based methods advance this process to the index-building stage, significantly reducing the time overhead of video queries. However, the data to query often presents a long-tail distribution, and different types of queries are sensitive to different parts of the distribution. Since the index-based methods cannot predict the queries, they can only provide ad-hoc proxy score generating strategies. This paper proposes a query-aware video analysis framework, QaVA, to improve query performance further. QaVA retains the time-consuming, query-independent semantic extraction process during the index-building stage and employs a tunable lightweight adapter network to accurately and quickly focus on the data parts most relevant to the query after it arrives. Meanwhile, QaVA can automatically tune the training strategy of the adapter network by analyzing the data access pattern of historical queries, thus meeting the needs of general users. Experimental results demonstrate that QaVA can significantly reduce the cost of various queries across multiple datasets, and can speed up query processing by up to$9.2\times$compared to the most advanced index-based method. Our code is available: https://github.com/InkosiZhong/QaVA.
Tianxiong Zhong, Zhiwei Zhang 0002, Yihang Fu, Guo Lu, Ye Yuan 0001, Guoren Wang
ICDE4
2025 Differentiable VMAF: A trainable metric for optimizing video compression codec
abstract
Video Multi-method Assessment Fusion (VMAF) is a widely used objective evaluation metric that has shown a stronger correlation with human visual system than other metrics. Some studies have already integrated VMAF as a measure of perceptual quality in the rate-distortion optimization of hybrid codecs. However, VMAF has long been considered non-differentiable, limiting its application in learning-based video compression. In this work, we successfully achieve a differentiable VMAF through mathematical approximations. Our implementation of differentiable VMAF maintains the performance of the original VMAF while ensuring accurate gradient calculations. By integrating it into the loss function for optimizing learning-based video compression, our experimental results demonstrate that this approach significantly enhances VMAF performance and improves perceptual quality at all bitrates.
Jiangchuan Li, Chuqin Zhou, Yunuo Chen 0002, Guo Lu
ISCAS4
2025 Towards a New Paradigm of Visual Signal Compression
abstract
Ultra-low bitrate image compression is a challenging and demand- ing topic. With the development of Large Multimodal Models (LMMs), a Cross Modality Compression (CMC) paradigm of Image-Text- Image has emerged. Compared with traditional codecs, this semantic- level compression can reduce image data size to 0.1% or even lower, which has strong potential applications. However, CMC has cer- tain defects in consistency with the original image and perceptual quality. To inspire insights into such a problem, we introduce CMC- Bench, a benchmark of the cooperative performance of Image-to- Text (I2T) and Text-to-Image (T2I) models for image compression. This benchmark covers 18,000 and 40,000 images respectively to verify 6 mainstream I2T and 12 T2I models, including 160,000 sub- jective preference scores annotated by human experts. At ultra-low bitrates, it proves that the combination of some I2T and T2I models has surpassed the most advanced visual signal codecs; meanwhile, it highlights where LMMs can be further optimized toward the compression task. We encourage LMM developers to participate in this test to promote the evolution of visual signal codec protocols.
Chunyi Li 0001, Xiele Wu, Haoning Wu 0001, Donghui Feng 0003, Guo Lu, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin
ACM Multimedia6
2025 H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting
abstract
Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity. This approach decomposes a dynamic scene into a static representation in a canonical space and time-varying scene motion. Scene motion is defined as the collective movement of all Gaussian points, and for compactness, existing approaches commonly adopt implicit neural fields or sparse control points. However, these methods predominantly rely on gradient-based optimization for all motion information. Due to the high degree of freedom, they struggle to converge on real-world datasets exhibiting complex motion. To preserve the compactness of motion representation and address convergence challenges, this paper proposes heterogeneous 3D control points, termed \textbf{H3D control points}, whose attributes are obtained using a hybrid strategy combining optical flow back-projection and gradient-based methods. This design decouples directly observable motion components from those that are geometrically occluded. Specifically, components of 3D motion that project onto the image plane are directly acquired via optical flow back projection, while unobservable portions are refined through gradient-based optimization. Experiments on the Neu3DV and CMU-Panoptic datasets demonstrate that our method achieves superior performance over state-of-the-art deformable 3D Gaussian splatting techniques. Remarkably, our method converges within just 100 iterations and achieves a per-frame processing speed of 2 seconds on a single NVIDIA RTX 4070 GPU.
Yunuo Chen 0002, Guo Lu, Cheems Wang, Qunshan Gu, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
NeurIPS3
2025 A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001
PCS6
2025 Neural Hamiltonian Deformation Fields for Dynamic Scene Rendering
abstract
Representing and rendering dynamic scenes with complex motions remains challenging in computer vision and graphics. Recent dynamic view synthesis methods achieve high-quality rendering but often produce physically implausible motions. We introduce NeHaD, a neural deformation field for dynamic Gaussian Splatting governed by Hamiltonian mechanics. Our key observation is that existing methods using MLPs to predict deformation fields introduce inevitable biases, resulting in unnatural dynamics. By incorporating physics priors, we achieve robust and realistic dynamic scene rendering. Hamiltonian mechanics provides an ideal framework for modeling Gaussian deformation fields due to their shared phase-space structure, where primitives evolve along energy-conserving trajectories. We employ Hamiltonian neural networks to implicitly learn underlying physical laws governing deformation. Meanwhile, we introduce Boltzmann equilibrium decomposition, an energy-aware mechanism that adaptively separates static and dynamic Gaussians based on their spatial-temporal energy states for flexible rendering. To handle real-world dissipation, we employ second-order symplectic integration and local rigidity regularization as physics-informed constraints for robust dynamics modeling. Additionally, we extend NeHaD to adaptive streaming through scale-aware mipmapping and progressive optimization. Extensive experiments demonstrate that NeHaD achieves physically plausible results with a rendering quality-efficiency trade-off. To our knowledge, this is the first exploration leveraging Hamiltonian mechanics for neural Gaussian deformation, enabling physically realistic dynamic scene rendering with streaming capabilities.
Hai-Long Qin, Sixian Wang, Guo Lu, Jincheng Dai
SIGGRAPH Asia3
2025 VCIP 2025 Ultra Low-Bitrate Video Compression Challenge
abstract
This report presents the VCIP 2025 Grand Challenge on Ultra Low-Bitrate Video Compression, which aims to promote research progress in perceptually optimized and computationally efficient video compression under extreme bandwidth constraints. The challenge focuses on scenarios such as emergency communication, remote monitoring, and low-power transmission, where conventional codecs like HEVC and AV1 struggle to maintain acceptable perceptual quality. Two benchmark tracks are introduced to assess the trade-off between compression ratio, visual quality, and complexity: Track 1 (50 kbps) targets extreme low-bitrate conditions, while Track 2 (200 kbps) allows moderately higher bitrates under real-time constraints. Participants were required to satisfy strict limits on encoding and decoding complexity, evaluated in kMac per pixel and per-frame runtime. The evaluation combined both objective metrics (PSNR, SSIM, VMAF, LPIPS) and subjective human perception scoring to comprehensively assess quality and efficiency. The results reveal two complementary trends for the future of video compression: (1) efficiency-oriented codec architecture design enabling low-latency deployment on edge devices, and (2) perceptual enhancement through post-decoding restoration guided by temporal and semantic priors. Together, these directions signal a paradigm shift from traditional rate–distortion optimization toward a broader rate–perception–complexity trade-off. The insights gained from this challenge are expected to inspire future standards and generative compression models for perceptually-driven, adaptive, and bandwidth-efficient video communication.
Guo Lu, Jing Wang 0194, Yunuo Chen 0002, Chuqin Zhou, Yibo Shi
VCIP1
2025 VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming With Edge Computing
abstract
Free-view video (FVV) allows users to explore immersive video content from multiple views. However, delivering FVV poses significant challenges due to the uncertainty in view switching, combined with the substantial bandwidth and computational resources required to transmit and decode multiple video streams, which may result in frequent playback interruptions. Existing approaches, either client-based or cloud-based, struggle to meet high Quality of Experience (QoE) requirements under limited bandwidth and computational resources. To address these issues, we propose VARFVV, a bandwidth- and computationally-efficient system that enables real-time interactive FVV streaming with high QoE and low switching delay. Specifically, VARFVV introduces a low-complexity FVV generation scheme that reassembles multiview video frames at the edge server based on user-selected view tracks, eliminating the need for transcoding and significantly reducing computational overhead. This design makes it well-suited for large-scale, mobile-based UHD FVV experiences. Furthermore, we present a popularity-adaptive bit allocation method, leveraging a graph neural network, that predicts view popularity and dynamically adjusts bit allocation to maximize QoE within bandwidth constraints. We also construct an FVV dataset comprising 330 videos from 10 scenes, including basketball, opera, etc. Extensive experiments show that VARFVV surpasses existing methods in video quality, switching latency, computational efficiency, and bandwidth usage, supporting over 500 users on a single edge server with a switching delay of 71.5ms. Our code and dataset are available at https://github.com/qianghu-huber/VARFVV.
Qiang Hu 0003, Qihan He, Houqiang Zhong, Guo Lu, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001
IEEE J. Sel. Areas Commun.4
2025 MISC: Ultra-Low Bitrate Image Semantic Compression Driven by Large Multimodal Model
abstract
With the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, all existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. During recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC.
Chunyi Li 0001, Guo Lu, Donghui Feng 0003, Haoning Wu 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin, Wenjun Zhang 0001
IEEE Trans. Image Process.2
2025 Instance-Adaptive Spatial-Temporal Enhancement for Efficient Video Compression
abstract
Efficiently compressing HD/UHD content has long been challenging due to high bitrate costs. Instance-adaptive enhancement methods try to tackle this issue by compressing a video at reduced resolution and enhancing it using a neural model specifically overfitted for this video. However, existing methods focus solely on spatial super-resolution (SR) and under-utilize the videos' temporal redundancy. Their limited management of the model's updated parameters also causes excessive overfitting overheads. Therefore, this paper introduces IASTE, the first instance-adaptive enhancement method based on spatial-temporal enhancement (STE), and incorporates low-rank adaptation (LoRA) for efficient model overfitting. Specifically, we downscale videos spatially and temporally to reduce the data volume and achieve efficient video compression. Then, we overfit a specific STE model for each video and use it to enhance the decoded video's spatiotemporal resolution. Leveraging the video swin transformer's strong capability in capturing spatiotemporal correlations, we design a lightweight and efficient model to implement video STE. The model is overfitted for each video using LoRA. By freezing the pre-trained model and selectively updating a few low-rank matrices, the bitrate overhead for model storage can be mitigated. Experiments prove that compared to directly compressing high-frame-rate (HFR), high-resolution (HR) videos, our method achieves around 30% BD-Rate gains on the CTC and UVG datasets, about 15% gains on the YoutubeUGC dataset, and about 10% gains on the ultra-long videos in the Xiph dataset.
Yan Zhao 0041, Zhengxue Cheng, Jiangchuan Li, Donghui Feng 0003, Qunshan Gu, Cheems Wang, Guo Lu, Li Song 0001
IEEE Trans. Image Process.7
2025 Implicit-Explicit Integrated Representations for Multi-View Video Compression
abstract
With the increasing consumption of 3D displays and virtual reality, multi-view video has become a promising format. However, its high resolution and multi-camera shooting result in a substantial increase in data volume, making storage and transmission a challenging task. To tackle these difficulties, we propose an implicit-explicit integrated representation for multi-view video compression. Specifically, we first use the explicit representation-based 2D video codec to encode one of the source views. Subsequently, we propose employing the implicit neural representation (INR)-based codec to encode the remaining views. The implicit codec takes the time and view index of multi-view video as coordinate input and generates the corresponding implicit reconstruction frames. To enhance the compressibility, we introduce a multi-level feature grid embedding and a fully convolutional architecture into the implicit codec. These components facilitate coordinate-feature and feature-RGB mapping, respectively. To further enhance the reconstruction quality from the INR codec, we leverage the high-quality reconstructed frames from the explicit codec to achieve inter-view compensation. Finally, the compensated results are fused with the implicit reconstructions from the INR to obtain the final reconstructed frames. Our proposed framework combines the strengths of both implicit neural representation and explicit 2D codec. Extensive experiments conducted on public datasets demonstrate that the proposed framework can achieve comparable or even superior performance to the latest multi-view video compression standard MIV and other INR-based schemes in terms of view compression and scene modeling. The source code can be found at https://github.com/zc-lynen/MV-IERV.
Guo Lu, Rong Xie 0004, Li Song 0001
IEEE Trans. Image Process.2
2025 DiFace: Cross-Modal Face Recognition through Controlled Diffusion
abstract
Diffusion probabilistic models (DPMs) have exhibited exceptional proficiency in generating visual media of outstanding quality and realism. Nonetheless, their potential in non-generative domains, such as face recognition (FR), has yet to be thoroughly investigated. Meanwhile, despite the extensive development of multi-modal FR methods, their emphasis has predominantly centered on visual modalities. In this context, FR through textual description presents a unique and promising solution that not only transcends the limitations from application scenarios but also expands the potential for research in the field of cross-modal FR. It is regrettable that this avenue remains underutilized, a consequence from the challenges mainly associated with three aspects: 1) the intrinsic imprecision of verbal descriptions; 2) the significant gaps between texts and images; and 3) the immense hurdle posed by insufficient databases. To tackle this problem, we present DiFace, an end-to-end solution that effectively achieves FR via text through a controllable diffusion process, by establishing its theoretical connection with probability transport. Our approach not only unleashes the potential of DPMs across a wide range of tasks but also showcases a remarkable improvement in accuracy for text-image FR, as evidenced by our experiments on verification and identification.
Guo Lu, Shibao Zheng
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Task-Aware Encoder Control for Deep Video Compression
abstract
Prior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DvC decoder. Empirical evidence demonstrates that our method is applica-ble across multiple tasks with various existing pre-trained Dv'Cs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bi-trate for different tasks, with only one pre-trained decoder.
Xingtong Ge, Jixiang Luo, Tongda Xu, Guo Lu, Dailan He, Yan Wang 0105, Jun Zhang 0004, Hongwei Qin
CVPR5
2024 Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression
Yuan Tian 0017, Guo Lu, Guangtao Zhai
ECCV (49)2
2024 GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
Xingtong Ge, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin, Guo Lu, Jun Zhang 0004
ECCV (9)7
2024 Neural Rate Control for Learned Video Compression
abstract
The learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control technique. Rate control can precisely control the coding bitrate with optimal compression performance, which is a critical technique in practical deployment. To address this issue, we present a fully neural network-based rate control system for learned video compression methods. Our system accurately encodes videos at a given bitrate while enhancing the rate-distortion performance. Specifically, we first design a rate allocation model to assign optimal bitrates to each frame based on their varying spatial and temporal characteristics. Then, we propose a deep learning-based rate implementation network to perform the rate-parameter mapping, precisely predicting coding parameters for a given rate. Our proposed rate control system can be easily integrated into existing learning-based video compression methods. The extensive experimental results show that the proposed method achieves accurate rate control on several baseline methods while also improving overall rate-distortion performance.
Guo Lu, Yunuo Chen 0002, Shen Wang 0013, Yibo Shi, Jing Wang 0194, Li Song 0001
ICLR2
2024 Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion Optimization
abstract
Volumetric videos, benefiting from immersive 3D realism and interactivity, hold vast potential for various applications, while the tremendous data volume poses significant challenges for compression. Recently, NeRF has demonstrated remarkable potential in volumetric video compression thanks to its simple representation and powerful 3D modeling capabilities, where a notable work is ReRF. However, ReRF separates the modeling from compression process, resulting in suboptimal compression efficiency. In contrast, in this paper, we propose a volumetric video compression method based on dynamic NeRF in a more compact manner. Specifically, we decompose the NeRF representation into the coefficient fields and the basis fields, incrementally updating the basis fields in the temporal domain to achieve dynamic modeling. Additionally, we perform end-to-end joint optimization on the modeling and compression process to further improve the compression efficiency. Extensive experiments demonstrate that our method achieves higher compression efficiency compared to ReRF on various datasets.
Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Anni Tang, Qiang Hu 0003, Li Song 0001
ICME2
2024 FreeFlow: A Unified Viewpoint on Diffusion Probabilistic Models via Optimal Transport and Fluid Mechanics
Guo Lu, Shibao Zheng
ICONIP (3)2
2024 Rate-aware Compression for NeRF-based Volumetric Video
abstract
The neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing solutions typically compress these NeRF representations after the training stage, leading to a separation between representation training and compression. In this paper, we try to directly learn a compact NeRF representation for volumetric video in the training stage based on the proposed rate-aware compression framework. Specifically, for volumetric video, we use a simple yet effective modeling strategy to reduce temporal redundancy for the NeRF representation. Then, during the training phase, an implicit entropy model is utilized to estimate the bitrate of the NeRF representation. This entropy model is then encoded into the bitstream to assist in the decoding of the NeRF representation. This approach enables precise bitrate estimation, thereby leading to a compact NeRF representation.Furthermore, we propose an adaptive quantization strategy and learn the optimal quantization step for the NeRF representations. Finally, the NeRF representation can be optimized by using the rate-distortion trade-off. Our proposed compression framework can be used for different representations and experimental results demonstrate that our approach significantly reduces the storage size with marginal distortion and achieves state-of-the-art rate-distortion performance for volumetric video on the HumanRF and ReRF datasets. Compared to the previous state-of-the-art method TeTriRF, we achieved an approximately -80% BD-rate on the HumanRF dataset and -60% BD-rate on the ReRF dataset.
Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Zhengxue Cheng, Anni Tang, Li Song 0001
ACM Multimedia2
2024 AsymLLIC: Asymmetric Lightweight Learned Image Compression
abstract
Learned image compression (LIC) methods often employ symmetrical encoder and decoder architectures, evitably increasing decoding time. However, practical scenarios demand an asymmetric design, where the decoder requires low complexity to cater to diverse low-end devices, while the encoder can accommodate higher complexity to improve coding performance. In this paper, we propose an asymmetric lightweight learned image compression (AsymLLIC) architecture with a novel training scheme, enabling the gradual substitution of complex decoding modules with simpler ones. Building upon this approach, we conduct a comprehensive comparison of different decoder network structures to strike a better trade-off between complexity and compression performance. Experiment results validate the efficiency of our proposed method, which not only achieves comparable performance to VVC but also offers a lightweight decoder with only 51.47 GMACs computation and 19.65M parameters. Furthermore, this design methodology can be easily applied to any LIC models, enabling the practical deployment of LIC techniques.
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Guo Lu, Li Song 0001, Wenjun Zhang 0001
VCIP4
2024 Coarse-to-fine Transformer For Lossless 3D Medical Image Compression
abstract
The rapid advancements in medical imaging have led to a growing demand for high-performance lossless compression of large 3D medical image datasets. Unlike natural images, medical images typically feature three-dimensional structures, and high bit-depth, necessitating specialized compression techniques. Based on a decoder-only transformer, we propose a learnable dual-decoder model for lossless compression of 3D medical images. Our approach packs voxels into patches, which are processed by a patch-level decoder to extract the patch feature. The voxels, along with the patch feature, are subsequently fed into a voxel-level decoder to model each voxel. This coarse-to-fine modeling strategy reduces the computational time for each voxel and enables long-range modeling dependencies. Experimental results demonstrate that our proposed model achieves state-of-the-art compression performance, with an approximately 15% improvement in compression performance over the traditional JP3D benchmark on various datasets.
Guo Lu, Donghui Feng 0003, Zhengxue Cheng, Guosheng Yu, Li Song 0001
VCIP2
2024 Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing System
abstract
In streaming media services, video transcoding is a common practice to alleviate bandwidth demands. Unfortunately, traditional methods employing a uniform rate factor (RF) across all videos often result in significant inefficiencies. Content-adaptive encoding (CAE) techniques address this by dynamically adjusting encoding parameters based on video content characteristics. However, existing CAE methods are often tightly coupled with specific encoding strategies, leading to inflexibility. In this paper, we propose a model that predicts both RF-quality and RF-bitrate curves, which can be utilized to derive a comprehensive bitrate-quality curve. This approach facilitates flexible adjustments to the encoding strategy without necessitating model retraining. The model leverages codec features, content features, and anchor features to predict the bitrate-quality curve accurately. Additionally, we introduce an anchor suspension method to enhance prediction accuracy. Experiments confirm that the actual quality metric (VMAF) of the compressed video stays within ±1 of the target, achieving an accuracy of 99.14%. By incorporating our quality improvement strategy with the rate-quality curve prediction model, we conducted online A/B tests, obtaining both +0.107% improvements in video views and video completions and +0.064% app duration time. Our model has been deployed on the Xiaohongshu App.
Shibo Yin, Zhiyu Zhang 0010, Peirong Ning, Qiubo Chen, Guo Lu, Li Song 0001
VCIP7
2024 Efficient Bitrate Ladder Construction for Per-Shot Adaptive Encoding
abstract
HTTP adaptive streaming (HAS) constructs bitrate ladders to deliver videos with the best possible quality under varying network conditions. Though per-shot content adaptive encoding (CAE) largely improves the compression efficiency by constructing the optimal bitrate ladder for each video shot, it suffers from excessive encoding complexity as all the points in the operating space (typically resolution × bitrate) need to be encoded and compared. To address this issue, this paper proposes an efficient bitrate ladder construction method that encodes only a subset of operating points, then uses curve fitting and inter-curve prediction to estimate other points’ RD performance. The proposed method enables low-complexity ladder construction even for high-dimension operating spaces that incorporate dimensions like encoding presets. Experiments show that this method can achieve RD performance comparable to the original per-shot CAE with only 42% encoding points. Even when minimizing the encoding points to 3.6% of the original CAE, it achieves 15% BD-Rate improvements compared to using the fixed bitrate ladder.
Yan Zhao 0041, Zhengxue Cheng, Guo Lu, Rong Xie 0004, Li Song 0001
VCIP3
2024 A unified efficient deep image compression framework and its application on human-centric Task
Xueyuan Chen, Guo Lu
Multim. Tools Appl.3
2024 A Coding Framework and Benchmark Towards Low-Bitrate Video Understanding
abstract
Video compression is indispensable to most video analysis systems. Despite saving the transportation bandwidth, it also deteriorates downstream video understanding tasks, especially at low-bitrate settings. To systematically investigate this problem, we first thoroughly review the previous methods, revealing that three principles, i.e., task-decoupled, label-free, and data-emerged semantic prior, are critical to a machine-friendly coding framework but are not fully satisfied so far. In this paper, we propose a traditional-neural mixed coding framework that simultaneously fulfills all these principles, by taking advantage of both traditional codecs and neural networks (NNs). On one hand, the traditional codecs can efficiently encode the pixel signal of videos but may distort the semantic information. On the other hand, highly non-linear NNs are proficient in condensing video semantics into a compact representation. The framework is optimized by ensuring that a transportation-efficient semantic representation of the video is preserved w.r.t. the coding procedure, which is spontaneously learned from unlabeled data in a self-supervised manner. The videos collaboratively decoded from two streams (codec and NN) are of rich semantics, as well as visually photo-realistic, empirically boosting several mainstream downstream video analysis task performances without any post-adaptation procedure. Furthermore, by introducing the attention mechanism and adaptive modeling scheme, the video semantic modeling ability of our approach is further enhanced. Fianlly, we build a low-bitrate video understanding benchmark with three downstream tasks on eight datasets, demonstrating the notable superiority of our approach. All codes, data, and models will be open-sourced for facilitating future research.
Yuan Tian 0017, Guo Lu, Yichao Yan, Guangtao Zhai, Li Chen 0021
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Preprocessing Enhanced Image Compression for Machine Vision
abstract
Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.
Guo Lu, Xingtong Ge, Tianxiong Zhong, Qiang Hu 0003, Jing Geng 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 A Character Position-Aware Compression Framework for Screen Text Image
abstract
Text patterns typically exhibit distinct boundaries and sparse color histograms. However, in current hybrid codec frameworks, the positions of coding units are often misaligned with the text patterns, resulting in prediction and color mapping tools consuming a large number of bits to indicate these patterns. Nowadays, some text detection and recognition methods have been proposed to accurately locate and analyze the text regions in screen images. Combined with these techniques, we propose a character position-aware compression framework for screen text image. On the encoder side, a low-complexity detection method is adopted to locate the text characters. Then it copies the detected characters to the position aligned with the coding unit (CU) grid to form a text layer. This text-layer representation can further increase the efficiency of existing screen content coding tools such as Intra Block Copy (IBC). Moreover, we design several compression tools based on this representation. We extend the two Motion Vector (MV) prediction modes: Adaptive Motion Vector Prediction (AMVP) and Merge. We modify the MV encoding syntax according to the layout characteristics of the text layer. We present a Gradient-guided In-loop Filter (GIF) to sharpen the text lines using a convolutional network. Experiments conducted on VVC reference software VTM all_intra configuration show that the proposed framework can achieve an average bitrate savings of 4.6% and 3.6% under the w/ GIF and w/o GIF versions, with a corresponding increase in CPU encoding complexity of 72% and 10%.
Guo Lu, Huanbang Chen, Donghui Feng 0003, Shen Wang 0013, Yan Zhao 0041, Rong Xie 0004, Li Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic Compression
abstract
Most video compression methods aim to improve the decoded video visual quality, instead of particularly guaranteeing the semantic-completeness, which deteriorates downstream video analysis tasks, e.g., action recognition. In this paper, we focus on a novel unsupervised video semantic compression problem, where video semantics is compressed in a downstream task-agnostic manner. To tackle this problem, we first propose a Semantic-Mining-then-Compensation (SMC) framework to enhance the plain video codec with powerful semantic coding capability. Then, we optimize the framework with only unlabeled video data, by masking out a proportion of the compressed video and reconstructing the masked regions of the original video, which is inspired by recent masked image modeling (MIM) methods. Although the MIM scheme learns generalizable semantic features, its inner generative learning paradigm may also facilitate the coding framework memorizing non-semantic information with extra bit costs. To suppress this deficiency, we explicitly decrease the non-semantic information entropy of the decoded video features, by formulating it as a parametrized Gaussian Mixture Model conditioned on the mined video semantics. Comprehensive experimental results demonstrate the proposed approach shows remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets.
Yuan Tian 0017, Guo Lu, Guangtao Zhai
ICCV2
2023 Content Adaptive Checkerboard Context Model for Learned Image Compression
abstract
Learned image compression methods are becoming popular and have achieved excellent performance, of which joint context and hyperprior architectures are the mainstream. In order to avoid the time-consuming serial decoding pipeline introduced by the autoregressive context model, the checkerboard context model (CCM) is proposed to implement fast two-pass coding. However, CCM sets half of the latents as anchors to extract spatial context for the other non-anchors, which is rough and redundant. We propose a more precise and flexible content adaptive checkerboard context model to decrease the numbers and bit consumption of anchors. By introducing pseudo-anchors for simple regions in latents, our method can preserve the capability of fast two-pass coding and outperform CCM in Rate-Distortion performance on several baseline models with negligible computational overhead.
Guo Lu, Donghui Feng 0003, Li Song 0001
ISCAS2
2023 High-Fidelity Free-View Talking Head Synthesis for Low-Bandwidth Video Conference
abstract
As video conferencing becomes an indispensable part of human’s daliy life, how to achieve a high-fidelity calling experience under low bandwidth has been a popular and challenging issue. Deep generative models have great potential in low-bandwidth facial video compression due to the excellent generation capability based on abridged information. Nevertheless, exsiting deep generation-based compression methods tend to handle motion information in pure 2D or pseudo 3D space, causing facial distortion when large head poses are encountered. In this paper, we propose a 3D-aware high-fidelity facial video conferencing system based on a parameterized NeRF-based face model. Through the compression of the parameterized face model and the transmisstion of extracted facial parameters, we implement high-fidelity talking head synthesis for video conferencing at an ultra-low bitrate. Additionally, the 3D perception capability of the system allows for viewpoint control over the head, achieving higher interactivity and practicability. Extensive experiments verify the effectiveness of the proposed 3D-aware high-fidelity free-view facial video conferencing system.
Zhiyu Zhang 0010, Anni Tang, Guo Lu, Rong Xie 0004, Li Song 0001
VCIP4
2023 FVC: An End-to-End Framework Towards Deep Video Compression in Feature Space
abstract
Deep video compression is attracting increasing attention from both deep learning and video processing community. Recent learning-based approaches follow the hybrid coding paradigm to perform pixel space operations for reducing redundancy along both spatial and temporal dimentions, which leads to inaccurate motion estimation or less effective motion compensation. In this work, we propose a feature-space video coding framework (FVC), which performs all major operations (i.e., motion estimation, motion compression, motion compensation and residual compression) in the feature space. Specifically, a new deformable compensation module, which consists of motion estimation, motion compression and motion compensation, is proposed for more effective motion compensation. In our deformable compensation module, we first perform motion estimation in the feature space to produce the motion information (i.e., the offset maps). Then the motion information is compressed by using the auto-encoder style network. After that, we use the deformable convolution operation to generate the predicted feature for motion compensation. Finally, the residual information between the feature from the current frame and the predicted feature from the deformable compensation module is also compressed in the feature space. Motivated by the conventional codecs, in which the blocks with different sizes are used for motion estimation, we additionally propose two new modules called resolution-adaptive motion coding (RaMC) and resolution-adaptive residual coding (RaRC) to automatically cope with different types of motion and residual patterns at different spatial locations. Comprehensive experimental results demonstrate that our proposed framework achieves the state-of-the-art performance on three benchmark datasets including HEVC, UVG and MCL-JCV.
Dong Xu 0001, Guo Lu, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 ParVoro++: A scalable parallel algorithm for constructing 3D Voronoi tessellations based on kd-tree decomposition
Guoqing Wu 0002, Hongyun Tian, Guo Lu
Parallel Comput.3
2023 TVM: A Tile-based Video Management Framework
abstract
With the exponential growth of video data, there is a pressing need for efficient video analysis technology. Modern query frameworks aim to accelerate queries by reducing the frequency of calls to expensive deep neural networks, which often overlook the overhead associated with video decoding and retrieval. Furthermore, video storage frameworks optimize video retrieval through video partition or caching, often relying on prior information about the query workload. To further accelerate queries, this study introduces a novel tile-based video management framework, called TVM, which leverages the semantic information embedded in videos, without being dependent on specific query workloads. By constructing a tile-based semantic index for newly ingested videos, TVM effectively reduces the size of decoded and processed video data. To achieve this, TVM introduces an optimal index construction algorithm that utilizes cost function and pseudo-labels. Additionally, the framework proposes a query-driven tile parallel decoding algorithm and resource caching algorithms, which further expedite the retrieval of video frames. Experimental results demonstrate that TVM can significantly enhance the throughput of various query tasks, achieving a notable speedup of more than 5.6×.
Tianxiong Zhong, Zhiwei Zhang 0002, Guo Lu, Ye Yuan 0001, Guoren Wang
Proc. VLDB Endow.3
2023 CBANet: Toward Complexity and Bitrate Adaptive Deep Image Compression Using a Single Network
abstract
In this work, we propose a new deep image compression framework called Complexity and Bitrate Adaptive Network (CBANet) that aims to learn one single network to support variable bitrate coding under various computational complexity levels. In contrast to the existing state-of-the-art learning-based image compression frameworks that only consider the rate-distortion trade-off without introducing any constraint related to the computational complexity, our CBANet considers the complex rate-distortion-complexity trade-off when learning a single network to support multiple computational complexity levels and variable bitrates. Since it is a non-trivial task to solve such a rate-distortion-complexity related optimization problem, we propose a two-step approach to decouple this complex optimization task into a complexity-distortion optimization sub-task and a rate-distortion optimization sub-task, and additionally propose a new network design strategy by introducing a Complexity Adaptive Module (CAM) and a Bitrate Adaptive Module (BAM) to respectively achieve the complexity-distortion and rate-distortion trade-offs. As a general approach, our network design strategy can be readily incorporated into different deep image compression methods to achieve complexity and bitrate adaptive image compression by using a single network. Comprehensive experiments on two benchmark datasets demonstrate the effectiveness of our CBANet for deep image compression. Code is released at https://github.com/JinyangGuo/CBANet-release.
Jinyang Guo 0002, Dong Xu 0001, Guo Lu
IEEE Trans. Image Process.3
2022 LSVC: A Learning-based Stereo Video Compression Framework
abstract
In this work, we propose the first end-to-end optimized framework for compressing automotive stereo videos (i.e., stereo videos from autonomous driving applications) from both left and right views. Specifically, when compressing the current frame from each view, our framework reduces temporal redundancy by performing motion compensation using the reconstructed intra-view adjacent frame and at the same time exploits binocular redundancy by conducting disparity compensation using the latest reconstructed cross-view frame. Moreover, to effectively compress the introduced motion and disparity offsets for better compensation, we further propose two novel schemes called motion residual compression and disparity residual compression to respectively generate the predicted motion offset and disparity offset from the previously compressed motion offset and disparity offset, such that we can more effectively compress residual offset information for better bit-rate saving. Overall, the entire framework is implemented by the fully-differentiable modules and can be optimized in an end-to-end manner. Our comprehensive experiments on three automotive stereo video benchmarks Cityscapes, KITTI 2012 and KITTI 2015 demonstrate that our proposed framework outperforms the learning-based single-view video codec and the traditional hand-crafted multi-view video codec.
Guo Lu, Shan Liu 0001, Wei Jiang 0001, Dong Xu 0001
CVPR2
2022 Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode Prediction
abstract
The previous deep video compression approaches only use the single scale motion compensation strategy and rarely adopt the mode prediction technique from the traditional standards like H.264/H.265 for both motion and residual compression. In this work, we first propose a coarse-to-fine (C2F) deep video compression framework for better motion compensation, in which we perform motion estimation, compression and compensation twice in a coarse to fine manner. Our C2F framework can achieve better motion compensation results without significantly increasing bit costs. Observing hyperprior information (i.e., the mean and variance values) from the hyperprior networks contains discriminant statistical information of different patches, we also propose two efficient hyperprior-guided mode prediction methods. Specifically, using hyper-prior information as the input, we propose two mode prediction networks to respectively predict the optimal block resolutions for better motion coding and decide whether to skip residual information from each block for better residual coding without introducing additional bit cost while bringing negligible extra computation cost. Comprehensive experimental results demonstrate our proposed C2F video compression framework equipped with the new hyperprior-guided mode prediction methods achieves the state-of-the-art performance on HEVC, UVG and MCL-JCV datasets.
Guo Lu, Jinyang Guo 0002, Shan Liu 0001, Wei Jiang 0001, Dong Xu 0001
CVPR2
2022 Learning based Multi-modality Image and Video Compression
abstract
Multi-modality (i.e., multi-sensor) data is widely used in various vision tasks for more accurate or robust perception. However, the increased data modalities bring new challenges for data storage and transmission. The existing data compression approaches usually adopt individual codecs for each modality without considering the correlation between different modalities. This work proposes a multi-modality compression framework for infrared and visible image pairs by exploiting the cross-modality redun-dancy. Specifically, given the image in the reference modality (e.g., the infrared image), we use the channel-wise alignment module to produce the aligned features based on the affine transform. Then the aligned feature is used as the context information for compressing the image in the current modality (e.g., the visible image), and the corresponding affine coefficients are losslessly compressed at negligible cost. Furthermore, we introduce the Transformer-based spatial alignment module to exploit the correlation between the intermediate features in the decoding procedures for different modalities. Our framework is very flexible and easily extended for multi-modality video compression. Experimental results show our proposed framework outperforms the traditional and learning-based single modality compression methods on the FLIR and KAIST datasets.
Guo Lu, Tianxiong Zhong, Jing Geng 0002, Qiang Hu 0003, Dong Xu 0001
CVPR1
2022 Content Adaptive Latents and Decoder for Neural Image Compression
Guanbo Pan, Guo Lu, Dong Xu 0001
ECCV (18)2
2022 Position-based Motion Vector Prediction for Textual Image Coding
abstract
Textual content is becoming increasingly important in video conferencing, while existing screen content encoding tools still produce a high bitrate in text regions. The main coding tool Intra Block Copy (IBC) inherits the MV prediction mechanism in inter-frame coding, but the adjacent text characters typically have irrelevant MVs, making it inefficient to predict MV using only neighbor MVs. To solve the problem, we propose the Position-based Motion Vector Prediction, to cache IBC AMVP PU positions as predictors. One character can find the previously encoded position to construct a good MV prediction. Experiment results show the effectiveness of the proposed prediction scheme.
Donghui Feng 0003, Guo Lu, Li Song 0001
PCS3
2022 Perceptual Video Coding Based on Semantic-Guided Texture Detection and Synthesis
abstract
Visually insensitive texture regions consume a large number of bitrate in hybrid video coding, leading to the waste of bandwidth resources. For this, we propose a semantic-guided texture synthesis framework (STSF). At encoder, high-level semantic information is adopted as texture features to detect texture regions and is sent to the decoder. Detected texture regions are coarsely encoded by hybrid codec. To generate realistic texture patterns, we design a multi-model semantic-guided texture synthesis generative adversarial network (STSGAN) at decoder, which works in a divide-and-conquer manner that semantically different texture regions are synthesized by different submodels in it. Experimental results show that STSF can achieve a −17.2% MOS BD-rate under the lowdelay_P configuration, compared with VVC.
Guo Lu, Rong Xie 0004, Li Song 0001
PCS2
2022 Learning 3D Human Shape and Pose From Dense Body Parts
abstract
Reconstructing 3D human shape and pose from monocular images is challenging despite the promising results achieved by the most recent learning-based methods. The commonly occurred misalignment comes from the facts that the mapping from images to the model space is highly non-linear and the rotation-based pose representation of the body model is prone to result in the drift of joint positions. In this work, we investigate learning 3D human shape and pose from dense correspondences of body parts and propose a Decompose-and-aggregate Network (DaNet) to address these issues. DaNet adopts the dense correspondence maps, which densely build a bridge between 2D pixels and 3D vertexes, as intermediate representations to facilitate the learning of 2D-to-3D mapping. The prediction modules of DaNet are decomposed into one global stream and multiple local streams to enable global and fine-grained perceptions for the shape and pose predictions, respectively. Messages from local streams are further aggregated to enhance the robust prediction of the rotation-based poses, where a position-aided rotation feature refinement strategy is proposed to exploit spatial relationships between body joints. Moreover, a Part-based Dropout (PartDrop) strategy is introduced to drop out dense information from intermediate representations during training, encouraging the network to focus on more complementary body parts as well as neighboring position features. The efficacy of the proposed method is validated on both indoor and real-world datasets including Human3.6M, UP3D, COCO, and 3DPW, showing that our method could significantly improve the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://hongwenzhang.github.io/dense2mesh.
Hongwen Zhang 0001, Jie Cao 0002, Guo Lu, Wanli Ouyang, Zhenan Sun
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Spatial Temporal Video Enhancement Using Alternating Exposures
abstract
High-speed video acquisition under poor illumination conditions is a challenging task. Imaging using long exposure can ensure brightness and suppress noise. However, the captured images may be blurry due to fast object movements or camera shakes. Imaging with short exposure can record sharp textures, but the high camera gain may cause noticeable noise. To alleviate this dilemma, we design a camera system using alternating exposures, where frames expose cyclically in a short-long way. The system consists of restoration and interpolation modules to reconstruct sharp, noise-reduced, high-frame-rate frames from low-frame-rate alternate-exposed input images. We design an optical-flow-based alternate-complementary alignment architecture for spatial enhancement, which effectively aligns the short-exposed and long-exposed images in a two-stage progressive way. Moreover, it explores complementary information from short-exposed and long-exposed inputs to ensure consistency between outputs. We propose a flow-enhanced frame interpolation module for temporal enhancement, which refines the intermediate flows and reconstructs the intermediate images based on the restored images of the alignment network and warped input neighboring frames. The whole network with two modules is end-to-end jointly learnable. We first evaluate the algorithm on simulation data. To demonstrate practicality, we then test it on real data by setting up a prototype camera. We propose an effective spatial degradation regularization strategy to reduce the domain gap between simulation and real data. Besides, we extend our method by integrating multi-frame exposure fusion technology to reduce overexposure areas in real scenarios. Experimental results show that our method performs favorably against state-of-the-art methods on both synthetic data and real-world data.
Wang Shen, Guo Lu, Guangtao Zhai, Li Chen 0021, Muhammad Salman Asif
IEEE Trans. Circuits Syst. Video Technol.3
2022 Exploiting Intra-Slice and Inter-Slice Redundancy for Learning-Based Lossless Volumetric Image Compression
abstract
3D volumetric image processing has attracted increasing attention in the last decades, in which one major research area is to develop efficient lossless volumetric image compression techniques to better store and transmit such images with massive amount of information. In this work, we propose the first end-to-end optimized learning framework for losslessly compressing 3D volumetric data. Our approach builds upon a hierarchical compression scheme by additionally introducing the intra-slice auxiliary features and estimating the entropy model based on both intra-slice and inter-slice latent priors. Specifically, we first extract the hierarchical intra-slice auxiliary features through multi-scale feature extraction modules. Then, an Intra-slice and Inter-slice Conditional Entropy Coding module is proposed to fuse the intra-slice and inter-slice information from different scales as the context information. Based on such context information, we can predict the distributions for both intra-slice auxiliary features and the slice images. To further improve the lossless compression performance, we also introduce two new gating mechanisms called Intra-Gate and Inter-Gate to generate the optimal feature representations for better information fusion. Eventually, we can produce the bitstream for losslessly compressing volumetric images based on the estimated entropy model. Different from the existing lossless volumetric image codecs, our end-to-end optimized framework jointly learns both intra-slice auxiliary features at different scales for each slice and inter-slice latent features from previously encoded slices for better entropy estimation. The extensive experimental results indicate that our framework outperforms the state-of-the-art hand-crafted lossless volumetric image codecs (e.g., JP3D) and the learning-based lossless image compression method on four volumetric image benchmarks for losslessly compressing both 3D Medical Images and Hyper-Spectral Images.
Shuhang Gu, Guo Lu, Dong Xu 0001
IEEE Trans. Image Process.3
2022 TSA-SCC: Text Semantic-Aware Screen Content Coding With Ultra Low Bitrate
abstract
Due to the rapid growth of web conferences, remote screen sharing, and online games, screen content has become an important type of internet media information and over 90% of online media interactions are screen based. Meanwhile, as the main component in the screen content, textual information averagely takes up over 40% of the whole image on various commonly used screen content datasets. However, it is difficult to compress the textual information by using the traditional coding schemes as HEVC, which assumes strong spatial and temporal correlations within the image/video. State-of-the-art screen content coding (SCC) standard as HEVC-SCC still adopts a block-based coding framework and does not consider the text semantics for compression, thus inevitably blurring texts at a lower bitrate. In this paper, we propose a general text semantic-aware screen content coding scheme (TSA-SCC) for ultra low bitrate setting. This method detects the abrupt picture in a screen content video (or image), recognizes textual information (including word, position, font type, font size and font color) in the abrupt picture based on neural networks, and encodes texts with text coding tools. The other pictures as well as the background image after removing texts from the abrupt picture via inpainting, are encoded with HEVC-SCC. Compared with HEVC-SCC, the proposed method TSA-SCC reduces bitrate by up to 3× at a similar compression quality. Moreover, TSA-SCC achieves much better visual quality with less bitrate consumption when encoding the screen content video/image at ultra low bitrates.
Ling Li 0001, Ruizhi Chen, Haochen Li 0002, Guo Lu, Limin Cheng
IEEE Trans. Image Process.6
2021 FVC: A New Framework Towards Deep Video Compression in Feature Space
abstract
Learning based video compression attracts increasing attention in the past few years. The previous hybrid coding approaches rely on pixel space operations to reduce spatial and temporal redundancy, which may suffer from inaccurate motion estimation or less effective motion compensation. In this work, we propose a feature-space video coding network (FVC) by performing all major operations (i.e., motion estimation, motion compression, motion compensation and residual compression) in the feature space. Specifically, in the proposed deformable compensation module, we first apply motion estimation in the feature space to produce motion information (i.e., the offset maps), which will be compressed by using the auto-encoder style network. Then we perform motion compensation by using deformable convolution and generate the predicted feature. After that, we compress the residual feature between the feature from the current frame and the predicted feature from our deformable compensation module. For better frame reconstruction, the reference features from multiple previous reconstructed frames are also fused by using the nonlocal attention mechanism in the multi-frame feature fusion module. Comprehensive experimental results demonstrate that the proposed framework achieves the state-of-the-art performance on four benchmark datasets including HEVC, UVG, VTL and MCL-JCV.
Guo Lu, Dong Xu 0001
CVPR2
2021 VoxelContext-Net: An Octree Based Framework for Point Cloud Compression
abstract
In this paper, we propose a two-stage deep learning framework called VoxelContext-Net for both static and dynamic point cloud compression. Taking advantages of both octree based methods and voxel based schemes, our approach employs the voxel context to compress the octree structured data. Specifically, we first extract the local voxel representation that encodes the spatial neighbouring context information for each node in the constructed octree. Then, in the entropy coding stage, we propose a voxel context based deep entropy model to compress the symbols of non-leaf nodes in a lossless way. Furthermore, for dynamic point cloud compression, we additionally introduce the local voxel representations from the temporal neighbouring point clouds to exploit temporal dependency. More importantly, to alleviate the distortion from the octree construction procedure, we propose a voxel context based 3D coordinate refinement method to produce more accurate reconstructed point cloud at the decoder side, which is applicable to both static and dynamic point cloud compression. The comprehensive experiments on both static and dynamic point cloud benchmark datasets(e.g., ScanNet and Semantic KITTI) clearly demonstrate the effectiveness of our newly proposed method VoxelContext-Net for 3D point cloud geometry compression.
Zizheng Que, Guo Lu, Dong Xu 0001
CVPR2
2021 Self-Conditioned Probabilistic Learning of Video Rescaling
abstract
Bicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a self-conditioned probabilistic framework for video rescaling to learn the paired downscaling and upscaling procedures simultaneously. During the training, we decrease the entropy of the information lost in the downscaling by maximizing its probability conditioned on the strong spatial-temporal prior information within the downscaled video. After optimization, the downscaled video by our framework preserves more meaningful information, which is beneficial for both the upscaling step and the downstream tasks, e.g., video action recognition task. We further extend the framework to a lossy video compression system, in which a gradient estimator for non-differential industrial lossy codecs is proposed for the end-to-end training of the whole system. Extensive experimental results demonstrate the superiority of our approach on video rescaling, video compression, and efficient action recognition tasks.
Yuan Tian 0017, Guo Lu, Xiongkuo Min, Zhaohui Che, Guangtao Zhai, Guodong Guo
ICCV2
2021 Deep Learning for Visual Data Compression
abstract
In this paper, we will introduce the recent progress in deep learning based visual data compression, including image compression, video compression and point cloud compression. In the past few years, deep learning techniques have been successfully applied to various computer vision and image processing applications. However, for the data compression task, the traditional approaches (i.e., block based motion estimation and motion compensation, etc.) are still widely employed in the mainstream codecs. Considering the powerful representation capability of neural networks, it is feasible to improve the data compression performance by employing the advanced deep learning technologies. To this end, the deep leaning based compression approaches have recently received increasing attention from both academia and industry in the field of computer vision and signal processing.
Guo Lu, Shenlong Wang, Shan Liu 0001, Radu Timofte
ACM Multimedia1
2021 Guest Editorial: Special Issue on Deep Learning for Video Analysis and Compression
Dong Xu 0001, Rama Chellappa, Luc Van Gool, Guo Lu
Int. J. Comput. Vis.4
2021 An End-to-End Learning Framework for Video Compression
abstract
Traditional video compression approaches build upon the hybrid coding framework with motion-compensated prediction and residual transform coding. In this paper, we propose the first end-to-end deep video compression framework to take advantage of both the classical compression architecture and the powerful non-linear representation ability of neural networks. Our framework employs pixel-wise motion information, which is learned from an optical flow network and further compressed by an auto-encoder network to save bits. The other compression components are also implemented by the well-designed networks for high efficiency. All the modules are jointly optimized by using the rate-distortion trade-off and can collaborate with each other. More importantly, the proposed deep video compression framework is very flexible and can be easily extended by using lightweight or advanced networks for higher speed or better efficiency. We also propose to introduce the adaptive quantization layer to reduce the number of parameters for variable bitrate coding. Comprehensive experimental results demonstrate the effectiveness of the proposed framework on the benchmark datasets.
Guo Lu, Xiaoyun Zhang 0001, Wanli Ouyang, Li Chen 0021, Dong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Improving Deep Video Compression by Resolution-Adaptive Flow Coding
Dong Xu 0001, Guo Lu, Wanli Ouyang, Shuhang Gu
ECCV (2)4
2020 Content Adaptive and Error Propagation Aware Deep Video Compression
Guo Lu, Chunlei Cai, Xiaoyun Zhang 0001, Li Chen 0021, Wanli Ouyang, Dong Xu 0001
ECCV (2)1
2020 Learned image and video compression with deep neural networks
abstract
This tutorial aims at reviewing the recent progress in the deep learning based data compression, including image compression and video compression. In the past years, deep learning techniques have been successfully applied to a large number of computer vision and image processing tasks. However, for the data compression task, the traditional approaches (i.e., block based motion estimation and motion compensation, etc.) are still widely employed in the mainstream codecs. Considering the powerful representation capability, it is possible to improve the data compression performance by employing the advanced deep learning technologies. To this end, deep leaning based compression approaches have recently received significant attention from both academia and industry in the field of computer vision and image/video compression. In this tutorial, we will introduce the related deep learning techniques for image compression and video compression. Specifically, in this tutorial, we will first introduce the basic pipeline for the traditional codecs, such as JPEG, H.264 and HEVC. Then, we will discuss the common network architectures for visual data compression and analyse different learning based entropy models. Based on these techniques, we will describe several widely used end-to-end optimized frameworks for visual data compression. In summary, our tutorial will cover both the traditional data coding techniques and the popular learning based visual data compression algorithms, which will help the audiences with different backgrounds learn the recent progresses in this emerging research area.
Dong Xu 0001, Guo Lu, Radu Timofte
VCIP2
2020 Deep Non-Local Kalman Network for Video Compression Artifact Reduction
abstract
Video compression algorithms are widely used to reduce the huge size of video data, but they also introduce unpleasant visual artifacts due to the lossy compression. In order to improve the quality of the compressed videos, we proposed a deep non-local Kalman network for compression artifact reduction. Specifically, the video restoration is modeled as a Kalman filtering procedure and the decoded frames can be restored from the proposed deep Kalman model. Instead of using the noisy previous decoded frames as temporal information, the less noisy previous restored frame is employed in a recursive way, which provides the potential to generate high quality restored frames. In the proposed framework, several deep neural networks are utilized to estimate the corresponding states in the Kalman filter and integrated together in the deep Kalman filtering network. More importantly, we also exploit the non-local prior information by incorporating the spatial and temporal non-local networks for better restoration. Our approach takes the advantages of both the model-based methods and learning-based methods, by combining the recursive nature of the Kalman model and powerful representation ability of neural networks. Extensive experimental results on the Vimeo-90k and HEVC benchmark datasets demonstrate the effectiveness of our proposed method.
Guo Lu, Xiaoyun Zhang 0001, Wanli Ouyang, Dong Xu 0001, Li Chen 0021
IEEE Trans. Image Process.1
2019 DVC: An End-To-End Deep Video Compression Framework
abstract
Conventional video compression approaches use the predictive coding architecture and encode the corresponding motion information and residual information. In this paper, taking advantage of both classical architecture in the conventional video compression method and the powerful non-linear representation ability of neural networks, we propose the first end-to-end video compression deep model that jointly optimizes all the components for video compression. Specifically, learning based optical flow estimation is utilized to obtain the motion information and reconstruct the current frames. Then we employ two auto-encoder style neural networks to compress the corresponding motion and residual information. All the modules are jointly learned through a single loss function, in which they collaborate with each other by considering the trade-off between reducing the number of compression bits and improving quality of the decoded video. Experimental results show that the proposed approach can outperform the widely used video coding standard H.264 in terms of PSNR and be even on par with the latest standard H.265 in terms of MS-SSIM. Code is released at https://github.com/GuoLusjtu/DVC.
Guo Lu, Wanli Ouyang, Dong Xu 0001, Xiaoyun Zhang 0001, Chunlei Cai
CVPR1
2019 DaNet: Decompose-and-aggregate Network for 3D Human Shape and Pose Estimation
abstract
Reconstructing 3D human shape and pose from a monocular image is challenging despite the promising results achieved by most recent learning based methods. The commonly occurred misalignment comes from the facts that the mapping from image to model space is highly non-linear and the rotation-based pose representation of the body model is prone to result in drift of joint positions. In this work, we present the Decompose-and-aggregate Network (DaNet) to address these issues. DaNet includes three new designs, namely UVI guided learning, decomposition for fine-grained perception, and aggregation for robust prediction. First, we adopt the UVI maps, which densely build a bridge between 2D pixels and 3D vertexes, as an intermediate representation to facilitate the learning of image-to-model mapping. Second, we decompose the prediction task into one global stream and multiple local streams so that the network not only provides global perception for the camera and shape prediction, but also has detailed perception for part pose prediction. Lastly, we aggregate the message from local streams to enhance the robustness of part pose prediction, where a position-aided rotation feature refinement strategy is proposed to exploit the spatial relationship between body parts. Such a refinement strategy is more efficient since the correlations between position features are stronger than that in the original rotation feature space. The effectiveness of our method is validated on the Human3.6M and UP-3D datasets. Experimental results show that the proposed method significantly improves the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://github.com/HongwenZhang/DaNet-3DHumanReconstrution .
Hongwen Zhang 0001, Jie Cao 0002, Guo Lu, Wanli Ouyang, Zhenan Sun
ACM Multimedia3
2019 A Novel Deep Progressive Image Compression Framework
abstract
In Internet applications, compressing the image without perceptually distinguishable distortions and loading the images without notable delays in the client end can significantly improve the user experience. Compressing the image at high bit rates can maintain the high quality of the decoded image but in cost of long transmitting and decoding time, resulting in bad user experience. The progressive coding scheme can resolve the conflict between the high quality requirement and the large loading delay. This paper proposes a novel efficient progressive image coding framework based on deep convolutional neural networks. The proposed framework is composed of a uniform encoder network and two progressive decoder networks. The encoder network decomposes the input image into two scales of representations, that can be transmitted and reconstructed progressively into a basic quality preview image and a high-quality image by two individual decoder networks respectively. All the networks are jointly learned when achieving the rate distortion optimization of both scales. Experiments results show that the proposed method has much better coding performance than the commercial codecs WebP and JPEG, which are commonly used in Internet applications. Meanwhile, the proposed codec consumes much less time to load the image compared with WebP.
Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001, Guo Lu
PCS4
2018 Deep Kalman Filtering Network for Video Compression Artifact Reduction
Guo Lu, Wanli Ouyang, Dong Xu 0001, Xiaoyun Zhang 0001, Ming-Ting Sun
ECCV (14)1
2018 Novel Integration of Frame Rate Up Conversion and HEVC Coding Based on Rate-Distortion Optimization
abstract
Frame rate up conversion (FRUC) can improve the visual quality by interpolating new intermediate frames. However, high frame rate videos by FRUC are confronted with more bitrate consumption or annoying artifacts of interpolated frames. In this paper, a novel integration framework of FRUC and high efficiency video coding (HEVC) is proposed based on rate-distortion optimization, and the interpolated frames can be reconstructed at encoder side with low bitrate cost and high visual quality. First, joint motion estimation (JME) algorithm is proposed to obtain robust motion vectors, which are shared between FRUC and video coding. What's more, JME is embedded into the coding loop and employs the original motion search strategy in HEVC coding. Then, the frame interpolation is formulated as a rate-distortion optimization problem, where both the coding bitrate consumption and visual quality are taken into account. Due to the absence of original frames, the distortion model for interpolated frames is established according to the motion vector reliability and coding quantization error. Experimental results demonstrate that the proposed framework can achieve 21% ~ 42% reduction in BDBR, when compared with the traditional methods of FRUC cascaded with coding.
Guo Lu, Xiaoyun Zhang 0001, Li Chen 0021
IEEE Trans. Image Process.1