VLDB 2026 Research / reviewers in the wild / expert
Zhengxue Cheng
dblp:179/1018
· DBLP profile ↗
41ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint VideosabstractFree-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent compression and customized motion formats, limiting generalization and standardization. To address this, we propose D-FCGS, a novel Feedforward Compression framework for Dynamic Gaussian Splatting. Key innovations include: (1) a standardized Group-of-Frames (GoF) structure with I-P coding, leveraging sparse control points to extract inter-frame motion tensors; (2) a dual prior-aware entropy model that fuses hyperprior and spatial-temporal priors for accurate rate estimation; (3) a control-point-guided motion compensation mechanism and refinement network to enhance view-consistent fidelity. Trained on Gaussian frames derived from multi-view videos, D-FCGS generalizes across diverse scenes in a zero-shot fashion. Experiments show that it matches the rate-distortion performance of optimization-based methods, achieving over 40 times compression compared to the baseline while preserving visual quality across viewpoints. This work advances feedforward compression of dynamic 3DGS, facilitating scalable FVV transmission and storage for immersive applications. Yan Zhao 0041, Qiang Wang 0061, Zhixin Xu, Li Song 0001, Zhengxue Cheng |
AAAI | 6 |
| 2026 | Diff-Band:Bandwidth Estimation in RTC via Diffusion-Based Offline Reinforcement Learning
Bingcong Lu, Zhengxue Cheng, Li Song 0001, Bingnan Duan, Jintao Fang |
ISCAS | 3 |
| 2026 | Unified Multimodal Retrieval Framework for Multimodal RAG
Tianyi Feng, Ruiyan Wang, Fei Huang 0002, Zhengxue Cheng, Rong Xie 0004, Li Song 0001 |
PAKDD (4) | 6 |
| 2026 | OmniScaleSR: Unleashing Scale-Controlled Diffusion Prior for Faithful and Realistic Arbitrary-Scale Image Super-ResolutionabstractArbitrary-scale super-resolution (ASSR) overcomes the limitation of traditional super-resolution (SR) that works only at a fixed scale (e.g., ×4), enabling a single model to achieve arbitrary-scale SR. Most ASSR methods explicitly incorporate implicit neural representation (INR) to achieve ASSR, but INR’s inherently regression-driven feature extraction and aggregation nature restricts their capacity to synthesize meticulous details, leading to low realism. Recently, diffusion-based realistic image super-resolution (Real-ISR) methods leverage the pre-trained diffusion prior and have shown promising results at ×4 scale. We find that they could also achieve ASSR because the powerful pre-trained diffusion prior implicitly employs SR scale adaptation by encouraging the model to always generate high-realism images. However, due to the lack of explicit SR scale controls, the model fails to effectively manage the diffusion behavior according to different SR scales, causing either excessive hallucination or blurry results, especially for ultra-high magnification. To address these limitations, we proposeOmniScaleSR, a novel diffusion-based realistic arbitrary-scale super-resolution (Real-ASSR) method to achieve both high fidelity and high-realism ASSR. We introduce explicit diffusion-native SR scale controls, which could be elegantly coupled with the implicit scale adaptation, unleashing scale-controlled diffusion prior to dynamically managing the diffusion behavior in a content- and scale-aware manner. Furthermore, we incorporate multi-domain fidelity enhancement designs to achieve more faithful reconstruction. Extensive experiments on both bicubic degradation benchmarks and real-world datasets demonstrate that OmniScaleSR consistently outperforms state-of-the-art methods in terms of both fidelity and perceptual realism, with especially strong performance under high-magnification scenarios. Codes will be at https://github.com/chaixinning/OmniScaleSR. Xinning Chai, Zhengxue Cheng, Hengsheng Zhang, Yingsheng Qin, Yucai Yang, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Distilling Complexity-Scalable Learned Image Compression Models via Neural Architecture Search
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Cheems Wang, Qunshan Gu, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Diff-Restorer: Unleashing Visual Prompts for Diffusion-Based Universal Image RestorationabstractImage restoration aims to recover high-quality images from degraded observations, yet real-world degradations are complex, coupled, and difficult to model. Existing task-specific methods struggle to generalize beyond predefined degradation types, while recent all-in-one or prompt-based methods still face three key challenges: (1) they rely on task-specific training or fixed prompt pools, limiting adaptability to real-world and mixed degradations; (2) human-instruction or implicit-prompt mechanisms make them difficult to use in practice; and (3) they often fail to balance structural fidelity and perceptual realism. To address these issues, we propose Diff-Restorer, a diffusion-based universal image restoration framework that unifies diverse degradation handling within a single model. Diff-Restorer adaptively extracts decoupled visual prompts from a visual-language model (CLIP), including clear semantic and degradation embeddings. The clear semantic embeddings serve as content prompts to guide the diffusion model for generation, improving perceptual quality. The degradation embeddings as the task identifier modulate the Image-guided Control Module to generate structure control, ensuring faithfulness. Furthermore, we design a Task-aware Decoder to perform structural correction and convert the latent code to the pixel domain. Extensive experiments on various single, real-world, and mixed degradation tasks show that Diff-Restorer outperforms state-of-the-art methods in terms of generality, realism, and fidelity. Hengsheng Zhang, Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | VRVVC: Variable-Rate NeRF-Based Volumetric Video CompressionabstractNeural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets. Qiang Hu 0003, Houqiang Zhong, Zihan Zheng, Xiaoyun Zhang 0001, Zhengxue Cheng, Li Song 0001, Guangtao Zhai, Yanfeng Wang 0001 |
AAAI | 5 |
| 2025 | L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text CompressionabstractLearning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity design while maintaining compression performance. We introduce a novel Learned Lossless Low-complexity Text Compression method (L3TC). Specifically, we conduct extensive experiments demonstrating that RWKV models achieve the fastest decoding speed with a moderate compression ratio, making it the most suitable backbone for our method. Second, we propose an outlier-aware tokenizer that uses a limited vocabulary to cover frequent tokens while allowing outliers to bypass the prediction and encoding. Third, we propose a novel high-rank reparameterization strategy that enhances the learning capability during training without increasing complexity during inference. Experimental results validate that our method achieves 48% bit saving compared to gzip compressor. Besides, L3TC offers compression performance comparable to other learned compressors, with a 50x reduction in model parameters. More importantly, L3TC is the fastest among all learned compressors, providing real-time decoding speeds up to megabytes per second. Junxuan Zhang, Zhengxue Cheng, Yan Zhao 0041, Dajiang Zhou, Guo Lu, Li Song 0001 |
AAAI | 2 |
| 2025 | Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image CompressionabstractNeural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects for a fixed neural image codec. Specifically, we introduce a plug-and-play module at the decoder side that leverages a latent diffusion process to transform the decoded features, enhancing either low distortion or high perceptual quality without altering the original image compression codec. Our approach facilitates fusion of original and transformed features without additional training, enabling users to flexibly adjust the balance between distortion and perception during inference. Extensive experimental results demonstrate that our method significantly enhances the pretrained codecs with a wide, adjustable distortion-perception range while maintaining their original compression capabilities. For instance, we can achieve more than 150% improvement in LPIPS-BDRate without sacrificing more than 1 dB in PSNR. Chuqin Zhou, Guo Lu, Jiangchuan Li, Zhengxue Cheng, Li Song 0001, Wenjun Zhang 0001 |
AAAI | 5 |
| 2025 | Linear Attention Modeling for Learned Image CompressionabstractRecent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKV-based Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https://github.com/sjtu-medialab/RwkvCompress. Donghui Feng 0003, Zhengxue Cheng, Shen Wang 0013, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song 0001 |
CVPR | 2 |
| 2025 | Stable Diffusion is a Natural Cross-Modal Decoder for Layered AI-Generated Image CompressionabstractRecent advances in Artificial Intelligence Generated Content (AIGC) triggered an increasing need to transmit and compress the vast number of AI-generated images (AIGIs). However, there is a noticeable deficiency in research focused on compression methods for AIGIs. To address this critical gap, we advocate that Stable Diffusion serves as a natural cross-modal decoder by leveraging rich and scalable priors, and introduce a scalable cross-modal compression framework that incorporates multiple human-comprehensible modalities. As illustrated in Fig. 1(a), the proposed framework encodes images into a layered bitstream: a semantic prior that delivers high-level semantic information through text prompts; a structural prior that captures spatial details using edge or skeleton maps; and a texture prior that preserves local textures via a colormap. Utilizing Stable Diffusion as the backend, the decoder leverages multi-modal scalable priors to generate images with different levels of fidelity. Experiments show our method preserves realistic details and semantic fidelity at an extremely low bitrate (< 0.02 bpp), comparable with recent perceptual coding approaches and outperforming VVC. The R-D performance also demonstrate the scalability of our proposed multi-layered bitstream since image fidelity incrementally improves with structure and texture priors provided during decoding. Additionally, as illustrated in Fig. 1(b), our framework facilitates downstream editing applications such as Structure Manipulation, Texture Synthesis, and Object Erasing, without requiring full decoding, thereby paving a new direction for future research in AIGI compression. Ruijie Chen, Qi Mao 0002, Zhengxue Cheng |
DCC | 3 |
| 2025 | Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-ResolutionabstractRecent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks. Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001 |
ICME | 4 |
| 2025 | A Lightweight 3-axis Permanent Magnetic Sponge-based Self-Adapting Tactile SensorabstractTactile sensors are indispensable in robotic systems because they deliver vital contact information during environmental interactions. In our work, we leverage the variable compliance of a porous material—where different interaction forces induce varying degrees of compliance—to achieve self-adapting tactile sensing. This distinctive non-linear characteristic allows its sensitivity to be automatically tuned over a range from 0.008 mT/N to 0.045 mT/N. After coating with a magnetic polymer, the porous material functions as a 3-axis magnetic sensing medium. Its length and width are set at 30 mm and 35 mm respectively to accommodate the printed circuit board. To preserve the overall measuring range, it is designed with a thickness of 15 mm. This thickness enables monitoring of the volumetric changes due to the enhanced compliance, which is suitable for three-dimensional shape recognition. In this work, we present the design, fabrication, experimental characterization, and applications of an lightweight 3-axis magnetic sponge sensor with overall dimensions of 30 mm (width) × 35 mm (length) × 17 mm (height) and a detection range of 60 N. Notably, the sensing material weighs only 2 g, thanks to its porous structure. Devesh Abhyankar, Yuhiro Iwamoto, Zhengxue Cheng, Ruotong Zhao, Shigeki Sugano, Mitsuhiro Kamezaki |
IROS | 4 |
| 2025 | Rate-Aware Learned Speech CompressionabstractThe rapid rise of real-time communication and large language models has significantly increased the importance of speech compression. Deep learning-based neural speech codecs have outperformed traditional signal-level speech codecs in terms of rate-distortion (RD) performance. Typically, these neural codecs employ an encoder-quantizer-decoder architecture, where audio is first converted into latent code feature representations and then into discrete tokens. However, this architecture exhibits insufficient RD performance due to two main drawbacks: (1) the inadequate performance of the quantizer, challenging training processes, and issues such as codebook collapse; (2) the limited representational capacity of the encoder and decoder, making it difficult to meet feature representation requirements across various bitrates. In this paper, we propose a rate-aware learned speech compression scheme that replaces the quantizer with an advanced channel-wise entropy model to improve RD performance, simplify training, and avoid codebook collapse. We employ multi-scale convolution and linear attention mixture blocks to enhance the representational capacity and flexibility of the encoder and decoder. Experimental results demonstrate that the proposed method achieves state-of-the-art RD performance, obtaining 53.51% BD-Rate bitrate saving in average, and achieves 0.26 BD-VisQol and 0.44 BD-PESQ gains. Zhengxue Cheng, Guangchuan Chi, Yuelin Hu, Li Song 0001 |
ISCAS | 2 |
| 2025 | MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene ReconstructionabstractMulti-view egocentric dynamic scene reconstruction holds significant research value for applications in holographic documentation of social interactions. However, existing reconstruction datasets focus on static multi-view or single-egocentric view setups, lacking multi-view egocentric datasets for dynamic scene reconstruction. Therefore, we present MultiEgo, the first multi-view egocentric dataset for 4D dynamic scene reconstruction. The dataset comprises five canonical social interaction scenes: meetings, performances, and a presentation. Each scene provides five authentic egocentric videos captured by participants wearing AR glasses. We design a hardware-based data acquisition system and processing pipeline, achieving sub-millisecond temporal synchronization across views, coupled with accurate pose annotations. Experiment validation demonstrates the practical utility and effectiveness of our dataset for free-viewpoint video (FVV) applications, establishing MultiEgo as a foundational resource for advancing multi-view egocentric dynamic scene reconstruction research. Bate Li, Houqiang Zhong, Zhengxue Cheng, Qiang Hu 0003, Qiang Wang 0061, Li Song 0001, Wenjun Zhang 0001 |
ACM Multimedia | 3 |
| 2025 | SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garmentsabstract3D digital garment generation and editing play a pivotal role in fashion design, virtual try-on, and gaming. Traditional methods struggle to meet the growing demand due to technical complexity and high resource costs. Learning-based approaches offer faster, more diverse garment synthesis based on specific requirements and reduce human efforts and time costs. However, they still face challenges such as inconsistent multi-view geometry or textures and heavy reliance on detailed garment topology and manual rigging. We propose SemanticGarment, a 3D Gaussian-based method that realizes high-fidelity 3D garment generation from text or image prompts and supports semantic-based interactive editing for flexible user customization. To ensure multi-view consistency and garment fitting, we propose to leverage structural human priors for the generative model by introducing a 3D semantic clothing model, which initializes the geometry structure and lays the groundwork for view-consistent garment generation and editing. Without the need to regenerate or rely on existing mesh templates, our approach allows for rapid and diverse modifications to existing Gaussians, either globally or within a local region. To address the artifacts caused by self-occlusion for garment reconstruction based on single image, we develop a self-occlusion optimization strategy to mitigate holes and artifacts that arise when directly animating self-occluded garments. Extensive experiments are conducted to demonstrate our superior performance in 3D garment generation and editing. Ruiyan Wang, Zhengxue Cheng, Zonghao Lin, Jun Ling, Yanru An, Rong Xie 0004, Li Song 0001 |
ACM Multimedia | 2 |
| 2025 | PA-HOI: A Physics-Aware Human and Object Interaction Dataset
Ruiyan Wang, Lin Zuo, Zonghao Lin, Qiang Wang 0061, Zhengxue Cheng, Rong Xie 0004, Jun Ling, Li Song 0001 |
ACM Multimedia | 5 |
| 2025 | A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001 |
PCS | 2 |
| 2025 | AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse ViewsabstractThe demand for semantically rich 3D models of indoor scenes is rapidly growing, driven by applications in augmented reality, virtual reality, and robotics. However, creating them from sparse views remains a challenge due to geometric ambiguity. Existing methods often treat semantics as a passive feature painted on an already-formed, and potentially flawed, geometry. We posit that for robust sparse-view reconstruction, semantic understanding instead be an active, guiding force. This paper introduces AlignGS, a novel framework that actualizes this vision by pioneering a synergistic, end-to-end optimization of geometry and semantics. Our method distills rich priors from 2D foundation models and uses them to directly regularize the 3D representation through a set of novel semantic-to-geometry guidance mechanisms, including depth consistency and multifaceted normal regularization. Extensive evaluations on standard benchmarks demonstrate that our approach achieves state-of-the-art results in novel view synthesis and produces reconstructions with superior geometric accuracy. The results validate that leveraging semantic priors as a geometric regularizer leads to more coherent and complete 3D models from limited input views. Our code is avaliable at https://github.com/MediaX-SJTU/AlignGS. Yijie Gao, Houqiang Zhong, Tianchi Zhu, Zhengxue Cheng, Qiang Hu 0003, Li Song 0001 |
VCIP | 4 |
| 2025 | Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video ConferenceabstractThe demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D video compression techniques fail to preserve fine-grained geometric and appearance details, while implicit neural rendering methods like NeRF suffer from prohibitive computational costs. To address these challenges, we propose a lightweight, high-fidelity, low-bitrate 3D talking face compression framework that integrates FLAME-based parametric modeling with 3DGS neural rendering. Our approach transmits only essential facial metadata in real time, enabling efficient reconstruction with a Gaussian-based head model. Additionally, we introduce a compact representation and compression scheme, including Gaussian attribute compression and MLP optimization, to enhance transmission efficiency. Experimental results demonstrate that our method achieves superior rate-distortion performance, delivering high-quality facial rendering at extremely low bitrates, making it well-suited for real-time 3D video conferencing applications. Jianglong Li, Bingcong Lu, Zhengxue Cheng, Hongwei Hu, Ronghua Wu, Li Song 0001 |
VCIP | 4 |
| 2025 | SSP-IR: Semantic and Structure Priors for Diffusion-Based Realistic Image RestorationabstractRealistic image restoration is a crucial task in computer vision, and diffusion-based models for image restoration have garnered significant attention due to their ability to produce realistic results. Restoration can be seen as a controllable generation conditioning on priors. However, due to the severity of image degradation, existing diffusion-based restoration methods cannot fully exploit priors from low-quality images and still have many challenges in perceptual quality, semantic fidelity, and structure accuracy. Based on the challenges, we introduce a novel image restoration method, SSP-IR. Our approach aims to fully exploit semantic and structure priors from low-quality images to guide the diffusion model in generating semantically faithful and structurally accurate natural restoration results. Specifically, we integrate the visual comprehension capabilities of Multimodal Large Language Models (explicit) and the visual representations of the original image (implicit) to acquire accurate semantic prior. To extract degradation-independent structure prior, we introduce a Processor with RGB and FFT constraints to extract structure prior from the low-quality images, guiding the diffusion model and preventing the generation of unreasonable artifacts. Lastly, we employ a multi-level attention mechanism to integrate the acquired semantic and structure priors. The qualitative and quantitative results demonstrate that our method outperforms other state-of-the-art methods overall on both synthetic and real-world datasets. Our project page ishttps://zyhrainbow.github.io/projects/SSP-IR. Hengsheng Zhang, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Instance-Adaptive Spatial-Temporal Enhancement for Efficient Video CompressionabstractEfficiently compressing HD/UHD content has long been challenging due to high bitrate costs. Instance-adaptive enhancement methods try to tackle this issue by compressing a video at reduced resolution and enhancing it using a neural model specifically overfitted for this video. However, existing methods focus solely on spatial super-resolution (SR) and under-utilize the videos' temporal redundancy. Their limited management of the model's updated parameters also causes excessive overfitting overheads. Therefore, this paper introduces IASTE, the first instance-adaptive enhancement method based on spatial-temporal enhancement (STE), and incorporates low-rank adaptation (LoRA) for efficient model overfitting. Specifically, we downscale videos spatially and temporally to reduce the data volume and achieve efficient video compression. Then, we overfit a specific STE model for each video and use it to enhance the decoded video's spatiotemporal resolution. Leveraging the video swin transformer's strong capability in capturing spatiotemporal correlations, we design a lightweight and efficient model to implement video STE. The model is overfitted for each video using LoRA. By freezing the pre-trained model and selectively updating a few low-rank matrices, the bitrate overhead for model storage can be mitigated. Experiments prove that compared to directly compressing high-frame-rate (HFR), high-resolution (HR) videos, our method achieves around 30% BD-Rate gains on the CTC and UVG datasets, about 15% gains on the YoutubeUGC dataset, and about 10% gains on the ultra-long videos in the Xiph dataset. Yan Zhao 0041, Zhengxue Cheng, Jiangchuan Li, Donghui Feng 0003, Qunshan Gu, Cheems Wang, Guo Lu, Li Song 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Rate-aware Compression for NeRF-based Volumetric VideoabstractThe neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing solutions typically compress these NeRF representations after the training stage, leading to a separation between representation training and compression. In this paper, we try to directly learn a compact NeRF representation for volumetric video in the training stage based on the proposed rate-aware compression framework. Specifically, for volumetric video, we use a simple yet effective modeling strategy to reduce temporal redundancy for the NeRF representation. Then, during the training phase, an implicit entropy model is utilized to estimate the bitrate of the NeRF representation. This entropy model is then encoded into the bitstream to assist in the decoding of the NeRF representation. This approach enables precise bitrate estimation, thereby leading to a compact NeRF representation.Furthermore, we propose an adaptive quantization strategy and learn the optimal quantization step for the NeRF representations. Finally, the NeRF representation can be optimized by using the rate-distortion trade-off. Our proposed compression framework can be used for different representations and experimental results demonstrate that our approach significantly reduces the storage size with marginal distortion and achieves state-of-the-art rate-distortion performance for volumetric video on the HumanRF and ReRF datasets. Compared to the previous state-of-the-art method TeTriRF, we achieved an approximately -80% BD-rate on the HumanRF dataset and -60% BD-rate on the ReRF dataset. Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Zhengxue Cheng, Anni Tang, Li Song 0001 |
ACM Multimedia | 4 |
| 2024 | AsymLLIC: Asymmetric Lightweight Learned Image CompressionabstractLearned image compression (LIC) methods often employ symmetrical encoder and decoder architectures, evitably increasing decoding time. However, practical scenarios demand an asymmetric design, where the decoder requires low complexity to cater to diverse low-end devices, while the encoder can accommodate higher complexity to improve coding performance. In this paper, we propose an asymmetric lightweight learned image compression (AsymLLIC) architecture with a novel training scheme, enabling the gradual substitution of complex decoding modules with simpler ones. Building upon this approach, we conduct a comprehensive comparison of different decoder network structures to strike a better trade-off between complexity and compression performance. Experiment results validate the efficiency of our proposed method, which not only achieves comparable performance to VVC but also offers a lightweight decoder with only 51.47 GMACs computation and 19.65M parameters. Furthermore, this design methodology can be easily applied to any LIC models, enabling the practical deployment of LIC techniques. Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Guo Lu, Li Song 0001, Wenjun Zhang 0001 |
VCIP | 2 |
| 2024 | Coarse-to-fine Transformer For Lossless 3D Medical Image CompressionabstractThe rapid advancements in medical imaging have led to a growing demand for high-performance lossless compression of large 3D medical image datasets. Unlike natural images, medical images typically feature three-dimensional structures, and high bit-depth, necessitating specialized compression techniques. Based on a decoder-only transformer, we propose a learnable dual-decoder model for lossless compression of 3D medical images. Our approach packs voxels into patches, which are processed by a patch-level decoder to extract the patch feature. The voxels, along with the patch feature, are subsequently fed into a voxel-level decoder to model each voxel. This coarse-to-fine modeling strategy reduces the computational time for each voxel and enables long-range modeling dependencies. Experimental results demonstrate that our proposed model achieves state-of-the-art compression performance, with an approximately 15% improvement in compression performance over the traditional JP3D benchmark on various datasets. Guo Lu, Donghui Feng 0003, Zhengxue Cheng, Guosheng Yu, Li Song 0001 |
VCIP | 4 |
| 2024 | Efficient Bitrate Ladder Construction for Per-Shot Adaptive EncodingabstractHTTP adaptive streaming (HAS) constructs bitrate ladders to deliver videos with the best possible quality under varying network conditions. Though per-shot content adaptive encoding (CAE) largely improves the compression efficiency by constructing the optimal bitrate ladder for each video shot, it suffers from excessive encoding complexity as all the points in the operating space (typically resolution × bitrate) need to be encoded and compared. To address this issue, this paper proposes an efficient bitrate ladder construction method that encodes only a subset of operating points, then uses curve fitting and inter-curve prediction to estimate other points’ RD performance. The proposed method enables low-complexity ladder construction even for high-dimension operating spaces that incorporate dimensions like encoding presets. Experiments show that this method can achieve RD performance comparable to the original per-shot CAE with only 42% encoding points. Even when minimizing the encoding points to 3.6% of the original CAE, it achieves 15% BD-Rate improvements compared to using the fixed bitrate ladder. Yan Zhao 0041, Zhengxue Cheng, Guo Lu, Rong Xie 0004, Li Song 0001 |
VCIP | 2 |
| 2020 | Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention ModulesabstractImage compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
CVPR | 1 |
| 2020 | Learned Lossless Image Compression with A Hyperprior and Discretized Gaussian Mixture LikelihoodsabstractLossless image compression is an important task in the field of multimedia communication. Traditional image codecs typically support lossless mode, such as WebP, JPEG2000, FLIF. Recently, deep learning based approaches have started to show the potential at this point. HyperPrior is an effective technique proposed for lossy image compression. This paper generalizes the hyperprior from lossy model to lossless compression, and proposes a L2-norm term into the loss function to speed up training procedure. Besides, this paper also investigated different parameterized models for latent codes, and propose to use Gaussian mixture likelihoods to achieve adaptive and flexible context models. Experimental results validate our method can outperform existing deep learning based lossless compression, and outperform the JPEG2000 and WebP for JPG images. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
ICASSP | 1 |
| 2020 | Scalable Learned Image Compression With A Recurrent Neural Networks-Based HyperpriorabstractRecently learned image compression has achieved many great progresses, such as representative hyperprior and its variants based on convolutional neural networks (CNNs). However, CNNs are not fit for scalable coding and multiple models need to be trained separately to achieve variable rates. In this paper, we incorporate differentiable quantization and accurate entropy models into recurrent neural networks (RNNs) architectures to achieve a scalable learned image compression. First, we present an RNN architecture with quantization and entropy coding. To realize the scalable coding, we allocate the bits to multiple layers, by adjusting the layer-wise lambda values in Lagrangian multiplier-based rate-distortion optimization function. Second, we add an RNN-based hyperprior to improve the accuracy of entropy models for multiple-layer residual representations. Experimental results demonstrate that our performance can be comparable with recent CNN-based hyperprior methods on Kodak dataset. Besides, our method is a scalable and flexible coding approach, to achieve multiple rates using one single model, which is very appealing. Rige Su, Zhengxue Cheng, Heming Sun, Jiro Katto |
ICIP | 2 |
| 2020 | End-To-End Learned Image Compression With Fixed Point Weight QuantizationabstractLearned image compression (LIC) has reached the traditional hand-crafted methods such as JPEG2000 and BPG in terms of the coding gain. However, the large model size of the network prohibits the usage of LIC on resource-limited embedded systems. This paper presents a LIC with 8-bit fixed-point weights. First, we quantize the weights in groups and propose a non-linear memory-free codebook. Second, we explore the optimal grouping and quantization scheme. Finally, we develop a novel weight clipping fine tuning scheme. Experimental results illustrate that the coding loss caused by the quantization is small, while around 75% model size can be reduced compared with the 32-bit floating-point anchor. As far as we know, this is the first work to explore and evaluate the LIC fully with fixed-point weights, and our proposed quantized LIC is able to outperform BPG in terms of MS-SSIM. Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
ICIP | 2 |
| 2020 | Energy Compaction-Based Image Compression Using Convolutional AutoEncoderabstractImage compression has been an important research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and its use in image compression has gradually been increasing. In this paper, we present an energy compaction-based image compression architecture using a convolutional autoencoder (CAE) to achieve high coding efficiency. Our main contributions include three aspects: 1) we propose a CAE architecture for image compression by decomposing it into several down(up)sampling operations; 2) for our CAE architecture, we offer a mathematical analysis on the energy compaction property and we are the first work to propose a normalized coding gain metric in neural networks, which can act as a measurement of compression capability; 3) based on the coding gain metric, we propose an energy compaction-based bit allocation method, which adds a regularizer to the loss function during the training stage to help the CAE maximize the coding gain and achieve high compression efficiency. The experimental results demonstrate our proposed method outperforms BPG (HEVC-intra), in terms of the MS-SSIM quality metric. Additionally, we achieve better performance in comparison with existing bit allocation methods, and provide higher coding efficiency compared with state-of-the-art learning compression methods at high bit rates. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
IEEE Trans. Multim. | 1 |
| 2020 | Enhanced Intra Prediction for Video Coding by Using Multiple Neural NetworksabstractThis paper enhances the intra prediction by using multiple neural network modes (NM). Each NM serves as an end-to-end mapping from the neighboring reference blocks to the current coding block. For the provided NMs, we present two schemes (appending and substitution) to integrate the NMs with the traditional modes (TM) defined in high efficiency video coding (HEVC). For the appending scheme, each NM is corresponding to a certain range of TMs. The categorization of TMs is based on the expected prediction errors. After determining the relevant TMs for each NM, we present a probability-aware mode signaling scheme. The NMs with higher probabilities to be the best mode are signaled with fewer bits. For the substitution scheme, we propose to replace the highest and lowest probable TMs. New most probable mode (MPM) generation method is also employed when substituting the lowest probable TMs. Experimental results demonstrate that using multiple NMs will improve the coding efficiency apparently compared with the single NM. Specifically, proposed appending scheme with seven NMs can save 2.6%, 3.8%, and 3.1% BD-rate for Y, U, and V components compared with using single NM in the state-of-the-art works. Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
IEEE Trans. Multim. | 2 |
| 2019 | Learning Image and Video Compression Through Spatial-Temporal Energy CompactionabstractCompression has been an important research topic for many decades, to produce a significant impact on data transmission and storage. Recent advances have shown a great potential of learning based image and video compression. Inspired from related works, in this paper, we present an image compression architecture using a convolutional autoencoder, and then generalize image compression to video compression, by adding an interpolation loop into both encoder and decoder sides. Our basic idea is to realize spatial-temporal energy compaction in learning image and video compression. Thereby, we propose to add a spatial energy compaction-based penalty into loss function, to achieve higher image compression performance. Furthermore, based on temporal energy distribution, we propose to select the number of frames in one interpolation loop, adapting to the motion characteristics of video contents. Experimental results demonstrate that our proposed image compression outperforms the latest image compression standard with MS-SSIM quality metric, and provides higher performance compared with state-of-the-art learning compression methods at high bit rates, which benefits from our spatial energy compaction approach. Meanwhile, our proposed video compression approach with temporal energy compaction can significantly outperform MPEG-4, and is competitive with commonly used H.264. Both our image and video compression can produce more visually pleasant results than traditional standards. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
CVPR | 1 |
| 2019 | Perceptual Quality Study on Deep Learning Based Image CompressionabstractRecently deep learning based image compression has made rapid advances with promising results based on objective quality metrics. However, a rigorous subjective quality evaluation on such compression schemes have rarely been reported. This paper aims at perceptual quality studies on learned compression. First, we build a general learned compression approach, and optimize the model. In total six compression algorithms are considered for this study. Then, we perform subjective quality tests in a controlled environment using high-resolution images. Results demonstrate learned compression optimized by MS-SSIM yields competitive results that approach the efficiency of state-of-the-art compression. The results obtained can provide a useful benchmark for future developments in learned image compression. Zhengxue Cheng, Pinar Akyazi, Heming Sun, Jiro Katto, Touradj Ebrahimi |
ICIP | 1 |
| 2019 | Dual Learning-based Video Coding with Inception Dense BlocksabstractIn this paper, a dual learning-based method in intra coding is introduced for PCS Grand Challenge. This method is mainly composed of two parts: intra prediction and reconstruction filtering. They use different network structures, the neural network-based intra prediction uses the full-connected network to predict the block while the neural network-based reconstruction filtering utilizes the convolutional networks. Different with the previous filtering works, we use a network with more powerful feature extraction capabilities in our reconstruction filtering network. And the filtering unit is the block-level so as to achieve a more accurate filtering compensation. To our best knowledge, among all the learning-based methods, this is the first attempt to combine two different networks in one application, and we achieve the state-of-the-art performance for AI configuration on the HEVC Test sequences. The experimental result shows that our method leads to significant BD-rate saving for provided 8 sequences compared to HM-16.20 baseline (average 10.24% and 3.57% bitrate reductions for all-intra and random-access coding, respectively). For HEVC test sequences, our model also achieved a 9.70% BD-rate saving compared to HM-16.20 baseline for all-intra configuration. Chao Liu 0027, Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
PCS | 4 |
| 2018 | Light-Weight Video Coding Based on Perceptual Video Quality for Live StreamingabstractIn video streaming on the internet, effective encoding recipes (i.e. bitrate-resolution pairs) are a main obstacle to deliver high-quality video streams. We developed a method to generate an encoding recipe that considers subjective visual quality with one just-noticeable difference (JND) distance. However, this method requires excessive computation time, which is not directly applicable for live streaming. In this paper, in order to provide a light-weight method for live streaming, we developed three acceleration techniques: resolution extrapolation, VMAF skipping and sampled objective measure calculation. These techniques are heuristic, but greatly contribute to reducing computational cost. Experimental results demonstrate that the proposed method achieves a significant reduction in computation time without significant effects on rate-JND characteristics. Yusuke Sakamoto, Shintaro Saika, Masaru Takeuchi, Tatsuya Nagashima, Zhengxue Cheng, Kenji Kanai, Jiro Katto, Kaijin Wei, Ju Zengwei |
ISM | 5 |
| 2018 | Deep Convolutional AutoEncoder-based Lossy Image CompressionabstractImage compression has been investigated as a fundamental research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and is gradually being used in image compression. In this paper, we present a lossy image compression architecture, which utilizes the advantages of convolutional autoencoder (CAE) to achieve a high coding efficiency. First, we design a novel CAE architecture to replace the conventional transforms and train this CAE using a rate-distortion loss function. Second, to generate a more energy-compact representation, we utilize the principal components analysis (PCA) to rotate the feature maps produced by the CAE, and then apply the quantization and entropy coder to generate the codes. Experimental results demonstrate that our method outperforms traditional image coding algorithms, by achieving a 13.7% BD-rate decrement on the Kodak database images compared to JPEG2000. Besides, our method maintains a moderate complexity similar to JPEG2000. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
PCS | 1 |
| 2018 | Perceptual Quality Driven Adaptive Video Coding Using JND EstimationabstractWe introduce a perceptual video quality driven video encoding solution for optimized adaptive streaming. By using multiple bitrate/resolution encoding like MPEG-DASH, video streaming services can deliver the best video stream to a client, under the conditions of the client's available bandwidth and viewing device capability. However, conventional fixed encoding recipes (i.e., resolution-bitrate pairs) suffer from many problems, such as improper resolution selection and stream redundancy. To avoid these problems, we propose a novel video coding method, which generates multiple representations with constant Just-Noticeable Difference (JND) interval. For this purpose, we developed a JND scale estimator using Support Vector Regression (SVR), and designed a pre-encoder which outputs an encoding recipe with constant JND interval in an adaptive manner to input video. Masaru Takeuchi, Shintaro Saika, Yusuke Sakamoto, Tatsuya Nagashima, Zhengxue Cheng, Kenji Kanai, Jiro Katto, Kaijin Wei, Ju Zengwei |
PCS | 5 |
| 2017 | A low-cost approximate 32-point transform architectureabstractThis paper presents an area-efficient approximate method for 32-point transform which is one of the most area-consuming parts in High Efficiency Video Coding (HEVC) applications. Compared to prior literatures, this work reduces the hardware cost of transform by 1) eliminating all the arithmetic operations of 6 least significant bits (LSB), 2) presenting a low-delay method for generating carry propagation from the remaining 5 LSBs and 3) truncating the most significant bits (MSB) according to the position of component. In the implementation of a 32-point forward transform, the experimental results show that 27% area consumption can be saved and the coding efficiency loss aroused by the approximation is only 0.044% compared with the origin. Heming Sun, Zhengxue Cheng, Amir Masoud Gharehbaghi, Shinji Kimura, Masahiro Fujita 0004 |
ISCAS | 2 |
| 2017 | A Pre-Saliency Map Based Blind Image Quality Assessment via Convolutional Neural NetworksabstractIn recent years, various approaches have been investigated towards blind image quality assessment (IQA) with high accuracy and low complexity. In this paper we develop a pre-saliency map based blind IQA method, which takes advantage of saliency information in prior of quality prediction for performance enhancement by two steps. 1) We split the image into patches and design a convolution neural network (CNN) to predict the patch-wise quality score. Then we explore the relation between image saliency information and CNN prediction error to present a statistical analysis. 2) Based on the analysis, we propose a patch quality aggregation algorithm by removing non-salient patches which are likely to bring large prediction error and assigning large weights for salient patches. Experimental results validate that our method can achieve high accuracy (0.978) with subjective quality scores, which outperforms existing IQA methods. Meanwhile, the proposed method can reduce 52.7% computational time than the IQA without pre-saliency map. Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
ISM | 1 |
| 2015 | Merge mode based fast inter prediction for HEVCabstractThe latest High Efficiency Video Coding (HEVC/H.265) obtains 50% bit rate reduction than H.264/AVC standard with comparable quality, but at the cost of high computational complexity. Inter prediction accounts for large complexity and merge mode is one of the most important new features introduced in HEVC. To address this issue, this paper utilizes the merge mode to accelerate inter prediction by three fast mode decision methods. 1) A merge candidate decision is proposed to select the best merge mode by Sum of Absolute Transformed Difference (SATD) cost to reduce the merge time. 2) An early merge termination is presented still based on SATD cost with more than 90% accuracy. 3) Based on efficient merge mode, symmetric motion partition (SMP) modes can be disabled for non-8 × 8 code units (CUs). Experimental results demonstrate that our work can achieve 53.1%-54.2% time reduction on average with 1.57%-2.30% BD-rate increment. Besides, our method achieves an improvement of 18%-30% time reduction with 0.89%-2.85% BD-rate increment when combined with other existing approaches. Zhengxue Cheng, Heming Sun, Dajiang Zhou, Shinji Kimura |
VCIP | 1 |