Peilin Chen 0001

dblp:193/9238-1 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0001-6636-522XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 27 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing
abstract
Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on the critical challenge of how to restore surrogate-driven protected data in diverse MLLM scenarios. We first bridge this research gap by contributing the SPPE (Surrogate Privacy Protected Editable) dataset, which includes a wide range of privacy categories and user instructions to simulate real MLLM applications. This dataset offers protected surrogates alongside their various MLLM-edited versions, thus enabling the direct assessment of privacy recovery quality. By formulating privacy recovery as a guided generation task conditioned on complementary multimodal signals, we further introduce a unified approach that reliably reconstructs private content while preserving the fidelity of MLLM-generated edits. The experiments on both SPPE and InstructPix2Pix further show that our approach generalizes well across diverse visual content and editing tasks, achieving a strong balance between privacy protection and MLLM usability.
Yibing Liu, Peilin Chen 0001, Yung-Hui Li, Shiqi Wang 0001, Sam Kwong
AAAI3
2026 TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval
abstract
Yuhang Zhang, Keyan Ding, Peilin Chen, Han Liu, Can Lin, Ruixi Chen, Shiqi Wang, Qi Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuhang Zhang 0030, Keyan Ding, Peilin Chen 0001, Can Lin, Ruixi Chen, Shiqi Wang 0001, Qi Song 0004
ACL (1)3
2026 DRFC: An End-to-End Deep Dynamic RF Signal Compression Framework
abstract
Radio frequency (RF) signals have gained widespread adoption in intelligent perception systems due to their unique advantages, including non-line-of-sight propagation capability, robustness in low-light environments, and inherent privacy preservation. However, their substantial data volumes, generated by the dual-polarization direction characteristic, result in significant challenges to data storage and transmission. To address this, we propose the first end-to-end deep dynamic RF signal compression (DRFC) framework, which primarily focuses on exploiting cross-directional correlation in dynamic RF signals. The proposed framework incorporates four key innovations: (1) a mask-guided RF motion estimation module that leverages Doppler shifts and electromagnetic noise characteristics to identify regions of significant motion using a threshold-based mask, significantly improving motion estimation accuracy; (2) a cross-directional RF motion entropy model that utilizes cross-directional RF motion latent priors to refine the probability distribution for motion entropy coding; (3) a cross-directional RF context mining module that predicts RF contexts from temporal and cross-directional reference signals, adaptively fusing these contexts with confidence maps to maximize complementary information utilization; and (4) a cross-directional RF contextual entropy model that incorporates cross-directional RF contextual latent priors to optimize contextual entropy modeling. Experimental results demonstrate the superiority of our framework over existing codecs. Our DRFC framework achieves significant bitrate savings on benchmark datasets, establishing a strong baseline for future research in this field.
Xihua Sheng, Peilin Chen 0001, Shiqi Wang 0001, Dapeng Oliver Wu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Optimizing Fidelity-Perception Tradeoff via Large Vision-Language Model Prior for Image Compression
abstract
Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity.
Yudong Mao, Peilin Chen 0001, Lingyu Zhu 0006, Yung-Hui Li, Shiqi Wang 0001
IEEE Trans. Image Process.2
2026 Unfolding High-Order Correlations for Interpretable Multi-Contrast MRI Super-Resolution
abstract
Deep unfolding network has gained significant attention for magnetic resonance imaging super-resolution (MRI SR) due to its performance and interpretability. However, 1) existing methods predominantly focus on cross-contrast correlations while neglecting high-order correlations embedded within spatially adjacent slices in volumetric MRI data. 2) Their degradation models are optimized via the proximal gradient algorithm (PGA) that relies on manually designed hyperparameters (e.g., step size), often leading to overshooting or suboptimal solutions. To solve these limitations, we propose HocMRI, a deep unfolding multi-contrast MRI SR framework, which seamlessly integrates dual-prior modeling and hyperparameter-free PGA for enhanced reconstruction. Specifically, we first design a novel degradation model based on the dual-prior mechanism: an explicit prior based on low-rank tensor factorization to capture intra- and inter-slice dependencies, and an implicit prior leveraging a Mamba-based network with a novel 3D scanning strategy to further exploit high-order correlations across slices. Then, we derive a hyperparameter-free PGA to boost the traditional PGA, which employs a hyperbolic tangent function to dynamically control the gradient descent step, eliminating manual tuning while ensuring stable convergence with theoretical proofs. Based on the hyperparameter-free PGA, we develop an efficient iterative optimization algorithm to solve the degradation model and unfold it into a multi-stage deep network. Numerous experimental results from widely used MRI datasets demonstrate that our HocMRI achieves superior performance with enhanced efficiency compared to the state-of-the-art methods.
Qiangqiang Shen, Xuanqi Zhang, Peilin Chen 0001, Zhiwei Zhong 0001, Howard Leung, Shiqi Wang 0001
IEEE Trans. Image Process.3
2026 Rate Control for 360$^{\circ }$ Versatile Video Coding Based on Visual Gaze Mechanism
abstract
In the past few years, 360° video has started to infiltrate various aspects of daily life. Although there have been significant developments in 360° video coding technology, understanding of the human visual gaze mechanism has been somewhat overlooked. In this paper, we propose a rate control scheme for 360° Versatile Video Coding (VVC) based on a human visual gaze mechanism, targeting at improving the coding performance and bitrate accuracy. More specifically, based on the Equi-rectangular Projection (ERP) format, latitude information is systematically analyzed and a stripe-level bit allocation scheme is established, to better mitigate the projection distortion. Subsequently, the Lagrange parameter λ is further optimized with distortion dependency and identification of the visual gaze guided key Coding Tree Units (CTUs). The proposed rate control scheme is implemented on the VVC Test Model for 360° video. Experimental results show that the proposed rate control scheme can achieve BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) under the various configurations, respectively. Meanwhile, a healthier buffer status and better visual quality can be observed, further demonstrating the advantages of the proposed scheme.
Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Peilin Chen 0001, Xiaohai He, Shiqi Wang 0001
IEEE Trans. Multim.4
2025 Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration
Yudong Mao, Zhiwei Zhong 0001, Peilin Chen 0001, Zhijiang Zhang, Shiqi Wang 0001
CVPR4
2025 Compact Feature Representation in Bird View for V2X Communication-Efficient Collaborative Analysis
abstract
Sensor data analysis is a crucial task for environmental cognition in smart traffic systems. Recently, vehicle-to-everything (V2X) collaborative analysis has leveraged intermediate feature communication between vehicles and infrastructure to achieve superior analysis performance compared to single-vehicle approaches. However, due to the limited bandwidth of V2X communication links, directly transmitting features can be inefficient, resulting in significant delays that are unacceptable for real-time decision-making. To address this challenge, we propose a compact feature representation method in the bird's eye view (BEV) space for communication-efficient collaborative analysis. As shown in Fig. 1, the proposed method can be viewed as a task-aware distributed coding approach with decoder side information. First, the ego vehicle and the networked infrastructure convert raw LiDAR data into BEV features using a shared PointPillars feature extractor. The infrastructure then applies the proposed BEV codec to transform these BEV features into a compact representation, encoding them into a binary bitstream through entropy coding based on the estimated distribution. The received features are subsequently warped and fused with the ego vehicle's features using a bidirectional attention fusion module, and processed by a single-shot detector to perform 3D object detection. Experimental results on the DAIR-V2X-C dataset demonstrate that the proposed framework achieves more than 1000 times compression compared to directly transmitting floating-point features, while maintaining high analysis performance in real-world V2X scenarios.
Linfeng Zheng, Peilin Chen 0001, Shiqi Wang 0001, Dapeng Oliver Wu
DCC2
2025 An Information-Theoretic Regularizer for Lossy Neural Image Compression
abstract
Lossy image compression networks aim to minimize the latent entropy of images while adhering to specific distortion constraints. However, optimizing the neural network can be challenging due to its nature of learning quantized latent representations. In this paper, our key finding is that minimizing the latent entropy is, to some extent, equivalent to maximizing the conditional source entropy, an insight that is deeply rooted in information-theoretic equalities. Building on this insight, we propose a novel structural regularization method for the neural image compression task by incorporating the negative conditional source entropy into the training objective, such that both the optimization efficacy and the model's generalization ability can be promoted. The proposed information-theoretic regularizer is interpretable, plug-and-play, and imposes no inference overheads. Extensive experiments demonstrate its superiority in regularizing the models and further squeezing bits from the latent representation across various compression structures and unseen domains.
Yingwen Zhang, Meng Wang 0017, Xihua Sheng, Peilin Chen 0001, Li Zhang 0006, Shiqi Wang 0001
ICCV4
2025 Efficient Image Compression through Extreme Image Rescaling
abstract
In this paper, we propose a generative image compression scheme for extremely low bitrate representation and high visual quality reconstruction. This method decomposes images into ultra-low-resolution thumbnails and text descriptions, achieving high compression rates while maintaining human-perceptible thumbnails for better previewing and understanding. To this end, we integrate an arbitrary-scale image rescaling model with a pre-trained conditional diffusion model, enhancing both rescaling flexibility and visual quality. Specifically, the high-resolution image is downscaled into a thumbnail for transmission or storage, then decoded by upscaling it to its original resolution, followed by a diffusion-based generative process for quality enhancement. To better utilize the generative priors of the pretrained diffusion model, the upscaled images are aligned with the original input in the latent space of the diffusion model. Leveraging these generative priors, thumbnails at extreme scales can be reconstructed to their original resolution with high fidelity and perceptual quality. Additionally, text descriptions extracted from the original image are used to condition the diffusion model, improving semantic consistency in the reconstruction. Extensive experimental results demonstrate that our method can achieve notable compression efficiency and visually pleasing reconstruction results at extremely low bitrates.
Jiancong Chen, Peilin Chen 0001, Shiqi Wang 0001, Zhu Li 0001
ISCAS3
2025 Multiple Glancing at Quality: Benchmark Dataset and Objective Quality Assessment Metric for Low-light Image Enhancement
abstract
Quality metrics play a crucial role in guiding the development of image enhancement algorithms, which have consistently sought effective quality assessment methodologies and comprehensive datasets. To address this need, we first built a large-scale dataset for low-light enhanced image quality assessment and gathered the corresponding subjective evaluation scores. Recently, vision-language pre-training models have demonstrated considerable potential in the realm of quality assessment. However, its efficacy is limited by the fine-grained perception in low-level quality assessment. As such, we further propose a novel quality assessment framework using contrastive prompt learning, which harnesses the robust priors of vision-language pre-training models to improve the perceptual capacity of deep networks for low-level quality features. Experiments on the proposed RSLE dataset show that our method outperforms existing SOTA image quality assessment methods. Our database and the source code will be made publicly available.
Yudong Mao, Peilin Chen 0001, Zhao Wang 0004, Qiuping Jiang, Shiqi Wang 0001
ISCAS2
2025 Learning Spatio-Temporal Resolutions for Deep Video Compression
abstract
We propose a spatio-temporal adaptive deep video compression scheme, which is capable of intelligently adjusting the spatial resolution and temporal frame rate for content adaptive compression, with the aim of pursuing enhanced rate-distortion performance. In particular, a neural network-based spatio-temporal adaptation network is integrated into the deep video coding paradigm, enabling the adaptive determination of the optimal rescaling ratios for compression, leading to the further reduction of spatial and temporal redundancies. Moreover, learning-based modules for rescaling parameter determination are incorporated into the spatio-temporal adaptation network. The proposed scheme can be easily plugged into, and seamlessly collaborate with the existing deep video coding frameworks. Experimental results demonstrate that, compared to the original neural video codecs, the proposed method achieves significant bitrate savings in terms of both PSNR and MS-SSIM.
Jiancong Chen, Meng Wang 0017, Peilin Chen 0001, Shiqi Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 MPSol: A Multimodal Prompt Learning Framework for Protein Solubility Prediction
abstract
Protein solubility is a critical determinant of biologic candidates' developability, stability, and therapeutic efficacy. However, accurate solubility prediction remains a central challenge in computational protein engineering due to the inherent complexity within protein sequences. In this work, we propose a multimodal prompt learning framework, called MPSol, for protein solubility prediction that integrates complementary representations derived from primary sequences, structural proxies, and textual descriptions generated by large language models (LLMs). MPSol is built upon a unified multimodal backbone with a dedicated cross-modal fusion module that captures fine-grained interactions across modalities. In addition, we design label-aware prompts that encode solubility-specific semantic cues associated with each class. These prompts provide semantic supervision, guiding the alignment of fused protein representations to promote semantic consistency. Extensive experiments demonstrate that MPSol achieves state-of-the-art performance, reaching an accuracy of 0.815, AUC of 0.867 and MCC of 0.642 on the standard PDBSol test set, and generalizes well to the external out-of-distribution test dataset with an accuracy of 0.632, AUC of 0.653 and MCC of 0.332. These results underscore the potential of prompt-driven multimodal learning for interpretable and effective protein property prediction.
Yuhang Zhang 0030, Peilin Chen 0001, Keyan Ding, Shiqi Wang 0001, Qi Song 0004
IEEE J. Biomed. Health Informatics2
2025 RCNet: Deep Recurrent Collaborative Network for Multi-View Low-Light Image Enhancement
abstract
Scene observation from multiple perspectives would bring a more comprehensive visual experience. However, in the context of acquiring multiple views in the dark, the highly correlated views are seriously alienated, making it challenging to improve scene understanding with auxiliary views. Recent single image-based enhancement methods may not be able to provide consistently desirable restoration performance for all views due to the ignorance of potential feature correspondence among different views. To alleviate this issue, we make the first attempt to investigate multi-view low-light image enhancement. First, we construct a new dataset called Multi-View Low-light Triplets (MVLT), including 1,860 pairs of triple images with large illumination ranges and wide noise distribution. Each triplet is equipped with three different viewpoints towards the same scene. Second, we propose a deep multi-view enhancement framework based on the Recurrent Collaborative Network (RCNet). Specifically, in order to benefit from similar texture correspondence across different views, we design the recurrent feature enhancement, alignment and fusion (ReEAF) module, in which intra-view feature enhancement (Intra-view EN) followed by inter-view feature alignment and fusion (Inter-view AF) is performed to model the intra-view and inter-view feature propagation sequentially via multi-view collaboration. In addition, two different modules from enhancement to alignment (E2A) and from alignment to enhancement (A2E) are developed to enable the interactions between Intra-view EN and Inter-view AF, which explicitly utilize attentive feature weighting and sampling for enhancement and alignment, respectively. Experimental results demonstrate that our RCNet significantly outperforms other state-of-the-art methods. All of our dataset, code, and model will be available athttps://github.com/hluo29/RCNet.
Baoliang Chen, Lingyu Zhu 0006, Peilin Chen 0001, Shiqi Wang 0001
IEEE Trans. Multim.4
2025 Comprehensive Action Quality Assessment Through Multi-Branch Modeling
abstract
Action Quality Assessment (AQA) aims to evaluate and score human actions in videos accurately. Existing approaches involve extracting features from the input video and implementing regression based on those features. However, representations derived from a single branch often lack the necessary diversity and flexibility to capture the complexity of human actions effectively. This work addresses these limitations by introducing a multi-branch architecture designed to capture a broad spectrum of video dynamics at varying levels of granularity. Specifically, we enhance video representation in the flow-guided branch by integrating optical flow with video features. This combination of multimodal features offers a more comprehensive context of global motion. Meanwhile, the moment-focused branch is tailored to extract frame-specific features, constructing two distinct quality-based representations with different focuses on moments, which achieves adaptive clues aggregation. Furthermore, the detail-aware branch leverages multiscale deep embeddings from a hierarchy convolutional neural network to capture fine-grained spatial information, which is useful when objects have complex spatial changes. Finally, a post-fusion strategy is employed to merge outputs from all branches, contributing to the comprehensive action quality assessment. Experimental evaluations on three benchmark datasets, FineDiving, MTL-AQA, and AQA-7, demonstrate the superiority of our model in providing reliable assessments of action quality.
Peilin Chen 0001, Meng Wang 0017, Shiqi Wang 0001, Hong Yan 0001, Sam Kwong
IEEE Trans. Multim.2
2025 HNR-ISC: Hybrid Neural Representation for Image Set Compression
abstract
Image set compression (ISC) refers to compressing the sets of semantically similar images. Traditional ISC methods typically aim to eliminate redundancy among images at either signal or frequency domain, but often struggle to handle complex geometric deformations across different images effectively. Here, we propose a new Hybrid Neural Representation for ISC (HNR-ISC), including an implicit neural representation for Semantically Common content Compression (SCC) and an explicit neural representation for Semantically Unique content Compression (SUC). Specifically, SCC enables the conversion of semantically common contents into a small-and-sweet neural representation, along with embeddings that can be conveyed as a bitstream. SUC is composed of invertible modules for removing intra-image redundancies. The feature level combination from SCC and SUC naturally forms the final image set. Experimental results demonstrate the robustness and generalization capability of HNR-ISC in terms of signal and perceptual quality for reconstruction and accuracy for the downstream analysis task.
Shiqi Wang 0001, Meng Wang 0017, Peilin Chen 0001, Wenhui Wu 0001, Xu Wang 0006, Sam Kwong
IEEE Trans. Multim.4
2025 EIN: Exposure-Induced Network for Single-Image HDR Reconstruction
abstract
Reconstructing high dynamic range (HDR) images from standard dynamic range (SDR) ones has received growing attention in recent years. A predominant problem of this task lies in the absence of texture and structural information in under/over-exposed regions. In this article, we propose an efficient and stable single-image HDR reconstruction method, namely exposure-induced network (EIN). More specifically, a dynamic range expansion branch (DB) is designed to expand the global dynamic range of the input SDR image. Moreover, two exposure-gated detail recovering branches for local over- (OB) and under- (UB) exposed regions are proposed to interact with the DB to progressively infer the texture and structural details with the learned confidence maps to resolve challenging ambiguities in such regions. The features from these three interactional branches are adaptively fused in the joint global–local decoder to reconstruct the final HDR image. The proposed network is trained based upon a large-scale dataset constructed with diverse content. Extensive experimental results demonstrate that the proposed model achieves consistent visual quality improvement for input SDR images with different exposures compared with state-of-the-art methods. The source code is available at: https://github.com/Yliu724/EIN .
Zhangkai Ni, Peilin Chen 0001, Shiqi Wang 0001, Xinfeng Zhang 0001, Hanli Wang, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Learned Image Compression for Both Humans and Machines via Dynamic Adaptation
abstract
Recent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs.
Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001
ICIP4
2024 Generative Visual Compression: A Review
abstract
Artificial Intelligence Generated Content (AIGC) is leading a new technical revolution for the acquisition of digital content and impelling the progress of visual compression towards competitive performance gains and diverse functionalities over traditional codecs. This paper provides a thorough review on the recent advances of generative visual compression, illustrating great potentials and promising applications in ultra-low bitrate communication, user-specified reconstruction/filtering, and intelligent machine analysis. In particular, we review the visual data compression methodologies with deep generative models, and summarize how compact representation and high-quality reconstruction could be actualized via generative techniques. In addition, we generalize related generative compression technologies for machine vision with different-domain analysis. Finally, we discuss the fundamental challenges on generative visual compression techniques and envision their future research directions.
Shanzhi Yin, Peilin Chen 0001, Shiqi Wang 0001, Yan Ye 0003
ICIP3
2024 Reveal Fluidity Behind Frames: A Multi-Modality Framework for Action Quality Assessment
abstract
Assessing the quality of a player's performance, such as in diving events, requires precise measurement of subtle action details and overall fluidity. Existing methods primarily utilize appearance information from RGB frames, often neglecting crucial motion information that could contribute to a more comprehensive assessment. In response to this limitation, this paper introduces a novel Multi-Modality Network for Action Quality Assessment (AQA). The proposed method first employs a self-attention based module to foster interaction between optical flow and appearance clues, facilitating the extraction of discriminative features from each modality. Subsequently, a pairwise cross-attention mechanism is designed to comprehensively capture subtle differences via both intra-modality and inter-modality relationships between the query and exemplar video. Finally, to enhance the robustness and achieve accurate score prediction, an adaptive clip aggregation module is introduced to weigh the reliability of each patch based on multi-modal difference features. Experimental results on two benchmarks, FineDiving and MTL-AQA, validate the effectiveness of the proposed model.
Peilin Chen 0001, Meng Wang 0017, Shiqi Wang 0001, Sam Kwong
MMSP2
2024 Exploiting Bidirectional Quality Impulse for Reference Picture Resampled Gaming Video Coding
abstract
Recent years have witnessed a variety of applications of gaming video coding, while how to improve the coding efficiency has been relatively under-explored. The state-of-the-art video coding standard, Versatile Video Coding (VVC), adopts the Reference Picture Resampling (RPR) which allows the variation of the frame resolutions in encoding/decoding. The great flexibility supported by RPR motivates us to develop a bidirectional quality impulse based guidance scheme, in an effort to fully exploit the potential of RPR in gaming video coding. The design philosophy involves reducing the data volume on the encoder side through selective downsampling, and enhancing reconstruction by harnessing quality conveyance from neighboring frames. More specifically, a new RPR structure is developed based on the underlying philosophy that the periodic quality impulse could promisingly boost the quality of the whole sequence. On top of the developed structure, we propose a bidirectional guidance model that faithfully enhances video quality by resorting to frames with quality impulse. Experimental results exhibit the proposed scheme can achieve significant bit-rate savings for gaming videos.
Xiaohan Fang, Peilin Chen 0001, Meng Wang 0017, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Revisiting All-Zero Block Detection for Versatile Video Coding
abstract
The Versatile Video Coding (VVC) standard adopts a series of new coding tools in transform and quantization, including multiple transform selection, low-frequency non-separable transform, and trellis quantization. These new technologies, which bring significant coding gain, create daunting challenges to optimizing the VVC codec. In this work, we propose a new all-zero block (AZB) detection scheme tailored for VVC, with the collaboration of genuine all-zero block (GAZB) and pseudo all-zero block (PAZB) detection. First, to accommodate the multiple transform sizes in VVC, we develop a GAZB detection method that is apt for square and non-square residual blocks. Meanwhile, a theoretical upper bound is derived to locate the last significant coefficient and detect the potential frequency domain GAZB. Subsequently, a method tailored for trellis-coded quantization in VVC is devised for detecting PAZB. Finally, the GAZB and PAZB detection methods are collaboratively employed for AZB detection in VVC. The proposed method is implemented on the VVC codec Versatile Video Encoder (VVenC), and extensive experimental results show that the proposed method achieves promising time savings for test sequences of different resolutions with negligible rate-distortion performance loss.
Zhenhao Sun, Meng Wang 0017, Peilin Chen 0001, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2024 Geometric Prior Based Deep Human Point Cloud Geometry Compression
abstract
The emergence of digital avatars has prompted an exponential increase in the demand for human point clouds with realistic and intricate details. The compression of such data becomes challenging due to massive amounts of data comprising millions of points. Herein, we leverage the human geometric prior in the geometry redundancy removal of point clouds to greatly promote compression performance. More specifically, the prior provides topological constraints as geometry initialization, allowing adaptive adjustments with a compact parameter set that can be represented with only a few bits. Therefore, we propose representing high-resolution human point clouds as a combination of a geometric prior and structural deviations. The prior is first derived with an aligned point cloud. Subsequently, the difference in features is compressed into a compact latent code. The proposed framework can operate in a plug-and-play fashion with existing learning-based point cloud compression methods. Extensive experimental results show that our approach significantly improves the compression performance without deteriorating the quality, demonstrating its promise in serving a variety of applications.
Xinju Wu, Meng Wang 0017, Peilin Chen 0001, Shiqi Wang 0001, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2024 Video Quality Assessment for Spatio-Temporal Resolution Adaptive Coding
abstract
Spatio-temporal resolution adaptive (STRA) coding has been repeatedly proven to be a promising way to improve coding efficiency and reduce coding complexity. The wide consensus is that the optimal subsampled resolution and frame rate should be governed by so- called generalized rate-distortion performance based on the ultimately perceived distortion. However, it is non-trivial to accurately predict the quality of reconstructed videos due to the fact that the distortion originates from both subsampling and compression. To address this issue, we propose a novel video quality assessment model that is fully aware of the information available in downsampled videos for compression, such as resolution and frame rate. More specifically, the proposed model relies on quality-aware spatial features that are extracted by an image quality fine-tuned backbone. Subsequently, the spatio-temporal quality is modeled based on the transformer encoder, which is adaptive to the downsampling spatial and temporal resolutions. This enables the transformer encoder to produce discriminative features that capture long-range temporal dependencies related to the current context. The quality score, which is the output of the transformer encoder, thus reflects both the influence of the subsampling and compression. We conduct extensive experiments that demonstrate the superiority of the proposed model over state-of-the-art methods on four subsampling and compression video quality datasets. Furthermore, we apply the proposed model to bitrate ladder optimization, leading to a perceptual-aware spatial and temporal downsampling strategy that yields promising bitrate savings. The source codes of the proposed model will be publicly available athttps://github.com/h4nwei/STRA-VQA.
Hanwei Zhu, Baoliang Chen, Lingyu Zhu 0006, Peilin Chen 0001, Linqi Song, Shiqi Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 2AFC Prompting of Large Multimodal Models for Image Quality Assessment
abstract
While abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models (LMMs), their image quality assessment (IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice (2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posteriori estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available athttps://github.com/h4nwei/2AFC-LMMs.
Hanwei Zhu, Xiangjie Sui, Baoliang Chen, Xuelin Liu, Peilin Chen 0001, Yuming Fang 0001, Shiqi Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Occupancy Map Guided Attributes Artifacts Removal for Video-Based Point Cloud Compression
abstract
Point clouds offer realistic 3D representations of objects and scenes at the expense of large data volumes. To represent such data compactly in real-world applications, Video-Based Point Cloud Compression (V-PCC) converts their texture into 2D attributes and occupancy maps before applying lossy video compression. Unfortunately, the coding artifacts introduced in the decoded attribute maps eventually degrade the quality of the reconstructed point cloud, thereby influencing its immersive experience. This article proposes a deep learning-based attribute map enhancement method that fully leverages the occupancy map's guidance. The design philosophy is that the cross-modality guidance from occupancy can be leveraged as critical information to enhance the attribute. Therefore, instead of treating attribute and occupancy as two separate sources of signals, occupancy serves as an indispensable auxiliary, such that the proposed framework explicitly provides the model with abundant clues by conducting local feature modification and global dependencies aggregation. In particular, the proposed framework is compatible with existing V-PCC bitstreams and can be feasibly incorporated into the standardized decoder pipeline. Extensive evaluations show the effectiveness of the proposed framework in attribute enhancement, with equivalently 6.0% Bjontegaard Delta-rate (BD-rate) savings obtained.
Peilin Chen 0001, Shiqi Wang 0001, Zhu Li 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Deep Shape-Texture Statistics for Completely Blind Image Quality Evaluation
abstract
Opinion-Unaware Blind Image Quality Assessment (OU-BIQA) models aim to predict image quality without training on reference images and subjective quality scores. Thereinto, image statistical comparison is a classic paradigm, while the performance is limited by the representation ability of visual descriptors. Deep features as visual descriptors have advanced IQA in recent research, but they are discovered to be highly texture-biased and lack shape-bias. On this basis, we find out that image shape and texture cues respond differently toward distortions, and the absence of either one results in an incomplete image representation. Therefore, to formulate a well-rounded statistical description for images, we utilize the shape-biased and texture-biased deep features produced by Deep Neural Networks (DNNs) simultaneously. More specifically, we design a Shape-Texture Adaptive Fusion (STAF) module to merge shape and texture information, based on which we formulate quality-relevant image statistics. The perceptual quality is quantified by the variant Mahalanobis distance between the inner and outer Deep Shape-Texture Statistics (DSTS), wherein the inner and outer statistics respectively describe the quality fingerprints of the distorted image and natural images. The proposed DSTS delicately utilizes shape-texture statistical relations between different data scales in the deep domain and achieves state-of-the-art (SOTA) quality prediction performance on images with artificial and authentic distortions.
Peilin Chen 0001, Hanwei Zhu, Keyan Ding, Leida Li, Shiqi Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Occupancy Map Guided Attributes Deblocking for Video-based Point Cloud Compression
abstract
Point clouds offer the realistic three-dimensional (3-D) representation of objects or scenes at the expense of high data volume. To compactly represent such data in real-world applications, Video-based Point Cloud Compression (V-PCC) converts them into two-dimensional (2-D) attribute maps before lossy compression. However, the coding artifacts introduced in the decoded attribute maps eventually bring texture degradation in the reconstructed point cloud. In this paper, we propose a deep-learning based attribute map enhancement method by fully leveraging the guidance of the occupancy map in local feature modification and non-local attention for capturing long-range spatial correlations.
Peilin Chen 0001, Shiqi Wang 0001, Zhu Li 0001
DCC1
2023 Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine Perception
abstract
We propose a novel face compression framework that leverages the external priors for joint human and machine perception under ultra-low bitrate scenarios. The proposed framework leverages the semantic richness of face images by representing the faces into sketches and thumbnails, resulting in improved bitrate utility for both human and machine vision. At the decoder side, the framework introduces a two-stage generative reconstruction, which faithfully enhances the reconstructed image via semi-parametric modeling and retrieved guidance from the external database. In particular, this coarse-to-fine strategy also results in improved identity consistency and analysis performance of the reconstructed image. Extensive evaluations of the proposed method have been conducted on the public face dataset by comparing it with end-to-end image compression techniques as well as traditional image compression standards. The experimental results demonstrate the effectiveness of the proposed method via superior perceptual and analytical performance under ultra-low bitrate conditions.
Yudong Mao, Peilin Chen 0001, Shurun Wang, Shiqi Wang 0001, Dapeng Oliver Wu
ACM Multimedia2
2022 Appearance Matters, So Does Audio: Revealing the Hidden Face via Cross-Modality Transfer
abstract
Recently, there has been an exponential increase in the security concerns raised by faking face (e.g., deepfake), which automatically changes the identity with a specifically learned deep generative model. With numerous approaches proposed to identify the fake content, much less work has been dedicated to automatically revealing the authentic one that is originally acquired. Here, we propose a new paradigm that seeks to reveal the authentic face hidden behind the fake one by leveraging the joint information of face and audio. More specifically, given the fake face as well as the audio segment, the cross-modality transferable capability is exploited by learning to generate the feature of the authentic face, based on the underlying clues from the audio as well as the fake face appearance. The effectiveness of the proposed scheme is validated through a series of evaluations, and experimental results show that the proposed model achieves promising face reconstruction performance in revealing the hidden faces, in terms of reconstruction quality, as well as identity and face attribute inference accuracy.
Chenqi Kong, Baoliang Chen, Wenhan Yang, Haoliang Li, Peilin Chen 0001, Shiqi Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Compressed Domain Deep Video Super-Resolution
abstract
Real-world video processing algorithms are often faced with the great challenges of processing the compressed videos instead of pristine videos. Despite the tremendous successes achieved in deep-learning based video super-resolution (SR), much less work has been dedicated to the SR of compressed videos. Herein, we propose a novel approach for compressed domain deep video SR by jointly leveraging the coding priors and deep priors. By exploiting the diverse and ready-made spatial and temporal coding priors (e.g., partition maps and motion vectors) extracted directly from the video bitstream in an effortless way, the video SR in the compressed domain allows us to accurately reconstruct the high resolution video with high flexibility and substantially economized computational complexity. More specifically, to incorporate the spatial coding prior, the Guided Spatial Feature Transform (GSFT) layer is proposed to modulate features of the prior with the guidance of the video information, making the prior features more fine-grained and content-adaptive. To incorporate the temporal coding prior, a guided soft alignment scheme is designed to generate local attention off-sets to compensate for decoded motion vectors. Our soft alignment scheme combines the merits of explicit and implicit motion modeling methods, rendering the alignment of features more effective for SR in terms of the computational complexity and robustness to inaccurate motion fields. Furthermore, to fully make use of the deep priors, the multi-scale fused features are generated from a scale-wise convolution reconstruction network for final SR video reconstruction. To promote the compressed domain video SR research, we build an extensive Compressed Videos with Coding Prior (CVCP) dataset, including compressed videos of diverse content and various coding priors extracted from the bitstream. Extensive experimental results show the effectiveness of coding priors in compressed domain video SR.
Peilin Chen 0001, Wenhan Yang, Meng Wang 0017, Kangkang Hu, Shiqi Wang 0001
IEEE Trans. Image Process.1
2020 When Bitstream Prior Meets Deep Prior: Compressed Video Super-resolution with Learning from Decoding
abstract
The standard paradigm of video super-resolution (SR) is to generate the spatial-temporal coherent high-resolution (HR) sequence from the corresponding low-resolution (LR) version which has already been decoded from the bitstream. However, a highly practical while relatively under-studied way is enabling the built-in SR functionality in the decoder, in the sense that almost all videos are compactly represented. In this paper, we systematically investigate the SR of compressed LR videos by leveraging the interactivity between decoding prior and deep prior. By fully exploiting the compact video stream information, the proposed bitstream prior embedded SR framework achieves compressed video SR and quality enhancement simultaneously in a single feed-forward process. More specifically, we propose a motion vector guided multi-scale local attention module that explicitly exploits the temporal dependency and suppresses coding artifacts with substantially economized computational complexity. Moreover, a scale-wise deep residual-in-residual network is learned to reconstruct the SR frames from the multi-scale fused features. To facilitate the research of compressed video SR, we also build a large-scale dataset with compressed videos of diverse content, including ready-made diversified kinds of side information extracted from the bitstream. Both quantitative and qualitative evaluations show that our model achieves superior performance for compressed video SR, and offers competitive performance compared to the sequential combinations of the state-of-the-art methods for compressed video artifacts removal and SR.
Peilin Chen 0001, Wenhan Yang, Shiqi Wang 0001
ACM Multimedia1
2017 Feature based inter prediction optimization for non-translational video coding in cloud
abstract
Visual features of images and video frames have become pervasive and maturely developed in extensive research fields such as computer vision and visual search. In more and more cases, the visual feature becomes necessary information which needs to be transmitted and stored at server side in cloud. Among visual features, the local feature descriptors extracted by SIFT can represent both translational and non-translational motion, such as orientation and zooming. On the other hand, only translational motion can be represented by the Motion Vector (MV) in current MV based block video coding standard. Inspired by these properties, a method that utilizes the available feature to optimize inter prediction video coding is proposed in this paper. In this method, the localization, orientation and scale parameters of matching features extracted by SIFT are delivered to inter prediction to provide non-translational motion estimation (ME) and optimized merge mode. Experimental results have shown that the proposed method can efficiently improve the coding performance according to the accurate feature-matching.
Xuelin Shen, Jun Wang 0015, Peilin Chen 0001, Fan Liang 0001
VCIP4