Feifeng Wang

dblp:314/2835 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Breaking Redundancy via 3D Sparse Geometry: 3D-aware Neural Compression for Multi-View Videos
Shiwei Wang 0005, Liquan Shen, Jimin Xiao, Zhaoyi Tian, Feifeng Wang, Xiangyu Hu 0003, Yao Zhu 0006, Guorui Feng
Int. J. Comput. Vis.5
2026 SCVQENet: Quality enhancement for compressed screen content video
Zhaoyi Tian, Shiwei Wang 0005, Feifeng Wang, Liquan Shen
Signal Process. Image Commun.4
2026 DBA-PCGC: Dual-Domain Boundary Aware for Task-Friendly Point Cloud Geometry Compression
abstract
Compressed point clouds are increasingly used in machine vision tasks, which rely on key semantic regions of the point cloud such as geometric details and structural boundaries. However, existing point cloud compression methods for machine vision lack explicit awareness of geometrically induced semantic boundaries, causing semantic ambiguity in certain boundary regions during compression, thereby degrading machine vision performance. To address this issue, we propose a Dual-domain Boundary Aware Point Cloud Geometry Compression (DBA-PCGC) method that explicitly preserves semantic geometric boundaries from complementary spatial and frequency perspectives, enabling beneficial for machine vision tasks. Specifically, a Structure Aware Transform Module (SATM) exploits Gram matrix traces on local graphs to capture structural variations and highlight high-variation boundary regions, while compactly encoding smooth areas. In parallel, a Frequency Aware Transform Module (FATM) applies Chebyshev high-pass filtering to enhance high-frequency components corresponding to semantic geometric boundaries and suppress redundant low-frequency content. Experimental results on point cloud machine vision tasks demonstrate that our method achieves superior performance compared with existing compression approaches.
Minjian Chen, Liquan Shen, Qi Teng, Shiwei Wang 0005, Feifeng Wang
IEEE Signal Process. Lett.5
2026 ESHIC: Efficient Learning-Based Scalable HDR Image Compression With Hybrid Structural-Tonal Prior Modeling
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Feifeng Wang, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.5
2026 DSCVC: Deep Screen Content Video Compression
abstract
Different from natural videos, screen content videos (SCVs) often exhibit homogeneous regions, abrupt content changes, and high prevalence of repetitive patterns. Existing deep learning (DL)-based video compression methods inadequately address the unique characteristics of SCVs, resulting in suboptimal compression performance. Therefore, in this paper, a dedicated deep screen content video compression (DSCVC) framework is proposed based on the motion and content characteristics of SCVs, which includes superpixel-constrained a motion estimation (SCME) module and inter and intra context aggregation (I2CA) module. The SCME is designed to construct a superpixel-based representation of homogeneous regions, leveraging the global correlations among superpixels to effectively capture large-scale motions, which efficiently improves the compression performance. I2CA is developed to jointly utilize inter and intra contexts, which employs a gating mechanism for content-aware context fusion, dynamically aggregating more similar contexts within SCVs. This allows for flexible adaptation to both contiguous and abrupt content changes within SCVs. Furthermore, by leveraging both learnable window and pixel displacements, a displacement-guided window attention mechanism is implemented in I2CA for precise long range repetitive feature localization, thereby reducing redundancy caused by repetitive patterns. To the best of our knowledge, it is the first DL-based video compression framework specifically designed for SCVs. Extensive experimental results demonstrate that the proposed DSCVC significantly outperforms existing methods in terms of compression performance, achieving a bitrate saving of 26.82% compared to VVC and a bitrate saving of 12.30% compared to SOTA DL-based methods.
Feifeng Wang, Liquan Shen, Zhaoyi Tian, Shiwei Wang 0005, Qi Teng, Yao Zhu 0006, Chengtao Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2026 Sketch-Based Extreme Underwater Image Compression Network
abstract
Underwater applications such as exploration and salvage operations require capturing underwater images (UWIs) to evaluate attributes such as the shape and structural integrity of submerged targets. However, underwater image transmission faces significant challenges due to the limited wireless acoustic channel available in underwater communication systems. Existing image compression algorithms struggle with limited compression ratios, which leads to a loss of crucial structural information and poor reconstruction quality, making them unsuitable for underwater practical applications. To overcome these limitations, we propose a sparse Sketch-based Extreme Underwater Compression framework (SEUCN), which mainly includes two sub-networks: Sparse Sketch Generation Network (SSGN) and Underwater Prior-guided Reconstruction Network (UPRN). To reduce redundancy and ensure effective compression at extremely low bitrates, the SSGN is designed to generate a compression-friendly sparse structural sketch through two ways. Firstly, it focuses on extracting important structural information to support analysis tasks within the constraints of limited bitrates. Secondly, it incorporates an underwater imaging model to focus on learning critical texture information for visual reconstruction. To restore the information lost during compression and achieve high-quality reconstruction, UPRN is designed to enhance structure details, restore underwater style, and enrich texture information during the reconstruction of UWIs from the decoded sketches, by effectively integrating multiple sources of prior knowledge. Specially, considering the high similarity of semantics and texture across different UWIs with common targets, the Dictionary-guided Texture Recovery Module (DTRM) leverages a universal underwater multi-scale feature dictionary as texture prior knowledge to supplement missing texture details. Extensive experiments show that our SEUCN demonstrates outstanding performance in retaining significant structural information to assist underwater practical tasks, and achieves superior visual quality compared to existing methods.
Liquan Shen, Shiwei Wang 0005, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.6
2025 High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithm
abstract
Recently, learned video compression (LVC) is undergoing a period of rapid development. However, due to absence of large and high-quality high dynamic range (HDR) video training data, LVC on HDR video is still unexplored. In this paper, we are the first to collect a large-scale HDR video benchmark dataset, named HDRVD2K, featuring huge quantity, diverse scenes and multiple motion types. HDRVD2K fills gaps of video training data and facilitate the development of LVC on HDR videos. Based on HDRVD2K, we further propose the first learned bit-depth scalable video compression (LBSVC) network for HDR videos by effectively exploiting bit-depth redundancy between videos of multiple dynamic ranges. To achieve this, we first propose a compression-friendly bit-depth enhancement module (BEM) to effectively predict original HDR videos based on compressed tone-mapped low dynamic range (LDR) videos and dynamic range prior, instead of reducing redundancy only through spatio-temporal predictions. Our method greatly improves the reconstruction quality and compression performance on HDR videos. Extensive experiments demonstrate the effectiveness of HDRVD2K on learned HDR video compression and great compression performance of our proposed LB-SVC network. Code and dataset will be released in https://github.com/sdkinda/HDR-Learned-Video-Coding.
Zhaoyi Tian, Feifeng Wang, Shiwei Wang 0005, Yao Zhu 0006, Liquan Shen
CVPR2
2025 Bézier Surface-Guided Sampling Space Constraint for Neural Point Cloud Geometry Compression
abstract
Implicit Neural Representations (INRs), capable of learning the mapping between sampling coordinates and occupancy statuses to effectively represent 3D content, are increasingly adopted in Point Cloud Geometry Compression (PCGC). Existing INR-based PCGC methods suffer from substantial sampling space redundancy, leading to limited compression performance on point clouds. To address this issue, we propose Bezier ´Surface-Guided Sampling Space Constraint (BSGSSC), a novel method that significantly reduces sampling space redundancy via Bezier surface guidance. The key contributions of our work are: ´ 1) Leveraging the inherent compactness and distortion robustness of Bezier surfaces, a Low Bitrate Robust Surface Approximation ´ (LBRSA) is presented to outline the surface of dense point clouds with low transmission overhead; 2) An Oriented Bounding Box-based Sampling Space Constraint (OBBSSC) is designed to effectively constrain sampling space using multiple Oriented Bounding Boxes generated by Bezier surfaces. Experimental ´ results show that our method can enhance INR-based PCGC by 0.89 dB in terms of Bjøntegaard Delta Peak Signal-to-Noise Rate (BP).
Qi Teng, Liquan Shen, Feifeng Wang, Minjian Chen
IEEE Signal Process. Lett.3
2025 DRLN: Disparity-Aware Rescaling Learning Network for Multi-View Video Coding Optimization
abstract
Efficient compression of multi-view video data is a critical challenge for various applications due to the large volume of data involved. Although multi-view video coding (MVC) has introduced inter-view prediction techniques to reduce video redundancies, further reduction can be achieved by encoding a subset of views at a lower resolution through asymmetric rescaling, achieving higher compression efficiency. However, existing network-based rescaling approaches are designed solely for single-viewpoint videos. These methods neglect inter-view characteristics inherent in multi-view videos, resulting in suboptimal performance. To address this issue, we first propose a Disparity-aware Rescaling Learning Network (DRLN) that integrates disparity-aware feature extraction and multi-resolution adaptive rescaling to enhance MVC efficiency by minimizing both self- and inter-view redundancies. On the one hand, during the encoding stage, our method leverages the non-local correlation of multi-view contexts and performs adaptive downscaling with an early-exit mechanism, resulting in substantial multi-view bitrate savings. On the other hand, during the decoding stage, a dynamic aggregation strategy is proposed to facilitate effective interaction with inter-view features, utilizing the inter-view and cross-scale information to reconstruct fine-grained multi-view videos. Extensive experiments show that our network achieves a significant 26.31% BD-Rate reduction compared to the 3D-HEVC standard baseline, offering state of-the-art coding performance.
Shiwei Wang 0005, Liquan Shen, Peiying Wu, Zhaoyi Tian, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.5
2024 DSCIC: Deep Screen Content Image Compression
abstract
Existing deep learning-based image compression methods overlook the unique properties of screen content images (SCIs), like limited color values and abundant repetitive patterns, leading to limited compression performance on SCIs. Therefore, a specialized framework, deep screen content image compression (DSCIC) is proposed, which contains a color context generator (CCG) and a region-based block aggregation (RBA) module. The CCG is designed to generate compression-friendly color contexts based on main color components, embedded in the encoder-decoder to remove color representation redundancy. Furthermore, to effectively reduce repetitive block redundancy in SCIs, the RBA captures repetitive patterns and enables adaptive aggregation in the latent space. It leverages region-based block matching and block content-aware aggregation to utilize repetitive features for further improving compression performance. Extensive experimental results demonstrate that the proposed DSCIC outperforms the most advanced traditional codec VVC-SCC, and is significantly superior to other learning-based image compression methods. Using VVC as the anchor, DSCIC exhibits further BD-Rate savings of 12.185% and 4.889% compared to VVC-SCC and the SOTA deep learning-based method, respectively.
Feifeng Wang, Liquan Shen, Qi Teng, Zhaoyi Tian
IEEE Trans. Circuits Syst. Video Technol.1
2024 Multi-Prior Driven Resolution Rescaling Blocks for Intra Frame Coding
abstract
Deep learning techniques are increasingly integrated into rescaling-based video compression frameworks and have shown great potential in improving compression efficiency. However, existing methods achieve limited performance because 1) they treat context priors generated by codec as independent sources of information, ignoring potential interactions between multiple priors in rescaling, which may not effectively facilitate compression; 2) they often employ a uniform sampling ratio across regions with varying content complexities, resulting in the loss of important information. To address the above two issues, this paper proposes a spatial multi-prior driven resolution rescaling framework for intra-frame coding, called MP-RRF, consisting of three sub-networks: a multi-prior driven network, a downscaling network, and an upscaling network. First, the multi-prior driven network employs complexity and similarity priors to smooth the unnecessarily complicated information while leveraging similarity and quality priors to produce high-fidelity complementary information. This interaction of complexity, similarity and quality priors ensures redundancy reduction and texture enhancement. Second, the downscaling network discriminatively processes components of different granularities to generate a compact, low-resolution image for encoding. The upscaling network aggregates a complementary set of contextual multi-scale features to reconstruct realistic details while combining variable receptive fields to suppress multi-scale compression artifacts and resampling noise. Extensive experiments show that our network achieves a significant 23.84% Bjøntegaard Delta Rate (BD-Rate) reduction under all-intra configuration compared to the codec anchor, offering the state-of-the-art coding performance.
Peiying Wu, Shiwei Wang 0005, Liquan Shen, Feifeng Wang, Zhaoyi Tian
IEEE Trans. Multim.4
2022 Spatial-frequency HEVC multiple description video coding with adaptive perceptual redundancy allocation
Feifeng Wang, Jing Chen 0001, Huanqiang Zeng, Canhui Cai
J. Vis. Commun. Image Represent.1