Shiwei Wang 0005

dblp:76/6446-5 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
13since 2021 · last 2026
0009-0002-6721-3521ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Breaking Redundancy via 3D Sparse Geometry: 3D-aware Neural Compression for Multi-View Videos
Shiwei Wang 0005, Liquan Shen, Jimin Xiao, Zhaoyi Tian, Feifeng Wang, Xiangyu Hu 0003, Yao Zhu 0006, Guorui Feng
Int. J. Comput. Vis.1
2026 SCVQENet: Quality enhancement for compressed screen content video
Zhaoyi Tian, Shiwei Wang 0005, Feifeng Wang, Liquan Shen
Signal Process. Image Commun.2
2026 DBA-PCGC: Dual-Domain Boundary Aware for Task-Friendly Point Cloud Geometry Compression
abstract
Compressed point clouds are increasingly used in machine vision tasks, which rely on key semantic regions of the point cloud such as geometric details and structural boundaries. However, existing point cloud compression methods for machine vision lack explicit awareness of geometrically induced semantic boundaries, causing semantic ambiguity in certain boundary regions during compression, thereby degrading machine vision performance. To address this issue, we propose a Dual-domain Boundary Aware Point Cloud Geometry Compression (DBA-PCGC) method that explicitly preserves semantic geometric boundaries from complementary spatial and frequency perspectives, enabling beneficial for machine vision tasks. Specifically, a Structure Aware Transform Module (SATM) exploits Gram matrix traces on local graphs to capture structural variations and highlight high-variation boundary regions, while compactly encoding smooth areas. In parallel, a Frequency Aware Transform Module (FATM) applies Chebyshev high-pass filtering to enhance high-frequency components corresponding to semantic geometric boundaries and suppress redundant low-frequency content. Experimental results on point cloud machine vision tasks demonstrate that our method achieves superior performance compared with existing compression approaches.
Minjian Chen, Liquan Shen, Qi Teng, Shiwei Wang 0005, Feifeng Wang
IEEE Signal Process. Lett.4
2026 UCSMC: An Underwater Compressed Sensing With Measurement Compression Framework
abstract
Thriving ocean applications demand efficient underwater image compression over bandwidth-limited acoustic channels. Recent works combine compressed sensing with measurement compression to improve compression ratios. However, as underwater attenuation weakens structural cues, sampling methods tend to overlook structural information and yield poor reconstructions. Meanwhile, sampling leaves discrete measurements with weak intra-image correlations, making it difficult for entropy models within measurement compression to predict accurate probability distributions. In this paper, we propose an Underwater Compressed Sensing with Measurement Compression (UCSMC) framework including Sketch-Assisted Sampling (SAS) and Spatial-Dictionary-based Mixture Entropy Coding (SDMEC) for low-bit-rate reconstruction. Specifically, in sampling, we incorporate sketch with underwater priors to drive the sampling process, steering more measurements toward critical structural regions and ultimately improving reconstruction quality. Additionally, we introduce a learnable spatial dictionary storing per-location entropy statistical characteristics in the underwater domain, which indicates local estimation difficulty and guides adaptive attention allocation in the entropy model, thereby improving probability estimation accuracy. Experimental results show our method outperforms previous schemes in reconstruction quality and measurement compression efficiency.
Liquan Shen, Shiwei Wang 0005, Minjian Chen
IEEE Signal Process. Lett.4
2026 ESHIC: Efficient Learning-Based Scalable HDR Image Compression With Hybrid Structural-Tonal Prior Modeling
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Feifeng Wang, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.6
2026 DSCVC: Deep Screen Content Video Compression
abstract
Different from natural videos, screen content videos (SCVs) often exhibit homogeneous regions, abrupt content changes, and high prevalence of repetitive patterns. Existing deep learning (DL)-based video compression methods inadequately address the unique characteristics of SCVs, resulting in suboptimal compression performance. Therefore, in this paper, a dedicated deep screen content video compression (DSCVC) framework is proposed based on the motion and content characteristics of SCVs, which includes superpixel-constrained a motion estimation (SCME) module and inter and intra context aggregation (I2CA) module. The SCME is designed to construct a superpixel-based representation of homogeneous regions, leveraging the global correlations among superpixels to effectively capture large-scale motions, which efficiently improves the compression performance. I2CA is developed to jointly utilize inter and intra contexts, which employs a gating mechanism for content-aware context fusion, dynamically aggregating more similar contexts within SCVs. This allows for flexible adaptation to both contiguous and abrupt content changes within SCVs. Furthermore, by leveraging both learnable window and pixel displacements, a displacement-guided window attention mechanism is implemented in I2CA for precise long range repetitive feature localization, thereby reducing redundancy caused by repetitive patterns. To the best of our knowledge, it is the first DL-based video compression framework specifically designed for SCVs. Extensive experimental results demonstrate that the proposed DSCVC significantly outperforms existing methods in terms of compression performance, achieving a bitrate saving of 26.82% compared to VVC and a bitrate saving of 12.30% compared to SOTA DL-based methods.
Feifeng Wang, Liquan Shen, Zhaoyi Tian, Shiwei Wang 0005, Qi Teng, Yao Zhu 0006, Chengtao Zhou
IEEE Trans. Circuits Syst. Video Technol.4
2026 Sketch-Based Extreme Underwater Image Compression Network
abstract
Underwater applications such as exploration and salvage operations require capturing underwater images (UWIs) to evaluate attributes such as the shape and structural integrity of submerged targets. However, underwater image transmission faces significant challenges due to the limited wireless acoustic channel available in underwater communication systems. Existing image compression algorithms struggle with limited compression ratios, which leads to a loss of crucial structural information and poor reconstruction quality, making them unsuitable for underwater practical applications. To overcome these limitations, we propose a sparse Sketch-based Extreme Underwater Compression framework (SEUCN), which mainly includes two sub-networks: Sparse Sketch Generation Network (SSGN) and Underwater Prior-guided Reconstruction Network (UPRN). To reduce redundancy and ensure effective compression at extremely low bitrates, the SSGN is designed to generate a compression-friendly sparse structural sketch through two ways. Firstly, it focuses on extracting important structural information to support analysis tasks within the constraints of limited bitrates. Secondly, it incorporates an underwater imaging model to focus on learning critical texture information for visual reconstruction. To restore the information lost during compression and achieve high-quality reconstruction, UPRN is designed to enhance structure details, restore underwater style, and enrich texture information during the reconstruction of UWIs from the decoded sketches, by effectively integrating multiple sources of prior knowledge. Specially, considering the high similarity of semantics and texture across different UWIs with common targets, the Dictionary-guided Texture Recovery Module (DTRM) leverages a universal underwater multi-scale feature dictionary as texture prior knowledge to supplement missing texture details. Extensive experiments show that our SEUCN demonstrates outstanding performance in retaining significant structural information to assist underwater practical tasks, and achieves superior visual quality compared to existing methods.
Liquan Shen, Shiwei Wang 0005, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.5
2025 High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithm
abstract
Recently, learned video compression (LVC) is undergoing a period of rapid development. However, due to absence of large and high-quality high dynamic range (HDR) video training data, LVC on HDR video is still unexplored. In this paper, we are the first to collect a large-scale HDR video benchmark dataset, named HDRVD2K, featuring huge quantity, diverse scenes and multiple motion types. HDRVD2K fills gaps of video training data and facilitate the development of LVC on HDR videos. Based on HDRVD2K, we further propose the first learned bit-depth scalable video compression (LBSVC) network for HDR videos by effectively exploiting bit-depth redundancy between videos of multiple dynamic ranges. To achieve this, we first propose a compression-friendly bit-depth enhancement module (BEM) to effectively predict original HDR videos based on compressed tone-mapped low dynamic range (LDR) videos and dynamic range prior, instead of reducing redundancy only through spatio-temporal predictions. Our method greatly improves the reconstruction quality and compression performance on HDR videos. Extensive experiments demonstrate the effectiveness of HDRVD2K on learned HDR video compression and great compression performance of our proposed LB-SVC network. Code and dataset will be released in https://github.com/sdkinda/HDR-Learned-Video-Coding.
Zhaoyi Tian, Feifeng Wang, Shiwei Wang 0005, Yao Zhu 0006, Liquan Shen
CVPR3
2025 DRLN: Disparity-Aware Rescaling Learning Network for Multi-View Video Coding Optimization
abstract
Efficient compression of multi-view video data is a critical challenge for various applications due to the large volume of data involved. Although multi-view video coding (MVC) has introduced inter-view prediction techniques to reduce video redundancies, further reduction can be achieved by encoding a subset of views at a lower resolution through asymmetric rescaling, achieving higher compression efficiency. However, existing network-based rescaling approaches are designed solely for single-viewpoint videos. These methods neglect inter-view characteristics inherent in multi-view videos, resulting in suboptimal performance. To address this issue, we first propose a Disparity-aware Rescaling Learning Network (DRLN) that integrates disparity-aware feature extraction and multi-resolution adaptive rescaling to enhance MVC efficiency by minimizing both self- and inter-view redundancies. On the one hand, during the encoding stage, our method leverages the non-local correlation of multi-view contexts and performs adaptive downscaling with an early-exit mechanism, resulting in substantial multi-view bitrate savings. On the other hand, during the decoding stage, a dynamic aggregation strategy is proposed to facilitate effective interaction with inter-view features, utilizing the inter-view and cross-scale information to reconstruct fine-grained multi-view videos. Extensive experiments show that our network achieves a significant 26.31% BD-Rate reduction compared to the 3D-HEVC standard baseline, offering state of-the-art coding performance.
Shiwei Wang 0005, Liquan Shen, Peiying Wu, Zhaoyi Tian, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Perceptual Quality Assessment of High-Dynamic-Range Image: A Benchmark Dataset and a No Reference Method
abstract
High dynamic range (HDR) imaging technology has received increasing attention in recent years, and HDR image quality assessment (IQA) metrics are indispensable during the capturing, processing and displaying of HDR images. However, existing HDR-IQA datasets and methods neglect complex distortions during the HDR image processing schemes, leading to limited generalization performance on practical application. In this work, to facilitate the development of HDR-IQA dataset, we present HDRQAD, a large-scale HDR Quality Assessment Dataset, which possesses diversified distortions during HDR imaging technologies, abundant scenes and considerable quantity. Specifically, the HDRQAD dataset contains 1409 HDR images, which are derived from source scenes with six types of distortions during the HDR imaging schemes. In contrast to existing datasets that contain only compression artifacts, the HDRQAD includes Under-exposure, Over-exposure, Motion blur and Ghosting in HDR images achieved with multi-exposure fusion technology, conversion artifacts in HDR images achieved with single image reconstruction technology and compression artifacts during the transmission of HDR images. Furthermore, during the process of constructing the dataset, we identified three key challenges in HDR-IQA tasks: 1) dynamic range variations, 2) HDR visual artifacts with large overall gap, 3) inter-regional non-uniform image quality. Based on these observations, we propose a new end-to-end network for HDR-IQA tasks, which consists of a Distortion-aware Representation Learning (DRL) module and an Inter-Regional Quality Interaction (IRQI) module. The DRL learns the representations of dynamic range variations and HDR visual artifacts, enhancing the reliability of prior information extraction. The IRQI captures inter-regional quality dependencies with interacting and fusing intermediate distortion features for more accurately predicting image quality. Extensive experiments prove the superiority of proposed HDRQAD and demonstrate that the proposed network achieves state-of-the-art performance. The Dataset and Code will be made publicly available at HDR-IQA-Dataset.
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.6
2024 Spatial-Temporal Inter-Layer Reference Frame Generation Network for Spatial SHVC
abstract
In the current spatial Scalable High Efficiency Video Coding (SHVC) standard, the main techniques involve exploiting the correlation between pixel values of different layers to achieve inter-layer prediction samples, allowing the enhancement layer (EL) to predict samples from the upsampled base layer (BL) frame and remove temporal redundancy. However, existing network-based methods cannot effectively handle multi-layer compressed images with different resolutions to generate reference frame in spatial SHVC. Meanwhile, spatial SHVC only uses traditional interpolation filters to upsample the BL frame for EL frame sample prediction, which cannot handle different structures and contents. Therefore, considering the high correlation of multi-scale distortion characteristics across different layers, this article proposes a spatial-temporal inter-layer reference frame generation network (ST-ILR) for spatial SHVC, which can generate a high-fidelity reference frame for efficient inter-prediction and insert it into the EL reference picture list. The proposed method consists of two modules: a multi-scale motion restoration (MMR) module and a guided multi-scale feature reconstruction (GMFR) module. The MMR model is designed to accurately predict the motion trend of the EL based on the BL motion information, while implicitly compensating for previous EL frames. This is achieved by dynamically modeling the current EL motion information from the BL, capturing compression downsampling differences of prior motion vectors across different layers. The GMFR module adaptively super-resolves compressed BL frames and selectively aggregates high-frequency information from aligned EL features to preserve precise spatial detail, fusing abundant features from different layers to achieve better ILR frame quality performance. Extensive experiments show that our network achieves a 13.6% BD-rate (Bjøntegaard Delta Rate) reduction in random access configuration compared to the SHVC baseline, which offers state-of-the-art coding performance.
Shiwei Wang 0005, Liquan Shen, Jingyue Liu 0002
IEEE Trans. Multim.1
2024 Multi-Prior Driven Resolution Rescaling Blocks for Intra Frame Coding
abstract
Deep learning techniques are increasingly integrated into rescaling-based video compression frameworks and have shown great potential in improving compression efficiency. However, existing methods achieve limited performance because 1) they treat context priors generated by codec as independent sources of information, ignoring potential interactions between multiple priors in rescaling, which may not effectively facilitate compression; 2) they often employ a uniform sampling ratio across regions with varying content complexities, resulting in the loss of important information. To address the above two issues, this paper proposes a spatial multi-prior driven resolution rescaling framework for intra-frame coding, called MP-RRF, consisting of three sub-networks: a multi-prior driven network, a downscaling network, and an upscaling network. First, the multi-prior driven network employs complexity and similarity priors to smooth the unnecessarily complicated information while leveraging similarity and quality priors to produce high-fidelity complementary information. This interaction of complexity, similarity and quality priors ensures redundancy reduction and texture enhancement. Second, the downscaling network discriminatively processes components of different granularities to generate a compact, low-resolution image for encoding. The upscaling network aggregates a complementary set of contextual multi-scale features to reconstruct realistic details while combining variable receptive fields to suppress multi-scale compression artifacts and resampling noise. Extensive experiments show that our network achieves a significant 23.84% Bjøntegaard Delta Rate (BD-Rate) reduction under all-intra configuration compared to the codec anchor, offering the state-of-the-art coding performance.
Peiying Wu, Shiwei Wang 0005, Liquan Shen, Feifeng Wang, Zhaoyi Tian
IEEE Trans. Multim.2
2022 Effective QTMT Partition Decision Algorithm for VVC Intercoding
abstract
Aiming at the problem of high complexity of Versatile Video Coding (VVC) inter coding, this paper proposes a QTMT partition decision algorithm based on a multi-level decision framework. Specifically, the multi-level decision framework decomposes the multi-mode partition decision problem into multiple independent single-mode partition decision problems and then adopts a classification method based on machine learning to predict result of each single mode partition. To design more efficient classification features, fast motion estimation on 4×4 blocks is first performed to construct a motion field, and features on global/local motion and global/local consistency of its corresponding residuals are designed to measure motion activity and texture homogeneity. Furthermore, a misclassification protection mechanism is designed to decrease influences of misclassification on coding performance loss. Experimental results show that the proposed effective QTMT partition decision algorithm achieves a computational complexity reduction more than 51%, while incurring 1.65% BDBR increase compared with that of the original coding in the test model of VVC(VTM).
Liquan Shen, Hao Yang 0008, Shiwei Wang 0005
MMSP3