Zhaoyi Tian

dblp:391/8169 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0002-7239-2640ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Breaking Redundancy via 3D Sparse Geometry: 3D-aware Neural Compression for Multi-View Videos
Shiwei Wang 0005, Liquan Shen, Jimin Xiao, Zhaoyi Tian, Feifeng Wang, Xiangyu Hu 0003, Yao Zhu 0006, Guorui Feng
Int. J. Comput. Vis.4
2026 SCVQENet: Quality enhancement for compressed screen content video
Zhaoyi Tian, Shiwei Wang 0005, Feifeng Wang, Liquan Shen
Signal Process. Image Commun.1
2026 ESHIC: Efficient Learning-Based Scalable HDR Image Compression With Hybrid Structural-Tonal Prior Modeling
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Feifeng Wang, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.3
2026 DSCVC: Deep Screen Content Video Compression
abstract
Different from natural videos, screen content videos (SCVs) often exhibit homogeneous regions, abrupt content changes, and high prevalence of repetitive patterns. Existing deep learning (DL)-based video compression methods inadequately address the unique characteristics of SCVs, resulting in suboptimal compression performance. Therefore, in this paper, a dedicated deep screen content video compression (DSCVC) framework is proposed based on the motion and content characteristics of SCVs, which includes superpixel-constrained a motion estimation (SCME) module and inter and intra context aggregation (I2CA) module. The SCME is designed to construct a superpixel-based representation of homogeneous regions, leveraging the global correlations among superpixels to effectively capture large-scale motions, which efficiently improves the compression performance. I2CA is developed to jointly utilize inter and intra contexts, which employs a gating mechanism for content-aware context fusion, dynamically aggregating more similar contexts within SCVs. This allows for flexible adaptation to both contiguous and abrupt content changes within SCVs. Furthermore, by leveraging both learnable window and pixel displacements, a displacement-guided window attention mechanism is implemented in I2CA for precise long range repetitive feature localization, thereby reducing redundancy caused by repetitive patterns. To the best of our knowledge, it is the first DL-based video compression framework specifically designed for SCVs. Extensive experimental results demonstrate that the proposed DSCVC significantly outperforms existing methods in terms of compression performance, achieving a bitrate saving of 26.82% compared to VVC and a bitrate saving of 12.30% compared to SOTA DL-based methods.
Feifeng Wang, Liquan Shen, Zhaoyi Tian, Shiwei Wang 0005, Qi Teng, Yao Zhu 0006, Chengtao Zhou
IEEE Trans. Circuits Syst. Video Technol.3
2025 High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithm
abstract
Recently, learned video compression (LVC) is undergoing a period of rapid development. However, due to absence of large and high-quality high dynamic range (HDR) video training data, LVC on HDR video is still unexplored. In this paper, we are the first to collect a large-scale HDR video benchmark dataset, named HDRVD2K, featuring huge quantity, diverse scenes and multiple motion types. HDRVD2K fills gaps of video training data and facilitate the development of LVC on HDR videos. Based on HDRVD2K, we further propose the first learned bit-depth scalable video compression (LBSVC) network for HDR videos by effectively exploiting bit-depth redundancy between videos of multiple dynamic ranges. To achieve this, we first propose a compression-friendly bit-depth enhancement module (BEM) to effectively predict original HDR videos based on compressed tone-mapped low dynamic range (LDR) videos and dynamic range prior, instead of reducing redundancy only through spatio-temporal predictions. Our method greatly improves the reconstruction quality and compression performance on HDR videos. Extensive experiments demonstrate the effectiveness of HDRVD2K on learned HDR video compression and great compression performance of our proposed LB-SVC network. Code and dataset will be released in https://github.com/sdkinda/HDR-Learned-Video-Coding.
Zhaoyi Tian, Feifeng Wang, Shiwei Wang 0005, Yao Zhu 0006, Liquan Shen
CVPR1
2025 DRLN: Disparity-Aware Rescaling Learning Network for Multi-View Video Coding Optimization
abstract
Efficient compression of multi-view video data is a critical challenge for various applications due to the large volume of data involved. Although multi-view video coding (MVC) has introduced inter-view prediction techniques to reduce video redundancies, further reduction can be achieved by encoding a subset of views at a lower resolution through asymmetric rescaling, achieving higher compression efficiency. However, existing network-based rescaling approaches are designed solely for single-viewpoint videos. These methods neglect inter-view characteristics inherent in multi-view videos, resulting in suboptimal performance. To address this issue, we first propose a Disparity-aware Rescaling Learning Network (DRLN) that integrates disparity-aware feature extraction and multi-resolution adaptive rescaling to enhance MVC efficiency by minimizing both self- and inter-view redundancies. On the one hand, during the encoding stage, our method leverages the non-local correlation of multi-view contexts and performs adaptive downscaling with an early-exit mechanism, resulting in substantial multi-view bitrate savings. On the other hand, during the decoding stage, a dynamic aggregation strategy is proposed to facilitate effective interaction with inter-view features, utilizing the inter-view and cross-scale information to reconstruct fine-grained multi-view videos. Extensive experiments show that our network achieves a significant 26.31% BD-Rate reduction compared to the 3D-HEVC standard baseline, offering state of-the-art coding performance.
Shiwei Wang 0005, Liquan Shen, Peiying Wu, Zhaoyi Tian, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Perceptual Quality Assessment of High-Dynamic-Range Image: A Benchmark Dataset and a No Reference Method
abstract
High dynamic range (HDR) imaging technology has received increasing attention in recent years, and HDR image quality assessment (IQA) metrics are indispensable during the capturing, processing and displaying of HDR images. However, existing HDR-IQA datasets and methods neglect complex distortions during the HDR image processing schemes, leading to limited generalization performance on practical application. In this work, to facilitate the development of HDR-IQA dataset, we present HDRQAD, a large-scale HDR Quality Assessment Dataset, which possesses diversified distortions during HDR imaging technologies, abundant scenes and considerable quantity. Specifically, the HDRQAD dataset contains 1409 HDR images, which are derived from source scenes with six types of distortions during the HDR imaging schemes. In contrast to existing datasets that contain only compression artifacts, the HDRQAD includes Under-exposure, Over-exposure, Motion blur and Ghosting in HDR images achieved with multi-exposure fusion technology, conversion artifacts in HDR images achieved with single image reconstruction technology and compression artifacts during the transmission of HDR images. Furthermore, during the process of constructing the dataset, we identified three key challenges in HDR-IQA tasks: 1) dynamic range variations, 2) HDR visual artifacts with large overall gap, 3) inter-regional non-uniform image quality. Based on these observations, we propose a new end-to-end network for HDR-IQA tasks, which consists of a Distortion-aware Representation Learning (DRL) module and an Inter-Regional Quality Interaction (IRQI) module. The DRL learns the representations of dynamic range variations and HDR visual artifacts, enhancing the reliability of prior information extraction. The IRQI captures inter-regional quality dependencies with interacting and fusing intermediate distortion features for more accurately predicting image quality. Extensive experiments prove the superiority of proposed HDRQAD and demonstrate that the proposed network achieves state-of-the-art performance. The Dataset and Code will be made publicly available at HDR-IQA-Dataset.
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.4
2024 EUICN: An Efficient Underwater Image Compression Network
abstract
Thriving ocean applications bring explosive growth of underwater images (UWIs), which urgently demand to be compressed efficiently for transmission in the narrow underwater acoustic channel. However, existing image compression networks achieve suboptimal performance on UWIs. More efficient UWI compression can be achieved by utilizing characteristics of UWIs: (1) Within an UWI, the details distribution is associated with the underwater imaging transmission map (T-map); (2) Different UWIs have higher correlation than terrestrial images because they often present gauzy-covered indistinct appearance and share some universal ocean objects that widely appear in different underwater scenes. This paper fully exploits the two characteristics in terms of quantization and entropy coding, two key components of image compression network. Specifically, we propose an efficient underwater image compression network (EUICN) including underwater T-map-based quantization (UTMQ) and mixture entropy coding (MEC). In which, UTMQ extracts the imaging features from T-map, which are integrated with latent features by a novel dual-spatial attention module (DSAM) to generate a feature reserved mask, adaptively reserving reasonable numbers of features for different regions. For more efficient entropy coding, MEC is designed, which includes three correlation information extraction modules (i.e., hyperprior, local and novel universal information) and a probability prediction module. Especially, the universal information extraction module utilizes a comprehensive underwater feature dictionary, which covers various universal ocean objects, to match with latent features to select the universal correlation features. After that, the probability prediction module is designed to consolidate the hyperprior, local, and universal information to predict more accurate probability of latent features. Extensive experiments show that our EUICN achieves better performance than SOTA learned and conventional codecs in terms of PSNR and MS-SSIM.
Liquan Shen, Zhaoyi Tian
IEEE Trans. Circuits Syst. Video Technol.4
2024 DSCIC: Deep Screen Content Image Compression
abstract
Existing deep learning-based image compression methods overlook the unique properties of screen content images (SCIs), like limited color values and abundant repetitive patterns, leading to limited compression performance on SCIs. Therefore, a specialized framework, deep screen content image compression (DSCIC) is proposed, which contains a color context generator (CCG) and a region-based block aggregation (RBA) module. The CCG is designed to generate compression-friendly color contexts based on main color components, embedded in the encoder-decoder to remove color representation redundancy. Furthermore, to effectively reduce repetitive block redundancy in SCIs, the RBA captures repetitive patterns and enables adaptive aggregation in the latent space. It leverages region-based block matching and block content-aware aggregation to utilize repetitive features for further improving compression performance. Extensive experimental results demonstrate that the proposed DSCIC outperforms the most advanced traditional codec VVC-SCC, and is significantly superior to other learning-based image compression methods. Using VVC as the anchor, DSCIC exhibits further BD-Rate savings of 12.185% and 4.889% compared to VVC-SCC and the SOTA deep learning-based method, respectively.
Feifeng Wang, Liquan Shen, Qi Teng, Zhaoyi Tian
IEEE Trans. Circuits Syst. Video Technol.4
2024 Multi-Prior Driven Resolution Rescaling Blocks for Intra Frame Coding
abstract
Deep learning techniques are increasingly integrated into rescaling-based video compression frameworks and have shown great potential in improving compression efficiency. However, existing methods achieve limited performance because 1) they treat context priors generated by codec as independent sources of information, ignoring potential interactions between multiple priors in rescaling, which may not effectively facilitate compression; 2) they often employ a uniform sampling ratio across regions with varying content complexities, resulting in the loss of important information. To address the above two issues, this paper proposes a spatial multi-prior driven resolution rescaling framework for intra-frame coding, called MP-RRF, consisting of three sub-networks: a multi-prior driven network, a downscaling network, and an upscaling network. First, the multi-prior driven network employs complexity and similarity priors to smooth the unnecessarily complicated information while leveraging similarity and quality priors to produce high-fidelity complementary information. This interaction of complexity, similarity and quality priors ensures redundancy reduction and texture enhancement. Second, the downscaling network discriminatively processes components of different granularities to generate a compact, low-resolution image for encoding. The upscaling network aggregates a complementary set of contextual multi-scale features to reconstruct realistic details while combining variable receptive fields to suppress multi-scale compression artifacts and resampling noise. Extensive experiments show that our network achieves a significant 23.84% Bjøntegaard Delta Rate (BD-Rate) reduction under all-intra configuration compared to the codec anchor, offering the state-of-the-art coding performance.
Peiying Wu, Shiwei Wang 0005, Liquan Shen, Feifeng Wang, Zhaoyi Tian
IEEE Trans. Multim.5