EDBT 2026 Demo / reviewers in the wild / expert
Saiping Zhang
dblp:199/8382
· DBLP profile ↗
13ranked-venue papers
7as first author
10since 2021 · last 2024
0000-0002-2442-1104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A slimmable framework for practical neural video compressionabstractDeep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs. Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz |
Neurocomputing | 7 |
| 2022 | DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed VideoabstractIn this paper, we propose a deformable convolution-based generative adversarial network (DCNGAN) for perceptual quality enhancement of compressed videos. DCNGAN is also adaptive to the quantization parameters (QPs). Compared with optical flows, deformable convolutions are more effective and efficient to align frames. Deformable convolutions can operate on multiple frames, thus leveraging more temporal information, which is beneficial for enhancing the perceptual quality of compressed videos. Instead of aligning frames in a pairwise manner, the deformable convolution can process multiple frames simultaneously, which leads to lower computational complexity. Experimental results demonstrate that the proposed DCNGAN outperforms other state-of-the-art compressed video quality enhancement algorithms. Saiping Zhang, Luis Herranz, Marta Mrak, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
ICASSP | 1 |
| 2022 | Random Forest Accelerated CU Partition for Inter Prediction in H.266/VVCabstractH.266/VVC offers a large reduction in bitrate compared to H.265/HEVC while maintaining the same quality. This is largely attributed to the introduction of the quadtree with the nested multitype tree (QTMTT), which enables the encoder of H.266/VVC to be more adaptive to video sequences with complex textures but also causes a significant increase in the complexity of the coding unit (CU) partition. To address this problem, we propose a fast CU partition algorithm for inter prediction based on random forest in this paper. Using the time-domain information and texture complexity information of CU as features, random forest models are introduced for early termination of the partition. The proposed algorithm is applied to the open and efficient encoder of H.266/VVC, named VVenC. Experimental results have demonstrated that our proposed algorithm achieves an average 15.27% reduction in encoding time in Random Access (RA), while BD-rate (SSIM) only decreased by 0.07%. Saiping Zhang |
ICME | 2 |
| 2022 | End-to-End Quality Controllable Image CompressionabstractQuality control is an important topic in many application scenarios such as medical image compression. To achieve quality control in image compression network, in this paper, we propose a quality controllable image compression network, Quality Controllable Variational Autoencoder (QCVAE). QC-VAE consists of the Quality-Feature-Level (QFL) model we proposed and the Hyperprior Continuously Variable Rate (HCVR) image compression network which can adapt to multiple target qualities with only one single model. Note that even if the target quality is the same, different images should be quantized with different levels. With the help of the QFL model, we can obtain the estimated corresponding quantization level of the input image under the target quality, the selected level will be used to control the quantization loss. By this means, the QC-VAE can adapt to the target quality with high accuracy. Experimental results have shown that compared with the HCVR model, the proposed QC-VAE achieves accurate quality control without rate-distortion (RD) performance loss, indicating its superiority. Luge Wang, Xionghui Mao, Saiping Zhang, Fuzheng Yang 0001 |
PCS | 3 |
| 2022 | Fast CU Partition Method Based on Extra Trees for VVC Intra CodingabstractIn this paper, we propose a method to skip unnecessary CU encoding modes for VVC Intra coding based on the extra trees model. Two extra tree models with calculated features are used to simplify the encoding process, where the first model determines whether to early terminate the partition and the best partition direction and the second model selects the better partition mode between the binary and ternary partition modes. Experimental results show that our proposed method can save encoding time from 34.68% to 46.70%with only from 0.81% to 1.65% increase of BDBR compared to VVC reference software (VTM 10.0). Besides, the method gets a great tradeoff when applied on VVenC 1.0, an efficient encoder ofVVC, at both preset slower and preset medium. Saiping Zhang |
VCIP | 3 |
| 2022 | Fast Inter Prediction Mode Decision Method Based On Random Forest For H.266/VVCabstractIn H.266/VVC, many new tools are introduced in the inter prediction process. These new techniques enable the encoder in H.266/VVC to predict motion vectors more accurately, but inevitably increase the coding complexity. To solve this problem, in this paper, we propose an early termination algorithm for inter prediction based on random forest which is characterized by the information provided by the temporal co-located block and the spatial adjacent block of the current Coding Unit (CU). Specifically, the random forest is used to predict whether the inter prediction process of the current CU will be terminated in advance. Our proposed algorithm is implemented on Fraunhofer Versatile Video Encoder (VVenC). Experimental results have shown that, in the Random Access (RA) mode, the encoding time of VVenC is reduced by 7.71 % on average while Bjontegaard Delta Bit Rate (BDBR) increases by 1.48 %. Kundan Xie, Jianquan Zhou, Saiping Zhang |
VCIP | 3 |
| 2022 | Rate Controllable Learned Image Compression Based on RFL ModelabstractIn this paper, we propose a rate controllable image compression framework, Rate Controllable Variational Autoencoder (RC-VAE), based on the Rate-Feature-Level (RFL) model established through our exploration on the correlation among target rates, image features and quantization levels. Considering that, when meeting the same target rate, different images should be quantized in different levels, we focus on jointly utilizing the target rate and the extracted features of the image to predict the corresponding quantization level and propose the RFL model. Combining the proposed RFL model with a Hyperprior Continuously Variable Rate (HCVR) image compression network, we further propose the RC-VAE. By controlling information loss in quantization process, the RC-VAE can work at the target rate. Experimental results have demonstrated that one single RC-VAE model can adapt to multiple target rates with higher rate control accuracy and better R-D performance compared with the state-of-the-art rate controllable Image compression networks. Saiping Zhang, Luge Wang, Xionghui Mao, Fuzheng Yang 0001, Shuai Wan |
VCIP | 1 |
| 2022 | A GCN-based fast CU partition method of intra-mode VVC
Saiping Zhang, Shixuan Feng, Jingwu Chen, Chunjie Zhou, Fuzheng Yang 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Relative Pose Estimation for Light Field Cameras Based on LF-Point-LF-Point Correspondence ModelabstractIn this paper, we propose a relative pose estimation algorithm for micro-lens array (MLA)-based conventional light field (LF) cameras. First, by employing the matched LF-point pairs, we establish the LF-point-LF-point correspondence model to represent the correlation between LF features of the same 3D scene point in a pair of LFs. Then, we employ the proposed correspondence model to estimate the relative camera pose, which includes a linear solution and a non-linear optimization on manifold. Unlike prior related algorithms, which estimated relative poses based on the recovered depths of scene points, we adopt the estimated disparities to avoid the inaccuracy in recovering depths due to the ultra-small baseline between sub-aperture images of LF cameras. Experimental results on both simulated and real scene data have demonstrated the effectiveness of the proposed algorithm compared with classical as well as state-of-art relative pose estimation algorithms. Saiping Zhang, Dongyang Jin, Yuchao Dai, Fuzheng Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | DVC-P: Deep Video Compression with Perceptual OptimizationsabstractRecent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual op-timizations (DVC-P), which aims at increasing perceptual quality of decoded videos. Our proposed DVC-P is based on Deep Video Compression (DVC) network, but improves it with perceptual optimizations. Specifically, a discriminator network and a mixed loss are employed to help our network trade off among distortion, perception and rate. Furthermore, nearest-neighbor interpolation is used to eliminate checkerboard artifacts which can appear in sequences encoded with DVC frameworks. Thanks to these two improvements, the perceptual quality of decoded sequences is improved. Experimental results demonstrate that, compared with the baseline DVC, our proposed method can generate videos with higher perceptual quality achieving 12.27% reduction in a perceptual BD- rate equivalent, on average. Saiping Zhang, Marta Mrak, Luis Herranz, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
VCIP | 1 |
| 2020 | Rate-distortion-complexity optimization for x265
Saiping Zhang, Fuzheng Yang 0001, Shuai Wan |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Stretching Schemes for Coding Frames of Panoramic Videos in Craster Parabolic ProjectionabstractPanoramic videos are spherical in nature, which further brings great challenges to deal with them. Usually they are projected to planar domain and processed as planar perspective videos. Craster parabolic projection (CPP), as a sphere-to-plane projection format, achieves approximately uniform sampling on the sphere. Without redundant pixels, it can store and represent panoramic videos effectively. However, frames in CPP format are no longer rectangular, which further violates the off-the-shelf video coding standards. In this paper, four stretching schemes are proposed for coding frames of panoramic videos in CPP. For introducing as few pixels as possible, strips are regard as the basic units. Strips in frames in different areas are stretched into rectangles in different sizes for coding. Spherical continuity, planar continuity, nearest-neighbour interpolation and Lanczos interpolation are considered in stretching respectively. Experimental results demonstrate that, compared with strips in Equi-rectangular projection (ERP) format, the proposed schemes can achieve BD-rate reductions up to 30.68% for Y, 32.68% for U and 34.13% for V, and that different schemes are well adapted for different strips. Saiping Zhang, Mengpin Qiu, Fuzheng Yang 0001, Shuai Wan |
VCIP | 1 |
| 2017 | Rate-Distortion Optimization for Video Coding under Given Computational ComplexityabstractRate-distortion optimization (RDO) is widely applied in video coding, which aims at minimizing the coding distortion under a target coding rate. Conventionally, RDO in video coding does not take into account the coding complexity. However, because of the diversity of video applications, the video encoders in different applications may have different requirements of or limitation on the computational complexity. Therefore, it is desirable for video encoders to perform RDO in flexible computational complexity. In this paper, we propose a novel RDO scheme under the given computational complexity for the latest H.265/HEVC standard. A model for prediction of the rate-distortion cost (RD cost) is first established based on a pre-searching process. Then according to the predicted RD cost, the rate-distortion-complexity (R-D-C) characteristics of different coding tree units (CTUs) are analyzed. Finally, the total complexity budget is properly allocated to different CTUs according to their R-D-C characteristics. Experimental results demonstrate that, compared with x265, the proposed algorithm can reduce, on average, the BD-rate by 18.8% under the same requirements of encoding speed. Junkai Feng, Saiping Zhang, Fuzheng Yang 0001, Shuai Wan |
DCC | 2 |