VLDB 2026 Research / reviewers in the wild / expert
Zhiyu Zhang 0010
dblp:45/6271-10
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0001-1728-4754ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Scheme for Face Video CompressionabstractWith the rapid development of social media, the amount of face video data has grown rapidly, making face video compression a hot research topic. Traditional video coding techniques do not discriminate video content and compress all videos in the same way, while talking head video compression should have more potential. Existing generative compression methods mostly adopt static reference frames, resulting in a decrease in fidelity caused by dynamic background or large pose change. In this article, we propose a hybrid compression scheme for face videos which combines traditional coding with generative compression. On the one hand, we sample and encode key frames with traditional codecs to provide dynamic reference frames which contain real-time background and motion information. On the other hand, we devise a deep video generation model to synthesize smooth video frames according to the extracted sparse keypoints. Combining the pixel-level recovery capability of traditional coding with the detail generation capability of deep generative models, our proposed hybrid scheme is able to implement high-fidelity face video compression at low bitrate in real time. Additionally, we also devise a Portrait Recovery module to recover the low-quality key frames, improving the reconstruction quality in low-bitrate scenarios. Extensive experiments show that our method has advantages over traditional codecs and existing generative compression methods in terms of both rate-distortion performance and coding complexity. Anni Tang, Zhiyu Zhang 0010, Jun Ling, Rong Xie 0004, Li Song 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion OptimizationabstractVolumetric videos, benefiting from immersive 3D realism and interactivity, hold vast potential for various applications, while the tremendous data volume poses significant challenges for compression. Recently, NeRF has demonstrated remarkable potential in volumetric video compression thanks to its simple representation and powerful 3D modeling capabilities, where a notable work is ReRF. However, ReRF separates the modeling from compression process, resulting in suboptimal compression efficiency. In contrast, in this paper, we propose a volumetric video compression method based on dynamic NeRF in a more compact manner. Specifically, we decompose the NeRF representation into the coefficient fields and the basis fields, incrementally updating the basis fields in the temporal domain to achieve dynamic modeling. Additionally, we perform end-to-end joint optimization on the modeling and compression process to further improve the compression efficiency. Extensive experiments demonstrate that our method achieves higher compression efficiency compared to ReRF on various datasets. Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Anni Tang, Qiang Hu 0003, Li Song 0001 |
ICME | 1 |
| 2024 | Rate-aware Compression for NeRF-based Volumetric VideoabstractThe neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing solutions typically compress these NeRF representations after the training stage, leading to a separation between representation training and compression. In this paper, we try to directly learn a compact NeRF representation for volumetric video in the training stage based on the proposed rate-aware compression framework. Specifically, for volumetric video, we use a simple yet effective modeling strategy to reduce temporal redundancy for the NeRF representation. Then, during the training phase, an implicit entropy model is utilized to estimate the bitrate of the NeRF representation. This entropy model is then encoded into the bitstream to assist in the decoding of the NeRF representation. This approach enables precise bitrate estimation, thereby leading to a compact NeRF representation.Furthermore, we propose an adaptive quantization strategy and learn the optimal quantization step for the NeRF representations. Finally, the NeRF representation can be optimized by using the rate-distortion trade-off. Our proposed compression framework can be used for different representations and experimental results demonstrate that our approach significantly reduces the storage size with marginal distortion and achieves state-of-the-art rate-distortion performance for volumetric video on the HumanRF and ReRF datasets. Compared to the previous state-of-the-art method TeTriRF, we achieved an approximately -80% BD-rate on the HumanRF dataset and -60% BD-rate on the ReRF dataset. Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Zhengxue Cheng, Anni Tang, Li Song 0001 |
ACM Multimedia | 1 |
| 2024 | Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing SystemabstractIn streaming media services, video transcoding is a common practice to alleviate bandwidth demands. Unfortunately, traditional methods employing a uniform rate factor (RF) across all videos often result in significant inefficiencies. Content-adaptive encoding (CAE) techniques address this by dynamically adjusting encoding parameters based on video content characteristics. However, existing CAE methods are often tightly coupled with specific encoding strategies, leading to inflexibility. In this paper, we propose a model that predicts both RF-quality and RF-bitrate curves, which can be utilized to derive a comprehensive bitrate-quality curve. This approach facilitates flexible adjustments to the encoding strategy without necessitating model retraining. The model leverages codec features, content features, and anchor features to predict the bitrate-quality curve accurately. Additionally, we introduce an anchor suspension method to enhance prediction accuracy. Experiments confirm that the actual quality metric (VMAF) of the compressed video stays within ±1 of the target, achieving an accuracy of 99.14%. By incorporating our quality improvement strategy with the rate-quality curve prediction model, we conducted online A/B tests, obtaining both +0.107% improvements in video views and video completions and +0.064% app duration time. Our model has been deployed on the Xiaohongshu App. Shibo Yin, Zhiyu Zhang 0010, Peirong Ning, Qiubo Chen, Guo Lu, Li Song 0001 |
VCIP | 2 |
| 2023 | High-Fidelity Free-View Talking Head Synthesis for Low-Bandwidth Video ConferenceabstractAs video conferencing becomes an indispensable part of human’s daliy life, how to achieve a high-fidelity calling experience under low bandwidth has been a popular and challenging issue. Deep generative models have great potential in low-bandwidth facial video compression due to the excellent generation capability based on abridged information. Nevertheless, exsiting deep generation-based compression methods tend to handle motion information in pure 2D or pseudo 3D space, causing facial distortion when large head poses are encountered. In this paper, we propose a 3D-aware high-fidelity facial video conferencing system based on a parameterized NeRF-based face model. Through the compression of the parameterized face model and the transmisstion of extracted facial parameters, we implement high-fidelity talking head synthesis for video conferencing at an ultra-low bitrate. Additionally, the 3D perception capability of the system allows for viewpoint control over the head, achieving higher interactivity and practicability. Extensive experiments verify the effectiveness of the proposed 3D-aware high-fidelity free-view facial video conferencing system. Zhiyu Zhang 0010, Anni Tang, Guo Lu, Rong Xie 0004, Li Song 0001 |
VCIP | 1 |
| 2022 | Generative Compression for Face Video: A Hybrid SchemeabstractAs the latest video coding standard, versatile video coding (VVC) has shown its ability in retaining pixel quality. To excavate more compression potential for video conference scenarios under ultra-low bitrate, this paper proposes a bitrate-adjustable hybrid compression scheme for face video. This hybrid scheme combines the pixel-level precise recovery capability of traditional coding with the generation capability of deep learning based on abridged information, where Pixel-wise Bi-Prediction, Low-Bitrate-FOM and Lossless Keypoint Encoder collaborate to achieve PSNR up to 36.23 dB at a low bitrate of 1.47 KB/s. Without introducing any additional bi-trate, our method has a clear advantage over VVC under a completely fair comparative experiment, which proves the effectiveness of our proposed scheme. Moreover, our scheme can adapt to any existing encoder/configuration to deal with different encoding requirements, and the bitrate can be dynamically adjusted according to the network condition. Anni Tang, Yan Huang 0033, Jun Ling, Zhiyu Zhang 0010, Rong Xie 0004, Li Song 0001 |
ICME | 4 |