VLDB 2026 Research / reviewers in the wild / expert
Qi Zhang 0042
dblp:52/323-42
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-1189-8755ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Predicting Satisfied User and Machine Ratio for Compressed Images: A Unified ApproachabstractNowadays, high-quality images are pursued by both humans for better viewing experience and by machines for more accurate visual analysis. However, images are usually compressed before being consumed, decreasing their quality. It is meaningful to predict the perceptual quality of compressed images for both humans and machines, which guides the optimization for compression. This issue remains underexplored. In this paper, we propose a unified approach to fill this gap. Specifically, we create a deep learning-based model to predict Satisfied User Ratio (SUR) and Satisfied Machine Ratio (SMR) of compressed images simultaneously. We first pre-train a feature extractor network on a large-scale SMR-annotated dataset with human perception-related quality labels generated by diverse image quality models, which simulates the acquisition of SUR labels. Then, we leverage and fuse the extracted multi-layer features to predict SUR and SMR with a more capable network architecture. We propose a Difference Feature Residual Learning (DFRL) module to learn more discriminative difference features. We also design a Multi-Head Attention Aggregation and Pooling (MHAAP) layer to aggregate difference features and reduce their redundancy. We further introduce an MLP-Mixer module to integrate global spatial and channel information for subsequent SUR and SMR regression. Experimental results indicate that the proposed model significantly outperforms state-of-the-art SUR and SMR prediction methods by up to 48% on several challenging datasets. Moreover, our joint learning scheme of human and machine perceptual quality prediction tasks is effective at improving the performance of both. Qi Zhang 0042, Shanshe Wang, Xinfeng Zhang 0001, Xiandong Meng, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Point Cloud-Assisted Neural Image CompressionabstractHigh-efficient image compression is a critical requirement. In several scenarios where multiple modalities of data are captured by different sensors, the auxiliary information from other modalities are not fully leveraged by existing image-only codecs, leading to suboptimal compression efficiency. In this paper, we increase image compression performance with the assistance of point cloud, which is widely adopted in the area of autonomous driving. As depicted in Figure 1 (a), we have unified the digital representation of image and point cloud, and propose the point cloud-assisted neural image codec (PCA-NIC) to enhance the preservation of image texture and structure by utilizing the high-dimensional point cloud information. As depicted in Figure 1 (b), we further introduce a multi-modal feature fusion transform module (MMFFT) to capture more representative image features, remove redundant information between channels and modalities that are not relevant to the image content. Ziqun Li, Qi Zhang 0042, Xiaofeng Huang, Zhao Wang 0004, Siwei Ma 0001 |
DCC | 2 |
| 2025 | LL-ICM: Image Compression for Low-Level Machine Vision via Large Vision-Language ModelabstractImage Compression for Machines (ICM) aims to compress images for machine vision tasks, while current methods mostly focus on the demands for high-level tasks. However, the quality of original images is usually not guaranteed in the real world, leading to even worse downstream task performance after compression. Thus, lowlevel (LL) restoration tasks should also be considered in ICM. In this paper, we propose the first ICM framework for LL machine vision tasks, namely LL-ICM, which optimizes the compression and LL processing performance simultaneously. Moreover, LL-ICM leverages large vision-language model (VLM) to solve different LL task within a single model, which is particularly useful when the distortion type of the original image is uncertain. As illustrated in Fig. 1(a), LL-ICM consists of a neural image codec and a VLM-based LL processing module. Given an original image with distortions, LL-ICM firstly compress it as$\hat{\mathbf{X}}$. Then, we extract a generalized feature F from$\hat{\mathbf{X}}$, which is then encoded as two representations, distortion type$\varphi$and caption$\sigma$. After that, the LL processing module receives$\hat{\mathbf{X}}$and its representations to generate the restored version of$\hat{\mathbf{X}}$, i.e.,$\hat{\mathbf{X}}_{\mathbf{H}}$. Qi Zhang 0042, Chuanmin Jia, Shiqi Wang 0001 |
DCC | 2 |
| 2025 | HGC-Avatar: Hierarchical Gaussian Compression for Streamable Dynamic 3D AvatarsabstractRecent advances in 3D Gaussian Splatting (3DGS) have enabled fast, photorealistic rendering of dynamic 3D scenes, showing strong potential in immersive communication. However, in digital human encoding and transmission, the compression methods based on general 3DGS representations are limited by the lack of human priors, resulting in suboptimal bitrate efficiency and reconstruction quality at the decoder side, which hinders their application in streamable 3D avatar systems. We propose HGC-Avatar, a novel Hierarchical Gaussian Compression framework designed for efficient transmission and high-quality rendering of dynamic avatars. Our method disentangles the Gaussian representation into a structural layer, which maps poses to Gaussians via a StyleUNet-based generator, and a motion layer, which leverages the SMPL-X model to represent temporal pose variations compactly and semantically. This hierarchical design supports layer-wise compression, progressive decoding, and controllable rendering from diverse pose inputs such as video sequences or text. Since people are most concerned with facial realism, we incorporate a facial attention mechanism during StyleUNet training to preserve identity and expression details under low-bitrate constraints. Experimental results demonstrate that HGC-Avatar provides a streamable solution for rapid 3D avatar rendering, while significantly outperforming prior methods in both visual quality and compression efficiency. Haocheng Tang, Ruoke Yan, Xinhui Yin, Qi Zhang 0042, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia |
ACM Multimedia | 4 |
| 2024 | Low-Complexity 3D-Vision Conferencing System based on Accelerated RIFE ModelabstractRecent advancements in telecommunication technologies have exceeded the requirements of numerous video-based Real-Time Communication (RTC) applications. Mean-while, in light of the growing demand for immersive 3D visual experience, researchers are currently focusing on developing next-generation telepresence systems. This paper presents a novel immersive conferencing system that offers a smooth, high-fidelity, and life-size autostereoscopic display of remote user portraits. To generate the binocular stereo vision in real time, an adaptive low-complexity view synthesis method based on an accelerated Real-time Intermediate Flow Estimation (RIFE) model is employed, which performs direct cross-view generation based on decoded multi-view videos and tracked eye positions of the watching user. Thanks to eliminating the complex 3D modeling procedure that relies on depth images, the proposed system requires significantly fewer computational resources and lower video transmission bandwidth compared to existing immersive conferencing systems. Therefore, the proposed system is low-cost and flexible when accommodating diverse conferencing environments. Hongyue Huang, Hongbo Ning, Haopeng Lu, Qi Zhang 0042, Yanpeng Liang, Wanjun Lyu, Chuanmin Jia, Xinfeng Zhang 0001, Liuxin Zhang, Siwei Ma 0001 |
PCS | 5 |
| 2024 | Pose-Driven Compression for Dynamic 3D Human via Human Prior ModelsabstractTo cost-effectively transmit high-quality dynamic 3D human images in immersive multimedia applications, efficient data compression is crucial. Unlike existing methods that focus on reducing signal-level reconstruction errors, we propose the first dynamic 3D human compression framework based on human priors. The layered coding architecture significantly enhances the perceptual quality while also supporting a variety of downstream tasks, including visual analysis and content editing. Specifically, a high-fidelity pose-driven Avatar is generated from the original frames as the basic structure layer to implicitly represent the human shape. Then, human movements between frames are parameterized via a commonly-used human prior model, i.e., the Skinned Multi-Person Linear Model (SMPL), to form the motion layer and drive the Avatar. Furthermore, the normals are also introduced as an enhancement layer to preserve fine-grained geometric details. Finally, the Avatar, SMPL parameters, and normal maps are efficiently compressed into layered semantic bitstreams. Extensive qualitative and quantitative experiments show that the proposed framework remarkably outperforms other state-of-the-art 3D codecs in terms of subjective quality with only a few bits. More notably, as the size or frame number of the 3D human sequence increases, the superiority of our framework in perceptual quality becomes more significant while saving more bitrates. Ruoke Yan, Qian Yin 0002, Xinfeng Zhang 0001, Qi Zhang 0042, Gai Zhang, Siwei Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Perceptual Video Coding for Machines via Satisfied Machine Ratio ModelingabstractVideo Coding for Machines (VCM) aims to compress visual signals for machine analysis. However, existing methods only consider a few machines, neglecting the majority. Moreover, the machine's perceptual characteristics are not leveraged effectively, resulting in suboptimal compression efficiency. To overcome these limitations, this paper introduces Satisfied Machine Ratio (SMR), a metric that statistically evaluates the perceptual quality of compressed images and videos for machines by aggregating satisfaction scores from them. Each score is derived from machine perceptual differences between original and compressed images. Targeting image classification and object detection tasks, we build two representative machine libraries for SMR annotation and create a large-scale SMR dataset to facilitate SMR studies. We then propose an SMR prediction model based on the correlation between deep feature differences and SMR. Furthermore, we introduce an auxiliary task to increase the prediction accuracy by predicting the SMR difference between two images in different quality. Extensive experiments demonstrate that SMR models significantly improve compression performance for machines and exhibit robust generalizability on unseen machines, codecs, datasets, and frame types. Qi Zhang 0042, Shanshe Wang, Xinfeng Zhang 0001, Chuanmin Jia, Zhao Wang 0004, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | RealVR: Efficient, Economical, and Quality-of- Experience-Driven VR Video System Based on MPEG OMAFabstractRecent years have witnessed the explosion of virtual reality (VR) videos and applications. This new form of media grants us the precious freedom we never had before to look at any directions of the video content. With such privilege, our desires for higher video resolution and frame-rate, better visual quality, and smoother watching experience have remarkably risen. However, VR videos occupy an astronomical amount of data, which poses unprecedented challenges to computation efficiency, deployment cost, and user experience of the system. In this paper, we propose RealVR, an end-to-end tile-based VR video system that faces and tackles such challenges. With all essential procedures from the initial capturing to the ultimate rendering included, the system is designed, implemented, and configured specifically to achieve the best balance among efficiency, economy, and quality-of-experience (QoE). It leverages the promising international VR standard, MPEG OMAF, to process and deliver 8K 60 fps VR video, but is also improved for much better performance and user experience than the pristine standard, especially for reducing the motion-to-high-quality (MtHQ) latency. It does not rely on additional expensive edge servers to offer an immersive user experience, which makes it more economical than many other works. Through extensive experiments, it is proven that RealVR can significantly improve the MtHQ experience and save bandwidth consumption without compromising on encoding efficiency or application cost, maximizing the user‘s freedom. Qi Zhang 0042, Jianchao Wei, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Just Recognizable Distortion for Machine Vision Oriented Image and Video Coding
Qi Zhang 0042, Shanshe Wang, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001 |
Int. J. Comput. Vis. | 1 |
| 2020 | A Novel Visual Analysis Oriented Rate Control Scheme for HEVCabstractRecent years have witnessed an explosion of machine visual intelligence. While impressive performance on visual analysis has been achieved by powerful Deep-Learning-based models, the texture and feature distortion caused by image and video coding is becoming a challenge in practical situations. In this paper, a new rate control scheme is proposed to improve visual analysis performance on coded video frames. Firstly, a new kind of visual analysis distortion is introduced to build a Rate-Joint-Distortion model. Secondly, the Rate-Joint-Distortion Optimization problem is solved by using Lagrange multiplier method, and the relationship between rate and Lagrange multiplier λ is described by a hyperbolic model. Thirdly, a logarithmic λ - QP model is established to achieve minimum Rate-Joint-Distortion cost for given λs. The experimental results show that the proposed scheme can improve visual analysis performance with stable bits used for coding. Qi Zhang 0042, Shanshe Wang, Siwei Ma 0001 |
VCIP | 1 |