VLDB 2026 Research / reviewers in the wild / expert
Yuhuai Zhang
dblp:166/5210
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPC-NeRF: Spatial Predictive Compression for Voxel-Based Radiance FieldabstractRepresenting the Neural Radiance Field (NeRF) with the Explicit Voxel Grid (EVG) is a promising direction for improving NeRFs. However, the EVG representation is not efficient for storage and transmission because of the tremendous memory cost. Existing methods for compressing EVG mainly inherit the methods designed for neural network compression, such as pruning and quantization, which do not take full advantage of the spatial correlation in the voxel grid. Inspired by the prosperous digital image compression techniques, this article proposes SPC-NeRF, a novel framework applying spatial predictive coding in EVG NeRF compression. The proposed framework can remove spatial redundancy efficiently for better compression performance. Our framework contains a progressive coding procedure to realize adaptive quantization precision according to the different importance of the voxels. Moreover, we model the coding bitrate of our framework and design a novel form of the loss function. With the loss function, we can jointly optimize the compression ratio and the rendering distortion to achieve higher coding efficiency. Extensive experiments demonstrate that our method can achieve 32% bit saving compared to the benchmark method VQRF on multiple representative test datasets, with comparable training time. Zetian Song, Jiaqi Zhang 0007, Wenhong Duan, Yuhuai Zhang, Xinfeng Zhang 0001, Siwei Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Enhanced Decoder-Side Secondary Transform Derivation for Video Coding Beyond AVS3abstractThe Decoder-side Secondary Transform Derivation (DSTD) method has been adopted by the exploration software of the fourth generation Audio Video coding Standard (AVS4), Exploration Video Model (EVM). Although DSTD can significantly improve the coding performance, it increases the encoding complexity simultaneously. To reduce the encoding complexity of DSTD, three enhanced DSTD methods are proposed in this paper, which consists of Deleting Mode Optimization (DMO), Intra Mode Dependent Optimization (IMDO) and Interleaved Intra Mode Dependent Optimization (IIMDO). The process of DMO is similar to that of the original DSTD method, but all the contents related to diagonal flipping have been removed. IMDO divides intra prediction modes into three areas based on the angle of the intra prediction mode, as shown in Fig. 1. Each area will correspondingly reduce a secondary transform type. IIMDO divides intra prediction modes into four areas, as shown in Fig. 2. The experimental results demonstrate that the proposed methods can effectively reduce the encoding complexity with negligible coding performance loss. The first enhanced method, Deleting Mode Optimization, has been adopted by EVM. Yuhuai Zhang, Jiaqi Zhang 0007, Weijia Jiang, Siwei Ma 0001 |
DCC | 2 |
| 2025 | Joint Structure-Texture Scan-Order for Point Cloud Attribute Compression Using Affine TransformationabstractExisting geometry-based Point Cloud Compression (PCC) frameworks are typically designed to code the geometric coordinates first, followed by compressing the attributes (e.g., colors, reflectances) according to the order derived from geometric structures, such as the Morton codes. Although geometry-based reordering methods can eliminate the redundancy of attributes, the errors caused by dramatic variations of attributes in the non-smooth areas potentially limit the efficiency of the point cloud attribute coding. To tackle this challenge, a novel joint structure-texture scan-order and coding scheme is proposed, which aims to explore a better attribute coding order from the viewpoint of improving the geometry-attribute consistency. Specifically, we formulate the attribute reordering problem as a geometry-attribute alignment task, and utilize the affine-transform model to find the optimal correspondence between geometry and attribute information by minimizing attributed prediction residuals. Then, the Morton codes-based point cloud reordering is conducted on the transformed point cloud. Note that our residual-based mode decision scheme implicitly embodies that the proposed reordering method further incorporates attribute textures based on the geometric structure. However, the exhaustive search for the optimal transformation space introduces the extremely high complexity to the encoder. Therefore, we also propose a fast pruning algorithm to narrow the search space for the approximate solution as an alternative. Experimental results conducted on the various benchmark datasets have illustrated that our proposed method outperforms the MPEG standard G-PCC with gains of up to 2%, 9%, and 7% in luma, chroma, and reflectance, respectively. Qian Yin 0002, Xinfeng Zhang 0001, Ruoke Yan, Yuhuai Zhang, Shanshe Wang, Siwei Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Decoder-side Secondary Transform Derivation for Video Coding beyond AVS3abstractSecondary transform was adopted into the third generation Audio Video coding Standard (AVS3) to improve the intra-coded residual coding by applying a 4×4 secondary transform kernel. However, the adaptability of the single 4×4 transform kernel is limited for various residual data. In order to achieve higher residual coding gains, we propose a Decoder-side Secondary Transform Derivation (DSTD) method. Specifically, DSTD expands the maximum range of secondary transform from 4×4 to 8×8, where an 8×8 size transform kernel is introduced to further enhance the capability of compacting residuals. In particularly, three flipped secondary transform types are employed to extend transform candidates, including horizontal, vertical and diagonal flipping types. The boundary continuity is utilized to derive the transform type. Experimental results show that the proposed method can achieve 0.51% and 0.18% BD-rate savings on average under All Intra (AI) and Random Access (RA) configurations, respectively. DSTD has been adopted into the Exploration Video Model (EVM) for AVS4. Yuhuai Zhang, Huiwen Ren, Shiqi Wang 0001, Siwei Ma 0001 |
DCC | 1 |
| 2024 | Image Encryption and Compression Based on Reversed Diffusion ModelabstractNowadays, as critical conduits of communication, the information security of images and videos is particularly important. The existing encryption techniques usually transform images into high-frequency content that resembles noise, pre-senting significant challenges in achieving efficient compression. This paper presents an innovative collaborative approach that integrates image encryption and compression using a reversed diffusion model. This method, by reversing the typical process of diffusion models, adeptly changes encrypted high-frequency content into a domain that is more amenable to compression. Leveraging the reversible nature of the Denoising Diffusion Implicit Models (DDIM), our framework ensures the high-fidelity restoration of information. Our experimental findings demonstrate that this approach not only effectively encrypts images but also compresses the encrypted high-frequency noise content, outperforming Video Versatile Coding (VVC) in compression performance. Jianhui Chang, Yuhuai Zhang, Jian Zhang 0018, Siwei Ma 0001 |
PCS | 3 |
| 2024 | Diffusion-Based Hypotheses Generation and Joint-Level Hypotheses Aggregation for 3D Human Pose EstimationabstractTo combine the advantages of deterministic and probabilistic 3D human pose estimation methods, we decompose pose estimation into two processes: hypotheses generation and hypotheses aggregation. For hypotheses generation, we propose a novel Diffusion-based 3D Pose generation (D3DP) method. D3DP generates a diversified group of plausible 3D pose hypotheses from a single 2D keypoint observation. Utilizing a diffusion process, it gradually transforms ground-truth 3D poses towards a random distribution, subsequently employing a conditioned denoiser guided by the observed keypoints to recover the uncorrupted 3D poses. Moreover, D3DP is compatible with existing deterministic 3D pose estimators and allows users to optimize the trade-off between computational efficiency and pose accuracy via two adjustable parameters. For hypotheses aggregation, we propose two alternative approaches: a Reprojection-Based Selection (RBS) method and a Hypotheses Selection Network (HSN). These methods adopt the joint-level strategy to assemble multiple hypotheses generated by D3DP into a single 3D pose for practical use. Specifically, RBS reprojects 3D pose hypotheses to the 2D camera plane, and selects the best hypothesis based on the reprojection errors. HSN evaluates each hypothesis and selects the hypothesis with the highest confidence score as the output. Then these selected joints are combined into the final pose. The proposed methods implement a joint-by-joint aggregation strategy that capitalizes on the 2D prior and temporal information, both of which have been ignored by previous pose-level methods. Extensive experiments on two benchmarks highlight that the proposed method outperforms the state-of-the-art deterministic and probabilistic approaches. Wenkang Shan, Yuhuai Zhang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Flexible Hierarchical Parallel Processing for AVS3 Video Coding
Hannong Zheng, Yuhuai Zhang, Jian Zhang 0018, Hengyu Man, Siwei Ma 0001 |
ICIG (3) | 2 |
| 2022 | Implicitly Selected Transform for AVS3abstractTransform coding removes redundancy by de-correlating the residual data, which plays a crucial role in video coding. Since different transform cores can adapt to different video content and residual data, multiple cores transform has been proposed to improve the coding performance. However, the overhead for representing the transform type is inevitable. This paper intensively studies the statistical characteristics of transform coefficients. Theoretical analysis shows that the implicit method in the Implicitly Selected Transform (IST) is superior to the explicit signalling method. As such, we propose the parity adjustment scheme which seamlessly cooperates with the rate-distortion optimization quantization for IST. Furthermore, we propose two combination methods, size-based and number-based, to optimize IST. Moreover, a restriction region of IST is applied to reduce the complexity. Experimental results on HPM-6.0, which is the reference software of the AVS3 video coding standard, show that our proposed method can achieve 1.76% and 0.76% BD-Rate savings on average under AI and RA configurations, respectively, along with negligible decoding time variations. The comparison results between the proposed method and explicit signalling method illustrate that the proposed method can achieve better coding gain than the latter. Our method has been adopted in AVS3. Yuhuai Zhang, Kai Zhang 0007, Li Zhang 0006, Shanshe Wang, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Implicit Seleted Transform Skip Method For Avs3abstractAVS3 is an emerging video coding standard, and screen content coding is a very important feature of AVS3. This paper presents a method of Implicit-Selected Transform Skip (ISTS) to further improve the screen content coding performance. With ISTS, transform skip mode is introduced as an optional substitution to transform-coding on residual signals of blocks with intra-prediction. The indication of whether to apply transform skip is hidden into the Parity of the Number of Non-zero Coefficients (PNNC) of a residual block, instead of being signaled to the decoder. Moreover, the coefficients of an intra-coded block are reordered to make the coefficients more compact. Experimental results show that the proposed method can achieve 12.04%, 8.15% and 10.19% BD-rate savings on average under All Intra (AI), Low Delay (LD) and Random Access (RA) configurations, respectively, with the encoding time reduced by 5% to 8%. ISTS has been adopted into AVS3. Yuhuai Zhang, Kai Zhang 0007, Li Zhang 0006, Hongbin Liu 0004, Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 1 |
| 2020 | A consensus reaching process with quantum subjective adjustment in linguistic group decision making
Yuhuai Zhang |
Inf. Sci. | 3 |
| 2015 | Boundedness of Marcinkiewicz integral with rough kernel on Triebel-Lizorkin spacesabstractThis paper is a continuation of our previous work (Zhang and Chen, 2010b). Following the same general steps of the proof there, we make essential improvement on our previous theorem by recalculating a key inequality. Our result shows that the Marcinkiewicz integral, with a bounded radial function in its kernel, is still bounded on the Triebel-Lizorkin space. Fangfang Ren, Yuhuai Zhang, Guilian Gao |
Frontiers Inf. Technol. Electron. Eng. | 3 |