Zhuozhen Yu

dblp:283/1992 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0006-2987-7844ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Zero-shot Quantization for Large-kernels via Shape-based Distribution and Diversity Self-distillation
abstract
Zero-shot quantization (ZSQ) has emerged as an effective method to reduce model complexity and memory footprint without using original training data, thereby mitigating data privacy and security concerns during model deployment. Recently, Large-Kernel Convolutional Neural Networks (LKCNNs) have achieved state-of-the-art performance on various vision tasks, which introduce challenges in terms of increased parameters and network complexity, making them difficult to deploy on resource-constrained edge devices. Despite the success of ZSQ, existing methods fail to apply to LKCNNs due to architectural differences such as Batch Normalization (BN) layers in models and thus result in significant performance declines. In this paper, we propose a novel ZSQ framework tailored specifically for LKCNNs, considering their two key characteristics: the large receptive field and the reliance on shape bias. Correspondingly, we first employ an edge detection-based loss to optimize synthetic images that closely mimic the distribution of real images, and a diversity self-distillation loss to maintain consistency in feature representation to enable the generation of synthetic images. Afterward, we use these synthetic images to fine-tune the quantization parameters with a shape-enhance data augmentation strategy. Experiment results demonstrate the superiority of the proposed framework over existing methods, with significant improvements in maintaining accuracy after quantization across various quantization configurations on the ImageNet dataset.
Zhuozhen Yu, Xinrui Chen 0001, Shunzhou Wang, Wei Gao 0003
ICASSP2
2025 Zero-shot Quantization of Vision Transformers: Leveraging Multi-model Ensembles and Attention Mixup
abstract
Zero-shot quantization (ZSQ) shows promise in compressing and accelerating deep neural networks in scenarios where the original training data is inaccessible. Recently, ZSQ for vision transformers (ViTs) has been proposed to synthesize samples for ViT network quantization, the quality of which significantly impacts the performance of quantized models. Nonetheless, we observe that the synthetic samples produced by current ZSQ techniques exhibit severe bias and insufficient optimization, which deviate from real data and lead to substantial performance declines. On the one hand, the synthetic samples generated by a single ViT network are inaccurate and biased to the specific model. On the other hand, unlike convolutional neural networks (CNNs) with BatchNorm layers, ViTs do not store any training set statistics in the networks, hindering the generation of high-quality calibration samples. To address the above issues, we propose leveraging Multi-model Ensembles and Attention Mixup in ZSQ for ViTs (MMA-ViT). Specifically, MMA-ViT employs an ensemble of diverse pre-trained proxy CNN models to narrow the sample synthesizing space, utilizing their predictive capabilities and BatchNorm statistics to generate exact synthetic images that enhance generality. Additionally, MMA-ViT integrates a unique attention-driven mixup technique for accurate data augmentation during the sample synthesis process, avoiding over-fitting to the networks. The efficacy of MMA-ViT has been demonstrated through extensive experiments and ablation studies on the ImageNet dataset. For example, when Swin-B is quantized to W3/A4, our method achieves a 11.89% top-1 accuracy increase on ImageNet compared to state-of-the-art methods.
Xinrui Chen 0001, Zhuozhen Yu, Shunzhou Wang, Wei Gao 0003
ICME3
2024 When Dynamic Neural Network Meets Point Cloud Compression: Computation-Aware Variable Rate and Checkerboard Context
abstract
For exploring the Rate-Distortion-Complexity (RDC) optimization in point cloud compression, we propose a point cloud compressor with dynamic channel. In the transform process of the proposed compressor ( Fig 1.a ), we devise a sparse convolution operator, named AdaSConv, shown in Fig 1.b , to support RDC optimization, which ensures model capacity can adjust Rate-Distortion performance. What is more, to fill the blank of improved entropy model in point cloud feature compression, we design a 3D checkerboard entropy model. The 3D checkerboard divides points in the whole space into two parts: anchor and non-anchor, which will be compressed in sequence. As Fig 1.c illustrates, the compression of non-anchor will refer to the information in coded anchor points through Masked AdaSConv ( Fig 1.d ). We conduct floating point operations (FLOPs) computation, which reveals that the smallest rate point only consumes 7% of the FLOPs used by the full-width model. Besides, experiment results show 3D checkerboard has at most 13.86% gains of BD-Rate compared with factorized entropy model in the same experimental settings in Owill dataset with only slight extra time and computation.
Zhuozhen Yu, Wei Gao 0003
DCC1
2024 OpenDIC: An Open-Source Library and Performance Evaluation for Deep-learning-based Image Compression
abstract
Deep learning technologies have been popular in the image compression field for some time. An increasing number of deep-learning-based models are proposed to improve Rate-Distortion (RD) performance. Previous algorithms are implemented in the specific platform and can not be applied in cross-platform environments. In this paper, we present an open-source algorithm library called OpenDIC, which integrates a variety of end-to-end image compression methods in cross-platform environments. The contribution and details of the algorithms used in the library are described. To evaluate the performance of these algorithms, we conduct a comprehensive performance test. We compare and analyze each algorithm according to RD performance, running time, and GPU memory occupancy. The algorithm library has been released at https://openi.pcl.ac.cn/OpenDIC/.
Wei Gao 0003, Huiming Zheng, Kaiyu Zheng, Zhuozhen Yu, Yuan Li 0076, Yongchi Zhang
ACM Multimedia5
2024 ViewPCGC: View-Guided Learned Point Cloud Geometry Compression
abstract
With the rise of immersive media applications such as digital museums, virtual reality, and interactive exhibitions, point clouds, as a three-dimensional data storage format, have gained increasingly widespread attention. The massive data volume of point clouds imposes extremely high requirements on transmission bandwidth in the above applications, gradually becoming a bottleneck for immersive media applications. Although existing learning-based point cloud compression methods have achieved specific successes in compression efficiency by mining the spatial redundancy of their local structural features, these methods often overlook the intrinsic connections between point cloud data and other modality data (such as image modality), thereby limiting further improvements in compression efficiency. To address the limitation, we innovatively propose a view-guided learned point cloud geometry compression scheme, namely ViewPCGC. We adopt a novel self-attention mechanism and cross-modality attention mechanism based on sparse convolution to align the modality features of the point cloud and the view image, removing view redundancy through Modality Redundancy Removal Module (MRRM). Simultaneously, side information of the view image is introduced into the Conditional Checkboard Entropy Model (CCEM), significantly enhancing the accuracy of the probability density function estimation for point cloud geometry. In addition, we design a View-Guided Quality Enhancement Module (VG-QEM) in the decoder, utilizing the contour information of the point cloud in the view image to supplement reconstruction details. The superior experimental performance demonstrates the effectiveness of our method. Compared to the state-of-the-art point cloud geometry compression methods, ViewPCGC exhibits an average performance gain exceeding 10% on D1-PSNR metric.
Huiming Zheng, Wei Gao 0003, Zhuozhen Yu, Tiesong Zhao, Ge Li 0002
ACM Multimedia3
2021 Two-dimensional jointly sparse robust discriminant regression
Zhihui Lai 0001, Zhuozhen Yu, Heng Kong, LinLin Shen
Signal Process. Image Commun.2